Flow prediction method based on multi-view space patch self-attention

By adopting the multi-view space patch self-attention mechanism in network traffic prediction, the problem of difficulty in capturing space-time dependencies in the existing technology is solved, and more efficient traffic prediction and computing resource optimization are achieved.

CN120046110APending Publication Date: 2025-05-27SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510183178.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture space-time dependencies in network traffic prediction, and the computing complexity and memory consumption are high, making it difficult to deploy on edge devices.

Method used

The traffic prediction method based on multi-view space patch self-attention is adopted, and the traffic and time information are divided into patch blocks through patch operations. Combined with historical data graphs, depth correlation graphs and adaptive graphs, the multi-head segmentation conversion mechanism and self-attention mechanism are used for feature extraction and prediction.

Benefits of technology

Improves the accuracy of traffic prediction, significantly reduces computing complexity and memory usage, making the model more efficient in deploying on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046110A_ABST
    Figure CN120046110A_ABST
Patent Text Reader

Abstract

The invention discloses a flow prediction method based on multi-view space patch self-attention, and belongs to the field of graph neural networks. The method synthesizes data from different angles to construct a more comprehensive representation. Network traffic data of a plurality of time steps are divided into small blocks through patch operation, and space-time correlation is considered jointly by using space-patch block self-attention. According to the method, the network traffic and the network topology structure are used as input, the future traffic prediction value of each node can be output after iterative training, and compared with other prediction methods, the method has excellent prediction performance and calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph neural networks, and particularly to a traffic prediction method based on multi-view spatial patch self-attention. Background Art

[0002] In the fields of transportation and communication, traffic data often exhibits complex spatio-temporal dependencies, which are characterized by non-linear interactions and heterogeneity among spatial regions. Traditional time series prediction methods often struggle to capture these subtle and complex change patterns. However, with the development of cutting-edge technologies such as graph neural networks, their application in network traffic prediction has become increasingly widespread, greatly enhancing the effectiveness of network management and optimization. Nevertheless, there are still many challenges in the actual deployment process, such as the interpretability of models, the processing efficiency of large-scale data, and how to better integrate various types of data. Therefore, continuously exploring and optimizing time series prediction methods is crucial for improving the operating efficiency of transportation and communication systems.

[0003] Early research focused on constructing an adjacency matrix between nodes based on geographical distance and then using it as model input. However, this does not necessarily reflect the actual dependencies. For example, two adjacent nodes may not be connected to each other. To address this issue, recent research has proposed data-driven models to learn the temporal similarity between nodes. However, these methods mainly construct the adjacency matrix from a single perspective. Secondly, existing models are difficult to reveal local and global spatio-temporal dependencies, and high computational complexity and excessive memory consumption pose significant challenges to deployment on edge devices with limited computing power. The transformer model based on the attention mechanism has been widely used in recent years, but fully simulating the network data transfer operation has too high a computational complexity. Some improved algorithms use sparse spatio-temporal fully connected matrices, but the computational complexity is still proportional to the square of time, and only the attention between single time steps is simulated. Summary of the Invention

[0004] The present invention provides a traffic prediction method based on multi-view spatial patch self-attention to minimize traffic prediction error.

[0005] An embodiment of the present invention provides a traffic prediction method based on multi-view spatial patch self-attention, including the following steps:

[0006] Step 1, establish a network traffic prediction problem;

[0007] Step 2, split the input multiple node traffic into patch blocks through patch operation and input them into the traffic mapping layer to obtain traffic embeddings, and split the time information corresponding to the traffic into patch blocks through patch operation and input them into the timestamp mapping layer to obtain timestamp embeddings;

[0008] Step 3: Obtain the spatial embedding from the Laplacian eigenvectors of the network graph adjacency matrix, add the timestamp embedding to the spatial embedding to obtain the spatio-temporal patch embedding, add the spatio-temporal patch embedding to the traffic embedding result, and use the resulting tensor as the input for the next step;

[0009] Step 4: Pass the tensor obtained in Step 3 through the multi-head split transformation mechanism, respectively through the historical data spatial patch self-attention and the deep correlation spatial patch self-attention, concatenate and project the results to obtain the prior data self-attention;

[0010] Step 5: Pass the tensor obtained in Step 3 through the adaptive spatial patch self-attention mechanism and project it, and input it into the gated fusion mechanism with the prior data self-attention in Step 4. Use the resulting tensor as the input for the next step;

[0011] Step 6: Pass the result of Step 5 through a patch flattening layer and a linear transformation layer in sequence to obtain the final traffic prediction result.

[0012] Optionally, in an embodiment of the present invention, Step 1 specifically includes:

[0013] Considering multiple nodes in the network and the associations between the nodes, assuming that there is an edge when two nodes are related, then the network is modeled as a graph where represents the set of nodes and there are represents the set of edges and |ε| = E, A represents the adjacency matrix and

[0014] According to the collected traffic data The traffic at the t-th moment is expressed as The timestamp at the t-th moment is expressed as

[0015] The network traffic prediction problem is to infer the traffic data [X t-T+1 ,..., X t at the next T' moments based on T historical traffic data [X t-T+1 ,..., y t , timestamp information [y and the network graph t+1 ,..., X t+T' .

[0016] Optionally, in an embodiment of the present invention, Step 2 specifically includes:

[0017] For the traffic data and timestamp information of each node, use the patch operation to process them into patch blocks, and assume that the length of each patch block is P l, if the starting interval between two adjacent patch blocks is S, then the sequence corresponding to the i-th node after the patching operation is If S values are added to the end of the original sequence before the patching operation, the total number of patch blocks is After that, use a fully connected layer for the traffic patch sequence X p and the timestamp patch sequence Y p for mapping, and the specific formula is:

[0018]

[0019] where FC(·) represents the fully connected layer, and then we get d -dimensional traffic embedding and the timestamp embedding T emb .

[0020] Optionally, in an embodiment of the present invention, step 3 specifically includes:

[0021] To retain the spatial information in the network, calculate the normalized Laplacian matrix using the network graph adjacency matrix A, and the formula is as follows:

[0022] Δ = I - D -1 / 2 AD -1 / 2

[0023] where I is the identity matrix and D is the degree matrix;

[0024] Use eigenvalue decomposition, and the formula is as follows:

[0025]

[0026] where Λ is the eigenvalue matrix and U is the eigenvector matrix;

[0027] Linearly project the m smallest non-singular eigenvectors in U to obtain the spatial embedding

[0028] Add the spatial embedding and the timestamp embedding through the broadcast mechanism to obtain the spatial-temporal patch embedding Then add it to the traffic embedding to obtain the input tensor

[0029]

[0030] Optionally, in an embodiment of the present invention, in step 3, the adjacency matrix A of the network graph is calculated based on the distance d between node i and node j ij where σ 2 and ∈ are thresholds for controlling the sparsity of A, and the calculation formula is:

[0031]

[0032] Calculate the historical data adjacency matrix A between each node using the differential dynamic time warping algorithm hd = DDTW(X). For node i and node j with length k, construct matrix M, where element m i,j represents the distance between two data points and . The distance is measured by the squared difference between the estimated derivatives D. The formula for calculating the derivative of node i at time t 1 is as follows:

[0033]

[0034] The first and last elements of the sequence are obtained using the same calculation method as the second and penultimate elements respectively;

[0035] Based on the dynamic programming algorithm, determine the optimal alignment distance dis and

[0036] between. The recursive formula for dis ij (t 1 , t 2 ) is:

[0037]

[0038] dis′ ij (t 1 , t 2 ) = {dis ij (t 1 - 1, t 2 - 1), dis ij (t 1 - 1, t 2 ), dis ij (t 1 , t 2 - 1)}

[0039] Construct the historical data adjacency matrix A hd , (A hd ) ij = dis ij (k, k);

[0040] Use the Katz centrality algorithm to obtain the depth correlation adjacency matrix. Let the weight decay coefficient β = 0.1. The formula is as follows:

[0041] A dr = (I - βA) -1 - I.

[0042] Optionally, in an embodiment of the present invention, step 4 specifically includes:

[0043] Input the tensor into the multi-head segmentation conversion layer, and map it to d through the historical data graph fully connected layer and the depth correlation graph fully connected layer respectively hd and d dr , d hd is the dimension corresponding to the mapping by the historical data graph fully connected layer, d dr is the dimension corresponding to the mapping by the depth correlation graph fully connected layer. Use the learnable weight matrices W Q , W K , W V to obtain the corresponding query Q, key K, and value V, and there is d hd + d dr = d′, where d′ is the sum of the dimensions of the two tensors obtained after mapping the tensor using the historical data graph fully connected layer and the depth correlation graph fully connected layer. The formula is as follows:

[0044]

[0045]

[0046] Among them, is the tensor obtained after mapping the tensor using the historical data graph fully connected layer, is the operation of mapping the tensor using the historical data graph fully connected layer, is the tensor obtained after mapping the tensor using the depth correlation graph fully connected layer, is the operation of mapping the tensor using the depth correlation graph fully connected layer, Q hd is the query in the historical data graph self-attention, is the learnable weight matrix corresponding to the query in the historical data graph self-attention, K hd is the key in the historical data graph self-attention, is the learnable weight matrix corresponding to the key in the historical data graph self-attention, V hd is the value in the historical data graph self-attention, is the learnable weight matrix corresponding to the value in the historical data graph self-attention, Q dr is the query in the depth correlation graph self-attention, is the learnable weight matrix corresponding to the query in the depth correlation graph self-attention, K dr is the key in the depth correlation graph self-attention, is a learnable weight matrix corresponding to the key in the depth correlation map self-attention, V dr is the value in the depth correlation map self-attention, is a learnable weight matrix corresponding to the value in the depth correlation map self-attention;

[0047] For each Q at a specific patch time step, generate K and V a list of, then calculate the self-attention of each patch time step n times and average the results to obtain the spatial patch time self-attention at that moment. The formula is as follows:

[0048]

[0049] where, A ∈ {A hd , A dr}, ⊙ represents the Hadamard product, Q i is the i-th query in the spatial patch time self-attention, is the transpose of the p-th key in the spatial patch time self-attention, V p is the p-th value in the spatial patch time self-attention;

[0050] After operating on each patch step of the two graphs, the historical data self-attention and the depth correlation self-attention The historical data self-attention and the depth correlation self-attention are merged in the last dimension to obtain the self-attention corresponding to the prior data The formula is as follows:

[0051]

[0052] Optionally, in an embodiment of the present invention, step 5 specifically includes:

[0053] Multiply the matrices of two dimensions to learn the adaptive adjacency matrix. The specific formula is as follows:

[0054]

[0055] where, A adp is the adaptive adjacency matrix, (·) T is the transpose operation;

[0056] Use the tensor obtained in step 3 with the learnable weight matrix to obtain the corresponding query Q, key K, and value V, and there is d adp = d′, d adp is the tensor using the adaptive graph self-attention weight matrix The dimensions of the tensor after processing are as follows:

[0057]

[0058] where Q adp is the query in the adaptive graph self-attention, is the learnable weight matrix corresponding to the query in the adaptive graph self-attention, K adp is the key in the adaptive graph self-attention, is the learnable weight matrix corresponding to the key in the adaptive graph self-attention, V adp is the value in the adaptive graph self-attention, is the learnable weight matrix corresponding to the value in the adaptive graph self-attention;

[0059] Obtain the adaptive self-attention Take the adaptive self-attention and fuse it with the prior data self-attention in step 4 to obtain the output result. The specific formula is as follows:

[0060]

[0061] where are learnable parameters, σ(·) represents the sigmoid activation function, z is the gating weight, is the tensor after being processed by the l-th layer of spatial-patch temporal block;

[0062] Take Repeat the operations in step 4 and step 5. After k layers, the input for the next step is obtained

[0063] Optionally, in an embodiment of the present invention, step 6 specifically includes:

[0064] Use the Flatten(·) operation to flatten the embedding dimension and the number of patch blocks into one dimension, and use a fully connected layer to project it to the future time step to be predicted. The specific formula is as follows:

[0065]

[0066] where X is the predicted value and T′ is the predicted time step;

[0067] Use the mean squared error as the loss function to measure the difference between the predicted value X and the true value X. The average loss function under the T′ predicted time step and N nodes is as follows:

[0068]

[0069] The traffic prediction method based on multi-view spatial patch self-attention according to the embodiments of the present invention obtains richer network traffic semantic information than a single time step by using patch operations, can accurately capture the changes and trends of time series data, improves the prediction accuracy of the model, and significantly reduces the computational complexity and memory occupancy. The present invention uses historical data graphs and depth correlation graphs, and combines adaptive graphs to obtain many indirect relationships between nodes beyond the basis of the original network graph. Through collaborative analysis, the extraction of the underlying network structure can be enhanced to an unprecedented depth and breadth.

[0070] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0071] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, wherein:

[0072] Figure 1 is a flowchart of a traffic prediction method based on multi-view spatial patch self-attention provided by an embodiment of the present invention;

[0073] Figure 2 is a comparison graph of the complexity of spatial patch attention and traditional spatio-temporal attention in an embodiment of the present invention;

[0074] Figure 3 is a schematic diagram of the overall framework of the model proposed in an embodiment of the present invention;

[0075] Figure 4 is a comparison graph of the prediction results of the model proposed in an embodiment of the present invention and a comparison algorithm;

[0076] Figures 5(a) and 5(b) are comparison graphs of the accuracy of the model proposed in an embodiment of the present invention and a comparison algorithm in 12 prediction intervals;

[0077] Figures 6(a), 6(b), 6(c) and 6(d) are visualization comparison results of the historical data graph, depth correlation graph, adaptive graph and original traffic graph obtained after actual training of the method in an embodiment of the present invention;

[0078] Figure 7 is a comparison graph of the prediction accuracy results of the method in an embodiment of the present invention after removing each module;

[0079] Figure 8 is a comparison graph of the results of the method in an embodiment of the present invention and a comparison model in terms of calculation speed and video memory occupancy;

[0080] Figures 9(a), 9(b) and 9(c) are diagrams showing the influence of different patch lengths on the MAE metric of the method according to the embodiments of the present invention on three datasets;

[0081] Figures 10(a), 10(b) and 10(c) are diagrams showing the influence of different step sizes on the MAE metric of the method according to the embodiments of the present invention on three datasets. Detailed implementation manners

[0082] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0083] Figure 1 It is a flowchart of a traffic prediction method based on multi-view spatial patch self-attention provided by an embodiment of the present invention.

[0084] As Figure 1 shown, the traffic prediction method based on multi-view spatial patch self-attention includes the following steps:

[0085] Step 1, establish a network traffic prediction problem.

[0086] Considering several nodes in the network and their associations, assuming there is an edge when two nodes are related, then the network can be modeled as a graph where represents the set of nodes and there are nodes, represents the set of edges and |ε| = E, and A represents the adjacency matrix and According to the collected traffic data the traffic at the t-th moment can be expressed as

[0087] The time stamp at the t-th moment can be expressed as t-T+1 ,..., X t t-T+1 ,..., y t and the network graph t+1 ,..., X t+T'

[0088] Step 2: Split the input multiple node flows into patch blocks through patch operation and input them into the flow mapping layer to obtain flow embeddings. Split the time information corresponding to the flows into patch blocks through patch operation and input them into the timestamp mapping layer to obtain timestamp embeddings.

[0089] For the flow data and timestamp information of each node, use patch operation to process them into patch blocks respectively. Let the length of each patch block be P l , and the starting interval between two adjacent patch blocks be S. Then the sequence corresponding to the i-th node after patch operation is If S values are added to the end of the original sequence before patch operation, the total number of patch blocks is After that, use a fully connected layer to map the flow patch sequence X p and the timestamp patch sequence Y p . The specific formula is:

[0090]

[0091] where FC(·) represents the fully connected layer, and the d -dimensional flow embedding and the timestamp embedding T emb can be obtained.

[0092] Step 3: Obtain the spatial embedding from the Laplacian eigenvectors of the network graph adjacency matrix, add the timestamp embedding and the spatial embedding to obtain the spatio-temporal patch embedding, add the spatio-temporal patch embedding and the flow embedding result, and use the resulting tensor as the input for the next step.

[0093] Figure 2 shows the comparison of the complexity between spatial patch attention and traditional spatio-temporal attention. To retain the spatial information in the network, first calculate the normalized Laplacian matrix using the adjacency matrix A, and the formula is as follows:

[0094] Δ = I - D -1 / 2 AD -1 / 2

[0095] where I is the identity matrix and D is the degree matrix. Then use eigenvalue decomposition, and the formula is as follows:

[0096]

[0097] where Λ is the eigenvalue matrix and U is the eigenvector matrix. Perform a linear projection on the m smallest non-singular eigenvectors in U to obtain the spatial embedding

[0098] Add the spatial embedding and the timestamp embedding through the broadcast mechanism to obtain the spatio-temporal patch embedding Adding it to the flow embedding, the input tensor in Step 3 can be obtained.

[0099]

[0100] Step 4: Pass the tensor obtained in Step 3 through the multi-head splitting conversion mechanism. Pass the two parts through the historical data space patch self-attention and the depth correlation space patch self-attention respectively, and connect and project the results to obtain the prior data self-attention.

[0101] Network diagram The adjacency matrix A of the ij is calculated according to the distance d between nodes i and j 2 through the following formula, where σ

[0102]

[0103] and ∈ are thresholds to control the sparsity of A: hd To capture deeper relationships between nodes, the differential dynamic time warping algorithm is used to calculate the historical data adjacency matrix A between each pair of nodes i,j = DDTW(X). For nodes i and j of length k, a matrix M is constructed, where the element m and represent the distance between two data points 1 This distance is measured by the squared difference between the estimated derivatives D. The formula for calculating the derivative of node i at time t

[0104]

[0105] The first and last elements of the sequence are obtained using the same calculation method as the second and second-to-last elements respectively. Subsequently, based on the dynamic programming algorithm, the optimal alignment distance dis and between the sequences can be determined. The following recursive formula for dis ij (t 1 , t 2 ) is as follows:

[0106]

[0107] In this way, the historical data adjacency matrix A can be constructed more accurately hd , where (A hd ) ij = dis ij (k, k).

[0108] The depth correlation adjacency matrix is obtained using the Katz centrality algorithm. Let the weight decay coefficient β = 0.1, and the formula is as follows:

[0109] A dr =(I - βA) -1 -I。

[0110] The Katz algorithm can be used to measure the strength of the relationship between two nodes because it not only detects direct connections but also takes into account the number and length of all paths connecting them through other nodes. By introducing a decay factor β, this algorithm reduces the noise brought by overly long paths. It is both simple and highly robust. Therefore, it can integrate the information of node degrees and edges, examine the entire network from a global perspective, and capture the indirect influence between nodes.

[0111] Input the tensor into the multi-head split transformation layer, and map it to d hd and d dr respectively through two different fully connected layers. Then, use the learnable weight matrices W Q , W K , W V to obtain the corresponding query Q, key K, and value V, and there is d hd + d dr = d′, and the formula is as follows:

[0112]

[0113] Performing self-attention on the entire spatio-temporal patch graph requires a large amount of memory. Therefore, for each Q at a specific patch moment, a list of K and V is generated from adjacent patches. Then, calculate the self-attention of each patch time step n times and average the results to obtain the spatio-temporal self-attention of the patch at that moment. The formula is as follows:

[0114]

[0115] where A ∈ {A hd , A dr}, and ⊙ represents the Hadamard product. After operating on each patch step of the two graphs, the historical data self-attention and the depth correlation self-attention are obtained. Merging the two in the last dimension can obtain the self-attention of the corresponding prior data The formula is as follows:

[0116]

[0117] Step 5: Pass the tensor obtained in Step 3 through the adaptive spatio-temporal patch self-attention mechanism and project it. Input it into the gated fusion mechanism together with the prior data self-attention obtained in Step 4, and use the resulting tensor as the input for the next step.

[0118] Input two matrices with smaller dimensions Multiply to learn the adaptive adjacency matrix to reveal the correlations beyond the prior information. The specific formula is as follows:

[0119]

[0120] For the tensor obtained in step 3 Use the learnable weight matrix to obtain the corresponding query Q, key K, and value V, and there is d adp = d′. The formula is as follows:

[0121]

[0122] After that, use the same formula in step 4 to obtain the adaptive self-attention

[0123] For the adaptive self-attention Fuse it with the prior data self-attention in step 4 to obtain the output result of this layer. The specific formula is as follows:

[0124]

[0125] where are learnable parameters, and σ(·) represents the sigmoid activation function, which can Repeat the operations in steps 4 and 5. After k layers, the input for the next step can be obtained

[0126] Step 6: Pass the result of step 5 through a patch flattening layer and a linear transformation layer in sequence to obtain the final traffic prediction result.

[0127] First, use the Flatten(·) operation to flatten the embedding dimension and the number of patch blocks into one dimension, and then use the fully connected layer to project it to the future time steps to be predicted. The specific formula is as follows:

[0128]

[0129] Use the mean squared error (MSE) as the loss function to measure the difference between the predicted value X and the true value X. The average loss function at the T′ prediction time steps and N nodes is as follows:

[0130]

[0131] Figure 3 Fig. shows the overall framework schematic diagram of the traffic prediction method proposed by the present invention. Next, the superiority of the calculation effect of the algorithm is verified through a specific implementation.

[0132] 1. Experimental parameter settings

[0133] The model proposed in the present invention is abbreviated as MSPformer. The prediction effect of the method of the present invention is tested on three real public transportation datasets, namely PeMS04, PeMS07, and PeMS08. In the PeMS04 dataset, there are 307 nodes, 340 edges, and 16,992 time-step data. In the PeMS07 dataset, there are 883 nodes, 866 edges, and 28,224 time-step data. In the PeMS08 dataset, there are 170 nodes, 295 edges, and 17,856 time-step data. In these three public datasets, the data collection interval for each piece of data is 5 minutes, and the data is divided into a training set, a validation set, and a test set according to a ratio of 6:2:2.

[0134] Meanwhile, in order to verify the effectiveness of the method of the present invention, the embodiments of the present invention select classical machine learning algorithms and the current state-of-the-art algorithms for comparison, including VAR, SVR, DCRNN, STGCN, STSGCN, STFGNN, STGODE, GMAN, PatchTST, ASTTN, GWNET, and STGNCDE.

[0135] 2. The effect of the algorithm of the present invention in improving the prediction accuracy

[0136] To fully illustrate the effectiveness of the algorithm of the present invention, consistent with the previous algorithms, the data of the past hour (12 steps) is used to predict the network traffic of the next hour (12 steps). The embodiments of the present invention use three widely adopted metrics to measure the prediction results, namely MAE, MAPE, and RMSE. These three metrics represent the difference between the predicted value and the true value, and the lower the value, the closer the prediction result is to the true value. From Figure 4 It can be seen from the prediction result comparison table that the algorithm proposed in the present invention always obtains superior or the best results on all datasets, and significantly improves the MAE compared with the second-best reported performance.

[0137] Embodiments of the present invention also plot the accuracies of five methods, namely GWNET, STSGCN, ASTTN, PatchTST, and MSPformer, on 12 different prediction intervals of the PeMS04 dataset, as shown in Figures 5(a) and 5(b). The results show that, in most cases, the model of the embodiments of the present invention performs better than other models. Among the five models, PatchTST performs the worst, followed by STSGCN. Although MSPformer is not as good as GWNET in short-term prediction, it shows stronger stability as the prediction range expands. As the prediction time lengthens, the MAE and MAPE errors of MSPformer increase less significantly. This advantage comes from the patch mechanism, which effectively captures the local semantic information in the network flow, thus greatly improving the accuracy of long-term traffic prediction. This indicates that MSPformer is particularly suitable for long-term pattern prediction and can provide a powerful solution to the challenges in long-term network traffic prediction.

[0138] 3. Influence of Each Part of the Invention on Prediction Accuracy

[0139] As Figures 6(a)-6(d) and Figure 7 shown, to study the influence of each part of the proposed model in the invention, embodiments of the present invention evaluate the influence of the model's prediction results on the PeMS04 dataset after deleting the following key elements: historical data self-attention, depth correlation self-attention, adaptive self-attention, spatio-temporal patch embedding, and patch mechanism. Removing SPA hd results in a significant increase in all metrics. Similarly, removing SPA dr also leads to a slight performance decline. Adaptive graph SPA is also very important, and removing it results in the highest RMSE. The most significant performance decline occurs when removing the spatio-temporal patch embedding (SPE), which indicates that SPE plays a key role in capturing the spatial features of network traffic data. In addition, the model performance decreases when the patch mechanism is missing, which proves the importance of the patch operation in capturing temporal patterns. In contrast, the complete MSPformer model performs best in all configurations. These results highlight the importance of each component in optimizing the model's ability to reveal hidden spatio-temporal relationships and accurately predict network traffic dynamics. From the prediction effects after deleting these components from the model proposed in the present invention, it can be seen that the patch mechanism enables the model to accurately capture the changes and trends of time series data and improves the prediction accuracy of the model. Combining historical data self-attention, depth correlation self-attention, and adaptive self-attention can jointly reveal hidden spatio-temporal relationships.

[0140] 4. Operating Efficiency of the Method of the Present Invention

[0141] As Figure 8As shown, the embodiments of the present invention evaluate the training and inference speeds of MSPformer for each training cycle on the PeMS04 dataset. The experimental results show that MSPformer has achieved significant improvements in terms of speed and memory efficiency compared to the model without patch operations. The patch mechanism compresses 12 time steps into 6 patch time steps, significantly reducing the time complexity. This reduction not only speeds up the training and inference processes but also reduces GPU memory usage, enabling MSPformer to utilize resources more efficiently.

[0142] Compared with other benchmark models, MSPformer with patches performs better in multiple aspects. Models like STSGCN and ASTTN have higher memory consumption and longer training times, while MSPformer achieves the best balance between resource consumption and prediction accuracy. Although MSPformer has more trainable parameters than ASTTN, it reduces memory usage, demonstrating its ability to capture spatio-temporal relationships with a more efficient structure without imposing too much burden on hardware resources. Although GWNET has lower memory consumption and fewer trainable parameters, it is slower than MSPformer in terms of training time and inference time. Due to its excellent efficiency in speed and resource usage, MSPformer is more suitable for practical deployment, especially in large-scale applications and edge computing environments with limited computing resources.

[0143] 5. Influence of patch length and stride in the present invention

[0144] As Figures 9(a)-9(c) and Figures 10(a)-10(c) shown, the embodiments of the present invention also evaluate the influence of patch length and stride on the model's prediction performance. The embodiments of the present invention fix the stride at 2 and vary the patch length within the range of P l ={2, 3, 4, 5, 6, 7}. The results show that for the PeMS04 and PeMS07 datasets, the model performs best when the patch length is 3 or 4. However, overall, the change in patch length has a relatively small impact on the prediction results, indicating that the model of the embodiments of the present invention has high robustness in terms of patch length. In addition, the embodiments of the present invention compare the influence of different stride values S = {1, 2, 3, 4, 5, 6} on the final MAE of three datasets. The results show that when the patch length is fixed at 3, the setting with a stride of 2 is significantly better than other stride settings, especially performing best on these datasets. However, the ideal stride value should be determined according to the specific characteristics of the dataset and the task requirements of the selected patch length.

[0145] The traffic prediction method based on multi-view spatial patch self-attention proposed according to the embodiments of the present invention integrates data from different perspectives to construct a more comprehensive representation. The network traffic data of multiple time steps is segmented into small pieces through patch operations, and spatio-patch block self-attention is used to jointly consider spatio-temporal correlations. This method not only retains the inherent local semantic information in the data, but also greatly reduces the requirements for computing power and memory usage. The experimental results on three public transportation datasets show that the model of the present invention achieves superior performance and computational efficiency.

[0146] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0147] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0148] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention belong.

[0149] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0150] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

Claims

1. A traffic prediction method based on multi-view spatial patch self-attention, characterized in that: The following steps are involved: Step 1, establish the network traffic prediction problem; Step 2: divide the input multiple node flows into patch blocks through patch operations, and input them into the flow mapping layer to obtain flow embedding; divide the time information corresponding to the flow into patch blocks through patch operations, and input them into the timestamp mapping layer to obtain timestamp embedding; Step 3: Get the spatial embedding from the Laplace eigenvector of the network graph adjacency matrix, add the timestamp embedding to the spatial embedding to get the space-time patch embedding, add the space-time patch embedding to the flow embedding, and use the obtained tensor as the input for the next step; Step 4: The tensor obtained in step 3 is converted through the multi-head splitting mechanism, and the results are concatenated and projected through the historical data space patch self-attention and the deep correlation space patch self-attention to obtain the prior data self-attention. Step 5: The tensor obtained in step 3 is projected through the adaptive spatial patch self-attention mechanism and input gated fusion mechanism with the prior data self-attention input in step 4, and the obtained tensor is used as the input for the next step; In step 6, the result of step 5 is passed through a patch flattening layer and a linear transformation layer in sequence to obtain the final flow prediction result.

2. The method according to claim 1, characterized in that Step 1 specifically includes: Considering multiple nodes in the network and the relationship between nodes, it is assumed that there is an edge when two nodes are related, so the network is modeled as a graph in, Represents a set of nodes and has the number of nodes represents the set of edges and |ε|=E, A represents the adjacency matrix and Based on the collected traffic data The flow rate at time t is expressed as The timestamp at time t is expressed as The network traffic prediction problem is based on T historical traffic data [X t-T+1 ,...,X t ], timestamp information [y t-T+1 ,...,y t ] and network diagram To infer the traffic data [X t+1 ,...,X t+T' ].

3. The method according to claim 2, characterized in that Step 2 specifically includes: For the traffic data and timestamp information of each node, use the patch operation to process them into patch blocks. Let the length of each patch block be P l , the starting interval between two adjacent patch blocks is S, then the sequence corresponding to the i-th node after the patch operation is If S values ​​are added to the end of the original sequence before the patch operation, the total number of patch blocks is Then use the fully connected layer to process the flow patch sequence X p and the time-stamped patch sequence Y p For mapping, the specific formula is: Among them, FC(·) represents the fully connected layer, and the d-dimensional flow embedding is obtained and timestamp embedded in T emb .

4. The method according to claim 3, characterized in that Step 3 specifically includes: In order to preserve the spatial information in the network, the network graph adjacency matrix A is used to calculate the normalized Laplacian matrix, as follows: Δ=I-D -1 / 2 AD -1 / 2 Where I is the identity matrix and D is the degree matrix; Using eigenvalue decomposition, the formula is as follows: Among them, Λ is the eigenvalue matrix, U is the eigenvector matrix; Linearly project the smallest m non-singular eigenvectors in U to obtain the spatial embedding The spatial embedding and the timestamp embedding are added through the broadcast mechanism to obtain the spatial-temporal patch embedding Then add it to the flow embedding to get the input tensor 5. The method according to claim 4, characterized in that In step 3, the network diagram The adjacency matrix A is based on the distance d between node i and node j ij Calculated, where σ 2 and ∈ are the thresholds for controlling the sparsity of A, and the calculation formula is: Use the differential dynamic time warping algorithm to calculate the historical data adjacency matrix A between each node hd =DDTW(X), for nodes i and j with length k, construct a matrix M, where element m i,j Represents two data points and The distance between them is measured by the squared difference between the estimated derivatives D. The formula for calculating the derivative of node i at time t1 is as follows: The first and last elements of the sequence are obtained using the same calculation method as the second and penultimate elements, respectively; Based on the dynamic programming algorithm, determine the sequence and The optimal alignment distance dis ij The recursive formula for (t1, t2) is: away ij (t1,t2)={dis ij (t1-1,t2-1),dis ij (t1-1,t2),dis ij (t1,t2-1)} Construct the historical data adjacency matrix A hd , (A hd ) ij =dis ij (k,k); The Katz centrality algorithm is used to obtain the deep correlation adjacency matrix, and the weight decay coefficient β is set to 0.

1. The formula is as follows: TO dr =(I-βA) -1 -THE.

6. The method according to claim 5, characterized in that Step 4 specifically includes: The tensor Input multi-head segmentation conversion layer, mapped to d through the historical data graph fully connected layer and the deep correlation graph fully connected layer respectively. hd and d dr , d hd is the dimension corresponding to the fully connected layer mapping of the historical data graph, d dr To correspond to the dimension of the fully connected layer mapped to the deep correlation graph, a learnable weight matrix W is used Q ,W K ,W V To obtain the corresponding query Q, key K and value V, and there are d hd +d dr = d′, d′ is the tensor of the fully connected layer of the historical data graph and the fully connected layer of the deep correlation graph After the mapping operation is performed, the sum of the dimensions of the two tensors is obtained. The formula is as follows: in, To use the historical data graph fully connected layer to tensor The tensor obtained after the mapping operation, To use the historical data graph fully connected layer to tensor Perform mapping operations, To use the deep correlation graph fully connected layer on the tensor The tensor obtained after the mapping operation, To use the deep correlation graph fully connected layer on the tensor Perform mapping operation, Q hd is the query in the historical data graph self-attention, is the learnable weight matrix corresponding to the query in the historical data graph self-attention, K hd is the key in the historical data graph self-attention, is the learnable weight matrix corresponding to the key in the historical data graph self-attention, V hd is the value in the self-attention of the historical data graph, is the learnable weight matrix corresponding to the median of the self-attention of the historical data graph, Q dr is the query in the deep relevance graph self-attention, is the learnable weight matrix corresponding to the query in the deep correlation graph self-attention, K dr is the key in the deep correlation graph self-attention, is the learnable weight matrix corresponding to the key in the deep correlation graph self-attention, V dr is the value in the self-attention of the deep correlation graph, is the learnable weight matrix corresponding to the median of the self-attention of the deep correlation graph; For each Q at a specific patch moment, a list of K and V is generated from neighboring patches, and then the self-attention of each patch time step is calculated n times and the results are averaged to obtain the spatial patch time self-attention at that moment, as follows: Among them, A∈{A hd ,A dr }, ⊙ represents the Hadamard product, Q i is the i-th query in the spatial patch temporal self-attention, is the transpose of the pth key in the spatial patch temporal self-attention, V p is the pth value in the temporal self-attention of the spatial patch; After operating each patch step of the two images, we can get the historical data self-attention and Deep Correlation Self-Attention Attention to historical data and Deep Correlation Self-Attention Merge in the last dimension to get the self-attention corresponding to the prior data The formula is as follows:

7. The method according to claim 6, characterized in that Step 5 specifically includes: The two-dimensional matrix Multiply to learn the adaptive adjacency matrix. The specific formula is as follows: Among them, A adp is the adaptive adjacency matrix, (·) T is the transpose operation; The tensor obtained in step 3 Using a learnable weight matrix To obtain the corresponding query Q, key K and value V, and there are d adp =d′,d adp To use the adaptive graph self-attention weight matrix for the tensor The dimension of the tensor after processing is as follows: Among them, Q adp is the query in adaptive graph self-attention, is the learnable weight matrix corresponding to the query in adaptive graph self-attention, K adp is the key in adaptive graph self-attention, is the learnable weight matrix corresponding to the key in adaptive graph self-attention, V adp is the value in the adaptive graph self-attention, is the learnable weight matrix corresponding to the median of the adaptive graph self-attention; Getting Adaptive Self-Attention Adaptive Self-Attention Self-attention with prior data in step 4 Fusion, get the output result, the specific formula is as follows: in, is a learnable parameter, σ(·) represents the sigmoid activation function, z is the gate weight, is the tensor processed by l layers of space-patch time blocks; Will Repeat the operations in steps 4 and 5, and get the input for the next step after k layers.

8. The method according to claim 7, characterized in that Step 6 specifically includes: The Flatten(·) operation is used to flatten the embedding dimension and the number of patches into one dimension, and a fully connected layer is used to project it to the future moment that needs to be predicted. The specific formula is as follows: Among them, X is the predicted value, T′ is the prediction time step; Using mean squared error as the loss function to measure the difference between the predicted value X and the true value X, the average loss function under T′ prediction time steps and N nodes is as follows: