Multi-scale space-time fusion image feature extraction method based on traffic flow

Through dynamic graph generation and learning weighting mechanism combined with Transformer module, the problem that the fixed adjacency matrix cannot adapt to the time-varying characteristics of traffic flow in the traditional method is solved, and high-precision traffic flow prediction is achieved.

CN120336804APending Publication Date: 2025-07-18HANGZHOU DIANZI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510388860.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing traffic flow prediction methods are difficult to effectively capture the highly dynamic nature of traffic flow and the complex spatiotemporal interactions, especially the traditional fixed adjacency matrix cannot adapt to the time-varying characteristics of traffic flow.

Method used

A multi-scale spatiotemporal fusion graph feature extraction method based on traffic flow is adopted, and an adaptive adjacency matrix is generated through a dynamic graph generation module combining historical data, spatial embedding and temporal embedding. The learning weighting mechanism and the Transformer module are used to enhance the spatiotemporal modeling capabilities, and feature representation is optimized through a sequence feature mapper.

Benefits of technology

It significantly improves the accuracy and stability of traffic flow prediction, can flexibly capture static and dynamic traffic relationships, and improves the model's modeling ability and prediction accuracy of complex road network structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336804A_ABST
    Figure CN120336804A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent traffic, and discloses a multi-scale space-time fusion image feature extraction method based on traffic flow, which comprises the following steps of: 1, constructing a dynamic image generation module; a self-adaptive adjacency matrix is generated in combination with historical traffic data, spatial embedding and time embedding, the matrix is used for spatial-temporal feature extraction of a GCN layer, and the generated adjacency matrix can flexibly capture static and dynamic relationships; 2, constructing a learnable weighting module; according to the module, the weights of different features are adaptively adjusted, so that various feature information is effectively fused; 3, constructing a sequence feature mapper; through a recurrent neural network (RNN) and a sequence compression-expansion mechanism, dynamic features of time sequence signals are effectively extracted and enhanced. According to the method, adaptive adjustment of an adjacent matrix is realized through a dynamic graph generation module, the multi-feature fusion capability is improved in combination with a learnable weighting mechanism, and the long sequence modeling capability of an RNN layer is optimized by adopting a sequence compression and expansion strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent transportation, and particularly relates to a multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow. Background Art

[0002] Traffic flow prediction is a core issue in intelligent transportation systems, involving time series modeling, spatial dependence modeling, and multi-source information fusion. Traditional methods mainly rely on statistical models such as autoregressive integrated moving average (ARIMA) and support vector regression (SVR). These methods perform well in short-term prediction but cannot effectively capture complex spatio-temporal dependencies.

[0003] In recent years, deep learning methods have made significant progress in the field of traffic flow prediction. Recurrent neural networks (RNNs) and their variants, long short-term memory networks (LSTMs) and gated recurrent units (GRUs), have been widely used to model time dependencies. However, these methods are difficult to model spatial relationships, so some studies combine convolutional neural networks (CNNs) to capture spatial patterns. With the rise of graph neural networks (GNNs), researchers have proposed a series of traffic flow prediction methods based on graph neural networks. Spatio-temporal graph convolutional network (ST-GCN) combines graph convolutional network (GCN) with sequence models to model spatio-temporal features. Graph attention network (GAT) further introduces an attention mechanism to dynamically adjust the influence weights of the adjacency matrix, improving the prediction performance. However, existing GNN methods usually rely on a fixed adjacency matrix and are difficult to capture dynamic traffic changes.

[0004] In the real world, the spatial relationship of traffic flow is dynamically changing. Therefore, the limitations of static adjacency matrices have prompted researchers to explore dynamic graph modeling. In recent years, researchers have proposed various dynamic graph construction methods, such as dynamic graph convolutional networks based on reinforcement learning, dynamic graph learning based on attention mechanisms, and adaptive adjacency matrices combined with external factors. Among them, dynamic graph neural network (DGNN) introduces a learnable graph structure, enabling the adjacency matrix to be adaptively adjusted over time. Transformer-based dynamic graph modeling uses self-attention mechanisms to learn spatio-temporal dependencies between nodes globally. Although these methods have made significant progress in dynamic graph modeling, there are still problems such as high computational complexity and difficulty in training the model. Summary of the Invention

[0005] Aiming at the problem that traditional methods are difficult to effectively capture the highly dynamic nature and complex spatio-temporal interaction relationships of traffic flow, a multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow is proposed. High-precision traffic flow prediction is completed through a series of optimized neural network operations. Temporal embedding and node embedding are performed on historical data to extract monthly, weekly, and current temporal patterns. At the same time, spatial node features are combined to support subsequent feature extraction. The dynamic graph generation and learnable weighting mechanism are used, combined with graph convolution operations, to capture dynamic spatio-temporal relationships and generate high-quality feature representations. The feature mapper optimizes the feature representations, and the prediction module maps the optimized features to the prediction space through a fully connected layer, and finally outputs accurate traffic flow prediction results.

[0006] To achieve the above-mentioned invention purpose, the present invention adopts the following technical solutions:

[0007] A multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow, comprising the following steps:

[0008] Step 1: Construct a dynamic graph generation module; the dynamic graph generation module combines historical traffic data, spatial embedding E v and temporal embedding E t to generate an adaptive adjacency matrix A t , which is used for spatio-temporal feature extraction in the GCN layer. The generated adjacency matrix can flexibly capture static and dynamic relationships;

[0009] Step 2: Construct a learnable weighting module; this module effectively fuses various feature information by adaptively adjusting the weights of different features, thereby improving the expression ability of the model;

[0010] Step 3: Construct a sequence feature mapper; this module effectively extracts and enhances the dynamic features of time series signals through a recurrent neural network RNN and a sequence compression-expansion mechanism.

[0011]

[0012] Furthermore, the first step includes the generation of a spatial graph, the generation of a temporal graph, and the combination of the spatial-temporal graph, and finally optimizes the generation result through a Transformer module;

[0013] The generation of the spatial graph is calculated based on the mutual relationship between nodes. By performing linear transformation and non-linear activation on the node embedding, the feature representation of the nodes can be obtained. These feature representations calculate a dense inter-node relationship matrix through matrix multiplication. Each element of this matrix represents the similarity or association strength between two nodes. Non-linear mapping is performed on the node embedding, and two layers of learnable transformations are adopted to enhance the expression ability:

[0014] Point embedding is non-linearly mapped, and two layers of learnable transformations are used to enhance the expression ability:

[0015] Hv = σ(W1Ev + b1)

[0016] H′ v = σ(W1Hv + b2)

[0017] where is the initial node embedding matrix, N is the number of nodes, and d is the embedding dimension. is the learnable weight matrix, b1 and b2 are bias terms, σ(·) is the non - linear activation function, and H′ v is the final node representation;

[0018] Then, calculate the similarity between nodes, using Gaussian kernel similarity instead of simple dot - product calculation:

[0019]

[0020] where is the spatial adjacency matrix, and σ is the learnable smoothing parameter;

[0021] To retain the most important connection relationships in the graph, sparsify the adjacency matrix through the Top - k strategy:

[0022]

[0023]

[0024] where is the mask matrix generated through the Top - k strategy, and ⊙ represents the element - wise multiplication operation;

[0025] The generation of the temporal graph is based on temporal embedding calculation, modeling different time granularities separately:

[0026] T d = softmax(W d X t )

[0027] T w = softmax(W w X t )

[0028] A t = αT d Td T + βT w T w T

[0029] where is the temporal feature embedding matrix, X t is the temporal input, W d , W w are learnable transformation matrices, At is the temporal adjacency matrix;

[0030] Similarly, the Top-k strategy is also applied to sparsify the temporal graph:

[0031]

[0032] After the spatial graph and the temporal graph are respectively generated, the module adopts spectral transformation for fusion to enhance the dynamic modeling ability:

[0033]

[0034] wherein, represents the element-wise Hadamard product, is the Laplacian matrix, which enhances the smoothness of the graph, and γ, λ are learnable parameters.

[0035] Finally, the dynamic adjacency matrix is obtained:

[0036]

[0037] Furthermore, the first step also introduces a Transformer layer to enhance the spatio-temporal modeling ability by using multi-head attention:

[0038]

[0039] where Q = W Q H v K = W k H v V = W v H v are the query, key, and value matrices, W Q W k W v are learnable weights, and d k is the dimension normalization factor.

[0040] Furthermore, the second step includes the following steps:

[0041] Weighted node embedding: Let represent the node embedding, where N is the number of nodes, and d n is the embedding dimension. A learnable weight vector is introduced and the weighted node embedding is calculated after Softmax normalization:

[0042] W n = diag(Softmax(w n ))

[0043]

[0044] Weighted Time Embedding: Similar to node embedding, a learnable time weight is used to calculate the weighted time features:

[0045] W t = diag(Softmax(w t ))

[0046]

[0047] where and represent time embeddings generated based on time within a day and time within a week respectively, and d t is the dimension of the time embedding;

[0048] Historical Feature Mapping: Map the historical traffic state features through a trainable linear transformation to align with the weighted node embedding and time embedding:

[0049]

[0050] where, is the historical traffic state feature, is the trainable linear transformation matrix, and d m is the final feature dimension;

[0051] Feature Fusion: Finally, perform a weighted sum of features from different sources to obtain the fused comprehensive feature representation:

[0052]

[0053] Furthermore, step three includes the following steps:

[0054] Sequence Feature Extraction: Let the input time series signal be where B is the batch size, L is the time step, N is the number of nodes, and d h is the feature dimension. First, use a one-dimensional convolutional layer to compress the sequence to capture broader time context information:

[0055] X compressed = Conv1D(X)

[0056] where Conv1D(·) represents the one-dimensional convolution operation for local feature extraction in the time dimension, reducing the data dimension and improving the computational efficiency;

[0057] Temporal Modeling RNN Layer: After sequence compression, use a gated recurrent unit GRU to model the temporal dependency. Let the hidden state of GRU be h t , and its update formula is as follows:

[0058] z t = σ(W z X t + U z h t-1 + b z )

[0059] r t = σ(W r X t + U r h t-1 + b r )

[0060]

[0061] where σ(·) represents the Sigmoid activation function, ⊙ represents element-wise multiplication, z t is the update gate, r t is the reset gate, is the candidate hidden state;

[0062] In addition, layer normalization LayerNorm is applied at each time step to stabilize training, and Dropout is introduced to prevent overfitting:

[0063] h t = LayerNorm(h t )

[0064] h t = Dropout(h t )

[0065] Sequence recovery: One-dimensional transposed convolution ConvTranspose1D is used for sequence expansion:

[0066] X expanded = ConvTranspose1D(h t )

[0067] X final = Dropout(X expanded )

[0068] Finally, the enhanced temporal feature representation is obtained:

[0069] X output = X final .

[0070] Furthermore, step four includes the following steps:

[0071] All experiments were run on a server equipped with an NVIDIA RTX 4090 GPU, an Intel Xeon processor, and 128 GB of RAM. The deep learning framework used was PyTorch, the optimizer was Adam, and mixed precision was used to accelerate training. The learning rate was set to 0.002, the batch size was set to 16, the weight decay coefficient was 1e-5, 300 epochs of training were performed, and an early stopping strategy was used to prevent overfitting;

[0072] Standardization was used to normalize the data so that all feature values were normalized to the range [0, 1] to improve the convergence of the model; at the same time, the data was divided into a training set, a validation set

[0073] validation set, and a test set in a ratio of 6:2:2. The traffic data for the first 12 timesteps was used to predict the next 12 timesteps;

[0074] Three common error metrics were used for quantitative analysis: Mean Absolute Error, Root Mean Squared Error, and Mean Absolute Scaled Error. The definitions of these metrics are as follows:

[0075]

[0076] where, x t is the true value, is the predicted value, |S| is the size of the sample set, MSE penalizes large deviations through squared errors and is sensitive to outliers; MAPE is expressed as a percentage but is unstable for values close to zero; MAE measures the absolute deviation, is stable and has a consistent unit, but penalizes large errors less.

[0077] The multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow of the present invention has the following advantages:

[0078] 1. A dynamic graph generation module combining spatio-temporal relationships is proposed:

[0079] The present invention proposes a new dynamic graph generation module that combines spatio-temporal relationships. This module captures the dynamic change patterns in traffic flow data in real time through an adaptive adjacency matrix mechanism. The adaptive adjacency matrix is jointly driven by spatial topology information and time series features, and can adaptively adjust the network structure according to the traffic flow changes in different time periods, dynamically modeling the traffic flow dependence relationships at different time scales. Through this dynamic graph generation module, the generated adjacency matrix can flexibly capture static and dynamic traffic relationships, overcoming the limitation that traditional fixed adjacency matrices cannot adapt to the time-varying characteristics of traffic flow, thus significantly improving the model's ability to model complex road network structures and enhancing the accuracy and stability of traffic flow prediction.

[0080] 2. An adaptive weight learning module is designed:

[0081] The present invention proposes a new adaptive weight learning module. By adaptively adjusting the weights of various features, including node features, time features, historical features, etc., this method can ensure that the model can effectively capture information in different dimensions, improve the ability to model spatio-temporal relationships, and provide more accurate feature inputs for subsequent prediction tasks. Through ablation experiments, we found that the impact of this adaptive weight learning module on the model is the greatest.

[0082] 3. An efficient sequence feature mapper is designed:

[0083] The present invention designs an efficient sequence feature mapper. It is used to enhance the expression ability of traffic flow time series features and improve the accuracy and robustness of long-time series modeling. This feature mapper adopts a strategy that combines a convolutional neural network (CNN) and a recurrent neural network (RNN), making full use of the advantages of both. The CNN part is responsible for extracting high-order features within a local time window, using one-dimensional convolution (1D-CNN) to model traffic flow patterns on a short time scale, enhancing the local feature expression ability, and improving the computational efficiency at the same time. The RNN part uses a gated recurrent unit (GRU) or a long short-term memory network (LSTM) to capture global time series dependencies over a longer time span, thus enhancing the ability to predict long-term traffic flow trends. By combining convolution and recurrent neural networks, the long-time series modeling ability is enhanced and the prediction accuracy is improved.

[0084] 4. A sequence compression and expansion optimization strategy is designed:

[0085] In view of the challenges of long-time series modeling, the present invention proposes a sequence compression and expansion strategy to optimize the learning ability of the RNN layer for long-term dependence relationships, avoid the problem of gradient disappearance, and improve computational efficiency. In addition, the Squeeze-and-Excitation (SE) attention mechanism is introduced into the convolutional network to enhance the feature representation ability of the model. This strategy not only improves the modeling ability of the RNN layer for long-time series, but also strengthens the representation ability of key features through the SE mechanism, enabling the model to have stronger spatio-temporal correlation modeling ability in complex traffic scenarios, thereby improving the accuracy and stability of traffic flow prediction. Description of the Drawings

[0086] Figure 1 It is a model architecture diagram of the present invention.

[0087] Figure 2 It is a dynamic graph generation module of the present invention.

[0088] Figure 3 It is an adaptive weight learning module of the present invention.

[0089] Figure 4 It is a sequence feature mapper module of the present invention. Detailed Description of the Invention

[0090] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow of the present invention with reference to the drawings.

[0091] As Figure 1 shown, the multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow of the present invention includes the following steps:

[0092] Step 1: Construct a dynamic graph generation module.

[0093] The dynamic graph generation module combines historical traffic data, spatial embedding E v and temporal embedding E t to generate an adaptive adjacency matrix A t , which is used for spatio-temporal feature extraction of the GCN layer. The generated adjacency matrix can flexibly capture static and dynamic relationships, thereby enhancing the modeling ability of the model for complex network structures. This process is divided into three steps: generation of the spatial graph, generation of the temporal graph, and combination of the spatial-temporal graph, and finally the generation result is optimized through the Transformer module, as Figure 2 shown.

[0094] The generation of the spatial graph is calculated based on the mutual relationships between nodes. By performing linear transformation and non-linear activation on the node embeddings, we can obtain the feature representations of the nodes. These feature representations are used to calculate a dense matrix of relationships between nodes through matrix multiplication, where each element of the matrix represents the similarity or association strength between two nodes.

[0095] For non-linear mapping of the node embeddings, we adopt two layers of learnable transformations to enhance the expressive power:

[0096] H v = σ(W1E v + b1)

[0097] H′ v = σ(W1H v + b2)

[0098] where is the initial node embedding matrix, N is the number of nodes, and d is the embedding dimension. is the learnable weight matrix, b1 and b2 are bias terms. σ(·) is the non-linear activation function. H′ v is the final node representation.

[0099] Then, to calculate the similarity between nodes, we use Gaussian Kernel Similarity instead of simple dot product calculation:

[0100]

[0101] where is the spatial adjacency matrix, and σ is the learnable smoothing parameter.

[0102] To retain the most important connection relationships in the graph, we sparsify the adjacency matrix through the Top-k strategy:

[0103]

[0104] where is the mask matrix generated through the Top-k strategy, and ⊙ represents the element-wise multiplication operation.

[0105] The generation of the temporal graph is based on time embedding calculation. We model different time granularities (such as days, weeks) separately:

[0106] T d = softmax(W d X t )

[0107] T w = softmax(W w X t )

[0108] A t = αT d T d T + βT w T w T

[0109] Among them, is the time feature embedding matrix, X t is the time input, W d , W w is the learnable transformation matrix, A t is the time adjacency matrix.

[0110] Similarly, the Top-k strategy is applied to sparsify the time graph:

[0111]

[0112] After the spatial graph and the time graph are respectively generated, the module adopts spectral transformation for fusion to enhance the dynamic modeling ability:

[0113]

[0114] Among them, represents the element-wise Hadamard product, is the Laplacian matrix, which enhances the smoothness of the graph, and γ, λ are learnable parameters.

[0115] Finally, we obtain the dynamic adjacency matrix:

[0116]

[0117] To further enhance the representation ability of the dynamic graph, the module also introduces a Transformer layer to enhance the spatio-temporal modeling ability by using multi-head attention:

[0118]

[0119] Among them, Q = W Q H v , K = W k H v , V = W v H v are the query, key, and value matrices. W Q , W k , W v are the learnable weights, and d k is the dimension normalization factor.

[0120] Step 2: Construct a learnable weighting module.

[0121] In the traffic flow prediction task, the contributions of different temporal features, spatial node features, and historical traffic states to the prediction vary. Therefore, we designed a Learnable Weighted Module, as Figure 3 shown, which effectively fuses multiple feature information and improves the model's representational ability by adaptively adjusting the weights of different features, including node features, temporal features, historical features, etc.

[0122] Weighted Node Embedding: Let represent the node embedding, where N is the number of nodes and d n is the embedding dimension. We introduce a learnable weight vector and calculate the weighted node embedding after Softmax normalization:

[0123] W n = diag(Softmax(w n ))

[0124]

[0125] Weighted Temporal Embedding: Similar to node embedding, we use a learnable temporal weight to calculate the weighted temporal feature:

[0126] W t = diag(Softmax(w t ))

[0127]

[0128] where and represent the temporal embeddings generated based on the time of day and the time of week respectively, and d t is the temporal embedding dimension.

[0129] Historical Feature Mapping: Map the historical traffic state features through a trainable linear transformation to align with the weighted node embedding and temporal embedding:

[0130]

[0131] where, is the historical traffic state feature, is the trainable linear transformation matrix, and d m is the final feature dimension.

[0132] Feature Fusion: Finally, we perform a weighted sum of the features from different sources to obtain the fused comprehensive feature representation:

[0133]

[0134] This module adaptively adjusts the weights of each feature to ensure that the model can effectively capture information in different dimensions, improve the ability to model spatio-temporal relationships, and provide more accurate feature inputs for subsequent prediction tasks.

[0135] Step 3: Construct a sequence feature mapper.

[0136] In traffic flow prediction tasks, the long-term dependence relationship of time series data is crucial. Therefore, we designed a sequence feature mapper (Feature Mapper), as Figure 4 shown. This module effectively extracts and enhances the dynamic features of time series signals through a recurrent neural network (RNN) and a sequence compression-expansion mechanism.

[0137] Sequence feature extraction: Let the input time series signal be where B is the batch size, L is the time step, N is the number of nodes, and d h is the feature dimension. First, we use a one-dimensional convolutional layer to compress the sequence to capture broader time context information:

[0138] X compressed = Conv1D(X)

[0139] where Conv1D(·) represents a one-dimensional convolutional operation for local feature extraction in the time dimension, reducing the data dimension and improving computational efficiency.

[0140] Temporal modeling (RNN layer): After sequence compression, we use a gated recurrent unit (GRU) to model the temporal dependence relationship. Let the hidden state of the GRU be h t , and its update formula is as follows:

[0141] z t = σ(W z X t + U z h t-1 + b z )

[0142] r t = σ(W r X t + U r h t-1 + b r )

[0143]

[0144] where, σ(·) represents the Sigmoid activation function, ⊙ represents element-wise multiplication, and z t is the update gate, r t is the reset gate, is the candidate hidden state.

[0145] In addition, we apply layer normalization (LayerNorm) at each time step to stabilize the training, and introduce Dropout to prevent overfitting:

[0146] h t = LayerNorm(h t )

[0147] h t = Dropout(h t )

[0148] Sequence recovery (deconvolution expansion): After the features are processed by the GRU, we need to restore the original time series length to ensure data integrity. For this purpose, we use one-dimensional deconvolution (ConvTranspose1D) for sequence expansion:

[0149] X expanded = ConvTranspose1D(h t )

[0150] X final = Dropout(X expanded )

[0151] This operation can ensure that the dimension of the final output is consistent with the input, while ensuring that the time-dependent information of the features is completely retained.

[0152] Finally, the enhanced temporal feature representation is obtained:

[0153] X output = X final

[0154] Our feature mapper significantly improves the ability to model long sequence dependencies through a combination of convolutional compression + GRU + deconvolution expansion, while maintaining computational efficiency and the generalization ability of the model. This design enables the model to more accurately learn the patterns of traffic flow changes over time, thereby improving the accuracy of prediction.

[0155] Step 4: Model training

[0156] All experiments were run on a server equipped with an NVIDIA RTX 4090 GPU, an Intel Xeon processor, and 128 GB of RAM. The deep learning framework used was PyTorch, the optimizer was Adam, and Mixed Precision Training was used to accelerate training. The learning rate was set to 0.002, the batch size was set to 16, the weight decay coefficient was 1e-5, 300 epochs of training were performed, and the Early Stopping strategy was used to prevent overfitting.

[0157] To ensure the fairness of the experiments, we normalized the data using Min-Max Scaling to normalize all feature values to the interval [0,1] to improve the convergence of the model. At the same time, we divided the data into a training set, a validation set, and a test set in a ratio of 6:2:2, and used the traffic data of the first 12 timesteps to predict the next 12 timesteps.

[0158] To comprehensively evaluate the model performance, we used three common error metrics for quantitative analysis: Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and Mean Absolute Scaled Error (MAPE). The definitions of these metrics are as follows:

[0159]

[0160] where x t is the true value, is the predicted value, and |S| is the size of the sample set. MSE penalizes large deviations through squared errors and is sensitive to outliers; MAPE is expressed as a percentage, which is convenient for comparison, but is unstable for values close to zero; MAE measures the absolute deviation, is stable and has a consistent unit, but penalizes large errors less.

[0161] Step Five: Experimental Results

[0162] The following table shows the prediction performance of different models on four datasets. It includes the Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and Mean Absolute Percentage Error (MAPE).

[0163]

[0164] From which we can draw the following conclusions:

[0165] (1) The DGMFN outperforms other models on all datasets, indicating that the model has stronger generalization ability and prediction accuracy in traffic flow prediction tasks.

[0166] (2) The traditional time series methods (HA and ARIMA) perform the worst as they are difficult to capture complex spatio-temporal dependencies. As a simple time series modeling method, LSTM has improvement compared to traditional methods, but still lags significantly behind the models combined with graph neural networks, suggesting that relying solely on information in the time dimension is insufficient to accurately predict traffic flow.

[0167] (3) Among the models related to graph neural networks (GNN), methods such as STGCN, STSGCN, Z-GCNETs, and STFGNN perform prominently and are all better than LSTM and traditional methods, indicating that GNN can effectively model spatial relationships in traffic networks. Among them, Z-GCNET and STFGNN achieve relatively good results, while DGMFN further reduces the prediction error on this basis, showing that the mechanisms such as dynamic graph generation (DG) and learnable weighted fusion (LW) adopted by it play a key role in optimizing spatio-temporal feature extraction.

[0168] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow, characterized in that It includes the following steps: Step 1: Construct a dynamic graph generation module; The dynamic graph generation module combines historical traffic data, spatial embedding E v and temporal embedding E t to generate an adaptive adjacency matrix A t , which is used for spatio-temporal feature extraction in the GCN layer, and the generated adjacency matrix can flexibly capture static and dynamic relationships; Step 2: Construct a learnable weighting module; this module effectively fuses various feature information by adaptively adjusting the weights of different features, thereby improving the expressive ability of the model; Step 3: Construct a sequence feature mapper; This module effectively extracts and enhances the dynamic features of time series signals through a recurrent neural network (RNN) and a sequence compression-expansion mechanism.

2. The multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow according to claim 1, wherein The first step includes the generation of a spatial graph, the generation of a temporal graph, and the combination of the spatial-temporal graph, and finally optimizes the generation result through a Transformer module; The generation of the spatial graph is calculated based on the mutual relationship between nodes. By performing linear transformation and non-linear activation on the node embeddings, the feature representations of the nodes can be obtained. These feature representations calculate a dense inter-node relationship matrix through matrix multiplication. Each element of this matrix represents the similarity or association strength between two nodes. Non-linear mapping is performed on the node embeddings, and two layers of learnable transformations are adopted to enhance the expressive ability: H v = σ(W1E v + b1) H′ v = σ(W1H v + b2) Among them, is the initial node embedding matrix, N is the number of nodes, and d is the embedding dimension. is the learnable weight matrix, b1 and b2 are bias terms, σ(·) is the non-linear activation function, and H′ v is the final node representation; Then, calculate the similarity between nodes, and use Gaussian kernel similarity instead of simple dot product calculation: Among them, is the spatial adjacency matrix, and σ is the learnable smoothing parameter; To retain the most important connection relationships in the graph, sparsify the adjacency matrix through the Top-k strategy: Among them, is a mask matrix generated by the Top-k strategy, and ⊙ represents the element-wise multiplication operation; The generation of the temporal graph is calculated based on temporal embeddings, and different time granularities are modeled separately: T d = softmax(W d X t ) T w = softmax(W w X t ) A t = αT d T d T + βT w T w T Among them, T d , is the time feature embedding matrix, X t is the time input, W d , W v are learnable transformation matrices, and A t is the time adjacency matrix; Similarly, apply the Top-k strategy to sparsify the temporal graph: After the spatial graph and the temporal graph are generated respectively, the module uses spectral transformation for fusion to enhance the dynamic modeling ability: Among them, represents the element-wise Hadamard product, is the Laplacian matrix, which enhances the smoothness of the graph, and γ, λ are learnable parameters. Finally, obtain a dynamic adjacency matrix:

3. The multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow according to claim 2, wherein The first step also introduces a Transformer layer to enhance the spatio-temporal modeling ability using multi-head attention: where Q = W Q H v , K = W k H v , V = W v H v is the query, key, value matrix, W Q , W k , W v is the learnable weight, and d k is the dimensional normalization factor.

4. The method for extracting multi-scale spatio-temporal fusion graph features based on traffic flow according to claim 2, wherein The second step includes the following steps: Weighted Node Embedding: Let denote the node embedding, where N is the number of nodes and d n is the embedding dimension. Introduce a learnable weight vector After normalizing by Softmax, calculate the weighted node embedding: W n = diag(softmax(w n )) Weighted time embedding: Similar to node embedding, use a learnable time weight to calculate weighted time features: W t = diag(Softmax(w t )) wherein and respectively represent time embeddings generated based on the time within a day and the time within a week, and d t is the dimension of the time embedding; Historical feature mapping: Map the historical traffic state features through a trainable linear transformation to align with the weighted node embeddings and temporal embeddings: Among them, is the historical traffic state feature, is a learnable linear transformation matrix, and d m is the final feature dimension; Feature fusion: Finally, perform weighted summation on features from different sources to obtain a fused comprehensive feature representation:

5. The multi-scale spatio-temporal fusion graph feature extraction method based on traffic flow according to claim 2, wherein The third step includes the following steps: Sequence feature extraction: Let the input time series signal be where B is the batch size, L is the time step, N is the number of nodes, and d h is the feature dimension. First, use a one-dimensional convolutional layer to compress the sequence to capture more extensive time context information: X compressed = Conv1D(X) Among them, Conv1D(·) represents a one-dimensional convolution operation, which is used to perform local feature extraction in the time dimension, reduce the data dimension, and improve the calculation efficiency; Temporal Modeling RNN Layer: After sequence compression, a gated recurrent unit (GRU) is used to model temporal dependencies. Let the hidden state of the GRU be h t , and its update formula is as follows: z t = σ(W z X t + U z h t-1 + b z ) r t = σ(W r X t + U r h t-1 + b r ) where, σ(·) represents the Sigmoid activation function, ⊙ represents element-wise multiplication, and z t is the update gate, r t is the reset gate, is the candidate hidden state; In addition, apply layer normalization (LayerNorm) at each time step to stabilize the training, and introduce Dropout to prevent overfitting: h t = LayerNorm(h t ) h t = Dropout(h t ) Sequence recovery: Use one-dimensional transposed convolution (ConvTranspose1D) for sequence expansion: X expanded = ConvTranspose1D(h t ) X final = Dropout(X expanded ) Finally, obtain an enhanced temporal feature representation: X output = X final .

6. The method for extracting multi-scale spatio-temporal fusion graph features based on traffic flow according to claim 2, wherein The fourth step includes the following steps: All experiments are run on a server equipped with an NVIDIA RTX 4090 GPU, an Intel Xeon processor, and 128 GB of RAM. The deep learning framework uses PyTorch, the optimizer uses Adam, and mixed precision is used to accelerate training. The learning rate is set to 0.002, the batch size is set to 16, the weight decay coefficient is 1e-5, 300 rounds of training are performed, and an early stopping strategy is used to prevent overfitting; Normalize the data using standardization to normalize all feature values to the range [0, 1] to improve the convergence of the model; at the same time, divide the data into a training set, a validation set, and a test set in a ratio of 6:2:2, and use the flow data of the first 12 timesteps to predict the next 12 timesteps; Use three common error metrics for quantitative analysis: Mean Absolute Error, Root Mean Squared Error, and Mean Absolute Scaled Error. The definitions of these metrics are as follows: where x t is the true value, is the predicted value, |S| is the size of the sample set, MSE penalizes large deviations through squared errors and is sensitive to outliers; MAPE is expressed as a percentage but is unstable for values close to zero; MAE measures the absolute deviation, is stable and has consistent units, but penalizes large errors less.

Citation Information

Cited By

  • Sequence data modeling method and system

    CN120874904A

  • A sequence data modeling method and system

    CN120874904B

  • Tidal station water level prediction method based on hybrid adaptive graph structure

    CN121071615A