Multi-mode space-time traffic flow modeling method supporting large-scale road network real-time prediction
By designing a spatiotemporal prediction framework for dynamic traffic networks, using multimodal feature fusion and expert mixing mechanisms, the problem of insufficient prediction of existing models in high spatiotemporal heterogeneity scenarios is solved, and efficient and accurate traffic flow prediction is achieved.
Patent Information
- Application Number
- CN202510821855.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
When facing high spatiotemporal heterogeneity scenarios, existing traffic flow prediction models are difficult to dynamically reflect the topological changes of the traffic network, have high computational complexity, and lack prediction accuracy and stability in emergencies, which cannot meet the real-time prediction needs of large-scale road networks.
A spatiotemporal prediction framework for dynamic traffic networks is designed, including data embedding layer, spatiotemporal encoding module and deep modeling module based on expert hybrid mechanisms. It adopts time embedding, spectral domain spatial embedding, block-level sparse temporal attention module, spatial attention-messaging module and dynamic expert modeling module to achieve efficient spatiotemporal feature fusion and prediction.
It improves the response speed and prediction accuracy of the model in high heterogeneity scenarios, enhances cross-region adaptability, meets the needs of emergency response and traffic scheduling optimization, and has efficient computing power and robustness.
Smart Images

Figure CN120337795A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent transportation systems, and particularly relates to a multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks. Background Art
[0002] The core of the intelligent transportation system lies in the accurate prediction of traffic flow dynamics, which directly affects key functions such as traffic signal optimization, congestion warning, and emergency dispatch. In high spatio-temporal heterogeneity scenarios with sudden crowd gatherings (such as a sharp increase in tourist flow during holidays at scenic spots), instantaneous high-density pedestrian agglomerations (such as the period when a sports event ends), and multi-source flow coupling effects (such as the peak period of transfer at transportation hubs), traditional prediction models based on static spatio-temporal assumptions face severe attenuation of prediction efficiency.
[0003] Traditional traffic flow prediction technologies mostly rely on pre-set graph structures and fixed time-series modeling frameworks. For example, traffic flow prediction models based on graph convolutional neural networks have been widely applied (such as Chinese Patent CN117037491A). This method extracts spatial features through a predefined road adjacency matrix and models the evolution process of traffic states by combining a time-series embedding mechanism. The model proposed in this patent integrates graph convolution, time attention, and external factor modeling, which improves the prediction accuracy to a certain extent. However, the construction of its graph structure still relies on static graph embedding methods such as Node2vec and cannot dynamically reflect the spatial dependence adjustment caused by event disturbances or changes in traffic density during the actual operation of vehicle flows.
[0004] The self-attention mechanism for global modeling has become one of the mainstream directions in spatio-temporal sequence prediction research. For example, the multi-head self-attention mechanism proposed in Chinese Patent CN117037483A can be used to capture long-range dependencies. However, when faced with non-uniform sampling, multi-period perturbations, or discontinuous event sequences, there is a problem of discrete assignment of attention weights in this type of mechanism, making it difficult to effectively model the short-term impact brought by local mutations. In addition, due to the quadratic complexity (O(T²)) of its attention mechanism in the time dimension, it is prone to cause computational bottlenecks when applied to long sequences or high-frequency sampled traffic data, which is particularly unfavorable for real-time deployment on edge devices.
[0005] In terms of spatial modeling, existing technologies mostly adopt a graph convolution framework driven by a static adjacency graph (such as Chinese Patent CN117037491A), which fails to fully consider the dynamic topological changes of the traffic network. Especially in the morning and evening rush hours or traffic accident scenarios, the propagation paths of vehicle flows and the correlations between nodes will change rapidly. The static graph model lacks adaptability, which easily leads to distorted characterization of the linkage relationship between upstream and downstream nodes, and further causes the accumulation and amplification of prediction errors. Although some studies attempt to introduce a dynamic graph structure learning mechanism, problems such as increased computational complexity and decreased network interpretability often follow. For example, in Chinese Patent CN119443156A, a fully connected attention structure is used for spatial modeling. When facing a road network with a node scale of tens of thousands, its message passing mechanism with O(N²) complexity can hardly support the performance requirements of online inference.
[0006] In terms of spatio-temporal cascade modeling of traffic states, Chinese Patent CN119107799A proposes a graph convolution recurrent neural network considering data missing situations, emphasizing the impact of data integrity on traffic prediction accuracy. Although this solution enhances the robustness of the model to abnormal inputs by introducing a Gaussian mixture mechanism, it is still insufficient in modeling the propagation delay effect between multiple nodes. In an actual road network, congestion usually has the characteristics of cascade diffusion. Especially in the compact urban core area, its propagation paths show complex laws such as non-linear superposition and multi-scale time delays. Existing local models based on geometric neighborhoods or sliding windows (such as Chinese Patent CN119107799A) have rough detection area division and lack a fine description of the congestion evolution path, making it difficult to accurately capture the spatio-temporal dependencies and dynamic changes during the propagation process, resulting in error amplification. Especially when facing non-stationary disturbances such as a sharp increase in passenger flow during holidays or instantaneous congestion caused by large-scale events, the prediction performance of such sliding window-based models drops significantly and cannot meet the accurate response requirements for emergencies in practical applications.
[0007] Some studies attempt to introduce multi-source information to improve the generalization ability of the model. For example, Chinese Patent CN119541217A proposes to integrate external factors such as weather and holidays through feature splicing. However, most of these methods adopt a static fusion strategy and lack a learnable dynamic weighting mechanism. When the external environmental data noise is large or data sources are missing, it is easy to cause the prediction stability of the model to decline; especially when sensors are sparsely deployed or offline failures occur, the injection of redundant features without discrimination will further interfere with the model training and decision-making process.
[0008] In summary, the current traffic flow prediction technology still has the following significant deficiencies in dealing with complex traffic behaviors and sudden disturbances in the real urban environment: First, the static graph structure and decoupled spatio-temporal modeling method are difficult to characterize the evolution process of the topological structure and the time-delay characteristics of path propagation in the dynamic traffic network; Second, the cross-nested relationship between short-term disturbances and periodic patterns is not sufficiently mined at the time modeling level, resulting in unstable model performance in non-stationary traffic flow scenarios; Third, there are obvious bottlenecks in the overall modeling framework in terms of computational efficiency, feature fusion robustness, and cross-regional generalization ability, making it difficult to meet the real-time prediction and deployment requirements in large-scale road networks. Summary of the Invention
[0009] To solve the technical bottlenecks existing in the spatio-temporal coupling modeling, dynamic topology perception, computational efficiency, and scenario adaptability of the existing traffic flow prediction models, the present invention provides a multi-mode spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks. This method designs a spatio-temporal prediction framework for dynamic traffic networks, which can support multi-mode travel behavior modeling, has dynamic structure perception, efficient reasoning ability, and elastic feature fusion mechanism, significantly improving the response speed, prediction accuracy, and cross-regional adaptability of the model in highly heterogeneous scenarios, and meeting the actual application requirements of intelligent transportation systems in emergency response, traffic scheduling optimization, and long-term traffic flow regulation.
[0010] To achieve the above object, the present invention adopts the following technical solutions: A multi-mode spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks. This method first designs a spatio-temporal prediction framework for dynamic traffic networks, and then trains the designed spatio-temporal prediction framework to obtain the final prediction model; the designed spatio-temporal prediction framework includes a data embedding layer, a spatio-temporal encoding module, a deep modeling and output module based on the expert mixture mechanism. Among them, the data embedding layer is used to convert the original spatio-temporal data into a high-dimensional dense representation; the data embedding layer includes two parallel channels of time embedding and spectral domain spatial embedding and a spatio-temporal data fusion module. The time embedding channel is used to extract time semantic features based on local temporal features and periodic position encoding, and the spectral domain spatial channel obtains spatial frequency domain structure features through graph Laplacian spectral decomposition; the spatio-temporal data fusion module fuses the time semantic features and the spatial frequency domain structure features through element-wise addition operation to obtain the input data of the spatio-temporal encoding module. The spatio-temporal encoding module is used to perform temporal modeling and spatial modeling on the preliminary fusion feature tensor obtained by the data embedding layer, and includes a parallel block-level sparse temporal attention module, a spatial attention-message passing module, and a weighted fusion layer; the block-level sparse temporal attention module adopts a self-attention mechanism based on time blocks and introduces a dynamic masking mechanism, which slices the input data according to nodes, extracts the time series corresponding to each node from the input data, and performs local modeling and scale alignment processing on the temporal features; the spatial attention-message passing module fuses a potential dynamic graph structure, a dynamic graph attention mechanism approximated by a kernel function, and a spatio-temporal masking mechanism driven by adaptive dynamic time warping, which slices the input data according to time steps and models the dynamic spatial dependencies between nodes in the traffic network; the weighted fusion layer is used to align the features output by the block-level sparse temporal attention module and the spatial attention-message passing module, and perform weighted fusion through learnable weight coefficients to output a unified spatio-temporal feature representation after fusion; The deep modeling and output module based on the mixture-of-experts mechanism is used to model the unified spatio-temporal feature tensor and generate a prediction output; it includes a MoE dynamic expert modeling module, a fully connected mapping module, a skip connection layer, and an output layer; the MoE dynamic expert modeling module includes multiple independent expert sub-networks and a gating network, and the gating network is used to dynamically calculate the weight coefficients of the expert sub-networks to achieve weighted fusion of multi-expert outputs; the fully connected mapping module is used to perform a fully connected mapping on the features output by the MoE dynamic expert modeling module to obtain a preliminary prediction result; the skip connection layer is used to transfer the spatio-temporal fusion features to a linear mapping layer to retain low-order feature information; the output layer introduces a learnable fusion weight coefficient to balance the outputs of the skip connection and the fully connected mapping module and generate a final prediction result.
[0011] As a preference of the present invention, the time embedding channel constructs a fine-grained temporal representation of traffic flow by fusing local temporal features and explicit periodic semantics, and the specific data processing process is as follows: Step A. Perform a unified dimensionality projection on the input original tensor and map it to an embedding space with a fixed dimension using a fully connected layer (i.e., ) to obtain a standard feature embedding, and the expression is: ; Step B. Set two learnable embedding dictionaries, namely the week embedding dictionary and the intra-day timestamp embedding dictionary , and respectively construct the week embedding tensor and the intra-day timestamp embedding tensor of the original tensor ; Step C. Embed the standard features, week embeddings, and intra-day timestamp embeddings and concatenate them along the feature dimension. The concatenated time embeddings are input into a linearly learnable projection layer to obtain high-order time semantic features ; Step D. Define the position encoding matrix , and fuse the periodic semantics and position structure through element-wise addition; finally, the output of the time embedding channel .
[0012] Preferably, the data processing flow of the spectral domain space channel is as follows: Step A. Construct a symmetric normalized graph Laplacian matrix according to the loop-free adjacency matrix and the corresponding degree matrix of the traffic road network; Step B. Perform eigenvalue decomposition on the graph Laplacian matrix, remove the eigenvectors corresponding to the zero eigenvalues, and start from the eigenvectors corresponding to the second smallest eigenvalue, select the first eigenvectors to construct a low-dimensional spectral graph embedding; Step C. Use a learnable linear function to map the low-dimensional spectral graph embedding to a high-dimensional space to obtain spatial embedding features, that is, spatial frequency domain structure features .
[0013] Preferably, the data processing flow of the spatio-temporal data fusion module is as follows: Step A. Perform an element-wise addition operation on the output of the time embedding channel and the output of the spectral domain space channel to generate a preliminary spatio-temporal embedding tensor ; Step B. Introduce a learnable position encoding matrix , and add it to the preliminary spatio-temporal embedding tensor element-wise again to obtain the input feature tensor of the finally fused spatio-temporal encoding module .
[0014] Preferably, the data processing flow of the block-level sparse time attention module is as follows: Step A. Slice at the node level. For each node, extract its corresponding time series from the input data; then adopt a sliding window strategy to divide the time series corresponding to the node into multiple time blocks with overlapping regions along the time axis; Step B. Flatten the two-dimensional features in each time block into one-dimensional vectors to obtain a flattened time block sequence; Step C. The flattened time block sequence is input into a learnable linear transformation module to obtain an embedding tensor ; Step D. Introduce a multi-head self-attention mechanism. For each attention head, based on the obtained embedding tensor , generate query, key, and value vectors respectively through a learnable linear mapping; Step E. For each attention head, calculate the attention scores of the -th time block with respect to the -th time block based on the query vector of the -th time block and the key vector of the -th time block, obtaining an attention score matrix; Step F. At each query time block , only keep the indices of the key vector positions corresponding to the top of its attention scores. Perform softmax normalization on the retained index positions, and set the attention scores at the remaining positions to zero to achieve sparse connection, obtaining a sparse attention matrix ; Step G. Introduce a dynamic channel masking mechanism to construct a masking matrix , where each element of this matrix follows Bernoulli distribution sampling, used to indicate whether to retain the output of the -th channel at the -th time block; Step H. Element-wise multiply the sparse attention matrix with the dynamic masking matrix to obtain an attention output with an inhibitory effect ; Step I. Concatenate the outputs of all attention heads along the channel dimension and fuse them through a learnable linear layer to obtain an intermediate output of the block-level sparse temporal attention module ; Step J. The intermediate output is normalized through layer normalization operation to obtain a final normalized output; Step K. Introduce a one-dimensional transposed convolution operation to restore the final normalized output to the original temporal length, then all nodes output a tensor after passing through the block-level sparse temporal attention module.
[0015] Preferably, the data processing flow of the spatial attention-message passing module is as follows: Step A. Slice the input data in units of time steps to construct a T-frame temporal graph, where each frame of the graph contains all nodes; at the -th layer of the graph neural network, the features of the -th and -th nodes are and respectively; introduce three learnable linear mapping matrices, which are used to construct query vectors, key vectors, and value vectors respectively; Step B. Replace the dot-product form of attention similarity calculation with an approximation method of a positive definite kernel function. The positive definite kernel function , is a low-dimensional kernel feature mapping function, which is constructed based on random Fourier features; respectively represent the query vector of the -th node and the key vector of the -th node. represents the index of the central node currently being updated, represents the index of the candidate source node; Step C. Introduce a spatio-temporal masking mechanism driven by adaptive dynamic time warping. The masking coefficient corresponding to nodes in the spatio-temporal masking matrix is: ; Among them, represents that there is a direct geographical connection between nodes and in the traffic network, otherwise it is ; represents the connection strength under the influence of external constraint conditions between node pairs . ; is the temperature factor, which is used to adjust the influence degree of DTW similarity on the masking weight. represents the original temporal feature of node 's DTW distance, is defined the same as the input time step length; Step D. Use the spatio-temporal masking matrix as the prior weight to update the features of node in the -th layer of the graph neural network. The feature of node in the current layer is constructed from the query, key, and value vectors of the features in the previous layer and aggregates neighbor information in a kernel function weighted manner. The calculation of its attention weight and the feature update process are: ; Among them, the index represents all nodes used for normalization, represents the corresponding key vector, represents the value vector of node . The construction logic of the element in the masking matrix is the same as . It is the masking coefficient corresponding to nodes and in the spatio-temporal masking matrix. The kernel mapping operation in the formula can share the calculation results to avoid redundancy, thereby optimizing its computational complexity.
[0016] Preferably, in the weighted fusion layer, a layer normalization mechanism and a residual connection structure are introduced, that is, the fused feature tensor is subjected to layer normalization processing, and the result is added to the initial input features by residual connection as the final output.
[0017] As a further preference of the present invention, when the time embedding channel processes data, for each time step , the corresponding even dimensions use the sine function to calculate the position encoding matrix , and the corresponding odd dimensions use the cosine function to calculate the position encoding matrix .
[0018] Compared with the prior art, the advantages and beneficial effects of the present invention are: (1) Aiming at the limitation of the existing spatio-temporal embedding method in which the time pattern and the spatial structure are modeled separately, the present invention proposes a fusion spatio-temporal spectrum embedding mechanism to achieve a unified and high-dimensional structured representation of traffic flow data; in the time channel, the mechanism combines local temporal features with periodic semantic encoding to capture the intra-day periodicity and short-term local fluctuation features in the traffic flow sequence; in the spatial channel, the graph Laplacian spectrum decomposition and clustering strategy are introduced simultaneously to extract the low-frequency principal components of the road network structure, realizing the compact representation (compressed representation) of the spatial structure and feature redundancy removal; in addition, a learnable position encoding is introduced as a structural prior, and through explicit combination with the fused features, the effective perception of the implicit spatio-temporal structure information in the input sequence is realized, and finally a unified time-frequency joint representation space is constructed, comprehensively enhancing the model's coverage ability and modeling expressiveness for spatio-temporal coupling features in complex traffic scenarios, and having stronger spatio-temporal coupling modeling ability and generalization performance compared with traditional single-modal embedding methods.
[0019] (2) Aiming at the problems of high computational complexity and insufficient capture of local patterns in the attention mechanism for long-term time series prediction, the present invention innovatively designs a time series modeling mechanism based on block compression; the mechanism divides the initial time series into multiple sliding time segments, and performs local attention modeling within each time block to improve the ability to capture short-term fluctuations and periodic signals; to further reduce the computational burden, a block-level sparse time attention structure is constructed based on the Top-k strategy at the time block unit, which can effectively capture fine-grained local dynamics and periodic fluctuations while reducing the attention computational complexity to a sub-linear level, improving the computational throughput while maintaining the time series correlation accuracy. In addition, an adaptive masking mechanism is introduced to dynamically regulate the attention weights through Bernoulli sampling in the feature dimension, realizing feature-level perturbation injection and generalization enhancement; this mechanism improves the robustness and cross-scenario prediction ability of the model in non-stationary time series change scenarios (such as holidays, emergencies, frequent jumps, etc.).
[0020] (3) Aiming at the problems of the solidification of the spatial dependence structure in the traffic network, the difficulty in adapting to dynamic evolution, and the high computational overhead restricting real-time prediction, the present invention proposes an efficient spatial attention mechanism based on kernel function approximation and dynamic regularization; based on kernel function mapping, this mechanism effectively compresses the similarity calculation and information propagation process between nodes through learning the latent graph structure and the kernel approximation method based on random Fourier features, reducing the computational complexity of spatial attention from to . While maintaining the ability to model the latent topological relationship, it significantly improves the scalability and real-time processing ability on traffic graphs with a scale of tens of thousands of nodes (supporting real-time inference for a scale of tens of thousands of nodes), meeting the deployment requirements in edge computing scenarios.
[0021] (4) To solve the problem of spatial modeling deviation caused by asynchronous traffic evolution between nodes and irregular emergencies, the present invention first introduces a spatio-temporal masking mechanism driven by adaptive dynamic time warping. This mechanism uses the DTW algorithm to measure the temporal consistency of historical traffic flow sequences between nodes, and combines external constraints such as geographical connection relationships and road control to dynamically adjust the attention weights. This masking mechanism works in collaboration with the kernelized spatial attention to guide the model to focus on node pairs with similar path evolution trends and satisfying the actual road connectivity, thereby enhancing the expression ability of the propagation delay effect, and enhancing the stability, prediction accuracy of the model in peak hours and heterogeneous traffic environments, as well as the adaptability and robustness to sudden abnormal scenarios.
[0022] (5) Aiming at the problems of dimension mismatch and semantic gap in multi-modal feature fusion, the present invention proposes a dynamic weight spatio-temporal joint optimization framework; through a differentiable gating mechanism, it realizes the dynamic weighted coupling of spatio-temporal attention outputs, combines an expert mixture network to replace the traditional feed-forward network to achieve the weighted fusion of multiple expert outputs, and at the same time combines skip connections and an adaptive fusion strategy to retain the complementarity of low-order spatio-temporal features and deep semantic features, enhancing the model's joint modeling ability for traffic flow mutation patterns and long-range evolution laws, improving the model's expression ability and generalization performance in heterogeneous traffic behavior scenarios, and enhancing the adaptability to complex urban traffic management requirements.
[0023] (6) The model provided by the present invention shows excellent prediction accuracy and structural robustness on typical datasets such as PeMS, NYCTaxi, T-Drive, and CHIBike, demonstrating strong robustness and cross-scenario promotion ability in dense urban areas, main road corridors, and multi-site travel networks, meeting the actual application requirements in emergency response, traffic scheduling optimization, and long-term traffic regulation. Brief Description of the Drawings
[0024] Figure 1 Schematic diagram of the spatio-temporal prediction framework for dynamic traffic networks constructed by the present invention; Figure 2 Data processing flowchart of the spatio-temporal prediction framework for the dynamic traffic network constructed for the present invention; Figure 3 Processing flowchart of the block-wise temporal self-attention module based on dynamic sparse masks constructed for the present invention; Figure 4 Processing flowchart of the kernel function-based spatial attention-message passing module constructed for the present invention; Figure 5 Loss curve of the training and validation processes; Figure 6 Visualization of the data prediction results for the PeMSD8 dataset. Detailed implementation manners
[0025] To enable those skilled in the art to better understand the technical solutions and advantages of the present invention, the present application will be described in detail below with reference to the accompanying drawings, but is not intended to limit the protection scope of the present invention.
[0026] As Figures 1 to 4 shown, the present invention provides a multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks. The method first designs a spatio-temporal prediction framework for the dynamic traffic network, and then trains the designed spatio-temporal prediction framework to obtain a final prediction model; the designed spatio-temporal prediction framework consists of a data embedding layer, a spatio-temporal encoding module, and a deep modeling and output module based on the expert mixture mechanism; Among them, the data embedding layer is used to perform spatio-temporal feature mapping on the original traffic flow input data to construct a high-dimensional dense representation; the data embedding layer includes two parallel channels of time embedding and spectral domain spatial embedding and a spatio-temporal data fusion module. The model first constructs a unified high-dimensional dense representation through time embedding and frequency domain spatial embedding (graph spectrum embedding), and then inputs it into the spatio-temporal encoding module for joint modeling; The time modeling part (block-level sparse time attention module) of the spatio-temporal encoding module divides the time blocks by sliding windows for the data after embedding, uses the sparse multi-head attention mechanism based on the Top-k strategy to extract local and global dependencies, and combines the query adaptive mask to improve the robustness of the model to perturbation information; the spatial modeling part (spatial attention-message passing module) adopts the dynamic graph attention mechanism approximated by the kernel function, improves the modeling efficiency of the dynamic adjacency relationship by constructing a latent graph structure, and further introduces a mask mechanism based on dynamic time warping (DTW) considering external constraints to guide the spatial attention to focus on the temporal consistency between nodes; finally, through the spatio-temporal data fusion module for fusion, the fused spatio-temporal features are input into the deep modeling and output module based on the expert mixture mechanism; The deep modeling and output module based on the mixture of experts mechanism is used to model the unified spatio-temporal feature tensor and generate a prediction output; it includes a MoE dynamic expert modeling module, a fully connected mapping module, a skip connection layer, and an output layer; the MoE dynamic expert modeling module includes multiple independent expert sub-networks and a gating network, and the gating network is used to dynamically calculate the weight coefficients of the expert sub-networks to achieve the weighted fusion of multiple expert outputs; the fully connected mapping module is used to perform a fully connected mapping on the features output by the MoE dynamic expert modeling module to obtain a preliminary prediction result; the skip connection layer is used to transfer the spatio-temporal fusion features to the linear mapping layer to retain the low-order feature information; the output layer introduces a learnable fusion weight coefficient to balance the outputs of the skip connection and the fully connected mapping module and generate the final prediction result.
[0027] In the traffic flow prediction task, define to represent the traffic flow of nodes in the road network at each time step , where represents the dimension of the traffic flow. After combining the traffic flows on all time slices, it is denoted as the original tensor : ; The goal of traffic flow prediction is to predict the traffic flow of the traffic system at future times based on the given historical observation data. Formally, let the observed traffic flow tensor in the traffic system be , and this embodiment aims to learn a mapping function such that it can, according to the input traffic flow observations of time steps, combined with the traffic network graph structure time steps of traffic flow tensor, that is: ; To achieve the above prediction goal and effectively capture the complex spatio-temporal dynamic features in the traffic system, this embodiment proposes a data embedding layer, aiming to convert the original spatio-temporal data into a high-dimensional dense representation, so as to provide sufficient feature support for subsequent modules.
[0028] In this embodiment, the data embedding layer includes two parallel channels of time embedding and spectral domain space embedding and a spatio-temporal data fusion module. The two parallel channels respectively extract time semantic features based on local temporal features and periodic position encoding, and spatial frequency domain structure features obtained through graph Laplacian spectral decomposition; then, key prior knowledge is incorporated into the representation learning process through spatio-temporal data fusion to obtain preliminary fusion features , where It is the unified feature dimension input for the encoder (spatiotemporal encoding module) and serves as the unified representation space for subsequent processing by the spatiotemporal encoding module.
[0029] Specifically, the expressions of the two parallel channels included in the data embedding layer are:
[0030] Among them, represents the time embedding channel, which models the periodic patterns and short-term perturbations in the traffic flow sequence; represents the spectral domain space embedding channel, which combines the spectral decomposition of the graph structure and uses the eigenvectors of the road network Laplacian matrix to extract the patterns implicit in the traffic data in the frequency domain, thereby further revealing the internal connection between the spectral characteristics and the road network structure.
[0031] In this embodiment, the time embedding channel constructs a fine-grained temporal representation of traffic flow by fusing local temporal features and explicit periodic semantics. To retain as much of the intrinsic information in the original data as possible, first, the input original tensor is projected with a unified dimension, and a fully connected layer (i.e., ) is used to map it into an embedding space with a fixed dimension to obtain the standard feature embedding , which is expressed as: , where represents the hidden dimension set by the network structure (set as a global fixed value in this embodiment).
[0032] Secondly, to introduce explicit periodic time semantic information, this embodiment sets two learnable time embedding dictionaries, namely the week embedding dictionary and the intra-day timestamp embedding dictionary . Among them, represents the seven natural days in a week, represents the number of time periods in a day with a granularity of 5 minutes. At each time step , let the week index and the intra-day timestamp index , extract the corresponding embedding representations from the above embedding dictionaries, and construct the week embedding tensor and the timestamp embedding tensor respectively. Subsequently, to make the dimensions of the three embedding tensors consistent, and are broadcast and replicated on the node dimension N, and concatenated with the standard feature embedding tensor on the feature dimension to obtain the fused temporal representation tensor: , where the symbol Represents a concatenation operation in the feature dimension (i.e., the last dimension). Finally, the concatenated tensor is mapped to a unified high-order representation space to achieve temporal semantic fusion and feature dimension compression, obtaining .
[0033] To further strengthen the sequential information in the temporal features, this embodiment adopts the sine-cosine position encoding method, which can effectively capture the sequential structure in the sequence and has translational invariance and scalability. For each time step and each dimension , the calculation rule of the position encoding matrix is defined as follows: ; For each time step , the corresponding even dimensions (the -th dimension) are encoded using the sine function, and the corresponding odd dimensions (the -th dimension) are encoded using the cosine function. Finally, the position encoding matrix is extended to the same dimension as the temporal semantic tensor using the broadcasting mechanism to obtain the broadcasted position encoding tensor . Finally, feature integration is performed in an element-wise addition manner to obtain the final output tensor of the time embedding channel: .
[0034] In this embodiment, a spectral domain space embedding channel is set up. To model the spatial structural features of the traffic network, the spectral decomposition method of the graph Laplacian matrix is adopted.
[0035] Given the traffic road network graph structure 's loop-free adjacency matrix and the corresponding degree matrix , where N represents the number of nodes in the graph. In the adjacency matrix , the element represents whether there is an edge between node and node . If there is an edge, then , otherwise , where , represents the node set. Since the graph considered in the present invention is a loop-free graph, thus is satisfied. The corresponding graph degree matrix is a diagonal matrix, and its -th diagonal element represents the degree of node , which is defined as: Based on the above definitions, the present invention further constructs a symmetric normalized graph Laplacian matrix: , where represents the identity matrix. Since is a symmetric positive semi - definite matrix, its eigenvalues reflect the structural characteristics of the graph. To further extract the frequency - domain feature information of the graph spectrum, the present invention performs eigenvalue decomposition (i.e., spectral decomposition) on and obtains: , where is the diagonal eigenvalue matrix, is the eigenvector matrix, and the superscript represents the transpose.
[0036] To extract the low - frequency global information of the graph, the present invention excludes the eigenvectors corresponding to the zero eigenvalues , starts from the eigenvectors corresponding to the second - smallest eigenvalue, and selects the first non - trivial eigenvectors (in this embodiment ) to construct a low - dimensional spectral graph embedding: ; Finally, to obtain a unified feature format and facilitate subsequent spatio - temporal feature fusion, this embodiment projects the spectral graph embedding into a preset feature space , and performs broadcast replication along the time dimension at this stage to obtain a spatial embedding tensor .
[0037] In this embodiment, after extracting the time - embedding features and spatial - embedding features, the spatio - temporal data fusion module integrates the two types of features in an element - wise addition manner to generate a unified - format spatio - temporal feature tensor for subsequent modules.
[0038] Specifically, an element - wise addition operation is performed on the time embedding and the spatial embedding to generate a unified - format fusion tensor (preliminary spatio - temporal embedding tensor) ; In addition, to enhance the model's ability to model implicit spatio - temporal structural information in the input sequence, a learnable position - encoding matrix is introduced in the data embedding layer as a structural prior. This position - encoding matrix is replicated and extended along the time dimension through the broadcast mechanism to make its dimension consistent with , and through element - wise addition, the final fused spatio - temporal input feature tensor is obtained.
[0039] As Figure 3As shown, in this embodiment, a block-level sparse temporal attention module is constructed, and the block-level sparse temporal attention module is used for local modeling and scale alignment processing of temporal features.
[0040] Traffic flow data is highly non-linear and dynamic, and it is difficult for traditional point-by-point modeling methods to effectively capture its time-dependent features. Therefore, the present invention designs a patch-based temporal self-attention mechanism (PTSA) for extracting long-term and short-term dependencies in time series, and integrates an adaptive masking mechanism to enhance the robustness of the model and prevent over-reliance on a single time point.
[0041] Specifically, in this embodiment, the block-level sparse temporal attention module adopts a patch-based temporal self-attention mechanism and introduces a dynamic masking mechanism to form a block-based temporal self-attention mechanism based on dynamic sparse masks; the data processing flow is as follows: First, the embedded time series is divided into overlapping time blocks, and the context dependencies between time blocks are modeled through a Top-k sparse multi-head self-attention mechanism, and a query adaptive masking strategy based on the Bernoulli distribution is introduced to achieve dynamic selection of feature dimensions and enhanced perturbation robustness; then, layer normalization is combined to enhance the non-linear modeling ability, and one-dimensional transposed convolution is used to upsample the downsampled time features for subsequent spatio-temporal fusion to provide a unified input.
[0042] Specifically, the data processing flow of the block-level sparse temporal attention module is as follows: Step A. Overlapping time block generation and high-order temporal modeling: The present invention adopts a sliding window strategy to divide the multi-dimensional feature tensor after spatio-temporal encoding into multiple time blocks with overlapping regions along the time axis to enhance the model's ability to model local dynamic changes.
[0043] Let the input feature tensor be , where is the batch size, represents the number of nodes in the traffic road network, is the time step length, is the unified feature dimension input by the encoder (spatio-temporal encoding module).
[0044] First, slice by node level. For each node , extract its corresponding time series from the input feature tensor : ; for each node, divide its corresponding time series along the time axis into time blocks with a length of and a step size of . The total number of time blocks is: ; Among them, the symbol represents the floor operation, and the time blocks of each node form a four-dimensional tensor: ; To capture the global semantic structure within each time block, the two-dimensional features (i.e., of size ) in each time block are flattened into a one-dimensional vector to obtain the flattened time block sequence: , and the flattened time block sequence is a three-dimensional tensor; Subsequently, the above-mentioned flattened time block sequence is input into a learnable linear transformation module (linear) to map it to a unified high-order temporal embedding space. Let the embedding dimension after mapping be , where is the number of heads of the multi-head time attention set in step B (in this embodiment, is set to 4), is the subspace dimension of each head. The mapping result is the time block embedding tensor corresponding to node : .
[0045] Step B. Multi-head time attention and dynamic masking mechanism: To model the high-order temporal dependencies between overlapping time blocks, the present invention introduces a multi-head self-attention mechanism. For each attention head , the present invention is based on the node time block embedding tensor constructed in step A, and respectively generates query (Query), key (Key), and value (Value) vectors , , through the following linear mapping: ; Among them, are respectively the learnable projection matrices for the th attention head to construct query, key, and value vectors, ensuring that each head independently learns the temporal pattern in the subspace.
[0046] To further improve the computational efficiency and structural sparsity of the multi-head time attention, the present invention introduces a block-level Top-k sparse attention mechanism and a dynamic channel masking mechanism. Specifically, for each attention head , let the query vector of the th time block be , and the key vector of the th time block be , then the scaled dot product attention score between the two is defined as: . Stack all query vectors and key vectors into a matrix, where represents the The attention score of the th time block for the th time block, then the attention score matrix obtained is:
[0047] To improve efficiency, the present invention adopts, at each query position , a Top- selection strategy, that is, only the position indexes of the key vectors in the top of its attention scores are retained to construct a set . Only softmax normalization operations are performed at these index positions, and the attention scores at the remaining positions are set to zero to achieve sparse connections. Therefore, the sparse attention weight and the sparse attention matrix are calculated as follows: ; ; To enhance the robustness of the model and the ability to select feature channels, the present invention introduces a dynamic channel masking mechanism after the output of the multi-head temporal attention. This mechanism realizes the suppression of feature-level noise and the enhancement of effective information through a dynamic retention strategy for the channel dimension (i.e., the last dimension ) in the output of each attention head. Specifically, a dynamic masking matrix is constructed for each attention head, where the retention probability of the th channel is , which is dynamically generated using a linear increasing strategy, that is: ; where, and are hyperparameters, representing the minimum and maximum retention probabilities respectively, satisfying , to achieve an increasing design of the channel retention probability. Each element of the matrix follows a Bernoulli distribution: ; where the Bernoulli distribution function represents that the probability of the sampling value being is , that is, it means that this channel is retained in the time block , and is masked when it is 0. The sampled masking matrix is broadcast to be consistent with the dimension of the sparse attention output tensor . Finally, by multiplying the sparse attention matrix and the dynamic masking matrix element by element, the attention output with an inhibitory effect is obtained: . Among them, denotes the Hadamard product, i.e., element-wise multiplication. While maintaining the expressive power of attention, this mechanism introduces channel-level dynamic selection and suppression, effectively enhancing the expression robustness and feature selection ability of the model.
[0048] Subsequently, the outputs of all attention heads are concatenated along the channel dimension and fused through a learnable linear layer, where the learnable linear mapping matrix is , obtaining the intermediate output of the block-level sparse temporal attention module : ; To alleviate feature distribution drift and improve training stability, after obtaining the intermediate output, a layer normalization (LayerNorm) operation is introduced. This normalization process calculates the mean and standard deviation of each sample in units of the feature dimension, and maps them through standardization and learnable scaling and bias parameters to finally obtain the normalized representation .
[0049] Step C. Temporal dimension restoration and spatio-temporal alignment: Since the attention input adopts a sliding window structure, the output dimension is . To keep the temporal dimension consistent with the spatial attention, the present invention introduces a one-dimensional transposed convolution (i.e., TransposedConv1D( )) operation to restore the original temporal length ; ; In summary, the above temporal modeling process is independently executed for all nodes in the graph, and finally the complete temporal feature tensor can be obtained as: . This tensor is consistent with the output of the spatial attention module in the temporal axis and channel dimension, providing a structurally unified and semantically aligned input for the subsequent spatio-temporal fusion module.
[0050] As Figure 4 shown, in this embodiment, a kernel function-based spatial attention-message passing module is constructed to model the dynamic spatial dependencies between nodes in the traffic network.
[0051] In traditional graph neural networks (GNNs), graph structures are usually predefined and static; however, in dynamic traffic networks, the relationships between nodes have obvious time-varying properties. In order to more effectively model spatial dependencies, the present invention designs a spatial attention-message passing module that integrates latent graph structure, multi-head graph attention, and spatiotemporal masking mechanism. The data processing flow is as follows: constructing a latent dynamic graph structure using node features, and introducing a kernel function approximation mechanism to achieve efficient spatial message propagation without explicitly calculating the attention matrix, greatly reducing the computational complexity; then introducing an adaptive dynamic time regularization-driven spatiotemporal masking mechanism, combined with external structural constraints such as dynamic alignment similarity of historical traffic sequences between nodes, geographic adjacency, and road access restrictions, constructing an attention prior that integrates behavioral consistency and structural accessibility, and guiding spatial attention to focus on key node pairs with consistent temporal evolution trends and the possibility of physical connection, thereby significantly improving the model's expressiveness and predictive robustness in complex traffic scenarios.
[0052] Specifically, the data processing flow of the spatial attention-message passing module is: Step A. Constructing the latent dynamic graph structure: The present invention firstly uses the time step as the unit to input the feature tensor Slice and construct Frame timing diagram, each frame contains Node, The initial node features of the frame graph (i.e., the input of the 0th layer of the graph neural network) are defined as: ; This paper adopts the idea of layer-by-layer message propagation. In order to elaborate on the calculation process of the spatial attention-message passing module, the batch dimension is not considered below, and only the feature interaction relationship between any node pairs in a single frame graph is focused. Layer ( , this embodiment sets ), No. The feature representation of all nodes in the frame graph is composed of tensors , all node pairs The potential graph connection weights of are expressed as: .in, Consistent with the node set in the spatial feature embedding step, Respectively represent nodes and nodes In the The feature vector of the layer, Representation Node To Node The latent graph connection weights of Indicates the target node index to be updated. Denotes the candidate source node index. To construct the spatial attention relationship between nodes, the present invention introduces three learnable linear mapping matrices , which are respectively used to construct the query vector , the key vector and the value vector , and their calculation methods are as follows: ; Furthermore, a global attention network is defined above, which estimates the potential interactions between instance nodes and enables corresponding dense connection information transmission. The calculation of its attention weights and the feature update process are as follows: ; Among them, the index traverses all nodes in the graph, represents the spatial attention key vector corresponding to node . This module does not adopt the traditional multi-head attention structure, but retains the core idea of weighted aggregation of adjacent nodes in the graph attention mechanism in its design; at the same time, to maintain output consistency with the block-level sparse temporal attention module, the output dimension of this module is set to .
[0053] Step B. Attention acceleration by kernel function approximation; To effectively reduce the computational complexity existing in the traditional attention mechanism in the graph structure, the present invention introduces an efficient message passing mechanism based on kernel functions. This mechanism replaces the dot product form of attention similarity calculation through the approximation method of kernel functions, thereby realizing scalable approximate attention propagation.
[0054] Specifically, the present invention adopts a positive definite kernel function in the following form: . Among them, is a kernel function that satisfies symmetry and positive definiteness, represents a low-dimensional kernel feature mapping function, which is constructed based on Random Fourier Features (RFF) in this embodiment and is used to approximate the above kernel inner product operation in the Euclidean space to simulate common translation-invariant kernel functions.
[0055] Taking as an example of construction, the corresponding form of the kernel feature mapping function is defined as follows: ; Among them, , is a frequency vector independently sampled from the Gaussian distribution corresponding to the kernel function and is used to construct Fourier features; is the phase shift sampled from a uniform distribution, which is used to make the mapping results smoothly distributed within the cosine period. represents the mapping dimension (the mapping dimension in the above equation is adjustable, and is set in this embodiment ). Through this mapping, the attention mechanism of dot product followed by exponentiation can be transformed into a linear inner product operation in the vector space, thus significantly reducing the computational overhead.
[0056] Step C. Introduce a spatio-temporal masking mechanism driven by adaptive dynamic time warping: To enhance the perception ability of the spatial attention-message passing module for the differences in temporal evolution features in the traffic network, the present invention designs a spatio-temporal masking mechanism based on adaptive dynamic time warping (DTW) to improve the model's ability to model node temporal heterogeneity. This mechanism comprehensively considers the dynamic alignment degree of historical traffic sequences between nodes to measure their temporal consistency; at the same time, it integrates external structural constraints such as geographical connection relationships and road controls to jointly construct the spatio-temporal mask.
[0057] Specifically, the original tensor is processed to extract the original temporal features of any two nodes and , and their DTW distance is defined as: , where represents the set of legal alignment paths that satisfy the boundary conditions, monotonicity, and continuity constraints, is the feature of node at the th time step, and is the feature of node at the th time step.
[0058] To consider the structural prior information such as the spatial reachability and external restrictions between nodes, the present invention constructs a physical adjacency matrix and an external constraint matrix ; where indicates that there is a direct geographical connection between nodes and in the traffic network, otherwise it is . represents the connection strength under the influence of external restriction conditions between node pairs , where , such as the attenuation weight caused by human factors such as construction traffic restrictions. Combining the above physical connections, external constraints, and DTW similarity, the final mask coefficient corresponding to node and node in the spatio-temporal mask matrix is defined.is as follows: ; In the above formula, is the temperature factor, which is used to adjust the influence degree of the DTW similarity on the mask weight. At the same time, can map the time series distance to the similarity weight, and through normalizing the time length control the scale, and improve the consistency and stability of the similarity measurement under different lengths.
[0059] Furthermore, the present invention takes this mask matrix as the prior weight in the kernel function attention, and combines the decomposability of the random Fourier feature kernel mapping in step B to effectively simplify the calculation process; thus, the redefinition of the potential graph connection weight is in the following matrix form: ; In this embodiment, due to the decomposability of the above kernel mapping, the explicit construction of the fully connected attention matrix is avoided , and at the same time, the representation of the node at the current layer is obtained by aggregating the neighbor information in the weighted manner of the kernel function after constructing the query, key, and value vectors from the features of the nodes in the previous layer , and then normalizing, so that the computational complexity is reduced from significantly to , greatly improving the modeling efficiency and scalability on large-scale graph data.
[0060] Finally, all the node feature representations output by the spatial attention-message passing module at the L-th layer (final layer) are .
[0061] In this embodiment, a weighted fusion layer is constructed to align the features of the outputs of the block-level sparse time attention module and the spatial attention-message passing module, and achieve weighted fusion through learnable weight coefficients, and output the unified spatio-temporal feature representation after fusion.
[0062] To improve the stability and expression ability of the fusion process, a layer normalization mechanism and a residual connection structure are further introduced in this fusion layer, that is, the feature result after weighted fusion is normalized, and residual addition is performed with the original input feature to strengthen the feature propagation and optimize the training effect.
[0063] The time attention mechanism and the spatial attention mechanism share the same input tensor . To ensure the consistency of the time and spatial attention outputs, the output dimensions of the time attention and the spatial attention are set to . As can be seen from the above, the time attention output is , and the spatial attention output is .
[0064] Specifically, in this embodiment, the data processing flow of the weighted fusion layer is as follows: Step A. Feature dimension alignment: To achieve the effective fusion of the time modeling output and the space modeling output, the present invention first needs to map both to the hidden dimension of a unified network structure . Specifically, the present invention designs two independent fully connected layers (i.e., one-dimensional linear transformations) for the time branch and the space branch respectively to adjust the channel dimension, , representing the fully connected mapping functions acting on the time branch and the space branch respectively, and defining the spatio-temporal feature representations after dimension alignment as : ; Step B. Weighted fusion: After completing the dimension alignment of the time and space features, the present invention performs adaptive fusion on the two by introducing learnable fusion weights to fully integrate the time and space perception information.
[0065] Specifically, the present invention sets a learnable fusion coefficient , and the finally fused spatio-temporal feature tensor is expressed as: ; Step C. Normalization and residual connection: To further improve the expression stability of the fusion features and the convergence efficiency of training, after obtaining the fusion feature tensor , the present invention introduces a layer normalization mechanism and a residual connection structure. First, the fusion result is subjected to layer normalization processing to obtain .
[0066] Then, to maintain information continuity and enhance the gradient propagation effect, for the spatio-temporal embedding feature tensor a learnable linear mapping function is used to obtain the residual feature tensor , and the normalized fusion feature and the residual feature are element-wise added to achieve residual connection. The finally output fusion feature is: ; This structure not only achieves the deep fusion of time-space perception information, but also improves the expression ability and training stability of the model through the layer normalization and residual connection mechanisms, providing a more robust spatio-temporal feature representation for downstream modeling tasks.
[0067] In this embodiment, a deep modeling and output module based on the mixture of experts mechanism (MoE-FFN-OUT) is constructed to perform further non-linear modeling and output generation on the fused high-dimensional spatio-temporal features.
[0068] To achieve more expressive and generalizable deep modeling and output prediction for the fused high-dimensional spatio-temporal features, this embodiment proposes an integrated modeling module that combines the mixture of experts mechanism (MoE), fully connected mapping structure, skip connection path, and output layer design. This module not only replaces the single feed-forward network layer in the traditional Transformer but also enhances the model's ability to model multi-scale spatio-temporal dependence features through dynamic expert selection and skip connection mechanisms, and maps the final deep representation to the output space. The module mainly includes four components: the MoE dynamic expert modeling module, the fully connected mapping module, the skip connection layer, and the output layer, and the overall structure effectively improves the expressiveness, stability, and prediction accuracy of the model.
[0069] Specifically, in this embodiment, the data processing flow of the deep modeling and output module based on the mixture of experts mechanism is as follows: Step A. Introduce the MoE dynamic expert modeling module to achieve dynamic expert modeling: In this embodiment, the improved feed-forward network layer uses the MoE dynamic expert modeling module, which includes multiple independent expert sub-networks and a gating network. The gating network is used to dynamically calculate the weight coefficients of the expert sub-networks based on the input features, so as to adapt to the multi-modal spatio-temporal dependence relationships existing in the traffic flow and enhance the expression flexibility and generalization ability of the model; There are (in this embodiment = 4) independent expert sub-networks, and each expert sub-network corresponds to a feed-forward sub-module . Define the input data as , which comes from the spatio-temporal fusion feature tensor output by the spatio-temporal encoding module: . For each position in the tensor, all expert sub-networks process the input in parallel to generate expert outputs: , where the serial number .
[0070] In this embodiment, to implement the soft routing mechanism for multi-expert outputs, a gating network is introduced to adaptively generate the expert weight vector for each position according to the input features. The calculation method is as follows: ; Among them, are the learnable parameters of the gating network, and the softmax function ensures that the weights are normalized among all experts and satisfy: . Finally, the MoE dynamic expert modeling module performs a weighted combination of the outputs of all expert sub-networks to form the final feed-forward output : ; This mechanism not only improves the non-linear expression ability of the model, but also has stronger expert selectivity, which helps to capture diverse spatio-temporal dependence structures.
[0071] Step B. Fully connected mapping and skip path establishment: In this embodiment, to convert the fused spatio-temporal features into the final prediction result, first, the high-order semantic features output by the MoE dynamic expert modeling module in step A are subjected to a fully connected mapping to obtain a preliminary prediction result; the mapping process is as follows: ; where represents the fully connected mapping function, is the target variable dimension (consistent with the input). At the same time, to retain the straight-through information in the low-order spatio-temporal features, this embodiment designs a skip connection path, that is, a skip connection layer, to directly send the spatio-temporal fusion features into an independent linear mapping layer; where represents the linear projection function in the skip connection path, and we get: ; Step C. Output layer design: To balance the deep semantic features and the initial fusion representation, this embodiment introduces a learnable fusion weight coefficient to balance the outputs of the two paths; the final prediction result is expressed as follows: ; This design effectively enhances the adaptability and robustness of the model in the multi-duration prediction scenario by combining high-order semantic representations and low-order spatio-temporal features, and enhances the model's ability to model traffic flow mutation patterns and long-term evolution trends.
[0072] In this embodiment, to verify the effectiveness and generalization ability of the proposed block-based temporal self-attention prediction method, a unified experimental environment is constructed and a scientific and reasonable performance evaluation mechanism is set.
[0073] Experimental Environment Configuration: Considering that traffic prediction tasks involve large-scale graph data and long-sequence modeling, the experiments were conducted on a high-performance computing platform. In terms of hardware, an NVIDIA GeForce RTX 4090 graphics card (24GB GDDR6X video memory), a 12th-generation Intel Core i9-12900 processor (24 threads), and 128GB of DDR4 memory were used to ensure the computing and memory requirements of graph neural networks during large-scale training. In terms of software, the operating system was Ubuntu 18.04 LTS, the deep learning framework was PyTorch 2.0.0, and the CUDA version was 11.8 to balance operator optimization and system compatibility. The Adam optimizer was used for model training, the initial learning rate was set to 0.001, and the Cosine Annealing Scheduler was used for dynamic learning rate adjustment. The batch size was set to 256, the maximum number of training epochs was 200, and the early stopping strategy (patience = 10) was enabled to monitor the performance of the validation set MAE and prevent overfitting. All experiments were set with a fixed random seed (seed = 42), and 5-fold random initialization cross-validation was adopted to enhance the stability of the results. In terms of data processing, the traffic flow data was standardized using Z-score; in the construction of the graph structure, an adjacency matrix generation strategy with a threshold distance of 2km was used. The dataset division followed the general specification: the PEMS series datasets adopted a 7:2:1 training / validation / test ratio, and the Grid-based format datasets adopted a 7:1:2 division scheme.
[0074] Selection of Experimental Datasets: To comprehensively verify the generalization ability of the model, the present invention selects six representative spatio-temporal datasets, covering multiple scenarios such as highway traffic flow and urban traffic demand. Among them, the structured road network data comes from the PeMSD4 and PeMSD8 datasets released by the California Transportation Data Laboratory (Caltrans Performance Measurement System, PeMS), covering the San Francisco Bay Area and the Los Angeles and San Bernardino regions respectively. These data are collected by approximately 39,000 detectors deployed on the major metropolitan highways in California. Each data point contains three features: traffic flow, average speed, and average occupancy. The present invention focuses on the traffic flow feature, and the spatial adjacency matrix of each dataset is constructed based on the geographical distance of the actual road network. Urban traffic demand data includes the following typical representatives: the New York Taxi dataset NYCTaxi (publicly released by the New York City Taxi and Limousine Commission), the Chicago Bike Share dataset CHIBike (including bike ride records at multiple stations), and the Beijing T-Drive taxi trajectory dataset released by Microsoft Research Asia. The above data constitute an urban traffic prediction experimental environment with multi-time scale and multi-trip mode characteristics.
[0075] Design of performance evaluation mechanism: To comprehensively evaluate the performance of the proposed model in traffic flow prediction tasks, a multi-dimensional performance evaluation system is constructed, and the following statistical indicators are used: 1) Mean Absolute Error (MAE): Measures the average absolute value of the difference between the predicted value and the true value, used to evaluate the overall level of error; 2) Root Mean Square Error (RMSE): Based on MAE, it emphasizes the impact of larger prediction deviations more. After taking the weighted average of the squared errors and then taking the square root, it is suitable for application scenarios sensitive to outliers; 3) Mean Absolute Percentage Error (MAPE): Quantifies the proportion of the prediction error in the true value in percentage form, facilitating an intuitive comparison of the relative prediction errors of the model on data of different magnitudes. Selection of baseline models: To comprehensively evaluate the performance of the proposed model in traffic flow prediction tasks, the present invention compares it with current mainstream deep learning models on multiple standard datasets, covering various paradigms such as graph neural networks, sequence modeling structures, and attention mechanism-based ones. The present invention selects eight representative models, namely STGCN, DCRNN, Graph WaveNet, STSGCN, STFGNN, STGODE, Transformer, and GMAN, as comparison methods, covering different ideas such as static graph modeling, dynamic graph modeling, and spatio-temporal joint modeling.
[0076] Analysis of experimental results: The experimental results are shown in Table 1 and Table 2 below. On the two highway datasets of PeMSD4 and PeMSD8, the model of the present invention exhibits optimal performance in all evaluation metrics. For example, on PeMSD4, the model of the present invention achieved an MAE of 18.89, an RMSE of 29.71, and a MAPE of 12.84%, significantly outperforming Graph WaveNet, with improvements of approximately 24.1%, 25.1%, and 4.5% respectively; on PeMSD8, the present invention also led comprehensively with an MAE of 13.89, an RMSE of 23.32, and a MAPE of 9.08%, significantly outperforming strong baseline models such as GMAN, demonstrating good modeling capabilities and cross-regional generalization performance. In the complex urban traffic scenario, the model of the present invention achieved optimal performance in both the in-flow and out-flow directions of the NYCTaxi dataset. In the in-flow direction, the model reached an MAE of 13.83, an RMSE of 22.72, and a MAPE of 13.60%; in the out-flow direction, it further achieved an MAE of 12.11, an RMSE of 19.12, and a MAPE of 13.43%, comprehensively outperforming models such as GMAN, Transformer, and STFGNN, demonstrating strong adaptability and stability to irregular urban travel patterns.
[0077] Table 1: Performance comparison of different models on PeMSD4 and PeMSD8 datasets
[0078] Table 2: Performance comparison of different models on NYCTaxi dataset
[0079] Figure 5 Shows the training and validation loss curves of the model in four typical experiments, where Figure 5Among them, (a) and (b) respectively correspond to the training loss and validation loss of the PeMSD4 dataset, (c) and (d) are the training loss and validation loss of the PeMSD8 dataset, (e) and (f) are the training loss and validation loss of the CHIBike data, and (g) and (h) show the training loss and validation loss of the NYCTaxi data. By analyzing the training and validation loss curves of the 4 different experiments, it can be observed that the training loss curves of all experiments continuously decline as the number of iterations increases, indicating that the model is continuously improving its fitting ability to the training data. At the same time, most of the validation losses also show a trend of first decreasing and then stabilizing, reflecting that the model has good generalization performance and convergence in multiple different types of traffic scenarios.
[0080] Figure 6 The visualization curve comparing the traffic flow prediction results of the model of the present invention with the existing SOTA model and the true value on the PeMSD8 dataset is shown. The horizontal axis represents the time step, and the vertical axis represents the traffic flow value, covering the full-day complete time series range. By visualizing and comparing the differences between the model prediction and the actual flow, from the overall trend, the model of the present invention can closely fit the real traffic flow data curve at the vast majority of time steps, significantly superior to the existing SOTA model, showing higher prediction accuracy and robustness. To further demonstrate the advantages of the model, three representative local regions are specifically selected in the figure for magnification analysis, and the results are as follows: The early peak pre-fluctuation section (time step is about 20 - 60): The model of the present invention has a stronger fitting ability to sudden fluctuations in traffic flow and can accurately reflect the trend of small-scale rapid rise and fall; in contrast, the prediction results of the SOTA model show obvious lag and errors. The low valley transition stage (time step is about 120 - 160): The model of the present invention still maintains good prediction accuracy in the low traffic volume interval and accurately captures the recovery change after the bottom; the SOTA model has a large deviation and insufficient capture of the rebound trend in this stage; The late peak oscillation stage (time step is about 180 - 220): The model of the present invention has a more delicate and stable prediction effect on the rapid fluctuations of traffic flow during the late peak period, reducing the error accumulation; the SOTA model has increased prediction jitter and direction errors at multiple oscillations, with poor robustness.
[0081] To further verify the applicability and effectiveness of the traffic flow prediction model of the present invention in a real and complex urban road network environment, the traffic flow prediction results generated by the model of the present invention are now combined with the Beijing taxi trajectory dataset publicly available from Microsoft Beijing in 2008 to draw an urban-level heat map. In this embodiment, the heat map is obtained by spatially aggregating the traffic flow intensity predicted by the model within a given time window, performing rasterization processing in combination with geographical location coordinates, and finally mapping it onto the geographical base map of the central urban area of Beijing. The intensity of the color in the heat map represents the density of the traffic flow. The heat map generated based on the prediction results of the model of the present invention can accurately depict the phenomenon of concentrated traffic flow during peak hours on the main traffic arteries and hubs in Beijing (such as Chang'an Avenue, the Zhongguancun area, the traffic corridor between the Second Ring Road and the Fourth Ring Road in the east, etc.), which is highly consistent with the historical statistics of the actual taxi trajectories, indicating that the model of the present invention has strong spatial generalization ability and regional congestion pattern recognition ability.
[0082] To more precisely understand the contributions and importance of each core component in this model, the present invention designed a series of ablation experiments. These experiments aim to reveal the specific impact of each component on the overall performance of the model by gradually removing key components from the model. The specific ablation experiment settings and results are as follows: Remove the block-based input mechanism and compare the traditional multi-head temporal attention structure with the model design of the present invention. In this ablation experiment, the present invention removed the local temporal sparse modeling mechanism in the model (i.e., the block temporal self-attention mechanism based on a dynamic sparse mask) and replaced it with a traditional multi-head temporal attention mechanism. The multi-head temporal attention mechanism performs global attention calculation based on the entire sequence. Although it has certain advantages in modeling long-term dependencies, due to its neglect of local structures, it results in a large number of parameters and high computational overhead, especially more obvious in the case of long sequences or multi-head parallel settings. The experimental results show that in short-term prediction tasks, the MAPE of the traditional mechanism is significantly higher than that of the PTSA module with Patch introduced, indicating that it is difficult to effectively capture short-term fluctuations and emergencies in traffic flow. In addition, due to the high dependence of this structure on a complete and fixed-length input sequence, it has poor adaptability to variable-length time periods or data with missing values, showing the problem of insufficient generalization ability.
[0083] Replace the dynamic spatial attention mechanism with traditional graph convolution operations. Specifically, the spatial attention-message passing module based on the latent graph structure and kernel function in the original model is replaced by traditional static graph convolution operations. After the replacement, the node feature update only depends on the predefined static adjacency matrix. When using traditional GCN (Graph Convolutional Neural Network) to replace the dynamic spatial attention mechanism, the MAE increases by about 18%, and the RMSE and MAPE also increase significantly, indicating that static graph convolution is difficult to effectively capture the time-varying correlations between nodes. At the same time, since the dynamic spatial attention mechanism enhances the modeling ability of spatial dependence changes by introducing the latent graph structure and kernel function, while the fixed graph structure of traditional GCN limits the model's response ability to the dynamics of the real traffic network.
[0084] Remove the spatio-temporal masking mechanism based on DTW. After removing the DTW mask, the model performance slightly decreases, and the MAE and RMSE increase by about 7.3% and 6.0% respectively, and the MAPE increases by about 0.26%. Compared with the latent graph structure and kernelized attention mechanism, the DTW mask contributes slightly less to the performance. However, as a pluggable spatio-temporal alignment module, it improves the model's expressive ability for complex traffic scenarios. The DTW mask is not a core structure, but it supplements the temporal structure information in capturing spatio-temporal alignment features. This design has stronger adaptability to situations such as abnormal flows, holiday impacts, and local traffic sudden changes.
[0085] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the technical field to which the present invention pertains, based on the idea of the present invention, several simple deductions, deformations or replacements can also be made. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.
Claims
1. A multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks, characterized in that, The method first designs a spatio-temporal prediction framework, and then trains it to obtain a final prediction model; the spatio-temporal prediction framework includes a data embedding layer, a spatio-temporal encoding module, and a deep modeling and output module based on the mixture-of-experts mechanism; Among them, the data embedding layer is used to convert the original spatio-temporal data into a high-dimensional dense representation, including two parallel channels of time embedding and spectral domain spatial embedding and a spatio-temporal data fusion module. The time embedding channel is used to extract local temporal features and the temporal semantic features of periodic position encoding. The spectral domain spatial channel obtains spatial frequency domain structural features through graph Laplacian spectral decomposition; the spatio-temporal data fusion module is used to perform feature fusion; The spatio-temporal encoding module is used to perform temporal modeling and spatial modeling on the input tensor obtained by the data embedding layer, including a parallel block-level sparse time attention module, a spatial attention-message passing module, and a weighted fusion layer; the block-level sparse time attention module adopts a self-attention mechanism based on time blocks and introduces a dynamic masking mechanism. It slices the input data according to nodes, extracts the time series corresponding to each node from the input data, and performs local modeling and scale alignment processing on the temporal features; the spatial attention-message passing module fuses a potential dynamic graph structure, a dynamic graph attention mechanism approximated by a kernel function, and a spatio-temporal masking mechanism driven by adaptive dynamic time warping. It slices the input data according to time steps and models the dynamic spatial dependencies between nodes in the traffic network; the weighted fusion layer is used to perform feature alignment to achieve weighted fusion; The deep modeling and output module based on the mixture-of-experts mechanism is used to model the unified spatio-temporal feature tensor and generate a prediction output, including a MoE dynamic expert modeling module, a fully connected mapping module, a skip connection layer, and an output layer; the MoE dynamic expert modeling module includes multiple independent expert sub-networks and a gating network; the fully connected mapping module is used to perform a fully connected mapping; the skip connection layer is used to transfer the spatio-temporal fusion features to the linear mapping layer; the output layer is used to balance the outputs of the skip connection and the fully connected mapping module to generate a final prediction result.
2. The multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks according to claim 1, characterized in that The time embedding channel constructs a fine-grained temporal representation of traffic flow by fusing local temporal features and explicit periodic semantics. The specific data processing flow is as follows: Step A. Perform unified dimensionality projection on the input original tensor and map it to an embedding space of a fixed dimension using a fully connected layer to obtain a standard feature embedding; Step B. Set two learnable embedding dictionaries, namely, the week embedding dictionary and the intra-day timestamp embedding dictionary , and respectively construct the week embedding tensor and the intra-day timestamp embedding tensor of the original tensor ; Step C. Embed the standard feature, week embedding, and intraday timestamp embedding and concatenate them along the feature dimension. The concatenated time embedding is input into a learnable linear projection layer to obtain high-order time semantic features ; Step D. Define the position encoding matrix , fuse the periodic semantics and the position structure through element-wise addition; finally, the output of the time embedding channel .
3. The multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks according to claim 1, characterized in that, The data processing flow of the spectral domain spatial channel is as follows: Step A. Construct a symmetric normalized graph Laplacian matrix according to the loop-free adjacency matrix and the corresponding degree matrix of the traffic road network; Step B. Perform eigenvalue decomposition on the graph Laplacian matrix, remove the eigenvectors corresponding to the zero eigenvalues, and starting from the eigenvector corresponding to the second smallest eigenvalue, select the first eigenvectors to construct a low-dimensional spectral graph embedding; Step C. Use a learnable linear mapping function to embed the low-dimensional spectral graph into a high-dimensional space to obtain spatial embedding features, i.e., spatial frequency domain structure features .
4. The multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks according to claim 1, characterized in that The data processing flow of the spatio-temporal data fusion module is as follows: Step A. Perform an element-wise addition operation on the output of the temporal embedding channel and the output of the spectral domain spatial channel to generate a preliminary spatio-temporal embedding tensor ; Step B. Introduce a learnable positional encoding matrix and add it to the preliminary spatio-temporal embedding tensor element-wise to obtain the input feature tensor of the finally fused spatio-temporal encoding module .
5. The multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks according to claim 1, characterized in that The data processing flow of the block-level sparse time attention module is as follows: Step A. Slice according to the node level. For each node, extract its corresponding time series from the input data; then adopt a sliding window strategy to divide the time series corresponding to the node into multiple time blocks with overlapping regions along the time axis; Step B. Flatten the two-dimensional features in each time block into a one-dimensional vector to obtain a flattened time block sequence; Step C. The flattened time block sequence is input into a learnable linear transformation module to obtain an embedding tensor ; Step D. Introduce the multi-head self-attention mechanism. For each attention head, based on the obtained embedding tensor , generate query, key, and value vectors through a learnable linear mapping; Step E. For each attention head, calculate the attention scores of the -th time block for the -th time block based on the query vector of the -th time block and the key vector of the -th time block to obtain an attention score matrix; Step F. At each query time block , only retain the key vector position indices with the top attention scores. Perform a softmax normalization operation on the retained index positions, set the attention scores at the remaining positions to zero, implement sparse connections, and obtain a sparse attention matrix ; Step G. Introduce a dynamic channel masking mechanism to construct a masking matrix , where each element of the matrix follows a Bernoulli distribution sampling and is used to represent whether to retain the output of the -th channel at the -th time block; Step H. Element-wise multiply the sparse attention matrix with the dynamic mask matrix to obtain an attention output with an inhibitory effect ; Step I. Concatenate the outputs of all attention heads along the channel dimension and fuse them through a learnable linear layer to obtain the intermediate output of the block-level sparse temporal attention module ; Step J. The intermediate output is finally normalized after layer normalization operation; Step K. Introduce a one-dimensional transposed convolution operation to restore the original time length of the final normalized output, so that the output tensor after all nodes pass through the block-level sparse temporal attention module .
6. The multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks according to claim 1, characterized in that The data processing flow of the spatial attention-message passing module is as follows: Step A. Slice the input data in time steps to construct a temporal graph of T frames, where each frame contains all nodes; at the -th layer of the graph neural network, the features of the -th and -th nodes are and respectively; introduce three learnable linear mapping matrices, which are used to construct query vectors, key vectors, and value vectors respectively; Step B. Replace the dot-product form of attention similarity calculation with an approximation method of positive definite kernel function. The positive definite kernel function , is a low-dimensional kernel feature mapping function constructed based on random Fourier features; respectively represent the query vector of the -th node and the key vector of the -th node, represents the index of the currently updated central node, represents the index of the candidate source node; Step C. Introduce a spatio-temporal mask mechanism driven by adaptive dynamic time warping, and nodes and nodes In the spatio-temporal mask matrix, the corresponding mask coefficient is as follows: ; Among them, represents a node and has a direct geographical connection in the transportation network, otherwise it is ; represents the connection strength under the influence of external constraint conditions between node pairs ; ; is a temperature factor used to adjust the influence degree of DTW similarity on the mask weight, represents the original time series feature of node ; is the DTW distance of and is defined the same as the input time step length; Step D. Using the spatio-temporal mask matrix as the prior weight, update the features of the node in the layer of the graph neural network. The features of the node in the current layer are query, key, and value vectors constructed from the features of the previous layer, and the neighbor information is aggregated in a kernel function weighted manner. The calculation of its attention weight and the feature update process are as follows: ; Among them, the index represents all nodes for normalization, represents the corresponding key vector, represents the node value vector, is the node and the node is the mask coefficient corresponding to the spatio-temporal mask matrix.
7. The multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks according to claim 1, characterized in that, The layer normalization mechanism and the residual connection structure are introduced in the weighted fusion layer, that is, the fused feature tensor is subjected to layer normalization processing, and the result is added to the initial input features as a residual to obtain the final output.
8. The multi-modal spatio-temporal traffic flow modeling method for supporting real-time prediction of large-scale road networks according to claim 2, characterized in that When the time embedding channel processes data, for each time step , the sine function is used to calculate the position encoding matrix for the corresponding even dimensions , and the cosine function is used to calculate the position encoding matrix for the corresponding odd dimensions .
Citation Information
Patent Citations
Traffic flow prediction method based on multi-head attention mechanism
CN117037483A
Dynamic graph convolution traffic flow prediction method based on spatial-temporal characteristics
CN117037491A
Graph convolution recurrent neural network road traffic flow prediction method considering data missing
CN119107799A
Traffic flow prediction model for realizing co-mining of traffic point semantics and network topology
CN119443156A
Short-term traffic flow prediction method and system based on multi-feature map
CN119541217A
Cited By
Drainage port dynamic management system based on quantum parameter optimization and air-ground collaborative awareness
CN120877142A
A quantum parameter optimization and air-ground collaborative perception dynamic management system for a drainage outlet
CN120877142B
Prediction model and device for large B-cell lymphoma gene rearrangement
CN121121241A
Traffic flow prediction method based on deep learning
CN121281279A
An end-to-end spatio-temporal prediction method based on improved three-dimensional rotary position encoding
CN121366384B