A Traffic Flow Prediction Method Based on Mamba

By adopting a static graph convolution network and a dual-branch Mamba architecture in traffic flow prediction, combined with a space-time dual-token mechanism, the prediction accuracy and efficiency problems of the existing methods in high dynamic road network and multivariate coupling scenarios are solved, and more efficient traffic flow prediction and urban-level road network monitoring are achieved.

CN119942813BActive Publication Date: 2025-06-17TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510421465.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-17
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing traffic flow prediction methods are difficult to effectively capture the high-dimensional, nonlinear and spatial-temporal correlation of real traffic flow, resulting in limited prediction accuracy, especially in high dynamic road network environment, multivariable coupling scenarios and long-sequence computing efficiency.

Method used

A Mamba-based traffic flow prediction method is adopted to extract the spatial correlation of the road network through a static graph convolution network. The dual-branch Mamba architecture captures the time dependence and variable correlation of traffic flow, and decouples the influence of multi-source factors through the space-time dual token mechanism to achieve independent modeling and interaction of features.

Benefits of technology

It significantly improves the modeling robustness and interpretability in multivariate coupled scenarios, improves the accuracy and efficiency of traffic flow prediction, and is suitable for all-weather real-time monitoring and rapid response of urban-level road networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942813B_ABST
    Figure CN119942813B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of intelligent transportation, and specifically relates to a traffic flow prediction method based on Mamba, including the following steps: data set collection and preprocessing; static graph construction and spatial feature extraction; spatio-temporal dual-token decoupled embedding; dual-branch Mamba modeling; spatio-temporal-variable feature fusion; performing multi-step prediction and outputting prediction results, and finally, outputting the predicted values of the multi-step sequence through projection mapping. The static graph convolutional network of the present invention extracts spatial correlation from the inherent topology of the road network, while the dual-branch Mamba architecture captures the temporal dependence and variable correlation of traffic flow. The spatio-temporal dual-token mechanism of the present invention explicitly distinguishes the independent influence and synergistic effect between variables by decoupling the action boundaries of multi-source factors such as traffic flow, environmental variables and event interference, significantly improving the modeling robustness and interpretability in the multi-variable coupling scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent transportation, and particularly relates to a traffic flow prediction method based on Mamba. Background Art

[0002] With the increasing complexity and dynamics of urban traffic systems, accurate traffic flow prediction is crucial for intelligent traffic management. The traffic flow data at different stations show significant spatial heterogeneity due to geographical location and functional attributes, and traffic flow exhibits significant time dependence, such as periodic congestion during morning and evening rush hours, changes in travel patterns during holidays, and instantaneous traffic flow fluctuations caused by large-scale events or emergencies. At the same time, its changes are also affected by the complex influence of multi-dimensional variables, including weather conditions (such as rain, snow, haze), traffic accidents, road construction, the operating status of public transportation, and the relevance of surrounding area functions (such as commercial areas, schools). The non-linear coupling of these spatio-temporal dynamics and multi-source variables makes it difficult to fully extract the patterns hidden in traffic flow data through traditional models. High-precision traffic flow prediction can not only optimize signal timing, alleviate road congestion, but also provide decision-making support for urban emergency response and public transportation scheduling, thereby improving urban operation efficiency and residents' travel experience.

[0003] Early traffic flow prediction mainly relied on traditional methods such as historical average algorithms and autoregressive integrated moving average models (ARIMA). These methods usually only considered the historical data of a single traffic collection point itself and made predictions based on linear rules. However, real traffic flow has characteristics such as high-dimensionality, non-linearity, and spatio-temporal correlation, and traditional methods are difficult to capture the complex features of dynamic changes, resulting in limited prediction accuracy. With the development of intelligent transportation systems, machine learning methods such as long short-term memory networks (LSTM) and diffusion convolutional recurrent neural networks (DCRNN) extract time features through recurrent neural networks and capture spatial correlation by combining graph convolutional networks. Deep learning methods such as spatio-temporal graph convolutional networks (ST-GCN), through a hybrid structure of temporal convolutional networks (TCN) and graph convolutional networks (GCN), simultaneously model traffic flow to extract time features through recurrent neural networks, significantly improving the prediction effect. However, it relies on a large amount of computing resources and is difficult to meet actual needs.

[0004] With the increasing demand of intelligent transportation systems for dynamic spatio-temporal modeling, there are significant bottlenecks in the existing technologies in high-dynamic road network environments, multi-variable coupling scenarios, and long-sequence computing efficiency, which are specifically manifested in the following three core problems: insufficient capture of static graph convolution and dynamic spatio-temporal dependence, inefficient multi-variable coupling modeling and interaction, and imbalance between long-sequence computing efficiency and long-term dependence capture. Summary of the Invention

[0005] In view of the technical problems existing in the above-mentioned traditional traffic flow prediction, the present invention provides a traffic flow prediction method based on Mamba.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0007] A traffic flow prediction method based on Mamba, comprising the following steps:

[0008] S1. Dataset collection and preprocessing: Obtain traffic flow data, as well as exogenous variables such as weather, holiday information, and road construction information, and then perform data preprocessing, including missing value filling, data cleaning, division and standardization processing of the training set, validation set, and test set;

[0009] S2. Static graph construction and spatial feature extraction: Construct an adjacency matrix based on the physical connectivity of the road network, use a symmetric normalized Laplacian matrix to enhance the stability of graph convolution, and aggregate multi-order neighborhood information through stacking multiple layers of GCNs to extract the spatial correlation features of road segments;

[0010] S3. Spatiotemporal dual-token decoupled embedding: Decouple the spatial features into time tokens T-Token and variable tokens V-Token, and fuse them with positional encoding and node embedding through fully connected layers respectively to achieve independent modeling of temporal continuity and cross-variable correlation;

[0011] S4. Dual-branch Mamba modeling: Temporal branch: Use bidirectional Mamba to capture the long-term periodic dependence of traffic flow, and fuse the temporal states through forward and backward scans; Variable branch: Based on unidirectional Mamba, model the non-linear coupling relationship between speed and flow multi-variables to generate cross-variable interaction features;

[0012] S5. Spatiotemporal-variable feature fusion: Fuse the dual-branch features through a cross-attention mechanism, use the temporal features as Query, the variable features as Key and Value, and combine residual connections to retain the original spatial information and enhance multi-dimensional feature interaction;

[0013] S6. Perform multi-step prediction and output the prediction results. Finally, output the predicted values of the multi-step sequence through projection mapping.

[0014] The method for dataset collection and preprocessing in S1 is as follows: The traffic flow data is from 10,000 loop inductive sensors obtained from external data, continuously recording traffic flow, average speed, and time occupancy at a sampling interval of 30 seconds; In terms of exogenous variables, access the meteorological data with a 5-minute granularity and real-time traffic accident annotation information obtained from external data to increase the influence of exogenous variables. The granular meteorological data includes wind speed, precipitation, and visibility;

[0015] For missing values, a spatio-temporal two-dimensional interpolation strategy is adopted: for short-term missing values in the time dimension, linear interpolation is performed, and in the space dimension, weighted filling is based on the adjacent three sensors in the same direction:

[0016]

[0017] where: d ij is the sensor spacing, t is the time point to be filled, σ is the control time decay rate, μ is the timestamp of the data, σ and μ are dynamically calculated according to historical data in the same period, and the weight w j reflects the proximity of the data time of sensor j to the current time t. The smaller the time difference, the higher the weight; represents the interpolation estimate value of sensor i at time t; represents the observed value of sensor j at time t; N(i) represents the three sensors in the same direction as sensor i and closest in spatial position;

[0018] The data is normalized and denormalized. The original data is converted to a normalized value according to the mean μi and standard deviation σi. The normalized data will have a zero mean and unit variance; after the prediction is completed, the same mean and standard deviation are used for the inverse transformation to convert the prediction result back to the scale of the original data for evaluation; finally, the training set, validation set, and test set are divided according to the ratio of 7:1:2.

[0019] The method for static graph construction and spatial feature extraction in S2 is as follows:

[0020] S2.1. Graph construction and adjacency matrix definition: Each section in the high-speed road network is regarded as a node in the graph. If two nodes are physically directly connected, the corresponding element A in the adjacency matrix A ij = 1, otherwise A ij = 0;

[0021] S2.2. Laplacian matrix normalization: The symmetric normalized Laplacian matrix is used to enhance stability, where D is the degree matrix, and the diagonal matrix D of the degree matrix D ii = ∑ j A ij , I is the identity matrix, and self-loops are introduced to ensure the retention of node self-features;

[0022] S2.3. GCN propagation formula: The feature update formula for each GCN layer is where H (l) is the node feature matrix of the l-th layer, and the input layer N is the number of nodes, F is the total number of variables, W (l)is the trainable weight matrix of the l-th layer, and σ(·) is the non-linear activation function ReLU; by stacking K layers of GCNs, multi-order neighborhood information is gradually aggregated, and X spatial = GCN(X) = H (K) , H (K) is the node feature matrix of the K-th layer.

[0023] The method of spatio-temporal dual-token decoupled embedding in S3 is as follows:

[0024] S3.1. Temporal token embedding module: In the temporal embedding module, data at the same time step is embedded into one token, enabling each token to represent the complete information of a time point. The embedding step is X T = FC T (X spatial ) + PositionalEncoding(t), where X T represents the features in the time dimension, FC T represents the fully connected layer designed for temporal features, X spatial represents the original spatial features, and PositionalEncoding(t) represents the temporal positional encoding, which preserves the continuity of the time dimension and encodes the temporal order information. The output where N is the number of nodes, d is the feature dimension, and T is the number of time steps;

[0025] S3.2. Variable token embedding module: The variable embedding module generates variable tokens by transposing the temporal tokens, and embeds the entire time series of each variable independently into one token to focus on the global features of the variables in the sequence. The embedding step is X V = FC V (X spatial ) + NodeEmbedding(i), where X V represents the features in the node dimension, FC V represents the fully connected layer designed for node-dimensional features, X spatial represents the original spatial features, and NodeEmbedding(i) represents the embedding vector of node i, aggregating node features and encoding the variable correlations across variables. The output where N is the number of nodes, d is the feature dimension, and F is the total number of variables. Since the order of the sequence is implicitly stored in each variable token, positional embedding is no longer required during variable embedding.

[0026] The method of dual-branch Mamba modeling in S4 is as follows:

[0027] S4.1. Bidirectional Mamba extracts temporal dependencies: Extract the temporal features from the temporal embedding vectors obtained in S3.1 and apply bidirectional Mamba along the time dimension, including forward scanning and backward scanning Output fusion Y T = Concat(h forward , h backward ), where A T , B T respectively represent the discretized temporal Mamba parameters; to improve computational efficiency, continuous parameters are discretized by the zero-order hold rule, h(t) represents the state vector of the system, h′(t) represents the time derivative of the state vector, A represents the system matrix, B represents the input matrix, x n (t) represents the zero-order hold form of the discrete input signal, y(t) represents the output signal of the system, C represents the output matrix; and historical information is captured through forward propagation, future information is captured through backward propagation, and context dependencies between the past and the future are modeled simultaneously to enhance the temporal expression ability;

[0028] S4.2. Unidirectional Mamba extracts variable correlations: Apply unidirectional Mamba along the variable dimension, from key variables to auxiliary variables, Output Y V = h F , and the final hidden state represents the dependencies between variables, where: h f represents the hidden state of the f-th variable, h f-1 represents the hidden state of the (f - 1)-th variable, represents the input feature of the f-th variable, F represents the total number of variables, A V , B V are respectively the state space parameters of the variable Mamba.

[0029] The method of spatio-temporal - variable feature fusion in S5 is as follows:

[0030] S5.1. Cross-attention mechanism: Use temporal features as Query, and variable features as Key and Value, where: Y V represents the variable feature matrix, Y T represents the temporal feature matrix, W Q represents projecting the temporal feature Y T into the Query matrix, W K represents projecting the temporal feature Y T into the Key matrix, W V represents projecting the temporal feature Y T into the Value matrix, d represents the feature dimension, Softmax(·) represents the normalized attention weights, Y fuse represents the fused spatio-temporal joint features;

[0031] S5.2. Residual connection: Preserve the original spatio-temporal features, Y final= Y fuse + Conv(X spatial ), where: Y final represents the final output, Conv(·) represents the convolution operation, and X spatial represents the original spatial feature.

[0032] The fused feature obtained in S5 outputs a future multi-step sequence through linear projection to achieve short-term prediction of traffic flow.

[0033] The prediction method in S6 uses the mean absolute error MAE, root mean square error RMSE, symmetric mean absolute percentage error SMAPE, and correlation coefficient for evaluation.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] The static graph convolutional network of the present invention extracts spatial correlation from the inherent topology of the road network, while the dual-branch Mamba architecture captures the temporal dependence and variable correlation of traffic flow. The spatio-temporal dual-token mechanism of the present invention explicitly distinguishes the independent effects and synergistic effects among variables by decoupling the action boundaries of multi-source factors such as traffic flow, environmental variables, and event interference, significantly improving the modeling robustness and interpretability in multi-variable coupling scenarios. In addition, the lightweight design based on Mamba replaces the traditional Transformer architecture, and uses the compression characteristics of the state space and the selective feature transfer mechanism to achieve more efficient use of computing resources in the processing of ultra-long sequence data, providing a feasible technical path for all-weather real-time monitoring and rapid response of urban-level road networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.

[0037] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0038] Figure 1 It is a schematic flowchart of the method of the present invention;

[0039] Figure 2 This is the module diagram of the present invention;

[0040] Figure 3 This is the Mamba structure diagram of the present invention;

[0041] Figure 4 This is the cross-attention mechanism diagram of the feature fusion of the present invention. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. These descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0043] Next, with reference to the accompanying drawings and embodiments, the specific implementation manners of the present invention will be further described in detail. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0044] This embodiment provides a traffic flow prediction method based on Mamba, as Figure 1 、 Figure 2 shown, and includes the following steps:

[0045] Step 1, data set collection and preprocessing. Obtain traffic flow data and exogenous variables such as weather, holiday information, and road construction information. Then perform data preprocessing, including missing value filling, data cleaning, and partitioning and standardization processing of the training set, validation set, and test set.

[0046] In step 1, the traffic flow data is 10,000 loop inductance sensors obtained from external data acquisition, continuously recording core indicators such as traffic flow, average speed, and time occupancy at a sampling interval of 30 seconds. In addition, in terms of exogenous variables, 5-minute granularity meteorological data (wind speed, precipitation, visibility) and real-time traffic accident annotation information provided by external data are accessed to increase the influence of exogenous variables.

[0047] For missing values, a spatio-temporal two-dimensional interpolation strategy is adopted: linearly interpolate for short-term missing values (<1 hour) in the time dimension, and perform weighted filling based on the adjacent 3 same-direction sensors in the space dimension (the weight is inversely proportional to the square of the sensor spacing, and a Gaussian time decay factor is introduced).

[0048]

[0049] where: d ijis the sensor spacing, t is the time point to be filled, σ is the control time decay rate, μ is the timestamp of the data, and σ and μ are dynamically calculated based on historical data of the same period. The weight w j reflects the proximity of the data time of sensor j to the current time t. The smaller the time difference, the higher the weight; represents the interpolation estimate value of sensor i at time t; represents the observed value of sensor j at time t; N(i) represents the 3 sensors that are in the same direction as sensor i and have the closest spatial positions.

[0050] For better training and convergence of the model, the data is standardized and de-normalized. The original data is converted into standardized values according to the mean μi and standard deviation σi. The standardized data will have zero mean and unit variance. After the prediction is completed, the same mean and standard deviation are used for inverse transformation to convert the prediction results back to the scale of the original data for evaluation. Finally, the training set, validation set, and test set are divided according to the ratio of 7:1:2.

[0051] Step 2, static graph construction and spatial feature extraction. Based on the physical connectivity of the road network, an adjacency matrix is constructed. The symmetric normalized Laplacian matrix is used to enhance the stability of graph convolution. By stacking multiple layers of GCNs, multi-order neighborhood information is aggregated to extract the spatial correlation features of road segments.

[0052] Step 2-1: Graph construction and adjacency matrix definition: Each road segment in the highway network is regarded as a node in the graph. If two nodes (road segments) are directly physically connected (such as upstream and downstream relationships or intersection connections), the corresponding element A in the adjacency matrix A ij = 1, otherwise A ij = 0.

[0053] Step 2-2: Laplacian matrix normalization: The symmetric normalized Laplacian matrix is used to enhance the stability, where D is the degree matrix (diagonal matrix, D ii = ∑ j A ij ), I is the identity matrix, and self-loops are introduced to ensure the retention of node self-features.

[0054] Step 2-3: GCN propagation formula: The feature update formula for each GCN layer is where H (l) is the node feature matrix of the l-th layer, and the input layer N is the number of nodes, F is the total number of variables, W (l) is the trainable weight matrix of the l-th layer, and σ(·) is the non-linear activation function ReLU. By stacking K layers of GCNs, multi-order neighborhood information is gradually aggregated, and X spatial = GCN(X) = H (K) , H(K) is the node feature matrix of the Kth layer.

[0055] Step 3, Spatiotemporal Dual-Token Decoupled Embedding. Decouple the spatial features into temporal tokens (T-Tokens) and variable tokens (V-Tokens), and fuse them with positional encoding and node embedding through fully connected layers respectively to achieve independent modeling of temporal continuity and cross-variable correlation.

[0056] Step 3-1: Temporal Token Embedding Module: In the temporal embedding module, data at the same time step is embedded into one token, enabling each token to represent the complete information of a time point. The embedding step is X T = FC T (X spatial ) + PositionalEncoding(t), where X T represents the features in the time dimension, FC T represents the fully connected layer designed for temporal features, X spatial represents the original spatial features, and PositionalEncoding(t) represents the temporal positional encoding, which preserves the continuity of the time dimension and encodes the time order information, and the output N is the number of nodes, d is the feature dimension, and T is the number of time steps.

[0057] Step 3-2: Variable Token Embedding Module: The variable embedding module generates variable tokens by transposing the temporal tokens, and embeds the entire time series of each variable independently into one token to focus on the global features of the variables in the sequence. The embedding step is X V = FC V (X spatial ) + NodeEmbedding(i), where X V represents the features in the node dimension, FC V represents the fully connected layer designed for node dimension features, X spatial represents the original spatial features, and NodeEmbedding(i) represents the embedding vector of node i, aggregating node features and encoding the variable correlation across variables, and the output N is the number of nodes, d is the feature dimension, and F is the total number of variables. Since the order of the sequence is implicitly stored in each variable token, positional embedding is no longer required during variable embedding.

[0058] Step 4, Dual-Branch Mamba Modeling. As Figure 3 shown, Temporal Branch: Adopt bidirectional Mamba to capture the long-term periodic dependencies of traffic flow and fuse the temporal states through forward and backward scans. Variable Branch: Based on unidirectional Mamba, model the nonlinear coupling relationships among multiple variables such as speed and flow to generate cross-variable interaction features.

[0059] Step 4-1: Bidirectional Mamba extracts temporal dependencies: Extract the temporal features from the time-embedded vectors obtained in Step 3-1, and apply bidirectional Mamba along the time dimension, including forward scanning and backward scanning Output fusion Y T = Concat(h forward , h backward ), where A T , B T represent the discretized temporal Mamba parameters respectively. To improve computational efficiency, continuous parameters are discretized by the zero-order hold rule,

[0060] h(t) represents the state vector of the system, h ′ (t) represents the time derivative of the state vector, A represents the system matrix, B represents the input matrix, x n (t) represents the zero-order hold form of the discrete input signal, y(t) represents the output signal of the system, and C represents the output matrix. And historical information is captured through forward propagation, future information is captured through backward propagation, and context dependencies between the past and the future are modeled simultaneously to enhance the temporal expression ability.

[0061] Step 4-2: Unidirectional Mamba extracts variable correlations: Method: Apply unidirectional Mamba along the variable dimension (from key variables → auxiliary variables), Output Y V = h F (the final hidden state represents the dependencies between variables), where: h f represents the hidden state of the f-th variable, h f-1 represents the hidden state of the (f - 1)-th variable, represents the input feature of the f-th variable, F represents the total number of variables, A V , B V are the state space parameters of the variable Mamba respectively.

[0062] Step 5, spatio-temporal - variable feature fusion. As Figure 4 shown, the dual-branch features are fused through the cross-attention mechanism, using the temporal features as Query, the variable features as Key and Value, and combining residual connections to retain the original spatial information and enhance multi-dimensional feature interaction.

[0063] Step 5-1: Cross-attention mechanism: Using the temporal features as Query and the variable features as Key / Value, where: Y V represents the variable feature matrix, Y T represents the temporal feature matrix, WQ denotes projecting the temporal feature Y T onto a Query matrix, W K denotes projecting the temporal feature Y T onto a Key matrix, W V denotes projecting the temporal feature Y T onto a Value matrix, d represents the feature dimension, Softmax(·) represents the normalized attention weights, Y fuse represents the fused spatio-temporal joint feature.

[0064] Step 5-2: Residual connection: Preserve the original spatio-temporal feature, Y final = Y fuse + Conv(X spatial ), where: Y final represents the final output, Conv(·) represents the convolution operation, X spatial represents the original spatial feature. The obtained fused feature outputs a future multi-step sequence through a linear projection to achieve short-term prediction of traffic flow.

[0065] Step 6, perform multi-step prediction and output the prediction result. Finally, output the predicted values of the multi-step sequence through a projection mapping.

[0066] The prediction method of the present invention is evaluated using the mean absolute error MAE, root mean square error RMSE, symmetric mean absolute percentage error SMAPE, and correlation coefficient.

[0067] The loss function used is the cross-entropy loss, and the formula is:

[0068]

[0069] This method is implemented using Pytorch and trained using the Adam optimizer. After adjusting the parameters through multiple experiments, the present invention sets the learning rate to 0.0001, the batch size to 32, each experiment runs for 100 epochs, and early stop is allowed when the effect is good.

[0070] The static graph convolutional network of the present invention extracts spatial correlations from the inherent topology of the road network (such as intersection hierarchy and road traffic directions), while the dual-branch Mamba architecture captures the temporal dependence and variable correlations of traffic flow. Further, the spatio-temporal dual-token mechanism explicitly distinguishes the independent impacts and synergistic effects among variables by decoupling the action boundaries of multi-source factors such as traffic flow, environmental variables, and event interferences (for example, the differential inhibition of the traffic capacities of arterial roads and branch roads during heavy rain, or the coupling propagation law of traffic accidents and the surrounding road network density), significantly enhancing the modeling robustness and interpretability in multi-variable coupling scenarios. In addition, the lightweight design based on Mamba replaces the traditional Transformer architecture, and by utilizing the compression characteristics of the state space and the selective feature transfer mechanism, it achieves more efficient utilization of computing resources in the processing of ultra-long sequence data (such as traffic flow collected at high frequencies for consecutive days), providing a feasible technical path for the all-weather real-time monitoring and rapid response of urban-level road networks.

[0071] Only the preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and all such changes should be included within the protection scope of the present invention.

Claims

1. A traffic flow prediction method based on Mamba, characterized in that: The following steps are involved: S1. Dataset collection and preprocessing: Obtain traffic flow data and exogenous variables such as weather, holiday information, and road construction information, and then perform data preprocessing, including missing value filling, data cleaning, and the division and standardization of training sets, validation sets, and test sets; S2, static graph construction and spatial feature extraction: construct an adjacency matrix based on the physical connectivity of the road network, use the symmetric normalized Laplacian matrix to enhance the stability of graph convolution, aggregate multi-order neighborhood information through multi-layer GCN stacking, and extract the spatial correlation features of road sections; S3, spatiotemporal dual-token decoupling embedding: decouple spatial features into time tokens T-Token and variable tokens V-Token, which are fused with position encoding and node embedding through fully connected layers to achieve independent modeling of temporal continuity and cross-variable correlation; The method for decoupling and embedding the spatiotemporal dual tokens in S3 is: S3.

1. Time token embedding: In the time embedding module, the data of the same time step is embedded into a token, so that each token can represent the complete information of a time point. The embedding step is X. T =FC T (X spatial )+PositionalEncoding(t),X T Represents the characteristics of the time dimension, FC T represents the fully connected layer designed for time features, X spatial Represents the original spatial features, PositionalEncoding(t) represents the temporal position encoding, retains the continuity of the time dimension, encodes the time order information, and outputs N is the number of nodes, d is the feature dimension, and T is the number of time steps; S3.2, variable token embedding: The variable embedding module generates variable tokens by transposing the time tokens, and embeds the entire time series of each variable into a token independently to focus on the global features of the variables in the sequence; the embedding step is X V =FC V (X spatial )+NodeEmbedding(i),X V Represents the node dimension feature, FC V represents the fully connected layer designed for node dimension features, X spatial represents the original spatial features, NodeEmbedding(i) represents the embedding vector of node i, aggregates node features, encodes variable correlations across variables, and outputs N is the number of nodes, d is the feature dimension, and F is the total number of variables. Since the order of the sequence is implicitly stored in each variable token, position embedding is no longer required when embedding variables; S4, Dual-branch Mamba modeling: Time series branch: bidirectional Mamba is used to capture the long-term periodic dependence of traffic flow, and the time series state is integrated through forward and backward scanning; variable branch: based on the nonlinear coupling relationship between speed and flow multivariables of unidirectional Mamba modeling, cross-variable interaction features are generated; The method of double-branch Mamba modeling in S4 is: S4.1, bidirectional Mamba extracts temporal dependencies: extract temporal features from the temporal embedding vector obtained in S3.1, and apply bidirectional Mamba along the temporal dimension, including forward scanning and backward scanning Output fusion Y T =Concat(h forward ,h backward ), where A T , B T They represent the discrete time series Mamba parameters respectively; in order to improve the computational efficiency, the continuous parameters are discretized by the zero-order hold rule. h(t) represents the state vector of the system, h′(t) represents the time derivative of the state vector, A represents the system matrix, B represents the input matrix, and x n (t) represents the zero-order hold form of the discrete input signal, y(t) represents the output signal of the system, and C represents the output matrix; historical information is captured through forward propagation, future information is captured through back propagation, and the contextual dependency of the past and the future is modeled to enhance the temporal expression capability; S4.2, One-way Mamba to extract variable correlation: Apply one-way Mamba along the variable dimension, from key variables to auxiliary variables, Output Y V =h F , the final hidden state represents the dependency between variables, where: h f represents the hidden state of the fth variable, h f-1 represents the hidden state of the f-1th variable, represents the input feature of the fth variable, F represents the total number of variables, A V , B V They are the state space parameters of the variable Mamba; S5, spatiotemporal-variable feature fusion: The dual-branch features are fused through the cross-attention mechanism, with time series features as query and variable features as key and value, combined with residual connection to retain the original spatial information and enhance multi-dimensional feature interaction; S6. Perform multi-step prediction and output the prediction results. Finally, output the predicted values ​​of the multi-step sequence through projection mapping.

2. A traffic flow prediction method based on Mamba according to claim 1, characterized in that: The method of data set collection and preprocessing in S1 is as follows: traffic flow data is obtained from 10,000 ring-shaped inductive sensors acquired from external data, which continuously record traffic flow, average speed and time occupancy at a sampling interval of 30 seconds; In terms of exogenous variables, 5-minute granular meteorological data and real-time traffic accident annotation information obtained from external data are connected to increase the impact of exogenous variables. The granular meteorological data includes wind speed, precipitation and visibility; For missing values, a dual-dimensional interpolation strategy of time and space is adopted: linear interpolation is performed on short-term missing values ​​in the time dimension, and the spatial dimension is filled based on the weighted addition of three adjacent sensors in the same direction: Where: d ij is the sensor spacing, t is the time point to be filled, σ is the control time decay speed, μ is the timestamp of the data, σ and μ are dynamically calculated based on the historical data of the same period, and the weight w j Reflects the closeness between the data time of sensor j and the current time t. The smaller the time difference, the higher the weight; represents the interpolated estimated value of sensor i at time t; represents the observation value of sensor j at time t; N(i) represents the three sensors with the same direction and closest spatial position as sensor i; The data is standardized and de-standardized. The original data is converted into standardized values ​​according to the mean μi and standard deviation σi. The standardized data will have zero mean and unit variance. After the prediction is completed, the same mean and standard deviation are used for inverse transformation to convert the prediction results back to the scale of the original data for evaluation. Finally, the training set, validation set, and test set are divided in a ratio of 7:1:

2.

3. A traffic flow prediction method based on Mamba according to claim 1, characterized in that: The method of static image construction and spatial feature extraction in S2 is: S2.

1. Graph construction and adjacency matrix definition: Consider each road segment in the highway network as a node in the graph. If two nodes are directly connected physically, the corresponding element A in the adjacency matrix A ij =1, otherwise A ij =0; S2.2, Laplace matrix normalization: using symmetric normalized Laplace matrix Enhanced stability, where D is the degree matrix, the diagonal matrix D of the degree matrix D ii =∑ j A ij , I is the unit matrix, and the self-loop is introduced to ensure that the node’s own characteristics are preserved; S2.3, GCN propagation formula: The feature update formula of each GCN layer is Among them, H (l) is the node feature matrix of the lth layer, the input layer N is the number of nodes, F is the total number of variables, W is (l) is the trainable weight matrix of the lth layer, σ(·) is the nonlinear activation function ReLU; by stacking K layers of GCN, multi-order neighborhood information is gradually aggregated, X spatial =GCN(X)=H (K) , H (K) is the node feature matrix of the Kth layer.

4. A Mamba-based traffic flow prediction method according to claim 1, characterized in that: The method of spatiotemporal-variable feature fusion in S5 is: S5.

1. Cross-attention mechanism: Time series features are used as Query, and variable features are used as Key and Value. Where: Y V represents the variable feature matrix, Y T represents the time series feature matrix, W Q Indicates that the time series feature Y T Projected as Query matrix, W K Indicates that the time series feature Y T Projection is the Key matrix, W V Indicates that the time series feature Y T The projection is the Value matrix, d represents the feature dimension, Softmax(·) represents the normalized attention weight, and Y fuse Represents the fused spatiotemporal joint features; S5.2, residual connection: retain the original spatiotemporal features, Y final =Y fuse +Conv(X spatial ), where: Y final represents the final output, Conv(·) represents the convolution operation, and X spatial Represents the original spatial characteristics.

5. The Mamba-based traffic flow prediction method according to claim 1, characterized in that: The fusion features obtained in S5 are output as future multi-step sequences through linear projection to achieve short-term prediction of traffic flow.

6. The Mamba-based traffic flow prediction method according to claim 1, characterized in that: The prediction method in S6 is evaluated using mean absolute error MAE, root mean square error RMSE, symmetric mean percentage error SMAPE and correlation coefficient.

Citation Information

Patent Citations

  • Unmanned surface ship cluster trajectory prediction method and system in uncertain environment

    CN119179863A

  • Cellular base station network flow prediction method based on Mama and graph neural network

    CN119232600A