A traffic flow prediction method and system based on a dynamic perception expert network

CN121686770BActive Publication Date: 2026-09-11TONGJI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511896343.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-09-11
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

但是该方案提出的模型未能有效解耦时间层面和空间层面的预测,对于复杂路况的适应能力较弱,且缺乏一定的可解释性

Benefits of technology

(1)本发明通过:使用双路径时间编码器分别捕捉节点特有的时序特征向量与全局共享的交通动态特征向量;随后通过尺度路由器,根据输入的交通数据的特征自适应选择对应时间尺度的专家网络,并利用尺度自适应扩张卷积和尺度条件动态图生成器提取特定尺度的时空特征向量;最后对所有时空特征向量进行加权聚合,输出交通状态的预测结果;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121686770B_ABST
    Figure CN121686770B_ABST
Patent Text Reader

Abstract

The application relates to a traffic flow prediction method and system based on a dynamic perception expert network, which comprises the following steps: acquiring traffic observation data, constructing a traffic network graph, performing feature embedding, and generating an initial space-time feature vector; inputting the initial space-time feature vector into a double-path time encoder, independently extracting an individualized time sequence feature vector and a time sequence dynamic feature vector through a parallel channel independent path and a channel mixed path; performing vector fusion through a gating mechanism to obtain a time context feature vector and input the time context feature vector into a multi-scale mixed expert model to activate multiple most relevant time scale experts and generate corresponding routing weights; generating a space-dependent graph through a scale conditional dynamic graph generator, performing graph convolution to extract multiple space-time feature vectors, and inputting the weighted and aggregated space-time feature vectors into a prediction model to generate a traffic prediction value. Compared with the prior art, the model constructed by the application has higher prediction accuracy and stronger robustness in a complex traffic scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent transportation systems (ITS) and data mining technology, and in particular to a traffic flow prediction method and system based on a dynamic perception expert network. Background Technology

[0002] With the acceleration of global urbanization and the continuous growth of motor vehicle ownership, urban traffic congestion has become an increasingly serious problem, a key bottleneck restricting urban operational efficiency, increasing energy consumption, and causing environmental pollution. As the foundation and core function of Intelligent Transportation Systems (ITS), accurate traffic flow and speed prediction has extremely high practical application value. On the one hand, accurate prediction results can assist traffic management departments in conducting scientific traffic guidance, signal timing optimization, and emergency response to sudden events, thereby effectively alleviating congestion and improving road capacity and safety. On the other hand, it can also provide the public with real-time road condition information and optimal route planning suggestions, significantly improving the travel experience and reducing unnecessary commuting time. Therefore, how to utilize massive traffic perception data to achieve high-precision and high-stability traffic condition prediction has become a hot issue that urgently needs to be addressed by both academia and industry.

[0003] In recent years, with the rapid development of big data and artificial intelligence technologies, data-driven deep learning methods have gradually replaced traditional statistical methods (such as ARIMA) and become the mainstream technology in the field of traffic prediction. In particular, spatiotemporal graph neural networks (STGNNs), by combining graph convolutional networks (GCNs) to handle spatial dependencies in non-Euclidean spaces and temporal convolutional networks (TCNs) or recurrent neural networks (RNNs) to handle time series dependencies, have made significant progress in traffic prediction tasks.

[0004] Short-term traffic flow (e.g., 5-10 minutes) is mainly affected by sudden events such as traffic accidents or traffic lights, exhibiting localized diffusion; while long-term traffic flow (e.g., over 1 hour) shows cyclical trends such as morning and evening peak hours or macroscopic connections between functional areas. More importantly, spatial relationships also evolve over time. On a short timescale, spatial dependence is mainly manifested as traffic flow diffusion between physically adjacent road segments (high connectivity); while on a long timescale, spatial dependence is more reflected in the semantic connections between distant but functionally similar areas (e.g., residential areas and office areas).

[0005] Existing technologies have failed to effectively decouple these heterogeneous spatiotemporal patterns, making it difficult for models to simultaneously capture short-term mutations and predict long-term trends. Furthermore, they often lack sufficient interpretability, making it difficult for managers to understand the logic behind the model's decisions. Therefore, there is an urgent need to develop a new traffic forecasting method that can dynamically adjust spatiotemporal modeling strategies according to time scales, while also possessing high accuracy and interpretability.

[0006] The invention disclosed in CN119541195A presents a full-cycle traffic flow prediction method based on deep fusion of spatiotemporal features. This method aims to deeply integrate the complex spatiotemporal dependencies of traffic flow, forming a composite traffic situational awareness and achieving comprehensive traffic flow prediction across multiple cycles. Firstly, through comprehensive spatiotemporal information representation, it extracts temporal features of multiple time modalities and constructs concrete real spatial relationships. To further construct a digital urban traffic flow prediction method that integrates multi-scale and multi-granular comprehensive elements, this invention proposes a comprehensive traffic flow prediction model that deeply integrates short-term spatiotemporal dependencies for short-term traffic flow prediction. This improves the sensitivity to instantaneous flow changes and fully models the short-term dynamic changes of traffic flow at multi-functional airspace nodes. Furthermore, a model combining deep airspace deconstruction and temporal feature fusion is designed for medium- and long-term traffic prediction, deeply extracting complex traffic patterns. Simulation results show that the proposed method outperforms existing traffic flow prediction technologies. However, the proposed model fails to effectively decouple the predictions at the temporal and spatial levels, has weak adaptability to complex road conditions, and lacks interpretability. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a traffic flow prediction method and system based on dynamic perception expert network. By introducing a hybrid expert (MoE) architecture, the complex spatiotemporal prediction task is decomposed into different time scales, realizing refined and differentiated modeling of spatiotemporal patterns at different scales.

[0008] The objective of this invention can be achieved through the following technical solutions: A traffic flow prediction method based on a dynamic perception expert network, the method comprising the following steps: Step 1: Obtain historical observation data of the traffic network, construct a traffic network map, and perform feature embedding to generate an initial spatiotemporal feature vector; Step 2: Input the initial spatiotemporal feature vector into the dual-path time encoder. Extract the node-specific personalized temporal feature vector and the network-wide shared temporal dynamic feature vector through parallel channel-independent paths and channel-mixed paths, respectively. Then, fuse the personalized temporal feature vector and the temporal dynamic feature vector through a gating mechanism to obtain the temporal context feature vector. Step 3: Input the temporal context feature vector into the multi-scale hybrid expert model, activate multiple most relevant temporal scale experts, and generate corresponding routing weights; each activated temporal scale expert generates a spatial dependency graph of the corresponding scale through a scale-conditional dynamic graph generator; perform graph convolution operation on the spatial dependency graph to extract spatiotemporal feature vectors; use the corresponding routing weights to perform weighted aggregation of all spatiotemporal feature vectors, and then feed them into the pre-built prediction model to generate traffic state prediction values.

[0009] Furthermore, the feature embedding process specifically includes: concatenating the traffic network graph with the learnable node embedding vector, intraday time embedding vector, and intraweek time embedding vector in the feature dimension, and mapping it to the hidden dimension through a linear transformation to obtain the initial spatiotemporal feature vector.

[0010] Furthermore, the channel-independent paths employ a shared-parameter temporal convolutional network and introduce learnable node-specific scaling and bias parameters. By changing these node-specific scaling and bias parameters, the features of each node are individually modulated, and finally, a unique temporal feature vector specific to each node is extracted, represented as: in, This represents the initial spatiotemporal features. This indicates scaling tensors. This represents a time-series convolutional network with shared parameters. This is the bias tensor.

[0011] Furthermore, after receiving the initial spatiotemporal feature vector, the channel hybrid path rearranges the input dimensions and uses two-dimensional convolutional layers with different dilation rates to process both spatial and temporal dimensions simultaneously in order to capture global interaction information and ultimately generate the globally shared temporal dynamic feature vector of the road network. The gating mechanism specifically involves: calculating the fusion weights using the Sigmoid function, and then weighting and fusing the personalized temporal feature vector and the temporal dynamic feature vector based on these fusion weights. The calculation formula is as follows: in, For time context feature vectors, For gating weights, For globally shared feature vectors; Linear indicates a fully connected layer. Denotes a nonlinear activation function, wherein the preferred one is... Activation function.

[0012] Furthermore, a one-to-one mapping relationship between time scales and expert networks is established in the multi-scale hybrid expert model, defining multiple time scale experts corresponding to different time granularities. The scale router includes convolutional layers and global average pooling layers to calculate the matching scores of experts at each time scale for the current traffic scenario.

[0013] Each activated timescale expert first extracts the time feature vector of the corresponding scale using scale-adaptive dilated convolution, and then generates the spatial dependency graph of the corresponding scale through a scale-conditional dynamic graph generator.

[0014] Furthermore, the scale-conditional dynamic graph generator introduces a scale-dependent decay threshold function and employs an attention-based method to generate the spatial dependency graph at the corresponding scale; the calculation formula for generating the spatial dependency graph is as follows: in, For time scale, For temperature parameters; Represents a sparse dynamic graph. Indicates the first Each attention head corresponds to a query vector. Indicates the first The transpose of the key vector corresponding to each attention head, where d represents the feature dimension of the query vector and the key vector, is used to normalize the attention score. Represents the scale decay threshold function. The initial adjacency matrix, To indicate the first The intermediate adjacency matrix corresponding to each attention head. Represents the identity matrix. This represents an enhanced node representation. Represents a dynamic node vector. This indicates a parameterized convolution operation. This represents a scale-aware attention module. Represented as a time feature, Indicates input features, Represents the balance coefficient. This indicates static node embedding. This represents the scale-dependent linear projection matrix. Indicates the first The query vector projection matrix corresponding to each attention head. Indicates the first The key vector corresponding to each attention head.

[0015] Furthermore, the process of activating experts across multiple most relevant timescales specifically includes: The scale router dynamically calculates the matching scores of multiple predefined time-scale experts based on the real-time time features of the time context feature vector; based on the matching scores, it selects multiple target experts most suitable for the current traffic scenario and calculates their corresponding routing weights; the formula for calculating the matching scores of the predefined multiple time-scale experts is as follows: in, For Logits score, For global average pooling, Represents the ReLU nonlinear activation function. Represents the convolution kernel parameters. It is the representation after transposing the input features; MLP stands for Multilayer Perceptron. The formula for calculating the corresponding route weight is: in, To normalize routing weights, For the sample Middle node The selected set of target experts, For timescale experts The set of Top-K experts that are active at the current node, where Top-K represents the highest score. One expert.

[0016] Furthermore, step three also includes: After weighting and aggregating all spatiotemporal feature vectors using the corresponding routing weights to generate an aggregated feature vector, which is then fed into a pre-built prediction model, the prediction model performs a residual connection between the aggregated feature vector and the temporal context feature vector to obtain the fused final feature vector. Finally, the fused final feature vector is input into the prediction output layer to map and generate traffic state prediction values ​​for multiple future time steps. The formulas for generating traffic state predictions for multiple future time steps include: in, For predicted values, Indicates residual connection, For the fused feature representation, This is the weight matrix. For spatial enhancement features, The normalized adjacency matrix, It is a time-related feature.

[0017] Furthermore, the prediction output layer includes a linear projection layer, which maps the aggregated high-dimensional feature vector to traffic state predictions for multiple future time steps. The prediction model is trained by optimizing its parameters using a weighted combination of the prediction error loss function and the load balancing auxiliary loss function. The weighted loss function is as follows: in, Represents the actual traffic condition value. This represents the prediction error loss function, which measures the deviation between the prediction result and the true value. It is preferably the mean absolute error loss or the mean square error loss. For loss weighting coefficients, This indicates that the person is routed to the timescale expert in a training batch. The node ratio, where Var represents variance calculation.

[0018] The present invention also provides a system for a traffic flow prediction method based on a dynamic perception expert network, comprising a memory and a processor, wherein the memory stores a computer program, and the processor calls the computer program to execute the steps of any of the methods described above.

[0019] Compared with the prior art, the present invention has the following advantages: (1) This invention uses a dual-path time encoder to capture the node-specific temporal feature vector and the globally shared traffic dynamic feature vector respectively; then, through a scale router, it adaptively selects the expert network corresponding to the time scale according to the features of the input traffic data, and uses scale-adaptive dilated convolution and scale-conditional dynamic graph generator to extract spatiotemporal feature vectors of a specific scale; finally, it performs weighted aggregation on all spatiotemporal feature vectors and outputs the prediction result of traffic status. It achieves effective decoupling of spatiotemporal modeling and time scale, significantly improving the accuracy and robustness of the model in both short-term and long-term predictions under complex traffic scenarios. Furthermore, it enhances the interpretability of the model by visualizing expert activation weights, which can clearly reveal the intrinsic mechanism of traffic state evolution.

[0020] (2) This invention constructs a multi-scale hybrid expert model, and selects the K target experts and their corresponding routing weights most suitable for the current traffic scenario from multiple predefined scale experts by using the Top-K strategy. This realizes the dynamic selection of the most suitable expert combination according to the traffic congestion status (such as peak or off-peak hours), and enhances the model's adaptability to complex road conditions.

[0021] (3) By establishing a one-to-one mapping between time scales and expert networks, this invention enables the model to distinguish between short-term local diffusion and long-term functional association, significantly improving the accuracy of long-term prediction. The visualization of expert activation weights can reveal the time scales that the model focuses on at different time periods (such as activating short-scale experts during peak periods), providing a basis for traffic management decisions. It can also provide more accurate and reliable decision-making basis for urban traffic congestion relief, intelligent traffic light timing, and residents' travel planning. It has important application value for building an efficient and green intelligent transportation system. Attached Figure Description

[0022] Figure 1 This is a flowchart of a traffic flow prediction method based on a dynamic perception expert network provided in an embodiment of the present invention; Figure 2 This is a model structure diagram of a traffic flow prediction method based on a dynamic perception expert network provided in Embodiment 3 of the present invention; Figure 3 This is a schematic diagram of the structure of a dual-path time encoder for a traffic flow prediction method based on a dynamic perception expert network provided in Embodiment 3 of the present invention. Figure 4 This is a schematic diagram of the internal structure of an expert network-based traffic flow prediction method provided in Embodiment 3 of the present invention.

[0023] Figure 2 This is a flowchart illustrating the overall architecture of the model of this invention. For example... Figure 2 As shown, the data flow of this method has a clear hierarchical progression: first, the original data is mapped to a high-dimensional initial embedding representation; then, temporal context features are extracted through a dual-path encoder; next, a gating mechanism is used to select specific scale experts; the activated experts generate dynamic graph structures and extract spatiotemporal features; finally, all features are aggregated to generate prediction results. Figure 3 This is a schematic diagram of a dual-path time encoder, demonstrating the parallel processing and fusion of specific paths and global paths. Figure 4 This is a schematic diagram of the internal structure of the Scale Expert, illustrating the processes of temporal extraction, scale attention, and dynamic graph convolution. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0025] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0026] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0027] Definitions: The Sigmoid function is a widely used sigmoid curve function in mathematics and machine learning. It maps any real input to the interval (0,1), and the output value can be intuitively understood as a probability or normalized signal. Therefore, it is often used as an activation function in logistic regression, binary classification, and early neural networks. Its monotonically continuous and differentiable properties facilitate gradient calculation, but it is prone to gradient saturation problems and has been gradually replaced by functions such as ReLU in deep networks.

[0028] Example 1 like Figure 1 As shown, this embodiment provides a traffic flow prediction method based on a dynamic perception expert network. The method includes the following steps: S1: Obtain historical observation data of the traffic network, construct a traffic network map, and perform feature embedding to generate an initial spatiotemporal feature vector; and fuse node features, intraday time features, and intraweek time features for embedding and encoding to generate an initial spatiotemporal feature representation; Specifically, The process of constructing an initial feature representation rich in spatiotemporal context based on historical observation data of the traffic network is as follows: First, standardization is performed, and node feature embeddings, intraday time feature embeddings, and intraweek time feature embeddings are generated respectively. Then, the above embedding vectors are concatenated and fused with the original observation data in the feature dimension, and an initial high-dimensional feature representation rich in multidimensional context information is constructed through linear mapping, providing a unified data foundation for subsequent time-series feature extraction.

[0029] Specifically, embedding encoding involves concatenating the original traffic observation features with learnable node embeddings, time-of-day embeddings, and day-of-week embeddings, and then mapping them to the hidden dimensions through a linear transformation.

[0030] S2: Input the initial spatiotemporal feature vector into the dual-path time encoder. Extract the node-specific personalized temporal feature vector and the network-wide shared temporal dynamic feature vector through parallel channel-independent paths and channel-mixed paths respectively. Then, fuse the personalized temporal feature vector and the temporal dynamic feature vector through a gating mechanism to obtain the temporal context feature vector. Preferred, A hybrid temporal feature extraction based on a dual-path mechanism is performed. The specific steps include: inputting the initial high-dimensional feature representation obtained in step 1 into a dual-path temporal encoder. This encoder runs two paths in parallel: the channel-independent path uses node adaptive parameters to personalize the input features to capture the unique traffic patterns of each sensor; the channel-hybrid path uses two-dimensional convolution to process the input features to capture the global dynamic changes of the entire road network. Finally, an adaptive gating mechanism dynamically fuses the features extracted from these two paths to generate temporal context features that take into account both node individuality and road network commonality, serving as the core input for subsequent multi-scale modeling.

[0031] Specifically, The dual-path temporal encoder includes: a channel-independent path: employing a shared-parameter temporal convolutional network (TCN) and introducing learnable node-specific scaling and bias parameters to personalize the features of each node; a channel-blended path: rearranging the input dimensions and using two-dimensional convolutional layers with different dilation rates to simultaneously process spatial and temporal dimensions, capturing global interaction information; and a gating mechanism: using the Sigmoid function to calculate fusion weights and performing a weighted summation of the two outputs.

[0032] S3: Input the temporal context feature vector into the multi-scale hybrid expert model. Utilize a scale router based on the Top-K strategy to adaptively select and activate the K most relevant temporal scale experts according to the current traffic state, and generate corresponding routing weights. Each activated temporal scale expert generates a spatial dependency graph of the corresponding scale through a scale-conditional dynamic graph generator. Perform graph convolution operation on the spatial dependency graph to extract spatiotemporal feature vectors. Use the corresponding routing weights to perform weighted aggregation of all spatiotemporal feature vectors, and then feed them into the pre-built prediction model. After processing through residual connections and prediction output layers, traffic flow prediction results for multiple future time steps are generated.

[0033] S301: Scale-based expert routing decision based on real-time features. The temporal context features output from step 2 are input into the scale router of the multi-scale hybrid expert module. The scale router dynamically calculates the matching scores of multiple predefined "scale experts" (corresponding to different time scales such as 10 minutes, 20 minutes...60 minutes) according to the real-time time pattern of the current input data. Based on the matching scores, a Top-K strategy is used to select the K target experts most suitable for the current traffic scenario and their corresponding routing weights, thereby determining the specific path for subsequent feature processing.

[0034] S302: Scale-Adaptive Spatiotemporal Feature Deep Modeling. Only the K target experts selected in step 3 are activated, and the temporal context features output in step 2 are differentiated. Each activated expert performs the following operations: (1) Scale-adaptive temporal convolution is performed on the input features using a dilation rate that matches the scale to which the expert belongs; (2) A dynamic adjacency matrix is ​​generated based on the temporal scale attribute of the expert using a scale-conditional dynamic graph generator; (3) Graph convolution is performed on the temporal features based on the generated dynamic graph to extract the spatial enhancement features at that specific scale.

[0035] S303: Multi-expert Feature Aggregation and Prediction Generation. The spatial augmentation features output by all activated experts in step 4 are collected and weighted and summed according to the routing weights calculated in step 3, achieving adaptive aggregation of multi-scale spatiotemporal features. Subsequently, the aggregated features are residually connected with the features output in step 2 to preserve original information and stabilize gradients. Finally, the fused final features are input into a fully connected prediction layer to map and generate traffic state predictions for multiple future time steps.

[0036] Specifically, The multi-scale hybrid expert module establishes a one-to-one mapping relationship between time scales and expert networks, defining multiple scale experts corresponding to different time granularities; the scale router contains convolutional layers and global average pooling layers to calculate the matching score of each expert for the current input, and introduces Gaussian noise to promote load balancing during training.

[0037] Preferred, The scale-conditional dynamic graph generator constructs dynamic graphs using an attention-based method, characterized by the introduction of a scale-dependent decay threshold function. ,in For time scale, This refers to the temperature parameter.

[0038] Preferred, The prediction output layer includes a linear projection layer, which maps the aggregated high-dimensional features to the future. The predicted values ​​at each time step are used; during model training, a weighted combination of the prediction error loss function and the load balancing auxiliary loss function is used for parameter optimization.

[0039] Example 2 This embodiment provides a specific implementation plan for a traffic flow prediction method based on a dynamic perception expert network as shown in Embodiment 1, including: Step 1: Data preprocessing and multidimensional feature embedding transform the original traffic observation data into a high-dimensional feature representation rich in spatiotemporal semantics.

[0040] (1.1) Data definition: Obtain data from the target traffic network (such as highways or urban roads). One sensor in Historical observation data (such as traffic flow or speed) at each time step. Construct an observation matrix. ,in Historical time step (e.g., 12 steps). For the number of nodes, The data consists of features (including flow values, timestamps, etc.). Z-Score standardization is applied to the data to eliminate the influence of different units.

[0041] (1.2) Multidimensional Feature Embedding: To incorporate rich contextual information, this invention introduces three learnable embedding vectors: Node Embedding, ): This is used to characterize the static spatial attributes (such as geographic location and road type) of each sensor node. Time-of-Day Embedding... ): This is used to divide a day into several time slices (e.g., 288 five-minute segments) to capture daily cyclical patterns such as morning and evening rush hours. Day-of-Week Embedding... ): This is used to distinguish the differences in transportation patterns between weekdays and weekends.

[0042] (1.3) Feature fusion: combining the original input The three embedding vectors mentioned above are concatenated along the feature dimension and projected onto the hidden layer dimension of the model through a fully connected (Linear Layer). The initial spatiotemporal feature representation is obtained. (in (for batch size) Step 2: Dual-path temporal feature extraction: In order to take into account both the individualized patterns of nodes and the common patterns of the global road network, a parallel dual-path temporal encoder was designed in this step.

[0043] (2.1) Specific Path: Used to capture the unique traffic patterns of each node. A parameter-sharing Temporal Convolutional Network (TCN) is employed, but two node-specific learnable parameter tensors are introduced: a scaling tensor. and bias tensor The output of this path The calculation is as follows: in This indicates element-wise multiplication. This mechanism allows the model to adaptively adjust the features extracted by the shared TCN to suit the specificity of different nodes.

[0044] (2.2) Global Path: Used to capture the global dynamics of the entire road network. First, the input dimensions are rearranged, treating the feature dimensions as channels. Then, two-dimensional dilated convolutional layers with different dilation rates ((1,1) and (2,2)) are used to extract globally shared features. Convolution operations are performed simultaneously in the temporal and spatial dimensions. .

[0045] (2.3) Adaptive Gated Fusion: Gated weights generated using the Sigmoid activation function The two sets of features are dynamically weighted and fused to obtain the temporal context features. : Step 3: Scale expert routing decision based on real-time features: This step utilizes the "Scale Router" in the Hybrid Expert (MoE) architecture to select the most suitable time scale expert for the current input data.

[0046] (3.1) Matching score calculation: Calculate the matching score using a scale router and the input features. The matching degree of predefined scale experts. The router first uses 1D convolution and global average pooling ( Aggregate time information and then generate Logits scores using MLP. : To improve stability, Gaussian noise was added to the Logits. .

[0047] (3.2) Top-K sparse activation: only activate the highest-scoring activation. One expert ( ). No. Normalized routing weights of individual experts Calculation as follows : in In other cases, The value is 0. This step outputs the route weight. The activation index will directly control the expert module that performs the calculations in step 4.

[0048] Step 4: Scale-Adaptive Spatiotemporal Feature Depth Modeling: For each specific scale expert activated by the gating mechanism in Step 3 (denoted as the ), Each expert, corresponding time scale The following deep feature extraction operations are performed respectively: (4.1) Scale-Adaptive Temporal Feature Extraction: In order to capture the temporal dependencies that match the current scale, the first... Each expert uses a set of parallel 1D dilated convolution branches to process the input features. The dilation rate of the convolution kernel varies depending on the scale. By setting parameters (the larger the scale, the greater the expansion rate), a temporal feature representation specific to that scale can be obtained. : in This indicates a parameterized convolution operation.

[0049] (4.2) Dynamic node representation generation: In order to integrate dynamic temporal information with static node attributes, a scale-aware attention module is first used. Aggregate temporal features to generate dynamic node vectors Subsequently, with learnable static node embeddings By merging, an enhanced node representation is obtained. : in This is the balance coefficient.

[0050] (4.3) Generation of dynamic graph under scale conditions: This step uses the above node representation to construct a dynamic graph.

[0051] First, perform initial attention calculations: As the query and key, the original association strength between nodes is calculated using a multi-head attention mechanism. For the first... Size: The initial adjacency matrix is ​​obtained by averaging the values ​​of all heads. Then define the scale decay threshold function. : in The temperature parameter is learnable. Thresholding is used to... Filtering is performed to generate the final sparse dynamic graph. : Based on the generated dynamic graph Regarding time characteristics Perform graph convolution operations to extract spatial augmentation features. : in The normalized adjacency matrix, This is the weight matrix.

[0052] Step 5: Multi-expert feature aggregation and prediction generation (5.1) Route-weighted aggregation: Collect spatial augmentation features from all activated expert outputs. And use the routing weights calculated in step 3 We perform a weighted summation to obtain the fused feature representation. : This step enables adaptive integration of spatiotemporal patterns at different scales.

[0053] (5.2) Residual Connectivity and Prediction: Introducing Residual Connectivity Finally, the future is generated through a linear projection layer. Predicted value of step : (5.3) During training, a weighted sum of prediction error loss and load balancing loss is used: To verify the effectiveness of this method, comparative experiments were conducted on publicly available traffic datasets (PEMS03, PEMS04, PEMS07, PEMS08). The experiments were set with a historical input step size of 12 (60 minutes) and predictions for the next 12 steps (60 minutes). Experimental results show that this invention outperforms existing baseline models (such as STGCN, GraphWaveNet, AGCRN, etc.) in all error metrics. Taking the PEMS04 dataset as an example, the specific metrics of this method on the 60-minute prediction task are as follows: Mean Absolute Error (MAE): 18.49 (compared to the suboptimal model AGCRN's 19.83, a reduction of approximately 6.7%). Root Mean Square Error (RMSE): 30.08. Mean Absolute Percentage Error (MAPE): 12.25%. The data demonstrate that this method can significantly reduce prediction error and improve prediction accuracy.

[0054] Example 3 like Figure 2 , Figure 3 and Figure 4 As shown, this embodiment provides a system for traffic flow prediction based on a dynamic perception expert network, including a memory and a processor. The memory stores a computer program, and the processor calls the computer program to execute the steps of any of the methods in Embodiment 1.

[0055] In this invention, "historical observation data of the traffic network" is the only external raw input data. All other content (traffic network diagram, node embedding, time embedding, dynamic graph, expert routing, etc.) are intermediate representations constructed or generated internally by the system based on this data. It should be noted that the historical observation data of the traffic network, as the only external input data of the method of this invention, contains time information and spatial node information that can be used in subsequent processing to construct traffic network diagrams, time embedding vectors, node embedding vectors, and dynamic spatial dependencies. The above-mentioned traffic network structure, feature embedding, scale expert, and routing weights are all automatically generated by the model based on historical observation data during training or inference, without the need to introduce additional external data.

[0056] The beneficial effects of this invention are as follows: (1) Multi-scale decoupling modeling: A one-to-one mapping between time scale and expert network was established, which enabled the model to distinguish between short-term local diffusion and long-term functional association, significantly improving the accuracy of long-term prediction.

[0057] (2) Dynamic adaptability: Through the Top-K routing mechanism, the model can dynamically select the most suitable expert combination according to the traffic congestion status (such as peak or off-peak hours), which enhances the model's adaptability to complex road conditions.

[0058] (3) Balancing individuality and commonality: The dual-path time encoder effectively combines the unique traffic patterns of nodes with the overall traffic trends of the city, thereby improving the robustness of prediction.

[0059] (4) Interpretability: Visualization of expert activation weights can reveal the time scales that the model focuses on at different time periods (such as activating short-scale experts during peak periods), providing a basis for traffic management decisions.

[0060] This invention addresses the problem of existing traffic prediction models' inability to distinguish spatial dependencies across different time scales. It proposes a multi-scale modeling framework based on a hybrid expert (MoE) architecture. This framework not only effectively decouples spatiotemporal modeling from time scales, significantly improving the model's short- and long-term prediction accuracy and robustness in complex traffic scenarios, but also enhances the model's interpretability through visualized expert activation weights, clearly revealing the intrinsic mechanisms of traffic state evolution. This method can provide more accurate and reliable decision-making support for urban traffic congestion management, intelligent traffic light timing, and resident travel planning, and has significant application value for building efficient and green intelligent transportation systems.

[0061] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.

[0062] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A traffic flow prediction method based on a dynamic perception expert network, characterized in that, include: Step 1: Obtain historical observation data of the traffic network, construct a traffic network map, and perform feature embedding to generate an initial spatiotemporal feature vector; The feature embedding vectors include learnable node embedding vectors, intraday time embedding vectors, and intraweek time embedding vectors. Step 2: Input the initial spatiotemporal feature vector into the dual-path time encoder. Extract the node-specific personalized temporal feature vector and the network-wide shared temporal dynamic feature vector through parallel channel-independent paths and channel-mixed paths, respectively. Then, fuse the personalized temporal feature vector and the temporal dynamic feature vector through a gating mechanism to obtain the temporal context feature vector. Step 3: Input the time context feature vector into the multi-scale hybrid expert model, activate multiple most relevant time-scale experts, and generate corresponding routing weights; Each activated timescale expert generates a spatial dependency graph at the corresponding scale through a scale-conditional dynamic graph generator; the spatial dependency graph is subjected to graph convolution to extract spatiotemporal feature vectors; all spatiotemporal feature vectors are weighted and aggregated using the corresponding routing weights, and then fed into a pre-built prediction model; the prediction model performs a residual connection between the aggregated feature vectors and the temporal context feature vectors to obtain the fused final feature vector; finally, the fused final feature vector is input into the prediction output layer to map and generate traffic state prediction values ​​for multiple future time steps; The multi-scale hybrid expert model establishes a one-to-one mapping relationship between time scales and expert networks, and defines multiple time scale experts corresponding to different time granularities. The scale router dynamically calculates the matching scores of multiple predefined time-scale experts based on the real-time time features of the time context feature vector; based on the matching scores, it selects the multiple target experts most suitable for the current traffic scenario and calculates the corresponding routing weights. The scale router includes convolutional layers and global average pooling layers, which are used to calculate the matching scores of experts at each time scale for the current traffic scenario. Each activated timescale expert first extracts the time feature vector of the corresponding scale using scale-adaptive dilated convolution, and then generates the spatial dependency graph of the corresponding scale through a scale-conditional dynamic graph generator.

2. The traffic flow prediction method based on a dynamic perception expert network according to claim 1, characterized in that, The feature embedding process specifically includes: concatenating the traffic network graph with the learnable node embedding vector, intraday time embedding vector, and intraweek time embedding vector in the feature dimension, and mapping it to the hidden dimension through a linear transformation to obtain the initial spatiotemporal feature vector.

3. The traffic flow prediction method based on a dynamic perception expert network according to claim 1, characterized in that, The independent paths of the channels employ a shared-parameter temporal convolutional network, and introduce learnable node-specific scaling and bias parameters. By changing these node-specific scaling and bias parameters, the features of each node are individually modulated, and finally, a unique temporal feature vector specific to each node is extracted, represented as follows: in, This represents the initial spatiotemporal features. This indicates scaling tensors. This represents a time-series convolutional network with shared parameters. This is the bias tensor.

4. The traffic flow prediction method based on a dynamic perception expert network according to claim 1, characterized in that, After receiving the initial spatiotemporal feature vector, the channel hybrid path rearranges the input dimensions and uses two-dimensional convolutional layers with different dilation rates to process the spatial and temporal dimensions simultaneously in order to capture global interaction information and finally generate the globally shared temporal dynamic feature vector of the road network. The gating mechanism specifically involves: calculating the fusion weights using the Sigmoid function, and then weighting and fusing the personalized temporal feature vector and the temporal dynamic feature vector based on these fusion weights. The calculation formula is as follows: in, For time context feature vectors, For gating weights, For globally shared feature vectors; Linear indicates a fully connected layer. This represents the non-linear activation function Sigmoid.

5. The traffic flow prediction method based on a dynamic perception expert network according to claim 1, characterized in that, The scale-conditional dynamic graph generator generates a spatial dependency graph at the corresponding scale by introducing a scale-dependent decay threshold function and employing an attention-based method; the calculation formula for generating the spatial dependency graph is as follows: in, For time scale, For temperature parameters; Represents a sparse dynamic graph. Indicates the first The query vector corresponding to each attention head. Indicates the first The transpose of the key vector corresponding to each attention head, where d represents the feature dimension of the query vector and the key vector, is used to normalize the attention score. Represents the scale decay threshold function. The initial adjacency matrix, To indicate the first The intermediate adjacency matrix corresponding to each attention head. Represents the identity matrix. This represents an enhanced node representation. Represents a dynamic node vector. This indicates a parameterized convolution operation. This represents a scale-aware attention module. Represented as a time feature, Indicates input features, Represents the balance coefficient. This indicates static node embedding. This represents the scale-dependent linear projection matrix. Indicates the first The query vector projection matrix corresponding to each attention head. Indicates the first The key vector corresponding to each attention head.

6. The traffic flow prediction method based on a dynamic perception expert network according to claim 1, characterized in that, The formula for calculating the matching score of experts across multiple predefined timescales is as follows: in, For Logits score, For global average pooling, Represents the ReLU nonlinear activation function. Represents the convolution kernel parameters. It is the representation after transposing the input features; MLP stands for Multilayer Perceptron. The formula for calculating the corresponding route weight is: in, To normalize routing weights, For the sample Middle node The selected set of target experts, For timescale experts The set of Top-K experts that are active at the current node, where Top-K represents the highest score. One expert.

7. The traffic flow prediction method based on a dynamic perception expert network according to claim 1, characterized in that, The calculation formula for generating traffic state predictions for multiple future time steps. include: in, For predicted values, Indicates residual connection, For the fused feature representation, This is the weight matrix. For spatial enhancement features, The normalized adjacency matrix, It is a time-related feature.

8. The traffic flow prediction method based on a dynamic perception expert network according to claim 7, characterized in that, The prediction output layer includes a linear projection layer, which is used to map the aggregated high-dimensional feature vector into traffic state prediction values ​​for multiple future time steps. The prediction model is trained by optimizing its parameters using a weighted combination of the prediction error loss function and the load balancing auxiliary loss function. The weighted loss function is as follows: in, Represents the actual traffic condition value. This represents the prediction error loss function, used to measure the deviation between the predicted result and the true value; For loss weighting coefficients, This indicates that the person is routed to the timescale expert in a training batch. The node ratio, where Var represents variance calculation.

9. A system for traffic flow prediction based on a dynamic perception expert network, characterized in that, It includes a memory and a processor, the memory storing a computer program, the processor invoking the computer program to perform the steps of the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Full-period traffic flow prediction method based on spatial-temporal feature deep fusion

    CN119541195A

  • A Spark-Based Deep Learning Method for Data-Driven Traffic Flow Forecasting

    AU2020102350A4

  • Traffic flow prediction method and system based on long and short term expert attention network, medium and equipment

    CN118070167A