Multimodal Traffic Flow Prediction Method Based on Dynamic Spatiotemporal Hypergraph and Large Language Model
By adopting a multimodal fusion method of dynamic spatiotemporal hypergraph and large language model in traffic flow prediction, the problem of difficulty in dynamic modeling and long-term prediction in the existing technology is solved, and traffic flow prediction with higher accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510426857.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Existing traffic flow prediction methods are difficult to accurately model the dynamic changes in traffic flow, especially with challenges in emergency and long-term prediction.
The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model is adopted. The space-time dependence relationship of traffic flow is captured through the dynamic spatiotemporal hypergraph learning module (DSTHL). The text feature integration module extracts the semantic features of traffic event text, and uses the feature extraction module based on large language model (FEM-LLM) to integrate multi-scale spatiotemporal dependence and text semantic features.
It significantly improves the accuracy and robustness of traffic flow prediction, especially in emergency events and long-term prediction, which can more accurately characterize traffic conditions and capture complex spatiotemporal patterns.
Smart Images

Figure CN119942803B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer machine learning, and specifically to a multimodal traffic flow prediction method based on dynamic spatio-temporal hypergraphs and large language models. Background Art
[0002] As a core component of intelligent transportation systems, the evolution of traffic flow prediction technology has always centered around three major goals: enhancing spatio-temporal feature expression capabilities, improving the integration effect of multi-source data, and optimizing long-term prediction accuracy. The current state of technology development in each stage is as follows:
[0003] (1) Application of historical traffic flow data in traffic flow prediction
[0004] Traffic flow prediction methods have evolved from traditional statistical methods to deep learning and then to methods based on graph neural networks. Early research mainly used statistical models and machine learning methods, such as the Seasonal Autoregressive Integrated Moving Average (SARIMA) model proposed by Billy M. Williams and Lester A. Hoel (2003), and the use of Support Vector Regression (SVR) combined with historical data for prediction by Wu et al. (Chun-Hsin Wu and Jan-Ming Ho and Lee, D.T. 2004). Although these methods can capture the periodicity of time series, they are difficult to characterize complex non-linear spatio-temporal relationships.
[0005] With the development of deep learning, researchers began to use Recurrent Neural Networks (RNNs), Long Short-Term Memory Networks (LSTMs) (Ma et al. 2015), and Gated Recurrent Units (GRUs) (Junyoung Chung et al. 2014) for traffic flow prediction. Zhao et al. (Zhao, Jun and Zhu, Wen-Xing 2021) enhanced spatio-temporal feature extraction capabilities through Deep Convolutional GRU (DCGN), and Guo et al. (Guo, Shengnan et al. 2019) used ST-3DNet combined with 3D convolution to simultaneously model spatial and temporal features. However, these methods have limitations in dealing with irregular traffic network structures and are difficult to comprehensively characterize complex traffic flow changes.
[0006] To address this issue, researchers introduced graph neural networks (GNNs) and combined graph convolutional networks (GCNs) for traffic flow modeling. STGCN (Bing Yu and Haoteng Yin and Zhanxing Zhu 2018) uses graph convolution to extract spatial features and combines temporal convolution for temporal modeling. MTGNN (Wu, Zonghan et al. 2020) improves the model's adaptability to dynamic traffic relationships by learning implicit graph structures. STSGCN (Song, Chao et al. 2020) further adopts a modular design to capture multi-scale spatio-temporal features. However, existing methods generally rely on static graph structures and are difficult to accurately model the dynamic changes of traffic flow. Therefore, further exploration of adaptive dynamic modeling methods is still needed.
[0007] (2) Application of Text Information in Traffic Flow Prediction
[0008] Text data such as social media can provide real-time traffic condition information and become an important data source for assisting traffic flow prediction. Researchers have explored methods to enhance prediction capabilities using text data such as social media and news reports and achieved certain results.
[0009] Abidin et al. (Abidin, Ahmad Faisal et al. 2014) proposed combining trustworthy information from social networks to improve the accuracy of public transportation prediction. Ni et al. (Ni, Ming and He, Qing and Gao, Jing 2017) combined social media data with historical passenger flow data and used support vector regression to improve the subway passenger flow prediction accuracy under emergency situations. Aniekan et al. (Essien, Aniekan et al. 2021) integrated real-time traffic, weather, and social media tweets (such as Twitter) and constructed an urban traffic prediction model based on a bidirectional LSTM autoencoder. Yao et al. (Yao, Weiran and Qian, Sean 2021) inferred the daily routines of residents by analyzing Twitter information and studied its impact on traffic conditions. Maryam et al. (Shoaeinaeini, Maryam and Ozturk, Oktay and Gupta, Deepak 2022) designed an urban traffic prediction model that combines social network data and verified its effectiveness based on the California PeMS dataset.
[0010] Although text data has shown potential in assisting traffic flow prediction, most studies mainly extract statistical information from social media while neglecting the rich semantic content in the text. Therefore, how to effectively integrate text data with traffic data, especially using the semantic information of the text to improve prediction accuracy, remains a key challenge.
[0011] (3)Application of Large Language Models in Traffic Flow Prediction
[0012] The development of large language models (LLMs) has brought new opportunities and challenges to traffic flow prediction. Researchers have explored how to use these models to process complex spatio-temporal data and provide new perspectives on traffic data analysis through natural language understanding and generation.
[0013] UrbanGPT (Li, Zhonghang et al., 2024) models the complex spatio-temporal dependencies of the urban traffic system. GGT (Wang, Xuhong et al., 2023) combines graph neural networks and Transformer to accurately capture the dynamic changes of traffic flow. Traffic Transformer (Jin, KyoHoon et al., 2021) uses the self-attention mechanism to effectively model complex spatio-temporal patterns. TrafficGPT (Siyao Zhang et al., 2023) processes multi-modal data under the GPT framework and is applicable to various traffic tasks. TrafficBERT (Jin, KyoHoon et al., 2021) combines natural language processing techniques to analyze traffic text data and realizes the combination of NLP and traffic analysis. ST-LLM (Liu, Chenxi et al., 2024) models deep spatio-temporal dependencies through large language models, while STG-LLM (Lei Liu et al., 2024) introduces graph convolution to process the spatial structure of the traffic network. In addition, PBTR (Duan, Wenying et al., 2019) uses a bidirectional encoder to capture long-term dependence features and is applicable to traffic data sets with long-term dependence characteristics.
[0014] Although large language models have shown great potential in traffic flow prediction, they still face challenges in semantic alignment of multi-source data, especially the semantic information matching problem when integrating text and traffic data. In addition, LLMs still have limitations in deep time pattern mining and context information extraction, which restricts their application effects in traffic flow prediction.
[0015] Despite many advances, traffic prediction still faces the following key challenges:
[0016] (1)Unpredictability of emergencies: Emergencies such as traffic accidents, concerts, and sports events often disrupt the original traffic patterns, leading to drastic changes in traffic conditions. Traditional single data sources (such as traffic flow data) are difficult to accurately reflect the impact of these emergencies on traffic, resulting in a decline in prediction accuracy.
[0017] (2)Limitations of spatial correlation: Many traffic prediction methods rely on the road network structure and the spatial correlation between adjacent road segments. However, in some cases (such as similar traffic patterns in residential and commercial areas during peak hours), even if there is no direct spatial association between road segments, their traffic patterns may still be highly similar. Existing methods are difficult to effectively model these non-spatially related scenarios, thus limiting the prediction accuracy.
[0018] (3)Difficulties in long-term prediction: Existing time series models are difficult to accurately capture the long-term dependencies of traffic flow, making long-term traffic prediction a major challenge. In addition, the influence of external factors such as weather changes, special events, and road construction makes long-term prediction even more difficult.
[0019] With the continuous popularization of intelligent transportation systems, a large amount of traffic flow data and event text data provide new opportunities to solve the above problems. Therefore, how to effectively integrate traffic flow data with the text description of traffic events and improve the spatio-temporal feature expression and long-term prediction ability has become a key problem to be solved urgently. Summary of the Invention
[0020] The purpose of the present invention is to provide a multi-modal traffic flow prediction method based on a dynamic spatio-temporal hypergraph and a large language model to solve the problems proposed in the above background technology.
[0021] To achieve the above purpose, the present invention provides the following technical solutions:
[0022] A multi-modal traffic flow prediction method based on a dynamic spatio-temporal hypergraph and a large language model, comprising the following steps:
[0023] Obtain a traffic data set, align the timestamps of text descriptions and sensor data, and divide the training set, validation set, and test set;
[0024] Construct a dynamic spatio-temporal hypergraph to retain key spatio-temporal dependencies;
[0025] Construct a traffic flow prediction model, use the training set to train the model, monitor the training process with the validation set and adjust the hyperparameters, and finally use the test set to evaluate the prediction performance of the model;
[0026] Among them, the traffic flow prediction model includes a dynamic spatio-temporal hypergraph learning module DSTHL, a text feature integration module, a feature extraction module FEM-LLM based on a large language model, and regression prediction:
[0027] Dynamic Spatiotemporal Hypergraph Learning Module (DSTHL): Construct a dynamic hypergraph using the FastDTW algorithm and capture the spatiotemporal dependencies of traffic flow by combining with a graph convolutional network;
[0028] Text Feature Integration Module: Extract the semantic features of traffic event texts, and splice them into a multi-modal representation after aligning with spatiotemporal features;
[0029] Feature Extraction Module Based on Large Language Model (FEM-LLM): Model multi-scale spatiotemporal dependencies by partially freezing the Transformer architecture, and fuse and output deep spatiotemporal representations by combining with text semantic features;
[0030] Predict the future traffic flow through the regression layer using the deep spatiotemporal representation output by FEM-LLM to obtain the final prediction result.
[0031] The construction of the dynamic spatiotemporal hypergraph includes the following steps:
[0032] To effectively capture the inherent spatiotemporal dependencies in traffic flow data, we propose a dynamic spatiotemporal hypergraph , which is used to quantify the pairwise similarity between different traffic segments. Since the traffic flow at a specific location is affected by both time patterns and spatial interactions, by representing these relationships as a graph structure, we can efficiently encode these interdependencies, thus providing a solid foundation for subsequent spatiotemporal prediction models. First, apply the FastDTW algorithm to calculate the dynamic time warping (DTW) distance of traffic flow data, and the distance formula is as follows:
[0033]
[0034] where and represent the traffic flow sequences of segments i and j respectively, and represent the traffic flow values of segments i and j at time step t respectively, is the standard deviation of the traffic flow sequence of segment i, and T is the total number of time steps within the time window.
[0035] Using these pairwise distances, construct a symmetric N×N adjacency matrix G t , where N is the total number of segments. Each element of the matrix is defined as follows:
[0036]
[0037] Here, is a threshold corresponding to the smallest 5% in the DTW distance to ensure that only the most important relationships are retained. This dimensionality reduction process helps reduce noise while preserving key spatio-temporal dependencies. The resulting graph is the dynamic spatio-temporal hypergraph, which intuitively depicts the traffic flow dynamics. Here, nodes represent traffic segments, and edges represent higher-order temporal and spatial dependencies.
[0038] Construct the dynamic spatio-temporal hypergraph learning module DSTHL, which includes the following steps:
[0039] Obtain the dynamic spatio-temporal hypergraph After that, to ensure the stability of subsequent operations, perform normalization on it:
[0040] Among them, is the degree matrix, represents the negative semi-power of the degree matrix, which is a diagonal matrix, and the diagonal elements are the reciprocals of the square roots of the diagonal elements, and its diagonal elements are defined as .
[0041] This normalization can balance the influence of different nodes and ensure the stability of information propagation. Using the normalized hypergraph , two-layer GCN generates node embeddings from the input sequence X t :
[0042] Among them , is the learnable weight matrix, is the Sigmoid function, and ReLU is the intermediate activation function.
[0043] Traffic flow is strongly influenced by periodic patterns. For example, traffic flow is higher during morning and evening rush hours, and traffic patterns on weekends and holidays are different from weekdays. These periodic changes are mainly driven by human activities (such as going to work, school, and leisure). Therefore, capturing these periodic time patterns is crucial for improving prediction accuracy. To model these time features, we use linear projection to map the input sequence into a time embedding space based on hours and days of the week. Specifically, we generate position encodings for each time step in a day and each day of the week: Define the hour encoding and the day-of-week encoding , where represents the total number of hours in a day, represents the total number of days in a week, and N represents the number of road nodes. Convert the position encodings to a shared embedding space through a learnable weight matrix:
[0044]
[0045] Among them, and are learnable projection matrices that map hour and day features into a unified embedding space of dimension D. Subsequently, the resulting embeddings are combined to obtain a composite time embedding :
[0046]
[0047] The resulting combined time embedding E T can be used to capture the periodic changes in traffic flow, thereby enhancing the model's understanding ability of temporal patterns and improving the accuracy of traffic flow prediction.
[0048] Construct a text feature integration module, including the following steps:
[0049] Traffic flow is usually affected by emergency events, such as accidents, construction, etc. These events will significantly change the traffic conditions, resulting in inaccurate predictions relying solely on sensor data. Therefore, we introduce text description information to enhance the model's ability to capture emergency situations. Use a pre-trained language model (here is BERT) to extract text semantic embeddings from the input sequence Xt:
[0050]
[0051] To ensure the consistency between different feature types, we first apply Z-score normalization to the text embeddings. Subsequently, a text convolution operation (TConv) is used to convert the normalized embeddings into the final representation:
[0052]
[0053] where and are the mean and standard deviation respectively, is a small constant to prevent division by zero, and the result , at the same time, apply spatial convolution (SConv) to the input sequence to extract spatial features:
[0054]
[0055] Get , and finally , , and are concatenated along the feature dimension, and through the fusion convolutional layer Generate the final feature representation :
[0056]
[0057] Construct a feature extraction module FEM-LLM based on a large language model, including the following steps:
[0058] The partially frozen Transformer structure is designed to address the long-term prediction challenges in traffic flow prediction. Different from traditional large language model (LLM)-based methods, FEM-LLM focuses on the most critical features by strategically freezing specific components while improving computational efficiency.
[0059] As Figure 1 shown, the model selectively freezes components at different levels. In the initial F layer, both the multi-head self-attention (MHA) mechanism and the feed-forward network (FFN) remain frozen, thus retaining the pre-trained model's ability to capture local spatio-temporal patterns without updating parameters. This freezing strategy not only reduces computational costs but also prevents the loss of key feature abstraction capabilities. In the subsequent U layer, the MHA component is selectively unfrozen while the FFN remains fixed. This design strikes a balance between efficiency and adaptability, enabling FEM-LLM to capture both short-term fluctuations and long-term dependencies in traffic flow data. In addition, through the multi-head self-attention mechanism, FEM-LLM can allocate attention to different spatio-temporal scales, thus ensuring the simultaneous capture of local and global patterns. This module gradually transforms the input features into deep displays , which can be described as:
[0060]
[0061] The Transformer layer of this module incorporates a hierarchical feature extraction mechanism. The output of the layer is iteratively calculated through the following formula:
[0062]
[0063] where and respectively represent the multi-head self-attention mechanism and the feed-forward network of the layer.
[0064] The multi-head self-attention mechanism enables FEM-LLM to simultaneously capture local and global patterns by attending to different parts of the input sequence, and its definition is as follows:
[0065]
[0066] Among them, the calculation method of each attention head is as follows:
[0067]
[0068] Here are the query, key, and value matrices respectively, is the dimension of the key vector, is the learnable parameter matrix, The function ensures that the sum of the attention weights is 1, enabling the model to focus on the most relevant parts of the input sequence.
[0069] Feed-forward network Enhances the model's ability to capture complex patterns by introducing non-linearity, and its definition is as follows:
[0070]
[0071] Among them , , , are learnable parameter matrices, is the hidden dimension of the feed-forward layer.
[0072] To ensure training stability, we adopt layer normalization after both the multi-head self-attention mechanism and the feed-forward network to reduce internal covariate shift, thereby improving the training efficiency of the model. The combination of the above components enables FEM-LLM to extract deeper spatio-temporal and text features, which is crucial for accurate long-term traffic flow prediction.
[0073] The prediction result is obtained through the regression layer, including the following steps:
[0074] The deep features extracted by FEM-LLM Pass through a 1×1 two-dimensional convolutional layer to perform a linear transformation on the feature dimension while maintaining the input structure, obtaining the prediction result :
[0075]
[0076] This operation expands to:
[0077]
[0078] Among them and are the learnable weights and bias parameters of the regression layer.
[0079] The objective loss function of the multimodal traffic flow prediction method based on dynamic spatio-temporal hypergraph and large language model is summarized as follows:
[0080]
[0081] Among them, Y represents the real traffic flow, is the L1 loss between the predicted value and the real value, θ is the weight of the model, is the L2 regularization term, which is used to prevent the model parameters from being too large, and λ controls the strength of regularization. The model is trained by minimizing the above objective function, so as to improve the prediction accuracy and prevent overfitting, and make it have better generalization ability in different traffic scenarios.
[0082] Compared with the prior art, the beneficial effects of the present invention are as follows: In the present invention, the text feature fusion module innovatively mines and utilizes the implicit semantic information in emergency events by introducing a text description module in multi-modal data fusion, so as to more accurately depict the changes in traffic conditions under emergency events. At the same time, the dynamic spatio-temporal hypergraph learning structure module (DSTHL) constructed by the FastDTW algorithm effectively captures the dynamic spatio-temporal correlation and periodic characteristics between roads; in addition, the present invention also introduces a feature extraction module based on large language model (FEM-LLM), which gives full play to the advantages of LLMs in deep feature extraction and semantic understanding, breaks through the bottleneck of long-term traffic prediction, and realizes accurate modeling of complex spatio-temporal patterns and context relationships. A large number of experimental results on the first large-scale public text-traffic dataset - Beijing Text-Traffic Dataset (BjTT) show that the method of the present invention is significantly superior to the prior art in multiple evaluation indexes, demonstrating excellent prediction accuracy and robustness. Brief Description of the Drawings
[0083] Figure 1 It is a framework diagram of the feature extraction module based on large language model (FEM-LLM) in the present invention.
[0084] Figure 2 It is the multimodal traffic flow prediction method based on dynamic spatio-temporal hypergraph and large language model in the present invention. Detailed Embodiments
[0085] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0086] The multimodal traffic flow prediction method based on dynamic spatio-temporal hypergraph and large language model includes the following steps:
[0087] Obtain a traffic dataset, align the timestamps of the text descriptions and sensor data, and divide the dataset into training set, validation set, and test set;
[0088] Construct a dynamic spatio-temporal hypergraph to retain key spatio-temporal dependencies;
[0089] Construct a traffic flow prediction model, train the model using the training set, monitor the training process with the validation set and adjust the hyperparameters, and finally evaluate the prediction performance of the model using the test set;
[0090] Among them, the traffic flow prediction model includes a dynamic spatio-temporal hypergraph learning module DSTHL, a text feature integration module, a feature extraction module based on large language model FEM-LLM, and regression prediction:
[0091] Dynamic spatio-temporal hypergraph learning module (DSTHL): Use the FastDTW algorithm to construct a dynamic hypergraph, and combine graph convolutional network to capture the spatio-temporal dependencies of traffic flow;
[0092] Text feature integration module: Extract the semantic features of traffic event texts, and splice them into a multi-modal representation after aligning with spatio-temporal features;
[0093] Feature extraction module based on large language model (FEM-LLM): Build a model of multi-scale spatio-temporal dependencies by partially freezing the Transformer framework, and combine text semantic features to fuse and output deep spatio-temporal representations;
[0094] Pass the deep spatio-temporal representation output by FEM-LLM through a regression layer to predict future traffic flow, and obtain the final prediction result.
[0095] The construction of the dynamic spatio-temporal hypergraph includes the following steps:
[0096] To effectively capture the inherent spatio-temporal dependencies in traffic flow data, we propose a dynamic spatio-temporal hypergraph G t , which is used to quantify the pairwise similarity between different traffic segments. Since the traffic flow at a specific location is affected by both time patterns and spatial interactions, by representing these relationships as a graph structure, we can efficiently encode these interdependencies, thus providing a solid foundation for subsequent spatio-temporal prediction models. First, apply the FastDTW algorithm to calculate the dynamic time warping (DTW) distance of traffic flow data, and the distance formula is as follows:
[0097]
[0098] Among them, and represent the traffic flow sequences of segments i and j respectively, and represent the traffic flow values of road segments i and j at time step t, respectively. is the standard deviation of the traffic flow sequence of road segment i, and T is the total number of time steps within the time window.
[0099] Construct a symmetric N×N adjacency matrix G using pairwise distances t , where N is the total number of road segments, and each element is defined as follows:
[0100]
[0101] Here is a threshold corresponding to the smallest 5% in the DTW distance to ensure that only the most important relationships are retained. This dimensionality reduction helps reduce noise while retaining key spatio-temporal dependencies. The resulting graph is the dynamic spatio-temporal hypergraph, which intuitively depicts traffic flow dynamics, where nodes represent traffic road segments and edges represent higher-order temporal and spatial dependencies.
[0102] Construct the dynamic spatio-temporal hypergraph learning module DSTHL, including the following steps:
[0103] After obtaining the dynamic spatio-temporal hypergraph , for the sake of ensuring the stability of subsequent operations, perform normalization on it:
[0104]
[0105] where is the degree matrix, represents the negative semi-power of the degree matrix, which is a diagonal matrix, and the diagonal elements are the reciprocals of the square roots of the diagonal elements, and its diagonal elements are defined as .
[0106] This normalization can balance the influence of different nodes and ensure the stability of information propagation. Using the normalized hypergraph , two-layer GCN generates node embeddings from the input sequence X t :
[0107]
[0108] where , is the learnable weight matrix, is the Sigmoid function, and ReLU is the intermediate activation function.
[0109] Traffic flow is strongly influenced by periodic patterns. For example, traffic volume is higher during the morning and evening rush hours, while traffic patterns on weekends and holidays are different from weekdays. These periodic changes are mainly driven by human activities (such as going to work, school, and leisure), so capturing these periodic time patterns is crucial for improving prediction accuracy. To model these temporal features, we use linear projection to map the input sequence into a time embedding space based on hours and days of the week. Specifically, we generate position encodings for each time step of the day and each day of the week: define the hour encoding and the day-of-week encoding , where represents the total number of hours in a day, represents the total number of days in a week. The position encodings are transformed into a shared embedding space through learnable weight matrices:
[0110]
[0111] where, and are learnable projection matrices that map the hour and day features into a unified embedding space of dimension D. Subsequently, the resulting embeddings are merged to obtain a composite time embedding :
[0112]
[0113] The resulting combined time embedding can be used to capture the periodic changes in traffic flow, thereby enhancing the model's understanding of temporal patterns and improving the accuracy of traffic flow prediction.
[0114] Construct a text feature integration module, including the following steps:
[0115] Traffic flow is usually affected by emergency events, such as accidents, construction, etc. These events can significantly change traffic conditions, resulting in inaccurate predictions relying solely on sensor data. Therefore, we introduce text description information to enhance the model's ability to capture emergency situations. Use a pre-trained language model (here it is BERT) to extract text semantic embeddings from the input sequence Xt:
[0116]
[0117] To ensure consistency between different feature types, we first apply Z-score normalization to the text embeddings. Subsequently, a text convolution operation (TConv) is used to transform the normalized embeddings into the final representation:
[0118]
[0119] wherein and are the mean value and the standard deviation respectively, is a small constant to prevent zero, and the result , meanwhile, apply spatial convolution (SConv) to the input sequence to extract spatial features:
[0120]
[0121] obtain , finally, combine , , and along the feature dimension, and generate the final feature representation through the fusion convolutional layer :
[0122]
[0123] Construct a feature extraction module FEM-LLM based on the large language model, including the following steps:
[0124] The partially frozen Transformer structure aims to address the long-term prediction challenges in traffic flow prediction. Different from traditional large language model (LLM)-based methods, FEM-LLM focuses on the most critical features by strategically freezing specific components, while improving computational efficiency.
[0125] As Figure 1 shown, the model selectively freezes components at different levels. In the initial F layer, both the multi-head self-attention (MHA) mechanism and the feed-forward network (FFN) remain frozen, thus retaining the ability of the pre-trained model to capture local spatio-temporal patterns without updating parameters. This freezing strategy not only reduces the computational cost but also prevents the loss of key feature abstraction ability. In the subsequent U layer, the MHA component is selectively unfrozen while the FFN remains fixed. This design strikes a balance between efficiency and adaptability, enabling FEM-LLM to capture both short-term fluctuations and long-term dependencies in traffic flow data. In addition, through the multi-head self-attention mechanism, FEM-LLM can allocate attention to different spatio-temporal scales, thus ensuring the simultaneous capture of local and global patterns. This module gradually transforms the input feature into a deep display , which can be described as:
[0126]
[0127] The Transformer layer of this module incorporates a hierarchical feature extraction mechanism. The output of the th layer is iteratively calculated through the following formula:
[0128]
[0129] where and respectively represent the multi-head self-attention mechanism and the feed-forward network of the th layer.
[0130] The multi-head self-attention mechanism enables FEM-LLM to simultaneously capture local and global patterns by attending to different parts of the input sequence, and is defined as follows:
[0131]
[0132] where the calculation of each attention head is:
[0133]
[0134] Here are the query, key, and value matrices respectively, is the dimension of the key vector, is the learnable parameter, The function ensures that the sum of the attention weights is 1, enabling the model to focus on the most relevant parts of the input sequence.
[0135] The feed-forward network enhances the model's ability to capture complex patterns by introducing non-linearity, and is defined as follows:
[0136]
[0137] where , , , are learnable parameter matrices, is the hidden dimension of the feed-forward layer.
[0138] To ensure training stability, we adopt layer normalization after both the multi-head self-attention mechanism and the feed-forward network to reduce internal covariate shift, thereby improving the training efficiency of the model. The combination of the above components enables FEM-LLM to extract deeper spatio-temporal and text features, which is crucial for accurate long-term traffic flow prediction.
[0139] The prediction results are obtained through the regression layer, including the following steps:
[0140] The deep features extracted by FEM-LLM Pass through a 1×1 two-dimensional convolutional layer to perform a linear transformation on the feature dimension while maintaining the input structure, obtaining the prediction result :
[0141]
[0142] This operation expands to:
[0143]
[0144] where and are the learnable weights and bias parameters of the regression layer.
[0145] The objective loss function of the multi-modal traffic flow prediction method based on dynamic spatio-temporal hypergraph and large language model is summarized as:
[0146]
[0147] where Y represents the real traffic flow, is the L1 loss between the predicted value and the real value, θ is the weight of the model, is the L2 regularization term, which is used to prevent the model parameters from being too large, and λ controls the strength of regularization. The model is trained by minimizing the above objective function, so as to improve the prediction accuracy and prevent overfitting, making it have better generalization ability in different traffic scenarios.
[0148] The present invention has carried out experimental verification on the above method and achieved obvious effects, which are described as follows:
[0149] 1. Dataset
[0150] The Beijing Text-Traffic (BjTT) dataset is one of the largest and most diverse datasets currently used for traffic flow prediction. This dataset combines traffic sensor data with detailed traffic event text descriptions, covering more than 32,000 time series records collected from January to March 2022, involving 1,260 main roads within the Fifth Ring Road of Beijing. Each record contains numerical features (such as average vehicle speed, congestion level, etc.) and context text describing traffic-related events. These events include traffic accidents, road construction, weather anomalies, and social activities. During the data preprocessing process, we aggregated the traffic data to the road segment level and normalized the traffic features. The normalization uses the following formula:
[0151]
[0152] where X SDenote the original data, where μ and σ are the mean and standard deviation of this road section respectively. Subsequently, the dataset is divided into a 70% training set, a 10% validation set, and a 20% test set for model evaluation. The multimodal structure of the BjTT dataset can provide a more comprehensive perspective on urban traffic dynamics, facilitating the analysis of complex factors influencing traffic flow.
[0153] 2. Evaluation Metrics
[0154] To evaluate all methods from multiple perspectives, we selected four common regression evaluation metrics, including Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Mean Absolute Percentage Error (MAPE), and Weighted Absolute Percentage Error (WAPE). For all metrics, a lower score indicates better prediction performance.
[0155] 3. Implementation Details
[0156] We set the learning rate of DSTH-LLM to 0.001, the batch size to 8, the weight decay rate to 0.0001, and used the AdamW optimizer for parameter updates. During training, the selection of these hyperparameters aims to ensure the stable convergence of the model and reduce the risk of overfitting. All experiments were conducted on an NVIDIA RTX 4090 GPU with 24GB of video memory and a 16-core Intel Xeon Gold 6430 CPU.
[0157] 4. Comparison Methods
[0158] DCRNN combines diffusion convolution and RNN, using the graph structure to propagate information to capture the spatio-temporal dependencies of traffic data.
[0159] STGCN uses graph convolution to capture spatial dependencies and time convolution to model dynamic traffic changes.
[0160] GWN combines graph neural networks and dilated convolution to model spatial and temporal dependencies simultaneously, improving prediction accuracy.
[0161] GMAN uses the multi-head self-attention mechanism to model complex spatio-temporal dependencies, effectively focusing on key regions and key time periods.
[0162] DGCRN combines dynamic graph convolution and RNN to capture the time-varying spatio-temporal dependencies in traffic data.
[0163] AGCRN fuses graph convolution, attention mechanism, and RNN to model spatio-temporal dependencies while focusing on the most relevant traffic nodes.
[0164] GATGPT combines graph attention mechanism and Transformer architecture to effectively model local spatial dependencies and long-term time patterns.
[0165] The GCNGPT fuses the graph convolutional network for spatial learning and combines Transformer for temporal modeling to simultaneously handle short-term and long-term spatio-temporal dependencies.
[0166]
[0167] Table 1: Experimental Results
[0168] All data is evaluated. We mark the best and the second-best results in bold and underlined.
[0169] In summary, to achieve multi-modal traffic flow prediction, the present invention innovatively constructs a framework based on dynamic spatio-temporal hypergraph learning (DSTHL), which can accurately capture the complex and multi-scale spatio-temporal dependency relationships between roads. Traditional graph neural networks often struggle to effectively capture long-term dependency relationships when dealing with dynamic traffic data, while the hypergraph structure can model high-order associations, thus significantly enhancing the overall understanding and prediction ability of traffic patterns. On this basis, we introduce text features extracted by large language models (LLMs), making full use of the rich context information contained in the text, such as emergency events like accidents, construction, and weather. These information can effectively make up for the deficiencies of sensor traffic data in reflecting external environmental changes and provide more comprehensive data support for traffic flow prediction. Finally, we integrate the fusion feature extraction module (FEM-LLM) to achieve the deep fusion of structured traffic data and unstructured text information, giving full play to the advantages of multi-modal data.
[0170] In the present invention, we propose a multi-modal traffic flow prediction framework called the multi-modal traffic flow prediction method based on dynamic spatio-temporal hypergraph and large language model (DSTH-LLM), which includes three main modules: The text feature fusion module innovatively mines and utilizes the implicit semantic information in emergency events by introducing a text description module in multi-modal data fusion, thus more accurately depicting the changes in traffic conditions under emergency events. At the same time, the dynamic spatio-temporal hypergraph learning structure module (DSTHL) constructed using the FastDTW algorithm effectively captures the dynamic spatio-temporal correlation and periodic characteristics between roads; in addition, the present invention also introduces a feature extraction module based on large language models (FEM-LLM), giving full play to the advantages of LLMs in deep feature extraction and semantic understanding, breaking through the bottleneck of long-term traffic prediction, and achieving accurate modeling of complex spatio-temporal patterns and context relationships. A large number of experimental results on the first large-scale public text-traffic dataset - Beijing Text-Traffic Dataset (BjTT) show that the method of the present invention is significantly superior to the prior art in multiple evaluation metrics, demonstrating excellent prediction accuracy and robustness.
[0171] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0172] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model, characterized in that: The following steps are involved: Obtain a traffic dataset, align the text description with the timestamp of the sensor data, and divide it into training, validation, and test sets; The DTW distance is calculated using the FastDTW algorithm, and then the key dependencies are filtered using thresholds to construct a dynamic spatiotemporal hypergraph. Build a traffic flow prediction model, use the training set to train the model, use the validation set to monitor the training process and adjust the hyperparameters, and finally use the test set to evaluate the model prediction performance; The traffic flow prediction model includes: a dynamic spatiotemporal hypergraph learning module DSTHL, a text feature integration module, a feature extraction module FEM-LLM based on a large language model, and a regression prediction module; Dynamic spatiotemporal hypergraph learning module DSTHL: First, the dynamic hypergraph structure is normalized to balance the complex spatiotemporal correlation strength between nodes; second, traffic data is processed through a two-layer graph network to generate node embedding representations that integrate dynamic relationships; finally, the hour / week encoding is mapped to a unified embedding space through a learnable projection matrix to achieve dynamic weight allocation; Text feature integration module: The BERT language model is used to extract the semantic features of the text description of the traffic incident, and then the normalized embedding is converted into the final representation using the text convolution operation TConv. The spatial topological features obtained by the spatial convolution SConv are concatenated along the feature dimension, and the multimodal feature joint representation is generated through the fusion convolution FConv. FEM-LLM, a feature extraction module based on a large language model: It uses a Transformer architecture with a layered parameter unfreezing strategy. In the initial F layer, the multi-head self-attention MHA mechanism and the feedforward network FFN are frozen. In the subsequent U layer, the MHA component is dynamically unfrozen to achieve local parameter fine-tuning, while keeping the FFN fixed to maintain the stability of long-term dependency representation. Finally, the multimodal feature representation and multi-scale spatiotemporal features are integrated to output a deep spatiotemporal representation. Regression prediction module: The deep spatiotemporal representation output by the feature extraction module FEM-LLM based on the large language model is decoded through a two-dimensional convolutional layer to obtain the prediction result of future traffic flow. A composite loss function is designed to integrate the L1 loss and the L2 regularization term to minimize the prediction error while controlling the model complexity.
2. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 1 is characterized in that: The construction of the dynamic spatiotemporal hypergraph comprises the following steps: The FastDTW algorithm is used to calculate the dynamic time warping DTW distance between road segment pairs. The distance formula is as follows: Among them, X s,i and X s,j represent the traffic flow sequence of road sections i and j respectively, X s,i,t and X s,j,t Respectively represent the traffic flow values of road sections i and j at time step t, σ i is the standard deviation of the traffic flow series of road segment i, T is the total number of time steps in the time window; Using pairwise distances, construct a symmetric N×N adjacency matrix G t , where N is the total number of road segments, and each element G ij is defined as follows: Among them, τ top5% is a threshold corresponding to the smallest 5% of the DTW distance, and the resulting graph G t It is a dynamic space-time hypergraph.
3. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 2 is characterized in that: Constructing the dynamic spatiotemporal hypergraph learning module DSTHL includes the following steps: Get the dynamic spatiotemporal hypergraph G t Then, normalize it: Among them, D t is the degree matrix, represents the negative half power of the degree matrix, which is a diagonal matrix with diagonal elements D t The reciprocal of the square roots of the diagonal elements; Based on the normalized hypergraph G t , input sequence X t , X t For the representation of N nodes at T time steps, the node embedding E is generated through two layers of GCN N : Where W (1) , W (2) is the learnable weight matrix, σ(·) is the Sigmoid function, and ReLU is the intermediate activation function; Define hour codes based on traffic periodicity characteristics Day of the week code Where T h Indicates the total number of hours in a day, T d Represents the total number of days in a week, N represents the number of road nodes, and the position encoding is converted into a shared embedding space through a learnable weight matrix: HAVE BEEN hour =X hour ·W hour (5) HAVE BEEN day =X day ·W day (6) in, and is a learnable projection matrix that maps the hour and day features into a unified embedding space of dimension D. The resulting embeddings are then merged to obtain a composite time embedding E T ∈R N ×D : AND T =And hour +E day (7)。 4. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 3 is characterized in that: Constructing a text feature integration module includes the following steps: Integrate event text description into the model: Use pre-trained language model f LM (·) From the input sequence X t Extract semantic embedding E W : E W =f LM (X t ) (8) Let T be the number of time steps and D be the embedding dimension. After normalizing the text embedding by Z-Score, the text is convolved through T Conv Generate a canonical representation: Where μ(E W ) and σ(E W ) are the mean and standard deviation respectively, ∈ is a constant to prevent zero division, and the result is E′ W ∈R T×N×D , at the same time, for the input sequence X t ∈R T×N×d Apply spatial convolution SConv to extract spatial features: E R =SConv(X t ) (10) Get E R ∈R T×N×D Finally, E r , E′ W , E N and E R Splicing along the feature dimension, the final feature representation P is generated by fusing the convolutional layer FConv F : P F =FConv(Concat(E T ,AND' W ,AND N ,AND R )) (11) Among them, Concat represents a concatenation operation.
5. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 4 is characterized in that: Constructing a feature extraction module FEM-LLM based on a large language model includes the following steps: The feature extraction module FEM-LLM based on the large language model models the spatiotemporal dependencies in traffic flow prediction by partially freezing the Transformer architecture. F ∈R T×N×D' Gradually transformed into deep display H out ∈R T×N×D" , described as: The feature extraction module FEM-LLM based on the large language model consists of multiple Transformer layers, each of which contains a multi-head self-attention mechanism and a feedforward neural network. The output H of the lth layer l The iterative calculation is as follows: Among them, LayerNorm represents layer normalization, and They represent the multi-head self-attention mechanism and feed-forward network of the lth layer respectively; Multi-head self-attention mechanism The calculation of is as follows: Among them, Concat represents the concatenation operation, W O It is a learnable parameter matrix that concatenates the outputs of multiple attention heads and performs a linear transformation to integrate the information of multiple heads and generate the final multi-head self-attention output. In addition, each attention head i The calculation method is: here and are query, key, and value matrices, respectively, and d k is the dimension of the key vector, and It is a learnable parameter matrix. The Softmax function ensures the normalization of the attention weights, allowing the model to focus on the most relevant parts of the input sequence. Feedforward Network The calculation is as follows: in b2∈R D″ is the learnable parameter matrix, D ff is the dimension of the feed-forward layer, and CELU is the activation function.
6. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 5 is characterized in that: The prediction results are obtained through the regression layer, including the following steps: The deep features H extracted by the feature extraction module FEM-LLM based on the large language model out ∈R T×N×D Through a 1×1 two-dimensional convolution layer Conv2D (1.1) Mapping to prediction results Expands to: where W∈R D″×O is a learnable weight matrix used to map the input feature dimension to the output dimension O, W d,o represents the weight from input channel d to output channel o, b o ∈R o is the bias vector, used to translate the output; The objective loss function of the multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model can be summarized as: Among them, Y represents the actual traffic flow, is the L1 loss between the predicted value and the true value, θ is the weight of the model, is the L2 regularization term, which is used to prevent the model parameters from being too large, and λ controls the strength of regularization.
Citation Information
Patent Citations
Space-time Transform traffic flow prediction method based on dynamic correlation
CN116543554A
Multi-level embedded traffic flow space-time prediction method based on pre-training large language model
CN119252022A