Multi-modal traffic flow prediction method based on dynamic space-time hypergraph and large language model

By adopting a multimodal fusion method of dynamic spatiotemporal hypergraph and large language model in traffic flow prediction, the problem of difficulty in dynamic modeling and long-term prediction in the existing technology is solved, and higher prediction accuracy and robustness are achieved.

CN119942803AActive Publication Date: 2025-05-06BEIJING FORESTRY UNIVERSITY

Patent Information

Application Number
CN202510426857.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing traffic flow prediction methods are difficult to accurately model the dynamic changes in traffic flow, especially in terms of emergency and long-term prediction.

Method used

The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model is adopted. The dynamic spatiotemporal hypergraph learning module captures the space-time dependence relationship of traffic flow, the text feature integration module integrates traffic event text information, and the feature extraction module based on the large language model extracts deep spatiotemporal characteristics, and finally predicts traffic flow through the regression layer.

Benefits of technology

It significantly improves the accuracy and robustness of traffic flow prediction, especially in emergency events and long-term prediction, and can more accurately characterize traffic conditions and complex spatiotemporal patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942803A_ABST
    Figure CN119942803A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal traffic flow prediction method based on a dynamic space-time hypergraph and a large language model, and relates to the field of computer machine learning, and the method comprises the steps: obtaining a traffic data set, aligning the timestamps of text description and sensor data, and dividing a training set, a verification set and a test set; constructing a dynamic space-time hypergraph, and retaining a key space-time dependency relationship; and constructing a traffic flow prediction model, training the model by using the training set, monitoring the training process and adjusting hyper-parameters by using the verification set, and finally evaluating the prediction performance of the model by using the test set. The method is superior to an existing method, and the effectiveness of the method in a complex dynamic traffic environment is proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer machine learning, and in particular to a multimodal traffic flow prediction method based on a dynamic spatiotemporal hypergraph and a large language model. Background Art

[0002] As a core component of the intelligent transportation system, the evolution of traffic flow prediction technology has always been centered around the three goals of improving the ability to express spatiotemporal features, enhancing the fusion effect of multi-source data, and optimizing long-term prediction accuracy. The current status of technology development in each stage is as follows: (1) Application of historical traffic flow data in traffic flow prediction Traffic flow prediction methods have evolved from traditional statistical methods to deep learning and then to methods based on graph neural networks. Early studies mainly used statistical models and machine learning methods, such as the seasonal autoregressive integrated moving average (SARIMA) model proposed by Billy M. Williams and Lester A. Hoel (2003), and the support vector regression (SVR) combined with historical data for prediction by Wu et al. (Chun-Hsin Wu and Jan-Ming Ho and Lee, DT 2004). Although these methods can capture the periodicity of time series, they are difficult to characterize complex nonlinear spatiotemporal relationships.

[0003] With the development of deep learning, researchers have begun to use recurrent neural networks (RNN), long short-term memory networks (LSTM) (Ma et al. 2015) and gated recurrent units (GRU) (Junyoung Chung et al. 2014) for traffic flow prediction. Zhao et al. (Zhao, Jun and Zhu, Wen-Xing 2021) enhanced the ability to extract spatiotemporal features through deep convolutional GRU (DCGN), and Guo et al. (Guo, Shengnan et al. 2019) used ST-3DNet combined with 3D convolution to simultaneously model spatial and temporal features. However, these methods have limitations in dealing with irregular traffic network structures and it is difficult to fully characterize complex traffic flow changes.

[0004] To solve this problem, researchers introduced graph neural networks (GNNs) and combined them with graph convolutional networks (GCNs) for traffic flow modeling. STGCN (Bing Yu and Haoteng Yin and Zhanxing Zhu 2018) uses graph convolution to extract spatial features and combines it with temporal convolution for temporal modeling. MTGNN (Wu, Zonghan et al. 2020) improves the adaptability of the model to dynamic traffic relations by learning implicit graph structures. STSGCN (Song, Chao et al. 2020) further adopts a modular design to capture multi-scale spatiotemporal features. However, existing methods generally rely on static graph structures and it is difficult to accurately model the dynamic changes of traffic flow. Therefore, it is still necessary to further explore adaptive dynamic modeling methods.

[0005] (2) Application of text information in traffic flow prediction Text data such as social media can provide real-time traffic status information and become an important data source for auxiliary traffic flow prediction. Researchers have explored ways to enhance prediction capabilities using text data such as social media and news reports, and have achieved certain results.

[0006] Abidin et al. (Abidin, Ahmad Faisal et al. 2014) proposed combining credible information from social networks to improve the accuracy of public transportation forecasts. Ni et al. (Ni, Ming and He, Qing and Gao, Jing 2017) combined social media data with historical passenger flow data and used support vector regression to improve the accuracy of subway passenger flow forecasts under emergency events. Aniekan et al. (Essien, Aniekan et al. 2021) integrated real-time traffic, weather and social media tweets (such as Twitter) and built an urban traffic forecasting model based on a bidirectional LSTM autoencoder. Yao et al. (Yao, Weiran and Qian, Sean 2021) analyzed Twitter information to infer residents' daily routines and studied their impact on traffic conditions. Maryam et al. (Shoaeinaeini, Maryam and Ozturk, Oktay and Gupta, Deepak 2022) designed an urban traffic forecasting model that combined social network data and verified its effectiveness based on the California PeMS dataset.

[0007] Although text data has shown great potential in assisting traffic flow prediction, most studies mainly extract statistical information from social media, while ignoring the rich semantic content in text. Therefore, how to effectively integrate text data with traffic data, especially using the semantic information of text to improve prediction accuracy, remains a key challenge.

[0008] (3) Application of large language models in traffic flow prediction The development of large language models (LLMs) has brought new opportunities and challenges to traffic flow prediction. Researchers have explored how to use these models to process complex spatiotemporal data and provide new perspectives on traffic data analysis through natural language understanding and generation.

[0009] UrbanGPT (Li, Zhonghang et al., 2024) models the complex spatiotemporal dependencies of urban traffic systems. GGT (Wang, Xuhong et al., 2023) combines graph neural networks and Transformer to accurately capture the dynamic changes of traffic flow. Traffic Transformer (Jin, KyoHoon et al., 2021) uses a self-attention mechanism to effectively model complex spatiotemporal patterns. TrafficGPT (Siyao Zhang et al., 2023) processes multimodal data under the GPT framework and is suitable for a variety of traffic tasks. TrafficBERT (Jin, KyoHoon et al., 2021) combines natural language processing technology to analyze traffic text data, realizing the combination of NLP and traffic analysis. ST-LLM (Liu, Chenxi et al., 2024) models deep spatiotemporal dependencies through a large language model, while STG-LLM (Lei Liu et al., 2024) introduces graph convolution to process the spatial structure of the traffic network. In addition, PBTR (Duan, Wenying et al., 2019) adopts a bidirectional encoder to capture long-term dependency features and is suitable for traffic datasets with long-term dependency characteristics.

[0010] Although large language models have shown great potential in traffic flow prediction, they still face challenges in semantic alignment of multi-source data, especially the semantic information matching problem when fusing text and traffic data. In addition, LLMs still have limitations in deep temporal pattern mining and contextual information extraction, which limits their application effect in traffic flow prediction.

[0011] Despite all the progress, traffic forecasting still faces the following key challenges: (1) Unpredictability of emergencies: Emergency events such as traffic accidents, concerts, and sports events often disrupt existing traffic patterns and cause dramatic changes in traffic conditions. Traditional single data sources (such as traffic flow data) are difficult to accurately reflect the impact of these emergencies on traffic, resulting in reduced prediction accuracy.

[0012] (2) Limitations of spatial correlation: Many traffic prediction methods rely on the road network structure and the spatial correlation between adjacent road segments. However, in some cases (e.g., residential and commercial areas have similar traffic patterns during peak hours in the morning and evening), even if there is no direct spatial correlation between road segments, their traffic patterns may still be highly similar. Existing methods have difficulty effectively modeling these non-spatially correlated scenarios, which limits the prediction accuracy.

[0013] (3) Difficulty in long-term prediction: Existing time series models are difficult to accurately capture the long-term dependence of traffic flow, making long-term traffic prediction a major challenge. In addition, the influence of external factors such as weather changes, special events, and road construction makes long-term prediction even more difficult.

[0014] With the increasing popularity of intelligent transportation systems, massive amounts of traffic flow data and event text data provide new opportunities to solve the above problems. Therefore, how to effectively integrate traffic flow data with text descriptions of traffic events to improve the expression of spatiotemporal features and long-term prediction capabilities has become a key issue that needs to be solved urgently. Summary of the invention

[0015] The purpose of the present invention is to provide a multimodal traffic flow prediction method based on a dynamic spatiotemporal hypergraph and a large language model to solve the problems raised in the above background technology.

[0016] To achieve the above object, the present invention provides the following technical solutions: The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model includes the following steps: Obtain a traffic dataset, align the text description with the timestamp of the sensor data, and divide it into training, validation, and test sets; Construct a dynamic spatiotemporal hypergraph that preserves key spatiotemporal dependencies; Build a traffic flow prediction model, use the training set to train the model, use the validation set to monitor the training process and adjust the hyperparameters, and finally use the test set to evaluate the model prediction performance; The traffic flow prediction model includes a dynamic spatiotemporal hypergraph learning module DSTHL, a text feature integration module, a feature extraction module FEM-LLM based on a large language model, and regression prediction: Dynamic Spatiotemporal Hypergraph Learning Module (DSTHL): uses the FastDTW algorithm to build a dynamic hypergraph and combines it with a graph convolutional network to capture the spatiotemporal dependencies of traffic flow; Text feature integration module: extracts the semantic features of traffic event text, aligns them with spatiotemporal features, and then splices them into a multimodal representation; Feature Extraction Module based on Large Language Model (FEM-LLM): Modeling multi-scale spatiotemporal dependencies by partially freezing the Transformer architecture, and combining text semantic features to output deep spatiotemporal representations; The deep spatiotemporal representation output by FEM-LLM is used to predict future traffic flow through the regression layer to obtain the final prediction result.

[0017] The construction of the dynamic spatiotemporal hypergraph comprises the following steps: In order to effectively capture the inherent spatiotemporal dependencies in traffic flow data, we propose a dynamic spatiotemporal hypergraph , which is used to quantify the pairwise similarity between different traffic segments. Because the traffic flow at a specific location is affected by both temporal patterns and spatial interactions, by representing these relationships as a graph structure, we can efficiently encode these interdependencies, thereby providing a solid foundation for subsequent spatiotemporal prediction models. First, the FastDTW algorithm is applied to calculate the dynamic time warping (DTW) distance of traffic flow data. The distance formula is as follows: in, and represent the traffic flow sequences of road sections i and j respectively, and Respectively represent the traffic flow values ​​of road sections i and j at time step t, is the standard deviation of the traffic flow sequence of road segment i, and T is the total number of time steps in the time window.

[0018] Using these pairwise distances, construct a symmetric N×N adjacency matrix G t , where N is the total number of road segments. Each element of the matrix is defined as follows: here, is a threshold corresponding to the smallest 5% of the DTW distance to ensure that only the most important relationships are retained. This dimensionality reduction helps reduce noise while retaining key spatiotemporal dependencies. The resulting graph It is a dynamic space-time hypergraph that intuitively depicts the traffic flow dynamics, where nodes represent traffic sections and edges represent high-order temporal and spatial dependencies.

[0019] Constructing the dynamic spatiotemporal hypergraph learning module DSTHL includes the following steps: Get dynamic spatiotemporal hypergraph Finally, in order to ensure the stability of subsequent operations, it is normalized: in, is the degree matrix, represents the negative half power of the degree matrix, which is a diagonal matrix with diagonal elements The reciprocal of the square root of the diagonal elements, whose diagonal elements are defined as .

[0020] This normalization can balance the influence of different nodes and ensure the stability of information dissemination. , the two-layer GCN starts from the input sequence X t Generate node embeddings in: in , is the learnable weight matrix, is the Sigmoid function and ReLU is the intermediate activation function.

[0021] Traffic flow is strongly affected by periodic patterns. For example, traffic volume is higher during rush hour in the morning and evening, while traffic patterns on weekends and holidays are different from weekdays. These periodic changes are mainly driven by human activities (such as going to work, school, and leisure), so capturing these periodic temporal patterns is crucial to improve prediction accuracy. To model these temporal features, we use linear projection to map the input sequence into hour- and week-based temporal embedding spaces. Specifically, we generate position encodings for each time step of the day and each day of the week: Define the hour encoding Day of the week code ,in Indicates the total number of hours in a day. Represents the total number of days in a week, and N represents the number of road nodes. The position encoding is converted into a shared embedding space through a learnable weight matrix: in, and are learnable projection matrices that map hour and day features into a unified embedding space of dimension D. The resulting embeddings are then merged to obtain a composite time embedding : The resulting combined temporal embedding E T It can be used to capture the periodic changes in traffic flow, thereby improving the model's ability to understand time series patterns and improving the accuracy of traffic flow predictions.

[0022] Constructing a text feature integration module includes the following steps: Traffic flow is often affected by emergencies, such as accidents and construction, which can significantly change traffic conditions, making it impossible to accurately predict based on sensor data alone. Therefore, we introduce text description information to enhance the model's ability to capture emergencies. Using a pre-trained language model (Here is BERT) Extract text semantic embedding from the input sequence Xt: To ensure consistency between different feature types, we first apply Z-score normalization to the text embeddings. Subsequently, a text convolution operation (TConv) is used to convert the normalized embeddings into the final representation: in and are the mean and standard deviation, To prevent division by zero, the result is , and at the same time, for the input sequence Apply spatial convolution (SConv) to extract spatial features: get , and finally , , as well as Concatenate along the feature dimension by fusing convolutional layers Generate the final feature representation : Constructing a feature extraction module FEM-LLM based on a large language model includes the following steps: The partially frozen Transformer structure is designed to address the long-term prediction challenges in traffic flow prediction. Different from traditional large language model (LLM) based methods, FEM-LLM focuses on the most critical features by strategically freezing specific components while improving computational efficiency.

[0023] like Figure 1As shown in the figure, the model selectively freezes components at different levels. In the initial F layer, the multi-head self-attention (MHA) mechanism and the feedforward network (FFN) are both kept frozen, thereby retaining the pre-trained model's ability to capture local spatiotemporal patterns without updating parameters. This freezing strategy not only reduces the computational cost, but also prevents the loss of key feature abstraction capabilities. In the subsequent U layer, the MHA component is selectively unfrozen, while the FFN remains fixed. This design strikes a balance between efficiency and adaptability, enabling FEM-LLM to simultaneously capture short-term fluctuations and long-term dependencies in traffic flow data. In addition, through the multi-head self-attention mechanism, FEM-LLM is able to allocate attention to different spatiotemporal scales, thereby ensuring that local and global patterns are captured simultaneously. This module takes the input features Gradually transforming into deep display , which can be described as: The Transformer layer of this module incorporates a hierarchical feature extraction mechanism. Output of the layer Iterative calculation is performed using the following formula: in, and Respectively represent Layer-by-layer multi-head self-attention mechanism and feed-forward network.

[0024] Multi-head self-attention mechanism enables FEM-LLM to capture both local and global patterns by focusing on different parts of the input sequence, which is defined as follows: Among them, the calculation method of each attention head is: here are query, key, and value matrices respectively, is the dimension of the key vector, is the learnable parameter matrix, The function ensures that the attention weights sum to 1, allowing the model to focus on the most relevant parts of the input sequence.

[0025] Feedforward Network The ability of the model to capture complex patterns is enhanced by introducing nonlinearity, which is defined as follows: in , , , is the learnable parameter matrix, is the hidden dimension of the feed-forward layer.

[0026] To ensure training stability, we use layer normalization after the multi-head self-attention mechanism and the feedforward network to reduce internal covariate shift, thereby improving the training efficiency of the model. The combination of the above components enables FEM-LLM to extract deeper spatiotemporal and textual features, which is crucial for accurate long-term traffic flow prediction.

[0027] The prediction results are obtained through the regression layer, including the following steps: The deep features extracted by FEM-LLM Through a 1×1 two-dimensional convolution layer, the feature dimension is linearly transformed while maintaining the input structure to obtain the prediction result. : This operation expands to: in and are the learnable weights and bias parameters of the regression layer.

[0028] The objective loss function of the multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model can be summarized as: Among them, Y represents the actual traffic flow, is the L1 loss between the predicted value and the true value, θ is the weight of the model, is the L2 regularization term, which is used to prevent the model parameters from being too large, and λ controls the strength of regularization. The model is trained by minimizing the above objective function, thereby improving the prediction accuracy and preventing overfitting, so that it has better generalization ability in different traffic scenarios.

[0029] Compared with the prior art, the beneficial effects of the present invention are as follows: the text feature fusion module in the present invention innovatively mines and utilizes the implicit semantic information in emergency events by focusing on the introduction of the text description module in multimodal data fusion, thereby more accurately describing the changes in traffic conditions under emergency events. At the same time, the dynamic spatiotemporal hypergraph learning structure module (DSTHL) constructed using the FastDTW algorithm effectively captures the dynamic spatiotemporal correlation and periodic characteristics between roads; in addition, the present invention also introduces a feature extraction module based on a large language model (FEM-LLM), giving full play to the advantages of LLMs in deep feature extraction and semantic understanding, breaking through the bottleneck of long-term traffic prediction, and realizing the accurate modeling of complex spatiotemporal patterns and contextual relationships. A large number of experimental results on the first large-scale public text-traffic dataset, the Beijing Text-Traffic Dataset (BjTT), show that the method of the present invention is significantly superior to the prior art in multiple evaluation indicators, showing excellent prediction accuracy and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a framework diagram of the feature extraction module (FEM-LLM) based on the large language model in the present invention.

[0031] Figure 2 This is a multimodal traffic flow prediction method based on a dynamic spatiotemporal hypergraph and a large language model in the present invention. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model includes the following steps: Obtain a traffic dataset, align the text description with the timestamp of the sensor data, and divide it into training, validation, and test sets; Construct a dynamic spatiotemporal hypergraph that preserves key spatiotemporal dependencies; Build a traffic flow prediction model, use the training set to train the model, use the validation set to monitor the training process and adjust the hyperparameters, and finally use the test set to evaluate the model prediction performance; The traffic flow prediction model includes a dynamic spatiotemporal hypergraph learning module DSTHL, a text feature integration module, a feature extraction module FEM-LLM based on a large language model, and regression prediction: Dynamic Spatiotemporal Hypergraph Learning Module (DSTHL): uses the FastDTW algorithm to build a dynamic hypergraph and combines it with a graph convolutional network to capture the spatiotemporal dependencies of traffic flow; Text feature integration module: extracts the semantic features of traffic event text, aligns them with spatiotemporal features, and then splices them into a multimodal representation; Feature Extraction Module based on Large Language Model (FEM-LLM): Modeling multi-scale spatiotemporal dependencies by partially freezing the Transformer architecture, and combining text semantic features to output deep spatiotemporal representations; The deep spatiotemporal representation output by FEM-LLM is used to predict future traffic flow through the regression layer to obtain the final prediction result.

[0034] The construction of the dynamic spatiotemporal hypergraph comprises the following steps: In order to effectively capture the inherent spatiotemporal dependencies in traffic flow data, we propose a dynamic spatiotemporal hypergraph G t , which is used to quantify the pairwise similarity between different traffic segments. Because the traffic flow at a specific location is affected by both temporal patterns and spatial interactions, by representing these relationships as a graph structure, we can efficiently encode these interdependencies, thereby providing a solid foundation for subsequent spatiotemporal prediction models. First, the FastDTW algorithm is applied to calculate the dynamic time warping (DTW) distance of traffic flow data. The distance formula is as follows: in, and represent the traffic flow sequences of road sections i and j respectively, and Respectively represent the traffic flow values ​​of road sections i and j at time step t, is the standard deviation of the traffic flow sequence of road segment i, and T is the total number of time steps in the time window.

[0035] Using pairwise distances, construct a symmetric N×N adjacency matrix G t , where N is the total number of road segments, and each element is defined as follows: here is a threshold corresponding to the smallest 5% of the DTW distance to ensure that only the most important relationships are retained. This dimensionality reduction helps reduce noise while retaining key spatiotemporal dependencies. The resulting graph It is a dynamic space-time hypergraph that intuitively depicts the traffic flow dynamics, where nodes represent traffic sections and edges represent high-order temporal and spatial dependencies.

[0036] Constructing the dynamic spatiotemporal hypergraph learning module DSTHL includes the following steps: Get dynamic spatiotemporal hypergraph Finally, in order to ensure the stability of subsequent operations, it is normalized: in, is the degree matrix, represents the negative half power of the degree matrix, which is a diagonal matrix with diagonal elements The reciprocal of the square root of the diagonal elements, whose diagonal elements are defined as .

[0037] This normalization can balance the influence of different nodes and ensure the stability of information dissemination. , the two-layer GCN starts from the input sequence X t Generate node embeddings in: in , is the learnable weight matrix, is the Sigmoid function and ReLU is the intermediate activation function.

[0038] Traffic flow is strongly affected by periodic patterns. For example, traffic volume is higher during rush hour in the morning and evening, while traffic patterns on weekends and holidays are different from weekdays. These periodic changes are mainly driven by human activities (such as going to work, school, and leisure), so capturing these periodic temporal patterns is crucial to improve prediction accuracy. To model these temporal features, we use linear projection to map the input sequence into hour- and week-based temporal embedding spaces. Specifically, we generate position encodings for each time step of the day and each day of the week: Define the hour encoding Day of the week code ,in Indicates the total number of hours in a day. Represents the total number of days in a week. The positional encoding is transformed into a shared embedding space via a learnable weight matrix: in, and are learnable projection matrices that map hour and day features into a unified embedding space of dimension D. The resulting embeddings are then merged to obtain a composite time embedding : The resulting combined temporal embedding It can be used to capture the periodic changes in traffic flow, thereby improving the model's ability to understand time series patterns and improving the accuracy of traffic flow predictions.

[0039] Constructing a text feature integration module includes the following steps: Traffic flow is often affected by emergencies, such as accidents and construction, which can significantly change traffic conditions, making it impossible to accurately predict based on sensor data alone. Therefore, we introduce text description information to enhance the model's ability to capture emergencies. Using a pre-trained language model (Here is BERT) Extract text semantic embedding from the input sequence Xt: To ensure consistency between different feature types, we first apply Z-score normalization to the text embeddings. Subsequently, a text convolution operation (TConv) is used to convert the normalized embeddings into the final representation: in and are the mean and standard deviation, To prevent division by zero, the result is , and at the same time, for the input sequence Apply spatial convolution (SConv) to extract spatial features: get , and finally , , as well as Concatenate along the feature dimension by fusing convolutional layers Generate the final feature representation : Constructing a feature extraction module FEM-LLM based on a large language model includes the following steps: The partially frozen Transformer structure is designed to address the long-term prediction challenges in traffic flow prediction. Different from traditional large language model (LLM) based methods, FEM-LLM focuses on the most critical features by strategically freezing specific components while improving computational efficiency.

[0040] like Figure 1As shown in the figure, the model selectively freezes components at different levels. In the initial F layer, the multi-head self-attention (MHA) mechanism and the feedforward network (FFN) are both kept frozen, thereby retaining the pre-trained model's ability to capture local spatiotemporal patterns without updating parameters. This freezing strategy not only reduces the computational cost, but also prevents the loss of key feature abstraction capabilities. In the subsequent U layer, the MHA component is selectively unfrozen, while the FFN remains fixed. This design strikes a balance between efficiency and adaptability, enabling FEM-LLM to simultaneously capture short-term fluctuations and long-term dependencies in traffic flow data. In addition, through the multi-head self-attention mechanism, FEM-LLM is able to allocate attention to different spatiotemporal scales, thereby ensuring that local and global patterns are captured simultaneously. This module takes the input features Gradually transforming into deep display , which can be described as: The Transformer layer of this module incorporates a hierarchical feature extraction mechanism. Output of the layer Iterative calculation is performed using the following formula: in, and Respectively represent Layer-by-layer multi-head self-attention mechanism and feed-forward network.

[0041] Multi-head self-attention mechanism enables FEM-LLM to capture both local and global patterns by focusing on different parts of the input sequence, which is defined as follows: Among them, the calculation method of each attention head is: here are query, key, and value matrices respectively, is the dimension of the key vector, is a learnable parameter, The function ensures that the attention weights sum to 1, allowing the model to focus on the most relevant parts of the input sequence.

[0042] Feedforward Network The ability of the model to capture complex patterns is enhanced by introducing nonlinearity, which is defined as follows: in , , , is the learnable parameter matrix, is the hidden dimension of the feed-forward layer.

[0043] To ensure training stability, we use layer normalization after the multi-head self-attention mechanism and the feedforward network to reduce internal covariate shift, thereby improving the training efficiency of the model. The combination of the above components enables FEM-LLM to extract deeper spatiotemporal and textual features, which is crucial for accurate long-term traffic flow prediction.

[0044] The prediction results are obtained through the regression layer, including the following steps: The deep features extracted by FEM-LLM Through a 1×1 two-dimensional convolution layer, the feature dimension is linearly transformed while maintaining the input structure to obtain the prediction result. : This operation expands to: in and are the learnable weights and bias parameters of the regression layer.

[0045] The objective loss function of the multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model can be summarized as: Among them, Y represents the actual traffic flow, is the L1 loss between the predicted value and the true value, θ is the weight of the model, is the L2 regularization term, which is used to prevent the model parameters from being too large, and λ controls the strength of regularization. The model is trained by minimizing the above objective function, thereby improving the prediction accuracy and preventing overfitting, so that it has better generalization ability in different traffic scenarios.

[0046] The present invention has experimentally verified the above method and achieved obvious results, as described below: 1. Dataset The Beijing Text-Traffic (BjTT) dataset is one of the largest and most diverse datasets currently used for traffic flow prediction. This dataset combines traffic sensor data with detailed text descriptions of traffic events, covering more than 32,000 time series records collected from January to March 2022, involving 1,260 major roads within the Fifth Ring Road of Beijing. Each record contains numerical features (such as average vehicle speed, congestion level, etc.), as well as contextual text describing traffic-related events. These events include traffic accidents, road construction, weather anomalies, and social activities. During data preprocessing, we aggregate traffic data to the road segment level and normalize the traffic features. The normalization uses the following formula: Among them, X S represents the original data, μ and σ are the mean and standard deviation of the road section, respectively. Subsequently, the dataset is divided into 70% training set, 10% validation set and 20% test set for model evaluation. The multimodal structure of the BjTT dataset can provide a more comprehensive perspective on urban traffic dynamics and help analyze the complex factors affecting traffic flow.

[0047] 2. Evaluation indicators In order to evaluate all methods from multiple perspectives, we selected four common regression evaluation metrics, including mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and weighted absolute percentage error (WAPE). For all metrics, lower scores indicate better prediction performance.

[0048] 3. Implementation details We set the learning rate of DSTH-LLM to 0.001, the batch size to 8, the weight decay rate to 0.0001, and use the AdamW optimizer for parameter updates. During training, the selection of these hyperparameters is intended to ensure stable convergence of the model and reduce the risk of overfitting. All experiments were performed on an NVIDIA RTX 4090 GPU with 24GB of video memory and a 16-core Intel Xeon Gold 6430 CPU.

[0049] 4. Comparison method DCRNN combines diffuse convolution and RNN and uses graph structure to propagate information to capture the spatiotemporal dependencies of traffic data.

[0050] STGCN uses graph convolution to capture spatial dependencies and temporal convolution to model dynamic traffic changes.

[0051] GWN combines graph neural networks and dilated convolutions to simultaneously model spatial and temporal dependencies and improve prediction accuracy.

[0052] GMAN uses a multi-head self-attention mechanism to model complex spatiotemporal dependencies, and can effectively focus on key areas and key time periods.

[0053] DGCRN combines dynamic graph convolution and RNN to capture the time-varying spatiotemporal dependencies in traffic data.

[0054] AGCRN fuses graph convolution, attention mechanism, and RNN to model spatiotemporal dependencies while focusing on the most relevant traffic nodes.

[0055] GATGPT combines the graph attention mechanism and the Transformer architecture to effectively model local spatial dependencies and long-term temporal patterns.

[0056] GCNGPT integrates graph convolutional networks for spatial learning and Transformer for temporal modeling to simultaneously handle short-term and long-term spatiotemporal dependencies.

[0057] Table 1: Experimental results All data were evaluated. We marked the best and second best results in bold and underlined.

[0058] In summary, in order to achieve multimodal traffic flow prediction, this paper innovatively constructs a framework based on dynamic spatiotemporal hypergraph learning (DSTHL), which can accurately capture the complex and multi-scale spatiotemporal dependencies between roads. Traditional graph neural networks often find it difficult to effectively capture long-term dependencies when processing dynamic traffic data, while hypergraph structures can model high-order associations, thereby significantly improving the overall understanding and prediction capabilities of traffic patterns. On this basis, we introduce text features extracted by the large language model (LLM) to make full use of the rich contextual information contained in the text, such as accidents, construction, weather and other emergency events. This information can effectively make up for the shortcomings of sensor traffic data in reflecting changes in the external environment and provide more comprehensive data support for traffic flow prediction. Finally, we incorporate the fusion feature extraction module (FEM-LLM) to achieve a deep fusion of structured traffic data and unstructured text information, giving full play to the advantages of multimodal data.

[0059] In this invention, we propose a multimodal traffic flow prediction framework called the multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model (DSTH-LLM), which includes three main modules: the text feature fusion module innovatively mines and utilizes the implicit semantic information in emergency events by focusing on the introduction of text description modules in multimodal data fusion, so as to more accurately describe the changes in traffic conditions under emergency events. At the same time, the dynamic spatiotemporal hypergraph learning structure module (DSTHL) constructed by the FastDTW algorithm effectively captures the dynamic spatiotemporal correlation and periodic characteristics between roads; in addition, the invention also introduces a feature extraction module based on a large language model (FEM-LLM), which gives full play to the advantages of LLMs in deep feature extraction and semantic understanding, breaks through the bottleneck of long-term traffic prediction, and realizes the accurate modeling of complex spatiotemporal patterns and contextual relationships. A large number of experimental results on the first large-scale public text-traffic dataset, the Beijing Text-Traffic Dataset (BjTT), show that the method of the present invention is significantly superior to the existing technology in multiple evaluation indicators, showing excellent prediction accuracy and robustness.

[0060] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0061] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model, characterized in that: The following steps are involved: Obtain a traffic dataset, align the text description with the timestamp of the sensor data, and divide it into training, validation, and test sets; Construct a dynamic spatiotemporal hypergraph that preserves key spatiotemporal dependencies; Build a traffic flow prediction model, use the training set to train the model, use the validation set to monitor the training process and adjust the hyperparameters, and finally use the test set to evaluate the model prediction performance; The traffic flow prediction model includes: a dynamic spatiotemporal hypergraph learning module DSTHL, a text feature integration module, a feature extraction module FEM-LLM based on a large language model, and a regression prediction module; Dynamic spatiotemporal hypergraph learning module DSTHL: uses the FastDTW algorithm to build a dynamic hypergraph and combines it with a graph convolutional network to capture the spatiotemporal dependencies of traffic flow; Text feature integration module: extracts the semantic features of traffic event text, aligns them with spatiotemporal features, and then splices them into a multimodal representation; FEM-LLM, a feature extraction module based on a large language model: It models multi-scale spatiotemporal dependencies by partially freezing the Transformer architecture and outputs deep spatiotemporal representations by fusing text semantic features; Regression prediction module: The deep spatiotemporal representation output by the feature extraction module FEM-LLM based on the large language model is used to predict future traffic flow through the regression layer to obtain the final prediction result.

2. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 1 is characterized in that: The construction of the dynamic spatiotemporal hypergraph comprises the following steps: The FastDTW algorithm is used to calculate the dynamic time warping DTW distance between road segment pairs. The distance formula is as follows: (1) in, and represent the traffic flow sequences of road sections i and j respectively, and Respectively represent the traffic flow values ​​of road sections i and j at time step t, is the standard deviation of the traffic flow series of road segment i, T is the total number of time steps in the time window; Using pairwise distances, construct a symmetric N×N adjacency matrix G t , where N is the total number of road segments, and each element is defined as follows: (2) in, is a threshold corresponding to the smallest 5% of the DTW distance, and the resulting graph It is a dynamic space-time hypergraph.

3. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 2 is characterized in that: Constructing the dynamic spatiotemporal hypergraph learning module DSTHL includes the following steps: Get dynamic spatiotemporal hypergraph Then, normalize it: in, is the degree matrix, represents the negative half power of the degree matrix, which is a diagonal matrix with diagonal elements The reciprocal of the square roots of the diagonal elements; Based on normalized hypergraph , input sequence X t , X t For the representation of N nodes at T time steps, the node embedding is generated through two layers of GCN : in , is the learnable weight matrix, is the Sigmoid function, and ReLU is the intermediate activation function; Define hour codes based on traffic periodicity characteristics Day of the week code ,in Represents the total number of hours in a day, Represents the total number of days in a week, N represents the number of road nodes, and the position encoding is converted into a shared embedding space through a learnable weight matrix: in, and is a learnable projection matrix that maps hour and day features into a unified embedding space of dimension D, and then merges the obtained embeddings to obtain a composite time embedding : (7)。 4. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 3 is characterized in that: Constructing a text feature integration module includes the following steps: Integrate event text descriptions into the model: Use pre-trained language models From the input sequence Extracting semantic embeddings : Let T be the number of time steps and D be the embedding dimension. After normalizing the text embedding by Z-Score, we use text convolution Generate a canonical representation: in and are the mean and standard deviation, To prevent division by zero, the result , and at the same time, for the input sequence Applying spatial convolution Extract spatial features: (10) get , and finally , , as well as Concatenate along the feature dimension by fusing convolutional layers Generate the final feature representation : in, Represents a concatenation operation.

5. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 4 is characterized in that: Constructing a feature extraction module FEM-LLM based on a large language model includes the following steps: The feature extraction module FEM-LLM based on the large language model models the spatiotemporal dependencies in traffic flow prediction by partially freezing the Transformer architecture. Gradually transforming into deep display , described as: (12) The feature extraction module FEM-LLM based on the large language model consists of multiple Transformer layers, each of which contains a multi-head self-attention mechanism and a feedforward neural network. Output of the layer The iterative calculation is as follows: in, Representation layer normalization, and Respectively represent Multi-head self-attention mechanism and feed-forward network of the layer; Multi-head self-attention mechanism The calculation of is as follows: in, Represents a splicing operation, is a learnable parameter matrix that concatenates the outputs of multiple attention heads and performs a linear transformation to integrate the information of multiple heads and generate the final multi-head self-attention output. In addition, each attention head The calculation method is: here , and are query, key, and value matrices respectively, is the dimension of the key vector, is the learnable parameter matrix, The function ensures the normalization of the attention weights so that the model focuses on the most relevant parts of the input sequence; Feedforward Network The calculation is as follows: in , , , is the learnable parameter matrix, is the dimension of the feed-forward layer, is the activation function.

6. The multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model according to claim 5 is characterized in that: The prediction results are obtained through the regression layer, including the following steps: The deep features extracted by the feature extraction module FEM-LLM based on the large language model Through a 1×1 2D convolutional layer Mapping to prediction results : Expands to: in is a learnable weight matrix used to map the input feature dimension to the output dimension O, Indicates input channel To output channel The weight of is the bias vector, used to translate the output; The objective loss function of the multimodal traffic flow prediction method based on dynamic spatiotemporal hypergraph and large language model can be summarized as: in, Represents the actual traffic flow, is the L1 loss between the predicted value and the true value, θ is the weight of the model, is the L2 regularization term, which is used to prevent the model parameters from being too large, and λ controls the strength of regularization.

Citation Information

Patent Citations

  • Traffic flow prediction method based on Transform space-time diagram convolutional network

    CN114330671A

  • Space-time Transform traffic flow prediction method based on dynamic correlation

    CN116543554A

  • Traffic information prediction method and device, equipment and medium

    CN118522157A

  • Traffic flow prediction method based on dynamic multi-graph space-time synchronization network

    CN119007442A

  • Multi-level embedded traffic flow space-time prediction method based on pre-training large language model

    CN119252022A

Cited By

  • Traffic monitoring method and device based on multi-mode and hypergraph structure, and medium

    CN120220422A

  • Time sequence prediction method and system based on dynamic hypergraph and multi-scale coding

    CN120278037A

  • A time series prediction method and system based on dynamic hypergraph and multi-scale coding

    CN120278037B

  • Data processing method, device and equipment for acquiring airport delay information

    CN120338205A

  • Tranform-based hypergraph association privacy protection method

    CN120449212A