A traffic flow prediction method and device based on large language model semantic enhancement
Patent Information
- Application Number
- CN202610627816.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-09
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2046-05-09
AI Technical Summary
当前模型仍面临两大挑战:一是动态交通模式捕获不足,过度依赖静态拓扑图,忽视节点间时序相似性;二是缺失全局空间相关性,局限于局部依赖,无法刻画远距离节点跨区域交互
[0020] By adopting the above scheme, the present invention has the following advantages and beneficial effects: The traffic flow prediction method based on semantic enhancement of large language model in the embodiments of the present invention innovatively introduces large language model (LLM) to transform traffic flow statistical profiles into semantic embeddings, successfully constructing a global semantic association graph, breaking through the limitation of traditional models that only rely on static geographical distance, thereby being able to capture implicit functional associations and global dependencies across regions. At the same time, the model combines a dynamic adaptive fusion gating mechanism, which can adaptively switch between local physical propagation and global semantic collaboration; adopting a parameter sharing bottleneck structure of "dimensionality reduction-convolution-dimensionality increase" significantly reduces model complexity and memory usage. In addition, the robust normalization at the data level, combined with the composite loss function (Huber+L1) that dynamically decays with training rounds, greatly improves the robustness of the model against sudden traffic anomaly noise, thereby comprehensively improving the accuracy and stability of urban road network traffic flow prediction.
Smart Images

Figure CN122153844B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation systems, and in particular to a traffic flow prediction method and apparatus based on semantic enhancement of a large language model. Background Technology
[0002] Urban traffic flow is characterized by complexity, nonlinearity, and strong spatiotemporal dynamic interaction. Traffic flow prediction is a core foundation for building a green, low-carbon, safe, efficient, and intelligent transportation system, providing support for road network planning, dynamic management, and residents' travel, and helping to alleviate traffic congestion. Urban road networks consist of multiple interconnected traffic sensor nodes, and traffic flow states exhibit complex spatiotemporal dependencies.
[0003] Traditional deep learning methods, such as LSTM, are suitable for temporal modeling and can capture short-term dependencies, but they cannot effectively model the spatial relationships of road network graph structures. They also suffer from limitations in capturing long-term dependencies, low computational efficiency, and difficulty in parallelization. With the development of artificial intelligence and graph neural network technology, prediction methods based on spatiotemporal graph convolution, such as STGCN, ASTGCN, and Graph WaveNet, have received widespread attention. These models can integrate spatiotemporal convolution to uniformly model spatiotemporal dependencies. Current models still face two major challenges: first, insufficient capture of dynamic traffic patterns, over-reliance on static topological graphs, and neglect of temporal similarity between nodes; second, lack of global spatial correlation, limited to local dependencies, and unable to characterize cross-regional interactions between distant nodes. While some multi-view models attempt to overcome the limitations of single graphs, they suffer from structural complexity, high training costs, fine-tuning of parameters, and high data requirements. There is an urgent need for models with strong feature extraction and semantic understanding capabilities to achieve implicit modeling of cross-regional dynamic patterns and global relationships. Summary of the Invention
[0004] The technical problem this invention aims to solve is how to construct a model with strong feature extraction and semantic understanding capabilities to achieve implicit modeling of cross-regional dynamic patterns and global correlations, thereby enabling accurate traffic flow prediction. To solve this technical problem, this invention adopts the following technical solution: a traffic flow prediction method based on semantic enhancement of a large language model, comprising the following steps: S1. Acquire monitoring data and physical topology data from urban road network traffic sensors, and perform preprocessing; the preprocessing includes the following steps: S11. Obtain historical feature data from road network traffic sensors, wherein the historical feature data includes time-series traffic flow data and sensor distance matrix data; S12. The historical feature data is preprocessed to extract multidimensional statistical features, and a large language model is introduced to extract the semantic features of the nodes. The physical adjacency matrix and the global semantic association matrix are constructed as parameters of the future time step predicted traffic label tensor. S13. Robust normalization processing is performed on the time-series traffic data in the historical feature data, that is, the time-series traffic feature data of each channel is robustly scaled in the global range using the median and interquartile range to obtain normalized time-series data of a uniform scale. S14. Based on the historical time step T and the future prediction time step T', construct sample pairs using the normalized time series data obtained in step S13; specifically: extract the data segment of the historical time step T from the normalized time series data as the input feature of the model, and set the normalized flow value of the following future prediction time step T' as the corresponding future flow status label. S15. For the normalized flow values with set labels, adopt the equal interval sliding window segmentation strategy to extract multiple equal-length input subsequences and corresponding target subsequences from the normalized flow values with a fixed-length time window, and generate preprocessed monitoring data samples. S2. Input the preprocessed monitoring data samples into a pre-trained spatiotemporal graph convolutional traffic flow prediction model based on large language model semantic enhancement to obtain traffic flow prediction results for future continuous time steps. The spatiotemporal graph convolutional traffic flow prediction model based on large language model semantic enhancement is trained based on the following steps: the monitoring data samples are divided into a training set, a validation set, and a test set according to a predetermined proportion of time steps in chronological order; based on the training set and the validation set, a traffic flow prediction model constructed by multi-view feature extraction, dynamic gating fusion, and enhanced spatiotemporal perception network is trained to establish a nonlinear mapping relationship between historical traffic status and future traffic labels; according to the test set, the final prediction model selected by the validation set after training is tested until the model is qualified and a pre-trained model is obtained, otherwise the model is retrained.
[0005] Preferably, step S11, obtaining historical feature data from road network traffic sensors, specifically involves: obtaining time-series traffic flow data and sensor distance matrices of sensor nodes within a historical time step T; extracting the time-series traffic flow data to construct the original historical traffic flow tensor. Where T is the historical time step, N is the total number of sensor nodes, and C is the dimension of the input features. This represents the road network traffic state matrix at time step t. Represents the set of real numbers. This indicates that the tensor exists in a real space with dimensions of time step × number of nodes × feature dimension.
[0006] Preferably, the construction of the physical adjacency matrix and the global semantic association matrix in step S12 specifically includes: Construction of the physical adjacency matrix: Based on the geographical distance between sensors, a threshold Gaussian kernel function is used to construct the physical adjacency matrix. The weights are calculated as follows: like ,but Otherwise, it is 0, where For sensor nodes With sensor nodes The Euclidean distance between them This represents the element in the i-th row and j-th column of the physical adjacency matrix. Let be the standard deviation of the Gaussian kernel function. This is the distance truncation threshold; Global semantic association matrix construction: computation nodes A multidimensional statistical profile of the sensor, including mean, standard deviation, daily coefficient of variation, information entropy, and the ratio of morning to evening peak hours, is generated and converted into natural language descriptive text. ; Describe the text The input is fed into a large language model, and the mean of its last hidden state is extracted as the semantic embedding vector of the node. ,in For large language models, nonlinear feature mapping function, For mean pooling operation, Let be the dimension of the semantic embedding vector; calculate the cosine similarity between nodes based on the semantic embedding vector, and select the Z most similar nodes to construct a semantic association matrix. : .
[0007] Preferably, in step S13, robust normalization of the feature data of each channel is performed globally using the median and interquartile range. Specifically, this includes: normalizing each feature based on the resampled data using a robust scaling algorithm to eliminate the influence of outliers such as sudden congestion. The normalization formula is as follows: : in, The original historical traffic flow tensor obtained in step S11, The normalized data tensor obtained after robust scaling. The median of the feature data after sorting by numerical value in ascending order, which is the 50th percentile. and These are the 75th percentile and the 25th percentile of the characteristic sequence, respectively.
[0008] Preferably, in step S14, setting corresponding future traffic status labels for the normalized time-series data of the road network sensors specifically includes: for the normalized time-series data obtained in step S13, using continuous time steps of historical time step T as input historical tensors. The subsequent continuous time steps of length T' are used as the real flow label tensors for future time steps. The model learns a nonlinear mapping function. Achieving future time step prediction of traffic label tensors Real traffic label tensor for future time steps The approximation of is expressed mathematically as follows:
[0009] in, Predict traffic label tensors for future time steps. To input the history tensor, It is the physical adjacency matrix. This is the global semantic association matrix. This is the set of parameters that the model needs to learn.
[0010] Preferably, step S15 uses an equally spaced sliding window segmentation strategy to generate samples, specifically including: For the normalized time series in the dataset, a sliding window sampling strategy with historical time step T, future prediction time step T', and sliding interval of step 1 is used to extract multiple fixed-length sequence samples from the original normalized time series. The specific operation for sample construction is as follows: for a normalized data tensor with a total length of L in the time dimension... Traverse normalized data tensors Index segmented by time and In each iteration, a subsequence of the normalized data tensor is extracted. as input history tensor Extract a subsequence from another normalized data tensor. Tensor as the real traffic label of the future time step .
[0011] Preferably, in step S2, the above samples are divided into training set, validation set and test set according to the time steps in chronological order in a ratio of 7:1:2. End-to-end parameter learning is performed using the training set, and optimization is achieved using a dynamic composite loss function. During training, the training process is supervised by validation set metrics to determine early stopping and adaptive adjustment of the learning rate. The optimal model weights are selected based on the validation set. The final prediction model selected is evaluated using mean absolute error, root mean square error, mean absolute percentage error, and coefficient of determination to determine whether the model is qualified.
[0012] Preferably, in step S2, based on the training and validation sets, a traffic flow prediction model constructed from multi-view feature extraction, dynamic gating fusion, and an enhanced spatiotemporal perception network is trained to establish a nonlinear mapping relationship between historical traffic states and future traffic labels. Specifically, the prediction model includes a lightweight spatiotemporal module that performs dual-view feature extraction: graph convolution is performed on the physical graph and semantic graph respectively, using a bottleneck structure of dimensionality reduction-convolution-upgrading, and Chebyshev approximations of different orders share transformation parameters. The formula is:
[0013]
[0014] In the formula, For the first Chebyshev polynomials The highest order of Chebyshev polynomials For the physical adjacency matrix, For the first The node feature matrix of the layer input, For activation function, and These are the dimension reduction and dimension increase projection matrices, respectively. To share the transformation parameter matrix; Dynamic fusion gating: and Input a lightweight gating network and generate dynamic gating coefficients using Softmax. The fusion formula is ;in, For the fused spatiotemporal feature tensor; operators This represents vector concatenation operations; operators Hadamard product, which is the element-wise multiplication of a matrix; dynamic gating coefficient. The specific expression is: In the formula, This is a mapping function for a multilayer perceptron, used to learn the contribution weights of physical and semantic features. It is the bias vector; Enhanced spatiotemporal modeling module: First, it captures short-term trends through local one-dimensional convolution to obtain... The expression is:
[0015] in: It is a local temporal feature representation, that is, the feature after local one-dimensional convolution processing. For layer normalization, For the fused spatiotemporal feature tensor, This involves a local one-dimensional convolution operation; then, a multi-head self-attention mechanism is used to capture long-range time.
[0016] in For the final spatiotemporal feature output, It is a multi-head self-attention mechanism used to capture long-distance temporal dependencies; Finally, the high-dimensional spatiotemporal features extracted by multiple KA-ST blocks are fed into the output layer and mapped to the target dimension prediction result through two consecutive 1x1 convolutional layers. , ,in For hidden layer convolution, To modify the activation function of the linear unit, For output layer convolution; Composite Loss Function: To balance robustness against sudden outliers with prediction accuracy for stable flow rates, a dynamic composite loss function consisting of a weighted average of Huber Loss and L1 Loss is used for parameter optimization. Dynamic decay with training rounds: In the formula: This represents the total loss value during model training. For the real traffic label tensor of future time steps; Predict traffic label tensors for future time steps; The Huber loss function is used to reduce the drastic impact of outliers on the gradient in the early stages of training, thereby enhancing the robustness of the model. The mean absolute error loss function is used to improve the regression prediction accuracy of the model in stable flow ranges. The weighting coefficients are dynamically changed with each training round, and their specific meaning is: assigned weights in the initial training phase. The highest weights are used to quickly stabilize the model, and as the number of training epochs (e) increases, Gradual decay causes the model's center of gravity to shift. To achieve higher prediction accuracy.
[0017] Preferably, the specific calculation rules for the evaluation indicators are as follows: The test set is input into the trained prediction model, and the high-dimensional feature output is denormalized to restore the original traffic flow to its true scale. Let Y denote the denormalized overall true traffic label tensor and denot the overall predicted traffic label tensor as Y. Let the total number of data points in the test set participating in the evaluation be . Where m is the index of the data point, calculate the following metrics: Mean Absolute Error formula: Root mean square error formula: Mean absolute percentage error formula: Coefficient of determination formula:
[0018] In the formula, This represents the actual traffic flow value of the m-th data point in the overall actual traffic flow label tensor Y. For the overall predicted traffic label tensor The corresponding m-th predicted traffic flow value, The mean of all real traffic flows in the test set; for the calculation of MAPE, M' is the set of valid data points where the actual flow rate is greater than a set minimum threshold. The total number of data points in the data is used to avoid division by zero anomalies when traffic flow is zero.
[0019] A traffic flow prediction device based on semantic enhancement of a large language model includes a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement the traffic flow prediction method based on semantic enhancement of a large language model as described above.
[0020] By adopting the above scheme, the present invention has the following advantages and beneficial effects: The traffic flow prediction method based on semantic enhancement of large language model in the embodiments of the present invention innovatively introduces large language model (LLM) to transform traffic flow statistical profiles into semantic embeddings, successfully constructing a global semantic association graph, breaking through the limitation of traditional models that only rely on static geographical distance, thereby being able to capture implicit functional associations and global dependencies across regions. At the same time, the model combines a dynamic adaptive fusion gating mechanism, which can adaptively switch between local physical propagation and global semantic collaboration; adopting a parameter sharing bottleneck structure of "dimensionality reduction-convolution-dimensionality increase" significantly reduces model complexity and memory usage. In addition, the robust normalization at the data level, combined with the composite loss function (Huber+L1) that dynamically decays with training rounds, greatly improves the robustness of the model against sudden traffic anomaly noise, thereby comprehensively improving the accuracy and stability of urban road network traffic flow prediction. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the overall process of the present invention.
[0022] Figure 2 This is a flowchart of the training and testing steps for the traffic flow prediction model of the present invention.
[0023] Figure 3 This is the process trigger for the integration of the large model of this invention with the ASTGCN language and the overall processing stage.
[0024] Figure 4 This is a diagram of the main architecture of the LLMASTGCN (Large Language Model Enhanced Spatiotemporal Graph Traffic Flow Prediction Model) according to an embodiment of the present invention.
[0025] Figure 5 This is a comparison chart of the traffic flow prediction performance of a node on the public dataset PEMS04 before applying the Large Language Model Enhancement (ASTGCN) in an embodiment of the present invention.
[0026] Figure 6 This is a comparison chart of the traffic flow prediction results of a node on the public dataset PEMS04 after applying large language model enhancement (LLMASTGCN) in an embodiment of the present invention.
[0027] Figure 7 This is a comparison chart of the traffic flow prediction performance of a node on the public dataset PEMS08 before applying the Large Language Model Enhancement (ASTGCN) in an embodiment of the present invention.
[0028] Figure 8 This is a comparison chart of the traffic flow prediction results of a node on the public dataset PEMS08 after applying large language model enhancement (LLMASTGCN) in an embodiment of the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Reference manual attached Figures 1 to 8 As shown, this invention provides a traffic flow prediction method based on semantic enhancement of a large language model. It can be executed by a spatiotemporal graph convolutional traffic flow prediction device based on semantic enhancement of a large language model (hereinafter referred to as: prediction device). Specifically, it is executed by one or more processors in the prediction device to implement steps S1 to S2: S1. Acquire multi-source monitoring data from road network traffic sensors and perform preprocessing.
[0031] Specifically, the steps for preprocessing the monitoring data include steps S11 to S15.
[0032] S11. Acquire monitoring data and physical topology data from urban road network traffic sensors and perform preprocessing.
[0033] Acquire the time-series traffic flow data and sensor distance matrix of sensor nodes within a historical time step T; extract the time-series traffic flow data and construct the original historical traffic flow tensor. Where T is the historical time step, N is the total number of sensor nodes, and C is the dimension of the input features. This represents the road network traffic state matrix at time step t. Represents the set of real numbers. This indicates that the tensor exists in a real space with dimensions of time step × number of nodes × feature dimension.
[0034] S12. The historical feature data is preprocessed to extract multidimensional statistical features, and a large language model is introduced to extract the semantic features of the nodes. The physical adjacency matrix and the global semantic association matrix are constructed as parameters of the future time step predicted traffic label tensor. Constructing the physical adjacency matrix and the global semantic association matrix specifically includes: Construction of the physical adjacency matrix: Based on the geographical distance between sensors, a threshold Gaussian kernel function is used to construct the physical adjacency matrix. The weights are calculated as follows: like ,but Otherwise, it is 0, where For nodes With nodes The Euclidean distance between them Let be the standard deviation of the Gaussian kernel function. This is the distance truncation threshold; Global semantic association matrix construction: computation nodes A multidimensional statistical profile of the sensor, including mean, standard deviation, daily coefficient of variation, information entropy, and the ratio of morning to evening peak hours, is generated and converted into natural language descriptive text. ; Describe the text The input is fed into a large language model, and the mean of its last hidden state is extracted as the semantic embedding vector of the node. ,in For large language models, nonlinear feature mapping function, For mean pooling operation, Let be the dimension of the semantic embedding vector; calculate the cosine similarity between nodes based on the semantic embedding vector, and select the Z most similar nodes to construct a semantic association matrix. : .
[0035] To overcome the limitations of traditional static maps that rely too heavily on geographical distance and struggle to capture dynamic relationships between nodes with similar functional attributes across regions (such as business centers in different regions), this step employs a large language model-driven node semantic profiling technology.
[0036] S13. Robust normalization processing is performed on the time-series traffic data in the historical feature data, that is, the time-series traffic feature data of each channel is robustly scaled in the global range using the median and interquartile range to obtain normalized time-series data of a uniform scale. Specifically, because urban traffic flow is highly susceptible to extreme values due to unforeseen events, a robust scaler based on the median and interquartile range is used instead of traditional max-min normalization. This is achieved by calculating the median of the training set data. (Right now and the 75th percentile and the 25th percentile The following formula is used for global normalization:
[0037] Where X represents the original time-series feature data; when the denominator To avoid computational anomalies, the scaling result of this feature is then... The data was uniformly set to 0. The processed data was scaled to a reasonable range, significantly enhancing the model's ability to resist outlier interference.
[0038] S14. Based on the historical time step T and the future prediction time step T', construct sample pairs using the normalized time series data obtained in step S13; specifically: extract the data segment of the historical time step T from the normalized time series data as the input feature of the model, and set the normalized flow value of the following future prediction time step T' as the corresponding future flow status label.
[0039] The normalized time-series data from the road network sensors is labeled with corresponding future traffic status tags. Specifically, this includes: for the normalized time-series data obtained in step S13, using continuous time steps of historical time step T as input historical tensors. The subsequent continuous time steps of length T' are used as the real flow label tensors for future time steps. The model learns a nonlinear mapping function. Achieving future time step prediction of traffic label tensors Real traffic label tensor for future time steps The approximation of is expressed mathematically as follows:
[0040] in, Predict traffic label tensors for future time steps. To input the history tensor, It is the physical adjacency matrix. This is the global semantic association matrix. This is the set of parameters that the model needs to learn.
[0041] S15. Divide the generated sequence samples into training, validation, and test sets according to a predetermined ratio. To prevent leakage of time-series data, strictly divide the processed samples into training, validation, and test sets in a 7:1:2 ratio according to the time sequence.
[0042] S2. Input the preprocessed monitoring data into the pre-trained spatiotemporal graph convolutional traffic flow prediction model (LLMASTGCN) based on semantic enhancement of a large language model to obtain the future traffic flow prediction status.
[0043] Specifically, the pre-trained spatiotemporal graph convolutional traffic flow prediction model has established a nonlinear mapping relationship between "physical spatial topological features," "implicit global semantic association features," and "future traffic flow labels." By inputting preprocessed monitoring data into the model, it can output traffic flow prediction tensors for multiple future time steps, thereby accurately predicting the future continuous road network status based on historical monitoring data and providing scientific decision support for urban road network planning and dynamic management.
[0044] Preferably, the spatiotemporal graph convolutional traffic flow prediction model based on semantic enhancement of a large language model is trained as follows: I. Based on the training and validation sets, train a dual-view spatiotemporal graph convolutional prediction model (LLMASTGCN) that combines dynamic adaptive fusion gating and lightweight spatiotemporal blocks.
[0045] A multi-view spatiotemporal graph neural network is constructed, consisting of an input projection layer, multiple stacked lightweight KA-ST blocks, and an output prediction layer. Each lightweight KA-ST block performs a deep extraction and fusion of spatiotemporal features, specifically including the following sub-modules: (1) Physics-semantics dual-view graph structure fusion module (multi-view graph convolution): In the physical adjacency matrix semantic association matrix Feature aggregation is performed on top of the model. To reduce the number of parameters under high concurrency, the model adopts a bottleneck structure of "dimensionality reduction-convolution-dimensionality increase," and the feature transformation parameters are shared between the physical perspective's Chebyshev polynomial approximations and the semantic perspective. :
[0046]
[0047] in, For the first Chebyshev polynomials The Laplace matrix of the physics graph. For the first The input features of the layer and These are the dimension reduction and dimension increase projection matrices, respectively.
[0048] (2) Dynamic adaptive gating fusion mechanism: To overcome the limitations of fixed fusion ratios, a lightweight gating network is designed. Dual-view features are processed through a multilayer perceptron (MLP) and a softmax activation function to autonomously generate mutually exclusive dynamic gating coefficients. :
[0049] The fused features are represented as follows:
[0050] (3) Enhanced spatiotemporal modeling module: Fusion features The data is then input into the serial spatiotemporal module. First, short-term trends are captured using local one-dimensional convolution (Conv1D) to obtain... Then, long-term temporal dependencies are captured through a multi-head self-attention mechanism:
[0051]
[0052] Finally, the high-dimensional spatiotemporal features extracted by multiple KA-ST blocks are fed into the output layer and mapped to the target dimension prediction result through two consecutive 1x1 convolutional layers. .
[0053] In the end-to-end parameter learning phase, a dynamic composite loss function is adopted. Balancing robustness against outliers with prediction accuracy during stationary periods:
[0054] in, Huber loss (truncated threshold) ), For L1 loss. Dynamic weighting coefficients. Decreasing with each training epoch, such as .
[0055] The optimizer uses AdamW with decoupled weight decay, combined with gradient accumulation and ReduceLROnPlateau cosine annealing learning rate scheduler. Early stopping is triggered when the validation set error fails to decrease for multiple consecutive times to prevent overfitting and preserve the optimal model weights.
[0056] 2. Input the test set into the validated and output the best prediction model, destandardize the output results and evaluate them.
[0057] The original data scale of the predicted values is restored using the inverse standardization function, and the mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and coefficient of determination (R²) are calculated. 2 The evaluation model is assessed using indicators such as ( ) and ( ). The specific formula is as follows: Mean Absolute Error Formula:
[0058] Root mean square error formula:
[0059] Mean absolute percentage error formula:
[0060] Formula for coefficient of determination: .
[0061] When the metrics meet the preset accuracy threshold, the final spatiotemporal graph convolutional traffic flow prediction model based on semantic enhancement of a large language model is determined. The specific criteria for model qualification are: whether the model's prediction performance on the validation set meets the preset training convergence criterion; if it does, the iteration stops and the current parameters are saved as the trained prediction model; otherwise, the next round of training continues. The preset training convergence criterion specifically includes any one of the following: 1. The composite loss function value calculated by the model on the validation set. No decrease was observed in P consecutive training rounds, where P is the set early stopping patience threshold; 2. The total number of training rounds of the model reaches the set maximum number of iterations threshold.
[0062] To verify the effectiveness of the proposed model and its solution, two publicly available benchmark datasets, PEMS03 and PEMS08, in the field of highway traffic flow were selected for comprehensive experimental evaluation.
[0063] The PEMS04 dataset contains 358 traffic sensor nodes and 340 edges.
[0064] The PEMS08 dataset contains 170 traffic sensor nodes and 295 edges.
[0065] Both datasets use a traffic sampling frequency of once every 5 minutes, with an input historical time step T=12 (i.e., 1 hour of data) and a predicted future time step T'=12 (the next hour). The data is divided into training, validation, and test sets in a 7:1:2 ratio according to the time series. Examples of feature data for some nodes and edges are shown in Table 1.
[0066] Table 1. Examples of original nodes and topological features in the traffic flow dataset:
[0067] The hyperparameter settings for the deep learning network model of this invention are shown in Table 2 below: Table 2 Experimental parameter settings:
[0068] Thanks to the lightweight spatiotemporal blocks and parameter-sharing bottleneck structure of dimensionality reduction-convolution-upgrading employed in this invention, the model significantly reduces computational complexity and memory usage. In one specific embodiment, the prediction model of this invention can be successfully deployed and run on conventional consumer-grade computing devices equipped with only 4GB of video memory, overcoming the excessive reliance of traditional multi-view large models on high-performance computing resources, and possessing extremely high engineering implementation and promotion value.
[0069] The proposed spatiotemporal graph convolutional traffic flow prediction algorithm (LLMASTGCN) based on semantic enhancement of a large language model was used to train and test the model. The model was compared with other cutting-edge mainstream models (such as LSTM, GraphWaveNet, STSGCN, ASTGCN, and MVSTGCN) to observe its performance on two public datasets.
[0070] The specific macroscopic test results are shown in Table 3.
[0071] Table 3 shows the experimental comparison results of each algorithm on the PEMS04 and PEMS08 datasets:
[0072] In summary, the traffic flow prediction model proposed in this invention can automatically extract physical and implicit semantic spatiotemporal features from multi-source monitoring data of the road network and provide high-precision future traffic flow predictions. Addressing the problems of existing traffic flow prediction methods that over-rely on static geographical distance, struggle to capture cross-regional implicit functional relationships, and neglect global spatial dependencies, this method innovatively introduces large language models such as DeepSeek as a "semantic understander" for the road network. It transforms the multidimensional statistical profile of traffic flow into high-dimensional semantic embeddings, constructing a global semantic similarity graph that reflects real functions. Simultaneously, leveraging a dynamic gating fusion mechanism and a lightweight spatiotemporal bottleneck structure with Chebyshev shared weights, the model can adaptively integrate local micro-features from a physical perspective with global macro-dependencies from a semantic perspective. Thus, the model successfully achieves a leap from "local physical perception" to "global semantic collaboration" in modeling, significantly improving prediction accuracy and robustness under complex road networks (such as PEMS04 and PEMS08) and extreme traffic fluctuations.
[0073] Example 2: This embodiment of the invention provides a spatiotemporal graph convolutional traffic flow prediction device based on large language model semantic enhancement, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a traffic flow prediction method based on large language model semantic enhancement as described in any paragraph of Example 1.
[0074] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A traffic flow prediction method based on semantic enhancement of a large language model, characterized in that, Includes the following steps: S1. Acquire monitoring data and physical topology data from urban road network traffic sensors, and perform preprocessing; the preprocessing includes the following steps: S11. Obtain historical feature data from road network traffic sensors, wherein the historical feature data includes time-series traffic flow data and sensor distance matrix data; S12. The historical feature data is preprocessed to extract multidimensional statistical features, and a large language model is introduced to extract the semantic features of the nodes. The physical adjacency matrix and the global semantic association matrix are constructed respectively. Constructing the global semantic association matrix includes: calculating sensor nodes A multidimensional statistical profile, including mean, standard deviation, daily coefficient of variation, information entropy, and the ratio of morning to evening peak hours, is generated and transformed into natural language descriptive text. ; Describe the text The input is fed into a large language model, and the mean of its last hidden state is extracted as the semantic embedding vector of the node. ,in For large language models, nonlinear feature mapping function, For mean pooling operation, Let be the dimension of the semantic embedding vector; calculate the cosine similarity between nodes based on the semantic embedding vector, and select the Z most similar nodes to construct a semantic association matrix. : ; S13. Robust normalization processing is performed on the time-series traffic data in the historical feature data, that is, the time-series traffic data of each channel is robustly scaled in the global range using the median and interquartile range to obtain normalized time-series data of a uniform scale. S14. Based on the historical time step T and the future prediction time step T', construct sample pairs using the normalized time series data obtained in step S13; specifically: extract the data segment of the historical time step T from the normalized time series data as the input feature of the model, and set the normalized flow value of the next future prediction time step T' as the corresponding future flow status label according to the physical adjacency matrix and the global semantic association matrix. S15. For the normalized flow values with set labels, adopt the equal interval sliding window segmentation strategy to extract multiple equal-length input subsequences and corresponding target subsequences from the normalized flow values with a fixed-length time window, and generate preprocessed monitoring data samples. S2. Input the preprocessed monitoring data samples into the pre-trained spatiotemporal graph convolutional traffic flow prediction model based on large language model semantic enhancement to obtain the traffic flow prediction results for future continuous time steps; the spatiotemporal graph convolutional traffic flow prediction model based on large language model semantic enhancement is trained based on the following steps: based on the training set and validation set, train the traffic flow prediction model constructed by multi-view feature extraction, dynamic gating fusion and enhanced spatiotemporal perception network, and use a dynamic composite loss function composed of Huber Loss and L1 Loss weights for parameter optimization to establish a nonlinear mapping relationship between historical traffic status and future traffic labels.
2. The traffic flow prediction method based on semantic enhancement of a large language model according to claim 1, characterized in that, Step S11: Obtain historical feature data from road network traffic sensors, specifically: obtain time-series traffic flow data and sensor distance matrix of sensor nodes located in the road network within a historical time step T; extract the time-series traffic flow data and construct the original historical traffic flow tensor. Where N is the total number of sensor nodes, and C is the dimension of the input features. This represents the road network traffic state matrix at time step t. Represents the set of real numbers. This indicates that the tensor exists in a real space with dimensions of time step × number of nodes × feature dimension.
3. The traffic flow prediction method based on semantic enhancement of a large language model according to claim 1, characterized in that, Step S12 involves constructing the physical adjacency matrix and the global semantic association matrix, specifically including: Construction of the physical adjacency matrix: Based on the geographical distance between sensors, a threshold Gaussian kernel function is used to construct the physical adjacency matrix. The calculation process is as follows: like ,but Otherwise, it is 0, where For sensor nodes With sensor nodes The Euclidean distance between them This represents the element in the i-th row and j-th column of the physical adjacency matrix. Let be the standard deviation of the Gaussian kernel function. This is the distance truncation threshold.
4. The traffic flow prediction method based on semantic enhancement of a large language model according to claim 1, characterized in that, In step S13, the time-series traffic data of each channel is robustly normalized globally using the median and interquartile range. Specifically, this includes normalizing the time-series traffic data of each channel using a robust scaling algorithm to eliminate the impact of outliers such as sudden congestion. The normalization formula is as follows: in, The original historical traffic flow tensor obtained in step S11, The normalized data tensor obtained after robust scaling. This represents the median of the time-series flow data, sorted in ascending order of numerical value, at the 50th percentile. and These are the 75th and 25th percentiles of the time-series traffic data, respectively.
5. The traffic flow prediction method based on semantic enhancement of a large language model according to claim 3, characterized in that, Step S14 involves setting corresponding future traffic status labels for the normalized time-series data of the sensor nodes. Specifically, this includes: for the normalized time-series data obtained in step S13, using continuous time steps of historical time step T as input historical tensors. The subsequent continuous time steps of length T' are used as the real flow label tensors for future time steps. The model learns a nonlinear mapping function. Achieving future time step prediction of traffic label tensors Real traffic label tensor for future time steps The approximation of is expressed mathematically as follows: in, Predict traffic label tensors for future time steps. To input the history tensor, It is the physical adjacency matrix. This is the global semantic association matrix. This is the set of parameters that the model needs to learn.
6. The traffic flow prediction method based on semantic enhancement of a large language model according to claim 5, characterized in that, Step S15 specifically includes: For the normalized flow values with set labels, a sliding window sampling strategy with historical time step T, future prediction time step T', and sliding interval of step 1 is used to extract multiple fixed-length sequence samples from the original normalized flow values. The specific operation for sample construction is as follows: for a normalized data tensor with a total length of L in the time dimension... Traverse normalized data tensors Index segmented by time and In each iteration, a subsequence of the normalized data tensor is extracted. as input history tensor Extract a subsequence from another normalized data tensor. Tensor as the real traffic label of the future time step .
7. The traffic flow prediction method based on semantic enhancement of a large language model according to claim 1, characterized in that, Step S2 divides the monitoring data samples into training set, validation set and test set according to the time steps in chronological order in a ratio of 7:1:
2. End-to-end parameter learning is performed using the training set, and optimization is achieved using a dynamic composite loss function. During training, the training process is supervised by validation set metrics to determine early stopping and adaptive adjustment of the learning rate. The optimal model weights are selected based on the validation set. The final prediction model selected is evaluated using mean absolute error, root mean square error, mean absolute percentage error, and coefficient of determination to determine whether the model is qualified.
8. The traffic flow prediction method based on semantic enhancement of a large language model according to claim 7, characterized in that, Step S2 involves training a traffic flow prediction model based on the training and validation sets. This model is constructed using multi-view feature extraction, dynamic gating fusion, and an enhanced spatiotemporal awareness network to establish a nonlinear mapping relationship between historical traffic conditions and future traffic labels. Specifically, the prediction model includes a lightweight spatiotemporal module that internally performs dual-view feature extraction: graph convolution is performed on both the physical and semantic graphs, employing a bottleneck structure of dimensionality reduction-convolution-upgrading, and Chebyshev approximations of different orders share transformation parameters. The formula is: In the formula, For the first Chebyshev polynomials This is the highest order of the Chebyshev polynomial. For the physical adjacency matrix, For the first The node feature matrix of the layer input, For activation function, and These are the dimension reduction and dimension increase projection matrices, respectively. To share the transformation parameter matrix; Dynamic fusion gating: and Input a lightweight gating network and generate dynamic gating coefficients using Softmax. The fusion formula is ;in, For the fused spatiotemporal feature tensor; operators This represents vector concatenation operations; operators Hadamard product, which is the element-wise multiplication of a matrix; dynamic gating coefficient. The specific expression is: In the formula, This is a mapping function for a multilayer perceptron, used to learn the contribution weights of physical and semantic features. It is the bias vector; Enhanced spatiotemporal modeling module: First, it captures short-term trends through local one-dimensional convolution to obtain... The expression is: in: It is a local temporal feature representation, that is, the feature after local one-dimensional convolution processing. For layer normalization, For the fused spatiotemporal feature tensor, This involves a local one-dimensional convolution operation; then, a multi-head self-attention mechanism is used to capture long-range time. in For the final spatiotemporal feature output, It is a multi-head self-attention mechanism used to capture long-distance temporal dependencies; Finally, the high-dimensional spatiotemporal features extracted by multiple KA-ST blocks are fed into the output layer and mapped to the target dimension prediction result through two consecutive 1x1 convolutional layers. , ,in For hidden layer convolution, To modify the activation function of the linear unit, For output layer convolution; Composite Loss Function: To balance robustness against sudden outliers with prediction accuracy for stable flow rates, a dynamic composite loss function consisting of a weighted average of Huber Loss and L1 Loss is used for parameter optimization. Dynamic decay with training rounds: In the formula: This represents the total loss value during model training. For the real traffic label tensor of future time steps; Predict traffic label tensors for future time steps; The Huber loss function is used to reduce the drastic impact of outliers on the gradient in the early stages of training, thereby enhancing the robustness of the model. The mean absolute error loss function is used to improve the regression prediction accuracy of the model in stable flow ranges. The weighting coefficients are dynamically changed with each training round, and their specific meaning is: assigned weights in the initial training phase. The highest weights are used to quickly stabilize the model, and as the number of training epochs (e) increases, Gradual decay causes the model's center of gravity to shift. To achieve higher prediction accuracy.
9. A traffic flow prediction method based on semantic enhancement of a large language model according to claim 7, characterized in that, The specific calculation rules for the evaluation indicators used are as follows: The test set is input into the trained prediction model, and the high-dimensional feature output is denormalized to restore the original traffic flow to its true scale. Let Y denote the denormalized overall true traffic label tensor and denot the overall predicted traffic label tensor as Y. Let the total number of data points in the test set participating in the evaluation be . Where m is the index of the data point, calculate the following metrics: Mean Absolute Error formula: Root mean square error formula: Mean absolute percentage error formula: Coefficient of determination formula: In the formula, This represents the actual traffic flow value of the m-th data point in the overall actual traffic flow label tensor Y. For the overall predicted traffic label tensor The corresponding m-th predicted traffic flow value, The mean of all real traffic flows in the test set; for the calculation of MAPE, M' is the set of valid data points where the actual flow rate is greater than a set minimum threshold. The total number of data points in the data is used to avoid division by zero anomalies when traffic flow is zero.
10. A traffic flow prediction device based on semantic enhancement of a large language model, characterized in that, It includes a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement a traffic flow prediction method based on semantic enhancement of a large language model as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Transportation business quick response processing system based on large language model technology
CN119274350A
Traffic flow prediction method based on large language model
CN121999615A