Traffic flow prediction method and system based on improved space-time coupling KA-LSTM

Through the improved space-time coupled KA-LSTM model, the traditional traffic flow prediction method has solved the shortcomings in space-time modeling accuracy and dynamic adaptability, and achieved high-precision, real-time and efficient prediction of traffic flow, which is suitable for complex traffic scenarios.

CN120260264AActive Publication Date: 2025-07-04HUAIYIN INSTITUTE OF TECHNOLOGY

Patent Information

Application Number
CN202510269770.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-04
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

Traditional traffic flow prediction methods have shortcomings in spatial and temporal modeling accuracy, dynamic adaptability and computational efficiency, and it is difficult to effectively capture the nonlinear dynamic characteristics and complex spatial and temporal relationships of traffic flow, especially in complex dynamic scenarios, with limited generalization ability.

Method used

Using the improved spatiotemporal coupled KA-LSTM model, the forget gate, input gate and memory update mechanism are dynamically optimized by integrating Kolmogorov-Arnold network (KANs) and LSTM, and combining multi-head attention mechanisms and gated fusion technology, the modeling ability of the spatiotemporal characteristics of traffic flow is enhanced, and the spatial and temporal consistency loss function is introduced to optimize the model performance.

Benefits of technology

It significantly improves the ability to capture long-term dependence on traffic flow, dynamically balances short-term fluctuations with long-term periodicity, enhances the prediction accuracy and robustness of the model, adapts to complex traffic scenarios, and meets the needs of real-time and efficient computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260264A_ABST
    Figure CN120260264A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic flow prediction method and system based on improved space-time coupling KA-LSTM, and belongs to the technical field of intelligent traffic and artificial intelligence. According to the method, for the problems that a traditional model is poor in adaptability, low in prediction precision and insufficient in calculation efficiency in a complex traffic scene, a KA-LSTM model is constructed, a forgetting gate, an input gate and a memory updating mechanism of LSTM are dynamically optimized through KANs, the capturing capacity for long-period dependence is enhanced, and a system comprises a data collection module, a model training module, an anomaly detection module and a real-time optimization module. Experiments show that the indexes such as MAE and RMSE are reduced by more than 40.1% compared with those of traditional LSTM, the training efficiency is remarkably improved, and the method is suitable for urban traffic management and optimization. The method has the core advantages that through dynamic gating optimization, multi-modal data fusion and an online updating strategy, the method has good performance effects in prediction precision, real-time performance and complex scene adaptability, and an efficient and reliable solution is provided for traffic flow prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the research field at the intersection of intelligent transportation systems and artificial intelligence technology, and specifically relates to a traffic flow prediction method and system for improving spatio-temporal coupled KA-LSTM. Background Art

[0002] Traffic flow prediction, as a core task of intelligent transportation systems (ITS), aims to accurately predict the traffic flow, speed, and density of future road networks using historical traffic data, providing key support for traffic management, congestion mitigation, and resource optimization. With the acceleration of urbanization and the surge in traffic demand, the performance of existing prediction methods still faces significant challenges in complex dynamic scenarios.

[0003] Traditional time series models (such as ARIMA) rely on linear assumptions and are difficult to capture the non-linear dynamic characteristics in traffic flow. Although recurrent neural networks (RNNs) and their improved long short-term memory networks (LSTMs) have shown advantages in time series modeling, their ability to capture long-term dependencies is still limited by the static design of the gating mechanism. For example, the forget gate and input gate of traditional LSTM update the memory state through fixed weight parameters, making it difficult to adapt to the periodic fluctuations across time scales in traffic flow (such as the differences between morning and evening rush hours and weekend patterns), resulting in insufficient modeling accuracy for complex time series patterns.

[0004] The traffic flow changes at nodes in the traffic road network have strong spatial dependencies. For example, the congestion propagation effect between adjacent road sections or the interaction between main roads and branch roads. Existing methods (such as models based on convolutional neural networks (CNNs) or single LSTM structures) usually use fixed adjacency matrices or local convolutional kernels to model spatial relationships, and are unable to effectively represent the implicit and non-linear dynamic interactions between nodes. In addition, static spatial modeling is difficult to adapt to road network topology changes (such as construction or temporary closures), restricting the generalization ability of the model in complex scenarios.

[0005] Most methods separately model time and space features (such as first using LSTM to extract time series features and then processing spatial relationships through graph convolution), and fail to fully exploit the synergistic effects of spatio-temporal dimensions. The spatio-temporal evolution of traffic flow has a high degree of coupling (such as the rapid diffusion of traffic flow during peak hours in specific areas), and traditional step-by-step modeling strategies are prone to losing cross-dimensional dynamic correlation information, resulting in prediction results deviating from the actual distribution.

[0006] The spatio-temporal distribution of traffic flow is affected by multiple factors such as weather and emergencies, and has significant dynamic characteristics. However, the parameter weights of traditional LSTM models are fixed and cannot dynamically adjust the feature importance according to real-time data (such as the priority of sudden traffic flow increases or accident nodes). In addition, existing methods usually update the memory state using a single time scale, making it difficult to balance the contributions of short-term fluctuations and long-term trends, further restricting the prediction accuracy. Summary of the Invention

[0007] Object of the Invention: The present invention aims to solve the deficiencies of traditional traffic flow prediction methods in spatio-temporal modeling accuracy, dynamic adaptability, and computational efficiency, and proposes a traffic flow prediction method and system based on improved spatio-temporal coupled KA-LSTM. By improving the LSTM structure and introducing a spatio-temporal joint mechanism, the prediction performance is enhanced.

[0008] Technical Solution: The present invention discloses a traffic flow prediction method based on improved spatio-temporal coupled KA-LSTM, including the following steps:

[0009] Step 1: Obtain multi-dimensional traffic data from road network sensors.

[0010] Step 2: Deeply fuse the Kolmogorov-Arnold network (KANs) with the long short-term memory network (LSTM), improve the forgetting gate, input gate, memory update, and output gate mechanisms of the traditional LSTM, and construct a KA-LSTM model, thereby enhancing the dynamic modeling ability of spatio-temporal characteristics of traffic flow.

[0011] Step 3: Through the multi-head attention mechanism and gating fusion technology, fuse the hidden state generated by KA-LSTM with spatio-temporal features, further enhancing the model's ability to model the dynamic spatio-temporal relationship of traffic flow.

[0012] Step 4: Train and predict the KA-LSTM model

[0013] Further, in the above Step 1, the collected traffic data is preprocessed by using spatio-temporal weighted interpolation method to complete missing data, removing outliers based on the Z-score method, standardization, and performing Z-score standardization for each feature dimension.

[0014] Further, a three-dimensional spatio-temporal tensor is generated Where, T: time step (such as a time window with a granularity of 5 minutes); N: number of road network nodes (such as the number of road segments covered by traffic sensors); F: feature dimension (such as flow, speed, density, etc.).

[0015] Further, a road network adjacency matrix is generated by using data such as road network topology structure data and road network connection relationship data

[0016] Among them, the use of road network topology structure data includes: 1) Node information: identifiers and geographical coordinates of all nodes in the network (such as intersections, starting and ending points of road segments, etc.). 2) Edge information: describing the connection relationship between nodes, that is, which nodes are directly connected by roads. Usually includes the starting node, ending node of the edge, and possible attributes.

[0017] The road network connection relationship data includes: 1) Road connectivity: It clarifies which nodes are directly connected and which are not. This can be obtained through the connectivity data of the road network, such as extracting from Geographic Information System (GIS) data. 2) Road directionality: If there are one-way roads in the road network, the directionality information of the roads needs to be recorded to ensure that the adjacency matrix can accurately reflect the actual traffic flow direction.

[0018] Furthermore, in step 2, the KA-LSTM model is specifically designed into initialization and input definition, forget gate improvement, input gate dynamic weighting, multi-scale memory fusion, and output gate optimization.

[0019] Furthermore, in step 2, initialize the hidden state and memory state both with dimension d h , and the initial value is set to a zero vector or determined through pre-training. Then, based on the road network adjacency matrix and the input at each time step t extracted step by step on the time dimension is Generate global spatial features through graph convolution:

[0020]

[0021] where, is the adjacency matrix with self-loops, is the degree matrix, is the trainable weight, d s is the spatial feature dimension, which is consistent with the subsequent KANs input.

[0022] Furthermore, in step 2, the KANs sub-network KAN f receives the global spatial feature the previous moment's hidden state and the original input feature of node i at the current time step t to generate a dynamic forget bias: Furthermore, it is fused into the forget gate to obtain the following new forget gate:

[0023]

[0024] where, is the weight matrix, is the bias, σ is the sigmoid activation function, is the output of the forget gate of node i at time step t, indicating the proportion of retaining the memory of the previous moment .

[0025] Furthermore, in step 2, the KANs sub-network KANβ receives the original input features of node i at the current time step t to generate a feature importance vector Furthermore, for weighted input calculation is performed where ⊙ is element-wise multiplication, is the weighted input feature. Furthermore, is fused into the input gate as the weighted input feature to obtain a new input gate formula:

[0026]

[0027] where is the weight matrix, is the bias, is the input gate output, representing the acceptance degree of new information.

[0028] Furthermore, in step 2, the candidate memory state introduces the enhanced candidate memory vector generated by the global spatial feature through the KANs sub-network KAN C to generate a new candidate memory state:

[0029]

[0030] Furthermore, the long-term hidden state generates the long-term memory weight that dynamically balances short-term fluctuations and long-term periodicity through the KAN sub-network KAN γ : Furthermore, the memory state formula is updated to:

[0031]

[0032] where is the current memory state, which fuses the short-term update and the long-term memory

[0033] Furthermore, in step 2, the KANs sub-network KAN f receives the global spatial feature to generate a dynamic forgetting bias Furthermore, it is fused into the output gate to obtain the following new output gate and hidden state output:

[0034] 1) Output gate: where and are the weights and biases, and KAN oIt is the KANs sub-network, which outputs a dynamic bias with a dimension of d h .

[0035] 2) Hidden state output: Among them, is the hidden state of node i at time t.

[0036] Furthermore, in step 3, the hidden state output by the KA-LSTM model in step 2 is received (the hidden state of each node i at time t, d h is the hidden state dimension) and the spatial feature (generated by the graph convolution in step 2, d s is the spatial feature dimension).

[0037] Furthermore, in step 3, the multi-head attention mechanism is used to extract the relevance of spatio-temporal features from different subspaces, and the expression ability of the model is enhanced through the following design.

[0038] Furthermore, the hidden state output by the KA-LSTM model is used (the hidden state of each node i at time t, d h is the hidden state dimension) and the spatial feature (generated by the graph convolution in step 2, d s is the spatial feature dimension). Queries, keys, and values are generated, and vectors required for multi-head attention are generated based on the hidden state and the spatial feature :

[0039]

[0040] Among them:

[0041] is the query weight matrix, W K , are the key and value weight matrices respectively;

[0042] d k = d h / H is the dimension of each attention head, and H is the number of attention heads;

[0043] are the query vector of the node and the key and value vectors of node j respectively, where j ∈ {1,..., N} traverses all nodes.

[0044] Furthermore, and are used to calculate the attention weights, and the similarity between nodes is calculated and normalized for each head of attention:

[0045]

[0046] Wherein:

[0047] h = 1, …, H represents the h-th attention head;

[0048] and are the query and key vectors of the h-th head;

[0049] is a scaling factor to avoid excessive numerical values;

[0050] is the attention weight of node i to node j, reflecting spatial correlation.

[0051] Furthermore, using and the generated attention weights and perform weighted aggregation of the features of all nodes to obtain the attention output:

[0052]

[0053] Wherein:

[0054] is the value vector of the h-th head; is the multi-head attention output of the node, generated by concatenating the results of H heads and linear transformation:

[0055]

[0056] Among them, is the output weight matrix.

[0057] Further, in the step 3, the original hidden state and the attention output are dynamically balanced through gate control to generate a gated unit:

[0058]

[0059] Wherein:

[0060] is the gated weight matrix, is the bias;

[0061] is the concatenation of the hidden state and the attention output;

[0062] is the gated vector, and each dimension independently controls the fusion ratio of the corresponding feature.

[0063] Furthermore, combining the gated unit the original hidden state And attention output Generate the final hidden state:

[0064]

[0065] Wherein: is the final hidden state of node i at time t; Retain the original temporal information, and incorporate the spatial attention result.

[0066] Furthermore, in step 4, by designing a joint loss function, an optimization strategy, and a prediction output process, train the KA-LSTM model and its spatio-temporal attention mechanism, and use the fully connected layer to output the prediction result.

[0067] The present invention discloses a traffic flow prediction system based on an improved spatio-temporal coupled KA-LSTM, including the following sub-modules:

[0068] (1) Data preprocessing module: used to collect and process road network traffic data to generate a spatio-temporal tensor;

[0069] (2) KA-LSTM prediction module: based on the LSTM structure improved by KANs and the spatio-temporal attention mechanism, generate traffic flow prediction results;

[0070] (3) Attention weight distribution recognition module: recognize traffic anomalies based on the attention weight distribution and adjust the prediction strategy;

[0071] (4) Real-time optimization interface module: output the prediction result to the traffic management system to support signal timing optimization.

[0072] Furthermore, for the traffic flow prediction system based on the improved spatio-temporal coupled LSTM, the data preprocessing module includes the following sub-modules: (1) Data collection module: automatically collect traffic flow, speed, and density data from road network sensors, with a time granularity of 5 minutes. (2) Missing value processing module: use spatio-temporal weighted interpolation method to fill in the missing data. (3) Outlier detection module: remove outliers based on the Z-score method, with the threshold set to 3. (4) Data normalization module: perform Z-score normalization on each feature dimension to generate a three-dimensional spatio-temporal tensor.

[0073] Furthermore, for the traffic flow prediction system based on the improved spatio-temporal coupled LSTM, the KA-LSTM prediction module includes the following sub-modules: (1) Forgetting gate sub-module: generate dynamic forgetting weights using KANs; (2) Input gate sub-module: achieve dynamic feature weighting through KANs; (3) Multi-scale memory sub-module: fuse short-term and long-term features to update the memory state; (4) Spatio-temporal attention sub-module: optimize the hidden state output based on multi-head attention and gating fusion.

[0074] Furthermore, in the traffic flow prediction system based on the improved spatio-temporal coupled LSTM, the attention weight distribution recognition module includes the following sub-modules: (1) Abnormality recognition module: Recognize traffic abnormal events based on the attention weight distribution. (2) Prediction strategy adjustment module: Automatically adjust the prediction strategy according to the abnormal situation to ensure the reliability and timeliness of the prediction results.

[0075] Furthermore, in the traffic flow prediction system based on the improved spatio-temporal coupled LSTM, the real-time optimization interface module includes the following sub-modules: (1) Prediction result output module: Transmit the prediction results to the traffic management system to support functions such as signal timing optimization. (2) Real-time optimization module: Make real-time traffic management decisions according to the prediction results.

[0076] Furthermore, in the traffic flow prediction system based on the improved spatio-temporal coupled LSTM, the system function design part specifically includes the following parts:

[0077] (1) Data collection and preprocessing: Automatically collect traffic flow, speed, and density data from road network sensors with a time granularity of 5 minutes. Use spatio-temporal weighted interpolation method to complete the missing data, eliminate outliers through the Z-score method, and perform Z-score standardization on each feature dimension. Finally, generate a three-dimensional spatio-temporal tensor to provide high-quality data for subsequent model training.

[0078] (2) Traffic flow prediction: Use the LSTM structure improved based on KANs and spatio-temporal attention mechanism to deeply analyze the preprocessed data. The system can accurately predict the traffic flow in future time periods and provide the predicted flow value and confidence interval (such as 95% confidence level) to meet the prediction needs in different scenarios.

[0079] (3) Abnormality detection and adaptive adjustment: Real-time detect traffic abnormal events based on the change of the attention weight distribution. Once an abnormality is found, the system automatically increases the short-term feature weight to improve the response ability to emergencies and ensure the reliability and timeliness of the prediction results.

[0080] (4) Multi-modal data fusion: Receive multi-source information such as weather data and traffic event data, generate a dynamic weighted vector through KANs, and integrate multi-modal features into the model to enhance the adaptability of the model to complex traffic scenarios and further improve the prediction accuracy.

[0081] (5) Online update: When new data is input, the system updates the KANs sub-network parameters through incremental learning, adopting a sliding window strategy (window size is W, update frequency is every M time steps), so that the model can timely adapt to the changes in traffic conditions and maintain good prediction performance.

[0082] (6) Real-time optimization interface: Output the prediction results to the traffic management system, providing data support for traffic management decisions such as signal timing optimization, helping to alleviate traffic congestion and improve traffic operation efficiency.

[0083] Furthermore, for a traffic flow prediction system based on improved spatio-temporal coupled LSTM, its system architecture design part specifically includes the following parts:

[0084] (1) Data preprocessing module: Responsible for collecting and processing road network traffic data. This module realizes functions such as data collection, missing value processing, outlier detection, and standardization, and converts the original data into a three-dimensional spatio-temporal tensor suitable for model input.

[0085] (2) KA-LSTM prediction module: Consists of a forget gate sub-module, an input gate sub-module, a multi-scale memory sub-module, and a spatio-temporal attention sub-module. The forget gate sub-module uses KANs to generate dynamic forgetting weights; the input gate sub-module realizes dynamic feature weighting through KANs; the multi-scale memory sub-module fuses short-term and long-term features to update the memory state; the spatio-temporal attention sub-module optimizes the hidden state output based on multi-head attention and gating fusion. Each sub-module works together to generate accurate traffic flow prediction results.

[0086] (3) Attention weight distribution recognition module: Recognize traffic anomalies based on the attention weight distribution, and adjust the prediction strategy according to the anomaly situation to ensure that the prediction results can timely reflect the changes in traffic conditions.

[0087] (4) Real-time optimization interface module: Transmit the prediction results to the traffic management system, support functions such as signal timing optimization, and realize the application of the prediction results in actual traffic management.

[0088] Furthermore, for a traffic flow prediction system based on improved spatio-temporal coupled LSTM, its system usage method part specifically includes the following parts:

[0089] (1) Ensure that the road network sensors work properly and continuously collect data. The system will automatically obtain the data and perform preprocessing, and users do not need to manually intervene in the data collection and preprocessing process.

[0090] (2) After the system is initialized, it will automatically use historical data for model training. After training is completed, the system performs traffic flow prediction based on the input real-time data. Users can view the prediction results on the system interface, including traffic flow values and confidence intervals for different future time periods.

[0091] (3) During the operation of the system, traffic anomalies are monitored in real time. When an anomaly is detected, the user can learn about the anomaly information in the system prompt. The system will automatically adjust the prediction strategy without the need for additional user operations.

[0092] (4) The system automatically receives multi-source data and performs fusion processing, and at the same time conducts online updates according to the set update strategy. Users do not need to manually operate on multi-modal data fusion and online updates, but can adjust relevant parameters in the system settings.

[0093] (5) The prediction results will be automatically transmitted to the traffic management system. Traffic managers can perform operations such as signal timing optimization based on the prediction data provided by the system to improve traffic conditions.

[0094] Furthermore, in the described traffic flow prediction system based on improved spatio-temporal coupled LSTM, the system maintenance and update part specifically includes the following parts:

[0095] (1) Regular inspection: Regularly check the working status of road network sensors to ensure the accuracy and integrity of data collection. Check the operation of each module of the system, and promptly discover and solve potential problems.

[0096] (2) Data update: Timely update historical data to ensure the effectiveness of model training. Regularly clean invalid or incorrect data to improve data quality.

[0097] (3) Model optimization: Optimize and upgrade the KA-LSTM model according to new research results and actual application feedback. For example, adjust the model structure, improve algorithm parameters, etc., to enhance the prediction performance of the system.

[0098] (4) System upgrade: Regularly perform software upgrades on the system to repair vulnerabilities, enhance the stability and security of the system. At the same time, according to user needs and technological development, add new functional modules to improve the practicality and adaptability of the system.

[0099] Beneficial effects:

[0100] (1) The present invention utilizes KANs to dynamically optimize the forgetting gate, input gate, and memory update mechanism of LSTM, significantly enhancing the ability to capture long-term dependencies in traffic flow. Specifically, the dynamic forgetting weights generated by KANs avoid the static linear assumption of the traditional LSTM forgetting gate, enhancing the selective retention of historical information by the model.

[0101] (2) The present invention adopts a multi-scale memory fusion technology, dynamically balances short-term fluctuations and long-term periodicity, and adapts to the complex dynamic changes of traffic flow. Specifically, the candidate memory state fuses short-term and long-term features, enhancing the model's ability to capture multi-time scale patterns. This improvement makes the model's performance in long-term prediction significantly better than that of traditional LSTM, further improving the prediction accuracy.

[0102] (3) The present invention designs a multi-head attention mechanism to extract the relevance of spatio-temporal features from different subspaces, and dynamically capture the spatial dependence and time evolution trend between nodes. Specifically, the multi-head attention mechanism generates query, key, and value vectors, calculates the similarity between nodes and normalizes it, generates attention weights, and thus weighted aggregates the features of all nodes. This improvement significantly enhances the model's expressive ability, enabling the model to more accurately capture the dynamic spatio-temporal relationship of traffic flow, and further improving the prediction accuracy.

[0103] (4) The present invention utilizes a gated fusion technology to dynamically balance the original hidden state and the attention output, and optimize the spatio-temporal contribution of the final representation. Specifically, the gated unit generates a gated vector by concatenating the hidden state and the attention output, and each dimension independently controls the fusion ratio of the corresponding feature. This improvement enables the model to more flexibly adjust feature fusion, retain the original temporal information while integrating the spatial attention result, and further improve the prediction accuracy and the robustness of the model.

[0104] (5) The present invention introduces a joint loss function (MAE and spatio-temporal consistency loss), which improves the robustness and generalization ability of the model. Specifically, the joint loss function jointly optimizes the prediction error and time consistency, constrains the spatial feature similarity of adjacent nodes, and ensures the appropriate representation intensity of the fused image. This improvement makes the model's performance more stable in complex traffic scenarios, significantly improving the model's generalization ability and prediction accuracy. Description of the Drawings

[0105] Figure 1 : Overall architecture diagram of the KA-LSTM algorithm

[0106] Figure 2 : Structure diagram of KA-LSTM

[0107] Figure 3 : Data processing flow chart of KA-LSTM

[0108] Figure 4 : Schematic diagram of spatio-temporal attention mechanism

[0109] Figure 5 : Comparison chart of predicted values and actual values of ablation experiment

[0110] Figure 6:Comparison Chart of Predicted Values and Actual Values of Different Models

[0111] Figure 7 :Line Chart of Comparison of Ablation Experiment Indexes

[0112] Figure 8 :Bar Chart of Comparison of Ablation Experiment Indexes

[0113] Figure 9 :Line Chart of Comparison of Experiment Indexes in Comparative Experiments of Different Models

[0114] Figure 10 :Bar Chart of Comparison of Experiment Indexes in Comparative Experiments of Different Models

[0115] Figure 11 :System Architecture Diagram Detailed Implementation Manner

[0116] For a better understanding of the present invention, the present invention will be further described below in conjunction with the accompanying drawings in the embodiments of the present invention, but it is not a limitation of the present invention. Without departing from the design concept of the present invention, various modifications and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention.

[0117] As Figures 1-11 shown, a traffic flow prediction method and system based on an improved spatio-temporal coupled KA-LSTM are as follows:

[0118] Step 1: Data collection and preprocessing

[0119] (1) Data acquisition: The PEMS-D7 dataset contains traffic flow data for a specific area of Highway 7 in California. Traffic flow, speed, and density data for N sections within the time window T are extracted from this dataset, with a time granularity of 5 minutes. Since some data in the PEMS-D7 dataset are missing and abnormal, preprocessing is required.

[0120] (2) Preprocessing process:

[0121] 1) Missing value processing: The spatio-temporal weighted interpolation method is used to fill in the missing data:

[0122]

[0123] where w ij is the spatial distance weight between nodes i and j (based on the Gaussian kernel function), and α ∈ [0, 1] is the spatio-temporal balance factor, which is default set to 0.7.

[0124] 2) Outlier detection: Outliers are removed based on the Z-score method (the threshold is set to 3).

[0125] (3) Standardization: Perform Z - score standardization for each feature dimension:

[0126]

[0127] (3) Output format: Generate a three - dimensional spatio - temporal tensor (F is the feature dimension).

[0128] Step 2: KA - LSTM model design

[0129] In this step, by deeply integrating the Kolmogorov - Arnold network (KANs) with the long short - term memory network (LSTM), the forgetting gate, input gate, memory update, and output gate mechanisms of the traditional LSTM are improved, and the KA - LSTM model is constructed to enhance the dynamic modeling ability of spatio - temporal traffic flow features. The specific design is divided into the following sub - steps:

[0130] 2.1 Initialization and input definition

[0131] (1) Input data: Receive the standardized three - dimensional spatio - temporal tensor output from Step 1 (T is the time step, N is the number of nodes, F is the feature dimension, such as traffic flow, speed, density).

[0132] (2) Hidden state and memory initialization: Initialize the hidden state and memory state both with dimension d h (default 128), and the initial value is set to a zero vector or determined through pre - training.

[0133] (3) Spatial feature extraction: Based on the road network adjacency matrix and the current input X t , generate global spatial features through graph convolution:

[0134]

[0135] where, is the adjacency matrix with self - loops, is the degree matrix, is the trainable weight, d s is the spatial feature dimension (such as 64), which is consistent with the subsequent KANs input.

[0136] 2.2 Improvement of the forgetting gate

[0137] (1) KANS introduces non - linear dynamic adjustment through basis function decomposition, avoiding the static linear assumption of the traditional LSTM forgetting gate, and enhancing the ability to capture long - term dependencies. Dynamically adjust the forgetting weight through KANS to enhance the model's ability to selectively retain historical information.

[0138] (2) Implementation method:

[0139] 1) The KANs sub-network KAN f Receives the global spatial feature The hidden state at the previous moment And the current input To generate a dynamic forgetting bias:

[0140]

[0141] Among them, KAN f Is a two-layer structure: the first layer is the basis function decomposition layer (using a 3rd-order spline function, with 10 grid numbers), and the second layer is the linear transformation layer, with an output dimension of d h (Consistent with )

[0142] 2) The forgetting gate update formula:

[0143]

[0144] Among them, Is the weight matrix, Is the bias, σ is the sigmoid activation function, Is the output of the forgetting gate of node i at time t, indicating the proportion of retaining the memory at the previous moment Of

[0145] (3) Design details:

[0146] KAN f Captures the non-linear spatial dependence in Through basis function decomposition, and the number of parameters is about 10·F·d h (Grid number × input dimension × output dimension)

[0147] During training, W f And b f Are updated through backpropagation, and the parameters of KAN f Are optimized independently

[0148] 2.3 Dynamic weighting of the input gate

[0149] (1) The dynamic weighting mechanism enables the model to adapt to the spatio-temporal changes of the input features, and is more flexible than the fixed weights of the traditional LSTM. The input features are dynamically weighted through KANs to highlight the contribution of key features (such as traffic mutations).

[0150] (2) Implementation method:

[0151] 1) The KANs sub-network KANβ receives the current input To generate a feature importance vector:

[0152]

[0153] Among them, KAN β is a single-layer network that uses a learnable spline function (order 3, number of grids 5), with an output dimension of F and a value range of [0, 1], representing the relative importance of each feature.

[0154] 2) Weighted input calculation:

[0155]

[0156] Among them, ⊙ is element-wise multiplication, is the weighted input feature.

[0157] 3) Input gate update formula:

[0158]

[0159] Among them, is the weight matrix, is the bias, is the output of the input gate, representing the acceptance degree of new information.

[0160] (3) Design details:

[0161] The dynamics of makes the model adapt to feature changes, such as increasing the weight of traffic features during sudden traffic surges. KAN β has a relatively small number of parameters (about 5·F2) and low computational overhead.

[0162] 2.4 Multi-scale memory fusion

[0163] (1) Multi-scale memory fusion dynamically balances short-term fluctuations and long-term periodicity through KANs, and is more adaptable to the complex dynamics of traffic flow than the single time scale of traditional LSTM. By fusing short-term and long-term memories, multi-time scale patterns of traffic flow are captured.

[0164] (2) Implementation method:

[0165] 1) Candidate memory state calculation:

[0166]

[0167] Among them, and are the weights and biases, KAN C is the KANS sub-network that receives spatial features with an output dimension of d h , enhancing the spatial context awareness of the candidate state.

[0168] 2) Long-term memory weight generation:

[0169]

[0170] Among them, KAN γ is a single-layer network, using a spline function (order 3, number of grids 5), and outputs a scalar is the long-term hidden state (Δt is the long-period span, such as 12 steps, representing 1 hour).

[0171] 3) Memory state update:

[0172]

[0173] Among them, is the current memory state, integrating short-term update and long-term memory

[0174] (3) Design details:

[0175] Δt can be adjusted according to the application scenario (for example, set to 6 steps for highway networks and 12 steps for urban road networks). KAN C and KAN γ dynamically adjust memory update through spatial and temporal features, which is more flexible than traditional LSTM.

[0176] 2.5 Output gate optimization

[0177] (1) The output gate introduces spatial feature adjustment to make the hidden state more accurately reflect the road network dynamics. Optimize the hidden state output by combining spatial context to improve the prediction accuracy.

[0178] (2) Implementation method:

[0179] 1) Output gate calculation:

[0180]

[0181] Among them, and are weights and biases, and KAN o is the KANs sub-network, which outputs a dynamic bias with a dimension of d h .

[0182] 2) Hidden state output:

[0183]

[0184] Among them, is the hidden state of node i at time t.

[0185] (3) Technical details: KAN oStructure and KAN f Similar to enhance the sensitivity of the output to spatial dependence.

[0186] 2.6 Parameters and Computational Optimization

[0187] Parameters of KANs: Each sub-network of KANs (KAN f 、KAN β 、KAN c 、KAN γ 、KAN o ) adopts a two-layer structure. The first layer is a spline function (order 3, number of grids 5 - 10), and the second layer is a linear transformation. The total number of parameters accounts for about 25% of LSTM.

[0188] Regularization: Apply L1 regularization (coefficient 0.01) to the weights of KANs to reduce overfitting.

[0189] Computational efficiency: Supports node-level parallel computing, and the time complexity is Step 3: Spatiotemporal Attention Mechanism Fusion

[0190] In this step, through the multi-head attention mechanism and gating fusion technology, the hidden states generated by KA-LSTM and spatiotemporal features are fused to further enhance the model's ability to model the dynamic spatiotemporal relationships of traffic flow. The specific implementation is divided into the following sub-steps.

[0191] 3.1 Input and Target Definition

[0192] (1) Input data: Receive the hidden states output by the KA-LSTM model in Step 2 (the hidden state of each node i at time t, d h is the dimension of the hidden state, such as 128) and spatial features (generated by the graph convolution in Step 2.1, d s is the dimension of the spatial feature, such as 64).

[0193] (2) Design goal: Dynamically capture the spatial dependence and time evolution trend between nodes through the attention mechanism to generate an enhanced hidden state representation Improve the prediction accuracy.

[0194] 3.2 Multi-Head Attention Computation

[0195] (1) Design goal: Use the multi-head attention mechanism to extract the correlation of spatiotemporal features from different sub-spaces and enhance the model's expressive ability.

[0196] (2) Implementation method:

[0197] 1) Query, key, and value generation: Based on the hidden state and spatial features Generate the vectors required for multi-head attention:

[0198]

[0199] Where:

[0200] is the query weight matrix, W K , are the key and value weight matrices respectively;

[0201] d k = d h / H is the dimension of each attention head, and H is the number of attention heads;

[0202] are the query vector of the node and the key and value vectors of node j respectively, and j ∈ {1,...,

[0203] N} traverses all nodes.

[0204] 2) Attention weight calculation: Calculate and normalize the similarity between nodes for each head of attention:

[0205]

[0206] Where:

[0207] h = 1,..., H represents the h-th attention head;

[0208] and are the query and key vectors of the h-th head;

[0209] is the scaling factor to avoid numerical values being too large;

[0210] is the attention weight of node i to node j, reflecting spatial correlation.

[0211] 3) Attention output: Weightedly aggregate the features of all nodes:

[0212]

[0213] Where:

[0214] is the value vector of the h-th head; is the multi-head attention output of the node, generated by concatenating the results of H heads and performing a linear transformation:

[0215]

[0216] Where, is the output weight matrix.

[0217] 3.3 Gated Fusion

[0218] (1) Design Goal: Dynamically balance the original hidden state and the attention output through gating to optimize the spatio-temporal contribution of the final representation.

[0219] (2) Implementation Method:

[0220] 1) Gating Unit Calculation:

[0221]

[0222] Where:

[0223] is the gating weight matrix, is the bias;

[0224] is the concatenation of the hidden state and the attention output;

[0225] is the gating vector, and each dimension independently controls the fusion ratio of the corresponding feature.

[0226] 2) Generation of the Final Hidden State:

[0227]

[0228] Where:

[0229] is the final hidden state of node i at time t; Preserve the original temporal information, Integrate the spatial attention result.

[0230] (3) Technical Details: The dimension design of allows feature-level fusion, which is more refined than scalar gating. The computational complexity of gating is It can be optimized by dimensionality reduction (such as projecting to d h / 2).

[0231] 3.4 Parameter Configuration and Optimization

[0232] (1) Attention Parameters: W Q , W K , W V Are initialized to a normal distribution (mean 0, variance 0.01) and updated through training. H is default 3, d k = d h / H to ensure dimension matching.

[0233] (2) Gating Parameters: W g and b gInitialize to a zero vector to avoid initial bias. Add L2 regularization (coefficient 0.001) to prevent overfitting.

[0234] (3) Optimization strategy: Support sparse pruning of attention weights, set to 0 when , reducing invalid computations. Compute each head of attention in parallel, reducing the time overhead to O(N 2 ·d h / H).

[0235] 3.5 Output and Verification

[0236] (1) Output result: Generate the final hidden state as the input for subsequent predictions.

[0237] (2) Verification method: Check the distribution of attention weights to verify spatial correlation (e.g., higher weights for neighboring nodes). Compare the with prediction performance to ensure the effectiveness of the attention mechanism.

[0238] (3) Technical details: The output dimension d h is the same as that of KA-LSTM to ensure module compatibility. Visualize the attention matrix for debugging and interpretive analysis.

[0239] Step 4: Model Training and Optimization

[0240] In this step, the KA-LSTM model and its spatio-temporal attention mechanism are trained by designing a joint loss function, optimization strategy, and prediction output process to ensure prediction accuracy and model robustness. The specific implementation is divided into the following sub-steps.

[0241] 4.1 Data Preparation and Partitioning

[0242] (1) Input data: Use the standardized three-dimensional spatio-temporal tensor generated in Step 1 (where T is the time step, N is the node dimension, F is the feature dimension, e.g., 3), and the corresponding true traffic labels (traffic values for the next K steps).

[0243] (2) Dataset partitioning: Training set: The first 70% of the total data, used for parameter learning. Validation set: The middle 20%, used for hyperparameter tuning. Test set: The last 10%, used for performance evaluation.

[0244] (3) Technical details: The data is partitioned in chronological order to avoid information leakage. Each batch input is a time series segment (e.g., length 24 steps, 2 hours), and the batch size is default 32.

[0245] 4.2 Loss Function Design

[0246] (1) Design objective: Jointly optimize the prediction error and temporal consistency to improve the accuracy and stability of the model.

[0247] (2) Implementation method

[0248] Adopt a weighted combination of MAE (Mean Absolute Error) and spatio-temporal consistency loss:

[0249]

[0250] Where:

[0251] is the true flow value of node i at time t;

[0252] is the model prediction value;

[0253] The first term is the MAE loss, which measures the prediction accuracy;

[0254] is the spatial feature of node i at time t (generated by step 2.1);

[0255] The second term is the spatio-temporal consistency loss, which constrains the similarity of spatial features of adjacent nodes. is the square of the L2 norm;

[0256] λ is the regularization coefficient, with a default value of 0.1, which balances the contributions of the two losses.

[0257] (3) Technical details:

[0258] 1) ∑ i , j Traverse all node pairs, and through the sparsification of the adjacency matrix A, only calculate adjacent node pairs (reducing the complexity to O(E·d s ))).

[0259] 2) In the MAE calculation, if predicting K steps into the future, it is extended to:

[0260]

[0261] where k = 1,..., K (e.g., K = 12, predicting 1 hour).

[0262] 4.3 Optimization strategy

[0263] (1) Design objective: Through efficient optimization algorithms and regularization means, accelerate convergence and prevent overfitting.

[0264] (2) Implementation method:

[0265] 1) Optimizer: Use the Adam optimizer with the following parameter settings:

[0266] Initial learning rate η = 0.001; momentum parameters β1 = 0.9, β2 = 0.999;

[0267] Exponential decay strategy: The learning rate is halved every 20 epochs, i.e., the learning rate at the e-th epoch is:

[0268]

[0269] 2) Gradient clipping: Apply a norm constraint to all parameter gradients with a threshold of 5.0:

[0270]

[0271] where is the gradient of the loss function with respect to the parameters, to avoid gradient explosion in the KANs sub-network.

[0272] (3) Technical details: 1) The default number of training epochs is 100. If the validation set loss does not decrease for 10 consecutive epochs, training stops early. 2) Add dropout regularization (dropout rate 0.2), applied to the hidden state of KA-LSTM and the attention output. 3) The parameters updated in each iteration include: W f , W i , W c , W o (weights of KA-LSTM), W Q , W K , W V , W g (attention weights) and the parameters of the KANS sub-network.

[0273] 4.4 Prediction Output

[0274] (1) Design goal: Generate predicted traffic values for the next K steps based on the trained model.

[0275] (2) Implementation method: Map the final hidden state to the predicted value through a fully connected layer:

[0276]

[0277] where:

[0278] is the final hidden state output in step 3 (the node dimension has been expanded to N·d h );

[0279] is the output weight matrix, is the bias;

[0280] The predicted traffic volume vector for the next K steps, reshaped by node and time step into

[0281] (2) Technical details:

[0282] 1) If predicting a single step (K = 1), then Output

[0283] 2) Add confidence interval estimation: Based on the predicted residual distribution, calculate the 95% confidence interval:

[0284]

[0285] where σ res is the standard deviation of the validation set residuals.

[0286] 4.5 Training Process and Validation

[0287] (1) Training process:

[0288] 1) Initialize the model parameters (normal distribution, mean 0, variance 0.01).

[0289] 2) Input the training batch, calculate the forward propagation output and the loss L.

[0290] 3) Update the parameters by backpropagation, perform gradient clipping and Adam optimization.

[0291] 4) Evaluate the MAE on the validation set in each epoch and record the best model.

[0292] (2) Validation method:

[0293] 1) Calculate the MAE, RMSE (Root Mean Square Error), and MAPE (Mean Absolute Percentage Error):

[0294]

[0295] 2) Check the spatio-temporal consistency loss to ensure that the difference in features between adjacent nodes is less than a threshold (e.g., 0.5).

[0296] (3) Technical details:

[0297] The training is executed on the GPU (e.g., NVIDIA RTX 4060), and the time consumption per epoch is about 5 seconds (N = 500, T = 24).

[0298] 4.6 Parameter and Optimization Configuration

[0299] (1) Hyperparameters:

[0300] Batch size: 32 (can be adjusted to 64 or 128).

[0301] Hidden state dimension \(d\) h = 128, spatial feature dimension \(d\) s = 64.

[0302] Regularization coefficient \(\lambda = 0.1\) (adjustable range \([0.01, 1.0]\)).

[0303] (2) Optimization details:

[0304] Warm-up stage: The learning rate linearly increases to 0.001 in the first 5 rounds of learning.

[0305] Learning rate scheduler: Cosine annealing is optional, with a period of 20 rounds.

[0306] (3) Computational complexity:

[0307] Single forward propagation (including KA-LSTM and attention mechanism), total training complexity (where \(E\) is the number of epochs and \(I\) is the number of iterations).

[0308] 4.4 Prediction process and output format

[0309] Input the preprocessed spatio-temporal tensor, and based on Predict the future \(K\)-step traffic through the fully connected layer:

[0310]

[0311] Furthermore, output the predicted traffic value and confidence interval (such as 95% confidence level).

[0312] To verify the effectiveness of each module improved and introduced in the present invention on the model improvement, under the same parameter conditions, the present invention conducted the following eight groups of ablation experiments for five innovation points. The results of the ablation experiments and comparative experiments are shown in Tables 1 and 2:

[0313] Table 1 Results of ablation experiments

[0314]

[0315] The basic LSTM is used as the baseline model, and its MAE (15.2 vehicles / minute), RMSE (22.5) and MAPE (12.8%) are all the highest values, exposing the inherent defects of traditional LSTM in complex spatiotemporal modeling. By introducing the spatiotemporal attention mechanism, the model performance is significantly optimized, and the MAE is reduced to 13.6 (a decrease of 10.5%), and the RMSE and MAPE are improved simultaneously, verifying the core role of the attention mechanism in dynamically capturing spatial associations and temporal evolution. After further integrating the dynamic gating optimization of KANs, the MAE is further compressed to 12.8 (a decrease of 15.8% compared with the baseline), and the training time is only slightly increased (4.0 seconds / round), indicating that KANs breaks through the linear assumption of traditional LSTM while taking into account computational efficiency. However, when multi-scale memory fusion is used alone, its improvement on long-term prediction is limited (MAE=14.1), but it has a significant synergy with KANs in the complete model, highlighting its dependence on modules such as dynamic weighting to fully capture the characteristics of multi-time scale features. The application of dynamic weighted input features reduces MAE to 13.3 (a decrease of 12.5%), especially in abnormal scenarios, which reflects its strong adaptability to burst traffic characteristics. The spatiotemporal consistency loss has limited effect when used alone (MAE = 14.5), but in the complete model, by constraining the similarity of adjacent node features, the robustness of the model is significantly improved. It is worth noting that the MAE of KA-LSTM (Experiment 7) with spatiotemporal consistency loss removed rises to 10.5, which is 13.3% higher than the complete model (Experiment 8), confirming the key role of this loss function in generalization ability. Finally, the complete KA-LSTM achieves a MAE as low as 9.1 (a decrease of 40.1% compared with the baseline) through the collaboration of various modules, and the RMSE and MAPE are 14.7 and 7.5% respectively. The comprehensive performance is leading in all aspects, and the training time (5.5 seconds / round) still meets the real-time requirements. This series of improvements shows that through the coordinated design of dynamic gating optimization, multi-scale feature fusion and spatiotemporal constraints, KA-LSTM has achieved comprehensive superiority over traditional and advanced models in terms of accuracy, efficiency and adaptability, providing a better solution for complex traffic scenarios.

[0316] Table 2 Comparative experimental results

[0317]

[0318] The MAE of the traditional linear model ARIMA is as high as 18.7 vehicles / minute, and the RMSE and MAPE are 26.3 and 16.5% respectively, which verifies its serious shortcomings in modeling nonlinear spatiotemporal features. Although its training time is the shortest (0.8 seconds / round), its accuracy cannot meet the needs of complex scenarios. The traditional LSTM reduces the MAE to 15.2 (18.7% lower than ARIMA) through the ability of temporal modeling, but its performance is still limited by the lack of spatial dependency modeling. The CNN-LSTM combined with the local features of CNN space further compresses the MAE to 14.0 (25.1% lower than ARIMA), but the fixed convolution kernel limits its capture of global dynamic spatial associations. The graph convolution-based STGCN significantly improves the MAE to 11.5 (38.5% lower than ARIMA) by explicitly modeling spatial dependencies, but it relies on a static adjacency matrix, is difficult to adapt to dynamic road network changes, and has a long training time (6.3 seconds / round). DCRNN introduces diffuse convolution to enhance spatial propagation modeling, and the MAE is reduced to 10.9 (41.7% lower than ARIMA), but the multi-step iterative calculation leads to a further increase in training time (7.1 seconds / round), which limits real-time performance. Transformer relies on the self-attention mechanism to capture long-term dependencies, with a MAE of 12.3 (34.2% lower than ARIMA), but the high computational complexity (training time 8.5 seconds / round) and insufficient response to short-term burst traffic become bottlenecks. In contrast, the complete KA-LSTM uses KANs to dynamically optimize the gating mechanism, multi-scale memory fusion and spatiotemporal attention collaborative design, with a MAE as low as 9.1 (51.3% lower than ARIMA), RMSE and MAPE of 14.7 and 7.5% respectively, and its comprehensive performance is leading in all aspects. Its training time (5.5 seconds / round) is significantly better than STGCN, DCRNN and Transformer, reflecting the efficiency advantage of KANs lightweight design. The dynamic adaptability of KA-LSTM is particularly prominent. Through dynamic feature weighting and spatiotemporal consistency loss, it can maintain stable predictions in abnormal events and multimodal data scenarios. Ultimately, it achieves a systematic surpassing of traditional and advanced models in terms of accuracy, efficiency and scenario adaptability, providing a better solution for highly dynamic urban traffic management.

[0319] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with the technology to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit of the present invention should be included in the present invention.

Claims

1. A traffic flow prediction method based on improved spatio-temporal coupled KA-LSTM, characterized in that, It includes the following steps: Step (1) Data collection and preprocessing: Obtain data and process the data to generate a standardized three-dimensional spatio-temporal tensor; Step (2) Improved LSTM structure: Use the Kolmogorov-Arnold network to dynamically optimize the forget gate, input gate, output gate, and memory update mechanism of the long short-term memory network to generate a spatio-temporal feature weighted representation; Step (3) Spatio-temporal attention mechanism fusion: Through multi-head attention calculation and gated fusion technology, fuse multi-scale spatio-temporal features to generate an enhanced hidden state representation; Step (4) Model training and prediction: Use a fully connected layer to train the model jointly with a loss function, and use the trained model to predict the traffic flow in the future time period.

2. The traffic flow prediction method based on improved spatio-temporal coupled KA-LSTM according to claim 1, characterized in that: The data collection and preprocessing in step (1) specifically include: (1.1) Collect traffic flow, speed, and density data within the time window T from N road network nodes, with a time granularity of 5 minutes; (1.2) Complement missing data using the spatio-temporal weighted interpolation method, and the calculation formula is: Among them, is the complement value of node i at time t, w ij is the spatial distance weight based on the Gaussian kernel function, α is the spatio-temporal balance factor, and its value range is [0, 1]; (1.3) Remove outliers using the Z-score method, with a threshold of 3; (1.4) Perform Z-score standardization on each feature dimension and output a three-dimensional spatio-temporal tensor where F is the number of features.

3. The traffic flow prediction method based on improved spatio-temporal coupled KA-LSTM according to claim 1, characterized in that: The improved LSTM structure in step (2) specifically includes: (2.1) Based on the road network adjacency matrix and the current input X t , generate global spatial features through graph convolution: Among them, is the adjacency matrix with self-loops, is the degree matrix, is the trainable weight, d s is the spatial feature dimension (such as 64), which is consistent with the input of subsequent KANs. (2.2) Improvement of the forget gate: Generate a dynamic forget weight through KANs and fuse it into the forget gate to obtain a new forget gate, and the calculation formula is: Among them, is the output of the forgetting gate of node i at time t, W f and b f are the weight matrix and bias respectively, is the output of the KANs sub-network based on the global spatial feature ; (2.3) Dynamic weighting of the input gate: Generate a feature weighted vector through KANs, and the calculation formula is: Fuse it into the input gate and update it to: Among them, is the feature weighted vector, ⊙ is element-wise multiplication, is the weighted input feature, is the weight matrix, is the bias, is the input gate output, indicating the acceptance degree of new information; (2.4) Multi-scale Memory Fusion: Global Spatial Features Through the KANs sub-network KAN C The enhanced candidate memory vectors generated Generate new candidate memory states: Furthermore, the long-term hidden state generates a long-term memory weight that dynamically balances short-term fluctuations and long-term periodicity through the sub-network KAN of KANs γ : Furthermore, update the memory state formula to Among them, is the current memory state, integrating short-term updates and long-term memory (2.5) Output Gate Optimization: KAN Sub-network KAN f Receive global spatial features Generate dynamic forgetting biases Fuse into the output gate to obtain: Among them, and are the weights and biases, and KAN o is the KANs sub-network, which outputs a dynamic bias with a dimension of d h . (2.5) Hidden state output: Among them, is the hidden state of node i at time t.

4. The traffic flow prediction method based on the improved spatio-temporal coupled KA-LSTM according to claim 1, characterized in that: The spatio-temporal attention mechanism fusion in step (3) specifically includes: (3.1) Multi-head attention calculation: Generate query, key, and value vectors based on the hidden state and spatial features, and the calculation formula is: Among them, H is the number of attention heads, and d k is the key vector dimension; (3.2) Gated fusion: Balance the state and attention output through a gated unit to generate a gated unit and the final hidden state, and the calculation formula is: Among them, is the final output hidden state, is the gating unit.

5. The traffic flow prediction method based on the improved spatio-temporal coupled KA-LSTM according to claim 1, characterized in that: The model training and prediction in step (4) specifically include: (4.1) Loss function design: Jointly optimize using MAE and spatio-temporal consistency loss, and the calculation formula is: Among them, is the true value, is the predicted value, and λ is the regularization coefficient; (4.2) Optimization strategy: Use the Adam optimizer, with an initial learning rate of 0.001, exponential decay, halving every 20 epochs, and a gradient clipping threshold of 5.0; (4.3) Prediction output: Generate future K-step prediction values through a fully connected layer, and the calculation formula is: Where: is the final hidden state; is the output weight matrix, is the bias.

6. A traffic flow prediction system based on an improved spatio-temporal coupled KA-LSTM, characterized in that, It includes: (6.1) Data preprocessing module, used to collect and process road network traffic data to generate a spatio-temporal tensor; (6.2) KA-LSTM prediction module, based on the LSTM structure improved by KANs and the spatio-temporal attention mechanism, generates traffic flow prediction results; (6.3) Attention weight distribution recognition module, based on the attention weight distribution, recognizes traffic anomalies and adjusts the prediction strategy; (6.4) Real-time optimization interface module, outputs the prediction results to the traffic management system to support signal timing optimization.

7. The traffic flow prediction system based on the improved spatio-temporal coupled KA-LSTM according to claim 6, wherein The data preprocessing module includes: (7.1) Data collection sub-module, automatically collects traffic flow, speed, and density data from road network sensors, with a time granularity of 5 minutes; (7.2) Missing value processing sub-module, uses the spatio-temporal weighted interpolation method to complement missing data; (7.3) Outlier detection sub-module, which eliminates outliers based on the Z-score method with a threshold set to 3; (7.4) Data normalization sub-module, which performs Z-score normalization on each feature dimension to generate a three-dimensional spatio-temporal tensor.

8. The traffic flow prediction system based on the improved spatio-temporal coupled KA-LSTM according to claim 6, wherein The KA-LSTM prediction module includes: (8.1) Forget gate sub-module, which generates dynamic forgetting weights using KANs; (8.2) Input gate sub-module, which realizes dynamic feature weighting through KANs; (8.3) Multi-scale memory sub-module, which fuses short-term and long-term features to update the memory state; (8.4) Spatio-temporal attention sub-module, which optimizes the hidden state output based on multi-head attention and gating fusion.

9. The traffic flow prediction system based on the improved spatio-temporal coupled KA-LSTM according to claim 6, characterized in that, The attention weight distribution recognition module includes: (9.1) Anomaly recognition sub-module, which recognizes traffic anomaly events based on the attention weight distribution; (9.2) Prediction strategy adjustment sub-module, which automatically adjusts the prediction strategy according to the anomaly situation to ensure the reliability and timeliness of the prediction results.

10. The traffic flow prediction system based on the improved spatio-temporal coupled KA-LSTM according to claim 6, wherein, The real-time optimization interface module includes: (10.1) Prediction result output sub-module, which transmits the prediction results to the traffic management system to support functions such as signal timing optimization; (10.2) Real-time optimization sub-module, which makes real-time traffic management decisions based on the prediction results.

Citation Information

Patent Citations

  • Traffic flow prediction method and system based on trend space-time diagram convolution, and medium

    CN116895157A

  • Traffic flow prediction method and system based on time-space synchronization dynamic graph attention network

    CN117671952A

  • Traffic flow prediction method based on time-varying fusion graph convolutional network

    CN118262517A

  • Short-term traffic flow prediction method based on causal gated-low-pass graph convolutional network

    US20240029556A1

Cited By

  • Molecular ratio optimization control model and prediction method under complex electrolyte system

    CN120808927A

  • Safety production standardization integrated management system and method

    CN121052811A

  • Multi-strategy improved traffic flow prediction method, equipment and medium

    CN121148156A

  • Intersection turning movement flow prediction method based on weight pre-balanced multivariate fusion

    CN122738213A

  • Integrated Data Collection System and Control Method for Road Risk Prediction Based on Fusion of GIS, Meteorological, and Traffic Data

    KR103007714B1