Traffic flow spatial-temporal feature adaptive extraction method based on dynamic Kolmogorov-Arnold network

Through the adaptive extraction method of the space-time feature of the traffic flow in dynamic Kolmogorov-Arnold network, the rigidity and delay problems of existing traffic flow prediction methods in the spatiotemporal coupled feature processing are solved, and efficient and real-time traffic flow prediction and decision support are achieved.

CN120164326AActive Publication Date: 2025-06-17HUAIYIN INSTITUTE OF TECHNOLOGY

Patent Information

Application Number
CN202510334572.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-17
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

When dealing with dynamic spatiotemporal features, existing traffic flow prediction methods have problems such as space-time coupling feature loss, model rigidity, and excessive smoothing, which is difficult to adapt to complex traffic scenarios, and the inference delay is high, which cannot meet the needs of real-time traffic control systems.

Method used

Adaptive extraction method of space-time features of traffic flow based on dynamic Kolmogorov-Arnold network is adopted. By dynamically adjusting the network structure and feature extraction granularity, combining graph convolution and timing convolution, the adaptive extraction and fusion of space-time features are realized, and prediction is performed using a lightweight decoder.

Benefits of technology

Improves the accuracy and robustness of traffic flow prediction, reduces inference delay, meets the needs of real-time traffic control systems, and supports auxiliary traffic management decisions through visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164326A_ABST
    Figure CN120164326A_ABST
Patent Text Reader

Abstract

The invention provides a traffic flow spatio-temporal feature adaptive extraction method based on a dynamic Kolmogorov-Arnold network, and aims to solve the defects of a traditional model in the aspects of dynamic spatio-temporal modeling, structural adaptability and feature expression efficiency. According to the method, a network structure is dynamically adjusted through differential topology search and a parameterized primary function library, a double-flow coupling architecture is designed to extract space and time dependent features respectively, and adaptive fusion of spatial and temporal features is realized by using a gating mechanism. The method comprises the following specific steps: performing space-time normalization and graph structure coding to generate a node feature matrix; extracting multi-scale spatial-temporal features in parallel by the dynamic graph convolution KANs and the time sequence convolution KANs; feature alignment and cooperative enhancement are realized based on a local correlation matrix and a bidirectional interactive attention mechanism; and the lightweight KAN decoder is combined with the dynamic basis function library to output a prediction result. Experiments show that the MAE is reduced to 1.78 vehicles per minute and the reasoning speed is improved by 1.8 times in a traffic accident emergency scene on a PeMS08 data set, and the method is suitable for an intelligent traffic management and vehicle-road cooperation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent transportation systems, and particularly relates to a method for adaptively extracting spatio-temporal characteristics of traffic flow based on a dynamic Kolmogorov-Arnold network. Background Art

[0002] With the acceleration of urbanization and the rapid development of intelligent transportation systems (ITS), traffic flow prediction has become one of the core technologies for optimizing traffic management, alleviating congestion, and enhancing road safety. However, traffic flow data has highly dynamic spatio-temporal coupling characteristics: in the spatial dimension, the correlation between nodes in the road network changes in real time with traffic states (such as peak hours, sudden accidents); in the time dimension, traffic flow exhibits multi-scale dependencies (such as minute-level fluctuations, hourly cycles, daily trends). Existing methods face the following key challenges when dealing with such complex spatio-temporal characteristics:

[0003] Traditional deep learning methods (such as convolutional neural network CNN, recurrent neural network RNN) need to process spatial and temporal features separately. For example, CNN extracts local spatial patterns through fixed convolutional kernels, but it is difficult to model the dynamic road network topology; although RNNs (such as LSTM, GRU) can capture time series dependencies, they ignore spatial correlations. Such separate modeling leads to the loss of spatio-temporal coupling characteristics of sudden traffic events (such as traffic accidents, temporary control), and the prediction error increases significantly. Research shows that on the PeMS dataset, the mean absolute error (MAE) of such models for sudden event scenarios is as high as 2.34 vehicles / minute, and the response delay is obvious.

[0004] Although classical Kolmogorov-Arnold networks (KANs) have the theoretical universal approximation ability, their fixed hierarchical structure and predefined activation functions (such as ReLU) are difficult to adapt to the spatio-temporal heterogeneity of traffic flow data. For example, the traffic patterns during morning and evening rush hours are significantly different from those during off-peak hours, and static networks cannot dynamically adjust the feature extraction granularity, resulting in insufficient modeling ability for long-tail distribution features (such as extreme congestion). In addition, existing graph neural network (GNN)-based models rely on static adjacency matrices (such as those constructed based on geographical distances) and cannot reflect the dynamic spatial dependencies under real-time traffic conditions (such as temporary path associations caused by accidents), severely restricting the generalization performance of the models.

[0005] The multi-layer GNN model aggregates neighbor information through message passing. However, the over-smoothing problem causes the node features in the deep network to tend to be homogenized, and the local fine-grained features (such as the short-term traffic flow mutation at intersections) are submerged. Experiments show that when the number of GNN layers exceeds 3 layers, the similarity of node features increases by more than 40%, directly affecting the model's ability to capture complex spatial dependencies in the dynamic road network. In addition, existing methods mostly adopt fixed-order graph convolution (such as 2-hop neighbors), making it difficult to adaptively adjust the spatial perception range, resulting in an imbalance in the trade-off between global dependencies and local features.

[0006] Existing methods mostly rely on black-box deep networks, lacking visualization support for the spatio-temporal feature propagation path, and it is difficult to assist traffic management decisions. At the same time, complex model structures (such as multi-layer stacked GCN+LSTM) lead to relatively high inference latency (for example, the inference time of the DCRNN model reaches 15.8 milliseconds per step), making it difficult to meet the low-latency requirements of real-time traffic control systems.

[0007] In summary, existing traffic flow prediction methods have significant defects in dynamic spatio-temporal modeling, structural adaptability, feature expression efficiency, and interpretability. Therefore, there is an urgent need for a new method that can dynamically adjust the network structure and feature extraction granularity and achieve spatio-temporal joint modeling to improve the prediction accuracy, efficiency, and decision-making support ability in complex traffic scenarios. Summary of the Invention

[0008] Objective of the Invention: Aiming at the above problems, the present invention provides a method for adaptively extracting spatio-temporal features of traffic flow based on dynamic KANs, which solves the problems of rigid traditional model structures, spatio-temporal feature separation, and over-smoothing, and improves the prediction accuracy and robustness in complex traffic scenarios.

[0009] Technical Solution: The present invention discloses a method for adaptively extracting spatio-temporal features of traffic flow based on dynamic Kolmogorov-Arnold networks (hereinafter referred to as KANs) (hereinafter referred to as DKAN-TF), including the following steps:

[0010] Step 1: Perform spatio-temporal normalization and graph structure encoding on the input data to generate a node feature matrix;

[0011] Step 2: Extract spatial and temporal features through dynamic KANs respectively;

[0012] Step 3: Dynamically adaptive feature alignment and spatio-temporal feature fusion with collaborative enhancement;

[0013] Step 4: Output the prediction result through a lightweight KANs decoder.

[0014] Further, in step 1, the collected traffic time series data is processed by zero-mean spatio-temporal normalization, and the missing values are filled using linear interpolation.

[0015] Further, in step 1, based on the topological structure of the traffic network, a static adjacency matrix is constructed The intermediate feature matrix of the normalized traffic time series data is combined with the static adjacency matrix to generate an initial node feature matrix

[0016] Further, in step 2, based on graph convolutional (GCN) dynamic KANs (GC-KANs), a spatial flow is constructed. According to the node feature matrix X of the current layer (l) and data of other layers, a dynamic adjacency matrix A is generated raw′ = W2·Z + b2. For A raw′ perform row-wise Softmax normalization to endow the edge weight with probabilistic meaning:

[0017] A dynamic = Softmax(A raw′ )

[0018] A dynamic (i, j) ∈ A dynamic represents the dynamic connection strength from node i to node j, and ∑ j A dynamic (i, j) = 1.

[0019] If a symmetric adjacency matrix is required (for undirected graphs), A raw can be symmetrized first:

[0020]

[0021] Further, in step 2, based on graph convolutional (GCN) dynamic KANs (GC-KANs), a spatial flow is constructed. Using the dynamic adjacency matrix A dynamic perform multi-order graph convolution to capture dependencies at different spatial scales:

[0022]

[0023] where is the input feature of the l-th layer (the current node feature matrix); (A dynamic ) k represents the k-th power of the adjacency matrix, capturing the influence of k-hop neighbors; is the convolutional weight matrix of the k-th order; K is the maximum order (e.g., K = 2 means considering two-hop neighbors).

[0024] Furthermore, in step 2, a spatial flow is constructed based on graph convolutional dynamic KANs (GC-KANs), and a gating mechanism is introduced to dynamically select the convolution order:

[0025] α k = Sigmoid(W gate ·Mean(X (l) ) + b gate )

[0026]

[0027] where α k is the contribution weight of the k-th order, W gate , b gate are the parameters of the gating unit, and Mean represents taking the average along the time dimension. This mechanism adaptively adjusts the convolution range according to the spatial characteristics of the input data.

[0028] Furthermore, in step 2, a spatial flow is constructed based on graph convolutional (GCN) dynamic KANs (GC-KANs). The traditional activation function is replaced by a library of basis functions characteristic of KANs, including B-spline functions, sine functions, Bessel functions, etc. The selection and combination of the basis functions are driven by the task:

[0029]

[0030] where φ i is the i-th basis function (e.g., φ1(x) = spline(x), φ2(x) = sin(x)); β i is the trainable coefficient representing the contribution of each basis function; M is the number of basis functions (e.g., M = 3).

[0031] Furthermore, in step 2, a spatial flow is constructed based on graph convolutional (GCN) dynamic KANs (GC-KANs). A lightweight selection network Ψ (single-layer MLP) is used to dynamically adjust βi according to the input of the current layer:

[0032]

[0033] Furthermore, in step 2, a spatial flow is constructed based on graph convolutional (GCN) dynamic KANs (GC-KANs). To ensure training stability, layer normalization is applied after each layer of convolution:

[0034]

[0035] Furthermore, in step 2, a temporal flow is constructed based on temporal convolutional (TCN) dynamic KANs (TC-KANs). Dilated convolutions are used to capture long- and short-term dependencies. The dilation coefficient d determines the receptive field of the convolution:

[0036] d = Round(Sigmoid(W d ·X (0) + b d )·d max )

[0037] where W d , b d are trainable parameters; d max is the maximum dilation coefficient (e.g., d max = 8); Round discretizes continuous values into integers. This dynamic selection mechanism adjusts the receptive field size according to the temporal characteristics of the input data (such as short-term fluctuations during peak periods or long-term trends during off-peak periods).

[0038] Furthermore, in step 2, a time flow is constructed based on Temporal Convolutional Networks (TCN) Dynamic KANs (TC-KANs), and one-dimensional dilated convolution is performed on the time dimension:

[0039]

[0040] where X (l) ∈ R N×T×D is the input of the l-th layer (the current node feature matrix); W TC ∈ R k×D×D is the convolutional kernel weight, k is the convolutional kernel size (e.g., k = 3); d is the dilation coefficient of the current layer; s is the stride (usually set to 1).

[0041] By stacking 3 convolutions, the receptive field is gradually enlarged, and the dilation coefficient of each layer increases exponentially (e.g., d = 1, 2, 4), but is constrained by the dynamic selection mechanism.

[0042] Furthermore, in step 2, a time flow is constructed based on Temporal Convolutional Networks (TCN) Dynamic KANs (TC-KANs). Similar to the spatial flow, a parametric basis function library is used to perform a non-linear transformation on the convolutional output:

[0043]

[0044] where ψ i is the basis function (e.g., ψ1(x) = spline(x), ψ2(x) = cos(x)), η i is the trainable coefficient.

[0045] Furthermore, in step 2, a time flow is constructed based on Temporal Convolutional Networks (TCN) Dynamic KANs (TC-KANs), and a selection network Ω (single-layer MLP) is used to adjust η i :

[0046]

[0047] Furthermore, in step 2, a time flow is constructed based on Temporal Convolutional Networks (TCN) with Dynamic KANs (TC-KANs). To preserve the original time information, a residual connection is added:

[0048]

[0049] Furthermore, in step 2, a time flow is constructed based on Temporal Convolutional Networks (TCN) with Dynamic KANs (TC-KANs), and layer normalization is applied:

[0050]

[0051] Furthermore, in step 2, a time flow is constructed based on Temporal Convolutional Networks (TCN) with Dynamic KANs (TC-KANs). After processing by the spatial flow and the time flow, spatial features and temporal features are obtained respectively, providing inputs for subsequent spatio-temporal feature fusion.

[0052] Furthermore, in step 3, a fusion strategy based on dynamic adaptive feature alignment and collaborative enhancement is proposed. By adaptively adjusting the correlation of spatio-temporal features and enhancing their complementarity, efficient fusion is achieved.

[0053] Furthermore, in step 3, based on the fusion strategy of dynamic adaptive feature alignment and collaborative enhancement, a local correlation matrix between the spatial features and the temporal features is calculated to measure the coupling degree between the two at each node and time step:

[0054] C i,t = CosineSimilarity(H spatial,i,t , H temporal,i,t )

[0055] where H spatial,i,t and H temporal,i,t are the spatial and temporal feature vectors of node i at time t respectively. Cosine similarity is used to capture the consistency of feature directions.

[0056] Furthermore, in step 3, based on the fusion strategy of dynamic adaptive feature alignment and collaborative enhancement, the alignment transformation uses a lightweight transformation network Φ (consisting of two layers of MLP). According to the correlation matrix, the spatial and temporal features are dynamically adjusted to align them in the semantic space:

[0057] H′ spatial = Φ spatial (H spatial , C) = W sp ·H spatial ·Sigmoid(C) + b sp

[0058] H′ temporal = Φ temporal (H temporal , C) = W tp ·H temporal ·Sigmoid(C) + b tp

[0059] Among them, W sp , b sp , W tp , b tp are trainable parameters, and Sigmoid(C) serves as a gating factor to control the alignment strength. This alignment mechanism ensures that spatial and temporal features remain consistent in high-correlation regions and retain independence in low-correlation regions.

[0060] Furthermore, in step 3, based on the fusion strategy of dynamic adaptive feature alignment and collaborative enhancement, a bidirectional interactive attention mechanism is introduced to enhance the complementary information flow between spatial and temporal features:

[0061]

[0062] Among them, W Q and W K are projection matrices for queries and keys, and D is the feature dimension.

[0063] Furthermore, in step 3, based on the fusion strategy of dynamic adaptive feature alignment and collaborative enhancement, update the features based on the attention weights:

[0064] H″ spatial = H′ spatial + A tp-to-sp ·(W V ·H′ temporal )

[0065] H″ temporal = H′ temporal + A sp-to-tp ·(W V ·H′ spatial )

[0066] Among them, W V is the value projection matrix. This bidirectional interaction enables spatial features to absorb temporal dependency information and temporal features to incorporate spatial topological information.

[0067] Furthermore, in step 3, based on the fusion strategy of dynamic adaptive feature alignment and collaborative enhancement, further enhance the expressiveness of the fused features through residual connection and basis function enhancement:

[0068]

[0069] Among them, φ i is a basis function (such as a spline function), γ i is a trainable coefficient, and LayerNorm ensures numerical stability.

[0070] Furthermore, in step 3, based on the fusion strategy of dynamic adaptive feature alignment and collaborative enhancement, the final fused feature incorporates the synergistic effect of spatio-temporal information.

[0071] Furthermore, in step 4, the spatio-temporal fused feature (where N is the number of nodes, T is the time step, and D is the feature dimension) generated in step 3 is decoded into the traffic flow prediction results for the next T' time steps (F is the original feature dimension, such as flow rate, speed, etc.).

[0072] Furthermore, in step 4, the input fused feature H fused has a relatively high hidden dimension D (for example, D = 64). To reduce the subsequent computational burden, a linear transformation is used for dimensionality reduction:

[0073]

[0074] Among them, is the dimensionality reduction weight matrix, D' < D (for example, D' = 16); is the bias vector; the output

[0075] Furthermore, in step 4, to generate the prediction for the next T' time steps, a lightweight fully connected layer is used to expand the time dimension:

[0076]

[0077] Among them, is the time expansion matrix, and T' is the number of prediction time steps (for example, T' = 12); is the bias; the output

[0078] Furthermore, in step 4, layer normalization is performed on the dimensionally reduced and expanded features to ensure numerical stability:

[0079]

[0080] Furthermore, in step 4, using the basis function property of KANs, a non-linear transformation is performed on the dimensionally reduced features to enhance the expressive power of the decoder. Define a parameterized basis function library {ψ1, ψ2,..., ψ M}(For example, ψ1(x) =}spline(x), ψ2(x) = sin(x), ψ3(x) = Bessel(x)), each basis function corresponds to a different type of non - linear pattern:

[0081]

[0082] Among them, λ i is the contribution coefficient of the i - th basis function, and the initial value is a trainable parameter; M is the number of basis functions (for example, M = 3); the output

[0083] Furthermore, in step 4, to make the basis functions adapt to different traffic flow prediction scenarios (such as short - term fluctuations or long - term trends), a dynamic weight generator Γ (a lightweight MLP) is introduced to adaptively adjust λ i :

[0084]

[0085] Among them, W Γ1 ,W Γ ,b Γ1 ,b Γ are MLP parameters.

[0086] Furthermore, in step 4, to retain the original feature information, a residual connection is added:

[0087]

[0088] Furthermore, in step 4, the decoded features are mapped to the target feature dimension F (for example, F = 2 represents flow and speed):

[0089]

[0090] Among them, is the output mapping matrix; is the bias; the output is the normalized prediction result.

[0091] Furthermore, in step 4, a dynamic adjustment module is introduced to adjust the output according to the historical prediction error. Define an error feedback vector (initially zero, updated during training), and correct the prediction through a gated unit:

[0092]

[0093] Among them, W gate ,b gate are the gated parameters; G is the adjustment weight, controlling the intensity of the error feedback.

[0094] Further, in step 4, the normalized prediction result Y norm is converted back to the original dimension:

[0095] Y pred = Y norm ·σ + μ

[0096] where μ and σ are the mean and standard deviation calculated in step 1.

[0097] Further, in step 4, by reducing the dimension (D' < D) and using a small number of basis functions (M = 3), the number of parameters of the decoder is significantly reduced. The single-layer basis function transformation and the dynamic adjustment module avoid the high computational cost of the multi-layer stacked network and are suitable for the real-time traffic prediction scenario.

[0098] Further, in step 4, the final output represents the traffic flow prediction values (such as flow, speed, etc.) for the future T' time steps and can be directly used for downstream tasks (such as congestion analysis or path planning).

[0099] Beneficial effects:

[0100] The traffic flow spatio-temporal feature adaptive extraction method based on the dynamic Kolmogorov - Arnold network proposed by the present invention has the following significant advantages:

[0101] (1) By generating a dynamic adjacency matrix in real time through a hypernetwork, the dynamic spatial dependencies in sudden traffic events are captured; combined with a gating mechanism, the graph convolution order is adaptively selected to balance the local and global feature perception ranges. In the sudden traffic event scenario of the PeMS08 dataset, the MAE drops to 1.78 vehicles / minute, a 23.7% reduction compared to traditional models (such as STGCN, DCRNN), and the RMSE and MAPE are optimized to 3.48 and 5.12% respectively.

[0102] (2) By dynamically simplifying the redundant connections of the network through differentiable topological search, combined with dimension reduction operations and single-layer basis function transformation, the computational complexity is reduced. The inference speed is increased by 1.8 times, and the single-step inference time is only 9.5 milliseconds, a 40.5% improvement compared to DCRNN (15.8 milliseconds / step), meeting the requirements of the real-time traffic control system.

[0103] (3) The adaptive combination of basis functions such as B-spline and sine is adopted to capture the non-linear patterns of traffic flow; the time receptive field is adjusted by dynamically dilating the convolution coefficient to adapt to the multi-scale temporal dependencies during peak / off-peak periods. In the traffic pattern mutation scenario (such as accidents, temporary control), the model robustness is improved, and the prediction error fluctuation is reduced by 18.3%.

[0104] (4) The heatmap of basis function contributions and the bidirectional interactive attention mechanism are introduced to reveal the spatio-temporal feature propagation paths; the spatio-temporal feature complementarity is quantified through attention weights. It assists traffic managers in analyzing the congestion evolution law, improves the transparency of model decision-making, and increases the accuracy of abnormal event location by 32%.

[0105] (5) The error feedback mechanism and the collaborative enhancement fusion strategy are combined: the historical error is used to dynamically correct the output to alleviate the problem of cumulative error; the semantic alignment and information complementarity of spatio-temporal features are enhanced through the bidirectional attention mechanism. In the scenario of cross-regional and multi-sensor heterogeneous data, the volatility of the model's MAE is reduced by 12.6%, the training time is shortened by 18.5% compared with the traditional KAN, and the convergence stability of the model is significantly improved. Description of the Drawings

[0106] Figure 1 : Overall flow chart of the adaptive spatio-temporal feature extraction method for traffic flow based on the dynamic Kolmogorov-Arnold network;

[0107] Figure 2 : Structure diagram of the dynamic graph convolutional KANs;

[0108] Figure 3 : Structure diagram of the dynamic temporal convolutional KANs;

[0109] Figure 4 : Flow chart of spatio-temporal feature fusion;

[0110] Figure 5 : Structure diagram of the lightweight KAN decoder;

[0111] Figure 6 : Column comparison chart of ablation experiment evaluation indicators;

[0112] Figure 7 : Column comparison chart of comparative experiment evaluation indicators;

[0113] Figure 8 : Result graph of the ablation experiment;

[0114] Figure 9 : Result graph of the comparative experiment.

[0115] Specific Embodiment Examples

[0116] For a better understanding of the present invention, the present invention will be further described below in conjunction with the drawings in the embodiments of the present invention, but it is not a limitation of the present invention. Without departing from the design concept of the present invention, various variations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention.

[0117] Such as Figure 1As shown in the figure, an adaptive extraction method for spatio-temporal characteristics of traffic flow based on a dynamic Kolmogorov-Arnold network is as follows:

[0118] Taking the traffic volume data of one day in the PeMS08 dataset (selecting the data of July 1, 2016) as an example, the implementation process of the adaptive extraction method for spatio-temporal characteristics of traffic flow based on dynamic Kolmogorov-Arnold networks (KANs) is described in detail. The implementation method includes the complete processes of data processing, feature extraction, fusion, and prediction output, strictly following the technical solutions of claims 1-6 of the patent. It is only for illustrative purposes and does not limit the protection scope of the present invention.

[0119] Step 1: Perform spatio-temporal normalization and graph structure encoding on the input data to generate a node feature matrix;

[0120] (1) Data preprocessing and normalization

[0121] The PeMS08 dataset records the traffic flow in a certain highway area in California, including 170 detection nodes (sensors), and the sampling frequency is once every 5 minutes. Assuming that we select the data of July 1, 2016, a total of 24×12 = 288 time steps are generated in 24 hours of a day. Therefore, the dimension of the selected input data is where N = 170 (the number of nodes, that is, 170 sensors); T = 288 (the number of time steps in a day, 24 hours x 12 sampling points per hour); F = 1 (the feature dimension, only traffic flow, unit: vehicles / 5 minutes).

[0122] For the traffic time series data of one day in the input PeMS08 dataset Perform spatio-temporal normalization processing to eliminate the dimension difference and enhance the model stability. Specifically, the spatio-temporal normalization method is adopted:

[0123]

[0124] where X raw (i, t, k): the k-th eigenvalue of node i at time t; μ k , σ k : the mean and standard deviation of the k-th feature over all nodes and times. The intermediate feature matrix of the normalized traffic time series data

[0125] Integrate the topological information of the static adjacency matrix A static into the intermediate features of the traffic time series data,

[0126] (1) Node degree: Degree(i) = ∑ j A static (i, j), indicating the number of connections of node i;

[0127] (2) Neighbor average feature: Based on A static Weighted aggregation of neighbors' X temp :

[0128]

[0129] (3) Feature concatenation:

[0130] X (0) = [X norm (i, t, k), Degree(i), X static (i, :)]

[0131] Concatenate the intermediate feature matrix of traffic time series data with the static graph features to form the initial node feature matrix F = F temp + F static ; F temp : Time feature dimension; F static : Static graph feature dimension.

[0132] Step 2: Input the current layer node feature matrix X (l) into the KANs network to generate a dynamic adjacency matrix. Replace the traditional activation function of GCN with the KANs basis function library to form a new GCN network. Input the dynamic adjacency matrix into the new GCN network to output spatial features. Replace the traditional activation function of TCN with the KANs basis function library and introduce the initial node feature matrix to improve the dilation coefficient to form a new TCN network. Input the initial node feature matrix into the new TCN network to output time features;

[0133] Among them: (1) Spatial flow: Based on graph convolution (GCN) dynamic KANs (GC-KANs)

[0134] The purpose of the spatial flow is to extract the spatial correlation features between nodes from the topological structure of the traffic network, which is realized by using graph convolution (GCN) operations combined with a dynamic adjustment mechanism.

[0135] 1) Dynamic adjacency matrix adjustment

[0136] Generate a dynamic adjacency matrix according to the current layer node feature matrix X (l) and data of other layers:

[0137] The KANs hypernetwork is usually a lightweight two-layer multi-layer perceptron (MLP) with learnable basis functions on its edges: The first layer: Map the static adjacency matrix a to an intermediate representation. The second layer: Output the unnormalized adjacency matrix A raw′ .

[0138] The mathematical formula is as follows:

[0139] Z = σ(W1·X (l) + b1)

[0140] A raw′ = W2·Z + b2

[0141] Where: W1 and b1 are the weights and biases of the first layer; σ is a non - linear activation function, which may be a learnable spline function in KANs rather than a fixed function such as ReLU; W2 and b2 are the weights and biases of the second layer; is the unnormalized adjacency matrix.

[0142] To ensure that the adjacency matrix is applicable to GCN and interpretable, perform row - by - row Softmax normalization on A raw to give the edge weights a probabilistic meaning:

[0143] A dynamic = Softmax(A raw′ )

[0144] A dynamic (i, j) represents the dynamic connection strength from node i to node j, and ∑ j A dynamic (i, j) = 1.

[0145] If a symmetric adjacency matrix (for undirected graphs) is required, A raw can be symmetrized first:

[0146]

[0147] The KANs hyper - network processes X at each time step or each layer of the model (e.g., during forward propagation). (l) . Since traffic data is usually streamed as input, X (l) is continuously updated, and A dynamic is recalculated in real - time accordingly.

[0148] The lightweight design of the KANs hyper - network (e.g., two - layer MLP with a moderate hidden layer dimension) ensures low computational overhead and is suitable for real - time applications (e.g., 9.5 milliseconds / step as mentioned in the patent).

[0149] Softmax ensures the normalization of the adjacency matrix, and the output This dynamic adjustment enables the adjacency matrix to reflect changes in real - time traffic patterns, such as strong correlations on certain sections during peak hours.

[0150] 2) Multi - order graph convolution (GCN) operation

[0151] Using the dynamic adjacency matrix A dynamicExecute multi - order graph convolutional network (GCN) to capture dependencies in different spatial scopes:

[0152]

[0153] Among them, is the input feature of the l - th layer (the current node feature matrix); (A dynamic ) k represents the power of the k - th order adjacency matrix, capturing the influence of k - hop neighbors; is the convolutional weight matrix of the k - th order; K is the maximum order (K = 2 means considering two - hop neighbors).

[0154] To avoid redundant calculations caused by a fixed order, a gating mechanism is introduced to dynamically select the convolutional order:

[0155] α k = Sigmoid(W gate ·Mean(X (l) )+b gate )

[0156]

[0157] Among them, α k is the contribution weight of the k - th order, W gate , b gate are the parameters of the gating unit, and Mean represents taking the average over 288 time steps along the time dimension. This mechanism adaptively adjusts the convolutional range according to the spatial characteristics of the input data.

[0158] 3) Activation of basis functions and enhancement of non - linearity

[0159] Traditional activation functions are replaced by a basis function library characteristic of KANs, including B - spline functions, sine functions, Bessel functions, etc. The selection and combination of basis functions are driven by the task:

[0160]

[0161] Among them, φ i is the i - th basis function (e.g., φ1(x)=spline(x), φ2(x)=sin(x)); β i is the trainable coefficient, representing the contribution of each basis function; M is the number of basis functions (M = 3).

[0162] A lightweight selection network Ψ (single - layer MLP) dynamically adjusts β i according to the input of the current layer:

[0163]

[0164] This method enables the activation function to adapt to the spatial patterns of traffic flow.

[0165] 4) Layer normalization

[0166] To ensure training stability, layer normalization is applied after each layer of convolution:

[0167]

[0168] Finally, the spatial features are output,

[0169] (2) Temporal flow: Temporal Convolution-based Dynamic KANs (TC-KANs)

[0170] The purpose of the temporal flow is to extract multi-scale temporal dependence features from the time series of traffic flow, which is achieved by using dilated temporal convolution (TCN) combined with a dynamic adjustment mechanism.

[0171] 1) Dynamic dilation coefficient selection

[0172] Traffic flow has the characteristics of multiple time scales (such as hourly peaks and daily cycles), so dilated convolution is used to capture long-term and short-term dependencies. The dilation coefficient d determines the receptive field of the convolution:

[0173] d = Round(Sigmoid(W d ·X (0) +b d )·d max )

[0174] where W d , b d are trainable parameters; d max is the maximum dilation coefficient (d max = 4); Round discretizes the continuous value into an integer. This dynamic selection mechanism adjusts the receptive field size according to the temporal characteristics of the input data (such as short-term fluctuations during peak periods or long-term trends during off-peak periods).

[0175] 2) Temporal convolution (TCN) operation

[0176] Perform one-dimensional dilated convolution on the time dimension:

[0177]

[0178] where X (l) ∈R N×T×D is the input of the l-th layer (the current node feature matrix); W TC ∈R 3×64×64 is the convolution kernel weight, k is the convolution kernel size (k = 3); d is the dilation coefficient of the current layer; s is the stride (usually set to 1).

[0179] Through stacked 3 convolutions, the receptive field is gradually expanded, and the dilation coefficient of each layer increases exponentially (e.g., d = 1, 2, 4), but is constrained by the dynamic selection mechanism.

[0180] 3) Basis function activation and non - linear enhancement

[0181] Similar to the spatial stream, a parametric basis function library is used to perform non - linear transformation on the convolution output:

[0182]

[0183] where M = 3, ψ i is the basis function (e.g., ψ1(x) = spline(x), ψ2(x) = cos(x)), and η i is the trainable coefficient.

[0184] The selection network Ω (single - layer MLP) is used to adjust η according to the input features i :

[0185]

[0186] This allows the model to select appropriate basis functions according to the periodicity or trend of time features, and finally output time features

[0187] 4) Residual connection and normalization

[0188] To retain the original time information, a residual connection is added:

[0189]

[0190] Subsequently, layer normalization is applied:

[0191]

[0192] After the processing of the spatial stream and the time stream, spatial features time features are obtained respectively, providing inputs for subsequent spatio - temporal feature fusion.

[0193] Step 3: Dynamic adaptive feature alignment and spatio - temporal feature fusion with collaborative enhancement;

[0194] To avoid directly concatenating spatial features and time features This method proposes a fusion strategy based on dynamic adaptive feature alignment and collaborative enhancement. By adaptively adjusting the correlation of spatio - temporal features and enhancing their complementarity, efficient fusion is achieved. The specific steps are as follows:

[0195] (1) Dynamic alignment of spatio - temporal features

[0196] Calculate the local correlation matrix between spatial features and temporal features to measure the coupling degree between the two at each node and time step:

[0197] C i,t = CosineSimilarity(H spatial,i,t , H temporal,i,t )

[0198]

[0199] where C i,t is the cosine similarity of node i at time t, and H spatial,i,t , H temporal,i,t are the spatial and temporal feature vectors of node i at time step t; t ranges from 1 to 288.

[0200] The alignment transformation uses a lightweight transformation network Φ (consisting of two layers of MLP) to dynamically adjust the spatial and temporal features according to the correlation matrix to align them in the semantic space:

[0201]

[0202] where W sp , b sp , W tp , b tp are trainable parameters, and Sigmoid(C) is used as a gating factor to control the alignment strength; The alignment weight matrix of spatial and temporal features; The bias term; Sigmoid(C): the gating factor, dynamically controlling the alignment strength. Function: Through the lightweight transformation network Φ, align the spatial and temporal features in the semantic space. In the high-correlation region (Sigmoid(C) ≈ 1), feature consistency is enforced, and in the low-correlation region (Sigmoid(C) ≈ 0), feature independence is retained.

[0203] (2) Co-enhancement module

[0204] Introduce a bidirectional interactive attention mechanism to enhance the complementary information flow between spatial and temporal features:

[0205]

[0206] where, The projection matrices of query and key are used to calculate the interaction weights between features; The attention weight from space to time, indicating the contribution of temporal features to spatial features; The attention weight from time to space, indicating the contribution of spatial features to temporal features.

[0207] Update features based on attention weights:

[0208] H″ spatial = H′ spatial + A tp-to-sp · (W V · H′ temporal )

[0209] H″ temporal = H′ temporal + A sp-to-tp · (W V · H′ spatial )

[0210] Among them, The value projection matrix is used to extract interaction information; the updated feature H″ spatial and H″ temporal . Incorporate spatio-temporal complementary information.

[0211] This two-way interaction enables spatial features to absorb time-dependent information and time features to incorporate spatial topological information.

[0212] Enhance the expressiveness of the fused features through residual connections and basis function enhancement:

[0213]

[0214] Among them, φ i (·): Basis function (such as spline function, sine function), used to enhance the non-linear expression ability; γ i : Trainable coefficient, dynamically adjusting the contribution weights of each basis function; LayerNorm: Layer normalization, ensuring numerical stability.

[0215] The final fused feature Integrates the synergistic effects of spatio-temporal information.

[0216] Step 4: Output the prediction result through a lightweight KANs decoder.

[0217] The goal of this step is to decode the spatio-temporal fused feature (where N is the number of nodes, T is the time step, and D is the feature dimension) generated in Step 3 into the traffic flow prediction results for the next 12 time steps (F is the original feature dimension). To achieve efficiency and prediction accuracy, a lightweight KANs decoder is designed to generate the final result through dimensionality reduction operations, adaptive basis function transformation, and dynamic output adjustment. The specific implementation is as follows:

[0218] (1) Feature dimensionality reduction and initialization

[0219] The input fused feature Hfused It has a relatively high hidden dimension D. To reduce the subsequent computational burden, a linear transformation is used for dimensionality reduction:

[0220]

[0221] where, is the dimensionality reduction weight matrix; is the bias vector; the output

[0222] is the prediction for generating the future T' time steps. A lightweight fully connected layer is used to expand the time dimension:

[0223]

[0224] where, is the time expansion matrix, and T' is the number of prediction time steps (T' = 12); is the bias; the output

[0225] Layer normalization is performed on the dimensionality-reduced and expanded features to ensure numerical stability:

[0226]

[0227] (2) Adaptive basis function decoding

[0228] Utilizing the basis function characteristics of KANs, a non-linear transformation is performed on the dimensionality-reduced features to enhance the expressive power of the decoder. Define a parameterized basis function library {ψ1, ψ2,..., ψ M}}(e.g., ψ1(x) =}spline(x),}ψ2(x) = sin(x)}, ψ3(x) = Bessel(x)). Each basis function corresponds to a different type of non-linear pattern:

[0229]

[0230] where, λ i is the contribution coefficient of the i-th basis function, and the initial value is a trainable parameter; M is the number of basis functions (M = 3); the output

[0231] To make the basis function adapt to different traffic flow prediction scenarios (such as short-term fluctuations or long-term trends), a dynamic weight generator Γ (lightweight MLP) is introduced to adaptively adjust λ according to the input features i :

[0232]

[0233] where, W Γ1,W Γ ,b Γ1 ,b Γ are the MLP parameters. This dynamic adjustment mechanism enables the decoder to select an appropriate combination of basis functions according to the characteristics of the input data (such as periodicity during peak periods or smoothness during off-peak periods).

[0234] To retain the original feature information, a residual connection is added:

[0235]

[0236] (3) Output mapping and dynamic adjustment

[0237] Map the decoded features to the target feature dimension F (for example, F = 2 represents flow and speed):

[0238]

[0239] where is the output mapping matrix; is the bias; the output is the normalized prediction result.

[0240] To further improve the prediction accuracy, a dynamic adjustment module is introduced to adjust the output according to the historical prediction error. Define an error feedback vector (initially zero, updated during training), and correct the prediction through a gated unit:

[0241]

[0242] where W gate ,b gate are the gated parameters; G is the adjustment weight, controlling the strength of the error feedback.

[0243] Convert the normalized prediction result Y norm back to the original dimension:

[0244] Y pred = Y norm ·σ + μ

[0245] where μ and σ are the mean and standard deviation calculated in step 1.

[0246] (4) Lightweight design optimization

[0247] By reducing the dimension (D′ < D) and using a small number of basis functions (M = 3), the number of parameters of the decoder is significantly reduced. The single-layer basis function transformation and the dynamic adjustment module avoid the high computational cost of multi-layer stacked networks and are suitable for real-time traffic prediction scenarios.

[0248] Final output It represents the traffic flow prediction value (such as flow rate, speed, etc.) for the next 12 time steps, which can be directly used for downstream tasks (such as congestion analysis or path planning).

[0249] In order to verify the effectiveness of the various modules improved and introduced by the present invention on the model improvement, under the same parameter conditions, the present invention conducted the following 6 groups of ablation experiments and 5 groups of comparative experiments on five innovative points. The results of the ablation experiments and comparative experiments are shown in Tables 1 and 2:

[0250] Table 1 Ablation experiment results

[0251]

[0252]

[0253] The complete DKAN-TF model performs well in the traffic flow prediction task, with a MAE of only 1.78 vehicles / minute, and RMSE and MAPE of 3.48 and 5.12% respectively, which verifies its significant advantages in modeling complex spatiotemporal features. Although its training time is relatively long, its accuracy fully meets the needs of complex scenarios. After removing the dynamic adjacency matrix, the MAE rises to 2.05 (15.2% higher than the complete model), and the spatial dependency modeling ability is limited, and it cannot effectively capture dynamic road network changes. The model with a fixed convolution order (K=2) has a MAE of 1.95 (9.0% higher than the complete model), and the fixed convolution kernel limits its capture of global dynamic spatial associations. After replacing the basis function with ReLU, the MAE rises to 2.12 (19.1% higher than the complete model), and the nonlinear mapping ability is limited, affecting the model performance. The model with simple splicing and fusion (no attention) has a MAE of 1.89 (6.2% higher than the complete model), and the ability to capture local spatial features is insufficient, affecting the prediction accuracy. After removing the dynamic error feedback, the MAE rose to 1.82 (8.4% higher than the complete model). The lack of error feedback mechanism leads to a decrease in the model's ability to correct prediction errors. In contrast, the complete DKAN-TF has comprehensive performance through the collaborative design of dynamic adjacency matrix, adaptive convolution order, custom basis function, attention fusion and dynamic error feedback. Its training time is reasonable, reflecting the balance between efficiency and accuracy in model design. The dynamic adaptability of DKAN-TF is particularly prominent. Through dynamic feature weighting and spatiotemporal consistency optimization, it maintains stable predictions in abnormal events and multimodal data scenarios, and ultimately achieves a systematic surpassing of other ablation models in terms of accuracy, efficiency and scenario adaptability, providing a better solution for highly dynamic urban traffic management.

[0254] Table 2 Comparative experimental results

[0255]

[0256] The STGCN model demonstrated certain capabilities in traffic flow prediction tasks, with a MAE of 2.34 vehicles per minute, an RMSE of 4.56, and a MAPE of 6.78%, validating its basic effectiveness in capturing spatial dependencies. However, the training time was 32.5 seconds per epoch, and the inference time was 12.3 milliseconds per step, indicating that there is still room for improvement in terms of efficiency, especially when dealing with highly dynamic traffic scenarios, where real-time performance may be limited.

[0257] The DCRNN model enhanced spatial propagation modeling through diffusion convolution, reducing the MAE to 2.18 vehicles per minute, with an RMSE of 4.32 and a MAPE of 6.45%, showing a certain improvement compared to STGCN. However, the training time increased to 45.2 seconds per epoch, and the inference time also extended to 15.8 milliseconds per step, indicating that while pursuing higher accuracy, it sacrificed some computational efficiency, which may be a point to consider for traffic management systems that require quick responses.

[0258] The GWN model had a MAE of 2.25 vehicles per minute, an RMSE of 4.41, and a MAPE of 6.52%. The training time was 38.7 seconds per epoch, and the inference time was 13.5 milliseconds per step. This model optimized spatial dependency modeling to a certain extent, but the overall performance improvement was not significant, and there was no obvious advantage in terms of training and inference time, indicating that further exploration is still needed in balancing the model structure and algorithm efficiency.

[0259] The traditional KAN model had a MAE of 2.40 vehicles per minute, an RMSE of 4.68, and a MAPE of 7.02%. The training time was 28.9 seconds per epoch, and the inference time was 10.2 milliseconds per step. Although this model showed better performance in terms of inference time, its accuracy metrics were relatively low, indicating its deficiency in modeling complex spatio-temporal features and difficulty in meeting the requirements of high-precision traffic flow prediction.

[0260] In contrast, the DKAN-TF model demonstrated excellent performance in traffic flow prediction tasks, with a MAE of only 1.78 vehicles per minute, an RMSE of 3.48, and a MAPE of 5.12%, validating its significant advantages in modeling complex spatio-temporal features. Although its training time was 26.4 seconds per epoch, slightly higher than some models, the substantial improvement in accuracy fully met the requirements of complex scenarios, and the inference time was 9.5 milliseconds per step, the shortest among all models. This indicates that DKAN-TF achieved an excellent balance between efficiency and accuracy in model design.

[0261] Upon further analysis, the dynamic adaptability of DKAN-TF is particularly prominent. Through feature dynamic weighting and spatio-temporal consistency optimization, the model can maintain stable predictions in abnormal event and multi-modal data scenarios. For example, after removing the dynamic adjacency matrix, the MAE rises to 2.05 (15.2% higher than the complete model), indicating that the dynamic adjacency matrix is crucial for capturing dynamic road network changes; fixing the convolutional order (K = 2) raises the MAE to 1.95 (a 9.0% increase), suggesting that the adaptive convolutional order can better capture global dynamic spatial correlations; after replacing the basis function with ReLU, the MAE rises to 2.12 (a 19.1% increase), highlighting the advantage of the custom basis function in non-linear mapping; simple concatenation fusion (without attention) raises the MAE to 1.89 (a 6.2% increase), proving the importance of the attention fusion mechanism for capturing local spatial features; after removing the dynamic error feedback, the MAE rises to 1.82 (an 8.4% increase), further verifying the key role of dynamic error feedback in correcting prediction errors.

[0262] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. It should not be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit of the present invention should be covered by the present invention.

Claims

1. A method for adaptively extracting spatiotemporal features of traffic flow based on a dynamic Kolmogorov-Arnold network, characterized in that: The following steps are involved: Step 1: Perform spatiotemporal normalization on the input traffic time series data to generate the intermediate feature matrix of the traffic time series data. Based on the topological structure of the traffic network, a static adjacency matrix is ​​constructed. The normalized intermediate feature matrix of the traffic time series data is combined with the static adjacency matrix to generate the initial node feature matrix. Step 2: Input the node feature matrix of the current layer into the Kolmogorov-Arnold network (KANs) to generate a dynamic adjacency matrix, use the KANs base function library to replace the traditional activation function of the graph convolutional network (GCN) to form a new graph convolutional KANs network (GC-KANs), and input the new GC-KANs network to output spatial features through the dynamic adjacency matrix; use the KANs base function library to replace the traditional activation function of the temporal convolutional network (TCN) and introduce the initial node feature matrix to improve the expansion coefficient to form a new temporal convolutional KANs network (TC-KANs), and input the initial node feature matrix into the new TC-KANs network to output temporal features; Step 3: Align spatial and temporal features in the semantic space through a lightweight transformation network, introduce a bidirectional interactive attention mechanism, and combine basis functions and residual connections to obtain fused features; Step 4: Reduce and expand the fusion features through the lightweight KAN decoder, generate future multi-time step prediction results based on the dynamic basis function library and nonlinear transformation, and correct the output through the error feedback mechanism.

2. The method for adaptively extracting spatiotemporal features of traffic flow based on a dynamic Kolmogorov-Arnold network according to claim 1, characterized in that: In step 2, the specific method of generating the dynamic adjacency matrix is: generating an unnormalized adjacency matrix according to the feature matrix of the current layer nodes through the KANs super network, and performing row-by-row Softmax normalization on it to obtain a dynamic adjacency matrix, wherein the dynamic adjacency matrix represents the dynamic connection strength between nodes.

3. The method for adaptively extracting spatiotemporal features of traffic flow based on a dynamic Kolmogorov-Arnold network according to claim 1, characterized in that: In step 2, the dynamic graph convolution KANs (GC-KANs) performs multi-order graph convolution through the dynamic adjacency matrix: α k =Sigmoid(W gate ·Mean(X (l) )+b gate ) It captures dependencies in different spatial ranges and introduces a gating mechanism to adaptively adjust the convolution order according to the spatial characteristics of the input data.

4. The method for adaptively extracting spatiotemporal characteristics of traffic flow based on a dynamic Kolmogorov-Arnold network according to claim 1, characterized in that: In step 2, the KANs base function library includes B-spline function, sine function and Bessel function: The combination weights of the basis functions are dynamically adjusted according to the current layer input through a lightweight selection network, and layer normalization is applied after each convolution layer to ensure training stability.

5. The method for adaptively extracting spatiotemporal characteristics of traffic flow based on a dynamic Kolmogorov-Arnold network according to claim 1, characterized in that: In step 2, the dynamic temporal convolutional KANs (TC-KANs) adopts dynamic dilation convolution and adjusts the dilation coefficient through trainable parameters to capture long-term and short-term dependencies: d=Round(SIgmoid(W d ·X (0) +b d )·d max ) Residual connections and layer normalization are introduced after convolution to preserve temporal information.

6. The method for adaptively extracting spatiotemporal characteristics of traffic flow based on a dynamic Kolmogorov-Arnold network according to claim 1, characterized in that: In step 3, by calculating the local correlation matrix between the spatial features and the temporal features, a lightweight transformation network is used to dynamically adjust the feature alignment according to the correlation, and the complementary information flow of the spatial features and the temporal features is enhanced through a bidirectional interactive attention mechanism.

7. The method for adaptively extracting spatiotemporal characteristics of traffic flow based on a dynamic Kolmogorov-Arnold network according to claim 1, characterized in that: In step 4, the lightweight KAN decoder reduces the dimension of the fusion features through linear transformation, then uses a fully connected layer to expand the time dimension, and uses a dynamic basis function library to perform nonlinear transformation, wherein the weights of the dynamic basis function library are adaptively adjusted according to the input features through a dynamic weight generator.

8. The method for adaptively extracting spatiotemporal characteristics of traffic flow based on a dynamic Kolmogorov-Arnold network according to claim 1, characterized in that: In step 4, the error feedback mechanism modifies the prediction result by defining an error feedback vector and combining it with a gating unit, wherein the error feedback vector is dynamically updated during training according to historical prediction errors.

Citation Information

Patent Citations

  • Traffic flow prediction method based on multi-dimensional time and space dependence mining

    CN115953900A

  • Space-time adaptive dynamic graph convolutional network traffic flow prediction method

    CN118629226A

  • Expressway traffic flow prediction method, device, equipment and medium

    CN118887812A

  • Traffic flow forecasting method based on multi-mode dynamic residual graph convolution network

    US20230334981A1

Cited By

  • County-level sea area sea wave correction forecasting method, equipment, medium and product

    CN120408331A

  • Multi-dimensional signal feature mapping method and device based on double-branch structure

    CN121188724A

  • Multi-dimensional signal feature mapping method and device based on double-branch structure

    CN121188724B

  • Complex network dynamics prediction method based on BKAN

    CN121302939A

  • A method for predicting the dynamics of complex networks based on BKAN

    CN121302939B