Transform and KAN network-based space-time air quality prediction method
Through the spatial and temporal air quality prediction method based on Transformer and KAN networks, the wavelet transformation and ProAttention modules are used to extract features in the time and space dimensions and fusion is solved, and the problem of difficult to obtain fine-grained results in the existing technology is solved, and the accuracy and robustness of air quality prediction are improved.
Patent Information
- Application Number
- CN202510262012.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-04
AI Technical Summary
The existing air quality prediction methods are difficult to obtain the correlation effect of spatial site data in spatiotemporal prediction, resulting in insufficient prediction accuracy.
The spatial and temporal air quality prediction method based on Transformer and KAN networks is adopted, and the main site features are extracted in the time dimension through wavelet transformation and ProAttention modules, divided into multiple sub-regions in the spatial dimension and processed independently, and finally feature fusion and fitting are performed in the spatial and temporal dimension.
It improves the accuracy and robustness of air quality prediction, can extract spatial site features more fine-grained, and improves the generalization ability of the prediction model.
Smart Images

Figure CN120256908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air quality prediction, and in particular to a spatio-temporal air quality prediction method based on Transformer and KAN network. Background Art
[0003] Carrying out air quality prediction is of great significance for protecting public health, improving quality of life, economic development and maintaining ecological balance.
[0004] At present, most cities monitor the air quality in real time by building air quality monitoring stations in key areas of the city, aiming to accurately obtain the air quality data of the city. For example, Beijing has built multiple monitoring stations, recording the concentrations of major gaseous pollutants per hour, especially PM2.5, PM10, NO2, CO, O3 and SO2. However, due to the limitedness of emission source information, the high uncertainty of air dynamic processes, sensor errors and environmental interference, and the complex kinetic characteristics of air pollutants, it is difficult to predict, and the above problems pose severe challenges to the accurate prediction of air quality.
[0005] At present, in terms of air quality prediction methods, there are mainly two categories: numerical physical calculation methods and data-driven methods. Traditional numerical physical methods are a kind of deterministic methods with good interpretability, such as Environment-Met, FLU-ENT, MISKAM and OSPM, etc. Since such models require the integrity of pollution source data, the generalization ability of the models is poor. Data-driven methods mainly include statistical methods, machine learning methods and deep learning methods. Statistical methods include autoregression, moving average, autoregressive moving average and autoregressive integrated moving average, etc. Machine learning methods include support vector regression (SVR) and random forest (RF), etc. In recent years, the data-driven deep learning method is becoming the mainstream of air quality prediction. The deep learning model fits the non-linear transformation from input to output by stacking multiple neural networks, realizing automatic feature learning. Utilizing the powerful complex feature extraction ability and learning ability of the deep learning model, it has higher prediction accuracy and better generalization ability.
[0006] In data-driven methods, there are mainly air quality prediction that only considers time features (time prediction) and air quality prediction that considers spatio-temporal features (spatio-temporal prediction). The advantages of time prediction are simple prediction model, high flexibility, fast response speed, etc., but it has disadvantages such as being greatly affected by outliers and relatively low prediction accuracy compared to spatio-temporal prediction. In spatio-temporal prediction, the model simultaneously fuses time and space dimension features, making the model have high prediction accuracy, strong robustness and generalization ability, and being more effective in dealing with non-linear complex dependence relationships. This method is the mainstream method of current prediction. However, in spatio-temporal prediction, many studies handle spatial features relatively simply. They simply merge the data of each station. Due to the influence of geographical and spatial environmental factors among stations, the influence of each spatial station on the prediction station is different. It is difficult to make good use of the information features of spatial stations in a fine-grained manner by simply merging spatial stations.
[0007] In summary, how to obtain the correlation influence of spatial station data on the main station data in a more fine-grained manner and further improve the prediction accuracy of air quality is an urgent problem to be solved at present. Summary of the Invention
[0008] In view of the above-mentioned shortcomings of the prior art, the present invention proposes a spatio-temporal air quality prediction method (ST-KAN-Former) based on Transformer and KAN network. The core technical idea is as follows: In terms of time feature extraction, first, use wavelet transform to decompose the features of the main station data to obtain the air quality features of the main station in a more fine-grained manner. Then, rely on the KAN network for embedding processing. Finally, use the ProAttention module to calculate the correlation features of the data to obtain the data features in the time dimension; in the spatial dimension, according to the distance parameter between the main station and the slave stations, divide the spatial region into 4 sub-regions, and independently process the data in each region to obtain the data features of the spatial stations in a more fine-grained manner; in the spatio-temporal dimension, fuse the time dimension data and the data features of each spatial region, use the Proattention attention module to extract the spatio-temporal data features, and then rely on the KAN network for processing to obtain the final output.
[0009] The technical solution adopted by the present invention is a spatio-temporal air quality prediction method based on Transformer and KAN network, which includes the following steps:
[0010] Step 1: Take the air quality monitoring station to be predicted as the main station, collect the historical feature data of the main station, and use the hybrid filling algorithm to fill the historical feature data.
[0011] Step 2: Divide all the monitoring stations within the area where the main station is located based on the relevance of the stations, obtaining multiple sub-regions, each of which contains several monitoring stations;
[0012] Step 3: Based on the KAN network, wavelet transform, and ProbAttention network, construct a time-dimensional feature extraction model to extract the time feature data of the main station in the time dimension;
[0013] Step 4: Based on the KAN network and ProAttention network, construct a spatial feature extraction model, and use this model to perform spatial autocorrelation calculations on the station data of each sub-region respectively to obtain fine-grained spatial feature data;
[0014] Step 5: Based on the ProbAttention network and KAN network, construct a spatio-temporal feature fusion model, fit the time feature data and the spatial feature data, and output spatio-temporal features;
[0015] Step 6: Use the spatio-temporal features to predict the air quality of the current station to be measured.
[0016] Furthermore, the data filling process of the hybrid filling algorithm in Step 1 includes:
[0017] Step 1.1: Find the column j with the largest missing rate in the dataset M to be processed max :
[0018]
[0019] where max(·) is the function to take the maximum value, is.na(·) is the function to count the number of observations with value NA in the j-th column variable, nrow(·) is the function to count the number of samples in M; M i,j , i ∈ [1, n], j ∈ [1, m] represents the element in the i-th row and j-th column of M, n represents the number of sample points, and m represents the total number of variables; P j , j ∈ [1, m] represents the proportion of missing values of the j-th column observed variable in M;
[0020] Step 1.2: Generate the initial dataset M′ to be filled through the following formula:
[0021]
[0022] In the formula, represents removing the j max column from M;
[0023] Step 1.3: Use the MICE.rf filling method in the Markov chain Monte Carlo method to perform data filling processing on M′ to obtain the complete dataset M c ′;
[0024] Step 1.4: Concatenate the data of M c ′ and column by column to obtain a new dataset M″ to be filled;
[0025] Step 1.5: Use the MICE.norm.predict filling method in the Markov chain Monte Carlo method to fill the data of M″ and obtain the complete dataset M c .
[0026] Furthermore, the specific steps for regional division in Step 2 include:
[0027] Step 2.1: Take the main site as the center and define the main site as G0;
[0028] Step 2.2: Take the area where all monitoring sites with a distance less than 20 km from the main site are located as the G1 area;
[0029] Step 2.3: Take the area where all monitoring sites with a distance between 20 km and 40 km from the main site are located as the G2 area;
[0030] Step 2.4: Take the area where all monitoring sites with a distance between 40 km and 60 km from the main site are located as the G3 area;
[0031] Step 2.5: Take the area where all monitoring sites with a distance greater than 60 km from the main site are located as the G4 area;
[0032] Step 2.6: Define S0, S1, S2, S3, S4 as the sets of sites in the G0, G1, G2, G3, G4 areas respectively, and n1, n2, n3, n4 as the number of sites in S1, S2, S3, S4 respectively. For the main site In other sub - regions, denote S i,j as the j - th site in the i - th area, where i ∈ (1, 2, 3, 4), 0 ≤ j ≤ w, w ∈ (n1, n2, n3, n4). Then represent the site data in each area as:
[0033]
[0034] Furthermore, the steps for extracting time - feature data in Step 3 include:
[0035] Step 3.1: On the time dimension of the main site, for the data of the main site at time t Use the data of the previous L in time - length before time t to predict the air quality at multiple consecutive future time points, and the data of the previous L in time - length before time t at the main site is:
[0036]
[0037] Among them, N P is the number of observation indicators, represents the data collected for the N P th indicator at time t;
[0038] Step 3.2: Decompose the main site data using wavelet transform to obtain the low-frequency and high-frequency components of each indicator at the main site:
[0039]
[0040] Among them, represents the low-frequency component of the main site data ; represents the high-frequency component of the main site data ; H 01 represents the data after wavelet transform;
[0041] Step 3.3: Fuse H 01 with the time data and perform Embeding processing using the KAN network to obtain the encoded data H 02 :
[0042] H 02 = Embeding(concat(H 01 , T1)) = KAN(concat(H 01 , T1))
[0043] Among them, concat represents the data fusion operation, and T1 represents the time period (t - L in+1 , t);
[0044] Step 3.4: Standardize H 02 to obtain the normalized data H 03 :
[0045]
[0046] Among them, LayerMorn represents the normalization process; h n represents the eigenvalue at time t, and Mean represents taking the mean of h n ;
[0047] Step 3.5: Calculate the autocorrelation of H 03 using the ProAttention attention mechanism:
[0048] Q = K = V = H 03
[0049]
[0050] Among them, Q, K, and V respectively represent the query, key, and value of the ProbAttention attention mechanism; is a sparse matrix of the same size as q, where q represents each row of Q, and the sparsity metric M(q i , K) is:
[0051]
[0052] In the formula, q i represents the i-th row in Q; d represents the dimension of the key (K) vector;
[0053] And, M(q i , K) can be approximately calculated as:
[0054]
[0055] Step 3.6: Standardize the data H 04 after self-correlation calculation to obtain the standardized data H 05 :
[0056] H 05 = LayerNorm(H 04 );
[0057] Step 3.7: Use a linear transformation to perform a dimensionality transformation on H 05 and output the time feature data in the time dimension
[0058]
[0059] where Linear represents the linear transformation.
[0060] Furthermore, the specific operation steps of Step 5 include:
[0061] Step 5.1: Fuse all the site data in the four sub-regions G1, G2, G3, and G4 to obtain the fused data of these four sub-regions
[0062]
[0063] where l1, l2, l3, and l4 respectively represent the distances from the center points of the G1, G2, G3, and G4 sub-regions to the main site, and ε1, ε2, ε3, and ε4 respectively represent the diffusion coefficients of the G1, G2, G3, and G4 regions; S 1,n1 represents the n1-th site in the G1 sub-region, S 2,n2Denote the n2-th site within the G2 sub-region, S 3,n3 Denote the n3-th site within the G3 sub-region, S 4,n4 Denote the n4-th site within the G4 sub-region;
[0064] Step 5.2: Fuse the fusion data of each region with the time data T, and perform Embeding processing on the fused data using the KAN network to obtain the encoded data of each region
[0065]
[0066] where i ∈ (1, 2, 3, 4) represents the numbers of 4 regions in space;
[0067] Step 5.3: Perform normalization processing on each and calculate the autocorrelation using the ProAttention attention mechanism to obtain the calculated data
[0068]
[0069]
[0070] where,
[0071] Step 5.4: Perform normalization processing on , and then use the Linear network to align the dimensionality of the normalized data features and output the spatial feature data of each region
[0072]
[0073] Furthermore, the specific steps of Step 6 include:
[0074] Step 6.1: Fuse the time feature data and the spatial feature data to form the fused spatio-temporal feature matrix H de1 :
[0075] H de1 = Concat(H out0 , H out1 , H out2 , H out3 , H out4 )
[0076] where, H out0 is the time feature data; H out1 , H out2, H out3 , H out4 are the spatial feature data of five regions G1, G2, G3, and G4 respectively;
[0077] Step 6.2: Use the ProAttention attention mechanism to calculate the correlation of spatio-temporal feature numbers, and obtain the calculated data H de2 :
[0078] H de2 = ProbAttention(Q de , K de , V de )
[0079] where Q de = K de = V de = H de1 ;
[0080] Step 6.3: Fit H de2 based on the KAN network, and output the fitted spatio-temporal feature H de3 :
[0081] H de3 = KAN(H de2 )
[0082] where KAN represents the KAN network;
[0083] Step 6.4: According to the spatio-temporal feature H de3 , output the prediction result Y:
[0084] Y = Projection(H de3 )
[0085] Therefore, the present invention adopts the above spatio-temporal air quality prediction method based on the Transformer and KAN networks, and has the following beneficial effects:
[0086] First, for the spatio-temporal prediction problem of air quality, the present invention proposes a method for dividing the spatial site area based on distance parameters. By dividing the spatial site into multiple sub-regions, designing a dedicated data feature extraction model for each sub-region, and introducing a diffusion coefficient according to the phenomenon that the concentration of air pollutants gradually decreases during the propagation process, the spatial site feature data can be extracted with a finer granularity;
[0087] Second, the present invention uses the KAN network as the mapping network for Embedding. Relying on the dynamic learning activation mode and flexible function fitting ability of the KAN network, the Embedding module can better map the spatio-temporal feature data into the continuous vector space, thereby capturing the potential relationships between data at a finer granularity.
[0088] Third, the present invention constructs a multivariate air quality prediction model based on spatio-temporal correlation. In the time dimension, wavelet transform, KAN network, and Proattention attention mechanism are used to extract time dimension feature data at a fine granularity. In the space dimension, the spatial sites are divided into regions, and independent feature extraction is performed on each sub-region to obtain spatial feature data at a fine granularity. Finally, the time dimension features and spatial dimension features are fused, and after spatio-temporal autocorrelation calculation using Proattention, the output is fitted by the KAN network.
[0089] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0090] Figure 1 It is the framework diagram of the method proposed by the present invention.
[0091] Figure 2 It is the spatial sub-region division of the air quality monitoring sites.
[0092] Figure 3 It is the structural diagram of the T-encoder network in the present invention.
[0093] Figure 4 It is the structural diagram of the S-encoder network in the present invention.
[0094] Figure 5 It is the structural diagram of the ST-Decoder network in the present invention.
[0095] Figure 6 It is the Embedding processing model.
[0096] Figure 7 It is the spatio-temporal relationship diagram of the sites.
[0097] Figure 8 It is the missing ratio of data indicators at the monitoring points.
[0098] Figure 9 It is the distribution of air quality sites.
[0099] Figure 10 (a)-(g) are the comparisons of the MAE performance indicators between the present invention and different baseline models.
[0100] Figure 11(a)-(b) shows the percentage of performance improvement of the present invention compared to the best results of other baseline models.
[0101] Figure 12 It is the percentage of performance improvement of each indicator after area division.
[0102] Figure 13 (a)-(d) respectively show the performance comparison results between the present invention and removing the spatio-temporal module, replacing the Linear network with the KAN network, removing the wavelet transform module, and replacing the ProAttention network with the FullAttention network.
[0103] Figure 14 (a)-(b) show the performance improvement results of each module for the MAE and RMSE indicators respectively.
[0104] Figure 15 (a)-(b) respectively show the percentage of performance improvement of E-1 relative to E-2 and E-3.
[0105] Figure 16 It is the percentage of performance improvement of the diffusion parameter for each indicator.
[0106] Figure 17 (a)-(b) respectively show the percentage of performance improvement of MAE and RMSE of K-1 relative to K-0, K-2, and K-3.
[0107] Figure 18 (a)-(b) respectively show the percentage of performance improvement of MAE and RMSE of B-32 relative to B-16, B-64, and B-128. Detailed implementation manners
[0108] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0109] 1. Related definitions
[0110] Spatio-temporal air quality prediction is to predict the air quality in the future for a period of time using the historical data of multiple stations. Let N S be the total number of monitoring stations, N P be the number of observed indicators, L in be the input sequence length of the model, L out be the prediction output length. The data collected for each indicator at each station according to the observation duration T is X = {x1, x2,..., x t}, where, xt For the observed value at time t, where t ∈ (1, T), let N P The multivariate time series data collected from N indices is defined as:
[0111] X S ={x p,1 , x p,2 ,..., x p,t} (1)
[0112] where p ∈ (1, N p ), S ∈ (1, N S ), S represents the monitoring station number, and x p,t represents the data collected for the p-th index at time t;
[0113] The collected multivariate time series data X S is represented by a matrix as:
[0114]
[0115] Then, the prediction problem of the present invention can be described as follows:
[0116]
[0117] The matrix form of Equation (3) is:
[0118]
[0119] 2. The prediction method proposed by the present invention
[0120] For the spatio-temporal prediction problem of air quality, the present invention proposes a spatio-temporal air quality prediction method based on the Transformer and KAN networks, as Figure 1 shown. This method includes the following steps:
[0121] Step 1: Using the air quality monitoring station to be predicted as the main station, collect the historical feature data of the main station, and fill the historical feature data using the hybrid filling algorithm;
[0122] Step 2: Divide the monitoring stations in the area where the main station is located based on the station correlation to obtain multiple sub-regions, and each sub-region contains several monitoring stations;
[0123] Step 3: Based on the KAN network, wavelet transform, and ProbAttention network, construct a time dimension feature extraction model (T-encoder) to extract the time feature data in the time dimension of the main station;
[0124] Step 4: Based on the KAN network and the ProAttention network, construct a spatial feature extraction model (S-Encoder), and use this model to perform spatial autocorrelation calculations on the site data of each sub-region respectively to obtain fine-grained spatial feature data;
[0125] Step 5: Based on the ProbAttention network and the KAN network, construct a spatio-temporal feature fusion model (ST-Decoder), and fit the time feature data and the spatial feature data to output spatio-temporal features;
[0126] Step 6: Use the spatio-temporal features output in Step 5 to predict the air quality of the current site to be measured.
[0127] The key technologies of the above steps are introduced separately below:
[0128] (1) Spatial sub-region division
[0129] Since air pollutants are dispersed in the geographical space, the air quality of a geographical location depends not only on its previous air quality but also on the air quality of the surrounding areas. In order to accurately predict the future concentration of each air quality monitoring site with finer granularity, the spatial sites are divided. The specific method is as Figure 2 shown. Assume that air quality monitoring stations (the dots in the figure represent monitoring stations, and different colors represent air pollution concentrations) are randomly distributed in the geographical space. Take the site to be predicted as the main site, and with the main site ( Figure 2 the black pentagram in it) as the center, all sites are divided into 5 regions according to the distance from the site to the main site. The main site is defined as G0, the area less than 20 km from the prediction site is the G1 region, the area from 20 km to 40 km is the G2 region, the area from 40 km to 60 km is the G3 region, and the area greater than 60 km is the G4 region.
[0130] According to the division rule, for Let S0, S1, S2, S3, S4 be the sets of sites in the G0, G1, G2, G3, G4 regions respectively, and n1, n2, n3, n4 be the numbers of sites in S1, S2, S3, S4 respectively. For the main site In other sub-regions, let S i,j represent the j-th site in the i-th region, where i ∈ (1, 2, 3, 4), 0 ≤ j ≤ w, w ∈ (n1, n2, n3, n4). The site data in each region is expressed as:
[0131]
[0132] (2) Time dimension feature extraction model T-cecoder
[0133] In the time dimension, the future air pollutant concentration at the main site is highly correlated with the historical air pollutant concentration. As Figure 3 shown, the present invention designs a T-Encoder model capable of finely extracting historical feature data of the main site. This model first performs wavelet transform on the air pollutant feature data to finely obtain the high-frequency component and low-frequency component features. Then, the feature data is fused with the time data, and Embeding processing is performed using the KAN network. Finally, the Proattention network is used for context self-correlation calculation to obtain the time dimension features. Specifically, it includes:
[0134] ① Wavelet transform
[0135] In the data of Fine Particles (PM2.5), Respirable Particulate Matter (PM10), Sulfur Dioxide (SO2), Nitrogen Dioxide (NO2), Ozone (O3), and Carbon Monoxide (CO) collected at the air quality site, the data of each pollutant fluctuates very frequently and violently. Such data belongs to non-stationary sequence data. In order to be able to more finely extract data features, the present invention uses wavelet transform to decompose the feature data of each pollutant to obtain the high-frequency component and high-frequency component of the data. Then, feature extraction is performed on the high-frequency component and low-frequency component respectively, and the data features in the time dimension can be obtained more finely, thereby further improving the prediction accuracy. Wavelet transform (WT) is a data preprocessing method with good data stability performance. The present invention uses Daubechies as the basis function for wavelet transform to decompose the original data into multiple sub-signals. The wavelet transform and decomposition process are as follows:
[0136]
[0137] Among them, a represents the scale parameter that controls the width of the wavelet transform, and τ represents the translation parameter that controls the position of the wavelet.
[0138] The original waveform can be expressed as after wavelet decomposition:
[0139]
[0140] Among them, A m is an approximate information set, representing the overall pattern and long-term trend characteristics of the original data, belonging to the low-frequency component of the data; D i is the high-frequency component information set, representing the detailed high-frequency fluctuations caused by various transient factors, belonging to the high-frequency component of the original data.
[0141] In the present invention, m is set to 3, so as to obtain one low-frequency component A3(t) and three high-frequency components D3(t), D2(t), D1(t), that is:
[0142] f(t) = A3(t) + D3(t) + D2(t) + D1(t) (8)
[0143] ② T-encoder model
[0144] On the time dimension of the main site, let the data of the main site at time t be Use the data of the previous L in time lengths before time t to predict the air quality at multiple consecutive future time points. The data of the previous L in time lengths before time t at the main site is:
[0145]
[0146] First, based on Equation (8), use wavelet transform to decompose the data of the main site to obtain the low-frequency component and high-frequency component of each index of the main site:
[0147]
[0148] Among them, represents the low-frequency component of the data of the main site ; represents the high-frequency component of the data of the main site ;
[0149] Secondly, fuse the data after wavelet transform and the time data, and perform Embeding processing using the KAN network (the specific method is described later) to obtain the encoded data H 02 :
[0150] H 02 = Embeding(concat(H 01 , T1)) = KAN(concat(H 01 , T1)) (11)
[0151] Among them, T1 represents the time period (t - L in+1 , t); concat represents the feature fusion operation;
[0152] Thirdly, perform normalization processing on H 02 to obtain the normalized data H 03 :
[0153]
[0154] Among them, LayerMorn represents normalization processing, and h n represents the eigenvalue at time t, and Mean represents taking the mean of h n .
[0155] Next, use the ProAttention attention mechanism to calculate the self-correlation of H 03 :
[0156] Q = K = V = H 03 (13)
[0157]
[0158] Among them, Q, K, and V respectively represent the query, key, and value of the ProAttention attention mechanism; is a sparse matrix of the same size as q, where q represents each row of Q, and only contains the most important q values under the sparse metric M(q i , K), and the M(q i , K) is:
[0159]
[0160] Among them, q i represents the i-th row in Q; d represents the dimension of the key (K) vector, which is used for scaling to avoid too large dot products;
[0161] Through the above formula, q vectors with larger attention can be selected from Q i ;
[0162] M(q i , K) can be approximately calculated as:
[0163]
[0164] Normalize the data H after self-correlation calculation 04 , and then use a linear transformation to perform dimensional transformation on the normalized data H 05 to output the data features in the time dimension
[0165] H 05 = LayerMorn(H 04 ) (17)
[0166]
[0167] Among them, Linear represents a linear transformation.
[0168] (3) Spatial feature extraction model S-encoder
[0169] As described above, in the spatial dimension, the master site and the slave sites have spatial correlation, and the closer to the master site, the higher the correlation. Therefore, the slave sites near the master site are divided into four sub-regions, G1, G2, G3, and G4, and the data of all sites in each region are fused. Each region uses one S-Encoder structure to calculate the autocorrelation of data features, and finally the spatial data features are obtained. The structure of the S-Encoder model is as Figure 4 shown. The specific steps to extract the spatial data features are as follows:
[0170] Step 1: Fuse the data in the four sub-regions G1, G2, G3, and G4 to obtain the fused data of each sub-region as:
[0171]
[0172] where l1, l2, l3, and l4 represent the distances from the center points of the G1, G2, G3, and G4 regions to the master site, and ε1, ε2, ε3, and ε4 represent the diffusion coefficients of the G1, G2, G3, and G4 regions; S 1,n1 represents the n1th site in the G1 sub-region, S 2,n2 represents the n2th site in the G2 sub-region, S 3,n3 represents the n3th site in the G3 sub-region, S 4,n4 represents the n4th site in the G4 sub-region;
[0173] Step 2: Fuse the fused data of each region with the time data T, and use the KAN network to perform Embeding processing on the fused data to obtain the encoded data
[0174]
[0175] where i ∈ (1, 2, 3, 4) represents the numbers of the 4 regions in space; T represents the time data recorded when the sensor collects data;
[0176] Step 3: Normalize the encoded data to obtain the normalized data Then use the ProAttention attention mechanism to perform autocorrelation calculation on to obtain the calculated data
[0177]
[0178]
[0179] Among them,
[0180] Step 4: Normalize and use the Linear network to align the dimensionality of the standardized data features to output the spatial features of each region
[0181]
[0182]
[0183] All the data in the four regions G1, G2, G3, and G4 are respectively used with the above S-encoder models (S-encoder1, S-encoder2, S-encoder3, S-encoder4) to calculate the autocorrelation features and obtain the spatial features of each region.
[0184] (4) ST-Decoder
[0185] As Figure 5 shown, in order to better fuse the time feature data and the spatial feature data, the present invention designs an ST-Decoder network that can simultaneously extract the time and spatial feature data. First, the time feature data and the spatial feature data are fused; secondly, the ProAttention attention mechanism is used to calculate the correlation of the spatio-temporal feature data; thirdly, the KAN network is used to further fit the spatio-temporal data to improve the prediction accuracy; finally, the model prediction result is output. Specifically, it includes:
[0186] ① KAN network
[0187] The Kolmogorov-Arnold network (KAN) is designed based on the Kolmogorov-Arnold representation theorem. The Kolmogorov-Arnold representation theorem states that any multivariate continuous function can be represented as a finite nested combination of a series of univariate continuous functions:
[0188] For any continuous function f: [0,1] t →R, there exist univariate continuous functions Φ q and such that:
[0189]
[0190] Among them, φ q,p is a mapping that maps each input variable (x pUnivariate function, such as φ q,p : [0, 1] → R, and
[0191] In Equation (25), the one-dimensional univariate function parameter φ q,p can be transformed into a B-spline curve, and the B-spline curve has learnable local B-spline basis function coefficients. Therefore, the KAN network with n-dimensional input and n-dimensional output can be defined as a matrix of one-dimensional functions:
[0192] Φ = {φ q,p}, p = 1, 2,..., n in , q = 1, 2,..., n out (26)
[0193] where φ q,p has trainable parameters;
[0194] Let the shape of KAN be represented by an integer array:
[0195] [n0, n1,..., n L
[0196] where n i is the number of nodes in the i-th layer of the computational graph.
[0197] Use (l, i) to represent the i-th neuron in the l-th layer, and use x l,i to represent the activation value of the (l, i) neuron. Between the l-th layer and the l + 1-th layer, there are n l n l +1 activation functions. The activation function connecting (l, i) and (l + 1, j) is:
[0198] φ l,j,i , l = 0,..., L - 1, i = 1,..., n l , j = 1,..., n l+1 (27)
[0199] The pre-activation of φ l,j,i is simply xl,i; the post-activation of φ l,j,i is expressed as: The activation value of the (l + 1, j) neuron is the sum of all incoming post-activations:
[0200]
[0201] Its matrix form is:
[0202]
[0203] where Φ l is the function matrix corresponding to the l-th layer.
[0204] A general KAN network consists of L layers: Given an input vector The output of the KAN network is:
[0205]
[0206] ②ST-Decoder
[0207] To more accurately predict the future air quality of the main site, first use the T-Encoder to extract the data features in the time dimension of the main site, and at the same time use the S-Encoder to extract the spatial dimension data features of the G1, G2, G3, and G4 regions respectively. Next, fuse the time dimension data features and the spatial dimension data features, and use the spatio-temporal feature extraction model ST-Decoder to predict the air quality of the main site for a period of time in the future. The structure of ST-Decoder is as Figure 5 shown, where L represents the length of the input sequence, D1 represents the total number of parameters after spatio-temporal feature merging, and D2 represents the number of parameters to be predicted. The specific process is as follows:
[0208] Step 1: Fuse the time data features and the spatial data features to form a fused spatio-temporal feature matrix H de1 :
[0209] H de1 = Concat(H out0 , H out1 , H out2 , H out3 , H out4 ) (31)
[0210] where H out0 is the data feature in the time dimension of the main site extracted by the T-Encoder; H out1 , H out2 , H out3 , H out4 are the spatial features of the G1, G2, G3, and G4 sub-regions extracted by the S-encoder respectively;
[0211] Step 2: Use the ProAttention attention mechanism network to calculate the correlation of H de1 :
[0212] H de2 = ProbAttention(Q de , K de , V de ) (32)
[0213] where Q de = Kde = V de = H de1 ;
[0214] Step 3: Use the KAN network to fit the spatio-temporal feature data H after correlation calculation de2 and output the spatio-temporal feature H after fitting de3 :
[0215] H de3 = KAN(H de2 );
[0216] Step 4: Based on the spatio-temporal feature H after fitting de3 , predict the air quality of the main site for a period of time in the future:
[0217] Y = Projection(H de3 )
[0218] where Y represents the prediction result; Projection represents the output layer, which is used to transform the dimension of the feature vector into the dimension of the feature label;
[0219] 3. Embedding processing
[0220] In the above T-encoder model, it finally uses the KAN network for Embedding processing. Through Embedding, the input vector can be converted into a feature vector with a length of d model = 512, enabling the model to better capture the periodic and non-periodic features of the sequence data, and at the same time being able to better express the dependency relationships in time series prediction. Different from the traditional Embedding method, the present invention relies on the KAN network as the conversion model. The KAN network uses a spline-parameterized univariate function instead of the traditional linear weights, enabling it to dynamically learn the activation pattern, so that the KAN network has a flexible function fitting ability and can more accurately capture complex non-linear relationships and periodic patterns in time series. The specific approach is as Figure 6 shown. In the figure, D represents data features, T represents time features, and L represents the sequence length. First, fuse the data features and time features, and then input the fused feature data into the KAN network. The number of layers of the KAN network is L layers; finally, obtain the feature data with a length of d model = 512 after the Embedding transformation.
[0221] 4. Spatio-temporal correlation
[0222] In the S-encoder, there is a spatial correlation between the main site and the slave sites. The closer to the main site, the higher the correlation. Based on the correlation, the spatial region is divided into several sub-regions. Such asFigure 7 As shown, the air quality of each site is highly correlated in time. In addition, it also has spatial correlation with surrounding sites. In the prior art, correlation analysis was performed on 34 air quality monitoring sites in Beijing, and the Spearman correlation coefficient was used to measure the correlation. The analysis results show that in terms of spatial correlation, as the distance increases, the correlation coefficient decreases significantly. The correlation coefficient of monitoring stations with a distance less than 20 km is relatively high (>0.9), and the correlation coefficient of monitoring stations with a distance less than 100 km is above 0.6. In terms of time correlation, when the delay time is 20 h, the correlation coefficient is above 0.8, and when the delay time is about 80 h, the correlation coefficient approaches 0. From the above two correlation analyses, it can be seen that air quality parameters have strong correlations in both space and time.
[0223] 5. Data filling
[0224] During the data collection of monitoring stations, due to the influence of factors such as the environment and equipment, the collected data contains outliers. The existence of outliers may affect the model performance, generalization ability, stability and reliability, making the model prone to overfitting and underfitting phenomena, and misleading the model decision. Removing outliers and using a filling algorithm to fill the data at the abnormal positions is one of the effective methods to handle outliers. Facing the missing ratios of different indicators in the dataset to be filled, adopting a more targeted data filling algorithm to preprocess the original data is beneficial to subsequent model training, and then ensuring the reliability of subsequent experimental results. Since there is a certain correlation between different indicators of air quality data, combined with the distribution of the missing value ratios of each indicator in the dataset to be processed, as Figure 8 shown in the missing value situations of 6 randomly selected sites, it can be seen from it that the distribution of missing values is relatively scattered, and the scattered range includes more than 15%, between 5% - 15% and less than 5%. For the missing ratios of each indicator, the linear regression method is used to fill the indicators with larger missing ratios, while other indicators are processed by the random forest filling method. Based on this, the present invention uses the MICE.rf filling method and the MICE.norm.predict filling method based on the Markov chain Monte Carlo method (MCMC) to preprocess the data indicators with a missing ratio lower than 5% and higher than 15% in the same dataset respectively, and proposes a Random forest imputations and Linear regression hybrid filling algorithm (abbreviated as MRFNP) based on the idea of multiple imputation.
[0225] This hybrid filling algorithm includes the following steps:
[0226] 1) Find the column with the largest missing rate:
[0227]
[0228] In the formula, max(·) is the maximum value function, is.na(·) is the function for counting the number of observations with NA in the j-th column variable, and nrow(·) is the function for counting the number of samples in M; M represents the dataset to be processed, M i,j , i ∈ [1, n], j ∈ [1, m] represents the element in the i-th row and j-th column of M, n represents the number of sample points, and m represents the total number of variables; P j , j ∈ [1, m] represents the proportion of missing values in the j-th column of observed variables in M;
[0229] 2) Generate the initial dataset M′ to be filled:
[0230]
[0231] In the formula, represents removing the j max th column from M;
[0232] 3) Use the MICE.rf filling method to perform data filling on M′ to obtain the complete dataset M c ′;
[0233] 4) Concatenate the data of M c ′ and column by column to obtain the new dataset M″ to be filled;
[0234] 5) Use the MICE.norm.predict filling method to perform data filling on M″ to obtain the complete dataset M c .
[0235] Embodiment
[0236] To verify the effectiveness of the air quality prediction method (hereinafter referred to as ST-KAN-Former) proposed by the present invention, a large number of experiments were carried out on a real-world dataset in this embodiment to evaluate the overall performance of this method, and the main modules and parameters affecting the model performance were experimentally verified.
[0237] 1. Dataset and experimental environment
[0238] (1) Dataset
[0239] As Figure 9As shown in the figure, experiments were conducted based on the real data of 28 air quality monitoring stations in Beijing. The dataset covers a wide range of areas, including urban and suburban areas. The dataset ranges from January 1, 2021 to December 31, 2023, with data recorded hourly. In each dataset, air quality observations include CO, NO2, O3, PM10, PM2.5, and SO2, and a total of 725,760 historical data were collected. The experimental goal is to predict the trends of various pollutants in the future for a period of time. The observed data of air quality in Beijing comes from the Chinese government website (https: / / www.cnemc.cn / en / ), and the data is divided into a training set, a validation set, and a test set in a ratio of 7:2:1.
[0240] (2) Experimental Environment Configuration and Model Parameters
[0241] Both the ST-KAN-Former model and all deep learning baseline models in this study were evaluated on a Windows server with NVIDIA A1000 GPUs. In terms of model parameters, the self-attention mechanism in Informer was used for attention, with the number of layers L = 1; the hidden layer of the KAN network was 1, with a size of 64; the model was trained using the stochastic gradient descent method with a learning rate lr = 0.0001, and the diffusion parameters ε1, ε2, ε3, and ε4 were 0.6, 0.6, 0.5, and 0.35 respectively.
[0242] (3) Evaluation Metrics
[0243] Two evaluation metrics, mean absolute error (MAE) and root mean square error (RMSE), were used to measure the performance of the prediction model. The ST-KAN-Former model was compared with other models. The smaller the MAE and RMSE values, the better the performance of the model. The calculation formulas for MAE and RMSE are as follows:
[0244]
[0245] (4) Baseline Models
[0246] The baseline models selected in this embodiment include:
[0247] ① The iTransformer model, which is a deep learning model that combines the Transformer architecture and the iterative learning mechanism. This model can effectively capture long-term dependencies in time series and is suitable for various time series analysis and prediction tasks.
[0248] ②The Dlinear model combines the advantages of linear transformation and deep network structure. By stacking multiple layers of linear transformation, this model can capture complex relationships and patterns in the data while maintaining computational efficiency and model interpretability.
[0249] ③The PatchTST model is an effective multivariate time series prediction and self-supervised representation learning model. It divides the time series into subsequences as inputs and has good prediction accuracy in multivariate time series prediction.
[0250] ④The FiLM model can effectively capture long-term dependence relationships by improving the Fourier transform and memory mechanism. It is a model focusing on long-short-term time series prediction.
[0251] ⑤The Informer model is an efficient sequence data processing model. This model can not only capture long dependencies and uncertainties in time series data, but also process large-scale time series datasets, effectively improving the accuracy and efficiency of prediction.
[0252] ⑥The Autoformer model combines the advantages of autoregressive models and Transformer structures. Through autocorrelation modeling and segmented prediction strategies, it effectively processes long sequence data and improves the prediction performance and efficiency of time series.
[0253] ⑦The Reformer model addresses the high computational complexity and huge memory consumption problems when processing long sequence data by introducing local sensitive hashing (LSH) and reversible network (ReversibleLayers) technologies, and has obvious advantages in processing long sequence data.
[0254] ⑧The LightTS model can simultaneously capture long-distance dependence relationships between sequences by focusing on different sequence parts in the input sequence, thus better understanding context information, and is particularly effective when dealing with long-period time series.
[0255] ⑨The TiDE model combines a time encoder and a deep neural network to capture long-term dependencies and non-linear patterns in time series data. Through a time-aware encoding mechanism and a deep learning structure, this model effectively improves the prediction accuracy for complex time series data.
[0256] ⑩The BiLSTM model is a classic RNN time series network. It combines forward and backward LSTM layers and can simultaneously capture bidirectional dependence information in the input sequence.
[0257] The GRU model can selectively remember or ignore information in the input data, better capture important features in the sequence, and thus improve the prediction efficiency and accuracy of the model.
[0258] 2. Experimental Results
[0259] Table 1 and Figure 10 are the experimental results with Guanyuan as the main station and the other stations as slave stations. According to the time correlation of the spread of air pollutants, the time correlation approaches 0 after 80 hours when the air pollutants arrive. Therefore, we use 72 hours as the longest historical duration to conduct prediction experiments on the air pollutant concentrations in the next 12, 24, 36, 48, 60, and 72 hours, and predict six pollutants including CO, NO2, O3, PM10, PM2.5, SO2 and Average pollutant collected by the monitoring stations. This method is compared with the above 9 advanced time series prediction models. In Table 1, the bold data indicates the best performance compared with other models, and the underlined indicates the second best.
[0260] It can be seen from Table 1 that when the prediction durations are 12 hours and 24 hours, this method has the best performance in terms of CO and Average compared with other methods, but it is not the best in terms of the indicators of NO2, O3, PM10 and PM2.5, and most of them are only the second best. As the prediction duration increases to 36 hours, among the 7 indicators, only the performance of the two indicators of O3 and SO2 is not the best, and the prediction performance of the other 5 indicators including the main pollutant PM2.5 has reached the best. When the prediction duration continues to increase to 48 hours, except for the performance of the mae of SO2 and O3 not being the best, the performance of other indicators is still the best. When the prediction time length reaches 60 and 72 hours, except for the SO2 indicator, the prediction performance of the other 6 indicators is the best.
[0261] To more intuitively observe the prediction situations of all models, as Figure 10 (a)-(g) shows, we visualized the MAE indicators of all models. Among them, the X-axis is the prediction sequence length, the Y-axis is different models, and the Z-axis is the predicted value. It can be very intuitively seen from the figure that in the prediction of 12 and 24 hours for the 7 indicators of CO, NO2, O3, PM10, PM2.5, SO2 and Average, the prediction performances of each model fluctuate greatly, and the performance of the ST-KAN-fomer model does not reach the best at this time. After 36 hours, the prediction performances of all models gradually tend to be stable, and the overall performance of the ST-KAN-fomer model is significantly better than that of other models.
[0262] Generally speaking, compared with other models, the overall performance of the present invention is better than that of other comparative models in terms of CO, NO2, O3, PM10, PM2.5 and Average indicators, except for the SO2 indicator. Especially in the main pollutants PM2.5 and the average indicator, the performance of the present invention is significantly better than that of other algorithms at other prediction durations, except that the prediction performance at 12 o'clock does not reach the best.
[0263] The possible reasons why some indicators do not reach the best in the 12-hour and 24-hour prediction durations of this method are as follows: During the process of air pollutants spreading from the sub-station to the main station, they are affected by distance factors, and there is a lag phenomenon in the pollutants spreading from the sub-station to the main station. In addition, the propagation rates of each pollutant are different, resulting in different arrival times of each pollutant at the main station. Only a small amount of air pollutants from the sub-station may have spread to the main station before 24 hours, so the prediction effects of some indicators are not as good as those of other models when using 12 hours and 24 hours as the prediction lengths.
[0264] The reason why the prediction performance of the SO2 indicator is worse than that of other methods is that SO2 is prone to chemical reaction with water vapor during the propagation process to form sulfuric acid mist. During the process of SO2 from the sub-station spreading to the main station, SO2 may gradually react with the water vapor along the way, continuously consuming SO2, making SO2 unable to effectively spread to the main station, thus affecting the prediction accuracy of the model.
[0265] Table 1 Overall performance of the model
[0266]
[0267] 3. Prediction results at different stations
[0268] In order to further verify the generalization ability of this method, the present invention predicts the air pollution concentration in the next 48 hours at three stations, namely Wanliu, Tiantan and Aotizhongxin. The experimental results are shown in Table 2. It can be seen from Table 2 that in the predictions at the three stations, except for the SO2 and O3 indicators which are not as good as other methods, the remaining 5 indicators are better than the other 11 comparative models. From the quantitative indicators, such as Figure 11As shown in the figure, in the prediction of Wanliu, compared with the other 9 models, the MAE of the five indicators of CO, NO2, PM10, PM2.5, and Average of ST-KAN-Former is 10.60%, 5.48%, 3.61%, 6.29%, and 4.43% higher, and the RMSE indicator is 13.06%, 5.76%, 3.52%, 7.61%, and 2.22% higher. In the prediction of Tiantan, compared with the other 9 models, the MAE of the five indicators of CO, NO2, PM10, PM2.5, and Average of ST-KAN-Former is 9.35%, 8.24%, 3.77%, 8.22%, and 12.21% higher, and the RMSE indicator is 10.06%, 8.76%, 1.11%, 7.63%, and 6.30% higher. In the prediction of Aotizhongxin, compared with the other 9 models, the MAE of the five indicators of CO, NO2, PM10, PM2.5, and Average of ST-KAN-Former is 6.28%, 0.78%, 0.47%, 6.89%, and 8.18% higher, and the RMSE indicators of CO, NO2, O3, PM10, PM2.5, and Average are 6.13%, 3.01%, 1.02%, 2.45%, 12.52%, and 4.20% higher. It can be seen that in the prediction of different stations, the present invention is better than other models in most indicators compared with the other 11 models, indicating that the method has good generalization ability.
[0269] Table 2 Comparison of the generalization ability of models
[0270]
[0271] 4. Verification of performance improvement after spatial region division
[0272] In order to verify the improvement of the model performance after dividing the spatial stations into four sub-regions G1, G2, G3, and G4, without changing other parameters and modules, without making any division of all stations in the spatial region, fusing all spatial station data and using it as the input of the S-encoder of the model, and comparing the two methods through experiments. The experimental results are shown in Table 3, where N-R represents that the spatial stations are not divided into regions.
[0273] As shown in Table 3, after dividing the spatial region into G1, G2, G3, and G4, the MAE and RMSE of other indicators except the SO2 indicator are better than those of N-R. For example Figure 12As shown in the figure, in terms of the performance improvement ratio, the MAE of CO, NO2, O3, PM10, PM2.5, and Average increased by 0.77%, 4.06%, 6.06%, 0.39%, 1.51%, and 5.08% respectively, and the RMSE increased by 0.43%, 3.53%, 5.54%, 1.24%, 1.94%, and 2.61% respectively. Only the MAE and RMSE of SO2 decreased by -1.12% and -0.27%. It can be seen that after dividing the space station into regions, the performance of the model has been improved to a certain extent.
[0274] Table 3 Comparison of the influence of regional division on model performance
[0275]
[0276] 5. Ablation experiment
[0277] To verify the effectiveness of each module in the ST-KAN-Former model, an ablation experiment was conducted. As shown in Table 4 and Figure 13 、 Figure 14 In the figure, N-S, N-W, K-L, and P-F in Table 4 represent removing the S-encoder module (N-S), removing the wavelet transform module (N-W), replacing the KAN network with a Linear network (K-L), and replacing the Proattention module with a Fullattention module (P-F) respectively, while keeping all other modules in the ST-KAN-Former model unchanged. In Figure 13 (a)-(d), N-S, N-W, K-L, and P-F represent the percentage of performance improvement after adding the S-encoder module, the percentage of performance improvement after using the wavelet transform module, the percentage of performance improvement of the KAN network compared to the Linear network, and the percentage of performance improvement of the Proattention module compared to the Fullattention module respectively. Specifically:
[0278] 1) In the case of N-S: As shown in Table 4 and Figure 13 (a), from a macroscopic perspective, without the S-encoder module, the performance of the model is greatly affected. Among all the predicted pollutants, only the mae of S02 is higher than that of the ST-KAN-Former model, while the mae and rmse of other indicators are lower than those of the ST-KAN-Former model. From the perspective of quantitative indicators, as Figure 14As shown in (a), after adding the S-encoder module, among the seven predicted indicators of CO, NO2, O3, PM10, PM2.5, SO2, and Average, only the mae of SO2 decreased by -0.56%, and the mae of the other six indicators increased by 11.58%, 7.24%, 5.52%, 11.15%, 15.24%, and 13.99% respectively. Unexpectedly, as Figure 14 shown in (b), the rmse of the seven indicators increased by 10.36%, 6.50%, 5.75%, 8.72%, 13.10%, 1.32%, and 7.31% respectively. It can be seen from this that adding the spatial feature extraction block module S-encoder improves the model performance to a certain extent. The addition of S-encoder enables the model to extract features from both the time and space dimensions, thereby capturing various patterns and trends from local to global, and further improving the model prediction accuracy.
[0279] 2) N-W case: As shown in Table 4 and Figure 13 (c), from a macroscopic perspective, without the wavelet transform module, the model performance is greatly affected. Similarly, among all the predicted pollutants, only the mae of S02 is higher than that of the ST-KAN-Former model, while the mae and rmse of other indicators are lower than those of the ST-KAN-Former model. From the quantitative indicators, as Figure 14 shown in N-W of (a)-(b), after adding the wavelet transform module, among the seven predicted indicators of CO, NO2, O3, PM10, PM2.5, SO2, and Average, only the mae and rmse of SO2 decreased by -1.88% and -0.81% respectively, while the mae of the other six indicators increased by 1.14%, 3.45%, 6.06%, 3.04%, 3.39%, and 5.44% respectively, and the rmse increased by 1.14%, 3.45%, 6.06%, 3.04%, 3.39%, and 5.44% respectively. It can be seen that the addition of the wavelet transform module improves the model performance to a certain extent. The wavelet transform module can extract features from the high-frequency and low-frequency components respectively, and obtain data features with finer granularity, thereby further improving the prediction accuracy.
[0280] (3) K-L case: As shown in Table 4 and Figure 13 (b), from a macroscopic perspective, after replacing the KAN network module with the Linear network module, the model performance is greatly affected. Among all the predicted pollutants, only the mae and rmse of O3 and S02 are higher than those of the ST-KAN-Former model, and the mae and rmse of the remaining five indicators are lower than those of the ST-KAN-Former model. From the quantitative indicators, as Figure 14As shown in K-L of (a)-(b), the performance using KAN is greatly improved compared to using Linear. Among the seven predicted indicators of CO, NO2, O3, PM10, PM2.5, SO2, and Average, the performance of five indicators, namely CO, NO2, PM10, PM2.5, and Average, is improved by 1.52%, 3.91%, 1.16%, 5.00%, and 3.45%, 2.26%, 5.53%, 3.16%, 4.51%, and 1.80% respectively in terms of mae and rmse compared to using the Linear module. Only the mae of O3 and SO2 decreased by -3.01% and -2.08%, and the rmse decreased by -2.65% and -1.08%. It can be seen that the addition of the KAN network improves the model performance to a certain extent. The KAN network uses a spline-parameterized univariate function to replace the traditional linear weights, enabling it to dynamically learn activation patterns, so that the KAN network has a flexible function fitting ability and can more accurately capture complex non-linear relationships and periodic patterns in time series.
[0281] (4) P-F case: As shown in Table 4 and Figure 13 (d), from a macroscopic perspective, after replacing Proattention with the Fullattention module, the model performance is greatly affected. Among all the predicted pollutants, only the mae, rmse of O3 and the mae of SO2 are higher than those of the ST-KAN-Former model, and the mae and rmse of the remaining indicators are lower than those of the ST-KAN-Former model. From the quantitative indicators, as Figure 14 shown in P-F of (a)-(b), the performance using Proattention is greatly improved compared to using Fullattention. Among the seven predicted indicators of CO, NO2, O3, PM10, PM2.5, SO2, and Average, only the mae and rmse of SO2 decreased by -1.50% and -0.27%, and the rmse of O3 decreased by -0.24%. The mae of the other six indicators increased by 0.57%, 0.81%, 0.31%, 0.78%, 2.15%, -1.50%, and 1.37% respectively, and the rmse of the other five indicators increased by 1.00%, 1.20%, 0.50%, 2.52%, and 0.70% respectively. It can be seen that the performance of the model using the Proattention module is better than that of the Fullattention module. ProbSparse self-attention optimizes the calculation by introducing probabilistic sparsity, reduces the time and space complexity, and improves the overall performance of the model.
[0282] Table 4 Influence of different modules on model performance
[0283]
[0284] 6. Parameter Experiment
[0285] The main parameters that have a greater impact on the performance of the ST-KAN-Former model include the number of spatio-temporal feature encoding blocks and decoding blocks, the number of hidden layers in the KAN network, Batch_Size, and diffusion parameters. The following experiments verify the impact of these parameters on the model respectively.
[0286] (1) Number of Feature Encoders
[0287] To verify the impact of the number of feature encoders on the performance of the ST-KAN-Former model, as shown in Table 5, we set the number of feature encoders to 1, 2, and 3 respectively, denoted as E-1, E-2, and E-3. Figure 15 (a)-(b) respectively show the performance improvement of E-1 relative to E-2 and E-3.
[0288] It can be easily seen from Table 5 that compared with E-2, for E-1, except that the performance of SO2 is lower than that of E-2, the other 6 indicators are better than E-2. Compared with E-3, the performance of E-1 is slightly better than that of E-3, and E-1 is better than E-3 in the main pollutants PM2.5 and the average index.
[0289] From the quantitative indicators, as Figure 15 (a) shows, compared with E-2, for E-1, only the mae and rmse of the SO2 index decreased by -2.66% and -0.95%, and the mae of the other 6 indicators of CO, NO2, O3, PM10, PM2.5, and Average increased by 0.76%, 1.60%, 1.21%, 0.78%, 2.15%, and 1.56%, and the RMSE index increased by 0.29%, 2.38%, 0.81%, 1.00%, 0.76%, and 0.84%; as Figure 15 (b) shows, compared with E-3, for E-1, only the MAE of CO decreased by -0.19%, and the MAE and RMSE of SO2 decreased by -2.85% and -0.95%. The MAE performance of NO2, O3, and PM10 is the same, and the MAE of PM2.5 and Average increased by 0.22% and 0.59%; while the RMSE of the 5 indicators of CO, NO2, PM10, PM2.5, and Average increased by 0.29%, 0.73%, 1.24%, 1.65%, and 0.28%. In addition, due to more layers, E-3 has a longer calculation time and lower efficiency compared to E-1. Therefore, the present invention selects E-1.
[0290] Table 5 Influence of Different Numbers of Encoders on Model Performance
[0291]
[0292] (2) Diffusion Parameters
[0293] To verify the influence of the diffusion parameters of the spatial encoder on the performance of the ST-KAN-Former model, as shown in Table 6, N-D indicates that the diffusion parameters are not added to the spatial encoder of the ST-KAN-Former model. From the overall performance, after adding the diffusion parameters to the spatial encoder, the performance is significantly better than the case without adding the diffusion parameters. From the quantization parameters, as Figure 16 shown, the performance of the vast majority of indicators has been improved. The highest percentage increase in performance is 4.65%, and only the rmse of O3 has decreased by -0.36%, and the mae of SO2 has decreased by -0.38%.
[0294] Table 6 Influence of Diffusion Parameters on Model Performance
[0295]
[0296] (3) Number of Hidden Layers in KAN
[0297] To verify the influence of the number of hidden layers in the KAN network on the performance of the ST-KAN-Former model, as shown in Table 7, we set the number of hidden layers in the KAN network to 0, 1, 2, and 3 respectively, denoted by K-0, K-1, K-2, and K-3. It is easy to see from the table that the performance of K-1 is significantly better than that of the other three cases.
[0298] From the quantization indicators, as Figure 17 (a)-(b) shown, K-1-0, K-1-2, and K-1-3 respectively represent the percentage increase in the performance of K-1 relative to K-0, K-2, and K-3. In K-1-0, only the mae of O3 and SO2 has decreased by -2.2% and -2.27%, and the rmse has decreased by -1.19% and -0.67%, and the performance of other indicators has been improved. More gratifyingly, the maximum increase in the mae and rmse of the main indicator PM2.5 is 5.98% and 3.81% respectively. In K-1-2, except for the mae indicator of SO2 decreasing by -0.56%, the performance of all indicators of K-1 is better than that of K-2, and the highest increase is in the most important indicator PM2.5, with the mae and rmse increasing by 9.34% and 5.87% respectively. In K-1-3, the performance of all indicators of K-1 is better than that of K-3, and the highest increase is in the most important indicator PM2.5, with the mae and rmse increasing by 16.8% and 18.9% respectively. Therefore, the present invention selects K-1.
[0299] Table 7 Influence of Different Hidden Layers of KAN Network on Model Performance
[0300]
[0301] (4)Batch Size
[0302] To verify the influence of batch size on the performance of the ST-KAN-Former model, as shown in Table 8, we set the batch size to 16, 32, 64, and 128 respectively, denoted as B-16, B-32, B-64, and B-128.
[0303] It is easy to see from the table that, except for some individual indicators, the performance of B-32 is significantly better than the other three cases. From the quantitative indicators, as shown in Figure 17 (a)-(b), 32-16, 32-64, and 32-128 respectively represent the percentage of performance improvement of B-32 relative to B-16, B-54, and B-128. In 32-16, only the mae of SO2 decreased by -0.19%, and the performance of other indicators of B-32 is better than that of B-16. More gratifyingly, for the main indicator PM2.5, the maximum improvements in mae and rmse are 5.39% and 3.1% respectively. In 32-64, except that the mae and rmse of the SO2 indicator decreased by -1.88% and -0.4%, and the mae indicator performance of PM10 is equal, the performance of other indicators of B-32 is better than that of B-64. In 32-128, except for the decrease in the mae and rmse of SO2, PM10, and the mae indicator of NO2, the performance of other indicators of B-32 is better than that of B-128. Therefore, we choose B-32 for batch size.
[0304] Table 7 Influence of Different Batch Sizes on Model Performance
[0305]
[0306] In summary, the present invention proposes a spatio-temporal prediction model ST-KAN-Former based on Transformer and KAN network. The above experimental results show that, compared with 11 advanced baselines, ST-KAN-Former achieves better overall performance. In the generalization experiment, we applied ST-KAN-Former to 3 different stations, and the effect is still better than the other 11 methods. In the regional division experiment, we kept other parameters and modules unchanged and compared the models with and without regional division. The experimental results show that the model with regional division has a better effect. In the ablation experiment, we verified the effectiveness of each module in ST-KAN-Former by keeping other modules unchanged and only changing one of them. The experiment shows that each module integrated in ST-KAN-Former is helpful for improving the performance of ST-KAN-Former. In the parameter experiment, we also adopted the method of keeping other parameters unchanged and only changing one other parameter to experiment with the number of spatio-temporal feature encoding blocks and decoding blocks, the number of hidden layers in the KAN network, batch_size, and diffusion parameters, and selected the optimal values of each parameter to be applied in ST-KAN-Former.
[0307] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements do not make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A spatio-temporal air quality prediction method based on Transformer and KAN network, characterized in that, It includes the following steps: Step 1: Take the air quality monitoring station to be predicted as the main station, collect the historical feature data of the main station, and fill the historical feature data using a hybrid filling algorithm; Step 2: Divide all the monitoring stations within the area where the main station is located based on the station correlation to obtain multiple sub-areas, and each sub-area contains several monitoring stations; Step 3: Based on the KAN network, wavelet transform, and ProbAttention network, construct a time dimension feature extraction model to extract the time feature data of the main station in the time dimension; Step 4: Based on the KAN network and ProAttention network, construct a spatial feature extraction model, and use this model to perform spatial autocorrelation calculations on the station data of each sub-area respectively to obtain fine-grained spatial feature data; Step 5: Based on the ProbAttention network and KAN network, construct a spatio-temporal feature fusion model, fit the time feature data and the spatial feature data, and output spatio-temporal features; Step 6: Use the spatio-temporal features to predict the air quality of the current station to be measured.
2. The spatio-temporal air quality prediction method based on the Transformer and KAN networks according to claim 1, wherein The data filling process of the hybrid filling algorithm in Step 1 includes: Step 1.1: Find the column j with the largest missing rate in the dataset M to be processed max : Among them, max(·) is the function for taking the maximum value, is.na(·) is the function for counting the number of observations with NA values in the j-th column variable, and nrow(·) is the function for counting the number of samples in M; M i,j , i ∈ [1, n], j ∈ [1, m] represents the element in the i-th row and j-th column of M, n represents the number of sample points, and m represents the total number of variables; P j , j ∈ [1, m] represents the proportion of missing values in the j-th column of observed variables in M; Step 1.2: Generate the initial dataset M′ to be filled through the following formula: wherein, represents removing column j from M max column; Step 1.3: Use the MICE.rf filling method in the Markov chain Monte Carlo method to perform data filling processing on M′ to obtain the complete dataset M c ′; Step 1.4: Concatenate the data of M c ' and to obtain a new data set M″ to be filled Step 1.5: Use the MICE.norm.predict filling method in the Markov chain Monte Carlo method to perform data filling processing on M″ to obtain the complete dataset M c .
3. The spatio-temporal air quality prediction method based on the Transformer and KAN networks according to claim 2, characterized in that, The specific steps for performing area division in Step 2 include: Step 2.1: Take the main station as the center and define the main station as G0; Step 2.2: Take the area where all the monitoring stations with a distance less than 20 km from the main station are located as the G1 area; Step 2.3: Take the area where all the monitoring stations with a distance between 20 km and 40 km from the main station are located as the G2 area; Step 2.4: Take the area where all the monitoring stations with a distance between 40 km and 60 km from the main station are located as the G3 area; Step 2.5: Take the area where all the monitoring stations with a distance greater than 60 km from the main station are located as the G4 area; Step 2.6: Define S0, S1, S2, S3, and S4 as the sets of sites in the G0, G1, G2, G3, and G4 regions respectively. Let n1, n2, n3, and n4 be the number of sites in S1, S2, S3, and S4 respectively. For the primary site In other sub-regions, denote S i,j as the j-th site in the i-th region, where i ∈ (1, 2, 3, 4). If 0 ≤ j ≤ w, w ∈ (n1, n2, n3, n4), then represent the station data within each area as:
4. The spatio-temporal air quality prediction method based on the Transformer and KAN networks according to claim 3, characterized in that, The steps for extracting time feature data in Step 3 include: Step 3.1: In the time dimension of the main site, for the data of the main site at time t Use the data of the previous L in time lengths before time t to predict the air quality at multiple consecutive future time points, and the data of the previous L in time lengths before time t of the main site are as follows: Among them, N P is the number of observation indicators, represents the data collected for the N P th indicator at time t; Step 3.2: Based on wavelet transform, decompose the data of the main site to obtain the low-frequency components and high-frequency components of each index of the main site: Among them, represents the low-frequency component of the main site data ; represents the high-frequency component of the main site data ; H 01 represents the data after wavelet transform; Step 3.3: Integrate H 01 with the time data, and perform Embeding processing using the KAN network to obtain the encoded data H 02 : H 02 = Embeding(concat(H 01 , T1)) = KAN(concat(H 01 , T1)) Among them, concat represents the data fusion operation, and T1 represents the time period (t-L in+1 ,t); Step 3.4: Normalize H 02 to obtain the normalized data H 03 : Among them, LayerMorn represents normalization processing; h n represents the eigenvalue at time t, and Mean represents taking the mean value of h n ; Step 3.5: Use the ProAttention attention mechanism to perform autocorrelation calculation on H 03 : Q = K = V = H 03 Among them, Q, K, and V respectively represent the query, key, and value of the ProbAttention attention mechanism; is a sparse matrix of the same size as q, where q represents each row of Q, and the sparsity metric M(q i , K) is: where q i represents the i-th row in Q; d represents the dimension of the key (K) vector; And, M(q i , K) can be approximately calculated as: Step 3.6: Standardize the data H 04 after autocorrelation calculation to obtain the standardized data H 05 : H 05 = LayerMorn(H 04 ); Step 3.7: Use linear transformation for H 05 to perform dimensionality transformation and output time feature data in the time dimension Among them, Linear represents a linear transformation.
5. A spatio-temporal air quality prediction method based on Transformer and KAN network according to claim 4, characterized in that, The specific operation steps in Step 5 include: Step 5.1: Perform fusion processing on all the site data in the four sub-regions G1, G2, G3, and G4 to obtain the fused data for these four sub-regions respectively Among them, l1, l2, l3, and l4 respectively represent the distances from the central points of the G1, G2, G3, and G4 sub-regions to the main site, and ε1, ε2, ε3, and ε4 respectively represent the diffusion coefficients of the G1, G2, G3, and G4 regions; S 1,n1 represents the n1th site within the G1 sub-region, S 2,n2 represents the n2th site within the G2 sub-region, S 3,n3 represents the n3th site within the G3 sub-region, S 4,n4 represents the n4th site within the G4 sub-region; Step 5.2: The fusion data of each region is fused with the time data T, and the fused data is processed by the KAN network for Embedding to obtain the encoded data of each region Among them, i ∈ (1, 2, 3, 4) represents the 4 area numbers in space; Step 5.3: For each perform standardization processing, and use the ProAttention attention mechanism to calculate self-correlation to obtain the calculated data Among them, Step 5.4: Normalize and then use the Linear network to align the dimensionality of the normalized data features and output the spatial feature data of each region after dimensional alignment 6. The spatio-temporal air quality prediction method based on the Transformer and KAN networks according to claim 5, wherein, The specific steps in Step 6 include: Step 6.1: Fuse the time feature data and the spatial feature data to form a fused spatio-temporal feature matrix H de1 : H de1 = Concat(H out0 , H out1 , H out2 , H out3 , H out4 ) Among them, H out0 is the time feature data; H out1 , H out2 , H out3 , H out4 are the spatial feature data of five regions G1, G2, G3, and G4 respectively; Step 6.2: Use the ProAttention attention mechanism to calculate the correlation of spatio-temporal feature numbers to obtain the calculated data H de2 : H de2 = ProbAttention(Q de , K de , V de ) Among them, Q de = K de = V de = H de1 ; Step 6.3: Based on the KAN network, fit H de2 and output the spatio-temporal feature H de3 after fitting: H de3 = KAN(H de2 ) Among them, KAN represents the KAN network; Step 6.4: Based on the spatio-temporal feature H de3 , output the prediction result Y: Y = Projection(H de3 ).
Citation Information
Cited By
Air quality mode meteorological field generation method and system
CN120671936A
Cross-modal ultrasound contrast image generation method and device based on diffusion model and readable storage medium thereof
CN121544731A