PM2.5 prediction method and system based on spatial-temporal characteristics

Through multi-scale decomposition and feature enhancement mapping, the trend and seasonal components of PM2.5 are extracted, combined with graph attention neural network and gated network, the uncertainty and information redundancy problems of the existing PM2.5 prediction model at hourly level prediction are solved, achieving higher prediction accuracy and robustness.

CN120030339AActive Publication Date: 2025-05-23TIANJIN NORMAL UNIVERSITY +1

Patent Information

Application Number
CN202510519360.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-23
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

There is uncertainty in the prediction of the existing PM2.5 prediction model at hourly level. Information redundancy and information loss affect the model performance and cannot fully reflect the changing trend of PM2.5 concentration.

Method used

The PM2.5 prediction method based on spatiotemporal features is adopted to extract trend components and seasonal components through multi-scale decomposition, feature enhancement mapping and linear prediction. The graph attention neural network focuses on the spatial relationship of data between monitoring sites, and dynamically fuses time and spatial prediction results through the gated network.

Benefits of technology

It significantly improves the accuracy of PM2.5 on the hour-level prediction task, enhances the ability to express spatiotemporal features, and improves the accuracy and robustness of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030339A_ABST
    Figure CN120030339A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of atmospheric pollutant prediction, and particularly discloses a PM2.5 prediction method and system based on spatial-temporal characteristics, and the method comprises the steps: obtaining atmospheric pollutant data and meteorological data of a target station and surrounding stations; processing the data of the peripheral sites, and then splicing the data with the data of the target site to obtain to-be-analyzed data; performing normalization processing on the data to be analyzed, decomposing normalized time series data into trend features and seasonal features, performing feature enhancement mapping and linear prediction respectively, and performing addition to obtain a time prediction result; according to the normalized time sequence data, constructing a graph structure by using a mutual information method, and obtaining a space prediction result through a graph attention neural network; and dynamically fusing the two prediction results through a gating network to obtain a predicted value of the PM2.5 concentration of the target station. According to the method, the precision of PM2.5 on an hour-level prediction task is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of atmospheric pollutant prediction, and in particular to a PM2.5 prediction method and system based on spatiotemporal characteristics. Background Art

[0002] Existing PM2.5 prediction models are often affected by multiple factors when dealing with hourly predictions, including meteorological conditions, geographical location, and time changes. The diversity of these factors leads to uncertainty in the prediction results. At the same time, information redundancy and information loss will also affect the performance of the model, making it unable to fully reflect the changing trend of PM2.5 concentration.

[0003] The Chinese patent application with the publication number CN114694767A discloses a PM2.5 concentration prediction method based on a space-time graph ordinary differential equation network. Through the Gaussian diffusion model, the adjacency matrix is ​​constructed in combination with the Euclidean distance and wind direction data of the monitoring station, and the gas monitoring stations in the industrial park are represented in the form of a graph. After processing the air humidity data, a space-time graph ordinary differential equation network model is constructed, and finally the PM2.5 concentration data, adjacency matrix and air humidity data are input into the model for training. This patent only combines wind direction data.

[0004] The Chinese patent application with the publication number CN114662791A discloses a long-time series PM2.5 prediction method and system based on spatiotemporal attention. First, the processed data is input into the feature extraction network for feature extraction, and then the features extracted from different sites are connected and fused using the spatial attention network. Then, the past features are obtained through a multi-layer bidirectional LSTM. Finally, the known future feature data of the time period to be predicted is extracted through a neural network, and the final prediction result is obtained after connection. Although this patent takes into account the impact of nearby meteorological data on the predicted area when predicting PM2.5, the method used does not consider the essential relationship between the data of nearby sites and the predicted site. Summary of the invention

[0005] The present invention aims to solve the above problems. To this end, the present invention provides a PM2.5 prediction method and system based on spatiotemporal features, which effectively extracts trend components and seasonal components through multi-scale decomposition, feature enhancement mapping and linear prediction, thereby capturing the change law of PM2.5 concentration; at the same time, a graph attention neural network is introduced to focus on the data spatial relationship between different monitoring sites, and improve the expression ability of spatial features; finally, the time prediction results and spatial prediction results are integrated through a gated network, which significantly improves the accuracy of PM2.5 in hourly level prediction tasks.

[0006] The present invention provides a PM2.5 prediction method based on spatiotemporal characteristics, and the technical solution adopted is as follows: Obtain air pollutant data and meteorological data of the target site and surrounding sites; The air pollutant data and meteorological data of the surrounding stations are processed using principal component analysis and geographic distance weighting, and then spliced ​​with the air pollutant data and meteorological data of the target station to obtain the data to be analyzed; Normalize the data to be analyzed to obtain normalized time series data; The normalized time series data is decomposed into trend features and seasonal features, and the trend component prediction results and seasonal component prediction results are obtained by feature enhancement mapping and linear prediction respectively, and the time prediction results are obtained by adding them together; Based on the normalized time series data, the mutual information method is used to construct the graph structure of the PM2.5 sequence, and the spatial prediction results are obtained through the graph attention neural network. The temporal prediction results and spatial prediction results are dynamically fused through the gating network, and after denormalization, the predicted value of PM2.5 concentration at the target site is obtained.

[0007] Furthermore, the atmospheric pollutant data include fine particulate matter, inhalable particulate matter, sulfur dioxide, nitrogen dioxide, carbon monoxide and ozone, and the meteorological data include temperature, air pressure, dew point temperature, wind direction and wind speed level.

[0008] Furthermore, in the principal component analysis, one principal component with the largest variance was selected; The calculation process of geographic distance weighting is: Calculate the geographic distance between the target site and each surrounding site; Calculate weights based on geographic distance; The principal component analysis results of each surrounding station are weighted and summed according to their weights to obtain the pollutant and meteorological data of the surrounding stations after dimensionality reduction.

[0009] Furthermore, the normalization process is: Calculate the mean and standard deviation of the data to be analyzed; Use the mean and standard deviation to normalize the data to be analyzed to obtain normalized data; Add noise to the normalized data; The normalized data with added noise is scaled and translated to obtain the normalized time series data.

[0010] Furthermore, the standard deviation of the noise is adjusted dynamically, and the calculation formula of the noise standard deviation is: in, represents the noise standard deviation, represents the value of the loss function, It means to find the maximum value, Indicates finding the minimum value.

[0011] Furthermore, the process of decomposing the normalized time series data into trend characteristics and seasonal characteristics is as follows: The normalized time series data are smoothed by moving average kernels of multiple scales. The moving average of each kernel is calculated and the mean of multiple moving averages is used as the trend feature. The normalized time series data and trend features are differentiated to obtain seasonal features. The kernel sizes are 17, 25, 33, 49, 65 and 73 respectively.

[0012] Furthermore, the calculation formula of the feature enhancement map is: in, represents the output tensor, represents the input tensor, represents the weight matrix, represents the bias matrix, represents the batch size, represents the time step, represents the input channel dimension, Indicates the number of output features mapped to each input channel, and c represents the cth dimension.

[0013] Furthermore, the prediction process of the spatial prediction result is: Based on the normalized time series data, the graph structure of PM2.5 series is constructed using the mutual information method; Use symmetric normalization to process the adjacency matrix of the graph structure to obtain a symmetric normalized adjacency matrix; The node feature matrix of the graph structure and the symmetrically normalized adjacency matrix are input into the graph attention neural network to predict the spatial prediction results.

[0014] Furthermore, in the prediction process of the graph attention neural network, the node feature matrix is ​​processed using sine and cosine position encoding.

[0015] The present invention also provides a PM2.5 prediction system based on spatiotemporal features, and the technical scheme adopted is as follows: comprising: a data acquisition module, a data processing module, a feature normalizer, a time prediction module, a space prediction module and a gated prediction module, A data acquisition module is used to obtain atmospheric pollutant data and meteorological data of the target site and surrounding sites; A data processing module is used to process the air pollutant data and meteorological data of the surrounding stations using principal component analysis and geographic distance weighting, and then splice them with the air pollutant data and meteorological data of the target station to obtain the data to be analyzed; Feature normalizer, used to normalize the data to be analyzed to obtain normalized time series data; The time prediction module is used to decompose the normalized time series data into trend features and seasonal features, and obtain the trend component prediction results and the seasonal component prediction results through feature enhancement mapping and linear prediction respectively, and add them to obtain the time prediction result; The spatial prediction module is used to construct the graph structure of the PM2.5 sequence based on the normalized time series data using the mutual information method, and predict the spatial prediction results through the graph attention neural network; The gated prediction module is used to dynamically fuse the temporal prediction results and the spatial prediction results through the gated network, and obtain the predicted value of the PM2.5 concentration of the target site after denormalization.

[0016] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects: 1. The present invention proposes a data fusion method based on principal component analysis (PCA) and geographic distance weighting to improve the prediction accuracy of PM2.5 concentration. First, the pollutant and meteorological data of multiple surrounding stations are synchronously sorted, and the PCA method is used to project the data to the direction with the largest variance, thereby effectively reducing the data dimension and retaining key information. Next, the longitude and latitude distances between each surrounding station and the target station are calculated, and these distances are used as weighting factors. Stations with closer distances are given higher weights, and stations with farther distances have lower weights. In this way, the influence of surrounding stations is adaptively adjusted according to spatial location. Finally, the weighted PCA results are fused with the data of the target station, and compressed to one-dimensional data through scaling processing to ensure data consistency and comparability. This method makes full use of the spatial information of surrounding stations, enhances the spatiotemporal prediction ability of the target station, and improves the accuracy and robustness of PM2.5 concentration prediction.

[0017] 2. The present invention designs a trend seasonal decomposer to perform multi-dimensional decomposition on the input normalized time series data, extracting long-term trends and seasonal changes layer by layer to adapt to the feature changes in different time dimensions. The process uses six convolution kernels of different sizes (17, 25, 33, 49, 65 and 73) to flexibly cope with the diversity of data. This not only effectively removes noise and retains valuable information, but also enables the model to accurately predict the changing trend of PM2.5 data in different dimensions.

[0018] 3. The present invention performs feature enhancement mapping on trend features and seasonal features and maps them to more dimensions. Then, by integrating various components, a comprehensive feature representation is obtained. Feature mapping integration can better capture complex nonlinear relationships and improve the accuracy of prediction.

[0019] 4. The present invention combines position encoding, graph attention neural network and pollutant data from multiple nearby surrounding areas to more accurately capture the mutual influence relationship between the characteristics affecting PM2.5 concentration in the region, thereby improving the accuracy of spatial prediction.

[0020] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 It is a flow chart of the method provided by the present invention.

[0023] Figure 2 It is a comparison result diagram of the predicted value and the true value provided by the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0025] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0026] Combine the following Figure 1 to Figure 2The present invention is further described in detail, and a PM2.5 prediction method and system based on spatiotemporal characteristics of the present invention are described: In this embodiment, Figure 1 As shown, a PM2.5 prediction method based on spatiotemporal characteristics is provided, comprising the following steps: Step 1: Obtain air pollutant data and meteorological data for the target site and surrounding sites.

[0027] In this embodiment, the atmospheric pollutant data includes fine particulate matter, inhalable particulate matter, sulfur dioxide, nitrogen dioxide, carbon monoxide and ozone, and the meteorological data includes temperature, air pressure, dew point temperature, wind direction and wind speed. In this embodiment, 7 surrounding stations around the target station are selected.

[0028] The atmospheric pollutant and meteorological data of the target site and the seven surrounding sites are all in the same time period, which ensures the temporal consistency of the data.

[0029] Step 2: Use principal component analysis and geographic distance weighting to process the air pollutant data and meteorological data of surrounding stations, and then splice them with the air pollutant data and meteorological data of the target station to obtain the data to be analyzed.

[0030] Step 2.1: Principal component analysis (PCA) is used to reduce dimension and extract key features, with the goal of reducing redundant information and projecting the data to the direction with the largest variance. In this way, the most useful variation information between stations is retained, making the data expression more concise while reducing the computational complexity of the model. For the atmospheric pollutants and meteorological data of each surrounding station, principal component analysis is performed separately, and the principal component with the largest variance is selected to retain the most critical features of the data. The principal component analysis results of each surrounding station are obtained.

[0031] Step 2.2: The calculation process of geographical distance weighting is: Calculate the geographical distance between the target site and each surrounding site. In this embodiment, the spherical distance between the target site and the surrounding sites is calculated using longitude and latitude.

[0032] The geographical distance is used as a weighting factor, and the weight is calculated based on the geographical distance. This embodiment adopts an inverse weight strategy, that is, the closer the distance, the greater the weight. The weight calculation formula is: in, It is The weight of surrounding sites, It is The geographical distance between the surrounding sites and the target site, is the weight decay coefficient.

[0033] After calculating the weights of each surrounding station, these weights are used to fuse the weighted principal component analysis results. Weighted PCA fusion: The principal component analysis results of each surrounding station are weighted and summed according to their weights to obtain the pollutant and meteorological data of the surrounding stations after dimensionality reduction. , that is, the weighted eigenvector, is calculated as: in, It is The principal component analysis results of the surrounding sites.

[0034] Step 2.3: Final feature fusion: The pollutant and meteorological data of the surrounding stations after dimensionality reduction are merged with the atmospheric pollutant data and meteorological data of the target station using the splicing method to form a new feature vector and obtain the data to be analyzed.

[0035] This embodiment uses multiple variables to predict a single variable to predict the PM2.5 concentration at the hourly level. The 11-dimensional data of the target site and the reduced-dimensional data obtained by principal component analysis and geographic distance weighting of the surrounding sites are used to construct the data to be analyzed and predict the PM2.5 concentration at the hourly level of the target site. This embodiment uses 18-dimensional data to construct the data to be analyzed, as shown in Table 1.

[0036] Table 1 Variable table

[0037] Step 3: Normalize the data to be analyzed to obtain normalized time series data.

[0038] This method designs a feature normalizer to normalize the data to be analyzed. The feature normalizer is an adaptive normalization layer that aims to address the limitations of traditional normalization methods when dealing with different time dimensions and data distribution changes in PM2.5 prediction. Traditional normalization methods (such as standardization and normalization) often have limited effects when facing multi-dimensional and complex data and cannot adapt to the dynamic changes of time series data. The feature normalizer adjusts the data to adapt to different time dimensions by dynamically calculating the mean and standard deviation of the input data. In particular, in PM2.5 data, the concentration of pollutants is affected by many factors and has significant time dependence. In addition, it can dynamically adjust the noise standard deviation during the normalization process to enhance the robustness of data processing. This method not only improves the normalization effect of PM2.5 time series data, but also uses the saved statistical information and affine parameters to restore the original distribution of the data during the denormalization process, ensuring the accuracy and stability of the model when dealing with complex time series data.

[0039] The process used in the feature normalizer is as follows: Step 3.1: Calculate the mean and standard deviation of the data to be analyzed.

[0040] For PM2.5 series data, accurately calculating the mean and standard deviation helps capture the overall trend of the data and reduce the impact of changes in the time dimension. The formulas for the mean and standard deviation are as follows: in, For the data to be analyzed, is the mean of the data to be analyzed, To find the mean, is the standard deviation of the data to be analyzed, is a hyperparameter, which is a small value used to prevent the denominator from being zero. is the i-th parameter of the data to be analyzed.

[0041] Step 3.2: Use the mean and standard deviation to normalize the data to be analyzed and obtain the normalized data.

[0042] The goal of the normalization process is to transform the data to be analyzed Convert to data with zero mean and unit variance. Using normalization can reduce the interference of noise and outliers in PM2.5 series data on model training and improve the convergence speed and stability of the model. The normalization process formula is as follows: in, Represents normalized data.

[0043] Step 3.3: Add noise to the normalized data.

[0044] Adding noise can prevent the model from overfitting, especially when processing PM2.5 data, which helps the model better generalize to unseen data. This can further enhance the robustness of the model. The formula for adding noise is as follows: in, represents the noise standard deviation, Representation and Standard normally distributed noise with the same shape, represents the normalized data with noise added.

[0045] During the normalization process, the noise standard deviation is dynamically adjusted according to the value of the loss function to adapt to the changes in PM2.5 data in different time periods, improve the adaptability of the model in different time dimensions, and enhance the robustness of the model. Dynamically adjust the standard deviation of the noise. The calculation formula for the noise standard deviation is: in, represents the value of the loss function, It means to find the maximum value, Indicates finding the minimum value.

[0046] The loss function refers to the error value calculated by the model during the training process. Commonly used loss functions include Mean Absolute Error (MAE) and Mean Squared Error (MSE). These two indicators are of great significance in measuring the accuracy and stability of the prediction model. In this embodiment, MAE is selected as the loss function.

[0047] The mean absolute error is used to measure the average absolute difference between the predicted value and the actual value. It is the average value of the prediction errors and reflects the degree of deviation of the prediction results. The smaller the MAE, the closer the model's prediction results are to the actual values ​​and the smaller the error. The mean square error (MSE) is used to measure the average squared difference between the predicted value and the actual value. It is the average of the sum of the squares of the prediction errors and reflects the degree of fluctuation of the prediction results. The smaller the MSE, the closer the model's prediction results are to the actual values ​​and the smaller the fluctuation. Compared with MAE, MSE is more sensitive to outliers because the errors are amplified after being squared.

[0048] The core idea of ​​dynamically adjusting noise is to control the size of noise through the value of the loss function. Specifically: When the loss is large, it means that the prediction effect of the model at the current stage is poor, and the noise standard deviation needs to be increased to prevent the model from being too dependent on certain fixed patterns, thereby improving the generalization ability of the model.

[0049] When the loss is small (i.e. the model performs better), the noise standard deviation can be reduced, allowing the model to focus more on the key features of the data and avoid excessive introduction of unnecessary interference.

[0050] A threshold can be set based on experience to determine whether the loss is large or small.

[0051] Step 3.4: Affine transformation: Affine transformation is used to scale and translate data. Affine transformation enables the model to flexibly adjust the data distribution and improve the effect of feature extraction, especially in PM2.5 prediction, which can better capture complex patterns. Scaling and translating the normalized data with added noise to obtain the normalized time series data The specific formula is as follows: in, represents the scaling factor, represents the amount of translation, and is a learnable parameter.

[0052] Step 4: Decompose the normalized time series data into trend features and seasonal features, and obtain the trend component prediction results and seasonal component prediction results through feature enhancement mapping and linear prediction respectively, and add them together to obtain the time prediction result.

[0053] Step 4.1: The normalized time series data is decomposed into trend features and seasonal features through multi-scale convolution kernels and difference processing.

[0054] The specific process is: The normalized time series data were smoothed using 6 moving average kernels of different scales. The moving average of each kernel was calculated and the mean of the 6 moving averages was used as the trend feature. The normalized time series data and trend features were differentiated to obtain seasonal features. The kernel sizes were 17, 25, 33, 49, 65 and 73, respectively.

[0055] This method designs a multi-scale seasonal trend decomposer. In the PM2.5 prediction task, the design of the multi-scale seasonal trend decomposer aims to enhance the model's understanding and prediction capabilities of complex data patterns through an efficient time series decomposition method. PM2.5 concentration data are usually affected by multiple factors, showing significant seasonal fluctuations, trend changes, and high noise characteristics. Traditional time series decomposition methods, such as the classic STL decomposition method and exponential smoothing method, although they perform well when processing stationary series, are difficult to effectively capture long-term trends and short-term seasonal fluctuations when faced with complex and noisy PM2.5 data, resulting in limited prediction accuracy. Therefore, it is particularly important to adopt a multi-scale decomposition method that adapts to the characteristics of complex data.

[0056] The multi-scale seasonal trend decomposer overcomes the shortcomings of traditional methods in capturing complex fluctuations by introducing the smoothing of multi-scale convolution kernels. In this embodiment, six convolution kernels of different sizes are used, with kernel sizes of 17, 25, 33, 49, 65, and 73, respectively, to perform multi-scale smoothing on the input normalized time series data. Smaller convolution kernels (17, 25, and 33) can remove high-frequency noise and extract short-term seasonal changes; larger convolution kernels (49, 65, and 73) are more suitable for extracting long-term trend components and revealing the overall change trend of PM2.5 concentration over a longer time span. By combining these convolution kernels of different scales, the multi-scale seasonal trend decomposer can efficiently extract the key features of data at multiple time scales, thereby achieving accurate decomposition of complex time series. In this embodiment, the normalized time series data is passed through 6 moving average kernels of different scales to obtain 6 moving averages, and then these moving averages are averaged to obtain the final trend characteristics.

[0057] After extracting the trend component of the normalized time series data, the multi-scale seasonal trend decomposer extracts the seasonal component from the data through a difference operation. The core idea of ​​this operation is to remove the trend component from the sequence and retain the periodic fluctuations. That is, subtract the trend feature from the normalized time series data to obtain the seasonal feature.

[0058] The difference operation can effectively eliminate the impact of long-term trends and highlight the periodic characteristics, thereby more accurately capturing the seasonal fluctuations of PM2.5 concentration data. Overall, the multi-scale seasonal trend decomposer significantly improves the performance of the model in complex time series analysis through the combination of multi-scale convolution and difference technology, especially in capturing the trend changes and seasonal characteristics of PM2.5 concentration data, showing strong adaptability and prediction effect. This separation of trend and seasonality enables the model to better understand the inherent laws in the data, thereby improving the prediction accuracy of PM2.5 concentration.

[0059] By adopting a multi-dimensional smoothing method, the multi-scale seasonal trend decomposer significantly improves the model's ability to understand time series data and enhances its ability to predict complex patterns. The decomposer can effectively deal with high noise, nonlinear characteristics and seasonal fluctuations in PM2.5 data, making the prediction results more accurate and reliable. The modular design ensures that the trend components and seasonal components extracted from different dimensions can provide more refined features for subsequent prediction models, improving the adaptability and robustness of the model. Therefore, the application of the multi-scale seasonal trend decomposer in PM2.5 concentration prediction not only optimizes the performance of the model in processing complex nonlinear data, but also demonstrates high efficiency and accuracy in actual scenarios, providing strong technical support for PM2.5 concentration prediction.

[0060] Step 4.2: This method designs a feature enhancement mapper to implement feature enhancement mapping. The feature enhancement mapper linearly transforms the input data through the weight matrix and the bias matrix, and then uses the Einstein sum convention (einsum), which can efficiently multiply the input features with the weight matrix to generate a new high-dimensional feature representation. The shape of the input data is (B, T, C), where B is the batch size, T is the time step, and C is the input channel dimension. Through matrix multiplication, the shape of the weight matrix is ​​(C, F), which is mapped to the F features corresponding to each channel, generating features, thus expanding the dimension of the input. This design enables the extraction of potential complex patterns from the raw data, which helps capture the long-term trend and short-term fluctuations of PM2.5 concentration. After the linear transformation, the Einstein summation convention is used to fuse the features generated by the mapping one by one to form a more comprehensive feature.

[0061] The calculation formula of feature enhancement map is: in, represents the output tensor, represents the input tensor, represents the weight matrix, represents the bias matrix, represents the batch size, represents the time step, represents the input channel dimension, Indicates the number of output features mapped to each input channel, and c represents the cth dimension.

[0062] In this embodiment, after the trend feature is subjected to feature enhancement mapping, a mapped trend feature is obtained; after the seasonal feature is subjected to feature enhancement mapping, a mapped seasonal feature is obtained.

[0063] In this embodiment, F is determined by the ratio of the sequence length of the input tensor to the sequence length of the output tensor. This mapping method can improve the expressiveness of features and enable the model to effectively extract information at different time scales. Since PM2.5 concentration data has nonlinear and time-varying characteristics, feature enhancement mappers are particularly important when integrating complex features. The calculation formula is: in, Indicates the sequence length of the input tensor. In this embodiment is 96; Indicates the sequence length of the output tensor. In this embodiment 1, 3, 6, 9, 12, 24; Indicates rounding down. The feature dimension can be adjusted to help the model adapt to data of different sizes. In this embodiment, is 64. This embodiment designs the formula to use input data to adjust the value of F. Generally, the value of F is an integer, and preferably, the value of F is 1-10.

[0064] Step 4.3: The output tensor output by the feature enhancement mapper is subjected to linear prediction to obtain the prediction result. That is, the trend feature after mapping is subjected to linear prediction to obtain the trend component prediction result; the seasonal feature after mapping is subjected to linear prediction to obtain the seasonal component prediction result.

[0065] In this embodiment, linear prediction is performed using a fully connected linear layer.

[0066] Step 4.4: Add the trend component prediction results and the seasonal component prediction results to obtain the time prediction results.

[0067] Step 5: Based on the normalized time series data, the mutual information method is used to construct the graph structure of the PM2.5 sequence, and the spatial prediction results are obtained through graph attention neural network prediction.

[0068] Step 5.1: First, convert the normalized time series data into a graph structure. The nodes of the graph structure represent sites, and the edges represent the correlation between sites. In this embodiment, the number of nodes is 8, corresponding to the target site and 7 surrounding sites. Based on the normalized time series data, the graph structure of the PM2.5 sequence is constructed using the mutual information method. Mutual information is a statistic that measures the correlation between random variables. In this model, the adjacency matrix of the graph structure is constructed based on mutual information, which captures the spatial dependencies between different sites. The adjacency matrix defines which nodes have strong correlations.

[0069] Step 5.2: Since the degrees of different nodes (the number of connections with other nodes) may vary greatly, directly using the adjacency matrix may lead to instability in the training process. Therefore, symmetric normalization is used to process the adjacency matrix of the graph structure to obtain a symmetric normalized adjacency matrix to ensure that the input signal of each node can be transmitted in a balanced manner. Symmetrically normalized adjacency matrix The calculation formula is: in, The adjacency matrix representing the graph structure, represents the degree matrix, is a diagonal matrix, the diagonal elements of the degree matrix are the nodes The degree, Representation and Node The number of directly connected edges, Degree matrix The inverse of the square root of Represents the nodes in the adjacency matrix and nodes edge.

[0070] Step 5.3: Input the node feature matrix of the graph structure and the symmetrically normalized adjacency matrix into the graph attention neural network to predict the spatial prediction results.

[0071] Sin-cosine position encoding is introduced into the graph attention neural network to add position encoding to the input features, thereby enhancing the expression ability of spatiotemporal information. After processing the node features, the adaptive attention mechanism can focus on the key information of the connected nodes, thereby effectively capturing the features that influence each other in space. Finally, through the weighted aggregation of the node features, the linear layer outputs the predicted value of PM2.5 concentration. Weighted aggregation dynamically adjusts the influence of neighboring nodes according to the dependencies between regions, improving the prediction performance of the model in complex environments.

[0072] First, the node feature matrix is ​​processed using sine and cosine position encoding. The purpose of introducing sine and cosine position encoding is to provide more contextual information for each node and help the model better understand the relationship between nodes. Sine and cosine position encoding not only provides rich position information, but also effectively captures the relative position relationship in the sequence data, further improving the model's ability to handle time series changes.

[0073] The formula for calculating the sine of an even dimension is: in, means selecting all columns with even dimensions in the positional encoding matrix, Column vector representing position indices, ranging from 0 to , Indicates the maximum sequence length. Represents the calculated scaling factor to ensure that the encoding amplitude of each dimension is appropriate, and its value decreases as the dimension increases. Represents the scaling factor used to adjust the amplitude of the sine value to fit the length of the entire position encoding.

[0074] The formula for calculating cosine of odd dimensions is: in, Indicates that all columns of odd dimensions are selected in the position encoding matrix. By combining the cosine function and the sine function, the encoding of each position is both periodic and variable in the high-dimensional space.

[0075] The node features after sine and cosine position encoding are mapped to the new space, and the attention weights are calculated based on the connection relationship defined by the symmetric normalized adjacency matrix. Using the attention weights, the features of neighboring nodes are weighted and aggregated, and the linear layer outputs the spatial prediction results (the output of the graph attention network).

[0076] In the spatial prediction of PM2.5, the adjacency matrix constructed based on mutual information provides a basis for the modeling of spatial relationships in graph neural networks. Specifically, the adjacency matrix clarifies which nodes in the graph (data from monitoring stations) have strong correlations, and GAL uses these adjacency relationships to dynamically assign weights and determine the influence of neighbor nodes on the central node. This method can capture the key correlations in the spatial distribution of PM2.5 and provide strong support for the prediction of PM2.5 concentrations between regions.

[0077] The attention mechanism further calculates the attention weights between each pair of connected region data , and perform weighted aggregation on the features of the neighbors, the formula is as follows: in, Is with the node The set of connected neighbor nodes, Representation Node The characteristic vector of Representation Node The feature vector of . Represents a join operation. is a learnable parameter vector, usually used to calculate the attention score, and T represents the transpose. is used The transpose of is multiplied by the concatenated eigenvector to obtain a scalar value, which is used to measure the node and nodes The similarity or correlation between them. Indicates the number of neighbor nodes of a node. represents a learnable parameter matrix that maps input features to a new representation space. Represents the activation function.

[0078] Step 6: Dynamically fuse the temporal prediction results and the spatial prediction results through the gating network, and after denormalization, obtain the predicted value of PM2.5 concentration at the target site.

[0079] In a gating network, the gating mechanism outputs values ​​between 0 and 1 (such as calculated by the sigmoid function). These values ​​are used to adjust the weight matrix in the model, such as amplifying, reducing, or selectively filtering the weights. Such a mechanism allows the model to dynamically adjust its weights at runtime, thereby adapting to the data more flexibly.

[0080] The calculation process of the gating mechanism: First, the gating network calculates the control coefficients through the temporal prediction results and the spatial prediction results : in, is the weight matrix of the gating network, which is responsible for extracting information from the temporal and spatial prediction results. is the time prediction result, is the spatial prediction result. σ is the sigmoid activation function, ensuring The output is between 0 and 1. is the bias term of the gating network.

[0081] when When it is close to 1, it means the time prediction result The importance is higher; when When it is close to 0, it means the spatial prediction result More important.

[0082] Secondly, after calculating the control coefficient Finally, the model fuses the temporal and spatial prediction results in a weighted manner to generate the final fusion result. : Among them, when When it is close to 1, the model relies more on the time prediction results. When it is close to 0, the model makes more use of spatial prediction results. The gating network adjusts the integration weights of temporal and spatial prediction results based on the input data.

[0083] Gated Network and It will be continuously updated with the back propagation during the model training process. Specifically, the model optimizes these weights and biases by minimizing the loss function during training, so that the model can dynamically adjust the fusion weights of time and space features according to different input data.

[0084] Time feature weight: In the period when seasonal changes are more significant, the time feature may be more important, so will be close to 1, and the model will rely more on the time prediction results. At this time, the gating network will strengthen the weight of the time feature, making the time prediction results Dominant in.

[0085] Spatial feature weights: When the air quality differences between monitoring sites (target site, surrounding sites) are large, spatial features may dominate, so will be close to 0, and the model will rely more on the spatial prediction results. At this time, the gating network adjusts the weights so that the spatial prediction results are occupies a dominant position in.

[0086] Adjust weights: In the actual PM2.5 prediction task, the gating network will dynamically adjust according to the characteristics of the input data For example, in some time periods, seasonal fluctuations make temporal features more important, while in other time periods, differences between monitoring sites lead to higher weights for spatial features. The gating network adjusts according to these dynamic changes. , so that the model can make predictions in the way that best suits the current situation at any moment.

[0087] After the denormalization process, the predicted value of PM2.5 concentration at the target site is obtained. The denormalization process is implemented in the opposite way of step 3.

[0088] In order to verify the effectiveness of this method, this example selects the DRSA dataset, which is hourly air pollutant data from 12 nationally controlled air quality monitoring points. This example takes the prediction of Guanyuan area as an example, and only uses the data of Guanyuan area and the data of the nearest 7 monitoring points in its surrounding geographical locations. In the experiment, BI-LSTM, Transformer, Autoformer, TEFN, TSMixer and other time series prediction models and the spatiotemporal model proposed by this method were selected for experiments, aiming to compare the performance of different models in PM2.5 concentration prediction. The indicators selected were MAE, RMSE (Root Mean Squared Error) and R 2 (R-squared, determination coefficient), the experimental results are shown in Table 2. The comparison results of the predicted value and the true value of this method in the next 116 time steps are shown in Figure 2 shown.

[0089] Table 2 Comparison of the results of spatiotemporal model and other time series prediction models

[0090] This embodiment also provides a PM2.5 prediction system based on spatiotemporal features, and the technical solution adopted is as follows: comprising: a data acquisition module, a data processing module, a feature normalizer, a time prediction module, a space prediction module and a gate prediction module, The data acquisition module is used to obtain the atmospheric pollutant data and meteorological data of the target site and surrounding sites.

[0091] The data processing module is used to process the air pollutant data and meteorological data of the surrounding stations using principal component analysis and geographic distance weighting, and then splice them with the air pollutant data and meteorological data of the target station to obtain the data to be analyzed.

[0092] The feature normalizer is used to normalize the data to be analyzed to obtain normalized time series data.

[0093] The time prediction module is used to decompose the normalized time series data into trend features and seasonal features, and obtain the trend component prediction results and the seasonal component prediction results through feature enhancement mapping and linear prediction respectively, and add them to obtain the time prediction results. The time prediction module mainly includes a multi-scale seasonal trend decomposer, a feature enhancement mapper and a fully connected linear layer. The multi-scale seasonal trend decomposer is used to decompose the normalized time series data into trend features and seasonal features. The feature enhancement mapper is used to perform feature enhancement mapping on the trend features and seasonal features respectively, and obtain the mapped trend features and the mapped seasonal features. The fully connected linear layer is used to perform linear prediction on the mapped trend features and the mapped seasonal features respectively, and obtain the trend component prediction results and the seasonal component prediction results.

[0094] The spatial prediction module is used to construct the graph structure of the PM2.5 sequence based on the normalized time series data using the mutual information method, and predict the spatial prediction results through the graph attention neural network.

[0095] The gated prediction module is used to dynamically fuse the temporal prediction results and the spatial prediction results through the gated network, and obtain the predicted value of the PM2.5 concentration of the target site after denormalization.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A PM2.5 prediction method based on spatiotemporal characteristics, characterized in that: include: Obtain air pollutant data and meteorological data of the target site and surrounding sites; The air pollutant data and meteorological data of the surrounding stations are processed using principal component analysis and geographic distance weighting, and then spliced ​​with the air pollutant data and meteorological data of the target station to obtain the data to be analyzed; Normalize the data to be analyzed to obtain normalized time series data; The normalized time series data is decomposed into trend features and seasonal features, and the trend component prediction results and seasonal component prediction results are obtained by feature enhancement mapping and linear prediction respectively, and the time prediction results are obtained by adding them together; Based on the normalized time series data, the mutual information method is used to construct the graph structure of the PM2.5 sequence, and the spatial prediction results are obtained through the graph attention neural network. The temporal prediction results and spatial prediction results are dynamically fused through the gating network, and after denormalization, the predicted value of PM2.5 concentration at the target site is obtained.

2. A PM2.5 prediction method based on spatiotemporal characteristics as claimed in claim 1, characterized in that: Atmospheric pollutant data include fine particulate matter, inhalable particulate matter, sulfur dioxide, nitrogen dioxide, carbon monoxide and ozone, and meteorological data include temperature, air pressure, dew point temperature, wind direction and wind speed.

3. A PM2.5 prediction method based on spatiotemporal characteristics as claimed in claim 1, characterized in that: In principal component analysis, select the principal component with the largest variance; The calculation process of geographic distance weighting is: Calculate the geographic distance between the target site and each surrounding site; Calculate weights based on geographic distance; The principal component analysis results of each surrounding station are weighted and summed according to their weights to obtain the pollutant and meteorological data of the surrounding stations after dimensionality reduction.

4. The PM2.5 prediction method based on spatiotemporal characteristics according to claim 1, characterized in that: The normalization process is: Calculate the mean and standard deviation of the data to be analyzed; Use the mean and standard deviation to normalize the data to be analyzed to obtain normalized data; Add noise to the normalized data; The normalized data with added noise is scaled and translated to obtain the normalized time series data.

5. A PM2.5 prediction method based on spatiotemporal characteristics as claimed in claim 4, characterized in that: Dynamically adjust the standard deviation of noise. The calculation formula for the standard deviation of noise is: in, represents the noise standard deviation, represents the value of the loss function, It means to find the maximum value, Indicates finding the minimum value.

6. The PM2.5 prediction method based on spatiotemporal characteristics according to claim 1, characterized in that: The process of decomposing normalized time series data into trend features and seasonal features is: The normalized time series data is smoothed by multiple moving average kernels of different scales, and the moving average of each kernel is calculated separately, and the mean of multiple moving averages is used as the trend feature; The normalized time series data and trend characteristics are used to obtain seasonal characteristics by difference; The kernel sizes are 17, 25, 33, 49, 65, and 73, respectively.

7. The PM2.5 prediction method based on spatiotemporal characteristics according to claim 1, characterized in that: The calculation formula of feature enhancement map is: in, represents the output tensor, represents the input tensor, represents the weight matrix, represents the bias matrix, represents the batch size, represents the time step, represents the input channel dimension, Indicates the number of output features mapped to each input channel, and c represents the cth dimension.

8. The PM2.5 prediction method based on spatiotemporal characteristics according to claim 1, characterized in that: The prediction process of spatial prediction results is: Based on the normalized time series data, the graph structure of PM2.5 series is constructed using the mutual information method; Use symmetric normalization to process the adjacency matrix of the graph structure to obtain a symmetric normalized adjacency matrix; The node feature matrix of the graph structure and the symmetrically normalized adjacency matrix are input into the graph attention neural network to predict the spatial prediction results.

9. A PM2.5 prediction method based on spatiotemporal characteristics as claimed in claim 8, characterized in that: During the prediction process of the graph attention neural network, the node feature matrix is ​​processed using sine and cosine position encoding.

10. A PM2.5 prediction system based on spatiotemporal characteristics, characterized in that: Used to execute a PM2.5 prediction method based on spatiotemporal features as described in any one of claims 1 to 9, comprising: a data acquisition module, a data processing module, a feature normalizer, a time prediction module, a space prediction module and a gated prediction module, A data acquisition module is used to obtain atmospheric pollutant data and meteorological data of the target site and surrounding sites; A data processing module is used to process the air pollutant data and meteorological data of the surrounding stations using principal component analysis and geographic distance weighting, and then splice them with the air pollutant data and meteorological data of the target station to obtain the data to be analyzed; Feature normalizer, used to normalize the data to be analyzed to obtain normalized time series data; The time prediction module is used to decompose the normalized time series data into trend features and seasonal features, and obtain the trend component prediction results and the seasonal component prediction results through feature enhancement mapping and linear prediction respectively, and add them to obtain the time prediction result; The spatial prediction module is used to construct the graph structure of the PM2.5 sequence based on the normalized time series data using the mutual information method, and predict the spatial prediction results through the graph attention neural network; The gated prediction module is used to dynamically fuse the temporal prediction results and the spatial prediction results through the gated network, and obtain the predicted value of the PM2.5 concentration of the target site after denormalization.

Citation Information

Patent Citations

  • PM2.5 concentration prediction method based on space-time diagram ordinary differential equation network

    CN114694767A

  • Long time sequence pm2.5 prediction method and system based on space-time attention

    CN114662791A

  • Regional air quality prediction model based on multi-scale dynamic synchronous graph mechanism

    CN117371571A

  • Deep learning quantile water quality prediction method based on trend decomposition

    CN117689081A

  • PM2.5 spatio-temporal change prediction system and method based on neural network

    CN118861543A

Cited By

  • VOC release prediction method and system for low-odor multi-layer composite adhesive tape

    CN122117137A