Multi-site multi-pollutant prediction method and device based on GATv2 and time feature extraction

Through the multi-site multi-pollutant prediction method based on GATv2 and time feature extraction, the problem of periodic fluctuations in the prior art cannot be extracted and the insufficient modeling of cross-site feature is solved, and higher prediction accuracy and stability are achieved.

CN120218320AActive Publication Date: 2025-06-27BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY

Patent Information

Application Number
CN202510276777.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The prior art has problems such as inability to extract periodic fluctuations, insufficient cross-site feature modeling, and model complexity in the prediction of atmospheric pollutants with multiple sites and multiple features.

Method used

The multi-site multi-pollutant prediction method based on GATv2 and temporal feature extraction is adopted. By obtaining the historical data of multi-site pollutant monitoring, an adjacency matrix between multiple sites is generated, spatial and temporal features are extracted, and they are spliced ​​into space-time features, and finally input into full connection layer for prediction.

Benefits of technology

It significantly improves the ability to capture the changes in the concentration of atmospheric pollutants, improves the prediction accuracy and stability, and enhances the scalability and real-timeness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218320A_ABST
    Figure CN120218320A_ABST
Patent Text Reader

Abstract

The invention relates to the related technical field of pollution early warning, in particular to a multi-site multi-pollutant prediction method and device based on GATv2 and time feature extraction. The method comprises the following steps: acquiring historical data of multi-site pollutant monitoring; generating an adjacent matrix among multiple sites based on the historical data; the adjacent matrix is used for representing the relationship and dependency between different sites; extracting a first feature based on the adjacent matrix and the historical data through a GATv2 model; based on a time feature extraction module, capturing time features in the historical data to obtain second features; splicing the first feature and the second feature to obtain a spatial-temporal feature; and inputting the spatial-temporal characteristics into a preset full connection layer to obtain a prediction result. Through the arrangement, more accurate prediction can be carried out on multiple pollutants at multiple sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of pollution warning, and specifically relates to a multi-site and multi-pollutant prediction method and device based on GATv2 and time feature extraction. Background Art

[0002] In recent years, with the rapid development of industrialization and urbanization, the problem of air pollution has increasingly become the focus of global attention. Atmospheric pollutants, such as fine particulate matter (PM2.5), nitrogen dioxide (NO2), ozone (O3), etc., have caused serious impacts on human health and the ecological environment. Therefore, accurately predicting the concentration changes of atmospheric pollutants is of great significance for environmental monitoring, pollution warning, and policy making. Summary of the Invention

[0003] In view of this, embodiments of this application are committed to providing a multi-site and multi-pollutant prediction method and device based on GATv2 and time feature extraction for pollution warning.

[0004] This application provides a multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction, including:

[0005] Obtain historical data on multi-site pollutant monitoring;

[0006] Generate an adjacency matrix between multi-sites based on the historical data; the adjacency matrix is used to characterize the relationships and dependencies between different sites;

[0007] Extract first features through a GATv2 model based on the adjacency matrix and the historical data;

[0008] Capture time features in the historical data based on a time feature extraction module to obtain second features;

[0009] Concatenate the first features and the second features to obtain spatio-temporal features;

[0010] Input the spatio-temporal features into a preset fully connected layer to obtain a prediction result.

[0011] In some embodiments, generating an adjacency matrix between multi-sites based on the historical data includes:

[0012] Based on the historical data, construct a first sub-adjacency matrix for geographical location features;

[0013] Based on the historical data, construct a second sub-adjacency matrix for meteorological condition similarity;

[0014] Based on the historical data, construct a third sub-adjacency matrix for pollution source distribution characteristics;

[0015] The first sub - adjacency matrix and the first sub - adjacency matrix are weighted and added to obtain an adjacency matrix.

[0016] In some embodiments, the GATv2 model is used to dynamically learn the weights of adjacency relationships through a multi - head attention mechanism and output the site features after multi - head concatenation or averaging processing.

[0017] In some embodiments, the time feature extraction module uses the Inception structure of fast Fourier transform and multi - kernel dilated convolution to capture time features.

[0018] In some embodiments, the time features include: lag features, rolling statistical features, and periodic features.

[0019] In some embodiments, the prediction results include the prediction results of the concentrations of various pollutants at multiple sites.

[0020] This application also provides a multi - site and multi - pollutant prediction device based on GATv2 and time feature extraction, including:

[0021] An acquisition module for acquiring historical data of multi - site pollutant monitoring;

[0022] A generation module for generating an adjacency matrix between multiple sites based on the historical data; the adjacency matrix is used to characterize the relationships and dependencies between different sites;

[0023] An extraction module for extracting first features based on the adjacency matrix and the historical data through the GATv2 model; capturing time features in the historical data through the time feature extraction module to obtain second features;

[0024] A splicing module for splicing the first feature and the second feature to obtain spatio - temporal features;

[0025] A prediction module for inputting the spatio - temporal features into a preset fully - connected layer to obtain prediction results.

[0026] This application also provides an electronic device, including:

[0027] A processor and a memory for storing the executable program of the processor;

[0028] The processor is used to implement the multi - site and multi - pollutant prediction method based on GATv2 and time feature extraction as described above by running the program in the memory.

[0029] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the processor is caused to execute the multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction as described above.

[0030] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction as described above.

[0031] A multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction provided by the present application first obtains historical data of multi-site pollutant monitoring; generates an adjacency matrix between multi-sites based on the historical data; the adjacency matrix is used to characterize the relationships and dependencies between different sites; through the GATv2 model, first features are extracted based on the adjacency matrix and the historical data; based on a time feature extraction module, time features in the historical data are captured to obtain second features; the first features and the second features are concatenated to obtain spatio-temporal features; the spatio-temporal features are input into a preset fully-connected layer to obtain a prediction result. Through the adjacency matrix, the spatial correlation between sites can be clearly represented, including factors such as geographical location, meteorological condition similarity, and pollution source distribution. This explicit modeling of the relationship enables the model to better understand the mutual influence between sites, thereby improving the accuracy of prediction. The time feature extraction module can extract features of multiple time scales from historical data, including short-term fluctuations, long-term trends, and periodic changes (such as daily cycles, seasonal cycles). This multi-scale feature capture ability enables the model to comprehensively understand the time dynamics of pollutant concentrations, making up for the deficiency of traditional single time series models such as LSTM in extracting periodic features. Through the comprehensive modeling of spatio-temporal features, this solution can more accurately capture the variation law of atmospheric pollutant concentrations and significantly improve the prediction accuracy. Description of the Drawings

[0032] By describing the embodiments of the present application in more detail with reference to the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0033] Figure 1 It is a schematic flowchart of a multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction provided by an embodiment of the present application.

[0034] Figure 2It is a partial process schematic diagram of the method provided by an embodiment of the present application.

[0035] Figure 3 It is a structural schematic diagram of a multi-site and multi-pollutant prediction device based on GATv2 and time feature extraction provided by an embodiment of the present application.

[0036] Figure 4 It is a structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Application Overview

[0039] In recent years, with the increasing attention to air pollution problems, the limitations of traditional atmospheric pollutant prediction methods have gradually emerged. For example, prediction methods based on a single model or a single site are difficult to capture complex spatio-temporal dependencies and cross-site information transmission. In addition, the change of pollutant concentration is affected by multiple factors, such as meteorological conditions, geographical location, pollution source distribution, and pollution impact coefficient, etc. These factors lead to significant differences in pollution characteristics between different regions. Therefore, a solution that can effectively handle spatio-temporal dependencies and enhance cross-site information transmission is needed for multi-site and multi-feature atmospheric pollutant prediction.

[0040] Traditional spatio-temporal prediction models usually model spatio-temporal characteristics simultaneously, resulting in complex models that are difficult to optimize, especially when there are significant differences in spatial dependence relationships and time series characteristics. There is a CNN-LSTM model for predicting the daily average concentration of PM2.5 in Beijing on the next day. This study combines the advantages of CNN and LSTM models. CNN is used to extract the inherent features required for PM2.5 concentration prediction from the input data, especially to extract spatial-specific features from the input features at different monitoring stations. LSTM is used to capture the long-term dependencies of time series data and can handle the long-term change trends in air quality and meteorological data. Regarding the CNN part, it includes 6 CNNs. Each CNN extracts features related to its corresponding monitoring station and is arranged in the order of the frequency of feature occurrence, and is processed from the target station to other stations. The input dimension of each CNN varies according to the performance of the features at different stations. This way enables CNN to extract unique features for each station. Regarding the LSTM part, the feature time series output by CNN is input into LSTM, and LSTM is used to integrate these features, reflecting the long-term historical process of time series data and providing long-term dependence information for the final PM2.5 concentration prediction. Generally speaking, this study effectively improves the performance of multi-station prediction of air pollutants based on the spatio-temporal dependence model, especially when considering air quality and meteorological data. Its method has certain potential for expansion in future applications in the prediction of other air pollutants and complex environmental modeling fields.

[0041] In the existing solutions, the methods studied mainly rely on the information transfer between the target station and its neighboring stations, and have weak feature extraction capabilities for monitoring stations in far regions or independent stations.

[0042] In the existing solutions, multiple sub-models are designed in the CNN part, and each sub-model focuses on the feature extraction of a specific station. This structure leads to too high dimensions of the input features, increasing the complexity of model training and optimization.

[0043] In the existing solutions, in time series modeling, it mainly relies on LSTM to capture long-term dependencies, but has weak feature extraction capabilities for periodic and seasonal fluctuation characteristics. For example, the concentration of PM2.5 is usually significantly affected by periodic factors such as seasonal changes and weekly cycles, and this model cannot effectively extract these periodic fluctuation characteristics, resulting in deficiencies in prediction accuracy.

[0044] The existing solutions fail to fully consider spatial dependence features (such as adjacency matrix, distance, geographical features, influence coefficients, etc.), resulting in the inability to effectively model the spatio-temporal dependence relationships of multi-features across stations, affecting the multi-station prediction performance and the expansion ability of the model.

[0045] In summary, existing solutions have problems in predicting atmospheric pollutants with multiple sites and multiple features, such as being unable to extract periodic fluctuations, insufficient cross-site feature modeling, and high model complexity.

[0046] To solve the above problems, in view of the shortcomings of traditional models in predicting atmospheric pollutants with multiple sites and multiple features, such as being unable to extract periodic fluctuations, lacking an adjacency matrix to model the spatial dependence between sites, not fully utilizing distance features and geographical features, and insufficient ability to handle feature differences between heterogeneous sites, the present invention proposes an atmospheric pollutant prediction model GATv2Net with spatio-temporal feature fusion. The present invention can:

[0047] 1. Fully model the spatial dependence of multiple sites: By means of an improved graph attention network (GATv2), a dynamic adjacency matrix is constructed using distance features, geographical features, and pollution influence coefficients between sites, effectively extracting complex spatial correlation relationships between multiple sites and overcoming the problem of insufficient dependence on the adjacency matrix in traditional methods.

[0048] 2. Capture periodic fluctuation features: With the multi-scale decomposition ability of the TimesNet model, periodic and trend features are extracted from historical time series, improving the model's prediction ability for long-term and short-term pollutant fluctuations and making up for the problem of insufficient extraction of periodic features by traditional LSTM models.

[0049] 3. Multi-feature fusion and heterogeneous data modeling: By dynamically modeling the feature differences between heterogeneous sites through GATv2 and combining with TimesNet to efficiently capture the non-linear relationships between multiple pollutant features (such as PM2.5, NO2, O3, etc.), collaborative prediction of multiple sites and multiple features is achieved.

[0050] 4. Improve prediction accuracy and stability: Considering spatio-temporal characteristics at the same time, using GATv2 to extract spatial characteristics and TimesNet to capture temporal characteristics, efficient combination of spatio-temporal features is achieved, thereby significantly improving the prediction accuracy and stability of atmospheric pollutant concentrations, especially in complex meteorological conditions and cross-regional prediction tasks.

[0051] 5. Enhance the scalability and real-time performance of the model: The present model proposes a new framework for predicting atmospheric pollutants with multiple sites and multiple features, which has good scalability and can flexibly adapt to new sites or newly added pollutant features. The computational cost is reduced through an efficient network structure design, supporting real-time prediction applications for multiple sites.

[0052] After introducing the basic principle of this application, various non-limiting embodiments of this application will be specifically introduced below with reference to the accompanying drawings.

[0053] Exemplary Method

[0054] Figure 1 This is a schematic flowchart of a multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction provided by an embodiment of the present application. As Figure 1 shown, the method includes the following content.

[0055] Step S110, obtaining historical data of multi-site pollutant monitoring;

[0056] Specifically, collect historical atmospheric pollutant concentration data from multiple monitoring sites, including but not limited to concentration values of pollutants such as PM2.5, PM10, NO2, SO2, O3, CO, etc., and at the same time collect relevant meteorological data (such as temperature, humidity, wind speed, wind direction, etc.).

[0057] In practical applications, standardized monitoring data can be obtained from national or local environmental monitoring departments, or data can be collected in real time through sensor networks. The data is usually stored in the form of time series, and each time point corresponds to the pollutant concentration and meteorological parameters of one or more sites. Clean the original data, including handling missing values, outliers, and aligning timestamps. This can provide the basic data for model training and provide input for subsequent feature extraction and prediction. The spatio-temporal information contained in the historical data is the key to achieving accurate prediction.

[0058] Step S120, generating an adjacency matrix between multiple sites based on the historical data; the adjacency matrix is used to characterize the relationships and dependencies between different sites;

[0059] Specifically, construct an adjacency matrix based on factors such as the geographical locations of sites, meteorological conditions, and pollutant source distributions in the historical data to characterize the relationships and dependencies between different sites.

[0060] Step S130, extracting first features based on the adjacency matrix and the historical data through the GATv2 model;

[0061] Specifically, the historical data of each site is used as node features, and the adjacency matrix is used as the graph structure for input. GATv2 dynamically calculates the attention weights between nodes through the multi-head attention mechanism, and can flexibly capture the complex relationships between sites. The output of GATv2 is a feature vector after spatial relationship modeling, and the output features of each node fuse the information of its neighbor nodes. The attention weights of GATv2 can be dynamically adjusted according to the input data, overcoming the limitations of static attention in traditional graph neural networks. In this way, extract the spatial dependence features between sites and enhance the model's ability to model cross-site information transmission. The generated spatial features provide the basis for subsequent time feature fusion.

[0062] Step S140, capturing the time features in the historical data based on the time feature extraction module to obtain second features;

[0063] Specifically, a time feature extraction module is used to extract features of the time series from historical data, including lag features, rolling statistical features, and periodic features.

[0064] Step S150: Concatenate the first feature and the second feature to obtain spatio-temporal features.

[0065] Specifically, the spatial features (the first feature) extracted by GATv2 and the time features (the second feature) obtained by the time feature extraction module are concatenated to generate spatio-temporal fusion features. The spatial features and the time features are concatenated by dimension to generate a comprehensive feature vector. Through the concatenation operation, the effective integration of spatial and time information is ensured, providing a richer feature representation for subsequent prediction tasks. The deep fusion of spatial features and time features enables the model to simultaneously consider the spatial dependence between sites and the dynamic changes of the time series. The modeling ability of the model for complex spatio-temporal data is improved, and the accuracy and stability of the prediction are enhanced.

[0066] Step S160: Input the spatio-temporal features into a preset fully connected layer to obtain a prediction result.

[0067] The fully connected layer can further process and extract features from the spatio-temporal features to generate the final prediction result. The prediction result can be the pollutant concentration value at a future time point or multiple time points. The model is trained and optimized through a loss function (such as mean squared error) to ensure the accuracy of the prediction result.

[0068] The above steps realize the efficient and accurate prediction of atmospheric pollutant concentration through the integration of multi-site data, the construction of the adjacency matrix, the extraction of spatial and time features, and the fusion of spatio-temporal features. Each step plays a key role in the overall process, jointly constituting a complete spatio-temporal feature modeling framework, significantly improving the prediction performance and generalization ability of the model.

[0069] In some embodiments, generating the adjacency matrix between multiple sites based on the historical data includes:

[0070] Based on the historical data, construct a first sub-adjacency matrix for the geographical location feature;

[0071] Based on the historical data, construct a second sub-adjacency matrix for the similarity of meteorological conditions;

[0072] Based on the historical data, construct a third sub-adjacency matrix for the distribution characteristics of pollution sources;

[0073] Perform weighted addition on the first sub-adjacency matrix, the first sub-adjacency matrix; the first sub-adjacency matrix to obtain the adjacency matrix.

[0074] Specifically, for the geographical location characteristics of multiple sites, calculate the Euclidean distance between each pair of sites, and construct the first sub - adjacency matrix based on the distance information. Extract the geographical coordinates (longitude, latitude, altitude) of each site from historical data. Distance calculation: Use the Euclidean distance formula to calculate the distance between any two sites. In some embodiments, a distance threshold θ can be set. For site pairs with a distance less than θ, a larger weight is assigned; for site pairs with a distance greater than θ, a smaller weight (such as 0) or directly set to 0 is assigned to reduce noise. The geographical location characteristics directly affect the intensity of the interaction between sites. Sites in close proximity usually have a stronger pollution transmission relationship. The first sub - adjacency matrix can effectively capture the spatial proximity between sites and provide a basis for subsequent spatial feature extraction.

[0075] For the similarity of meteorological conditions of multiple sites, calculate the similarity of each site in terms of meteorological characteristics, and construct the second sub - adjacency matrix based on the similarity information. Extract the meteorological characteristics of each site from historical data, such as temperature, humidity, wind speed, wind direction, etc. Similarity calculation: Use a correlation coefficient (such as the Pearson correlation coefficient) to calculate the similarity of any two sites in terms of meteorological characteristics. Assign weights to the adjacency matrix according to the similarity value. The higher the similarity, the greater the weight. The similarity of meteorological conditions reflects the interaction between sites under the meteorological background. Especially under the influence of wind direction and wind speed, the transmission path and intensity of pollutants will be significantly affected. The second sub - adjacency matrix can dynamically capture the meteorological correlation between sites and enhance the model's adaptability to complex meteorological conditions.

[0076] For the pollution source distribution characteristics of multiple sites, calculate the correlation of each site in terms of pollution source distribution, and construct the third sub - adjacency matrix based on the correlation information. Extract the pollution source distribution information around each site from historical data, such as industrial areas, traffic - intensive areas, construction sites, etc. Calculate the pollution source correlation between sites according to the type and distribution range of pollution sources. For example, if two sites are both located in traffic - intensive areas, a higher correlation weight is assigned. Assign weights to the adjacency matrix according to the correlation of pollution source distribution. The pollution source distribution characteristics directly affect the pollution transmission relationship between sites. Especially for sites in a similar pollution source environment, the changes in their pollutant concentrations often have similarities. The third sub - adjacency matrix can capture the pollution source correlation between sites and enhance the model's ability to model pollution source distribution.

[0077] Weightedly add the first sub - adjacency matrix, the second sub - adjacency matrix, and the third sub - adjacency matrix to generate the final adjacency matrix. According to the importance of geographical location characteristics, meteorological condition similarity, and pollution source distribution characteristics, assign different weights α, β, and γ to each sub - adjacency matrix respectively.

[0078] Adjacency matrix fusion:

[0079] A = α·A(1) + β·A(2) + γ·A(3)

[0080] Among them, A is the final adjacency matrix, and α + β + γ = 1.

[0081] A(1) is the first sub - adjacency matrix; A(2) is the second sub - adjacency matrix; A(3) is the third sub - adjacency matrix;

[0082] Normalize the final adjacency matrix to ensure that its weight values are within a reasonable range (such as between 0 and 1). By the way of weighted summation, the sub - adjacency matrices with different features are fused into a comprehensive adjacency matrix, which can comprehensively reflect the complex relationships between stations.

[0083] The distribution of weights can be adjusted according to the actual application scenario to enhance the model's emphasis on different features and improve the accuracy and generalization ability of prediction.

[0084] Through the above steps, three sub - adjacency matrices are constructed respectively based on geographical location features, meteorological condition similarity, and pollutant source distribution characteristics, and the final adjacency matrix is generated by weighted summation of these sub - adjacency matrices. This process can not only comprehensively capture the spatial relationships between stations, but also dynamically adjust the weights of the adjacency matrix to enhance the model's adaptability to complex spatio - temporal environments. The final adjacency matrix provides a solid foundation for subsequent spatial feature extraction (such as through the GATv2 model), significantly improving the performance of multi - station air pollutant prediction.

[0085] The GATv2 model is used to dynamically learn the weights of adjacency relationships through the multi - head attention mechanism and output the station features after multi - head concatenation or averaging processing.

[0086] GATv2 (Graph Attention Network v2) is an improved graph neural network model dedicated to modeling node relationships in graph - structured data. By introducing the multi - head attention mechanism, it can dynamically learn the relationship weights between nodes, thus more effectively capturing the complex structural information in the graph. In this technical solution, the GATv2 model is used to extract spatial features from multi - station air pollutant monitoring data, enhancing the model's ability to model the mutual influence between stations. The multi - head attention mechanism is one of the core technologies of GATv2, which allows the model to learn the relationships between nodes from multiple different perspectives, thereby improving the richness and flexibility of feature extraction. Through the above - mentioned multi - head attention mechanism and aggregation operation, the GATv2 model finally outputs the updated feature vectors of each station. These feature vectors integrate the features of each station itself and the feature information of its neighbor stations, and can more comprehensively reflect the spatial dependence relationships between stations.

[0087] In some embodiments, the time feature extraction module adopts the Inception structure of fast Fourier transform and multi-core dilated convolution to capture time features. The time feature extraction module is a key component in the multi-site multi-pollutant prediction method and is used to capture the dynamic characteristics of time series from historical data. This module realizes efficient time feature extraction through the fast Fourier transform (FFT) and the Inception structure of multi-core dilated convolution, which can significantly improve the model's understanding and prediction ability for time series data. Fast Fourier transform (FFT) FFT is used to transform time series data from the time domain to the frequency domain and extract its periodic features. Specifically, perform FFT transformation on the historical pollutant concentration sequence of each site and calculate its spectrum. Identify several dominant frequencies with the largest amplitudes in the spectrum, and these frequencies correspond to the main periodic features of the time series (such as daily cycle, seasonal cycle). Extract these dominant frequencies and their corresponding average amplitudes as the core periodic features of the time series.

[0088] The Inception structure of multi-core dilated convolution is used to capture the long-term and short-term dependencies of time series while maintaining a low computational cost. Specifically: Design multiple dilated convolution kernels, each with a different dilation rate, to expand the receptive field and capture features at different time scales. Input the two-dimensional time series data after FFT transformation into the multi-core dilated convolution layer, and perform parallel processing through convolution kernels with different dilation rates to extract local and global features. Use the Inception structure to fuse features at different scales to generate a comprehensive time feature representation.

[0089] Through FFT and multi-core dilated convolution, it is possible to capture the periodicity, short-term fluctuations, and long-term trends of time series simultaneously, significantly improving the model's ability to model complex time series. The design of the Inception structure enables the model to extract rich feature information while maintaining efficient computation. The time feature extraction module provides high-quality time features for spatio-temporal feature fusion, significantly improving the accuracy and stability of multi-site multi-pollutant prediction.

[0090] The time features include: lag features, rolling statistical features, and periodic features. Time features are a key component in time series modeling and are used to capture the dynamic change laws in time series. In multi-site multi-pollutant prediction, the extraction of time features can significantly improve the model's understanding and prediction ability for the changing trends of pollutant concentrations.

[0091] Lag Feature Definition: Shift the time series data forward or backward by a certain number of time steps to generate new feature columns. This is used to capture short-term dependencies in the time series. In some embodiments, an appropriate lag step (such as 1 hour, 6 hours, 24 hours) can be selected to shift the historical pollutant concentration data forward by the corresponding step. The generated lag features can be directly used as inputs to the model to help the model understand recent changes in pollutant concentrations.

[0092] Rolling Statistical Feature Definition: Calculate statistical values (such as mean, standard deviation, maximum, minimum, etc.) within a sliding window on the time series to capture the local trends and fluctuations of the time series. In some embodiments, an appropriate sliding window size (such as 3 hours, 6 hours, 24 hours) can be selected to calculate the statistical values within the window at each time point. The generated rolling statistical features can help the model understand the local change patterns of the time series.

[0093] Periodic Feature Definition: Extract periodic patterns (such as daily cycle, weekly cycle, seasonal cycle) in the time series through periodic transformation of time information. This is used to capture the periodic change patterns in the time series. In some embodiments, information such as year, month, day, hour, day of the week, etc. can be extracted from the timestamp. Sine and cosine functions are used to transform the time information into periodic features.

[0094] Thus, by extracting lag features, rolling statistical features, and periodic features, it is possible to comprehensively capture the short-term dependencies, local trends, and periodic changes of the time series, significantly improving the model's ability to model time series. These time features can help the model better adapt to prediction tasks at different time scales, performing well in both short-term (such as 1 hour) and long-term (such as 72 hours) predictions. The introduction of time features provides rich information for spatio-temporal feature fusion, significantly improving the accuracy and stability of multi-site multi-pollutant predictions.

[0095] The prediction results include: the prediction results of the pollutant concentrations at multiple sites. The prediction results are the final output of the multi-site multi-pollutant prediction method and are used to provide decision support for environmental monitoring and pollution warning. In this technical solution, the prediction results include the predicted values of the pollutant concentrations at multiple sites, which can comprehensively reflect the changing trends of pollutant concentrations at future time points.

[0096] The following combines Figure 2 , to further illustrate the solution provided by this application:

[0097] I. Conduct correlation analysis; Mutual Information (MI) is a statistical measure that quantifies the dependence between two random variables X and Y. In multi-site feature selection, MI can effectively evaluate the correlation between different sites. In multi-site data, due to differences in geographical location, meteorological conditions, or pollutant source distribution, different sites may exhibit complex correlations. Through MI, the degree of dependence between sites can be accurately quantified. If the MI between two sites is high, it indicates a strong correlation in pollutant concentration or meteorological changes between them.

[0098] II. Preprocessing of multi-site datasets

[0099] In the preprocessing of multi-site data, first, clean six major air pollutants (PM2.5, PM10, NO2, SO2, O3, CO) and meteorological data (temperature, humidity, wind speed, etc.), including handling missing values, outliers, and timestamp alignment. Then, use normalization or standardization methods to unify the feature dimensions to ensure the consistency of the input data. For time characteristics, construct lag features, rolling statistical features, and periodic features;

[0100] III. Generate the adjacency matrix between sites

[0101] When predicting air pollutants at multiple sites, constructing the adjacency matrix of the model is a crucial step. The adjacency matrix A is used to represent the relationships and dependencies between sites. When constructing the adjacency matrix, the following key factors need to be considered:

[0102] 1. Geographical location features. Calculate the Euclidean distance between sites through geographical coordinates. Sites that are closer usually have stronger mutual influence. Set a distance threshold threshold, and set pairs of sites beyond the threshold to zero to reduce noise. Geographical location feature A ij The formula is as follows:

[0103]

[0104] By introducing three-dimensional coordinates (latitude, longitude, altitude), the calculation of distance can more realistically represent the actual geographical distance, especially in areas with significant altitude differences. The Euclidean distance calculation formula is as follows.

[0105]

[0106] 2. Similarity of meteorological conditions. Use the correlation coefficient to quantify the similarity of meteorological characteristics between two sites. Pairs of sites with high similarity can be assigned larger weights. Since meteorological conditions vary with time, the adjacency matrix can be dynamically adjusted according to meteorological conditions at different time periods.

[0107] 3. Pollutant source distribution characteristics. Different stations may be within the influence range of similar pollutant sources, such as industrial areas and traffic-intensive areas. The pollutant source distribution characteristics are used to weight the adjacency matrix.

[0108] For the above three characteristics, each characteristic corresponds to a sub-adjacency matrix with a size of N×N, where N is the number of stations. The above three matrices are fused into an overall adjacency matrix A by weighted summation, where α, β, and γ are hyperparameters, and their weights reflect the importance of different characteristics. This can enable the GATv2 model to fully learn the spatial dependence relationship between stations and improve the accuracy of air pollutant prediction.

[0109] A = αA geo + βA meteo + γA source

[0110] A geo is the geographical location feature; A meteo is the meteorological condition similarity; A source is the pollutant source distribution characteristic.

[0111] IV. Model construction

[0112] 1. The data of each station not only depends on its own historical data but also is related to the air quality of surrounding stations. For the existing technology I (CNN-LSTM model) mentioned in 1.2, it ignores the mutual influence between stations and fails to effectively utilize the spatial characteristics described in the adjacency matrix. The original Graph Attention Network (GAT) is a deep learning model based on graph-structured data, which models the node relationships in the adjacency matrix through the self-attention mechanism. However, the GATConv layer of GAT has the problem of static attention, that is, the topological structure of the adjacency matrix is fixed during training and cannot be dynamically adjusted, thus limiting the modeling ability of node relationships. The GATv2 adopted in the present invention improves the original GAT, and each node can dynamically focus on any node in the graph.

[0113] By inputting the feature vector X of each station and the fused adjacency matrix, GATv2 dynamically learns the adjacency relationship weights through the multi-head attention mechanism, and the output features are processed by multi-head concatenation or averaging to generate updated station features.

[0114] 2. Based on the spatial features processed by GATv2, the sequential patterns of the data are extracted through temporal features. The temporal feature extraction module, based on the Inception structure of the fast Fourier transform (FFT) and multi-core dilated convolution, can efficiently capture the temporal features of multi-site time series. Using the fast Fourier transform, the one-dimensional time series data input is first transformed into the frequency domain to extract its main periodic features. The specific steps are to calculate the spectrum of the time series, identify several dominant frequencies with the largest amplitudes in the frequency components, and these frequencies correspond to the main periodic features of the time series. Subsequently, the main period values and the average amplitudes of their corresponding frequency components are returned as the core features of the input time series. On the extended two-dimensional time series data, multi-dilated convolution is used for processing. Dilated convolution expands the receptive field by inserting holes in the convolutional kernel, enabling the model to capture a larger range of context information while extracting local features. Multi-dilated convolution effectively improves the ability to capture the long-term and short-term dependencies of time series through different combinations of dilation rates, while maintaining a low computational cost. This processing method enhances the model's ability to understand complex time series data.

[0115] 3. In the spatio-temporal fusion stage of the model, GATv2 and the temporal feature module are integrated to improve the accurate prediction ability of multi-site pollutant concentrations. The spatial features extracted by GATv2 and the temporal features extracted by the temporal module are combined together, and this process is completed through feature concatenation to ensure the effective integration of spatial and temporal information. Subsequently, the fused features are input into the fully connected layer to output the prediction results of multi-site pollutant concentrations at future time points. This spatio-temporal fusion-based model design can simultaneously capture the spatial dependencies between monitoring sites and the dynamic change characteristics of time series, achieving accurate prediction of multi-site pollutant concentrations.

[0116] In summary, the present invention proposes a multi-site pollutant concentration prediction method combining GATv2 and a time feature module, which can effectively capture the spatial dependence and time series characteristics among sites. The GATv2 model overcomes the limitation of the fixed attention mechanism of the original GAT model by dynamically adjusting the adjacency matrix, improving the spatial feature extraction ability. The time feature module extracts the periodic changes of multi-site time series based on FFT and multi-core dilated convolution, further enhancing the model's grasp of long-term and short-term trends. This method can reduce the dependence on single-site historical data, improve the generalization ability and prediction accuracy of the model, and is suitable for multi-site pollutant prediction tasks in complex, multi-dimensional and heterogeneous data environments. By combining GATv2 and the time feature module, this technology can effectively capture complex spatio-temporal dependence and periodic changes, thus significantly improving the accuracy and stability of multi-site pollutant prediction. The GATv2 model accurately models the spatial relationship between sites by dynamically learning the adjacency matrix, while the time feature module uses FFT and multi-core dilated convolution to extract the long-term periodicity of pollutant concentration, improving the ability to capture long-term trends and seasonal changes. In addition, the modular design of the model enables it to be extended to multi-pollutant and multi-site tasks, with stronger generalization ability and interpretability, providing a flexible and efficient framework for spatio-temporal characteristic modeling.

[0117] Exemplary Device

[0118] The device embodiment of the present application can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0119] Figure 3 The block diagram of a multi-site multi-pollutant prediction device based on GATv2 and time feature extraction provided by an embodiment of the present application is shown. As Figure 4 shown, the device includes:

[0120] An acquisition module 31, configured to acquire historical data of multi-site pollutant monitoring;

[0121] A generation module 32, configured to generate an adjacency matrix between multi-sites based on the historical data; the adjacency matrix is used to characterize the relationship and dependence between different sites;

[0122] An extraction module 33, configured to extract first features based on the adjacency matrix and the historical data through a GATv2 model; capture time features in the historical data based on a time feature extraction module to obtain second features;

[0123] A splicing module 34, configured to splice the first feature and the second feature to obtain spatio-temporal features;

[0124] The prediction module 35 is configured to input the spatio-temporal features into a preset fully connected layer to obtain a prediction result.

[0125] Exemplary Electronic Device

[0126] Next, reference Figure 4 will be made to describe the electronic device according to an embodiment of the present application. Figure 4 The block diagram of the electronic device according to an embodiment of the present application is illustrated.

[0127] As Figure 4 shown, the electronic device 400 includes one or more processors 410 and a memory 420.

[0128] The processor 410 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.

[0129] The memory 420 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 410 may run the program instructions to implement the multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction of various embodiments of the present application described above and / or other desired functions. Various contents such as category correspondence relationships may also be stored in the computer-readable storage medium.

[0130] In one example, the electronic device 400 may further include: an input device 430 and an output device 440, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0131] In addition, the input device 430 may further include, for example, a keyboard, a mouse, an interface, etc. The output device 440 may output various information to the outside, including analysis results, etc. The output device 440 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0132] Of course, for simplicity, Figure 4 only some of the components related to the present application in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.

[0133] Exemplary Computer Program Product and Computer Readable Storage Medium

[0134] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction according to various embodiments of the present application described in the "Exemplary Method" section above of this specification.

[0135] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0136] Furthermore, an embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when run by a processor, cause the processor to execute the steps in the multi-site and multi-pollutant prediction method based on GATv2 and time feature extraction according to various embodiments of the present application described in the "Exemplary Method" section above of this specification.

[0137] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0138] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. A multi-site multi-pollutant prediction method based on GATv2 and time feature extraction, characterized in that: include: Obtain historical data on pollutant monitoring at multiple sites; generating an adjacency matrix between multiple sites based on the historical data; The adjacency matrix is ​​used to characterize the relationship and dependency between different sites; Extracting a first feature based on the adjacency matrix and the historical data through a GATv2 model; Based on the time feature extraction module, the time feature in the historical data is captured to obtain a second feature; Concatenating the first feature and the second feature to obtain a spatiotemporal feature; The spatiotemporal features are input into a preset fully connected layer to obtain a prediction result.

2. The multi-site multi-pollutant prediction method based on GATv2 and time feature extraction according to claim 1 is characterized in that: Generating an adjacency matrix between multiple sites based on the historical data includes: Based on the historical data and targeting the geographical location features, constructing a first sub-adjacency matrix; Based on the historical data, a second sub-adjacency matrix is ​​constructed for similarity of meteorological conditions; Based on the historical data and the distribution characteristics of pollution sources, a third sub-adjacency matrix is ​​constructed; The first sub-adjacency matrix, the first sub-adjacency matrix, and the first sub-adjacency matrix are weighted added to obtain an adjacency matrix.

3. The multi-site multi-pollutant prediction method based on GATv2 and time feature extraction according to claim 1 is characterized in that: The GATv2 model is used to dynamically learn the adjacency weights through a multi-head attention mechanism, and output site features after multi-head splicing or averaging.

4. The multi-site multi-pollutant prediction method based on GATv2 and time feature extraction according to claim 1 is characterized in that: The temporal feature extraction module uses the Inception structure of fast Fourier transform and multi-core dilated convolution to capture temporal features.

5. The multi-site multi-pollutant prediction method based on GATv2 and time feature extraction according to claim 1 is characterized in that: The time characteristics include: hysteresis characteristics, rolling statistical characteristics and periodic characteristics.

6. The multi-site multi-pollutant prediction method based on GATv2 and time feature extraction according to claim 1 is characterized in that: The prediction results include: prediction results of the concentrations of various pollutants at multiple sites.

7. A multi-site multi-pollutant prediction device based on GATv2 and time feature extraction, characterized in that: include: An acquisition module is used to obtain historical data of pollutant monitoring at multiple sites; A generation module, used for generating an adjacency matrix between multiple sites based on the historical data; The adjacency matrix is ​​used to characterize the relationship and dependency between different sites; An extraction module, configured to extract a first feature based on the adjacency matrix and the historical data through a GATv2 model; Based on the time feature extraction module, the time feature in the historical data is captured to obtain a second feature; A splicing module, used for splicing the first feature and the second feature to obtain a spatiotemporal feature; The prediction module is used to input the spatiotemporal features into a preset fully connected layer to obtain a prediction result.

8. An electronic device, characterized in that: include: A processor, and a memory for storing a program executable by the processor; The processor is used to implement the multi-site multi-pollutant prediction method based on GATv2 and time feature extraction as described in any one of claims 1 to 6 by running the program in the memory.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to execute the multi-site multi-pollutant prediction method based on GATv2 and time feature extraction as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, which, when executed by a processor, implements the multi-site multi-pollutant prediction method based on GATv2 and time feature extraction as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • PM2.5 concentration space-time change prediction method and system based on space-time diagram neural network

    CN113919231A

  • Air quality prediction method and system

    CN116611561A

  • Air quality prediction method for mining space-time attention mechanism based on multiple relations

    CN119204352A

  • Air pollutant concentration prediction method and related equipment thereof

    CN119474762A

Cited By

  • Water eutrophication monitoring and pollution early warning system

    CN120971678A