Model establishment method and application of fully residual deep network based on graph convolution

Through the full-residual deep network model based on graph convolution, the problem of insufficient fitting accuracy and generalization ability in the spatiotemporal modeling of air pollutant concentration is solved, and efficient irregular data processing and accurate prediction are achieved.

CN112712169BActive Publication Date: 2025-07-08INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202110021814.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-08
Publication Date
2025-07-08
Estimated Expiration
2041-01-08

AI Technical Summary

Technical Problem

The existing spatiotemporal modeling methods for air pollutant concentrations have problems such as limited fitting accuracy and insufficient generalization ability when dealing with sparse monitoring sites and complex spatiotemporal variations. The traditional methods consume too much computing resources in irregular data applications.

Method used

A full-residual deep network model based on graph convolution is adopted, and air pollution data is collected and sorted out, the ratio of spatial and temporal dimensional correlation coefficients is calculated, and a multi-level local graph convolution network is established by using k-NN nearest neighbor algorithm and hierarchical random sampling, and a multi-level local graph convolution network is trained and predicted by combining full-residual deep network.

Benefits of technology

It improves the accuracy and generalization of space-time estimation of air pollutant concentrations, avoids the requirements for continuous data input, reduces computing resource consumption, and improves the generalization ability and prediction efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112712169B_ABST
    Figure CN112712169B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for establishing a model of a fully residual deep network based on graph convolution and its application. The method is as follows: collect and preprocess air pollution data; find the nearest neighbor pairs in the spatial and temporal dimensions respectively, and calculate the ratio of the correlation coefficients of the two dimensions; perform regularization processing to determine the spatio-temporal dimensions; use the k-NN nearest neighbor algorithm to find all the k nearest neighbors and their distances; use the location-based stratified random sampling method to divide the training and test samples; establish and train a fully residual deep network model based on graph convolution; optimize the model parameters to obtain the optimal model. The application is: call the optimal model to predict the air pollution situation. The present invention reduces overfitting, greatly improves the accuracy and generalization ability (actual prediction ability) of the spatio-temporal estimation of air pollutant concentrations, connects the spatio-temporal variation output with the fully residual deep network, and improves the ability of the model to mine spatio-temporal variations and the generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for establishing a model and an application thereof, and particularly relates to a method for establishing a model based on a graph convolutional fully residual deep network and an application thereof. Background Art

[0002] The spatiotemporal modeling methods for air pollutant concentration include methods such as kriging, Bayesian maximum entropy (BME), and generalized additive model (GAM). However, in these methods, kriging and Bayesian maximum entropy require the spatiotemporal variation to satisfy the demand of stationary spatiotemporal random field. In the actual air pollution concentration distribution, it is greatly affected by various factors, and the spatiotemporal variation difference is extremely large, making it difficult to meet the premise of spatiotemporal stability; due to the influence of randomness, the fitting of spatiotemporal variation has a certain uncertainty and the fitting accuracy is limited; while the generalized additive model lacks the consideration of spatiotemporal variation and mainly conducts spatiotemporal modeling through coordinates and covariates of spatiotemporal changes. Compared with emerging machine learning, such as neural network models, the generalization ability of these modeling methods based on traditional statistical models is limited, which is reflected in the low accuracy in the independence test based on location and a large deviation in the time application scenario.

[0003] The time series models in air pollutant concentration prediction include a series of traditional methods such as moving average and autoregressive moving average. Recently, advanced long-short term memory network (LSTM) and methods based on convolutional neural network (CNN) have also been used for air pollution monitoring. For example, the existing technology proposes that continuous historical data input is required to obtain the output of the next stage of prediction. However, this method makes it difficult for the time series model to be applied to the application scenario of air pollution monitoring data with irregular spatiotemporal sample distribution.

[0004] In emerging deep learning technologies, in addition to time series methods such as LSTM and CNN, feedforward neural network models are also gradually applied to process irregular data. For example, the existing technology has proposed a fully residual link deep network, and its strong generalization ability also improves the prediction accuracy of air pollution compared with previous methods. The emerging graph neural network (GNN) technology can use technologies such as graph convolutional network (GCN) to process sample points with irregular distributions, providing new ideas for new spatio-temporal modeling methods. However, current graph networks are mainly used for classification and prediction of images and traffic data, such as for road network status monitoring, but lack relevant methods for spatio-temporal modeling of air pollution concentration with sparse monitoring stations and obvious spatio-temporal variations. In addition, the current popular full-graph node modeling has too much input data, often resulting in an overly large matrix in applications and being difficult to use for actual spatio-temporal modeling. Summary of the Invention

[0005] In order to solve the deficiencies of the above technologies, in the field of environmental science, due to limited data of current air pollutant detection stations and limited spatio-temporal modeling methods for air pollutant concentration, including problems such as complex spatio-temporal variation fitting and insufficient generalization of training models, the present invention proposes a model establishment method and application based on a fully residual deep network with graph convolution.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: a model establishment method based on a fully residual deep network with graph convolution, including the following steps:

[0007] Step 1: Collect, organize and preprocess air pollution and its covariate data;

[0008] Step 2: In the spatial dimension and time dimension, find the nearest adjacent point pairs for the sample points of the data and calculate the ratio of the correlation coefficients in the two dimensions;

[0009] Step 3: Regularize all the data and determine the spatio-temporal dimensions;

[0010] Step 4: Use the k-NN nearest neighbor algorithm to find the k nearest neighbors and their spatio-temporal distances of all sample points;

[0011] Step 5: Use the method of location-based stratified random sampling to divide the training and test samples;

[0012] Step 6: Establish a multi-level nearest neighbor network topology, where each level is equivalent to a lag;

[0013] Step 7: Establish a fully residual deep network model based on graph convolution, pre-train the graph convolutional network, and fuse the fully residual network for training;

[0014] Step 8: Optimize the model parameters to obtain a model with optimal parameters.

[0015] Further, in Step 1, after collecting the original air pollution data, organize it, determine the spatio-temporal resolution, and extract and preprocess the corresponding covariate factors; the covariate factors include coordinates, elevation, meteorology, land use, MERRA2 data with coarse resolution, and time index.

[0016] Further, in Step 2, calculate the Pearson correlation coefficients for the spatial dimension and the temporal dimension respectively, calculate the ratio of the two, and obtain the ratio of the correlation coefficients of the two dimensions.

[0017] Further, in Step 3, use the standardized data regularization technique for processing, as shown in Equation 1:

[0018] Equation 1

[0019] In the formula, represents the original input of the j-th feature variable, represents the mean of the original input, represents the variance of the original input, is the regularized variable.

[0020] Further, in Step 4, when calculating the spatio-temporal distance, use the ratio of the correlation coefficients of the two dimensions calculated in Step 2 to adjust the spatio-temporal distance calculation process, so that the time and space dimensions are statistically consistent when calculating the distance.

[0021] Further, in Step 5, divide the area according to administrative regions or climate zones to obtain the locations of spatial monitoring points, randomly select and divide the locations into M equal parts. Among them, 60 - 90% * M parts of the samples corresponding to the monitoring point locations are used for model training, and the samples corresponding to the remaining parts of the monitoring point locations are used for the independence test of the locations.

[0022] Further, in Step 6, use a graph network that can handle irregular sample points. One lag represents the influence of the neighbors of the nearest spatio-temporal point, and the next lag is for the nearest neighbors of the connected target node; use the autocorrelation function graph to determine the lag, and explore the lag step K, that is, the total number of layers, according to the actual data.

[0023] The multi-level lag spatio-temporal modeling process is as follows: From the outermost layer to the innermost layer to the target node, the influence of adjacent points is transmitted layer by layer and finally aggregated to the target node; in the three spatio-temporal dimensions, including the spatial dimensions x and y and the temporal dimension t, and according to the principle of closer being more relevant in the first law of geography, the nearest neighbors are determined. At the same time, due to the irregular distribution of spatio-temporal sample points, their distance weight coefficients are aggregated for the nearest neighbors, as shown in Equation 2:

[0024] Equation 2,

[0025] where k represents the index of the level, u is the target node, represents that v is an arbitrary adjacent point of u, is the output of the k-th layer of the adjacent node v, is the output of the (k - 1)-th layer of the adjacent node v, is the aggregation result of the target node u obtained from the nearest neighbor points, equivalent to the state adjacent points of u at the k-th layer is the output of the aggregation function of aggregate (k) represents the aggregation function of the adjacent points at the k-th layer;

[0026] The aggregation function is defined as a weighted sum or a pool operation;

[0027] For the weighted aggregation function, the output of the k-th layer state of the target node u is shown in Equation 3:

[0028] Equation 3,

[0029] where, represents the distance between the adjacent node v and the target node u, and WMEAN represents the inverse distance weighted summation;

[0030] For the pool operation, the following aggregation function is used, and the output of the k-th layer state of the target node u is shown in Equation 4:

[0031] Equation 4,

[0032] where σ represents the sigma non-linear activation function, W represents the weight coefficient matrix of the pool layer, and b represents the bias vector;

[0033] After k-layer convolution operations, the state of the target node u is updated as:

[0034] Equation 5,

[0035] represents the concatenation of two output vectors, represents the output of the k-th layer of node u when updating the state, represents the output of its own k-1 layer;

[0036] After K layers of nearest neighbor aggregation operations, the output of the target node u is obtained:

[0037] Equation 6,

[0038] where Z u Finally, the estimated value of the target node is to be estimated.

[0039] Furthermore, in step seven, the established model consists of two parts. One part is a graph convolutional network, and the other part is a fully residual deep network; the graph convolutional network module is a spatio-temporal GCN module that considers multi-layer local convolution. The output of this module is concatenated with the output of the original node as the input of the fully residual deep network;

[0040] The training process of the model is as follows: First, perform semi-supervised pre-training on the graph spatio-temporal network, and use the MSE loss function shown in Equation 7:

[0041] Equation 7,

[0042] where represents the coefficient set of the weight matrix W and bias vector b of the network layer, y represents the ground truth value, represents the estimation result of the graph network for the target variable y with respect to the input x using the parameter of, is the loss function of the graph, is the regularization term of the parameters of the graph network G, and N represents the number of training samples;

[0043] After initializing the training, by concatenating the input unit with the output unit matrix of the graph network, the graph network of the network is fused with the fully residual deep network to form a unified network model for model training, and the following total mean square error MSE loss function is used:

[0044] Equation 8,

[0045] where represents the coefficient set of the weight matrix W and bias vector b of the network layer, y represents the ground truth value, represents the estimation result of the network for the target variable y with respect to the input x using the parameter of, is the regularization term of the parameters of the graph network G, and N represents the number of training samples;

[0046] The output result of model training is subjected to anti-regularization operation with the previously saved regularization parameters to obtain the output result of the original input variable scale, and the corresponding test indexes are calculated.

[0047] Further, in step eight, after establishing the model and performing preliminary training, systematic training comparison is carried out for the key parts in model training, including the number of local graph convolutional layers K adopted, activation function, nodes and depth of each residual layer, and the size of each batch of training samples, to obtain the model with the optimal parameters.

[0048] An application of a method for establishing a model of a fully residual deep network based on graph convolution. The model is called to predict the air pollution situation: after the model training and parameter tuning are completed, the model is saved. Then, when applying the model to a new data set, parameter regularization is required first, and then the same k-NN is used to find the multi-layer nearest neighbor nodes. Then, the model is called to make a prediction to obtain the prediction result of the new data set, and the prediction result is the prediction result of the air pollution situation.

[0049] An application of a method for establishing a model of a fully residual deep network based on graph convolution: The model is called to predict the air pollution situation: after the model training and parameter tuning are completed, the model is saved. Then, when applying the model to a new data set, parameter regularization is required first, and then the same k-NN is used to find the multi-layer nearest neighbor nodes. Then, the model is called to make a prediction to obtain the prediction result of the new data set, and the prediction result is the prediction result of the air pollution situation.

[0050] The present invention discloses a method for establishing a model of a fully residual deep network based on graph convolution. By introducing the spatio-temporal dimension point pair correlation coefficient ratio and regularization technology, efficient spatio-temporal nearest neighbor extraction is realized according to the first law of geography, thus realizing a multi-layer locally weighted graph convolutional network, which enables the model: 1) Compared with the conventional time series method, it avoids the requirement of continuous data input, and the model can process irregularly distributed spatio-temporal sample points; 2) Compared with the conventional spatio-temporal modeling such as geostatistics, it avoids the prerequisite condition that is extremely difficult to meet in actual problem modeling, which is the spatio-temporal steady change required for spatio-temporal variation fitting, making the method proposed in this patent have a wider practical range; 3) Compared with the full graph convolutional network, the farther the distance, the smaller the influence, and the local graph convolution of this patent is applicable to dealing with similar problems, avoiding the large amount of computing resources consumed by the full graph participating in the operation, and improving the efficiency of graph convolution operation and the training and prediction of the entire model in terms of calculation.

[0051] The model constructed by the method of the present invention fully combines the advantages of spatio-temporal modeling of local graph convolutional networks and the modeling of full residual deep networks, overcomes the deficiencies of each model, enables the model to capture the main nearest neighbor information, and also gives full play to the non-linear prediction ability of its own features. Therefore, a higher accuracy is obtained in the location-based independence test, fully reflecting its good generalization ability. Finally, the fusion model realizes "end-to-end" optimized training, overcomes the deficiencies of the traditional spatio-temporal modeling steps being complex and requiring manual intervention with limited accuracy, and fully improves the efficiency of training and prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of the overall process of the present invention.

[0053] Figure 2 It is a schematic diagram of the multi-level local lag spatio-temporal modeling of the present invention.

[0054] Figure 3 It is a schematic diagram of the structure of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0056] The present invention proposes a spatio-temporal modeling method based on a full residual deep network of local graph convolution, which greatly improves the accuracy and generalization ability (actual prediction ability) of spatio-temporal estimation of air pollutant concentrations. This method adopts a multi-level local graph convolution method to simulate the lag modeling of time series methods, and fuses spatio-temporal modeling through the ratio of statistical correlation weights, efficiently capturing the spatio-temporal correlation of air pollutants; this method connects the spatio-temporal variation output with a full residual deep network, and compared with traditional spatio-temporal modeling methods, greatly improves the model's ability to mine spatio-temporal variations and the generalization ability of the model.

[0057] Figure 1 The process of the model establishment method of the present invention is shown, including the following steps:

[0058] Step 1: Collect, organize and preprocess air pollution and its covariate data;

[0059] After collecting the original air pollution data, it is necessary to organize it, including the synthesis of original measurement data; subsequently, determine the spatio-temporal resolution, extract and preprocess the corresponding covariate factors; the covariate factors include coordinates, elevation, meteorology, land use, coarse-resolution MERRA2 data, and time index; the preprocessing includes data denoising processing, etc., to improve the quality of the data.

[0060] Step 2: In the spatial dimension and time dimension, find the nearest adjacent point pairs for the sample points of the data, and calculate the ratio of the correlation coefficients in the two dimensions;

[0061] According to the coordinates and time indices, using the nearest neighbor method, find the nearest neighboring point pairs in the spatial dimension and the time dimension respectively, calculate the Pearson correlation coefficients of the spatial dimension and the time dimension respectively, calculate the ratio of the two, and obtain the ratio of the correlation coefficients of the two dimensions.

[0062] Step 3: Perform regularization processing on all data and determine the spatio-temporal dimensions;

[0063] All data includes the above-mentioned covariates and the target variable. The target variable of the present invention refers to the air pollution concentration. Regularization processing is performed on both the covariates and the target variable to eliminate the differences between different variable units and scales and improve the stability of neural network training. This patent adopts a standardized data regularization technique, as shown in Equation 1:

[0064] Equation 1

[0065] In the formula, represents the original input of the j-th feature variable, represents the mean of the original input, represents the variance of the original input, is the regularized variable.

[0066] For the target variable (air pollution concentration), perform regularization using a similar method and record its mean and variance , which is convenient for the scale recovery of the prediction results.

[0067] Step 4: Adopt the k-NN nearest neighbor algorithm to find the k nearest neighbors and their spatio-temporal distances of all sample points; when calculating the spatio-temporal distance, use the ratio of the correlation coefficients of the two dimensions calculated in Step 2 to adjust the spatio-temporal distance calculation process, so that the time and space dimensions are statistically consistent when calculating the distance. Thus, a method is proposed to calculate the ratio of the correlation coefficients of the spatio-temporal dimensions based on the nearest neighbor point pairs as the adjustment factor for spatio-temporal dimension fusion, obtaining a quantitative comparison of time and space in a statistical sense, which can overcome the differences in measurement units and scales between time and space. Through the correlation coefficient ratio correction factor, and the regularization technique eliminates the measurement unit differences in time and space. The combination of the two effectively fuses the time and space dimensions.

[0068] Step 5: Adopt the method of location-based stratified random sampling to divide the training and test samples;

[0069] For the stratification factor of the stratified random sampling method, existing zoning methods can be adopted, such as zoning according to administrative divisions or climate zones to obtain the locations of spatial monitoring points, randomly selecting and dividing the locations into M equal parts. Among them, 60-90% of the samples corresponding to the monitoring point locations are used for model training, and the samples corresponding to the remaining number of parts of the monitoring point locations are used for the independence test of the locations. For example, randomly select and divide the locations into 4 equal parts, among which, the samples corresponding to the locations of 3 monitoring points are used for model training, and the samples corresponding to the locations of 1 monitoring point are used for the independence test. Adopting the independence test based on location is an important step in verifying this algorithm. Using this method will make the training sample data no longer contain the time series data of the test samples, which is the main step in verifying the true prediction ability (i.e., generalization) of the model.

[0070] Step Six: Establish a multi-level nearest neighbor network topology, where each level is equivalent to a lag;

[0071] This patent adopts a graph network that can handle irregular sample points. One lag represents the influence of the neighbors of a nearest spatio-temporal point, and the next lag is for the nearest neighbors of the connected target node. As the number of lags increases, the influence gradually weakens. The steps to determine the lag can be determined using an autocorrelation function graph (autocorrelation function, abbreviated as ACF). Here, 3 can be used as the initial value for air pollution data, and the appropriate lag step size K, that is, the total number of layers, can be explored based on the actual data.

[0072] Figure 2 Shows the multi-level local lag spatio-temporal modeling. From the outermost layer to the inner layer to the target node, the influence of neighboring points is transmitted layer by layer and finally aggregated to the target node. Different from the previous SAGE network based on local convolution of ordinary feature variables, this patent only considers the three spatio-temporal dimensions (space: x and y, time: t), determines the nearest neighbors according to the principle of "the closer, the more relevant" of the first theorem of geography, and at the same time considers the irregular distribution of spatio-temporal sample points itself, and aggregates its distance weight coefficients for the nearest neighbors:

[0073] Equation 2,

[0074] In the formula, k represents the index of the layer, u is the target node, represents that v is any neighboring point of u, is the output of the k-th layer of the neighboring node v, is the output of the (k - 1)-th layer of the neighboring node v, is the aggregation result obtained by the target node u from the nearest neighbor points, which is equivalent to the state of the nearest neighbor points of u at the k-th layer the output of the aggregation function of, aggregate(k) An aggregation function for neighboring points in the k-th layer;

[0075] Define the aggregation function as a weighted sum or a pool operation;

[0076] For the weighted aggregation function, the state output of the target node u at the k-th layer is shown in Equation 3:

[0077] Equation 3,

[0078] In the formula, represents the distance between the neighboring node v and the target node u, and WMEAN represents the distance-inverse weighted sum;

[0079] For the pool operation, use the following aggregation function, and the state output of the target node u at the k-th layer is shown in Equation 4:

[0080] Equation 4,

[0081] In the formula, σ represents the sigma non-linear activation function, W represents the weight coefficient matrix of the pool layer, and b represents the bias vector;

[0082] After k layers of convolution operations, the state of the target node u is updated to:

[0083] Equation 5,

[0084] represents the concatenation of two output vectors, represents the output of node u at the k-th layer, represents the output of its own (k - 1)-th layer; represents the neighboring nodes of node u at the k-th layer after the completion of the aggregation operation, and the two are concatenated (concatenate, i.e., ) as the input of the k-th layer of u.

[0085] After the K-layer nearest neighbor aggregation operation, the output of the target node u:

[0086] Equation 6,

[0087] In the formula, Z u Finally, estimate the estimated value of the target node.

[0088] Step 7. Model establishment and training: Establish a model, pre-train the graph convolutional network, and fuse the fully residual network for training; as Figure 3As shown, the structure of the model is presented. The model consists of two parts. One part is the graph convolutional network, and the other part is the fully residual depth network. The graph convolutional network module is the spatio-temporal GCN module that considers multi-layer local graph convolution mentioned above. The output of this module is concatenated with the output of the original nodes as the input of the fully residual depth network. The fully residual depth network has strong learning ability but lacks the ability of spatio-temporal modeling. In this patent, the two are combined to enhance the overall training and generalization performance of the model.

[0089] For training, semi-supervised pre-training can be performed on the graph spatio-temporal network first, using the following mean squared error (MSE) loss function:

[0090] Equation 7,

[0091] In the formula, represents the set of coefficients of the weight matrix W and the bias vector b of the network layer, y represents the ground observation value, represents the estimation result of the graph network for the target variable y with respect to the input x using the parameter of, is the loss function of the graph, and in this patent, it is the MSE loss function, is the regularization term of the parameters of the graph network G, using Elastic net regularization, and N represents the number of training samples;

[0092] After the initialization training, through the concatenation (concat) of the input unit and the output unit matrix of the graph network, the graph network of the network is fused with the fully residual depth network to form a unified network model for model training, using the following total mean squared error MSE loss function:

[0093] Equation 8,

[0094] In the formula, represents the set of coefficients of the weight matrix W and the bias vector b of the network layer, y represents the ground observation value, represents the estimation result of the network for the target variable y with respect to the input x using the parameter of, is the regularization term of the parameters of the graph network G, using Elastic net regularization, and N represents the number of training samples;

[0095] The output result of the model training is subjected to anti-regularization operation with the previously saved regularization parameters to obtain the output result at the scale of the original input variables, and the corresponding test indicators are calculated. Through the anti-regularization process for reverse deduction, the accuracy of the training process is verified.

[0096] Step 8: Optimize the model parameters to obtain a model with optimal parameters.

[0097] After establishing the model and conducting preliminary training, for the key parts during model training, including the number of local graph convolutional layers K, activation function, nodes and depth of each residual layer, and the size of each batch of training samples, systematic training comparisons are carried out to obtain a model with optimal parameters. Moreover, to verify the optimization performance of the model, finally, the average measurement value is taken after multiple trainings as the evaluation index of the model.

[0098] Also as Figure 1 shown, after obtaining the optimal model, the model can be called to predict the air pollution situation.

[0099] Step 9: After the model training and parameter tuning are completed, save the model. When applying the model to a new dataset, it is necessary to first regularize the parameters, then use the same k-NN to find the nearest multi-layer nearest neighbor nodes, and then call the model for prediction to obtain the prediction results of the new dataset; the prediction results are the prediction results of the air pollution situation.

[0100] Compared with the existing technology, this patent mainly solves the following four problems:

[0101] (1) Through the weighted multi-layer local graph convolutional network (graph convolutional network, abbreviated as GCN), the lag modeling function of irregular spatio-temporal samples is realized, solving the defect that the conventional time series model requires dense sample input of continuous time series in air pollution estimation;

[0102] (2) Through the statistical correlation ratio and multi-layer local graph convolution, the spatio-temporal modeling function of spatio-temporal dimension fusion is realized. Through graph convolution operations, the spatio-temporal correlation of irregular samples of air pollution detection data is better extracted, effectively solving the deficiencies of spatio-temporal stationarity and variation fitting in the traditional geostatistics field;

[0103] (3) Adopting the hierarchical graph aggregation operation based on local graph convolution, avoiding the disadvantage that the full graph matrix calculation requires a large amount of computing resources, improving the efficiency of the graph convolution operation for extracting spatio-temporal information correlation, and thus enabling flexible fusion with the full residual network;

[0104] (4) By fusing the full residual depth network, combining the extracted spatio-temporal correlation with its own feature prediction ability, compared with the traditional spatio-temporal modeling method of air pollutants, the generalization function of the model is fully improved. Compared with other methods, high-precision prediction results are obtained in the location-based independence test.

[0105]

Embodiment

[0106] The following further elaborates in detail on the air pollution prediction method of the fully residual deep network based on graph convolution disclosed in the present invention in combination with specific embodiments. In this embodiment, the spatio-temporal prediction of the concentration of nitrogen dioxide (NO2), the main air pollutant in the monitored and predicted area in 2015, is taken as an example to illustrate the application and advantages of this invention patent.

[0107] Step 1: Figure 1 The main steps of the present invention are shown. As shown in the first step, the hourly NO2 monitoring data of ground monitoring stations in 2015 are adopted. These data are sorted to obtain average concentration data, and corresponding meteorological data (including air temperature, wind speed, relative humidity), MERRA2 (the Modern-Era Retrospective analysis for Research and Applications, Version 2) coarse-resolution satellite variables (including ozone and planetary boundary layer height (PBLH), cloud fraction data, coordinates (x, y, and xy), elevation at 500-meter resolution, OMI-NO2 variable of coarse-resolution satellite, main road network density, closest distance to the road, land use type of pollution sources, data related to POI (point of interest) and pollution sources, time indices (year-day, week, date), etc.) are collected. These data are preprocessed to delete invalid and missing values, and 16 prediction covariates and 1 target variable (ground NO2 concentration, unit: μg / m 3 ) are obtained. The processing and operation of the data mainly adopt relevant packages of Python language and R statistical software.

[0108] Step 2: Find the nearest neighbor point pairs in the spatial dimension and time dimension. For each sample point, find the nearest sample point along the space, and remove the repeated points. Finally, a set of spatial nearest neighbor point variable pairs is formed , where x nt is the nearest neighbor point of x st in space; find the nearest sample point along the time, and remove the repeated points. Finally, a set of variable pairs is formed , where x sn is the nearest neighbor point of x st in time. Calculate the Pearson correlation coefficient of the two sets of nearest neighbor point variable pairs in the spatial and time dimensions, calculate the ratio of the two, and obtain the correlation ratio of the two dimensions. In terms of the algorithm for nearest neighbor retrieval, the KD tree is adopted to achieve fast retrieval of the nearest neighbor. In the specific implementation, cKDTree of the scipy package of Python is applied to achieve fast retrieval.

[0109] Step 3: Regularize the predictive covariates and target variables of the data, save the regularization parameters (the mean and variance of each variable), and determine the spatio-temporal dimensions (x, y, t). In terms of implementation, save each relevant parameter into the data structure of a Python dictionary.

[0110] Step 4: Use the k-nearest neighbor (k-NN) algorithm to find the k nearest neighbors of each node in the sample points, and record their indices and spatio-temporal distances. Refer to the kNN algorithm in the Python package torch_cluster, and also return the spatio-temporal distances between the sample points for subsequent program calls.

[0111] Step 5: Use location-based stratified random sampling. Using the climate zone zoning map as the stratification factor, first divide the spatial locations into 4 equal parts. Among them, the spatio-temporal samples corresponding to 3 parts of the sample spatial locations are used for training, while the spatio-temporal samples corresponding to the remaining 1 equal part of the spatial locations are used for location-based independence testing.

[0112] Step 6: Establish a multi-level nearest neighbor local topology. Each level is equivalent to a time lag. Here, K = 3 is used as the number of levels according to the sensitivity analysis, and the network topology structure is saved in the inner data. In terms of implementation, refer to the NeighborSampler function in the network modeling package of the Python graph network package Torch Geometric, input the node nearest neighbor topology in Step 4, and realize the extraction of the nearest neighbor network nodes layer by layer.

[0113] Step 7: Model establishment. First, set the objective function as the output NO2 concentration, pre-train the graph model, and then link the graph model with the fully residual deep network for model training. The model is established based on the Python deep learning package Pytorch. The widely favored graph neural network deep learning package in the industry, namely the Pytorch Geometric Library, is used. The basic graph network model in it is used as the basis for our spatio-temporal algorithm, realizing the spatio-temporal module of the multi-level local graph network based on spatio-temporal and the fully residual network model module. Through the splicing operation of the input, the local multi-level graph spatio-temporal network is combined with the fully residual to realize the final model of this patent. The pre-trained model first performs pre-training on the graph network, and the output of the pre-trained graph network is embedded into the input, realizing the "end-to-end" network structure of the overall network and improving the efficiency of training, prediction, and application.

[0114] Step 8: Parameter Tuning. A training algorithm of grid search with multiple parameters is set, and the main parameters to be investigated are: the number of convolutional layers K, the activation function, the number of nodes and depth of each residual layer, and the size of each batch of training samples. Finally, the optimal NO2 concentration prediction model is obtained, and the optimized parameters are: K = 3; the activation function for the middle layer is ReLU, and the output layer is a linear output activation; the network topology of the encoding layer of the residual layer is: [512, 320, 256, 128, 96, 64, 32, 16, 12]; the optimal parameter for batch training is 1024.

[0115] To verify the optimized performance of the model, in this step, the model is trained 6 times to obtain the mean value of the test parameters as the evaluation index of the model; at the same time, other models are trained with the same data, and the average value is taken after training 6 times in the same way to compare and verify the advantages of this model. The test items include ordinary random tests and location-based random tests. The results are shown in Table 1.

[0116] Table 1

[0117]

[0118] The test data in Table 1 show that this invention patent has achieved the best performance (the highest test R 2 , the lowest RMSE) in the location-based independence verification compared with other traditional spatio-temporal modeling methods, including Kriging and GAM, as well as graph networks and fully residual deep networks. It can also be seen from the test results that in terms of ordinary random sample tests, the results obtained by the method of this patent are similar to those of the fully residual deep network model, but in the location-based independence test, the method of this patent has obtained better results, improving the R 2 by 9% and reducing the RMSE by 1.35 μg / m 3 . The location-based independence test shows that this patent has better generalization performance compared with other methods.

[0119] Step 9: Prediction and Application of the Model. Store the samples used for training, the obtained regularization parameters, and the trained model mentioned above. In actual application, the k-NN algorithm can be used to find the nearest neighbor points of the training samples first, and then the trained model is called to predict the air pollution situation to obtain the prediction results of the target nodes of the new data set. The prediction results are the prediction results of the air pollution situation.

[0120] Compared with the prior art, the method for establishing a model and its application of a fully residual deep network based on graph convolution disclosed in this invention have the following technological improvements:

[0121] (1) In terms of the fusion of spatio-temporal dimensions, a method is proposed to calculate the spatio-temporal dimension correlation coefficient ratio based on the nearest neighbor pairs as an adjustment factor for spatio-temporal dimension fusion. By this, a quantitative comparison of time and space in the statistical sense can be obtained, overcoming the differences in measurement units and scales between time and space. Through the correlation coefficient ratio correction factor and the regularization technique to eliminate the measurement unit differences in time and space, the two are combined to effectively fuse the time and space dimensions, laying a foundation for spatio-temporal modeling using graph convolutional networks.

[0122] (2) A local multi-layer weighted graph neural network is used to simulate the multi-lag effect in time series modeling, effectively enabling the model to handle irregularly distributed spatio-temporal samples, avoiding the need for continuous sample input in time series models, and effectively improving the efficiency and accuracy of spatio-temporal modeling. Through local graph convolutional operations, the problem of large matrix calculations caused by full graph networks in the case of massive data is avoided, effectively improving the computational efficiency of spatio-temporal modeling; and through local multi-layer weighted graph convolution, using the strong generalization ability of deep learning, the prerequisite of spatio-temporal stationarity required by variograms is effectively avoided, and the problems of difficult fitting or overfitting of spatio-temporal variograms are avoided.

[0123] (3) According to the principle of the first law of geography "things are more related to nearby things", an aggregation function for spatio-temporal proximity calculation is established. Different from the previous SAGE network that considers all features, this patent only considers spatio-temporal dimensions for spatio-temporal proximity point analysis, and establishes a function for feature aggregation to achieve spatio-temporal modeling.

[0124] (4) The spatial graph convolution of spatio-temporal modeling is fully combined with the full residual depth network, giving full play to the advantages of spatio-temporal feature extraction and deep residual learning. As shown by the tests, the combination of the two greatly improves the generalization ability of the model. If separated, spatio-temporal modeling overly emphasizes the influence of adjacent points and has limited predictive ability for its own features; while the residual depth network lacks the ability to capture adjacent point information and is prone to overfitting. Only by combining the two, fully considering the influence of its own features and adjacent points, can the optimal model generalization performance be obtained.

[0125] (5) Using location-based independence testing to measure the more real performance of the model is a key step in measuring this method. Using ordinary cross-validation methods cannot measure the more real generalization of the model. Samples at one location may have partially participated in data training, which may lead to untrue test performance; while using location-based sampling training to ensure that no time series samples at the test sample location participate in training, the real generalization performance can be tested.

[0126] (6) Using a modeling method based on spatio-temporal local graph models also improves the computational efficiency of prediction using graph networks. The system only needs to extract the nearest neighbor points in the K-layer spatio-temporal dimensions to complete the prediction of spatio-temporal points.

[0127] The above embodiments are not limitations on the present invention, nor is the present invention limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the technical solution of the present invention also fall within the protection scope of the present invention.

Claims

1. A method for establishing a model of a fully residual deep network based on graph convolution, characterized in that: It includes the following steps: Step 1: Collect, organize the original air pollution data and its covariate data and perform preprocessing; Step 2: In the spatial dimension and time dimension, find the nearest adjacent point pairs for the sample points of the preprocessed data and calculate the ratio of the correlation coefficients in the two dimensions; Step 3: Perform regularization processing on all data to determine the spatio-temporal dimension; all data includes the target variable and the above-mentioned covariates, and the target variable refers to the air pollution concentration; Step 4: Use the k-NN nearest neighbor algorithm to find the k nearest neighbors and their spatio-temporal distances of all sample points; Step 5: Use the method of location-based stratified random sampling, with administrative divisions or climate zone divisions as the stratification factors, to divide the training and test samples; Step 6: For air pollution data, establish a multi-level nearest neighbor network topology. From the outermost layer to the inner layer to the target node, the influence of adjacent points is transmitted layer by layer and finally aggregated to the target node; from the three spatio-temporal dimensions, including the spatial dimensions x and y, the time dimension t, and determine the nearest neighbors according to the principle of the first law of geography that the closer the more relevant. At the same time, due to the irregular distribution of spatio-temporal sample points itself, aggregate its distance weight coefficients for the nearest neighbors, define the aggregation function as a weighted sum, or a pool operation; after the convolution operation, update the state of the target node; finally, after the nearest neighbor aggregation operation of the total number of layers, obtain the output of the target node; Among them, each level is equivalent to a lag; Step 7: Establish an air pollution prediction model based on a graph convolutional fully residual deep network. The established model consists of two parts, one part is a graph convolutional network, and the other part is a fully residual deep network; pre-train the graph convolutional network and fuse the fully residual network for training; Step 8: Optimize the model parameters to obtain the model with the optimal parameters; Step 9: Call the model to predict the air pollution situation.

2. The method for establishing a model of the fully residual depth network based on graph convolution according to claim 1, wherein: In Step 1, after collecting the original air pollution data, organize it, determine the spatio-temporal resolution, extract and preprocess the corresponding covariate factors; the covariate factors include coordinates, elevation, meteorology, land use, MERRA2 data with a coarse resolution, and time index.

3. The method for establishing a model of the fully residual deep network based on graph convolution according to claim 1, wherein: In Step 2, calculate the Pearson correlation coefficients of the spatial dimension and the time dimension respectively, calculate the ratio of the two, and obtain the ratio of the correlation coefficients in the two dimensions.

4. The method for establishing a model of the fully residual depth network based on graph convolution according to claim 1, wherein: In Step 3, use the standardized data regularization technology for processing, as shown in Equation 1: Where x' j represents the original input of the j-th characteristic variable, μ(x' j ) represents the mean of the original input, σ(x' j ) represents the variance of the original input, and x j is the regularized variable.

5. The method for establishing a model of the fully residual deep network based on graph convolution according to claim 1, wherein: In Step 4, when calculating the spatio-temporal distance, use the ratio of the correlation coefficients in the two dimensions calculated in Step 2 to adjust the spatio-temporal distance calculation process to make the time and space dimensions statistically consistent when calculating the distance.

6. The method for establishing a model of the fully residual deep network based on graph convolution according to claim 1, wherein: In Step 5, divide the area according to administrative divisions or climate zones to obtain the locations of spatial monitoring points, randomly select and divide them into M equal parts. Among them, 60-90% of the sample points corresponding to the monitoring point locations in each part are used for model training, and the samples corresponding to the remaining monitoring point locations are used for the independence test of the location.

7. The method for establishing a model of the fully residual deep network based on graph convolution according to claim 1, wherein: In step six, a graph network capable of processing irregular sample points is adopted. One lag represents the influence of the neighbors of a nearest spatio-temporal point, and the next lag is for the nearest neighbors of the connected target node. The lag is determined using the autocorrelation function graph, and the lag step K, i.e., the total number of layers, is explored based on the actual data. The specific process of multi-level lag spatio-temporal modeling is as shown in Equation 2: where k represents the index of the layer, u is the target node, represents that v is a neighbor of u, is the output of the k-th layer of the neighboring node v, is the output of the (k - 1)-th layer of the neighboring node v, is the aggregation result of the target node u obtained from the nearest neighbors, equivalent to the k-th layer state neighbors of u is the output of the aggregation function, aggregate (k) represents the aggregation function of the neighbors in the k-th layer; The aggregation function is defined as a weighted sum or a pool operation. For the weighted aggregation function, the output of the k-th layer of the target node u is as shown in Equation 3: wherein, represents the distance between the neighboring node v and the target node u, and WMEAN represents the inverse distance weighted summation; For the pool operation, the following aggregation function is used, and the output of the k-th layer of the target node u is as shown in Equation 4: In the formula, σ represents the sigma non-linear activation function, W represents the weight coefficient matrix of the pool layer, and b represents the bias vector. After k-layer convolution operations, the state of the target node u is updated to: ∪ represents the concatenation of two output vectors, represents the output of node u at the k-th layer when updating the state, represents the output of its own (k - 1)-th layer; After K-layer nearest neighbor aggregation operations, the output of the target node u is obtained: where Z u represents the estimated value of the target node to be finally estimated.

8. The method for establishing a model of the fully residual depth network based on graph convolution according to claim 7, characterized in that: In step seven, the graph convolution network module is a spatio-temporal GCN module that considers multi-level local convolutions. The output of this module is concatenated with the output of the original node as the input to the fully residual depth network. The training process of the model is as follows: First, semi-supervised pre-training is performed on the graph spatio-temporal network, and the MSE loss function shown in Equation 7 is used: where, θ W,b represents the coefficient set of the weight matrix W and the bias vector b of the network layer, y represents the ground observation value, represents the estimation result of the target variable y for the input x by the graph network using the parameter θ W,b , Ω(θ W,b ) is the regularization term of the θ W,b parameters of the graph network G, and N represents the number of training samples; After initial training, the graph network of the network is fused with the fully residual depth network by concatenating the input unit with the output unit matrix of the graph network to form a unified network model for model training. The following total mean square error MSE loss function is used: where θ W,b represents the set of coefficients of the weight matrix W and the bias vector b in the network layer, y represents the ground observation value, represents the estimated result of the network for the target variable y with respect to the input x using the parameter θ W,b , Ω(θ W,b ) is the regularization term of the θ W,b parameters of the graph network G, and N represents the number of training samples; The output result of model training is anti-regularized with the previously saved regularization parameters to obtain the output result of the original input variable scale, and the corresponding test metrics are calculated.

9. The method for establishing a model of the fully residual deep network based on graph convolution according to claim 1, characterized in that: In step eight, after establishing the model and performing preliminary training, systematic training comparisons are made for the key parts during model training, including the number of local graph convolution layers K used, the activation function, the nodes and depths of each residual layer, and the size of each batch of training samples, to obtain the model with the optimal parameters.

10. The method for establishing a model of the fully residual depth network based on graph convolution according to claim 1, characterized in that: In step nine, after model training and parameter tuning are completed, the model is saved. Then, when applying the model to a new dataset, parameter regularization is required first, then the same k-NN is used to find the nearest multi-level nearest neighbor nodes, and then the model is called for prediction to obtain the prediction result of the new dataset, which is the prediction result of the air pollution situation.

Citation Information

Patent Citations

  • Pollution source identification method based on correlation coefficient and monitoring stationing method

    CN102967689A

  • City road traffic pollutant discharge monitoring early warning system and method

    CN104318315A

  • Multi-task urban space-time prediction method based on adversarial learning

    CN111626490A

  • Air quality prediction method for multi-task learning based on multi-dimensional secondary feature extraction

    CN111814956A

  • Intelligent air pollution monitoring system

    CN108333314A