An air quality forecasting method based on integrated deep learning

By adopting integrated deep learning methods in air quality forecasting, optimizing the weight parameters of the traditional mode, building graph neural networks and multi-layer perceptron models, the problem of high computing resource consumption in traditional methods is solved, and more efficient prediction and lower costs are achieved.

CN119474769BActive Publication Date: 2025-05-13SICHUAN GUOLAN ZHONGTIAN ENVIRONMENTAL TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510066210.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

Traditional numerical models operate for a long time in air quality forecasting and have high computing resource requirements, resulting in high storage, operation and maintenance and computing costs.

Method used

The air quality forecasting method based on integrated deep learning is adopted, and by using the output results of the traditional mode as the label of the model, the weight parameters of the traditional mode are learned and optimized, the graph neural network and multi-layer perceptron model are constructed to reduce the consumption of computing resources.

Benefits of technology

It significantly shortens the prediction time, improves prediction efficiency, reduces costs, provides faster and more accurate air quality data support, and improves the scientific nature of pollution control and the effectiveness of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474769B_ABST
    Figure CN119474769B_ABST
Patent Text Reader

Abstract

The present invention discloses an air quality forecasting method based on integrated deep learning, which obtains a sample data set of grids in a monitoring area, analyzes the correlation of data between grids based on the Pearson correlation coefficient, obtains the label correlation coefficient of the model training label and the feature correlation coefficient of the model training feature; uses each grid as each node of a neural network graph, judges the connection relationship between nodes according to the feature correlation coefficient and the label correlation coefficient between each grid, and generates a connection matrix list; uses the neural network graph structure and the connection matrix list as the input of the graph neural network layer, uses the output of the graph neural network layer as the input of the MLP fully connected layer, and obtains the air quality prediction result in the monitoring area after the nonlinear combination of the MLP fully connected layer. By using the output result of the traditional mode as the label of the model, learning and optimizing the weight parameters of the traditional mode, the prediction time can be significantly shortened and the prediction efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of air quality monitoring, and in particular to an air quality forecasting method based on integrated deep learning. Background Art

[0002] Air pollution is mainly caused by a variety of human activities such as industrial production, agricultural activities, daily life, transportation and urbanization. At the same time, meteorological conditions are also an important natural factor affecting air pollution. Air pollution is not only limited to a certain area, but may also cross regions and even the world. Traditional numerical models, such as the commonly used mesoscale meteorological forecast model WRF, have the characteristics of high resolution, flexibility, powerful physical process description, rich data input, etc., and can be applied to weather forecasting, climate simulation, disaster warning, agricultural climate services, etc. There is also a model CMAQ specifically used for air quality forecasting and assessment. Its biggest feature is the concept of one atmosphere (One-Atmosphere), which breaks through the traditional model of simulating single species and single-phase species, and comprehensively handles complex air pollution conditions such as tropospheric ozone, PM, toxic substances, acid deposition and visibility. It is used for multi-scale, multi-pollution air quality forecasting, assessment and decision-making research and other purposes. It is an important means to study the spatiotemporal distribution and component contribution of atmospheric pollutants.

[0003] These two models can help researchers understand key information such as the source, cause, degree of pollution, duration, main components and relative contribution of various factors of atmospheric pollution. However, these traditional numerical models take a long time to run and have extremely high demands on computing resources. For example, deploying a set of WRF climate field forecast models and CMAQ pollution field forecast models locally usually requires a computer with a 64-core processor and 128GB of memory to run around the clock. In addition, these models generate about 20GB of data per day, which has serious storage costs, operation and maintenance costs, and computing costs. Summary of the invention

[0004] The purpose of the present invention is to provide an air quality forecasting method based on integrated deep learning. In order to reduce the consumption of computing resources, the output results of the traditional model are used as the label of the model to learn and optimize the weight parameters of the traditional model, which can significantly shorten the prediction time and improve the prediction efficiency.

[0005] To achieve the above objectives, this application adopts the following scheme:

[0006] The present invention provides an air quality forecasting method based on integrated deep learning, which specifically comprises the following steps:

[0007] S1. Obtain a sample data set corresponding to each grid in the monitoring area, wherein the sample data set includes a model training feature based on WRF simulation and a model training label based on CMAQ simulation;

[0008] S2. Based on the Pearson correlation coefficient, the correlation between the data of each grid is analyzed to obtain the label correlation coefficient of the model training label and the feature correlation coefficient of the model training feature;

[0009] S3, constructing a neural network graph structure, taking each grid as each node of the neural network graph, judging the connection relationship between the nodes according to the feature correlation coefficient and the label correlation coefficient between the grids, and constructing a connection matrix list of the neural network graph structure according to the connection relationship between the nodes;

[0010] S4. Construct an integrated forecasting model. The integrated forecasting model includes a graph neural network layer and an MLP fully connected layer. The neural network graph structure and the connection matrix list are used as the input of the graph neural network layer, and the output of the graph neural network layer is used as the input of the MLP fully connected layer. After the nonlinear combination of the MLP fully connected layer, the air quality prediction results in the monitoring area are obtained.

[0011] In some specific embodiments, the specific process of step S1 includes:

[0012] S11, obtaining the WRF data set of the WRF simulation and the CMAQ data set of the CMAQ model, wherein the WRF data set includes the following fields: upward wave radiation flux GLW at the surface, terrain height HGT, boundary layer height PBLH, surface air pressure PSFC, 2-meter humidity at the ground Q2, precipitation RAINNC, 2-meter temperature at the ground T2, surface temperature TSK, 10-meter wind field latitude component U10, 10-meter wind field longitude component V10, grid ID, and release time;

[0013] The CMAQ dataset includes the following fields: PM25_TOT, PM10, O3, CO, NO2, SO2, grid ID, and release time;

[0014] S12. Convert the data formats of the WRF dataset and the CMAQ dataset to obtain the datasets. as well as ;

[0015] Convert the field values ​​of GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, and V10 to floating point numbers, the field value of grid ID to positive integers, and the field value of release time to time type;

[0016] Convert the field values ​​of PM25_TOT, PM10, O3, CO, NO2, and SO2 to floating point numbers, the field value of grid ID to positive integers, and the field value of release time to time type;

[0017] S13. Dataset and The data in are merged according to the grid ID field to obtain the sample data set DATA;

[0018] S14. Mining time series features according to the field value of the release time field in the sample data set DATA. The time series features include year, month, week, whether it is a holiday, day, and hour.

[0019] S15. Use GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, V10, and time series features as model training features, and PM25_TOT, PM10, O3, CO, NO2, and SO2 as model training labels.

[0020] In some specific implementation schemes, after obtaining the sample data set DATA in step S1, the missing values ​​in the PM25_TOT, PM10, O3, CO, NO2, and SO2 fields in the sample data set DATA are filled, and the filling methods are respectively:

[0021] Sampling linear regression fills the missing values ​​of PM25_TOT, PM10, and O3 fields. The calculation formula is:

[0022]

[0023] Among them, PM25 null Indicates missing value of PM25_TOT field, PM10 null Indicates missing values ​​in the PM10 field, O3 null Represents the missing value of the O3 field, PM10, PM25, hour, and weekday represent the field values ​​of PM25_TOT, PM10, hour, and weekday, W1, W2, and W3 represent the corresponding weights, Solar_radiation represents the solar radiation characteristics, and Temp represents the temperature characteristics; CO, NO2, and SO2 all use historical average values ​​to fill missing values.

[0024] In some specific implementation schemes, the specific process of step S3 determining the connection relationship between nodes is as follows:

[0025] Calculate the aggregation correlation coefficient between nodes based on the label correlation coefficient and feature correlation coefficient between grids;

[0026] When the aggregation correlation coefficient is greater than 0.7, it is judged that there is a bidirectional connection line between the nodes;

[0027] When the aggregation correlation coefficient is less than 0.7 and the label correlation coefficient is greater than 0.5, it is judged that there is a unidirectional connection line between the nodes.

[0028] In some specific implementation schemes, the method for constructing the connection matrix in step S3 is:

[0029] Each node is numbered, and for each node, a set of related connection matrices related to the current node is obtained according to the connection relationship between other nodes and the current node;

[0030] Count the relevant connection matrices of all nodes and get the connection matrix list of the neural network graph structure.

[0031] In some specific implementation schemes, the data format of the input graph neural network layer is three-dimensional data, where the three dimensions are the number of nodes, the number of samples corresponding to each node, and the number of model training features corresponding to each sample, wherein a fixed amount of data corresponding to each node obtained from the sample data set DATA is used as sample data, and each sample data contains 16 features, namely: GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, V10, year, month, week, whether it is a holiday, day, and hour.

[0032] In some specific implementation schemes, the graph neural network layer includes two layers of SAGEConv feature extraction layers, which expand the features corresponding to each sample in the three-dimensional data twice, and the calculation formula for the output of each layer of SAGEConv feature extraction layer is:

[0033] ;

[0034] Among them, L represents the number of model training features input to the current feature extraction layer, represents the rth model training feature value of the hth node, represents the weight parameter corresponding to the r-th model training feature value of the h-th node, represents the weight parameter corresponding to the model training feature value of the jth node connected to the hth node, It represents the rth model training feature value of the jth node connected to the hth node, and mean represents the average value calculation.

[0035] In some specific implementations, the MLP fully connected layer includes four linear connection layers, which are used to expand or compress features in three-dimensional data, wherein the calculation formula for the output of each linear connection layer is:

[0036]

[0037] Among them, M represents the number of model training features input to the current linear connection layer, represents the i-th model training feature value of the N-th node, Represents the weight parameter corresponding to the i-th model training feature value of the N-th node, It represents the bias corresponding to the i-th model training feature value of the N-th node, and exp represents the activation function.

[0038] In some specific embodiments, the MLP fully connected layer also includes a skip connection layer, the four linear connection layers include a first fully connected layer, a second fully connected layer, a third fully connected layer and a fourth fully connected layer, the skip connection layer is used to connect the first fully connected layer and the third fully connected layer, the output of the first fully connected layer is used as the input of the second fully connected layer, the output of the second fully connected layer is used as the input of the third fully connected layer, the input of the skip connection layer is the output of the first fully connected layer and the output of the third fully connected layer, the output of the skip connection layer is used as the input of the fourth fully connected layer, and the number of model training features output by the first fully connected layer and the third fully connected layer is the same.

[0039] In some specific embodiments, the output of the skip connection layer:

[0040] ;

[0041] in Skip Indicates the connection result. Feature 1 represents the output of the first fully connected layer, Feature 3 represents the output of the third fully connected layer.

[0042] The present invention has the beneficial effects:

[0043] This application makes full use of the multi-model fusion strategy, adopts the graph neural network (GNN) in deep learning as the feature extraction head, and deeply explores the feature relationship between nodes. On this basis, the multi-layer perceptron (MLP) is combined to extract multi-layer nonlinear combination features, and the skip connection is used to prevent the model from overfitting. In this way, the parameterized conversion from the traditional model to deep learning is effectively realized. The advantage of this conversion is that since pollution control is a business field with very strong timeliness, the traditional model has high costs in terms of running time, manpower costs and resource consumption. Through deep learning methods, prediction results can be quickly generated in a short time, which greatly reduces costs and improves efficiency. This not only provides fast and effective data support for urban air quality pollution, but also provides decision makers with a more scientific and accurate basis, improving the overall control level. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of an air quality forecasting method based on integrated deep learning provided by an embodiment of the present invention;

[0045] Figure 2 A schematic diagram of the structure of an integrated forecasting model provided by an embodiment of the present invention;

[0046] Figure 3 A schematic diagram of a neural network graph structure provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] The relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention unless specifically stated otherwise.

[0049] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0050] Additionally, descriptions of well-known structures, functions, and configurations may be omitted for clarity and conciseness.One of ordinary skill in the art will recognize that various changes and modifications may be made to the examples described herein without departing from the spirit and scope of the present disclosure.

[0051] Technologies, methods, and apparatus known to ordinary technicians in the relevant field may not be discussed in detail, but where appropriate, such technologies, methods, and apparatus should be considered part of the authorization specification.

[0052] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0053] Example 1

[0054] like Figure 1 As shown, this embodiment provides an air quality forecasting method based on integrated deep learning, which specifically includes the following steps:

[0055] S1. Obtain a sample data set corresponding to each grid in the monitoring area, wherein the sample data set includes a model training feature based on WRF simulation and a model training label based on CMAQ simulation;

[0056] The sample dataset is generated as follows:

[0057] S11, obtaining the WRF data set of the WRF simulation and the CMAQ data set of the CMAQ model, wherein the WRF data set includes the following fields: upward wave radiation flux GLW at the surface, terrain height HGT, boundary layer height PBLH, surface air pressure PSFC, 2-meter humidity at the ground Q2, precipitation RAINNC, 2-meter temperature at the ground T2, surface temperature TSK, 10-meter wind field latitude component U10, 10-meter wind field longitude component V10, grid ID, and release time;

[0058] The CMAQ dataset includes the following fields: PM25_TOT, PM10, O3, CO, NO2, SO2, grid ID, and release time;

[0059] S12. Convert the data formats of the WRF dataset and the CMAQ dataset to obtain the datasets. as well as ;

[0060] Convert the field values ​​of GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, and V10 to floating point numbers, the field value of grid ID to positive integers, and the field value of release time to time type;

[0061] Convert the field values ​​of PM25_TOT, PM10, O3, CO, NO2, and SO2 to floating point numbers, the field value of grid ID to positive integers, and the field value of release time to time type;

[0062] S13. Dataset and The data in are merged according to the grid ID field to obtain the sample data set DATA.

[0063] S14. Mining time series features according to the field value of the release time field in the sample data set DATA. The time series features include year, month, week, whether it is a holiday, day, and hour.

[0064] S15. Use GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, V10, and time series features as model training features, use the field values ​​of GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, V10, and time series features as model training feature values, and use PM25_TOT, PM10, O3, CO, NO2, and SO2 as model training labels.

[0065] In order to reduce the error caused by the traditional method of filling the missing values ​​of the fields with the average value, the linear regression model is used to predict the missing values ​​of the data when filling the missing values ​​of the model training feature values ​​in this embodiment, which can better retain the relationship between the features and maintain the trend and dynamic changes of the features themselves, thereby improving the overall accuracy and reliability of the model. When dealing with complex data such as air pollution, this strategy can ensure the quality of the data and provide a more reliable basis for subsequent analysis and prediction. Specifically, after obtaining the sample data set DATA, the missing values ​​in the PM25_TOT, PM10, O3, CO, NO2, and SO2 fields in the sample data set DATA are filled, and the filling methods are:

[0066] The PM25_TOT, PM10, and O3 missing value filling scheme is based on linear regression. This scheme is adopted because PM25 and PM10 data have a very strong linear characteristic relationship between themselves, and O3 and meteorological data temperature and solar radiation also have a very strong linear characteristic relationship.

[0067] Sampling linear regression fills the missing values ​​of PM25_TOT, PM10, and O3 fields. The calculation formula is:

[0068]

[0069] Among them, PM25 null Indicates missing value of PM25_TOT field, PM10 null Indicates missing values ​​in the PM10 field, O3 null Indicates the missing value of the O3 field, PM10, PM25, hour, and weekday indicate the field values ​​of PM25_TOT, PM10, hour, and weekday corresponding to the missing data, W1, W2, and W3 indicate the corresponding weights, Solar_radiation indicates solar radiation characteristics, and Temp indicates temperature characteristics;

[0070] CO, NO2, and SO2 all use historical average values ​​to fill missing values. For example, if the CO data at 3 p.m. is missing, the average value of all CO data at 3 p.m. in the past month is used to fill it.

[0071] S2. Based on the Pearson correlation coefficient, the correlation between the data of each grid is analyzed to obtain the label correlation coefficient of the model training label and the feature correlation coefficient of the model training feature;

[0072] Based on the Pearson correlation coefficient, the linear correlation between data is calculated. The connection matrix used to construct the graph neural network "The connection matrix is ​​a line segment that represents the relationship between two nodes". It can be understood that a feature correlation coefficient is calculated based on each model training feature, and then a label correlation coefficient is calculated based on each model training label. The feature correlation and label correlation can represent the intrinsic connection between grids, which is exactly the same as the statistical model of the graph neural network. Therefore, the grid can be used as a node of the graph neural network.

[0073] S3, constructing a neural network graph structure, taking each grid as each node of the neural network graph, judging the connection relationship between the nodes according to the feature correlation coefficient and the label correlation coefficient between the grids, and constructing a connection matrix list of the neural network graph structure according to the connection relationship between the nodes;

[0074] S31. Connection relationship judgment

[0075] Specifically, the grid data is converted into graph structure data, each grid is defined as a node, the connection matrix is ​​matched and calculated based on feature correlation and label correlation, and the aggregation correlation coefficient between nodes is calculated based on the label correlation coefficient and feature correlation coefficient between grids;

[0076] If the aggregate correlation coefficient is greater than 0.7, it proves that there is a bidirectional connection line between nodes. If the aggregate correlation coefficient is less than 0.7, it is necessary to continue to determine whether the label correlation is greater than 0.5. If it is greater than 0.5, it proves that there is a unidirectional connection line between nodes. In the feature correlation and label correlation screening, the unidirectional connection line determination is based on label correlation, which is based on the distribution of data results. Therefore, compared with the correlation between features and labels, since the model is result-oriented, the data correlation priority of labels is much higher.

[0077] The calculation formula is:

[0078]

[0079] Among them, Double represents a bidirectional connection, Single represents a unidirectional connection, and Feature corr Represents the summed average of the feature correlation coefficients of all model training features, Lable corr Represents the summed average of the label correlation coefficients for all model training labels.

[0080] S32, connection matrix construction

[0081] Each node is numbered, and for each node, a set of related connection matrices related to the current node is obtained according to the connection relationship between other nodes and the current node;

[0082] Count the relevant connection matrices of all nodes and get the connection matrix list of the neural network graph structure.

[0083] S4. Construct an integrated forecasting model. The integrated forecasting model includes a graph neural network layer and an MLP fully connected layer. The neural network graph structure and the connection matrix list are used as the input of the graph neural network layer, and the output of the graph neural network layer is used as the input of the MLP fully connected layer. After the nonlinear combination of the MLP fully connected layer, the air quality prediction results in the monitoring area are obtained.

[0084] The data format of the input graph neural network layer is three-dimensional data, and the three dimensions are the number of nodes, the number of samples corresponding to each node, and the number of model training features corresponding to each sample. Among them, a fixed number of data corresponding to each node obtained from the sample data set DATA is used as sample data. Each sample data contains 16 features, namely: GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, V10, year, month, week, whether it is a holiday, day, and hour.

[0085] like Figure 2 As shown in Figure 1, the integrated prediction model includes a graph neural network layer and an MLP fully connected layer. Specifically:

[0086] S41, the graph neural network layer includes two layers of SAGEConv feature extraction layers (SAGEConv1 and SAGEConv2 respectively). The two layers of SAGEConv feature extraction layers expand the features corresponding to each sample in the three-dimensional data twice. The calculation formula for the output of each layer of SAGEConv feature extraction layer is:

[0087]

[0088] Among them, L represents the number of model training features input to the current feature extraction layer, represents the rth model training feature value of the hth node, represents the weight parameter corresponding to the r-th model training feature value of the h-th node, represents the weight parameter corresponding to the model training feature value of the jth node connected to the hth node, It represents the rth model training feature value of the jth node connected to the hth node, and mean represents the average value calculation.

[0089] S42, the MLP fully connected layer includes four linear connection layers, which are used to expand or compress the features in the three-dimensional data. Each linear connection layer includes an input layer, a hidden layer and an output layer. The calculation formula for the output of each linear connection layer is:

[0090]

[0091] Among them, M represents the number of model training features input to the current linear connection layer, represents the i-th model training feature value of the N-th node, Represents the weight parameter corresponding to the i-th model training feature value of the N-th node, It represents the bias corresponding to the i-th model training feature value of the N-th node, and exp represents the activation function.

[0092] S43. In order to avoid the problems of gradient vanishing and gradient exploding due to the large model depth during deep learning, the present application adopts the design idea of ​​jump connection, including a jump connection layer (Jumping5) in the MLP fully connected layer, and the four linear connection layers include the first fully connected layer (LinearLayer1), the second fully connected layer (LinearLayer2), the third fully connected layer (LinearLayer3) and the fourth fully connected layer (LinearLayer4). The jump connection layer is used to connect the first fully connected layer and the third fully connected layer. The output of the first fully connected layer is used as the input of the second fully connected layer, and the output of the second fully connected layer is used as the input of the third fully connected layer. The input of the jump connection layer is the output of the first fully connected layer and the output of the third fully connected layer. The output of the jump connection layer is used as the input of the fourth fully connected layer. The number of model training features output by the first fully connected layer and the third fully connected layer is the same. The output of the jump connection layer is:

[0093]

[0094] in Skip Indicates the connection result. Feature 1 represents the output of the first fully connected layer, Feature 3 represents the output of the third fully connected layer.

[0095] In order to better understand the contents of the above steps S3-S4, an example is given below:

[0096] 1. Training data preparation for integrated forecasting model

[0097] Assume that the graph data corresponding to the training label of a certain model is constructed as follows Figure 3As shown in the figure, it can be seen that the assumed graph neural network includes 5 nodes, and each node is numbered in the order of 0, 1, 2, 3, and 4. The connection relationship between the 5 nodes is interspersed with unidirectional and bidirectional connection lines. The unidirectional connection indicates that there is a unilateral correlation, while the bidirectional connection indicates a mutual correlation. Assuming that each node has 100 samples and each sample has 16 features, the data representation is as follows:

[0098] Feature_data = {

[0099] 0: [100,16],

[0100] 1: [100, 16],

[0101] 2: [100, 16],

[0102] 3: [100, 16],

[0103] 4: [100, 16],

[0104] }

[0105] The data format input to the graph neural network layer is three-dimensional data = [5, 100, 16], where 5 represents the number of nodes, 100 represents the number of sample data, and there are 100 samples. The number of sample data is related to the time series, and 100 samples correspond to 100 time series data. 16 represents the number of model training features, and there are 16 model training features.

[0106] The connection matrix can be directly defined in the computer according to the node number, and the "order of front and back is not affected". From the figure, it can be seen that there are a total of 9 groups of data connection relationships between the one-way connection lines and the two-way connection lines between the nodes. A simple group is listed to explain. The relevant connection matrix related to node 0 is designed. Node 2 is connected to node 0 in a one-way direction and node 2 is connected in a two-way direction. According to the connection relationship between node 0 and node 1 and node 2, a group of data connection relationships (related connection matrix) between node 0 and node 1 is [0, 1], and a group of data connection relationships between node 0 and node 2 is [0,2]. Correspondingly, the 9 groups of data connection relationships can be obtained according to the nodes: node 0: [0,1], [0, 2], node 1: [1, 0], [1,3], [1, 4], node 2: [2,3], node 3: [3, 1], [3, 4], node 4: [4, 1];

[0107] Count the relevant connection matrices of all nodes into a two-dimensional connection matrix list Connection_data = [

[0108] Starting point of connecting line: [0,0,1,1,1,2,3,3,4],

[0109] Connecting line end point: [1,2,0,3,4,3,1,4,1] ]

[0111] The starting point of the connecting line represents the previous value in the relevant connection matrix, and the end point of the connecting line represents the next value in the relevant connection matrix.

[0112] 2. The integrated forecast model design of this embodiment adopts the Graph Neural Network (GNN) module, the Multi-Layer Perceptron (MLP) model and the Skip Connection component commonly used in computer vision in deep learning technology. Since air quality has spatial correlation and complexity, the pollution conditions in surrounding areas will affect each other. Therefore, by utilizing the characteristics of the graph neural network model and combining the connection matrix, the feature information of the surrounding area can be effectively captured and extracted. Furthermore, nonlinear combination is performed through the multi-layer perceptron to enhance the expression ability and generalization performance of the model. In addition, the introduction of the skip connection layer can effectively prevent the gradient explosion and gradient disappearance problems of the model during the training process, and ensure the stability and convergence of the model.

[0113] For the model training label corresponding to each pollutant, the regression label of the integrated prediction model will be set to the model training label that needs to be predicted currently before the neural network graph structure is input into the integrated prediction model, and finally the prediction result corresponding to the pollutant is obtained. In order to better illustrate the prediction results of pollutants in this application, taking a certain pollutant as an example, the prediction process of the model is shown in 2.1-2.3, and the prediction process of other pollutants is similar. The difference is that each time a prediction is made, the model training label of the integrated prediction model regression is changed, and the loss function corresponding to the current label is calculated according to the label of the integrated prediction model, and the loss corresponding to the current label is calculated according to the loss function. Then, the loss is back-propagated twice, and the chain is derived to obtain the gradient of the integrated prediction model under the current model training label, and the weight parameters in the following model are automatically updated to the weight parameters corresponding to the current model training label according to the gradient.

[0114] 2.1. Graph Neural Network: This application uses two layers of SAGEConv feature extraction layers. The calculation mode of this feature extraction layer is to assume that there is a connection matrix relationship with node 0, perform random sampling and weighted averaging of features to obtain connection information, and then add it to its own weighted features to obtain the graph neural network calculation output. The calculation formula is as follows.

[0115] The two-layer feature extraction layer expands the number of model training features twice. The calculation formula for the feature expansion of each SAGEConv feature extraction layer is:

[0116] The input of the first SAGEConv feature extraction layer is: (5, 100, 16)

[0117] The output of the first SAGEConv feature extraction layer is: (5, 100, SAGEConv1)

[0118]

[0119] Where L represents the number of input model training features (L is 16 in the first layer), represents the rth model training feature value of the hth node, represents the weight parameter corresponding to the r-th model training feature value of the h-th node, represents the weight parameter corresponding to the model training feature value of the jth node connected to the hth node, represents the rth model training feature value of the jth node connected to the hth node, It represents the set of nodes connected to the hth node, and mean represents the average value calculation.

[0120] The input of the second SAGEConv feature extraction layer is: (5, 100, SAGEConv1)

[0121] The output of the second SAGEConv feature extraction layer is: (5, 100, SAGEConv2)

[0122]

[0123] The calculation logic of SAGEConv2 is the same as that of SAGEConv1, and the parameter definitions are the same. The difference is that the value of L is SAGEConv1, and the other calculation processes are adaptively transformed.

[0124] 2.2. MLP network: This embodiment uses an MLP network. Each layer of the network is composed of an input layer, a hidden layer, and an output layer. Through nonlinear activation function approximation, it can fully and effectively mine multi-layer nonlinear combination relationships and mine multiple feature weights. The calculation formula is as follows.

[0125] The input of Linearlayer1 is: (5, 100, SAGEConv2)

[0126] The output of Linearlayer1 is: (5, 100, MLP1)

[0127]

[0128] Among them, M is the number of model training features input to the current linear connection layer (corresponding to the M hidden neurons contained in the hidden layer), that is, SAGEConv2, represents the i-th model training feature value of the N-th node, The weight parameter corresponding to the i-th model training feature value of the N-th node, represents the bias corresponding to the i-th model training feature value of the N-th node, and exp represents the activation function. The calculation formula for each layer below is the same, the difference is that the value of each parameter changes according to the input adaptability.

[0129] The input of Linearlayer2 is: (5, 100, MLP1)

[0130] The output of Linearlayer2 is: (5, 100, MLP2)

[0131]

[0132] The input of Linearlayer3 is: (5, 100, MLP2)

[0133] The output of Linearlayer3 is: (5, 100, MLP3)

[0134]

[0135] Where MLP3=MPL1

[0136] The input of Linearlayer4 is: (5, 100, Skip )

[0137] The output of Linearlayer4 is: (5, 100, MLP4)

[0138]

[0139] After the model training features are expanded twice in the feature extraction layer, they are expanded in the first fully connected layer, compressed in the second fully connected layer, expanded to the same output as the first fully connected layer in the third fully connected layer, and finally compressed to 1 (i.e., MLP4) in the fourth fully connected layer. The output three-dimensional data is then expanded to obtain the prediction result of the current model training label.

[0140] 2.3. Skip connection: This application uses the skip connection design idea, which was first applied to the computer vision ResNet network. Although deep learning can effectively extract features in long-term nonlinear combinations, the large depth of the model will lead to gradient vanishing and gradient explosion problems. Therefore, this design idea can effectively avoid the gradient problem, which can effectively avoid the gradient problem and improve the model depth and learning effect. The calculation formula is as follows.

[0141] in Skip Indicates the connection result. Feature 1 represents the number of model training features output by the first fully connected layer (MLP1), Feature 3 represents the number of model training features output by the third fully connected layer (MLP3).

[0142] It can be understood that the method of this embodiment has the following advantages:

[0143] 1. Model parameterization conversion: By using the output results of traditional models (WRF and CMAQ) as the labels of the model, learning and optimizing the weight parameters of the traditional model, the prediction time can be significantly shortened and the prediction efficiency can be improved. This conversion can not only effectively reduce the consumption of computing resources, but also enhance the real-time and accuracy of the model, thereby playing a greater role in air pollution control and improving the scientificity and effectiveness of decision-making.

[0144] 2. Multi-model fusion strategy: By combining feature extractors such as graph neural network (GNN), skip connection (SkipConnection) in computer vision, and multi-layer perception (MLP), we can make full use of the unique advantages of each model to achieve efficient learning and optimization of traditional model parameters.

[0145] 3. Automatic connection matrix setting: Automatically build the connection matrix between nodes by fusing feature correlation coefficients and label correlation coefficients. This connection matrix not only reflects the trend correlation between data, but also captures the complex dependencies between different nodes, providing a solid foundation for subsequent model feature extraction. This method can significantly improve the robustness and predictive ability of the model, ensuring that the model can more accurately capture key features and patterns when dealing with complex problems such as air pollution.

[0146] 4. Adaptive model filling strategy: By using the linear regression model to predict missing data values, the error caused by using the average value to fill in the traditional method can be effectively reduced. This method can not only better preserve the relationship between features, but also maintain the trend and dynamic changes of the features themselves, thereby improving the overall accuracy and reliability of the model. When dealing with complex data such as air pollution, this strategy can ensure the quality of the data and provide a more reliable basis for subsequent analysis and prediction.

[0147] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. According to the technical essence of the present invention, within the spirit and principles of the present invention, any simple modification, equivalent replacement and improvement made to the above embodiment still falls within the protection scope of the technical solution of the present invention.

Claims

1. An air quality forecasting method based on integrated deep learning, characterized in that: The specific steps include: S1. Obtain a sample data set corresponding to each grid in the monitoring area, wherein the sample data set includes a model training feature based on WRF simulation and a model training label based on CMAQ simulation; S2. Based on the Pearson correlation coefficient, the correlation between the data of each grid is analyzed to obtain the label correlation coefficient of the model training label and the feature correlation coefficient of the model training feature; S3, constructing a neural network graph structure, taking each grid as each node of the neural network graph, judging the connection relationship between the nodes according to the feature correlation coefficient and the label correlation coefficient between the grids, and constructing a connection matrix list of the neural network graph structure according to the connection relationship between the nodes; The specific process of step S3 determining the connection relationship between nodes is as follows: Calculate the aggregation correlation coefficient between nodes based on the label correlation coefficient and feature correlation coefficient between grids; When the aggregation correlation coefficient is greater than 0.7, it is judged that there is a bidirectional connection line between the nodes; When the aggregation correlation coefficient is less than 0.7 and the label correlation coefficient is greater than 0.5, it is judged that there is a one-way connection line between the nodes; The method for constructing the connection matrix in step S3 is: Each node is numbered, and for each node, a set of related connection matrices related to the current node is obtained according to the connection relationship between other nodes and the current node; Count the relevant connection matrices of all nodes to obtain the connection matrix list of the neural network graph structure; S4. Construct an integrated forecasting model. The integrated forecasting model includes a graph neural network layer and an MLP fully connected layer. The neural network graph structure and the connection matrix list are used as the input of the graph neural network layer, and the output of the graph neural network layer is used as the input of the MLP fully connected layer. After the nonlinear combination of the MLP fully connected layer, the air quality prediction results in the monitoring area are obtained.

2. The air quality forecasting method based on integrated deep learning according to claim 1, characterized in that: The specific process of step S1 includes: S11, obtaining the WRF data set of the WRF simulation and the CMAQ data set of the CMAQ model, wherein the WRF data set includes the following fields: upward wave radiation flux GLW at the surface, terrain height HGT, boundary layer height PBLH, surface air pressure PSFC, 2-meter humidity Q2 at the ground, precipitation RAINNC, 2-meter temperature T2 at the ground, surface temperature TSK, 10-meter wind field latitude component U10 at the ground, 10-meter wind field longitude component V10 at the ground, grid ID, and release time; The CMAQ dataset includes the following fields: PM25_TOT, PM10, O3, CO, NO2, SO2, grid ID, and release time; S12. Convert the data formats of the WRF dataset and the CMAQ dataset to obtain the datasets. as well as ; Convert the field values ​​of GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, and V10 to floating point numbers, the field value of grid ID to positive integers, and the field value of release time to time type; Convert the field values ​​of PM25_TOT, PM10, O3, CO, NO2, and SO2 to floating point numbers, the field value of grid ID to positive integers, and the field value of release time to time type; S13. Dataset and The data in are merged according to the grid ID field to obtain the sample data set DATA; S14. Mining time series features according to the field value of the release time field in the sample data set DATA. The time series features include year, month, week, whether it is a holiday, day, and hour. S15. Use GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, V10, and time series features as model training features, and PM25_TOT, PM10, O3, CO, NO2, and SO2 as model training labels.

3. The air quality forecasting method based on integrated deep learning according to claim 2, characterized in that: After obtaining the sample data set DATA in step S1, the missing values ​​in the PM25_TOT, PM10, O3, CO, NO2, and SO2 fields in the sample data set DATA are filled. The filling methods are: Sampling linear regression fills the missing values ​​of PM25_TOT, PM10, and O3 fields. The calculation formula is: Among them, PM25 null Indicates missing value of PM25_TOT field, PM10 null Indicates missing values ​​in the PM10 field, O3 null represents the missing value of the O3 field, PM10, PM25, hour, and weekday represent the field values ​​of PM10, PM25_TOT, hour, and weekday, W1, W2, and W3 represent the corresponding weights, Solar_radiation represents the solar radiation characteristics, and Temp represents the temperature characteristics; Historical average values ​​are used to fill missing values ​​for CO, NO2, and SO2.

4. The air quality forecasting method based on integrated deep learning according to claim 2, characterized in that: The data format of the input graph neural network layer is three-dimensional data, and the three dimensions are the number of nodes, the number of samples corresponding to each node, and the number of model training features corresponding to each sample. Among them, a fixed number of data corresponding to each node obtained from the sample data set DATA is used as sample data. Each sample data contains 16 model training features, namely: GLW, HGT, PBLH, PSFC, Q2, RAINNC, T2, TSK, U10, V10, year, month, week, whether it is a holiday, day, and hour.

5. The air quality forecasting method based on integrated deep learning according to claim 4, characterized in that: The graph neural network layer includes two layers of SAGEConv feature extraction layers. The two layers of SAGEConv feature extraction layers expand the model training features corresponding to each sample in the three-dimensional data twice. The calculation formula for the output of each layer of SAGEConv feature extraction layer is: ; Among them, L represents the number of model training features input to the current feature extraction layer, represents the rth model training feature value of the hth node, represents the weight parameter corresponding to the r-th model training feature value of the h-th node, represents the weight parameter corresponding to the model training feature value of the jth node connected to the hth node, It represents the rth model training feature value of the jth node connected to the hth node, and mean represents the average value calculation.

6. The air quality forecasting method based on integrated deep learning according to claim 4, characterized in that: The MLP fully connected layer includes four linear connection layers, which are used to expand or compress the features in the three-dimensional data. The calculation formula for the output of each linear connection layer is: ; Among them, M represents the number of model training features input to the current linear connection layer, represents the i-th model training feature value of the N-th node, Represents the weight parameter corresponding to the i-th model training feature value of the N-th node, It represents the bias corresponding to the i-th model training feature value of the N-th node, and exp represents the activation function.

7. The air quality forecasting method based on integrated deep learning according to claim 6, characterized in that: The MLP fully connected layer also includes a skip connection layer. The four linear connection layers include the first fully connected layer, the second fully connected layer, the third fully connected layer and the fourth fully connected layer. The skip connection layer is used to connect the first fully connected layer and the third fully connected layer. The output of the first fully connected layer is used as the input of the second fully connected layer, and the output of the second fully connected layer is used as the input of the third fully connected layer. The input of the skip connection layer is the output of the first fully connected layer and the output of the third fully connected layer. The output of the skip connection layer is used as the input of the fourth fully connected layer. The number of model training features output by the first fully connected layer and the third fully connected layer is the same.

8. The air quality forecasting method based on integrated deep learning according to claim 7, characterized in that: The output of the skip connection layer is: in Skip Indicates the connection result. Feature 1 represents the output of the first fully connected layer, Feature 3 represents the output of the third fully connected layer.

Citation Information

Patent Citations

  • Graph neural network prediction method for power regional load

    CN113822467A

  • Multi-model fusion pollutant emission distribution rapid correction method and system

    CN119207615A