Ozone pollution emission reduction strategy rapid identification method based on ConvLSTM

By using the ConvLSTM model, combining the advantages of CNN and LSTM, the dependence between time series and spatial sequence is captured, and the problem of long operation time of traditional models is solved, and the optimal precursor emission reduction strategy is achieved when quickly identifying ozone pollution, meeting the needs of pollution emergency prevention and control.

CN120015174APending Publication Date: 2025-05-16SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411905280.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The traditional atmospheric chemical transmission model has a long calculation time, making it difficult to quickly identify the precursor emission reduction strategies during ozone pollution, and cannot meet the needs of emergency pollution prevention and control.

Method used

The ConvLSTM model is adopted, combining the advantages of CNN and LSTM, and the dependence between time series and spatial sequences is captured, and the temporal changes and spatial distribution response of ozone concentration under different precursor emission reduction strategies are quickly identified.

Benefits of technology

The optimal precursor emission reduction strategy for quickly identifying ozone pollution has been achieved, the timeliness requirements for emergency prevention and control of ozone pollution have been met, and the public health has been guaranteed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015174A_ABST
    Figure CN120015174A_ABST
Patent Text Reader

Abstract

The invention discloses an ozone pollution emission reduction strategy rapid identification method based on ConvLSTM. The method comprises the following steps: constructing different emission scenes, and obtaining basic data under different emission scenes in a set time period of an area to be analyzed; constructing an original data set, performing dimension reduction processing, and dividing the original data set after dimension reduction processing into a training set and a test set; constructing a ConvLSTM model, and training the ConvLSTM model by using the training set; and performing ozone concentration simulation on the test set by using the trained ConvLSTM model to obtain the ozone concentration of an area needing to be analyzed under different emission scenes, and further analyzing to obtain a corresponding optimal emission reduction strategy. According to the method, the time change and spatial distribution response conditions of the ozone concentration under different precursor emission reduction strategies can be quickly identified, so that the optimal precursor emission reduction strategy in different areas at different time periods can be disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of atmospheric environmental science and pollution control, and more specifically, to a method for quickly identifying ozone pollution reduction strategies based on ConvLSTM. Background Art

[0002] Near-surface O3 is mainly composed of volatile organic compounds (VOCs) and nitrogen oxides (NO x ) and other precursors undergo complex photochemical reactions. The core issue of O3 pollution control is the reduction of precursor emissions. However, there is a nonlinear relationship between O3 and its precursor emissions, and unreasonable precursor emission reduction may cause an increase in ozone concentrations. Traditionally, atmospheric chemical transport models are generally used to explore feasible precursor emission reduction strategies. Atmospheric chemical transport models can set up multiple emission scenarios by adjusting the emissions of different precursors in the emission inventory, simulate the changes in ozone pollution under different scenarios, and thus determine the response relationship of ozone to different precursor emission reduction strategies. However, the operation time of the atmospheric chemical transport model is long, often taking hours or even days, which makes it difficult to meet the urgent need to quickly identify precursor emission reduction strategies when ozone pollution occurs or is about to occur, in order to carry out emergency pollution prevention and control.

[0003] Machine learning has been increasingly widely used in environmental science research due to its fast computing speed and good recognition of nonlinear data. In terms of ozone pollution, the application of machine learning is mainly focused on improving the accuracy of pollution predictions. For the key scientific and technological issues of precursor emission reduction strategy identification and effectiveness evaluation, which are closely related to ozone pollution prevention and control, the application of machine learning technology is still blank. The main reason is that traditional machine learning techniques, including decision trees, random forests and other models, cannot fully utilize the spatiotemporal variation characteristics of ozone and precursor concentrations and emission data, and it is difficult to accurately evaluate the impact of changes in precursor emissions on ozone concentrations.

[0004] In the prior art, the DeepRSM model is established by using the convolutional neural network algorithm in deep learning (Xing J., Zheng S., Ding D., et al. Deep Learning for Prediction of the Air Quality Response to Emission Changes [J]. Environmental Science & Technology, 2020, 54 (14): 8589-8600.). DeepRSM needs to use two air quality model scenario simulation results as input, and then quickly simulate the response relationship between future ozone and precursor emission changes, which can provide support for air quality management decisions. However, DeepRSM still needs to run the air quality model, which takes a certain amount of time. At the same time, it is difficult for convolutional neural networks to identify the time dependence of ozone concentration, and it is impossible to learn the past ozone concentration and its possible impact on the future. Summary of the invention

[0005] The purpose of this invention is to quickly identify the temporal changes and spatial distribution responses of ozone concentration under different precursor emission reduction strategies, so as to reveal the optimal precursor emission reduction strategy for different regions at different time periods, so as to more accurately serve the emergency prevention and control of ozone pollution and protect public health.

[0006] ConvLSTM (Convolutional Long Short-Term Memory) is a powerful framework in deep learning that combines the advantages of CNN (Convolutional Neural Network) and LSTM (Long Short-Term Memory Network). The ConvLSTM network expands the input data dimension of the traditional LSTM by introducing convolution operations in the LSTM unit. This extension enables the ConvLSTM network to capture the dependencies of both time series and spatial series. Specifically, in ConvLSTM, the input data is not just a one-dimensional vector, but a two-dimensional (or higher-dimensional) spatial image (such as video frames, meteorological fields, pollutant distribution fields, etc.). LSTM is a variant of recurrent neural network specifically designed to process time series data. It captures long-term dependencies in sequence data through a gating mechanism and can effectively avoid the problem of gradient vanishing or gradient exploding. LSTM contains input gates, forget gates, output gates, and cell states. These gating units allow the network to selectively remember or forget input information and generate output. ConvLSTM introduces convolution operations in the input, forget, and output gates of the LSTM unit, enabling it to operate on the spatial dimensions of the input data. This means that the ConvLSTM network can extract the spatial features of the input data like CNN, and capture the temporal information of the data like LSTM.

[0007] Therefore, ConvLSTM can give full play to the powerful capabilities of CNN (convolutional neural network) in processing spatial data and LSTM in processing long time series data. This will help to quickly identify the temporal changes and spatial distribution responses of ozone concentration under different precursor emission reduction strategies, so as to reveal the optimal precursor emission reduction strategies in different regions at different times, so as to more accurately serve the emergency prevention and control of ozone pollution and protect public health.

[0008] The purpose of the present invention is achieved by at least one of the following technical solutions.

[0009] The ConvLSTM-based rapid identification method for ozone pollution reduction strategies includes the following steps:

[0010] S1. Construct different emission scenarios and use the WRF / SMOKE / CMAQ model system to obtain basic data under different emission scenarios within the set time period in the required analysis area;

[0011] S2, constructing an original data set based on the basic data obtained in step S1 and performing dimensionality reduction processing, and dividing the original data set after dimensionality reduction processing into a training set and a test set;

[0012] S3. Construct a ConvLSTM model, train the ConvLSTM model using the training set, and evaluate the prediction effect of the trained ConvLSTM model using the test set.

[0013] S4. Use the trained ConvLSTM model to simulate the ozone concentration of the test set to obtain the ozone concentration in the required analysis area under different emission scenarios, and then analyze and obtain the corresponding optimal emission reduction strategy.

[0014] Furthermore, in step S1, the initial data of anthropogenic emissions in the area to be analyzed is obtained. The initial data of anthropogenic emissions in the area to be analyzed is based on the existing MEIC emission inventory. The MEIC emission inventory contains nine types of pollutants, namely VOCs, NO x , carbon monoxide, sulfur dioxide, coarse particulate matter, fine particulate matter, ammonia, black carbon, and organic carbon;

[0015] According to VOCs (Volatile Organic Compounds), NO x The emission reduction factor is used to construct the emission scenario, in which VOCs and NO x The emission reduction coefficients are all 0, that is, no control is done, and the baseline emission scenario is constructed; 30% is used as the VOCs and NO x The upper limit of emission reduction ratio is 10%, and 15 other emission reduction scenarios are constructed as the difference ratio between different emission reduction scenarios, with a total of 16 emission scenarios.

[0016] Furthermore, in the WRF / SMOKE / CMAQ model system, the Weather Research and Forecasting Model WRF (Weather Research and Forecasting Model) is first used to simulate the meteorological field of the required analysis area to obtain meteorological parameters; then, the Sparse Matrix Operator Kernel Emissions Model SMOKE (Sparse Matrix Operator Kernel Emissions) is used to perform spatiotemporal and species distribution processing on the initial data of anthropogenic emissions, and the initial data of anthropogenic emissions are processed into 1-hour time resolution data; the 1-hour time resolution emission data is divided into 1-hour time resolution emission data under multiple emission scenarios according to different emission scenarios; the 1-hour time resolution data under different emission scenarios and the meteorological parameters simulated by the Weather Research and Forecasting Model WRF are input into the Community Multiscale Air Quality Model CMAQ (Community Multiscale Air Quality Model), and the Community Multiscale Air Quality Model CMAQ simulates atmospheric chemical reactions based on the input data to obtain ozone concentrations under different emission scenarios;

[0017] The basic data under different emission scenarios include meteorological parameters, ozone concentrations under different emission scenarios, and 1-hour time resolution data under different emission scenarios.

[0018] Furthermore, in the WRF / SMOKE / CMAQ model system, the Weather Research and Forecasting Model WRF is used to provide meteorological field simulation, and its initial meteorological field uses the existing global meteorological reanalysis data with an accuracy of 1°×1°, and the underlying surface data uses the existing MODIS data;

[0019] The meteorological parameters output by the Weather Research and Forecasting Model WRF include temperature, solar radiation, 10-meter meridional wind speed, 10-meter zonal wind speed, rainfall and boundary layer height.

[0020] Furthermore, in the WRF / SMOKE / CMAQ model system, the sparse matrix operator kernel emission model SMOKE is used to process the initial data of anthropogenic emissions into 1-hour time resolution data using the allocation factors of the allocation database using the localization factors of the required analysis area.

[0021] Furthermore, in the WRF / SMOKE / CMAQ model system, the community multiscale air quality model CMAQ is used to simulate atmospheric chemical reactions, and its gas phase reaction mechanism and aerosol reaction mechanism adopt carbon bond mechanism and AE5 respectively;

[0022] The community multi-scale air quality model CMAQ includes the JPROC module for calculating the clear sky photolysis rate, the ICON module for providing the initial concentration field of pollutants to the simulation area, the BCON module for providing the boundary field concentration, and the CCTM module for photochemical transport;

[0023] The JPROC module uses the clear sky photolysis rate data released simultaneously with the community multiscale air quality model CMAQ; during the ICON module simulation process, the PROFILE initial concentration field provided by the community multiscale air quality model CMAQ is used on the first day of the simulation, and the simulation results of the previous day are used as the initial concentration field for subsequent days; in the BCON module, the three-layer grid area is simulated in a nested manner, and the grid resolutions of the first layer D1, the second layer D2, and the third layer D3 decrease successively, and the areas included in the third layer D3, the second layer D2, and the first layer D1 are nested from the inside to the outside, and the areas reported in the third layer D3 are the required analysis areas. The boundary field generated by PROFILE is used to simulate the first layer D1, and the simulation results of the second layer D2 and the third layer D3 boundary field data are used to simulate the previous layer respectively.

[0024] Further, in step S2, principal component analysis (PCA) is used to reduce the dimension of the original data set;

[0025] Among the basic data under different emission scenarios within a set time period, the basic data under some emission scenarios in some time periods constitute the training set, and the basic data under all emission scenarios in the remaining time periods constitute the test set.

[0026] Further, in step S3, the ConvLSTM model includes an encoding network and a prediction network;

[0027] The encoding network includes a plurality of ConvLSTM layers connected in sequence, each ConvLSTM layer is respectively provided with a plurality of hidden states for extracting key temporal variation features and key spatial variation features of the input data;

[0028] In the prediction network, the hidden state of each ConvLSTM layer is copied, and the hidden state of each ConvLSTM layer is input into a convolutional layer, and the predicted ozone concentration under various emission reduction strategies is output.

[0029] Furthermore, in step S3, the hyperparameters of the ConvLSTM model are set as follows:

[0030] The time step is 6, the number of ConvLSTM layers is 3, the number of hidden states in the ConvLSTM layer is [32, 16, 8], the convolution kernel size of each ConvLSTM is 3, the optimizer is RMSProp, and the learning rate is 0.001.

[0031] Furthermore, step S4 includes the following steps:

[0032] S4.1. Input the data in the test set into the trained ConvLSTM model to simulate the ozone concentration under 16 emission scenarios, and then obtain the ozone concentration of different grids in the required analysis area under different emission scenarios;

[0033] S4.2. Identify the emission scenario that can achieve the lowest ozone concentration in each grid. This emission scenario is the optimal emission reduction strategy for the grid.

[0034] S4.3. After completing the calculation of all grids, the optimal emission reduction strategy distribution map of the entire area to be analyzed is obtained to serve the refined prevention and control of ozone pollution.

[0035] Compared with the prior art, the advantages of the present invention are:

[0036] The present invention uses the ConvLSTM model to simulate ozone concentration and identify emission reduction strategies. ConvLSTM can give full play to the powerful capabilities of CNN (convolutional neural network) in processing spatial data and LSTM in processing long time series data. This will help to quickly identify the temporal changes and spatial distribution responses of ozone concentration under different precursor emission reduction strategies, so as to reveal the optimal precursor emission reduction strategies in different regions at different time periods, thereby more accurately serving the emergency prevention and control of ozone pollution and protecting public health. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flowchart of the steps of a method for rapidly identifying ozone pollution reduction strategies based on ConvLSTM in an embodiment of the present invention.

[0038] Figure 2 Schematic diagram of the structure of the ConvLSTM model in an embodiment of the present invention.

[0039] Figure 3 This is a spatial distribution map of ozone concentration in the Pearl River Delta from September 22 to September 30, 2022 in an embodiment of the present invention.

[0040] Figure 4 This is a spatial distribution map of the optimal emission reduction strategy in the Pearl River Delta from September 22 to September 30, 2022 in an embodiment of the present invention. DETAILED DESCRIPTION

[0041] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the specific implementation of the present invention is described in detail below with reference to the accompanying drawings and examples.

[0042] Example:

[0043] A fast identification method for ozone pollution reduction strategies based on ConvLSTM, such as Figure 1 As shown, the following steps are included:

[0044] S1. In one embodiment, the Pearl River Delta of my country is taken as the research area, and the 30 days and 720 hours in September 2022 when ozone pollution is frequent are taken as the research period. Different emission scenarios are constructed, and the basic data under different emission scenarios in the Pearl River Delta in September 2022 for a total of 30 days and 720 hours are obtained using the WRF / SMOKE / CMAQ model system;

[0045] Obtain the initial data of anthropogenic emissions in the Pearl River Delta. In one embodiment, the initial data of anthropogenic emissions in the Pearl River Delta adopts the MEIC emission inventory with the base year of 2020. The MEIC emission inventory contains nine categories of pollutants, namely VOCs (Volatile Organic Compounds), NO x , carbon monoxide, sulfur dioxide, coarse particulate matter, fine particulate matter, ammonia, black carbon, and organic carbon;

[0046] According to VOCs, NO x The emission reduction factor is used to construct the emission scenario, in which VOCs and NO x The emission reduction coefficients are all 0, that is, no control is done, and the baseline emission scenario is constructed; 30% is used as the VOCs and NO x The upper limit of emission reduction ratio is 10%, and 15 other emission reduction scenarios are constructed as the difference ratio between different emission reduction scenarios. There are 16 emission scenarios in total, as shown in Table 1, where No. 1 is the basic emission scenario, and No. 2 to 16 are emission reduction scenarios.

[0047] Table 1 VOCs and NO in 16 emission scenarios x Emission reduction coefficient setting table

[0048]

[0049] The WRF / SMOKE / CMAQ model system uses the Lambert conformal projection centered at 28.5°N and 114°E, with the two true projection latitudes at 15°N and 40°N being the basic projection coordinates.

[0050] In the WRF / SMOKE / CMAQ model system, the Weather Research and Forecasting Model WRF (Weather Research and Forecasting Model) is first used to simulate the meteorological field of the required analysis area to obtain meteorological parameters; then, the initial data of anthropogenic emissions are obtained, and the Sparse Matrix Operator Kernel Emissions (Sparse Matrix Operator Kernel Emissions) model is used to perform spatiotemporal and species distribution processing on the initial data of anthropogenic emissions, and the allocation factors in the localization factor allocation database of the Pearl River Delta are used to process the initial data of anthropogenic emissions into 1-hour time resolution data; the 1-hour time resolution emission data is divided into 1-hour time resolution emission data under multiple emission scenarios according to different emission scenarios; the 1-hour time resolution data under different emission scenarios and the meteorological parameters simulated by the Weather Research and Forecasting Model WRF are input into the Community Multiscale Air Quality Model CMAQ (Community Multiscale Air Quality Model), and the Community Multiscale Air Quality Model CMAQ simulates atmospheric chemical reactions based on the input data to obtain ozone concentrations under different emission scenarios;

[0051] The basic data under different emission scenarios include meteorological parameters, ozone concentrations under different emission scenarios, and 1-hour time resolution data under different emission scenarios.

[0052] In one embodiment, the weather research and forecasting model WRF uses WRF3.9 provided by the National Center for Atmospheric Research of the United States, and its initial meteorological field uses global meteorological reanalysis data with an accuracy of 1°×1° provided by the National Center for Environmental Prediction of the United States, and the underlying surface data uses MODIS data provided by NASA;

[0053] The meteorological parameters output by the Weather Research and Forecasting Model WRF include temperature, solar radiation, 10-meter meridional wind speed, 10-meter zonal wind speed, rainfall and boundary layer height.

[0054] In the WRF / SMOKE / CMAQ model system, the sparse matrix operator kernel emission model SMOKE is used to process the initial data of anthropogenic emissions in terms of time, space and species. The allocation factors of the localization factor allocation database of the Pearl River Delta are used to process the initial data of anthropogenic emissions into 1-hour time resolution data.

[0055] In one embodiment, the community multi-scale air quality model CMAQ adopts the CMAQ 5.0.2 version provided by the U.S. Environmental Protection Agency, and its gas phase reaction mechanism and aerosol reaction mechanism adopt the carbon bond mechanism and AE5 respectively;

[0056] The community multi-scale air quality model CMAQ includes the JPROC module for calculating the clear sky photolysis rate, the ICON module for providing the initial concentration field of pollutants to the simulation area, the BCON module for providing the boundary field concentration, and the CCTM module for photochemical transport;

[0057] The JPROC module uses the clear sky photolysis rate data released simultaneously with the community multiscale air quality model CMAQ; during the ICON module simulation process, the PROFILE initial concentration field provided by the community multiscale air quality model CMAQ is used on the first day of the simulation, and the simulation results of the previous day are used as the initial concentration field for subsequent days; in the BCON module, the three-layer grid area is simulated in a nested manner, and the grid resolutions of the first layer D1, the second layer D2, and the third layer D3 decrease successively, and the areas included in the third layer D3, the second layer D2, and the first layer D1 are nested from the inside to the outside, and the areas reported in the third layer D3 are the required analysis areas. The boundary field generated by PROFILE is used to simulate the first layer D1, and the simulation results of the second layer D2 and the third layer D3 boundary field data are used to simulate the previous layer respectively.

[0058] In one embodiment, the grid resolution of the first layer D1 is 27km, including Southeast Asia, East Asia and the northwest of the Pacific Ocean, the grid resolution of the second layer D2 is 9km, including the entire Guangdong region, and the third layer D3 is the Pearl River Delta region to be analyzed, including 9 administrative cities in Guangdong Province, with a grid resolution of 3km, and a total of 110×152 grids. In order to avoid the impact of meteorological boundary fields on the simulation of the community multi-scale air quality model CMAQ, the simulation area of ​​the Weather Research and Forecasting Model WRF is slightly larger than that of the community multi-scale air quality model CMAQ. In order to improve the accuracy and reliability of the simulation results, the outer layer simulation results are used as the initial field conditions for the inner layer simulation.

[0059] S2, constructing an original data set based on the basic data obtained in step S1 and performing dimensionality reduction processing, and dividing the original data set after dimensionality reduction processing into a training set and a test set;

[0060] In one embodiment, a total of 57 characteristic variables, including meteorological parameters input by the WRF / SMOKE / CMAQ model system, ozone concentrations under different emission scenarios, and emission data under different scenarios, form an original data set.

[0061] PCA is a commonly used unsupervised learning algorithm with the advantages of being simple, effective, and fast in dimensionality reduction. Its basic idea is to reduce the dimension of data by projecting the data into the direction with the largest variance. PCA maps the original data into a new coordinate system through linear transformation, maximizing the variance of the new data. In the new coordinate system, each dimension is a linear combination of the original features, called principal components. These principal components are arranged according to the size of the variance, and the first few principal components usually contain the most important information in the data. By maximizing the variance of the new data, PCA retains the most important information in the original data set and removes unnecessary redundant information. This can reduce the dimension of the original data and speed up the model training process without affecting the performance of the model.

[0062] PCA was used to reduce the 57-dimensional features contained in the original dataset to 2, 4, 8, 16, and 32 dimensions for experiments. The hyperparameters of ConvLSTM were uniformly set to 6 hours of time step, 2 ConvLSTM layers, 32 hidden layers per ConvLSTM layer, 3x3 convolution kernel size, Adam optimizer, and 0.001 learning rate. After 300 rounds of training, the coefficient of determination (R 2 ) as shown in Table 2. As the dimension increases, R 2 It keeps rising from 0.32 in 2D to 0.91 in 16D. After that, increasing the dimension does not improve the model R 2 The more dimensions there are, the more computing resources are required to train the model. In one embodiment, principal component analysis (PCA) is used to reduce the dimensionality of the original data set, specifically to 16 dimensions.

[0063] Table 2 Model performance of different PCA dimensionality reduction numbers

[0064]

[0065]

[0066] Among the basic data under different emission scenarios within a set time period, the basic data under some emission scenarios in some time periods constitute the training set, and the basic data under all emission scenarios in the remaining time periods constitute the test set.

[0067] In one embodiment, the data from September 1, 2022 to September 21, 2022 under the emission scenarios numbered 1, 2, 4, 5, 7, 8, 9, 10, 12, 13, 15, and 16 in Table 1 are used as training sets, and the data from September 22, 2022 to September 30, 2022 under all emission scenarios are used as test sets. The data under the emission scenarios numbered 3, 6, 11, and 14 are not included in the training set. The simulation results of the ConvLSTM model on them can verify the extrapolation ability of the ConvLSTM model.

[0068] S3. Construct a ConvLSTM model, train the ConvLSTM model using the training set, and evaluate the prediction effect of the trained ConvLSTM model using the test set.

[0069] The ConvLSTM model, such as Figure 2 As shown, it includes an encoding network and a prediction network;

[0070] The encoding network includes a plurality of ConvLSTM layers connected in sequence (Shi, Xingjian, Zhourong Chen, HaoWang, DYYeung, Wai-Kin Wong, and Wang-chun Woo. "Convolutional LSTM network: a machine learning approach for precipitation nowcasting." Neural Information Processing Systems (2015).), each ConvLSTM layer is respectively set with a plurality of hidden states for extracting key temporal variation features and key spatial variation features of the input data;

[0071] In the prediction network, the hidden state of each ConvLSTM layer is copied, and the hidden state of each ConvLSTM layer is input into a convolutional layer, and the predicted ozone concentration under various emission reduction strategies is output.

[0072] The hyperparameters that need to be set for the ConvLSTM model include the time step, the number of ConvLSTM layers, the number of hidden states of each ConvLSTM layer, the convolution kernel size of each ConvLSTM layer, the optimizer, and the learning rate. Among the hyperparameters of the ConvLSTM module, the time step size is determined first. The time step is a very important hyperparameter that directly determines whether the ConvLSTM module can learn the temporal characteristics of the response to changes in ozone and precursors. This paper sets time steps of 4, 6, 12, 24, and 48 for experiments. When determining the time step, other hyperparameters are uniformly set to 3 ConvLSTM layers, the number of hidden states of each ConvLSTM layer is 32, the convolution kernel size is 3x3, the Adam optimizer, and the learning rate is 0.001. After 300 rounds of training, the model R at each time step is 2 The index results are shown in Table 3.2 The highest value reached was 0.91, set to R 48 2 The lowest. It can be seen that the more time steps are set, the better. Because the ConvLSTM model shares the same weight at each time step, too many time steps will cause the ConvLSTM model to consider a lot of useless information when updating the weight, which will interfere with the ConvLSTM model learning the time characteristics of the response to the changes in ozone and precursors. Therefore, a time step of 6 is used for subsequent analysis.

[0073] Table 3 ConvLSTM model performance at different time steps

[0074]

[0075]

[0076] The next hyperparameters to be determined are the number of ConvLSTM layers, the number of hidden states in each ConvLSTM layer, and the size of the convolution kernel in each ConvLSTM layer. These three hyperparameters determine the network structure and network complexity of the ConvLSTM model. Generally speaking, the more complex the structure of the ConvLSTM model, the more it can capture complex patterns and features in the data. However, when it is too complex, it may overfit the noise of the training data, resulting in poor performance on new data. Therefore, it is necessary to select a suitable network structure based on the actual situation. Considering the computing resources, the value ranges of these three hyperparameters are set for grid search. The value range of the number of ConvLSTM layers is [1,2,3,4], the value range of the number of hidden layers in each ConvLSTM layer is [8,16,32,64], and the value range of the convolution kernel size in each ConvLSTM layer is [3,5,7].

[0077] Table 4 lists the model R with different hyperparameters after 300 rounds of training. 2 The results show that the structure of the ConvLSTM model is not as complex as possible. When the number of ConvLSTM layers is only one, the structure is too simple to accurately learn the response relationship between ozone and precursor changes. 2 Only 0.82, with the increase of ConvLSTM layers, R 2 Gradually increases, when the number of ConvLSTM layers is 3, R 2 The setting of the number of hidden layers will also affect the performance of the model. When the number of hidden states of each ConvLSTM layer is the same, the number of hidden states is set to 32, R 2 The maximum value has been reached. Continuing to increase the number of hidden states will not improve the model R. 2When the number of ConvLSTM layers is large, you can try various ways to set the number of hidden states, such as increasing the number of hidden states in each layer, decreasing the number of hidden states in each layer, or increasing the number of hidden states in each layer first and then decreasing it. ConvLSTM model R with the number of hidden states of [32,32,32] and [32,16,8] 2 The maximum value reached was 0.91. The size of the convolution kernel determines the spatial range of the ConvLSTM model to learn the response relationship between ozone and precursor changes. The larger the convolution kernel, the larger the spatial range considered and the more complex the spatial features learned by the model. ConvLSTM model R 2 As the convolution kernel increases, it decreases, because the impact of precursor emissions on ozone will weaken as the spatial distance between the two increases. If the convolution kernel is too large, the model will easily learn incorrect spatial features, reducing the prediction accuracy of the model. Considering the model performance of each hyperparameter and the required computing resources, the number of ConvLSTM layers is selected as 3, the number of hidden states of the ConvLSTM layer is [32, 16, 8], and the convolution kernel size of each ConvLSTM layer is 3.

[0078] Table 4 Performance of ConvLSTM models with different numbers of layers, hidden layers, and convolution kernel sizes

[0079]

[0080]

[0081] The last hyperparameters to be determined are the optimizer and learning rate. These two hyperparameters mainly affect the training speed, convergence and stability of the model performance. Only by setting the appropriate optimizer and learning rate can the model find the global optimal solution. Common optimizers for ConvLSTM models are SGD, RMSProp and Adam. Both RMSProp and Adam are adaptive learning rate optimization algorithms. They can effectively adjust the learning rate during the training process to adapt to the changes of different parameters. The difference is that RMSProp is relatively conservative in adjusting the learning rate, and RMSProp is usually more stable than Adam. Although Adam and RMSProp are generally considered to be better optimizers in deep learning, SGD is still an effective optimizer in some cases. A learning rate that is too large or too small is not conducive to model building. A learning rate that is too large may cause the parameters to fluctuate greatly during the training process, resulting in gradient explosion or gradient disappearance problems, making the training unstable. A learning rate that is too small may cause the model to oscillate near the local optimal solution, and the model has poor generalization ability on the test data. Therefore, experiments were conducted using SGD, RMSProp, and Adam with learning rates set to 0.01, 0.001, and 0.0001, respectively. SGD was the worst of the three optimizers, and the training error (loss) of SGD was significantly greater than that of RMSProp and Adam. For both RMSProp and Adam, the loss was the smallest when the learning rate was set to 0.001. After more than 200 training rounds, the loss of RMSProp and Adam was almost the same, but RMSProp was more stable than Adam. Therefore, RMSProp was selected as the optimizer with a learning rate set to 0.001.

[0082] In one embodiment, the hyperparameters of the ConvLSTM model are set as follows:

[0083] The time step is 6, the number of ConvLSTM layers is 3, the number of hidden states in the ConvLSTM layer is [32, 16, 8], the convolution kernel size of each ConvLSTM is 3, the optimizer is RMSProp, and the learning rate is 0.001.

[0084] The test set was used to test the performance of the ConvLSTM model from two aspects: first, to verify whether the ConvLSTM model can accurately fit the response relationship of ozone concentration to changes in precursor emissions in the future (time outside the training set); second, to verify the ConvLSTM model's ability to fit external emission reduction scenarios (emission reduction scenarios outside the training set).

[0085] S4. Use the trained ConvLSTM model to simulate ozone concentration and obtain the ozone concentration in the analysis area under different emission scenarios, and then analyze and obtain the corresponding optimal emission reduction strategy, including the following steps:

[0086] S4.1. Input the data in the test set into the trained ConvLSTM model to identify the optimal emission reduction strategies for different regions in the Pearl River Delta at different times in September 2022, simulate the ozone concentration under 16 emission scenarios, and then obtain the ozone concentration of 110×152 grids in the entire Pearl River Delta under different emission scenarios;

[0087] S4.2. Identify the emission scenario that can achieve the lowest ozone concentration in each grid. This emission scenario is the optimal emission reduction strategy for the grid.

[0088] S4.3. After completing the calculation of all grids, the optimal emission reduction strategy distribution map for the entire Pearl River Delta is obtained to serve the refined prevention and control of ozone pollution.

[0089] In one embodiment, ozone concentration simulations under different emission scenarios are performed as follows:

[0090] As a fast response model, the ConvLSTM model only takes 270 seconds (the computing device is an ordinary personal computer equipped with RTX 3060 12GB GPU, Intel i5-13400F CPU and 16GB running memory) to calculate the ozone concentration of the test set under 16 emission scenarios from September 22 to September 30. It takes at least 48 hours to run the traditional atmospheric chemical transport model. It can be seen that the ConvLSTM model can quickly simulate the response relationship between ozone and precursor emission changes, meeting the timeliness requirements of ozone emergency prevention and control.

[0091] Using R 2 The overall accuracy of the ConvLSTM model in simulating ozone concentration under different emission reduction scenarios is evaluated by using the mean absolute error (MAE). The performance of the ConvLSTM model in the test set is shown in Table 5. The ConvLSTM model has excellent simulation effect on each emission reduction scenario. Its R 2 The indexes were all greater than or equal to 0.91, and the MAE indexes were all less than 10 μg / m 3 , the ratio of MAE to the average ozone value is less than 10%, that is, the model simulation accuracy is greater than 90%. It is worth noting that the R of the ConvLSTM model simulating external emission reduction scenarios (emission reduction scenarios not included in the training set) is 2 All indicators are greater than or equal to 0.90, and are not significantly lower than the effects when simulating internal emission reduction scenarios, proving that the ConvLSTM model can accurately simulate ozone concentrations under different precursor emission levels and has excellent generalization capabilities and practical application value.

[0092] Table 5 R of ConvLSTM model simulating 16 scenarios 2 and MAE table

[0093]

[0094] In one embodiment, the optimal emission reduction strategy for the region is output as follows:

[0095] The distribution of the maximum daily 8-hour average ozone concentration (MDA8) in the Pearl River Delta from September 22 to September 30, 2022 is as follows: Figure 3 As shown in the figure, it can be found that the spatial distribution of ozone concentration in the Pearl River Delta was different every day during this period. The spatial distribution of ozone in the Pearl River Delta on the 22nd and 23rd was relatively similar. The ozone concentration in various regions of the Pearl River Delta was 160μg / m 3 From the 24th to the 26th, the Pearl River Delta region experienced a serious ozone pollution. The high ozone concentration was mainly distributed in Foshan and its surrounding areas, exceeding 200μg / m 3 , while the ozone MDA8 in other areas was close. Starting from the 27th, the high ozone areas in the Pearl River Delta were mainly distributed in the northwest, and the ozone concentration in the southeast began to meet the standard. By the 30th, only a small amount of high ozone was distributed at the junction of Zhaoqing City and Qingyuan City.

[0096] Then, the ConvLSTM model is used to output the optimal emission reduction strategy for each grid in the Pearl River Delta. Figure 4 As shown, it can be seen that on different dates, the optimal emission reduction strategy in most areas is 16. The second is strategy 13, and then strategy 4. The distribution of other strategies is relatively small (the emission reduction ratios corresponding to each strategy are shown in Table 1). At the same time, the spatial distribution of emission reduction strategies changes every day. If the same strategy is used every day, it is difficult to achieve the maximum ozone concentration reduction effect, which shows the necessity of implementing precise emission reduction. It can be seen that on the 22nd and 23rd, strategy 13 was distributed on the west side of Foshan and at the junction of Dongguan and Huizhou. Strategy 4 is more scattered in the outer northern part of the Pearl River Delta. From the 24th to the 26th, as mentioned above, a more serious ozone pollution occurred in the Pearl River Delta. Near the heavily polluted areas, southeastern Foshan, Zhongshan and northeastern Jiangmen, the optimal emission reduction strategy is 13. From the 27th to the 30th, the spatial position of strategy 13 on different dates will change. From the 27th, strategy 13 will be distributed in Zhuhai, southern Guangzhou, and northern and southwestern Foshan. Strategy 4 is distributed in the peripheral areas of the Pearl River Delta.

[0097] In general, the ConvLSTM model constructed by the present invention can provide the optimal emission reduction strategy for each grid in the region during ozone pollution, which is an effective method for achieving precise prevention and control.

[0098] The preferred embodiments of the present application disclosed above are only used to help understand the present invention and its core ideas. For those skilled in the art, according to the ideas of the present invention, there will be changes in specific application scenarios and implementation operations, and this specification should not be interpreted as limiting the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A fast identification method for ozone pollution reduction strategies based on ConvLSTM, characterized in that: The steps include: S1. Construct different emission scenarios and use the WRF / SMOKE / CMAQ model system to obtain basic data under different emission scenarios within the set time period in the required analysis area; S2, constructing an original data set based on the basic data obtained in step S1 and performing dimensionality reduction processing, and dividing the original data set after dimensionality reduction processing into a training set and a test set; S3. Construct a ConvLSTM model, train the ConvLSTM model using the training set, and evaluate the prediction effect of the trained ConvLSTM model using the test set. S4. Use the trained ConvLSTM model to simulate the ozone concentration of the test set to obtain the ozone concentration in the required analysis area under different emission scenarios, and then analyze and obtain the corresponding optimal emission reduction strategy.

2. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 1 is characterized in that: In step S1, the initial data of anthropogenic emissions in the area to be analyzed is obtained. The initial data of anthropogenic emissions in the area to be analyzed uses the existing MEIC emission inventory. The MEIC emission inventory contains nine major types of pollutants, namely VOCs, NO x , carbon monoxide, sulfur dioxide, coarse particulate matter, fine particulate matter, ammonia, black carbon, and organic carbon; According to VOCs, NO x The emission reduction factor is used to construct the emission scenario, in which VOCs and NO x The emission reduction coefficients are all 0, that is, no control is done, and the baseline emission scenario is constructed; 30% is used as the VOCs and NO x The upper limit of emission reduction ratio is 10%, and the remaining 15 emission reduction scenarios are constructed as the difference ratio between different emission reduction scenarios, with a total of 16 emission scenarios.

3. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 1, characterized in that: In the WRF / SMOKE / CMAQ model system, the Weather Research and Forecasting Model WRF is first used to simulate the meteorological field of the required analysis area to obtain meteorological parameters; then, the sparse matrix operator core emission model SMOKE is used to perform spatiotemporal and species distribution processing on the initial data of anthropogenic emissions, and the initial data of anthropogenic emissions are processed into 1-hour time resolution data; the 1-hour time resolution emission data is divided into 1-hour time resolution emission data under multiple emission scenarios according to different emission scenarios; the 1-hour time resolution data under different emission scenarios and the meteorological parameters simulated by the Weather Research and Forecasting Model WRF are input into the Community Multiscale Air Quality Model CMAQ, and the Community Multiscale Air Quality Model CMAQ simulates atmospheric chemical reactions based on the input data to obtain ozone concentrations under different emission scenarios; The basic data under different emission scenarios include meteorological parameters, ozone concentrations under different emission scenarios, and 1-hour time resolution data under different emission scenarios.

4. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 3 is characterized in that: In the WRF / SMOKE / CMAQ model system, the Weather Research and Forecasting Model WRF is used to provide meteorological field simulation. The initial meteorological field uses the existing global meteorological reanalysis data with an accuracy of 1°×1°, and the underlying surface data uses the existing MODIS data. The meteorological parameters output by the Weather Research and Forecasting Model WRF include temperature, solar radiation, 10-meter meridional wind speed, 10-meter zonal wind speed, rainfall and boundary layer height.

5. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 3 is characterized in that: In the WRF / SMOKE / CMAQ model system, the sparse matrix operator kernel emission model SMOKE is used to process the initial data of anthropogenic emissions in terms of time, space and species distribution. The allocation factors of the localization factor allocation database for the required analysis area are used to process the initial data of anthropogenic emissions into 1-hour time resolution data.

6. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 3 is characterized in that: In the WRF / SMOKE / CMAQ model system, the community multiscale air quality model CMAQ is used to simulate atmospheric chemical reactions, and its gas phase reaction mechanism and aerosol reaction mechanism adopt carbon bond mechanism and AE5 respectively; The community multi-scale air quality model CMAQ includes the JPROC module for calculating the clear sky photolysis rate, the ICON module for providing the initial concentration field of pollutants to the simulation area, the BCON module for providing the boundary field concentration, and the CCTM module for photochemical transport; The JPROC module uses the clear sky photolysis rate data released simultaneously with the community multiscale air quality model CMAQ; during the ICON module simulation process, the PROFILE initial concentration field provided by the community multiscale air quality model CMAQ is used on the first day of the simulation, and the simulation results of the previous day are used as the initial concentration field for subsequent days; in the BCON module, the three-layer grid area is simulated in a nested manner, and the grid resolutions of the first layer D1, the second layer D2, and the third layer D3 decrease successively, and the areas included in the third layer D3, the second layer D2, and the first layer D1 are nested from the inside to the outside, and the areas reported in the third layer D3 are the required analysis areas. The boundary field generated by PROFILE is used to simulate the first layer D1, and the simulation results of the second layer D2 and the third layer D3 boundary field data are used to simulate the previous layer respectively.

7. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 1, characterized in that: In step S2, principal component analysis is used to reduce the dimension of the original data set; Among the basic data under different emission scenarios within a set time period, the basic data under some emission scenarios in some time periods constitute the training set, and the basic data under all emission scenarios in the remaining time periods constitute the test set.

8. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 1, characterized in that: In step S3, the ConvLSTM model includes an encoding network and a prediction network; The encoding network includes a plurality of ConvLSTM layers connected in sequence, each ConvLSTM layer is respectively provided with a plurality of hidden states for extracting key temporal variation features and key spatial variation features of the input data; In the prediction network, the hidden state of each ConvLSTM layer is copied, and the hidden state of each ConvLSTM layer is input into a convolutional layer, and the predicted ozone concentration under various emission reduction strategies is output.

9. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 1, characterized in that: In step S3, the hyperparameter settings of the ConvLSTM model are as follows: The time step is 6, the number of ConvLSTM layers is 3, the number of hidden states in the ConvLSTM layer is [32, 16, 8], the convolution kernel size of each ConvLSTM is 3, the optimizer is RMSProp, and the learning rate is 0.

001.

10. The method for rapid identification of ozone pollution reduction strategies based on ConvLSTM according to claim 1, characterized in that: Step S4 includes the following steps: S4.

1. Input the data in the test set into the trained ConvLSTM model to simulate the ozone concentration under 16 emission scenarios, and then obtain the ozone concentration of different grids in the required analysis area under different emission scenarios; S4.

2. Identify the emission scenario that can achieve the lowest ozone concentration in each grid. This emission scenario is the optimal emission reduction strategy for the grid. S4.

3. After completing the calculation of all grids, the optimal emission reduction strategy distribution map of the entire area to be analyzed is obtained to serve the refined prevention and control of ozone pollution.

Citation Information

Patent Citations

  • Hot point network technology-based atmosphere pollutant diffusion path tracing method

    CN110673229A

  • Ozone concentration prediction method and device, electronic equipment and storage medium

    CN114861551A

  • Atmospheric multi-pollutant emission reduction optimization regulation and control method and system based on dynamic scene simulation

    CN117195585A