A method, system, device and medium for inverting near-surface nitrogen dioxide concentration

By building a two-layer machine learning model and combining multiple data sources for data correction and feature analysis, the shortcomings of ground stations and satellite remote sensing are overcome, and efficient and accurate inversion of near-ground nitrogen dioxide concentrations is achieved, which is suitable for environmental monitoring and air quality management.

CN119830014BActive Publication Date: 2025-09-26CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411924287.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-09-26
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately invert near-ground nitrogen dioxide concentrations. Ground station monitoring is costly and has limited coverage. Satellite remote sensing cannot provide direct near-ground data, and traditional machine learning models cannot effectively express complex spatiotemporal relationships.

Method used

A two-layer machine learning model was constructed using a lightweight gradient boosting model and an extreme random forest model. Combined with the European Centre for Medium-Range Weather Forecasts reanalysis data, tropospheric nitrogen dioxide column concentration data, MODIS satellite sensing data, and global population distribution data, data correction and feature correlation analysis were performed to construct a modeling dataset. The near-ground nitrogen dioxide concentration was inverted using the two-layer machine learning model.

Benefits of technology

The accuracy and reliability of near-surface nitrogen dioxide concentration inversion have been improved, and accurate GNO2 products can be obtained on different time scales, which are suitable for environmental monitoring and air quality management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830014B_ABST
    Figure CN119830014B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for inverting near-ground nitrogen dioxide concentration, comprising: obtaining a data set including European Centre for Medium-Range Weather Forecasts reanalysis data, tropospheric nitrogen dioxide column concentration data, MODIS satellite sensing data, global terrain elevation and population distribution. The tropospheric nitrogen dioxide data is corrected using European Centre for Medium-Range Weather Forecasts data to obtain near-ground concentration data. The above data set is spatially aligned with the near-ground nitrogen dioxide data to form a modeling data set. A lightweight gradient boosting machine and an extreme random forest model are used to construct a nitrogen dioxide concentration inversion model, and the model is trained using the modeling data set; the trained model is used to invert near-ground nitrogen dioxide concentration data of the test area at different time and space scales. The present invention can improve the accuracy of the inversion results and can obtain near-ground nitrogen dioxide concentration products on different time scales, and has strong applicability and authenticity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of nitrogen dioxide concentration inversion, and in particular relates to a near-ground nitrogen dioxide concentration inversion method, system, equipment and medium. Background Art

[0002] Nitrogen dioxide (NO2), a major atmospheric pollutant, is a toxic, reddish-brown, irritating gas. NO2 is chemically active and a key precursor of ozone (O3) and aerosol particulate matter. It can trigger photochemical reactions and photochemical smog, generating nitrate aerosols with strong radiative forcing. Furthermore, the nitric acid and nitrates formed by NO2 oxidation are the primary causes of nitric acid rain. Currently, with rapid economic and population development, NO2 emissions are increasing dramatically, leading to a continuous increase in atmospheric NO2 concentrations, posing serious risks to human health and the natural environment. Therefore, studying atmospheric NO2 concentrations and their changing trends is of great significance and urgent need for strengthening environmental governance and ecological environmental protection. Atmospheric NO2 primarily originates from natural and anthropogenic emissions. Natural sources include lightning oxidation, soil microbial decomposition, and biomass combustion, while anthropogenic sources include fossil fuel combustion and vehicle exhaust.

[0003] Currently, atmospheric NO2 concentration retrieval methods primarily rely on ground-based monitoring stations and satellite remote sensing. Ground-based monitoring stations utilize gas sensors or gas sampling to retrieve NO2 concentrations. A common method relies on chemical analysis, converting NO2 into measurable products such as nitrates, and then using spectrometers to measure the product concentrations. However, ground-based monitoring stations have drawbacks such as high cost, complex maintenance, fixed monitoring stations, and limited coverage. Compared to ground-based monitoring stations, satellite remote sensing, with its advantages of long-term, large-scale, real-time, dynamic, and continuous monitoring, is increasingly becoming the primary method for retrieving NO2 concentrations. Among these, the Tropospheric Ozone Meter (TROPOMI) instrument boasts the highest temporal and spatial resolution and most advanced technology since the Ozone Meter (OMI), providing superior data support. While satellite remote sensing can monitor the global distribution and changes of NO2, it can only provide total atmospheric NO2 column or tropospheric NO2 column data. While tropospheric NO2 columns are of great significance for global and regional meteorological and environmental research, surface-level NO2 concentrations (GNO2) are more directly linked to the ecological environment and human well-being. They can directly reflect the air quality conditions in cities and local areas, enabling timely detection and response to air pollution problems and the implementation of effective control measures. Therefore, GNO2 holds greater practical research value and application prospects. Attempts have been made to convert satellite-derived tropospheric NO2 column concentrations into GNO2 using various chemical transport, statistical, and machine learning models.

[0004] However, most previous studies simply selected relevant meteorological factors and topographic factors as training data for machine learning, and rarely considered the impact of the vertical distribution of NO2 on the process of converting the NO2 vertical column (XNO2 obtained from satellites) into the horizontal process of GNO2. As a black box algorithm, the internal implementation details of these traditional machine learning models cannot be understood. Therefore, if there is no high correlation between the training data and GNO2 in a practical sense, the training results obtained will not have high practical value. Therefore, adding the vertical hierarchical structure of NO2 to the retrieval process is an improvement to the model physics. In addition, traditional machine learning models find it difficult to fully express the complex spatiotemporal relationship between GNO2 and XNO2, which is also the reason why the current model results are quite different from the field measurement results.

[0005] Generally speaking, the inversion method of ground-based stations uses gas sensors or chemical analysis instruments to obtain real-time data to measure NO2 concentrations. It has high timeliness and accuracy, and can promptly detect and respond to air pollution incidents. However, since ground-based stations require a large amount of equipment and human resources to maintain and operate, and their spatial layout is limited by various objective factors such as geographical conditions, resulting in uneven distribution of stations, it is unable to provide concentration data over a wide range. While satellite remote sensing can provide global NO2 concentration data, it can only provide total atmospheric or tropospheric NO2 column concentration data and cannot directly provide GNO2 concentration data. Furthermore, due to the influence of weather conditions, cloud cover, and algorithm selection, there are also problems such as missing data and inaccurate measurement accuracy. Summary of the Invention

[0006] The purpose of the present invention is to provide a method, system, equipment and medium for inverting near-ground nitrogen dioxide concentration to solve the problems existing in the above-mentioned prior art.

[0007] To achieve the above-mentioned object, the present invention provides a method for inverting near-surface nitrogen dioxide concentration, comprising:

[0008] Acquiring a data set, wherein the data set includes European Centre for Medium-Range Weather Forecasts reanalysis data, tropospheric nitrogen dioxide column concentration data, MODIS satellite sensing data, global terrain elevation data, and global population distribution data;

[0009] Correcting the tropospheric nitrogen dioxide column concentration data based on the European Centre for Medium-Range Weather Forecasts reanalysis data to obtain surface nitrogen dioxide concentration data; spatially aligning the dataset with the surface nitrogen dioxide concentration data to obtain a modeling dataset;

[0010] A nitrogen dioxide concentration inversion model is constructed based on a lightweight gradient boosting machine model and an extreme random forest model, the nitrogen dioxide concentration inversion model is trained based on the modeling data set, and nitrogen dioxide concentration inversion is performed in the area to be inverted based on the trained nitrogen dioxide concentration inversion model to obtain nitrogen dioxide concentration data in the area to be inverted at different temporal and spatial scales.

[0011] Optionally, the correcting the tropospheric nitrogen dioxide column concentration data specifically includes:

[0012] The vertical variation data of nitrogen dioxide concentration in the troposphere are simulated based on the reanalysis data of the European Centre for Medium-Range Weather Forecasts, and the ratio of nitrogen dioxide concentration below the boundary layer to the overall concentration in the troposphere is calculated based on the vertical variation data;

[0013] The tropospheric nitrogen dioxide column concentration data is resampled, and the resampled tropospheric nitrogen dioxide column concentration data is height-corrected based on the ratio data to obtain near-ground nitrogen dioxide concentration data.

[0014] Optionally, the resampled tropospheric nitrogen dioxide column concentration data is height-corrected based on the ratio data, and the specific calculation formula is:

[0015]

[0016] Where, is the height correction value of the satellite-observed tropospheric NO2 column concentration (XNO2), For satellite observation of XNO2 data, is the NO2 concentration at different layers / heights (r), and BLH is the boundary layer height derived from the ERA5 reanalysis data.

[0017] Optionally, the process of constructing the modeling dataset specifically includes:

[0018] The dataset and the near-ground nitrogen dioxide concentration data are spatially aligned, and based on a feature correlation analysis method, data having a strong correlation with the nitrogen dioxide concentration is selected, and the modeling dataset is constructed based on the data selected based on the feature correlation analysis method.

[0019] Optionally, the training of the nitrogen dioxide concentration inversion model based on the modeling data set specifically includes:

[0020] The modeling data set is input into the nitrogen dioxide concentration inversion model for concentration inversion, and training is performed according to the physical change mechanism of near-ground nitrogen dioxide to obtain a trained nitrogen dioxide concentration inversion model.

[0021] Optionally, inputting the modeling data set into the nitrogen dioxide concentration inversion model to perform concentration inversion specifically includes:

[0022] The modeling data set is input into a lightweight gradient boosting machine model for concentration prediction to obtain a first concentration inversion result. The residual between the first concentration inversion result and the corresponding measured value is input into an extreme random forest model for optimization to obtain a second concentration inversion result. Based on the first concentration inversion result and the second concentration inversion result, a total nitrogen dioxide concentration inversion result is constructed to complete the concentration inversion.

[0023] Optionally, the total nitrogen dioxide concentration inversion result is constructed based on the first concentration inversion result and the second concentration inversion result. The specific calculation formula is:

[0024] GNO2(i,j)=stableGNO2(i,j)+randomGNO2(i,j)

[0025] where GNO2(i,j) is the near-ground nitrogen dioxide level of the specified pixel; stableGNO2(i,j) and randomGNO2(i,j) are the stable term and random term of the specified pixel.

[0026] A near-ground nitrogen dioxide concentration inversion system, comprising:

[0027] a data acquisition module for acquiring a data set, the data set including European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis data, tropospheric nitrogen dioxide column concentration data, MODIS satellite sensing data, global terrain elevation data, and global population distribution data; correcting the tropospheric nitrogen dioxide column concentration data based on the ECMWF reanalysis data, spatially aligning the data set with the corrected tropospheric nitrogen dioxide column concentration data to obtain a modeling data set;

[0028] The nitrogen dioxide concentration inversion module is used to construct a nitrogen dioxide concentration inversion model based on a lightweight gradient boosting machine model and an extreme random forest model, train the nitrogen dioxide concentration inversion model based on the modeling data set, and perform nitrogen dioxide concentration inversion in the area to be inverted based on the trained nitrogen dioxide concentration inversion model to obtain nitrogen dioxide concentration data in the area to be inverted at different temporal and spatial scales.

[0029] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform a method for inverting near-ground nitrogen dioxide concentration.

[0030] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method for inverting near-ground nitrogen dioxide concentration.

[0031] The technical effects of the present invention are:

[0032] The present invention integrates various factors that affect the near-ground nitrogen dioxide concentration (GNO2), such as meteorological data, vegetation index (such as NDVI), terrain data (such as DEM) and population data, so that the inversion results are more comprehensive and reliable. This embodiment constructs a two-layer machine learning model by combining LGBM (lightweight gradient boosting machine model) and ERF (extreme random forest), which can more accurately capture the stability effect and random effect in the time series of GNO2 concentration, thereby improving the accuracy of the inversion results. This solution provides an efficient, accurate and reliable near-ground nitrogen dioxide concentration inversion method through an advanced two-layer machine learning model, which can obtain GNO2 products on different time scales, more accurately reflect abnormal events, have strong applicability and authenticity, and have important application value for environmental monitoring and air quality management. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0034] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0035] Figure 1 Flowchart of a method for inverting near-surface nitrogen dioxide concentration based on a two-layer machine learning model in an embodiment of the present invention;

[0036] Figure 2 is a correlation coefficient diagram between different influencing factors in an embodiment of the present invention;

[0037] Figure 3 This figure shows the accuracy comparison between the two-layer machine learning model in an embodiment of the present invention and the traditional model;

[0038] Figure 4 This is a time series distribution diagram of GNO2 concentration data inverted by the model in an embodiment of the present invention in China and typical regions from 2018 to 2023;

[0039] Figure 5 2 is a structural diagram of a near-ground nitrogen dioxide inversion device in an embodiment of the present invention. DETAILED DESCRIPTION

[0040] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0041] It should be understood that the terms described herein are intended only to describe particular embodiments and are not intended to limit the present invention. In addition, for numerical ranges herein, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each smaller range between any intermediate value within a stated value or stated range and any other stated value or intermediate value within the stated range is also encompassed by the present invention. The upper and lower limits of these smaller ranges may be independently included or excluded within the scope.

[0042] The words “include,” “including,” “have,” “contain,” etc. used in this article are open-ended terms, meaning including but not limited to.

[0043] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0044] Example 1

[0045] like Figure 1 - Figure 5 As shown, this embodiment provides a near-ground nitrogen dioxide concentration inversion method, including: acquiring a data set, the data set including European Centre for Medium-Range Weather Forecasts reanalysis data, tropospheric nitrogen dioxide column concentration data, MODIS satellite sensor data, global terrain elevation data and global population distribution data; correcting the tropospheric nitrogen dioxide column concentration data based on the European Centre for Medium-Range Weather Forecasts reanalysis data to obtain near-ground nitrogen dioxide concentration data; spatially aligning the data set and the near-ground nitrogen dioxide concentration data to obtain a modeling data set; constructing a nitrogen dioxide concentration inversion model based on a lightweight gradient boosting model and an extreme random forest model, training the nitrogen dioxide concentration inversion model based on the modeling data set, performing nitrogen dioxide concentration inversion in the area to be inverted based on the trained nitrogen dioxide concentration inversion model, and obtaining nitrogen dioxide concentration data in the area to be inverted at different temporal and spatial scales.

[0046] This embodiment discloses a method, device, system, and storage medium for inverting near-ground nitrogen dioxide (GNO2) concentration based on a two-layer machine learning model. This method takes into account the impact of NO2's vertical layered structure on GNO2 and the difficulty of traditional machine learning models in fully expressing the complex spatiotemporal relationship between GNO2 and XNO2. By combining an advanced lightweight gradient boosting machine (LGBM) and an extreme random forest (ERF) model, a two-layer machine learning (DLML) model is established to invert China's GNO2 concentration from 2018 to 2023. The annual average GNO2 data obtained in this embodiment are basically consistent with the measured data at the monitoring stations. At the same time, GNO2 products on different time scales can be obtained, which more accurately reflect abnormal events and have strong applicability and authenticity.

[0047] The implementation steps of this embodiment include:

[0048] 1) Step 1: Analyze the sources and influencing factors of GNO2 concentration and perform altitude correction on the tropospheric NO2 column concentration data provided by TROPOMI;

[0049] 2) Step 2: Preprocess the data products such as the highly corrected NO2 data, meteorological data, topographic data, and population data and unify them to the same resolution for feature correlation analysis;

[0050] 3) Step 3: The processed variable data are combined into a modeling dataset, and a two-layer machine learning model is used to adjust the model parameters of the matching dataset to achieve the optimal model verification results;

[0051] 4) Step 4: Use the adjusted optimal prediction model to perform inversion to obtain the GNO2 concentration dataset at different spatial and temporal scales in the region.

[0052] Perform height correction on existing tropospheric NO2 column concentration data products: Obtain CAMS dataset and boundary layer height (BLH) data provided by the European Centre for Medium-Range Weather Forecasts (ECMWF) and perform height correction on tropospheric NO2 column concentration data. The CAMS dataset product has a spatial resolution of 40 km and a temporal resolution of 1 day. The boundary layer height (BLH) data has a spatial resolution of 0.25°×0.25°.

[0053] The existing tropospheric NO2 column concentration data product is obtained. The spatial resolution of this product is 5.5×3.5 km and the temporal resolution is 1 day. However, there are missing values ​​in the daily NO2 column concentration data. Therefore, the TROPOMI data are resampled to a 0.05°×0.05° grid using bilinear interpolation.

[0054] Using the NO2 concentration profile data provided by CAMS, we simulated the vertical variation of concentration and calculated the ratio of NO2 concentration below the boundary layer to the overall troposphere concentration. Based on this calculated ratio, we corrected the NO2 column concentration product of TRO POMI, i.e., we corrected the NO2 concentration of TROPOM in the troposphere to the surface concentration to obtain GNO2. The specific calculation formula is as follows:

[0055]

[0056] in, is the height correction value of the satellite-observed tropospheric NO2 column concentration (XNO2), For satellite observation of XNO2 data, is the NO2 concentration at different layers / heights (r), and BLH is the boundary layer height derived from the ERA5 reanalysis data.

[0057] Data products such as height-corrected NO2 data, meteorological data, topographic data, and population data were preprocessed and unified to the same resolution. Specifically, hourly meteorological data from ECMWF were obtained, including 2-m temperature (TEMP), relative humidity (RELH), surface pressure (PRES), 10-m meridional wind (u10), and 10-m longitudinal wind (v10), with a spatial resolution of 0.25°×0.25°.

[0058] Obtain NDVI data from MODIS, DEM data from SRTM, and the LandScan population dataset from EVC. The temporal and spatial resolutions of NDVI data are 16 days and 250M, respectively, the resolution of DEM data is 30m, and the resolution of LandScan is 1km.

[0059] The modeling dataset was obtained by spatially unifying the highly corrected NO2 data, meteorological data, and topographic data into a 0.05° × 0.05° grid. By resampling all data to a uniform spatial resolution (e.g., 0.05° × 0.05°), compatibility and comparability between different data sources were ensured, improving the consistency of the inversion dataset.

[0060] The processed variable data are combined into a modeling data set, and the matching data set is modeled using a two-layer machine learning model. Specifically, using the data set, a DLML model is established by combining advanced LGBM and ERF models to describe the stable effect and random effect of GNO2 concentration in the time series, respectively. The retrieval residuals of the first-layer LGBM model are regarded as random terms. Then, the ERF model in the second layer can estimate the random terms to correct the inversion level of GNO2. This embodiment decomposes the change in GNO2 concentration into a stable term and a random term, allowing the model to handle these two different changes separately, thereby improving the model's ability to identify and respond to different influencing factors. By combining with the two-layer model, the advantages of different models can be better utilized, thereby improving the overall retrieval performance.

[0061] It is assumed that the true GNO2 level consists of two components: a stable term and a random term. The stable term refers to the stability of GNO2 across years. Specifically, the variation in GNO2 concentration at the same time in different years can be ignored. The numerical deviations are usually small and show similar trends. The random term refers to the random fluctuations of GNO2. GNO2 is affected by meteorological conditions or human activities and can vary unpredictably at the same time in different years. It may manifest as a rapid change in the GNO2 concentration at a pixel every day. Therefore, according to this mechanism, the true GNO2 level can be decomposed into:

[0062] GNO2(i,j)=stableGNO2(i,j)+randomGNO2(i,j)

[0063] Among them, GNO2(i,j) refers to the true GNO2 level of the specified pixel (i,j); stableGNO2(i,j) and randomGNO2(i,j) are the stable term and random term of the specified pixel (i,j).

[0064] The stable term represents the long-term pattern of GNO2 levels. We choose to simulate the stable term of GNO2 by inversion of the LGBM model, which has more efficient training capacity and more stable training results. The LGBM model is used as the first layer for retrieval:

[0065]

[0066] in, is the stabilizing term for a given pixel (i, j); the explanatory variables of the LGBM model include CNO2, NDVI, DEM, POP, and other meteorological factors. The predictor variable of the LGBM model is the monthly average of ground station data.

[0067] At the same time, it is difficult to characterize the daily random variation of GNO2 levels using only the traditional LGBM model. Therefore, the ERF model is introduced as a second layer to further optimize the residual results of the LGBM model inversion, thereby reducing the negative impact of random effects on the single traditional machine learning model and inverting the true GNO2 random term. The ERF model is used to describe the random term of GNO2:

[0068]

[0069] in, The random term of GNO2 for a given pixel (i, j); CNO2, NDVI, DE M, POP, and other meteorological factors are input into the ERF model as a series of explanatory variables. In addition, the predictor variable of the ERF model is the overall residual of the LGBM model.

[0070] The established GNO2 concentration inversion model is used for prediction to obtain GNO2 concentration datasets at different spatial and temporal scales in the region. Specifically, all modeling data are processed into standard spatial grids, and the established two-layer machine learning model is used to invert the GNO2 value of each pixel in the region to obtain GNO2 concentration datasets at different spatial and temporal scales in the region, and respond to abnormal events.

[0071] This embodiment integrates multiple factors affecting GNO2 concentration, such as meteorological data, vegetation index (such as NDVI), terrain data (such as DEM) and population data, making the inversion results more comprehensive and reliable.

[0072] This example uses a two-layer machine learning model constructed by combining LGBM and ERF to more accurately capture the stability and random effects in the time series of GNO2 concentrations, thereby improving the accuracy of the inversion results. This solution, through an advanced two-layer machine learning model, provides an efficient, accurate, and reliable method for inverting near-surface nitrogen dioxide concentrations. It can obtain GNO2 products on different time scales, more accurately reflect abnormal events, and has strong applicability and authenticity, with important application value for environmental monitoring and air quality management.

[0073] The specific application process of this embodiment includes:

[0074] 1. Analyze the sources and influencing factors of GNO2 concentration, perform height correction on the tropospheric NO2 column concentration data provided by TROPOMI, and pre-process the meteorological data, topographic data, population data and other data products to the same resolution for feature correlation analysis.

[0075] Specifically, the tropospheric NO2 column concentration data used in this embodiment is obtained by the Tropospheric Observation Instrument (TROPOMI) on the Sentinel-5P satellite launched by the European Space Agency (ESA) in 2017. The TROPOMI observation width is about 2600km, which can cover the entire world every day. The spatial resolution has been increased from 7km×3.5km to 5.5km×3.5km since August 6, 2019. It is the atmospheric monitoring spectrometer with the highest spatial resolution today. The TROPOMI tropospheric NO2 column concentration is obtained by inversion based on the differential absorption spectroscopy (DOAS) algorithm in the spectral range of 405-465nm. At present, there are two levels of tropospheric NO2 column concentration data products that TROPOMI can provide, namely, the data formats are L1B and L2. This embodiment uses L2 data, and the time range is from January 1, 2022 to December 31, 2022. Since there are data gaps in the tropospheric NO2 column concentration data provided by satellites, the TR OPOMI data are resampled to a 0.05° × 0.05° grid using bilinear interpolation.

[0076] This example uses the NO2 stratification data provided by CAMS to simulate the vertical variation of tropospheric NO2 concentration and calculate the ratio of NO2 concentration below the boundary layer to the overall tropospheric concentration, thereby calibrating the TROPOMI tropospheric NO2 column concentration product. The specific formula is as follows:

[0077]

[0078] in, is the height correction value of the satellite-observed tropospheric NO2 column concentration (XNO2), For satellite observation of XNO2 data, is the NO2 concentration at different layers / heights (r), and BLH is the boundary layer height derived from the ERA5 reanalysis data.

[0079] The model input variables were selected according to the source and influencing factors of GNO2 concentration, namely meteorological data, vegetation index data, topographic data and population data. All data were uniformly resampled to a spatial resolution of 0.05°×0.05° for feature correlation analysis. The correlation between each data set and GNO2 concentration is shown in Figure 2. Figure 2 .

[0080] 2. The processed variable data are combined into a modeling data set, and the matching data set is modeled using a two-layer machine learning model. Specifically, a modeling data set is constructed by analyzing the correlation of variables. The model selected in this embodiment is a two-layer machine learning model. It is difficult for traditional machine learning models to fully express the complex spatiotemporal relationship between GNO2 and tropospheric NO2 column concentrations. To this end, this embodiment divides the real GNO2 into two components, namely a stable term and a random term. The stable term refers to the part whose changes at the same time in different years can be ignored, for example, the monthly GNO2 concentration of a specified pixel is similar to the value of other years. The random term refers to the part that changes unpredictably at the same time in different years, such as the rapid change in the daily concentration of GNO2 in a specified pixel. According to this mechanism, the real GNO2 can be decomposed into:

[0081] GNO2(i,j)=stableGNO2(i,j)+randomGNO2(i,j)

[0082] Among them, GNO2(i,j) refers to the true GNO2 level of the specified pixel (i,j); stableGNO2(i,j) and randomGNO2(i,j) are the stable term and random term of the specified pixel (i,j).

[0083] (1) Calculate stable GNO2. The stable term represents the long-term pattern of GNO2. Choosing a model that can represent the overall change trend of GNO2 is crucial for expressing GNO2. The LGBM model is an iterative boosting tree system based on the gradient boosting decision tree (GBDT) framework launched by Microsoft in 2017. LGBM uses an enhanced histogram decision tree algorithm to divide continuous feature values ​​into K intervals before training, and traverses the entire sample from the histogram containing K for statistical accumulation to find the optimal split point, which improves the computational execution efficiency by half compared with GBDT. Compared with the classic GBDT that generally only uses the loss function to fit the model, LGBM can find the optimization function through training, thereby integrating K regression trees to fit the final model, with more efficient training capabilities and more stable training results. The LGBM model is regarded as the first layer of inversion, and its formula is as follows:

[0084] stableGNO2(i,j)=LGBM[CNO2(i,j),NDVI(i,j),DEM(i,j),POP(i,j)...]

[0085] In the calculation formula of the first layer of the above inversion, is the stability term for the specified pixel (i, j); LGBM[CNO2(i, j), NDVI(i, j), DEM(i, j), POP(i, j)...] is a set of variables input to the reference LGBM model, including CNO2, NDVI, DEM, POP, and other meteorological factors.

[0086] (2) Calculate random GNO2. It is difficult to use only the traditional LGBM model to describe the random changes in GNO2 between different days. Therefore, the embodiment introduces the ERF model as the second layer to further optimize the inversion results of the first layer to overcome the negative impact of random effects on a single traditional machine learning model. ERF is an ensemble learning algorithm that uses a set of decision trees to improve the accuracy and stability of predictions. It is a variant of the random forest algorithm, but has stronger randomness during the training process. ERF randomly selects features for splitting at each node instead of using the best performing features based on criteria such as information gain. This random selection of the ERF model helps to maintain the random variation characteristics of the daily GNO2 level, thereby preventing overfitting and improving the generalization ability of the ensemble. Therefore, this embodiment uses the ERF model to describe the random terms in the equation, and the specific formula is as follows:

[0087]

[0088] in, is the random term of GNO2 at the specified pixel (i, j); CNO2, NDVI, DEM, POP and other meteorological factors are a series of variables input into the reference ERF model.

[0089] In addition to the direct fitting results after statistical model training, a ten-fold cross-validation method is also used to verify the model, which can avoid potential overfitting problems in the model. This embodiment divides the dataset into 10 equal parts, one of which is used for testing and nine for model fitting. By repeating this process 10 times, the optimized model parameters are selected to obtain the best model. In addition, this embodiment takes into account the influence of adjacent location information and selects sample-based cross-validation and location-based cross-validation to evaluate the generalization and spatial extrapolation of the model.

[0090] At the same time, this embodiment also selects the cross-validation determination coefficient (R 2 ), mean absolute error (MAE) and root mean square error (RMSE) are used as statistical indicators for accuracy evaluation. The calculation formula is as follows:

[0091]

[0092] Among them, y i is the actual value, is the predicted value, is the mean of the actual values ​​and n is the number of observations. 2 The closer the absolute value of is to 1, the better the model fit is. The closer the RMSE and MAE values ​​are to 0, the closer the predicted value is to the actual value.

[0093] The accuracy of the model is verified using the above accuracy indicators, and the results are as follows Figure 3 As shown in Figure 2, it can be seen that the overall accuracy of the DLML model is higher than that of the traditional LGBM model. 2 It can be as high as 0.97, which is 0.02 higher than the traditional LGBM model. Its MAE and RMSE are about 0.87μɡ / m lower than the LGBM model. 3 and 1.11 μɡ / m 3 It can be seen that the DLML model improves the accuracy and reliability of model prediction. The R 2 All are above 0.86, and the cross-sample validation R 2 It can reach 0.87, indicating that the DLML model has strong stability on different data sets. At the same time, compared with the traditional LGBM model, the spatiotemporal cross-validation R 2 The MAE and RMSE are reduced to about 4.5 μɡ / m 3 and 6.1 μɡ / m 3 The universality and effectiveness of the proposed DLML model in unknown data estimation are analyzed. The model can accurately retrieve GNO2 at non-participating sites and has good retrieval ability and spatial generalization ability.

[0094] 3. All modeling data are processed into a standard spatial grid, and the established two-layer machine learning model is used to invert the GNO2 concentration value at each pixel in the region, resulting in a comprehensive set of GNO2 data products covering different spatial and temporal scales in the region. However, due to the short life cycle of GNO2 concentration, analyzing only the annual average may obscure the details of these changes and fail to provide a specific understanding of the overall annual changes. Therefore, a spatial distribution map of the annual average GNO2 value from 2018 to 2023 is drawn, and the rate of change of GNO2 concentration over these five years is calculated, in order to provide a more comprehensive and scientific reference for governance policies.

[0095] The model was used to draw the time series distribution of GNO2 concentration data in China and typical regions from 2018 to 2023, as shown in the figure below: Figure 4 It can be seen that the GNO2 concentration distribution trend inverted in this embodiment is roughly the same as the annual average concentration distribution trend, and the model has good performance in the time dimension and can perform accurate inversion.

[0096] It should be noted that the modules involved in the embodiments of the present invention may be implemented in software or hardware, and the modules described may also be set in a processor. In some cases, the names of these modules do not constitute limitations on the modules themselves.

[0097] This embodiment provides a near-surface nitrogen dioxide concentration inversion system, including:

[0098] a data acquisition module for acquiring a data set, the data set including European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis data, tropospheric nitrogen dioxide column concentration data, MODIS satellite sensing data, global terrain elevation data, and global population distribution data; correcting the tropospheric nitrogen dioxide column concentration data based on the ECMWF reanalysis data, spatially aligning the data set with the corrected tropospheric nitrogen dioxide column concentration data to obtain a modeling data set;

[0099] The nitrogen dioxide concentration inversion module is used to construct a nitrogen dioxide concentration inversion model based on a lightweight gradient boosting machine model and an extreme random forest model, train the nitrogen dioxide concentration inversion model based on the modeling data set, and perform nitrogen dioxide concentration inversion in the area to be inverted based on the trained nitrogen dioxide concentration inversion model to obtain nitrogen dioxide concentration data in the area to be inverted at different temporal and spatial scales.

[0100] This embodiment provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the method for inverting near-ground nitrogen dioxide concentration.

[0101] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method for inverting near-ground nitrogen dioxide concentration is implemented.

[0102] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for inverting near-surface nitrogen dioxide concentration, characterized in that: include: Acquiring a data set, wherein the data set includes European Centre for Medium-Range Weather Forecasts reanalysis data, tropospheric nitrogen dioxide column concentration data, MODIS satellite sensing data, global terrain elevation data, and global population distribution data; The vertical variation data of nitrogen dioxide concentration in the troposphere are simulated based on the reanalysis data of the European Centre for Medium-Range Weather Forecasts, and the ratio of nitrogen dioxide concentration below the boundary layer to the overall concentration in the troposphere is calculated based on the vertical variation data; resampling the tropospheric nitrogen dioxide column concentration data; Based on the ratio data, the resampled tropospheric nitrogen dioxide column concentration data is height-corrected to obtain the near-ground nitrogen dioxide concentration data. The specific calculation formula is: Where, is the height correction value of the NO2 column concentration in the troposphere observed by satellite, For satellite observation of XNO2 data, is the NO2 concentration at different layers / heights r, BLH is the boundary layer height derived from the ERA5 reanalysis data; spatially aligning the dataset and the near-ground nitrogen dioxide concentration data to obtain a modeling dataset; A nitrogen dioxide concentration inversion model is constructed based on a lightweight gradient boosting machine model and an extreme random forest model, the nitrogen dioxide concentration inversion model is trained based on the modeling data set, and nitrogen dioxide concentration inversion is performed in the area to be inverted based on the trained nitrogen dioxide concentration inversion model to obtain nitrogen dioxide concentration data in the area to be inverted at different temporal and spatial scales.

2. The method for inverting near-surface nitrogen dioxide concentration according to claim 1, characterized in that: The process of constructing the modeling dataset specifically includes: The dataset and the near-ground nitrogen dioxide concentration data are spatially aligned, and based on a feature correlation analysis method, data having a strong correlation with the nitrogen dioxide concentration is selected, and the modeling dataset is constructed based on the data selected based on the feature correlation analysis method.

3. The method for inverting near-surface nitrogen dioxide concentration according to claim 1, characterized in that: The training of the nitrogen dioxide concentration inversion model based on the modeling data set specifically includes: The modeling data set is input into the nitrogen dioxide concentration inversion model for concentration inversion, and training is performed according to the physical change mechanism of near-ground nitrogen dioxide to obtain a trained nitrogen dioxide concentration inversion model.

4. The method for inverting near-surface nitrogen dioxide concentration according to claim 3, characterized in that: Inputting the modeling data set into the nitrogen dioxide concentration inversion model to perform concentration inversion specifically includes: The modeling data set is input into a lightweight gradient boosting machine model for concentration prediction to obtain a first concentration inversion result. The residual between the first concentration inversion result and the corresponding measured value is input into an extreme random forest model for optimization to obtain a second concentration inversion result. Based on the first concentration inversion result and the second concentration inversion result, a total nitrogen dioxide concentration inversion result is constructed to complete the concentration inversion.

5. The method for inverting near-surface nitrogen dioxide concentration according to claim 4, characterized in that: The total nitrogen dioxide concentration inversion result is constructed based on the first concentration inversion result and the second concentration inversion result. The specific calculation formula is: GNO2(i,j)=stableGNO2(i,j)+randomGNO2(i,j) where GNO2(i,j) is the near-ground nitrogen dioxide level of the specified pixel; stableGNO2(i,j) and randomGNO2(i,j) are the stable term and random term of the specified pixel.

6. A near-surface nitrogen dioxide concentration inversion system, characterized in that: include: a data acquisition module for acquiring a data set, the data set including European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis data, tropospheric nitrogen dioxide column concentration data, MODIS satellite sensing data, global terrain elevation data, and global population distribution data; simulating vertical variation data of tropospheric nitrogen dioxide concentration based on the ECMWF reanalysis data, and calculating data on the ratio of nitrogen dioxide concentration below the boundary layer to the overall tropospheric concentration based on the vertical variation data; resampling the tropospheric nitrogen dioxide column concentration data; Based on the ratio data, the resampled tropospheric nitrogen dioxide column concentration data is height-corrected to obtain the near-ground nitrogen dioxide concentration data. The specific calculation formula is: Where, is the height correction value of the NO2 column concentration in the troposphere observed by satellite, For satellite observation of XNO2 data, is the NO2 concentration at different layers / heights r, BLH is the boundary layer height derived from the ERA5 reanalysis data; spatially aligning the dataset and the corrected tropospheric nitrogen dioxide column concentration data to obtain a modeling dataset; The nitrogen dioxide concentration inversion module is used to construct a nitrogen dioxide concentration inversion model based on a lightweight gradient boosting machine model and an extreme random forest model, train the nitrogen dioxide concentration inversion model based on the modeling data set, and perform nitrogen dioxide concentration inversion in the area to be inverted based on the trained nitrogen dioxide concentration inversion model to obtain nitrogen dioxide concentration data in the area to be inverted at different temporal and spatial scales.

7. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform a near-ground nitrogen dioxide concentration inversion method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed by a processor, implements a near-ground nitrogen dioxide concentration inversion method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Atmospheric NO2 space-time prediction algorithm combining hyper-spectral satellite and artificial intelligence

    CN115730718A

  • Method for inverting near-surface nitrogen dioxide concentration based on satellite data

    CN115950832A