Method and System for Inverting Vertical Profile of CO Concentration Based on Multi-Source Remote Sensing Data Fusion

Through multi-source remote sensing data fusion and XGBoost machine learning, a CO concentration vertical profile inversion model was established, which solved the problems of low data coverage and lack of vertical information in satellite monitoring technology, and achieved high-precision CO vertical distribution analysis, supporting pollutant traceability and climate effect evaluation.

CN120234770BActive Publication Date: 2025-07-25CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510727505.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-25
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing satellite monitoring technology has low spatial coverage of CO vertical distribution information, and cannot accurately analyze the contribution of high-altitude CO., and traditional statistical models fail to effectively integrate meteorological parameters, resulting in insufficient dynamic characterization capabilities for CO vertical diffusion and chemical reactions.

Method used

The multi-source remote sensing data fusion method is used to filter the correlation sites through unary linear regression and multivariate linear regression, and a CO concentration vertical profile inversion model is established in combination with the XGBoost machine learning algorithm to fill in data loss and generate CO vertical profile distribution with high spatial coverage.

Benefits of technology

It significantly improves the spatiotemporal continuity and accuracy of CO vertical profile data, supports long-term continuous observation, accurately quantifies the contribution ratio of different height layers, and analyzes the cross-regional transmission paths and vertical diffusion mechanism of pollutants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234770B_ABST
    Figure CN120234770B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for retrieving the vertical profile of CO concentration based on multi-source remote sensing data fusion, including: obtaining the CO column concentration data of the TROPOMI satellite, the CO vertical profile mixing ratio data of the MOPITT satellite, and the ERA5 meteorological reanalysis data; constructing a training dataset through spatio-temporal matching and spatial resampling; using the XGBoost machine learning algorithm to establish an estimation model with the TROPOMI column concentration and ERA5 meteorological parameters as inputs and the MOPITT ten-layer vertical profile mixing ratio as the output target; using the trained model to estimate the concentration and fill in the data for the missing areas to generate a reconstructed dataset of the CO vertical profile with high spatial coverage. The present invention significantly improves the spatio-temporal continuity and accuracy of the CO vertical distribution data through multi-source data fusion and intelligent algorithm modeling, solves the technical problems of insufficient coverage of traditional satellite data and missing vertical information, and provides highly reliable data support for the tracing of air pollution sources, cross-regional transmission analysis, and environmental governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of environmental science and artificial intelligence, and particularly relates to a method and system for retrieving the vertical profile of CO concentration based on multi-source remote sensing data fusion. Background Art

[0002] Atmospheric carbon monoxide (CO) is an important trace gas in the Earth's atmosphere, and its sources mainly include incomplete combustion of fossil fuels, biomass burning, and non-combustion processes (such as oxidation of hydrocarbons). By consuming OH radicals in the atmosphere, CO significantly affects the oxidation capacity of the atmosphere and participates in photochemical reactions as a precursor of ozone, indirectly intensifying the greenhouse effect. Traditional monitoring methods (such as ground spectrometers and aircraft surveys) have problems such as limited spatial coverage, high costs, and difficulty in long-term continuous observations, and are difficult to meet the pollution monitoring requirements of large ranges and high timeliness.

[0003] The vertical distribution of atmospheric carbon monoxide (CO) is a key parameter for studying atmospheric chemical processes, pollution transmission mechanisms, and climate effects. The vertical profile information of CO can reveal the following core scientific issues:

[0004] 1. Source Apportionment: CO emitted from the ground (such as industry and transportation) is mainly distributed in the near-surface layer, while CO generated by biomass burning, sandstorms, etc. can affect the middle atmosphere through uplift effects and even be transported to the stratosphere over long distances;

[0005] 2. Chemical Process Research: The reaction rate of CO with OH radicals varies significantly with height, and the oxidation capacity of the boundary layer is weaker than that of the free convective layer, resulting in differences in the lifetime of CO at different heights (about 1 month near the ground and up to 2 - 3 months at high altitudes);

[0006] 3. Climate Effect Assessment: High-altitude CO (300 hPa - 100 hPa) contributes 3 - 5 times more to radiative forcing by affecting stratospheric ozone formation than near the ground.

[0007] In recent years, satellite remote sensing technology has provided new solutions for CO monitoring. Among them:

[0008] MOPITT (Measurements of Pollution in the Troposphere) satellite:

[0009] It can provide 10-layer CO vertical profile mixing ratio data from the surface (1000 hPa) to the upper troposphere (100 hPa). However, affected by factors such as satellite orbital gaps and cloud cover, its daily average spatial coverage is less than 5%, and data is severely missing in the Qinghai-Tibet Plateau and cloudy areas, seriously restricting the study of the distribution and transmission process of high-altitude CO.

[0010] Existing studies have shown that the coverage rate of MOPITT data in complex terrain areas is less than 1%, resulting in insufficient monitoring capabilities for sudden pollution events such as biomass burning.

[0011] TROPOMI (TROPOspheric Monitoring Instrument) satellite:

[0012] It has a high spatial resolution (0.05°×0.05°) and the ability to cover the globe daily, but only provides the total column concentration of CO, lacks vertical stratification information, and cannot distinguish between ground and upper-air contributions.

[0013] In extreme events (such as cross-border biomass burning), due to the inability of TROPOMI data to resolve the vertical distribution, it is difficult to accurately quantify the contribution of upper-level transport to ground pollution.

[0014] Limitations of existing estimation methods:

[0015] Traditional statistical models (such as simple linear regression) only rely on the simple association between TROPOMI column concentration and ground monitoring values, without integrating meteorological parameters and upper-air information, resulting in significant errors in the model during vertical transport events.

[0016] Existing technologies have not proposed an effective method to quantify the contribution of upper-level CO, and cannot reveal the cross-regional transport mechanism of pollutants.

[0017] Insufficient utilization of meteorological data:

[0018] Key parameters (such as vertical velocity, temperature) in meteorological reanalysis data such as ERA5 have not been fully integrated into the model, resulting in insufficient dynamic characterization ability for CO vertical diffusion and chemical reactions.

[0019] Summary of existing technical defects:

[0020] Low data coverage: There are a large number of missing values in the MOPITT vertical profile data in terms of time and space, making it difficult to support high-resolution research.

[0021] Lack of vertical information: TROPOMI only provides the total column concentration and cannot resolve the pollution contributions of different altitude layers.

[0022] Insufficient model accuracy: Traditional statistical methods do not integrate multi-source data and have significant errors under complex meteorological conditions.

[0023] Unclear upper-level transport mechanism: Lack of a quantitative analysis tool for the upper-air CO transport path and contribution ratio. Summary of the Invention

[0024] In view of the defects of the existing technology, the present invention provides a method and system for retrieving the vertical profile of CO concentration based on the fusion of multi-source remote sensing data.

[0025] To achieve the above invention object, the technical solution adopted by the present invention is as follows:

[0026] A method for retrieving the vertical profile of CO concentration based on multi-source remote sensing data fusion, comprising the following steps:

[0027] Step S1: Through the unary linear regression analysis of the CO concentration at the ground station and the TROPOMI CO column concentration, select the relevant stations, and divide the central data set and the extreme data set based on the ratio factor;

[0028] Introduce the MOPITT CO vertical profile mixing ratio data into the extreme data set, establish a multiple linear regression model, calculate the height factors of each height layer, and select the key height layers;

[0029] Step S2: Obtain the CO column concentration data of the TROPOMI satellite, the CO vertical profile mixing ratio data of the MOPITT satellite, and the ERA5 meteorological reanalysis data;

[0030] Step S3: Perform preprocessing of spatio-temporal matching and spatial resampling on the data in S2, and construct a training data set including the TROPOMI column concentration, ERA5 meteorological parameters, and MOPITT vertical profile data;

[0031] Step S4: Use the XGBoost machine learning algorithm to establish a CO concentration vertical profile inversion model, and use the TROPOMI column concentration and ERA5 meteorological parameters as input features, and the MOPITT ten-layer vertical profile mixing ratio as the output target for model training;

[0032] Step S5: Use the trained model to estimate the concentration in the area lacking MOPITT data, and generate a CO vertical profile reconstruction data set with high spatial coverage;

[0033] Step S6: Integrate the model estimation results with the original MOPITT data, fill in the missing areas, and output the complete spatio-temporal distribution data of the CO vertical profile.

[0034] Furthermore, in step S1, the formula of the unary linear regression model is as follows:

[0035] ,

[0036] In the formula, represents the current ground station number, is the TROPOMI CO column concentration value of the satellite grid corresponding to the ground station, is the CO concentration value monitored by the ground station, is the linear regression fitting parameter.

[0037] The ratio factor has the following formula:

[0038] ,

[0039] where, temperature at time ; and the monitored value of ground CO concentration at time time .

[0040] The established multiple linear regression model is as follows:

[0041] ,

[0042] In the formula, is the MOPITT CO column concentration data of the pixel point corresponding to the ground station, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 300 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 900 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 800 hPa to 300 hPa, and are the multiple linear regression fitting parameters, where is to unify the dimension level of the equation, represents temperature.

[0043] Furthermore: The calculation formula of the height factor in S1 is as follows:

[0044] ,

[0045] where, is the height factor, is the fitting coefficient of the th layer in the multiple regression model, is the CO concentration of the th layer, is the height layer index, and the value range is from 1 to 10.

[0046] When ≥ 0.3, it is determined as a significant area of high-level transportation, and a pollution transmission path map is generated.

[0047] Furthermore, in the step S2: The TROPOMI data is an L3-level offline product with a spatial resolution of 0.05°×0.05°;

[0048] The MOPITT data is a combined channel (TIR+NIR) level-3 product, including 10-layer vertical profile mixing ratio data with a spatial resolution of 1°×1°;

[0049] The ERA5 data includes temperature, meridional wind, zonal wind, and vertical velocity parameters at ten pressure levels from 1000 hPa to 100 hPa, with a spatial resolution of 0.25°×0.25°.

[0050] Further, the MOPITT data is resampled to a 0.25°×0.25° grid using bilinear interpolation;

[0051] The TROPOMI data is resampled to a 0.25°×0.25° grid using area weighting;

[0052] The MOPITT data is quality-filtered to retain valid observation data with a signal degree of freedom (DFS) value ≥1.5.

[0053] Further, the training dataset is divided into 5 subsets. Each time, 4 subsets are taken as the training set, and the remaining 1 subset is used as the validation set;

[0054] The XGBoost hyperparameters are optimized through random search, and the final parameter combination is determined as follows: the number of weak estimators n_estimators = 901, the maximum tree depth max_depth = 9, and the learning rate learning_rate = 0.14;

[0055] An early stopping mechanism is set to terminate training when the performance of the validation set does not improve for 100 consecutive iterations.

[0056] Further, the data reconstruction in step S5 includes:

[0057] Generate daily predicted CO concentration values for ten layers at each grid point;

[0058] Retain the original valid MOPITT observation data and only fill in the missing data areas with the model;

[0059] Mark the confidence level for the estimated results with missing features, and give priority to the measured values when there is original MOPITT data at the corresponding position.

[0060] Further, the data fusion in step S6 uses a spatio-temporal continuity verification method:

[0061] Establish a three-dimensional spatial correlation function for adjacent grid points:

[0062] ,

[0063] where, is the model-estimated CO concentration value of the i-th sample, is the original observed CO concentration value of the i-th sample. is the mean of the estimated value and the observed value, is the standard deviation of the estimated value and the observed value, is the total number of samples, is the spatial coordinate (longitude and latitude) and the time index, representing the calculation result of the correlation coefficient in a specific area and time period.

[0064] When the value < 0.7, start the manual review mechanism to verify the rationality of the transmission channel in combination with the meteorological field data.

[0065] The present invention also discloses a CO concentration vertical profile inversion system, which can be used to implement the above-mentioned CO concentration vertical profile inversion method. Specifically, it includes:

[0066] Data acquisition module: Regularly pull data from the GEE, NASA, and Copernicus platforms.

[0067] Preprocessing module: Call the HARP toolbox to complete resampling and quality filtering.

[0068] Model prediction module: Batch generate reconstructed data based on the trained XGBoost model.

[0069] Visualization output module: Generate the following products:

[0070] Dynamic atlas of the spatio-temporal distribution of the CO vertical profile;

[0071] Spatio-temporal evolution sequence diagram of high-level transportation events;

[0072] Thermal map of the pollution contribution degree in key areas, dividing the pollution source types by different height layers.

[0073] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned CO concentration vertical profile inversion method is realized.

[0074] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned CO concentration vertical profile inversion method is realized.

[0075] Compared with the prior art, the advantages of the present invention are:

[0076] Through the fusion of multi-source satellite data and machine learning modeling, the present invention overcomes the deficiencies of traditional monitoring technologies and realizes the following core advantages:

[0077] 1. Significantly fill the data gaps caused by satellite orbit gaps and cloud cover, enhance the monitoring ability in complex terrain and cloudy areas, and support long-term continuous observation and analysis.

[0078] 2. Use machine learning algorithms to establish a non-linear mapping relationship to achieve high-precision estimation of the CO vertical profile. The model performance is stable and reliable, superior to traditional statistical methods.

[0079] 3. Propose quantitative indicators to effectively analyze the contribution ratio of different altitude layers to the total column concentration, and accurately identify the cross-regional transport path and vertical diffusion mechanism of pollutants.

[0080] 4. Integrate satellite data and meteorological parameters to reveal the influence of meteorological conditions (such as wind field, temperature) on the CO distribution, and support the analysis and tracing of the formation mechanism of pollution events. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 is a flowchart of the method for inverting the CO concentration vertical profile in an embodiment of the present invention;

[0082] Figure 2 is a scatter plot of the prediction results of the five-fold cross-validation test set of the XGBoost model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0083] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further describes the present invention in detail with reference to the drawings and by way of examples.

[0084] As Figure 1 shown, a method for inverting the CO concentration vertical profile based on multi-source remote sensing data fusion includes the following steps:

[0085] The formula of the unary linear regression model is as follows:

[0086] ,

[0087] In the formula, represents the current ground station number, is the TROPOMI CO column concentration value of the satellite grid corresponding to the ground station, is the CO concentration value monitored by the ground station, is the linear regression fitting parameter.

[0088] The ratio factor, the formula is as follows:

[0089] ,

[0090] Among them, temperature at time , time is the ground CO concentration monitoring value at time;

[0091] The established multiple linear regression model is as follows:

[0092] ,

[0093] In the formula, is the MOPITT CO column concentration data of the pixel corresponding to the ground station, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 300 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 900 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 800 hPa to 300 hPa, and are the multiple linear regression fitting parameters. Among them is to unify the dimension level of the equation, represents temperature.

[0094] Furthermore: The calculation formula of the height factor in S1 is as follows:

[0095] ,

[0096] where is the height factor, is the fitting coefficient of the th layer in the multiple regression model, is the th layer CO concentration, is the height layer index, and the value range is from 1 to 10.

[0097] When ≥ 0.3, it is determined as a significant high-level transport area, and a pollution transport path map is generated.

[0098] I. Data Validation and Feature Screening

[0099] (1) Unary linear regression to verify the ground-column concentration relationship:

[0100] ,

[0101] In the formula, represents the current ground station number, is the TROPOMI CO column concentration value of the satellite grid point corresponding to the ground station, is the CO concentration value monitored by the ground station, is the linear regression fitting parameter.

[0102] Implementation operation:

[0103] For each ground station (a total of 1577 stations), a unary linear regression model is established respectively to calculate And the correlation coefficient R, select the stations with R≥0.6 (such as 218 stations including Beijing, Tianjin, etc.).

[0104] Divide the dataset by the ratio factor:

[0105] ,

[0106] Among them, The temperature at time ; The monitored value of ground CO concentration at time .

[0107] Implementation operation:

[0108] Set the threshold range 0.8≤ ≤1.2, and divide the data into:

[0109] Central dataset ( Within the threshold, there are 143 stations in total): directly used for model training;

[0110] Extreme dataset ( Beyond the threshold, there are 144 stations in total): need to introduce vertical stratification information for correction.

[0111] Quantify the contribution of vertical stratification by multiple linear regression:

[0112] ,

[0113] In the formula, is the MOPITT CO column concentration data of the pixel point corresponding to the ground station, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 300 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 900 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 800 hPa to 300 hPa, and are the fitting parameters of multiple linear regression. Among them, is to unify the dimension level of the equation, represents the temperature.

[0114] Implementation operation:

[0115] For the extreme dataset, calculate the height factor and select the stations with ≥0.3 (such as Henan, Shandong and other regions), indicating that the contribution of high-altitude transmission is significant;

[0116] Retain the corresponding height layer (800 - 300 hPa) as the key input feature for the subsequent XGBoost model.

[0117] II. Data Preparation and Preprocessing

[0118] (1) Data Sources and Acquisition:

[0119] TROPOMI Data: Obtain the L3-level CO column concentration data (spatial resolution 0.05°×0.05°) from July 1, 2018 to December 27, 2020 through the Google Earth Engine platform, and filter the valid data with a QA value greater than 0.5. The data format is NetCDF, and the time resolution is daily average.

[0120] MOPITT Data: Download the MOPITT V8 combined channel (TIR+NIR) CO vertical profile mixing ratio data (spatial resolution 1°×1°) from the NASA official website, which includes 10 layers of data from 1000hPa to 100hPa. Use the threshold of DFS≥1.5 for quality filtering.

[0121] ERA5 Meteorological Data: Download the hourly meteorological parameters from 1000hPa to 100hPa from the Copernicus platform, as shown in Table 1, including temperature (t), meridional wind (u), zonal wind (v), and vertical velocity (w), with a spatial resolution of 0.25°×0.25°.

[0122] Table 1 ERA5 Meteorological Parameters

[0123]

[0124] (2) Temporal and Spatial Matching and Resampling:

[0125] Temporal Matching: Unify the TROPOMI and MOPITT data to the daily average from 9:00-15:00 (UTC time 1:00-7:00), and match it with the satellite overpass time. Take the average value of the corresponding period for the ERA5 data.

[0126] Spatial Resampling:

[0127] MOPITT Data: Resample the 1°×1° data to the 0.25°×0.25° grid using the bilinear interpolation method.

[0128] TROPOMI Data: Use the HARP toolbox for area-weighted resampling to generate 0.25°×0.25° grid data.

[0129] Data Format Conversion: Store the processed data as a CSV file by date, with each row containing longitude and latitude, TROPOMI column concentration, ERA5 meteorological parameters (10 layers×4 parameters), and MOPITT ten-layer CO concentration.

[0130] (3)Dataset Division (statistical results are shown in the site number distribution in Table 2):

[0131] Training Dataset: Screen grid points with complete TROPOMI, MOPITT, and ERA5 data simultaneously, with a total of 616,324 samples.

[0132] Estimation Dataset: Includes all spatio-temporally matched grid points, allowing some features to be missing (such as missing MOPITT data), and is used for model prediction filling.

[0133] Table 2 Number of Sites for Accuracy Evaluation of Unary Linear Regression Model

[0134]

[0135] III. Construction and Training of XGBoost Model

[0136] (1)Input Features and Target Variables (parameter definitions are shown in Table 1):

[0137] Feature Variables (41 dimensions):

[0138] TROPOMI CO Column Concentration (1 dimension).

[0139] ERA5 Meteorological Parameters: Temperature (t), meridional wind (u), zonal wind (v), and vertical velocity (w) for each layer, a total of 10 layers × 4 parameters = 40 dimensions.

[0140] Target Variable: MOPITT Ten-Layer CO Vertical Profile Mixing Ratio (from 1000 hPa to 100 hPa, 10 dimensions).

[0141] (2)XGBoost Model Configuration:

[0142] Algorithm Selection: Use the XGBoost regression model, with the loss function being the mean squared error (MSE), and the goal being to minimize the deviation between the predicted value and the measured value.

[0143] Hyperparameter Optimization:

[0144] Search Space: n_estimators is an integer between 100 and 1000; max_depth is an integer between 3 and 10, and learning_rate is a decimal between 0.001 and 0.3.

[0145] Random Search Results: Finally, the optimal parameters are determined as n_estimators = 901, max_depth = 9, and learning_rate = 0.14.

[0146] Training Strategy:

[0147] Five-fold cross validation: The training set is randomly divided into 5 subsets, 4 subsets are used for training and 1 is used for validation. Repeat 5 times and take the average result.

[0148] Early stopping mechanism: When the RMSE of the validation set does not improve for 100 consecutive iterations, the training is terminated to prevent overfitting.

[0149] (3) Model performance evaluation:

[0150] In addition to using the XGBoost model to build a CO high-level concentration estimation model, this example also uses a common linear regression model (LR) as a control group to verify the performance of the XGBoost model. Table 3 shows the verification results of the five-fold cross-test set of different algorithm models. It can be seen that the CO high-level concentration estimation model built based on XGBoost shows better performance. The overall R 2 The LR model has a high accuracy of 0.96, RMSE of 5.79 ppbv, and MAE of 3.81 ppbv. 2 The R of the XGBoost model is only 0.28, the RMSE is 23.74ppbv, and the MAE is 18.32ppbv, which is far inferior to the XGBoost model. 2 The range is between 0.92-0.99, with high accuracy and stability. The RMSE range is 0.94-15.65ppbv, and the MAE range is 0.65-10.03ppbv. Due to the characteristics of different high-level data sets, the RMSE and MAE calculated based on the concentration values decrease with the increase of altitude, and the values are within the acceptable range. 2 The range is 0.02-0.58, with poor accuracy and very unstable, the range of RMSE is 12.46-57.98ppbv, and the range of MAE is 9.84-42.94ppbv, with large values and wide range of variation. Based on the above, it can be seen that the CO concentration estimation model constructed by XGBoost has good generalization ability, the model prediction results are highly accurate, and the RMSE and MAE are low, the model error is small, and it shows good performance.

[0151] Table 3 Five-fold cross validation results of different algorithm models

[0152]

[0153] Figure 2This is a scatter plot of the prediction results of the XGBoost model's five-fold cross-validation test set. The scatter plot can intuitively reflect the data distribution of the model's predicted values and the true values, as well as the gap between the two. The abscissa represents the true value, that is, the mixing ratio concentration of the CO vertical profile at different altitude levels observed by MOPITT; the ordinate represents the mixing ratio concentration of the CO vertical profile at different altitude levels predicted by the model, and the unit is ppbv for both; the color reflects the density of the data points at that coordinate, and from blue to red indicates an increase in data points. The black line in the figure is the 1:1 line with a slope of 1. If the data points exactly fall on this line, it means that the model's predicted value is equal to the observed value, which is the most ideal situation. The red line in the figure is the best fit line. It accurately reflects the actual relationship between the model's predicted value and the observed value by minimizing the vertical distance from all sample points to it. The closer the fit line is to the 1:1 line, the better the correlation between the model's predicted value and the observed value. From the data distribution of the scatter plot of the model's prediction results, the red data points in each layer are concentrated in the middle or the lower left area, indicating that the mixing ratio concentration of the CO vertical profile at different altitude levels is at or slightly below the average level of that layer. For this part of the data, the red fit line is very close to the black 1:1 line, and the red data points are also concentrated around the fit line, indicating that the predicted value and the observed value of the data model are relatively close here, the model accuracy is high, and at the same time, the prediction effect of the model at different altitude levels does not show a particularly large deviation, and the stability of the model is good. It can also be seen from the scatter plot that when the mixing ratio concentration of the CO vertical profile is at a high value in each layer, the fit line is below the 1:1 line, indicating that the model is prone to underestimation when the mixing ratio of the CO vertical profile is at a high concentration. This phenomenon exists in each altitude level. The underestimation phenomenon at the 1000 hPa altitude level mainly occurs when the mixing ratio concentration of the CO vertical profile is higher than about 450 ppbv; the underestimation phenomenon at the 900 hPa altitude level mainly occurs when the concentration is higher than about 400 ppbv; the underestimation phenomenon at the 800 hPa altitude level mainly occurs when the concentration is higher than about 250 ppbv; the underestimation phenomenon at the 700 hPa altitude level mainly occurs when the concentration is higher than about 200 ppbv; the underestimation phenomenon at the 600 hPa - 500 hPa altitude levels mainly occurs when the concentration is higher than about 130 ppbv; the underestimation phenomenon at the 400 hPa - 300 hPa levels mainly occurs when the concentration is higher than about 125 ppbv; the underestimation phenomenon at the 200 hPa level mainly occurs when the concentration is higher than about 140 ppbv; the underestimation phenomenon at the 100 hPa level is not obvious.When the CO vertical profile mixing ratio concentration is at the low value of the layer at different altitudes, the fitting line is above the 1:1 line, indicating that the model is prone to overestimation when the CO vertical profile mixing ratio concentration is low. This overestimation phenomenon is not obvious at the 1000hPa-600hPa altitude layer. At the 500hPa-300hPa altitude layer, when the CO vertical profile mixing ratio concentration is lower than about 60ppbv, overestimation will occur; the overestimation phenomenon at 200hPa mainly occurs when the concentration is lower than about 45ppbv, and the overestimation phenomenon at the 100hPa altitude layer is not obvious. Although the model has certain high-value underestimation and low-value overestimation phenomena, the amount of data in this part is small, and the slopes of the fitting line equations at each altitude are around 0.96, with only the slope of 300hPa being the smallest, but it also reaches 0.88. The slopes of the fitting lines at each altitude do not deviate much from the 1:1 line. Overall, the prediction results of the model are good.

[0154] 4. Reconstruction and verification of CO vertical profile data

[0155] (1) Model prediction and filling rules:

[0156] Forecast generation: Daily forecasts are made for each grid point in the estimation dataset to generate ten layers of CO concentration estimates.

[0157] Data fusion rules:

[0158] Priority rule: If original MOPITT data exists for a grid point, the original value is retained; otherwise, the model estimated value is used.

[0159] Outlier processing: Perform range verification on the model estimation results (e.g. trigger manual review when CO concentration at 1000hPa layer is >500 ppbv).

[0160] (2) Output dataset:

[0161] Format and content: Generate daily ten-layer CO vertical profile mixing ratio reconstruction data with a spatial resolution of 0.25°×0.25°, stored as a NetCDF file, including latitude and longitude, time, ten-layer concentration and data source mark (original / estimated).

[0162] Coverage improvement statistics:

[0163] The average daily coverage of the original MOPITT data was <5%, which increased to >15% after reconstruction (300-500 days in the northern region).

[0164] The regional coverage of the Qinghai-Tibet Plateau increased from <1% to ~20%.

[0165] (3) Senior management contribution assessment:

[0166] The height factor is calculated using the following formula:

[0167] ,

[0168] Among them, is the fitting coefficient of the th layer in the multiple regression model, is the th layer CO concentration.

[0169] Decision threshold: When ≥ 0.3, it is determined as a significant area of high-level transportation.

[0170] In another embodiment of the present invention, a CO concentration vertical profile inversion system is provided. This system can be used to implement the above-mentioned CO concentration vertical profile inversion method. Specifically, it includes:

[0171] Data acquisition module: Regularly pull data from the GEE, NASA, and Copernicus platforms.

[0172] Preprocessing module: Call the HARP toolbox to complete resampling and quality filtering.

[0173] Model prediction module: Batch generate reconstructed data based on the trained XGBoost model.

[0174] Visualization output module: Generate the following products:

[0175] Dynamic atlas of the spatio-temporal distribution of the CO vertical profile;

[0176] Spatio-temporal evolution sequence diagram of high-level transportation events;

[0177] Thermal map of the pollution contribution degree of key areas, dividing the pollution source types by different height layers.

[0178] Hardware configuration:

[0179] Server: 64-core CPU, 256GB memory, NVIDIA A100 GPU × 4, supporting large-scale parallel computing.

[0180] Storage system: Distributed storage ≥ 1PB, supporting efficient reading and writing of NetCDF big data.

[0181] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the CO concentration vertical profile inversion method.

[0182] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is the memory device in the terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. And in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.

[0183] One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the CO concentration vertical profile inversion method in the above embodiment; one or more instructions in the computer-readable storage medium are loaded and executed by the processor.

[0184] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0185] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0186] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0188] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the implementation methods of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A method for retrieving the vertical profile of CO concentration based on multi-source remote sensing data fusion, characterized in that, It includes the following steps: Step S1: Through the unary linear regression analysis of the CO concentration at the ground station and the TROPOMI CO column concentration, relevant stations are screened, and the central dataset and the extreme dataset are divided based on the ratio factor; For the extreme dataset, the MOPITT CO vertical profile mixing ratio data is introduced to establish a multiple linear regression model, calculate the height factors of each height layer, and screen the key height layers; The formula of the unary linear regression model is as follows: T i = α i · S i where \(i\) represents the current ground station number, \(T\) i is the TROPOMI CO column concentration value of the satellite grid point corresponding to the ground station, \(S\) i is the CO concentration value monitored by the ground station, \(\alpha\) i is the linear regression fitting parameter; The ratio factor, the formula is as follows: Among them, T t is the temperature at time t, S t is the monitored value of ground CO concentration at time t; The established multiple linear regression model is as follows: Where M is the MOPITT CO column concentration data of the pixel corresponding to the ground station, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 300 hPa, M r is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 900 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 800 hPa to 300 hPa, β i and γ i are the multiple linear regression fitting parameters, where is to unify the dimension level of the equation, and T represents temperature; The calculation formula of the height factor is as follows: Among them, H f is the height factor, and β h is the fitting coefficient of the h-th layer in the multiple regression model, is the CO concentration of the h-th layer, and i is the height layer index, with a value range from 1 to 10; When H f ≥ 0.3, it is determined as a significant area for high-level transportation, and a pollution transmission path map is generated; Step S2: Obtain the TROPOMI satellite's CO column concentration data, the MOPITT satellite's CO vertical profile mixing ratio data, and the ERA5 meteorological reanalysis data; Step S3: Perform spatio-temporal matching and spatial resampling preprocessing on the data in S2 to construct a training dataset; Step S4: Use the XGBoost machine learning algorithm to establish a CO concentration vertical profile inversion model, with the TROPOMI column concentration and ERA5 meteorological parameters as input features and the MOPITT ten-layer vertical profile mixing ratio as the output target for model training; Step S5: Use the trained model to estimate the concentration in the area lacking MOPITT data and generate a CO vertical profile reconstruction dataset with high spatial coverage; Step S6: Integrate the model estimation results and the original MOPITT data and output the complete spatio-temporal distribution data of the CO vertical profile.

2. The CO concentration vertical profile inversion method according to claim 1, characterized in that: In step S2: The TROPOMI data is an L3-level offline product with a spatial resolution of 0.05°×0.05°; The MOPITT data is a combined channel level-3 product, containing 10-layer vertical profile mixing ratio data, with a spatial resolution of 1°×1°; The ERA5 data contains temperature, meridional wind, zonal wind, and vertical velocity parameters of ten pressure layers from 1000 hPa to 100 hPa, with a spatial resolution of 0.25°×0.25°.

3. The method for inverting the vertical profile of CO concentration according to claim 1, wherein: The MOPITT data is resampled to a 0.25°×0.25° grid using the bilinear interpolation method; The TROPOMI data is resampled to a 0.25°×0.25° grid using the area weighting method; The MOPITT data is quality-filtered to retain the valid observation data with a signal degree of freedom (DFS) value ≥ 1.

5.

4. The method for inverting the vertical profile of CO concentration according to claim 1, characterized in that: The training dataset is divided into 5 subsets. Each time, 4 subsets are taken as the training set, and the remaining 1 subset is used as the validation set; The XGBoost hyperparameters are optimized through random search, and the final determined parameter combination is: the number of weak estimators n_estimators = 901, the maximum depth of the tree max_depth = 9, and the learning rate learning_rate = 0.14; An early stopping mechanism is set, and the training is terminated when the performance of the validation set does not improve for 100 consecutive iterations.

5. The CO concentration vertical profile inversion method according to claim 1, wherein: The data reconstruction in step S5 includes: Generate ten-layer CO concentration prediction values for each grid point day by day; Retain the original MOPITT valid observation data and only fill in the missing data area with the model. Perform confidence marking on the estimation results with missing features, and preferentially use the measured values when there is original MOPITT data at the corresponding position.

6. The method for inverting the vertical profile of CO concentration according to claim 1, wherein: The data fusion in step S6 adopts a spatio-temporal continuity verification method: Establish a three-dimensional spatial correlation function for adjacent grid points: Among them, is the model-estimated CO concentration value of the i-th sample, is the original observed CO concentration value of the i-th sample; μ est , μ obs are the means of the estimated value and the observed value, σ est , σ obs are the standard deviations of the estimated value and the observed value, N is the total number of samples, and x, u, t are the spatial coordinates and time indices, representing the calculation results of the correlation coefficient in a specific area and time period; When the R value < 0.7, start the manual review mechanism and verify the rationality of the transmission channel in combination with meteorological field data.

7. A CO concentration vertical profile inversion system, characterized in that: This system can be used to implement the CO concentration vertical profile inversion method described in any one of claims 1 to 6. Specifically, it includes: Data acquisition module: Regularly pull data from the GEE, NASA, and Copernicus platforms; Preprocessing module: Call the HARP toolbox to complete resampling and quality filtering; Model prediction module: Batch generate reconstructed data based on the trained XGBoost model; Visualization output module: Generate the following products: Dynamic atlas of the spatio-temporal distribution of the CO vertical profile; Sequence diagram of the spatio-temporal evolution of high-level transport events; Thermal map of the pollution contribution degree in key areas, and classify pollution source types according to different altitude layers.

8. A computer-readable storage medium, characterized in that: It stores a computer program, and when the program is executed by a processor, it implements the CO concentration vertical profile inversion method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Regional PM2.5 remote sensing inversion model fusing fine particulate matter concentration data

    CN111323352A

  • Remote measurement method for concentration of atmospheric greenhouse gas carbon dioxide and methane foundation column

    CN120009217A