CO concentration vertical profile inversion method and system based on multi-source remote sensing data fusion

Through multi-source remote sensing data fusion and XGBoost machine learning, the problem of insufficient data coverage of atmospheric CO vertical distribution is solved, and high-precision CO vertical profile inversion is achieved, supporting long-term continuous monitoring and cross-region transmission analysis.

CN120234770AActive Publication Date: 2025-07-01CHINA UNIV OF MINING & TECH

Patent Information

Application Number
CN202510727505.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing technology cannot effectively solve the problems of low data coverage, lack of vertical information and insufficient model accuracy of atmospheric carbon monoxide (CO) vertical distribution, especially in complex terrain and cloudy areas, it is difficult to achieve high-time, long-term continuous pollution monitoring and quantitative analysis of cross-region transmission mechanisms.

Method used

The multi-source remote sensing data fusion method is adopted, and TROPOMI, MOPITT satellite data and ERA5 meteorological data are used to screen the correlation site through unary linear regression and multivariate linear regression. Combined with the XGBoost machine learning algorithm, a CO concentration vertical profile inversion model is established to fill in data loss and improve model accuracy.

Benefits of technology

High-precision estimation of CO vertical profiles is achieved, monitoring capabilities of complex terrain and cloudy areas are enhanced, long-term continuous observation is supported, and contribution ratios of different height layers and cross-regional transmission paths of pollutants are accurately quantified, revealing the impact of meteorological conditions on CO distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234770A_ABST
    Figure CN120234770A_ABST
Patent Text Reader

Abstract

The invention discloses a CO concentration vertical profile inversion method and system based on multi-source remote sensing data fusion. The method comprises the steps that CO column concentration data of a TROPOMI satellite, CO vertical profile mixing ratio data of an MOPITT satellite and ERA5 meteorological reanalysis data are acquired; constructing a training data set through space-time matching and space resampling; an XGBoost machine learning algorithm is used for establishing an estimation model with the TROPOMI column concentration and the ERA5 meteorological parameters as input and the MOPITT ten-layer vertical profile mixing ratio as an output target; and carrying out concentration estimation and data filling on the missing region by adopting the trained model, and generating a CO vertical profile reconstruction data set with high space coverage. According to the method, through multi-source data fusion and intelligent algorithm modeling, the space-time continuity and precision of CO vertical distribution data are remarkably improved, the technical problems that traditional satellite data coverage is insufficient and vertical information is lost are solved, and high-reliability data support is provided for air pollution traceability, cross-regional transmission analysis and environmental governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of environmental science and artificial intelligence, and particularly relates to a method and system for inverting the vertical profile of CO concentration based on multi-source remote sensing data fusion. Background Art

[0002] Atmospheric carbon monoxide (CO) is an important trace gas in the Earth's atmosphere. Its sources mainly include incomplete combustion of fossil fuels, biomass burning, and non-combustion processes (such as oxidation of hydrocarbons). By consuming OH free radicals in the atmosphere, CO significantly affects the oxidation capacity of the atmosphere and participates in photochemical reactions as a precursor of ozone, indirectly exacerbating the greenhouse effect. Traditional monitoring methods (such as ground spectrometers and aircraft surveys) have problems such as limited spatial coverage, high cost, and difficulty in long-term continuous observation, and are difficult to meet the requirements of large-scale and high-timeliness pollution monitoring.

[0003] The vertical distribution of atmospheric carbon monoxide (CO) is a key parameter for studying atmospheric chemical processes, pollution transmission mechanisms, and climate effects. The vertical profile information of CO can reveal the following core scientific issues: 1. Source apportionment: CO emitted from the ground (such as industry and transportation) is mainly distributed in the near-surface layer, while CO generated by events such as biomass burning and sandstorms can affect the middle atmosphere through uplift, and even be transported to the stratosphere over long distances; 2. Chemical process research: The reaction rate of CO with OH free radicals varies significantly with height. The oxidation capacity of the boundary layer is weaker than that of the free convective layer, resulting in different CO lifetimes at different heights (about 1 month near the ground and up to 2 - 3 months at high altitudes); 3. Climate effect assessment: High-altitude CO (300 hPa - 100 hPa) contributes 3 - 5 times more to radiative forcing than near the ground by affecting stratospheric ozone formation.

[0004] In recent years, satellite remote sensing technology has provided new solutions for CO monitoring. Among them: MOPITT (Measurements of Pollution in the Troposphere) satellite: It can provide the mixing ratio data of the 10-layer CO vertical profile from the surface (1000 hPa) to the upper troposphere (100 hPa). However, affected by factors such as satellite orbit gaps and cloud cover, its daily average spatial coverage is less than 5%. Especially in the Qinghai-Tibet Plateau and cloudy areas, the data is severely missing, seriously restricting the research on the distribution and transmission process of high-altitude CO.

[0005] Existing research shows that the coverage rate of MOPITT data in complex terrain areas is less than 1%, resulting in insufficient monitoring capabilities for sudden pollution events such as biomass burning.

[0006] TROPOMI (TROPOspheric Monitoring Instrument) satellite: It has a high spatial resolution (0.05°×0.05°) and the ability to cover the globe daily, but only provides the total column concentration of CO, lacks vertical stratification information, and cannot distinguish between ground and high-altitude contributions.

[0007] In extreme events (such as cross-border biomass burning), it is difficult to accurately quantify the contribution of high-altitude transport to ground pollution due to the inability of TROPOMI data to resolve the vertical distribution.

[0008] Limitations of existing estimation methods: Traditional statistical models (such as simple linear regression) only rely on the simple correlation between TROPOMI column concentration and ground monitoring values, without integrating meteorological parameters and high-altitude information, resulting in significant errors in the model during vertical transport events.

[0009] Existing technologies have not proposed an effective method to quantify the contribution of high-altitude CO and cannot reveal the cross-regional transport mechanism of pollutants.

[0010] Insufficient utilization of meteorological data: Key parameters (such as vertical velocity, temperature) in meteorological reanalysis data such as ERA5 have not been fully integrated into the model, resulting in insufficient ability to dynamically characterize the vertical diffusion and chemical reactions of CO.

[0011] Summary of existing technical defects: Low data coverage: There are a large number of temporal and spatial gaps in MOPITT vertical profile data, making it difficult to support high-resolution research.

[0012] Lack of vertical information: TROPOMI only provides the total column concentration and cannot resolve the pollution contributions of different altitude layers.

[0013] Insufficient model accuracy: Traditional statistical methods do not integrate multi-source data and have significant errors under complex meteorological conditions.

[0014] Unclear high-altitude transport mechanism: Lack of a quantitative analysis tool for the high-altitude CO transport path and contribution ratio. Summary of the Invention

[0015] In view of the defects of the existing technology, the present invention provides a method and system for retrieving the vertical profile of CO concentration based on the fusion of multi-source remote sensing data.

[0016] In order to achieve the above invention objectives, the technical solutions adopted by the present invention are as follows: A method for retrieving the vertical profile of CO concentration based on the fusion of multi-source remote sensing data, which includes the following steps: Step S1: Through the unary linear regression analysis of the CO concentration at the ground station and the TROPOMI CO column concentration, screen the relevant stations, and divide the central dataset and the extreme dataset based on the ratio factor; Introduce the MOPITT CO vertical profile mixing ratio data into the extreme dataset, establish a multiple linear regression model, calculate the height factors of each height layer, and screen the key height layers; Step S2: Obtain the CO column concentration data of the TROPOMI satellite, the CO vertical profile mixing ratio data of the MOPITT satellite, and the ERA5 meteorological reanalysis data; Step S3: Perform preprocessing of spatio-temporal matching and spatial resampling on the data in S2, and construct a training dataset containing TROPOMI column concentration, ERA5 meteorological parameters, and MOPITT vertical profile data; Step S4: Use the XGBoost machine learning algorithm to establish a CO concentration vertical profile inversion model, and use the TROPOMI column concentration and ERA5 meteorological parameters as input features, and the MOPITT ten-layer vertical profile mixing ratio as the output target for model training; Step S5: Use the trained model to estimate the concentration in the area where MOPITT data is missing, and generate a CO vertical profile reconstruction dataset with high spatial coverage; Step S6: Integrate the model estimation results with the original MOPITT data, fill in the missing areas, and output the complete spatio-temporal distribution data of the CO vertical profile.

[0017] Further, in step S1, the formula of the unary linear regression model is as follows: , In the formula, represents the current ground station number, is the TROPOMI CO column concentration value of the satellite grid corresponding to the ground station, is the CO concentration value monitored at the ground station, is the linear regression fitting parameter.

[0018] The ratio factor, the formula is as follows: , Among them, time is the temperature at time time is the ground CO concentration monitoring value at time The established multiple linear regression model is as follows: , In the formula, is the MOPITT CO column concentration data of the corresponding pixel points of the ground station, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 300 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 900 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 800 hPa to 300 hPa, and are the multiple linear regression fitting parameters, where is to unify the dimension level of the equation, represents temperature.

[0019] Furthermore: The calculation formula of the height factor in S1 is as follows: , where is the height factor, is the fitting coefficient of the th layer in the multiple regression model, is the CO concentration of the th layer, is the height layer index, and the value range is from 1 to 10.

[0020] When ≥ 0.3, it is determined as a significant area of high-level transportation, and a pollution transmission path map is generated.

[0021] Furthermore, in the step S2: The TROPOMI data is an L3-level offline product with a spatial resolution of 0.05°×0.05°; The MOPITT data is a combined channel (TIR + NIR) level-3 product, containing 10 layers of vertical profile mixing ratio data with a spatial resolution of 1°×1°; The ERA5 data contains temperature, meridional wind, zonal wind, and vertical velocity parameters of ten pressure layers from 1000 hPa to 100 hPa, with a spatial resolution of 0.25°×0.25°.

[0022] Furthermore, the MOPITT data is resampled to a 0.25°×0.25° grid using bilinear interpolation; The TROPOMI data is resampled to a 0.25°×0.25° grid using area weighting; The MOPITT data is quality-filtered, and the effective observation data with a signal degree of freedom (DFS) value ≥ 1.5 is retained.

[0023] Furthermore, the training dataset is divided into 5 subsets. Each time, 4 subsets are taken as the training set, and the remaining 1 subset is taken as the validation set; Optimize the hyperparameters of XGBoost through random search, and finally determine the parameter combination as follows: the number of weak estimators n_estimators = 901, the maximum depth of the tree max_depth = 9, and the learning rate learning_rate = 0.14; Set the early stopping mechanism to terminate the training when the performance of the validation set does not improve for 100 consecutive iterations.

[0024] Furthermore, the data reconstruction in step S5 includes: Generate ten-layer CO concentration prediction values for each grid point day by day; Retain the original MOPITT valid observation data and only fill in the missing data areas with the model; Mark the confidence level of the estimation results with missing features, and give priority to the measured values when there is original MOPITT data at the corresponding position.

[0025] Furthermore, the data fusion in step S6 adopts the spatio-temporal continuity verification method: Establish a three-dimensional spatial correlation function for adjacent grid points: , where, is the model-estimated CO concentration value of the i-th sample, is the original observed CO concentration value of the i-th sample. is the mean of the estimated value and the observed value, is the standard deviation of the estimated value and the observed value, is the total number of samples, is the spatial coordinate (latitude and longitude) and time index, representing the calculation result of the correlation coefficient in a specific area and time period.

[0026] When the value < 0.7, start the manual review mechanism and verify the rationality of the transmission channel in combination with the meteorological field data.

[0027] The present invention also discloses a CO concentration vertical profile inversion system, which can be used to implement the above-mentioned CO concentration vertical profile inversion method. Specifically, it includes: Data acquisition module: Regularly pull data from the GEE, NASA, and Copernicus platforms.

[0028] Preprocessing module: Call the HARP toolbox to complete resampling and quality filtering.

[0029] Model prediction module: Batch generate reconstructed data based on the trained XGBoost model.

[0030] Visualization output module: Generate the following products: Dynamic Atlas of the Spatiotemporal Distribution of CO Vertical Profiles Sequence Diagram of the Spatiotemporal Evolution of High-Level Transport Events Thermal Map of Pollution Contribution Degrees in Key Areas, Dividing Pollution Source Types by Different Altitude Layers

[0031] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above CO concentration vertical profile inversion method is implemented.

[0032] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above CO concentration vertical profile inversion method is implemented.

[0033] Compared with the prior art, the advantages of the present invention are as follows: Through the fusion of multi-source satellite data and machine learning modeling, the present invention overcomes the deficiencies of traditional monitoring technologies and realizes the following core advantages: 1. Significantly fill the data gaps caused by satellite orbit gaps and cloud cover, enhance the monitoring ability in complex terrain and cloudy areas, and support long-term continuous observation and analysis.

[0034] 2. Adopt machine learning algorithms to establish a non-linear mapping relationship, realize high-precision estimation of CO vertical profiles, and the model performance is stable and reliable, superior to traditional statistical methods.

[0035] 3. Propose quantitative indicators to effectively analyze the contribution ratio of different altitude layers to the total column concentration, and accurately identify the cross-regional transmission paths and vertical diffusion mechanisms of pollutants.

[0036] 4. Integrate satellite data and meteorological parameters to reveal the influence of meteorological conditions (such as wind fields, temperatures) on the CO distribution, and support the analysis and tracing of the formation mechanisms of pollution events. Description of the Drawings

[0037] Figure 1 is the flowchart of the CO concentration vertical profile inversion method according to the embodiment of the present invention; Figure 2 is the scatter plot of the prediction results of the five-fold cross-validation test set of the XGBoost model according to the embodiment of the present invention. Detailed Embodiments

[0038] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further elaborates on the present invention with reference to the drawings and by way of examples.

[0039] As Figure 1 shown, a method for inverting CO concentration vertical profiles based on the fusion of multi-source remote sensing data includes the following steps: The formula of the unary linear regression model is as follows: , In the formula, represents the current ground station number, is the TROPOMI CO column concentration value of the satellite grid point corresponding to the ground station, is the CO concentration value monitored by the ground station, is the linear regression fitting parameter.

[0040] The ratio factor is as follows: , where time is the temperature at time time is the monitored ground CO concentration value at time The established multiple linear regression model is as follows: , In the formula, is the MOPITT CO column concentration data of the pixel point corresponding to the ground station, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 300 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 900 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 800 hPa to 300 hPa, and are the multiple linear regression fitting parameters. Among them is to unify the dimension level of the equation, represents the temperature.

[0041] Furthermore: The calculation formula of the height factor in S1 is as follows: , where is the height factor, is the fitting coefficient of the th layer in the multiple regression model, is the th layer CO concentration, is the height layer index, and the value range is from 1 to 10.

[0042] When ≥ 0.3, it is determined as a significant area of high-level transportation, and a pollution transmission path map is generated.

[0043] I. Data Verification and Feature Screening (1)Verify the relationship between ground and column concentration by simple linear regression: , where, represents the current ground station number, is the TROPOMI CO column concentration value of the satellite grid corresponding to the ground station, is the CO concentration value monitored at the ground station, are the simple linear regression fitting parameters.

[0044] Implementation operation: Establish a simple linear regression model for each ground station (a total of 1577 stations), calculate and the correlation coefficient R, and select the stations with R≥0.6 (such as 218 stations including Beijing and Tianjin).

[0045] Divide the dataset by the ratio factor: , where, time is the temperature at time time is the monitored ground CO concentration value at time Implementation operation: Set the threshold range 0.8≤ ≤1.2, and divide the data into: Central dataset ( within the threshold, a total of 143 stations): directly used for model training; Extreme dataset ( beyond the threshold, a total of 144 stations): vertical stratification information needs to be introduced for correction.

[0046] Quantify the contribution of vertical stratification by multiple linear regression: , where, is the MOPITT CO column concentration data of the pixel point corresponding to the ground station, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 300 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 900 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 800 hPa to 300 hPa, and are the multiple linear regression fitting parameters. Among them is to unify the dimension level of the equation, represents the temperature.

[0047] Implementation operation: For extreme datasets, calculate the height factor and filter stations with a value ≥ 0.3 (such as regions in Henan, Shandong, etc.), indicating a significant contribution from upper-air transport; Retain the corresponding height layer (800 - 300 hPa) as the key input feature for the subsequent XGBoost model.

[0048] II. Data Preparation and Preprocessing (1) Data Sources and Acquisition: TROPOMI data: Obtain L3-level CO column concentration data (spatial resolution 0.05° × 0.05°) from July 1, 2018, to December 27, 2020, through the Google Earth Engine platform, and filter valid data with a QA value greater than 0.5. The data format is NetCDF, and the time resolution is daily mean.

[0049] MOPITT data: Download MOPITT V8 combined-channel (TIR + NIR) CO vertical profile mixing ratio data (spatial resolution 1° × 1°) from the NASA official website, which includes 10 layers of data from 1000 hPa to 100 hPa. Use a threshold of DFS ≥ 1.5 for quality filtering.

[0050] ERA5 meteorological data: Download hourly meteorological parameters from 1000 hPa to 100 hPa from the Copernicus platform. As shown in Table 1, it includes temperature (t), meridional wind (u), zonal wind (v), and vertical velocity (w), with a spatial resolution of 0.25° × 0.25°.

[0051] Table 1 ERA5 Meteorological Parameters

[0052] (2) Temporal and Spatial Matching and Resampling: Temporal matching: Unify the TROPOMI and MOPITT data to the daily mean from 9:00 - 15:00 (UTC time 1:00 - 7:00), which is matched with the satellite overpass time. The ERA5 data takes the average value of the corresponding period.

[0053] Spatial resampling: MOPITT data: Resample the 1° × 1° data to a 0.25° × 0.25° grid using bilinear interpolation.

[0054] TROPOMI data: Use the HARP toolbox for area-weighted resampling to generate 0.25° × 0.25° grid data.

[0055] Data format conversion: Store the processed data as a CSV file by date. Each row contains longitude and latitude, TROPOMI column concentration, ERA5 meteorological parameters (10 layers × 4 parameters), and MOPITT ten-layer CO concentration.

[0056] (3)Dataset division (the statistical results are shown in the site quantity distribution in Table 2): Training dataset: Screen grid points with complete TROPOMI, MOPITT, and ERA5 data simultaneously, with a total of 616,324 samples.

[0057] Estimation dataset: Include all spatio-temporally matched grid points, allowing partial feature missing (such as missing MOPITT data), for model prediction filling.

[0058] Table 2 Number of sites for accuracy evaluation of the simple linear regression model

[0059] III. XGBoost Model Construction and Training (1)Input features and target variables (parameter definitions are shown in Table 1): Feature variables (41 dimensions): TROPOMI CO column concentration (1 dimension).

[0060] ERA5 meteorological parameters: Temperature (t), zonal wind (u), meridional wind (v), and vertical velocity (w) for each layer, a total of 10 layers × 4 parameters = 40 dimensions.

[0061] Target variable: MOPITT ten-layer CO vertical profile mixing ratio (from 1000 hPa to 100 hPa, 10 dimensions).

[0062] (2)XGBoost model configuration: Algorithm selection: Adopt the XGBoost regression model, with the loss function being the mean squared error (MSE), and the goal being to minimize the deviation between the predicted value and the measured value.

[0063] Hyperparameter optimization: Search space: n_estimators is an integer between 100 and 1000; max_depth is an integer between 3 and 10, and learning_rate is a decimal between 0.001 and 0.3.

[0064] Random search result: Finally, determine the optimal parameters as n_estimators being 901, max_depth being 9, and learning_rate being 0.14.

[0065] Training strategy: Five - fold cross - validation: Randomly divide the training set into 5 subsets. 4 subsets are used for training and 1 for validation, and repeat this 5 times to take the average result.

[0066] Early stopping mechanism: Terminate the training when the RMSE of the validation set shows no improvement for 100 consecutive iterations to prevent overfitting.

[0067] (3)Model performance evaluation: In this embodiment, in addition to using the XGBoost model to construct the CO different high - layer concentration estimation model, a common linear regression model (LR) is also used as a control group to verify the performance of the XGBoost model. Table 3 shows the five - fold cross - test set validation results of different algorithm models. It can be seen that the CO different high - layer concentration estimation model constructed based on XGBoost shows better performance. The overall R 2 of the XGBoost model is 0.96, the RMSE is 5.79 ppbv, and the MAE is 3.81 ppbv, with relatively high accuracy; while the overall R 2 of the LR model is only 0.28, the RMSE is 23.74 ppbv, and the MAE is 18.32 ppbv, and the effect is far less than that of the XGBoost model. From the validation results of different height layers, the R 2 of the XGBoost model ranges from 0.92 to 0.99, with high and relatively stable accuracy. The RMSE ranges from 0.94 to 15.65 ppbv, and the MAE ranges from 0.65 to 10.03 ppbv. Due to the characteristics of different high - layer datasets, the RMSE and MAE calculated according to the concentration values both decrease with the increase of height, and the values are within the acceptable range; the R 2 of the LR model ranges from 0.02 to 0.58, with poor and very unstable accuracy. The RMSE ranges from 12.46 to 57.98 ppbv, and the MAE ranges from 9.84 to 42.94 ppbv, with large values and a wide range of changes. Based on the above, it can be seen that the CO different high - layer concentration estimation model constructed by XGBoost has good generalization ability, the model prediction results have relatively high accuracy, and both the RMSE and MAE are relatively low, the model error is small, showing good performance.

[0068] Table 3 Five - fold cross - validation results of different algorithm models

[0069] Figure 2It is a scatter plot of the prediction results of the XGBoost model on the five-fold cross-validation test set. The scatter plot can intuitively reflect the data distribution of the model's predicted values and the true values, as well as the gap between the two. The abscissa is the true value, that is, the mixing ratio concentration of the CO vertical profile at different altitude levels observed by MOPITT; the ordinate is the mixing ratio concentration of the CO vertical profile at different altitude levels predicted by the model, and the unit is ppbv for both; the color reflects the density of the data points at that coordinate, and from blue to red indicates an increase in data points. The black line in the figure is the 1:1 line with a slope of 1. If the data points exactly fall on this line, it means that the model's predicted value is equal to the observed value, which is the most ideal situation. The red line in the figure is the best-fit line. It accurately reflects the actual relationship between the model's predicted value and the observed value by minimizing the vertical distance from all sample points to it. The closer the fit line is to the 1:1 line, the better the correlation between the model's predicted value and the observed value. From the data distribution of the scatter plot of the model's prediction results, the red data points in each layer are concentrated in the middle or the lower left area, indicating that the mixing ratio concentration of the CO vertical profile at different altitude levels is at or slightly below the average level of that layer. For this part of the data, the red fit line is very close to the black 1:1 line, and the red data points are also concentrated around the fit line, indicating that the predicted value and the observed value of the data model are relatively close here, the model accuracy is high, and at the same time, the prediction effect of the model at different altitude levels does not show a particularly large deviation, and the stability of the model is good. It can also be seen from the scatter plot that when the mixing ratio concentration of the CO vertical profile at different altitude levels is at a high value of that layer, the fit line is below the 1:1 line, indicating that the model is prone to underestimation when the CO vertical profile has a high mixing ratio concentration. This phenomenon exists in all altitude levels. The underestimation phenomenon at the 1000 hPa altitude level mainly occurs when the mixing ratio concentration of the CO vertical profile is higher than about 450 ppbv; the underestimation phenomenon at the 900 hPa altitude level mainly occurs when the concentration is higher than about 400 ppbv; the underestimation phenomenon at the 800 hPa altitude level mainly occurs when the concentration is higher than about 250 ppbv; the underestimation phenomenon at the 700 hPa altitude level mainly occurs when the concentration is higher than about 200 ppbv; the underestimation phenomenon at the 600 hPa - 500 hPa altitude levels mainly occurs when the concentration is higher than about 130 ppbv; the underestimation phenomenon at 400 hPa - 300 hPa mainly occurs when the concentration is higher than about 125 ppbv; the underestimation phenomenon at 200 hPa mainly occurs when the concentration is higher than about 140 ppbv; the underestimation phenomenon at 100 hPa is not obvious.When the mixing ratio concentration of the CO vertical profile is at a low value in different altitude layers, the fitting line is above the 1:1 line, indicating that the model is prone to overestimation when the mixing ratio of the CO vertical profile is at a low concentration. This overestimation phenomenon is not obvious in the altitude layer of 1000 hPa - 600 hPa. In the altitude layer of 500 hPa - 300 hPa, when the mixing ratio concentration of the CO vertical profile is lower than about 60 ppbv, the overestimation phenomenon will occur; the overestimation phenomenon at 200 hPa mainly appears when the concentration is lower than about 45 ppbv, and the overestimation phenomenon in the 100 hPa altitude layer is not obvious. Although there are certain phenomena of underestimation of high values and overestimation of low values in the model, the amount of data in this part is small, and the slopes of the fitting line equations for each altitude layer are all around 0.96. Only the slope at 300 hPa is the smallest, but it also reaches 0.88. The degree of deviation of the fitting line slopes for each altitude layer from the 1:1 line is not large. Generally speaking, the prediction results of the model are good.

[0070] IV. Reconstruction and Verification of CO Vertical Profile Data (1) Model Prediction and Filling Rules: Prediction Generation: Perform daily predictions for each grid point in the estimation dataset to generate estimated values of CO concentration for ten layers.

[0071] Data Fusion Rules: Priority Rule: If there is original MOPITT data for a certain grid point, retain the original value; otherwise, use the model estimated value.

[0072] Outlier Processing: Conduct range verification on the model estimation results (e.g., manual review is triggered when the CO concentration in the 1000 hPa layer > 500 ppbv).

[0073] (2) Output Dataset: Format and Content: Generate daily reconstructed data of the mixing ratio of the CO vertical profile for ten layers, with a spatial resolution of 0.25°×0.25°, stored as a NetCDF file, including longitude, latitude, time, ten-layer concentration, and data source markers (original / estimated).

[0074] Coverage Enhancement Statistics: The daily average coverage of the original MOPITT data < 5%, and after reconstruction, it is increased to > 15% (the northern region is increased by 300 - 500 days).

[0075] The coverage rate in the Qinghai-Tibet Plateau region is increased from < 1% to ~20%.

[0076] (3) Evaluation of the Contribution Degree of High Layers: Calculation of the altitude factor, the formula is as follows: , where, is the Fitting coefficient of the layer is the CO concentration of the

[0077] Determination threshold: When ≥ 0.3, it is determined as a significant area of high-level transportation.

[0078] In another embodiment of the present invention, a CO concentration vertical profile inversion system is provided. This system can be used to implement the above-mentioned CO concentration vertical profile inversion method. Specifically, it includes: Data acquisition module: Regularly pull data from GEE, NASA, and Copernicus platforms.

[0079] Preprocessing module: Call the HARP toolbox to complete resampling and quality filtering.

[0080] Model prediction module: Batch generate reconstructed data based on the trained XGBoost model.

[0081] Visualization output module: Generate the following products: Dynamic atlas of the spatio-temporal distribution of the CO vertical profile; Sequence diagram of the spatio-temporal evolution of high-level transportation events; Thermal map of the pollution contribution degree of key areas, dividing the pollution source types by different height layers.

[0082] Hardware configuration: Server: 64-core CPU, 256GB memory, NVIDIA A100 GPU × 4, supporting large-scale parallel computing.

[0083] Storage system: Distributed storage ≥ 1PB, supporting efficient reading and writing of NetCDF big data.

[0084] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the CO concentration vertical profile inversion method.

[0085] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is the memory device in the terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.

[0086] One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the CO concentration vertical profile inversion method in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor.

[0087] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0088] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0089] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0090] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0091] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the implementation methods of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A method for retrieving the vertical profile of CO concentration based on multi-source remote sensing data fusion, characterized in that It includes the following steps: Step S1: Through the unary linear regression analysis of the CO concentration at the ground station and the TROPOMI CO column concentration, the relevant stations are screened, and the central dataset and the extreme dataset are divided based on the ratio factor; For the extreme dataset, the MOPITT CO vertical profile mixing ratio data is introduced, a multiple linear regression model is established, the height factors of each height layer are calculated, and the key height layers are screened; Step S2: Obtain the TROPOMI satellite's CO column concentration data, the MOPITT satellite's CO vertical profile mixing ratio data, and the ERA5 meteorological reanalysis data; Step S3: Perform spatio-temporal matching and spatial resampling preprocessing on the data in S2 to construct a training dataset; Step S4: Use the XGBoost machine learning algorithm to establish a CO concentration vertical profile inversion model. The TROPOMI column concentration and ERA5 meteorological parameters are used as input features, and the MOPITT ten-layer vertical profile mixing ratio is used as the output target for model training; Step S5: Use the trained model to estimate the concentration in the area where MOPITT data is missing, and generate a CO vertical profile reconstruction dataset with high spatial coverage; Step S6: Integrate the model estimation results and the original MOPITT data, and output the complete spatio-temporal distribution data of the CO vertical profile.

2. The method for inverting the vertical profile of CO concentration according to claim 1, characterized in that: In step S1, the formula of the unary linear regression model is as follows: , In the formula, represents the current ground station number, is the TROPOMI CO column concentration value of the satellite grid corresponding to the ground station, is the CO concentration value monitored by the ground station, is the linear regression fitting parameter; The said ratio factor, the formula is as follows: , Among them, Time The temperature at that time, Time The monitored value of ground CO concentration at that time; The established multiple linear regression model is as follows: , In the formula, is the MOPITT CO column concentration data of the pixel corresponding to the ground station, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 300 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 1000 hPa to 900 hPa, is the value of the MOPITT CO vertical profile mixing ratio from 800 hPa to 300 hPa, and are the multiple linear regression fitting parameters, where is to unify the dimension level of the equation, represents temperature.

3. The method for inverting the vertical profile of CO concentration according to claim 1, characterized in that: The calculation formula of the height factor in S1 is as follows: , Among them, is the height factor, is the fitting coefficient of the th layer in the multiple regression model, is the CO concentration of the th layer, is the height layer index, and its value range is from 1 to 10; When ≥ 0.3, it is determined as a significant area for high-level transportation, and a pollution transmission path map is generated.

4. The method for inverting the vertical profile of CO concentration according to claim 1, characterized in that: In step S2: The TROPOMI data is an L3-level offline product with a spatial resolution of 0.05°×0.05°; The MOPITT data is a combined channel level-3 product, including 10-layer vertical profile mixing ratio data with a spatial resolution of 1°×1°; The ERA5 data includes temperature, meridional wind, zonal wind, and vertical velocity parameters of ten pressure layers from 1000 hPa to 100 hPa, with a spatial resolution of 0.25°×0.25°.

5. The CO concentration vertical profile inversion method according to claim 1, characterized in that: The MOPITT data is resampled to a 0.25°×0.25° grid by the bilinear interpolation method; The TROPOMI data is resampled to a 0.25°×0.25° grid by the area weighting method; Quality filtering is performed on the MOPITT data, and the effective observation data with a signal degree of freedom (DFS) value ≥1.5 is retained.

6. The method for inverting the vertical profile of CO concentration according to claim 1, characterized in that: The training dataset is divided into 5 subsets. Each time, 4 subsets are taken as the training set, and the remaining 1 subset is taken as the validation set; The XGBoost hyperparameters are optimized through random search, and the final determined parameter combination is: the number of weak estimators n_estimators = 901, the maximum depth of the tree max_depth = 9, and the learning rate learning_rate = 0.14; Set an early stopping mechanism to terminate the training when the performance of the validation set does not improve for 100 consecutive iterations.

7. The CO concentration vertical profile inversion method according to claim 1, characterized in that: The data reconstruction in step S5 includes: Generate ten-layer CO concentration prediction values for each grid point day by day; Retain the original MOPITT effective observation data, and only fill in the missing data area with the model; Confidence marking is performed on the estimation results with missing features, and the measured values are preferentially used when there is original MOPITT data at the corresponding positions.

8. The CO concentration vertical profile inversion method according to claim 1, characterized in that: The data fusion in step S6 adopts a spatio-temporal continuity verification method: A three-dimensional spatial correlation function is established for adjacent grid points: , Among them, is the model-estimated CO concentration value of the i-th sample, is the original observed CO concentration value of the i-th sample; is the mean of the estimated value and the observed value, is the standard deviation of the estimated value and the observed value, is the total number of samples, is the spatial coordinate and time index, representing the calculation result of the correlation coefficient in a specific area and time period; When When the value < 0.7, start the manual review mechanism and verify the rationality of the transmission channel in combination with meteorological field data.

9. A CO concentration vertical profile inversion system, characterized in that: This system can be used to implement the CO concentration vertical profile inversion method described in any one of claims 1 to 8. Specifically, it includes: Data acquisition module: Regularly pull data from the GEE, NASA, and Copernicus platforms; Preprocessing module: Call the HARP toolbox to complete resampling and quality filtering; Model prediction module: Batch generate reconstructed data based on the trained XGBoost model; Visualization output module: Generate the following products: Dynamic atlas of the spatio-temporal distribution of the CO vertical profile; Sequence diagram of the spatio-temporal evolution of high-level transport events; Thermal map of the pollution contribution degree in key areas, dividing the pollution source types according to different altitude layers.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, it implements the CO concentration vertical profile inversion method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Regional PM2.5 remote sensing inversion model fusing fine particulate matter concentration data

    CN111323352A

  • Satellite data temperature profile inversion method based on generalized ensemble learning

    CN116186486A

  • Remote measurement method for concentration of atmospheric greenhouse gas carbon dioxide and methane foundation column

    CN120009217A

Cited By

  • Cross-border area PM2.5 three-dimensional reconstruction method and device

    CN122134950A

  • A method and apparatus for three-dimensional reconstruction of PM2.5 in cross-border areas

    CN122134950B