XCO2 satellite product bias correction method and system based on multi-source ground-based data

By using time-series decomposition and adaptive weighting mechanisms of multi-source ground-based data, combined with the XGBoost model and SHAP analysis, the systematic bias and regional adaptability issues of satellite XCO2 products were resolved, achieving high-precision bias correction and improving data utilization and spatial coverage.

CN121659282BActive Publication Date: 2026-07-31CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNIV OF MINING & TECH
Filing Date
2025-12-10
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing XCO2 products retrieved from satellites suffer from systematic biases, regional differences, and nonlinear errors that are difficult to characterize. Machine learning models lack temporal decomposition, regional adaptability, and interpretability, resulting in low data utilization and insufficient spatial coverage.

Method used

A closed-loop framework combining multi-source ground-based data with temporal decomposition, machine learning, and adaptive weighting is adopted. Systematic errors are identified through state-space decomposition, bias correction is performed using the XGBoost model, and factor contribution is analyzed through SHAP to dynamically adjust regional weights to improve model adaptability and interpretability.

Benefits of technology

This improves the accuracy and consistency of XCO2 products, reduces systemic bias, and enables more accurate reflection of regional carbon source and sink distribution, providing high-quality data for global carbon budget assessment and climate model validation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121659282B_ABST
    Figure CN121659282B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for correcting XCO2 satellite product bias based on multi-source ground-based data, belonging to the field of remote sensing and atmospheric environment monitoring technology. First, XCO2 data is acquired and combined with multi-source auxiliary factors for spatial and temporal registration to construct a matching sample set. Based on a state-space decomposition model, the bias sequence between satellite and benchmark observations is decomposed into long-term trend terms, seasonal periodic terms, and random residual terms to characterize systematic errors and random disturbances. Regional labels are generated based on regional ecological and geographical characteristics. Subsequently, a bias correction model is established using machine learning algorithms, and the fitting effect in high-error areas is optimized through regional weighting. Finally, a corrected high-precision XCO2 product is output. This invention can significantly reduce systematic biases caused by atmospheric, surface, and sensor characteristics in satellite observations, improving the consistency and scientific usability of XCO2 products at regional and global scales.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing and atmospheric environment monitoring technology, and particularly relates to a method and system for XCO2 satellite product deviation correction based on multi-source ground-based data. It integrates multi-source factors and uses interpretable machine learning methods to perform deviation diagnosis and adaptive correction on satellite-retrieved XCO2 products. Background Technology

[0002] Carbon dioxide (CO2) is one of the main greenhouse gases caused by Earth's radiative forcing, and its concentration changes have a significant impact on the global climate system. For a long time, the scientific community has conducted high-precision in-situ CO2 measurements through ground-based benchmark observation networks (such as the Total Carbon Column Observing Network, TCCON), providing crucial evidence for understanding the terrestrial and marine carbon cycles. However, due to the sparse distribution of ground stations and limitations imposed by geographical and climatic conditions, their spatial representativeness is limited, making it difficult to comprehensively reflect the dynamic changes in carbon sources and sinks at global and regional scales.

[0003] With the advent of the space remote sensing era, satellite-based CO2 detection technology has developed rapidly, significantly enhancing the ability to observe the global carbon cycle. In 2002, the SCIAMACHY and AIRS satellites achieved the first global spectral measurement of atmospheric CO2. Subsequently, dedicated greenhouse gas observation satellites such as GOSAT (2009) and OCO-2 (2014) were launched, providing continuous and systematic data support for monitoring the global CO2 distribution. The column-mean dry air mole fraction (XCO2) obtained by satellite remote sensing is playing an increasingly important role in global carbon cycle research, climate change assessment, and anthropogenic emissions monitoring.

[0004] However, due to various factors such as instrument calibration errors, the incompleteness of atmospheric radiative transfer models, the complex influence of aerosols and clouds, and differences in surface reflectance characteristics, XCO2 products obtained from satellite inversion often contain systematic biases and random errors. For example, NASA's ACOS (Atmospheric CO2 Observations from Space) inversion algorithm inevitably generates errors related to co-inversion parameters (such as aerosol optical thickness, surface reflectance, and air pressure) when solving for XCO2 using Bayesian optimal estimation. If these biases are not addressed, they will affect the accuracy and stability of global carbon balance inversion.

[0005] To suppress systematic bias, existing technologies generally employ bias correction methods based on multiple linear regression (MLR). This method utilizes the empirical linear relationship between XCO2 and co-inversion parameters to linearly fit and correct the inversion results, thereby reducing overall bias. For example, both OCO-2 and OCO-3 satellite products use MLR models combined with empirical threshold quality control (Quality Filter) for systematic error correction. This method can significantly reduce the average bias of XCO2 and improve its consistency with ground-based TCCON data.

[0006] However, the MLR method inherently assumes a linear relationship between the bias and various influencing factors, and is only effective within a specific feature space during model training. When there is nonlinear coupling between the state parameters and the bias, or when regional differences are significant, the linear model struggles to fully capture the complex structure of the bias. Furthermore, to ensure the linear relationship holds, strict quality thresholds are typically set in the operational process, leading to the rejection of a large number of samples, thus reducing data utilization and spatial coverage. For special regions such as cities, polar regions, or areas with high cloud cover, a uniform global linear correction strategy often fails to effectively adapt to local characteristics, exhibiting significant regional errors.

[0007] In recent years, nonlinear modeling methods such as machine learning have shown great potential in remote sensing data error correction. For example, researchers have applied nonlinear machine learning techniques to the bias correction tasks of GOSAT, GOSAT-2, and TROPOMI, achieving better fitting accuracy and generalization ability than traditional MLR. These methods can automatically learn the complex nonlinear relationship between multi-source input features (such as aerosol properties, observation geometry, surface reflectance, temperature and humidity profiles, etc.) and biases, effectively improving the consistency and reliability of satellite products.

[0008] Nevertheless, current research on machine learning bias correction still has several shortcomings: (1) Most models do not consider the time evolution characteristics of bias and cannot distinguish between long-term system drift and seasonal periodic errors; (2) Regional differences are not fully quantified, and the models have poor adaptability in different ecological zones and areas with high intensity of human activity; (3) Most existing models are trained once and lack interpretable analysis of the importance of factors and dynamic optimization mechanisms.

[0009] Therefore, a novel XCO2 bias correction method is urgently needed, combining multi-source data, spatiotemporal feature decomposition, and adaptive learning mechanisms. This method should be able to extract multidimensional information from satellite inversion products, meteorological reanalysis data, and ecological and environmental factors; characterize systematic and random error features through time-series decomposition; and adaptively adjust model weights based on regional labels, thereby achieving high-precision bias correction at global and regional scales. Simultaneously, utilizing interpretable algorithms (such as SHAP) to analyze the contribution of different factors helps establish a closed-loop model of "diagnosis-optimization-retraining," improving the model's physical rationality and long-term maintainability.

[0010] In summary, while existing technologies can reduce the systematic errors of satellite XCO2 products to some extent, they still have significant limitations in areas such as linearity assumptions, regional adaptability, multi-source fusion, and model interpretability. Therefore, developing an XCO2 bias correction model based on multi-source ground-based data to achieve efficient and intelligent correction of complex regions and nonlinear biases has significant scientific and practical value. Summary of the Invention

[0011] This invention aims to address the problems of large systematic biases, regional differences, and difficulty in characterizing nonlinear errors in existing XCO2 remote sensing inversion products by providing a bias correction method and system for XCO2 satellite products based on multi-source ground-based data. This method comprehensively utilizes multi-source information such as satellite observation data, meteorological reanalysis data, ecological and environmental elements, and geographic regional labels, combining a closed-loop framework of time-series decomposition, machine learning, adaptive weighting, and interpretability analysis to achieve high-precision and interpretable bias correction of XCO2 products at global and regional scales.

[0012] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows:

[0013] A method for XCO2 deviation correction based on multi-source foundation data includes the following steps:

[0014] Step S1: Acquire XCO2 inversion products from the target satellite (such as OCO-2, OCO-3, or GOSAT), and simultaneously collect multi-source auxiliary data corresponding to the spatiotemporal region of the observation area, including meteorological reanalysis data (such as ERA5 temperature, air pressure, and humidity), vegetation index (NDVI), aerosol optical depth (AOD), topographic features, and anthropogenic emission indicators. Perform temporal and spatial registration on the multi-source data, remove outliers and cloud-polluted samples, and construct a sample set paired with satellite observations and ground-based reference observations.

[0015] Step S2: Calculate the deviation sequence ΔXCO2 between the satellite inversion XCO2 and the baseline observation. Use methods such as State Space Decomposition (STL) to decompose the deviation into long-term trend terms, seasonal periodic terms, and random residual terms.

[0016] Step S3: Generate regional label information based on the geographical location, climate zone, ecosystem type, and intensity of human activities of the observation points. For example, similar regions can be divided according to continental divisions, ecological zoning, or clustering algorithms.

[0017] Step S4: Using the trend term, periodic term, residual term, multi-source auxiliary factors and region labels obtained from time series decomposition as input features, establish an initial bias correction model using ensemble learning algorithms such as XGBoost or random forest.

[0018] Step S5: During model training, dynamically adjust sample weights based on the error performance of each region. For regions with larger deviations, assign higher training weights to enable the model to achieve stronger fitting capabilities in high-error regions.

[0019] Step S6: Use the optimized model to correct the deviation of the satellite XCO2 product and output the corrected result.

[0020] Step S7: Use the SHAP method to quantify the contribution of different factors to the deviation, forming a closed-loop iterative mechanism of "diagnosis-optimization-retraining".

[0021] Further, in step S2, let the deviation sequence be:

[0022]

[0023] Constructing a state-space model:

[0024] Equations of state:

[0025]

[0026] Observation equation:

[0027]

[0028] in: Let F be the state vector, representing the trend term (systematic deviation), seasonal term (periodic deviation), and random term (random error), respectively; F and H are the state transition matrix and observation matrix, respectively. , These are process noise and observation noise, respectively.

[0029] Through Kalman filtering and smoothing algorithms, recursive estimation can be achieved. This enables the dynamic decomposition of deviations.

[0030] Furthermore, in step S3,

[0031] Constructing region labels based on multi-source auxiliary data The methods include, but are not limited to:

[0032] The area is divided according to latitude and longitude grids or administrative divisions; high-vegetation and low-vegetation areas are distinguished based on the mean NDVI value; and high-human-activity and low-activity areas are distinguished based on nighttime light intensity. The sample can be represented as:

[0033]

[0034] in This represents multiple auxiliary factors, including meteorological and vegetation data. For region labels.

[0035] Furthermore, in step S4, the bias sequence components (trend term, seasonal term, and random residual term) obtained from state-space decomposition, multi-source auxiliary factors (meteorological, AOD, cloud mask, NDVI, nighttime light, DEM, land cover, etc.), and regional labels are used as input features. The output feature variable is the bias value between satellite and ground-based observations.

[0036] Algorithm selection: The XGBoost regression model is adopted, with the optimization objectives of minimizing the root mean square error (RMSE) and maximizing the coefficient of determination (R²) to achieve the modeling and correction of XCO2 bias.

[0037] Furthermore, in step S5, during initial training, the weights of all samples are set to... .

[0038] After completing the first round of training, calculate the mean absolute error for each region:

[0039]

[0040] In the next round of training, update the weights:

[0041]

[0042] Where: α is the adjustment coefficient;

[0043] This can increase the weight of samples in high-error regions, making the model training more sensitive to spatial heterogeneity.

[0044] Furthermore, in step S7, SHAP is used to interpret the model output:

[0045]

[0046] in This represents the contribution of the j-th feature to the prediction result. By comparing the SHAP distribution under different region labels, the sources of bias can be diagnosed and the model feature system can be optimized.

[0047] The XCO2 satellite product bias correction system based on multi-source ground-based data can be used in the aforementioned XCO2 bias correction method based on multi-source remote sensing data, and specifically includes the following modules:

[0048] Data acquisition module: used to automatically retrieve satellite XCO2 products and multi-source auxiliary data from platforms such as NASA GES DISC, Copernicus, and GEE at regular intervals. The auxiliary data includes meteorological parameters, topographic elevation, vegetation index, and surface albedo.

[0049] Preprocessing module: used to perform spatial resampling, temporal registration and outlier removal on the acquired data to ensure that data from different sources are consistent under the same spatial grid and temporal scale;

[0050] Time series decomposition module: Based on the state space decomposition method, it decomposes the deviation sequence between satellite and benchmark observation into long-term trend term, seasonal periodic term and random residual term to characterize the systematic and random deviation features.

[0051] Model prediction module: used to generate high-precision XCO2 products with bias correction based on trained machine learning models (such as XGBoost), fusion of time series decomposition results, multi-source auxiliary factors and region labels;

[0052] Adaptive weighting module: used to dynamically adjust the sample training weights according to the deviation characteristics of each region, so as to achieve key optimization and adaptive enhancement of the model in high error regions;

[0053] Visualization and Output Module: Used to output analysis and display products including the following: distribution map of XCO2 bias correction results at the regional scale; heat map of the contribution of characteristic factors in each region and SHAP interpretation results.

[0054] Compared with the prior art, the advantages of the present invention are as follows:

[0055] This invention addresses the problems of systematic bias, overly strong linear assumptions, and insufficient regional adaptability in existing XCO2 satellite products by proposing a bias correction model that integrates multi-source ground-based data with an adaptive learning mechanism. Compared with existing technologies, this invention has the following significant advantages:

[0056] (1) Introduce time series decomposition to identify systematic errors.

[0057] By using the state-space decomposition method, the deviation sequence is split into trend, seasonal and residual terms, which can effectively distinguish between long-term drift and short-term disturbances and enhance the model's ability to characterize time features.

[0058] (2) Regional adaptive weighting improves the generalization of the model.

[0059] Regional labels are generated using geographical location, climate zone, and intensity of human activity, and the weights are dynamically adjusted during training to enable the model to adapt well to different ecological and topographic regions.

[0060] (3) Interpretability analysis and closed-loop optimization improve model maintainability.

[0061] By quantifying the contribution of each factor to bias correction using the SHAP method, a closed-loop mechanism of "diagnosis-optimization-retraining" is constructed to improve the transparency and physical reliability of the model.

[0062] (4) Improve the precision and consistency of XCO2 products.

[0063] The XCO2 product, after being corrected by the model of this invention, shows significantly improved consistency with high-precision ground-based observation data such as TCCON. System bias is reduced and error distribution converges, enabling it to more accurately reflect the distribution of regional carbon sources and sinks, and providing high-quality basic data for global carbon budget assessment, emission monitoring and climate model validation. Attached Figure Description

[0064] Figure 1 This is a flowchart illustrating the implementation of the present invention. Detailed Implementation

[0065] The specific technical solutions of the present invention will be described with reference to the embodiments.

[0066] The following XCO2 satellite product bias correction system, based on multi-source ground-based data, is adopted, specifically including the following modules:

[0067] Data acquisition module: used to automatically retrieve satellite XCO2 products and multi-source auxiliary data from NASA GES DISC, Copernicus, and GEE platforms at regular intervals. The auxiliary data includes meteorological parameters, topographic elevation, vegetation index, and surface albedo.

[0068] Preprocessing module: used to perform spatial resampling, temporal registration and outlier removal on the acquired data to ensure that data from different sources are consistent under the same spatial grid and temporal scale;

[0069] Time series decomposition module: Based on the state space decomposition method, it decomposes the deviation sequence between satellite and benchmark observation into long-term trend term, seasonal periodic term and random residual term to characterize the systematic and random deviation features.

[0070] Model prediction module: Used to generate high-precision XCO2 products with bias correction based on a trained machine learning model, which integrates time series decomposition results, multi-source auxiliary factors and regional labels.

[0071] Adaptive weighting module: used to dynamically adjust the sample training weights according to the deviation characteristics of each region, so as to achieve key optimization and adaptive enhancement of the model in high error regions;

[0072] Visualization and Output Module: Used to output analysis and display products including the following: distribution map of XCO2 bias correction results at the regional scale; heat map of the contribution of characteristic factors in each region and SHAP interpretation results.

[0073] like Figure 1 The process shown includes the following parts:

[0074] I. Data Preparation and Preprocessing

[0075] (1) Data source and acquisition:

[0076] OCO-2 data: OCO-2 Level 2 XCO2 product downloaded from the NASA GES DISC official website, with a spatial resolution of approximately 1.3 km × 2.25 km.

[0077] TCCON ground-based observation data: CO2 column concentration observation data from various sites downloaded from the official TCCON database. This data is based on ground-based Fourier transform infrared spectroscopy (FTIR) inversion with a time resolution of approximately 2 minutes, providing a benchmark for verifying the accuracy of OCO-2 inversion.

[0078] ERA5 reanalysis meteorological data: The ERA5 reanalysis meteorological products downloaded from the Copernicus platform have a spatial resolution of 0.25° × 0.25° and a temporal resolution of 1 h, including surface temperature, 2-meter air temperature, 10-meter u / v wind field, boundary layer height, specific humidity, surface air pressure, mean sea level air pressure, total precipitation, total column water vapor, and 100-meter wind speed.

[0079] Aerosol optical thickness data: MODIS MYD04_L2 AOD product downloaded from NASA Earthdata platform, with a spatial resolution of 10 km × 10 km.

[0080] Cloud mask data: MODIS MYD35_L2 cloud mask product (Collection 6.1) downloaded from NASA Earthdata platform, with a spatial resolution of 1 km × 1 km.

[0081] Nighttime light data: VIIRS / DNB VNP46A2 nighttime light products downloaded from the NASA Earthdata platform, with a spatial resolution of 500 m × 500 m and a temporal resolution of daily, used to reflect the intensity of human activity on the Earth's surface.

[0082] Digital elevation data: SRTM GL1 v003 global digital elevation model downloaded from NASA Earthdata platform, with a spatial resolution of approximately 90 m.

[0083] Land cover type data: MODIS MCD12Q1 land cover product downloaded from NASA Earthdata platform, with a spatial resolution of 500 m × 500 m and a temporal resolution of one year.

[0084] (2) Data processing:

[0085] OCO-2 data gridding: OCO-2 observation data are resampled to a regular grid with a grid resolution of 0.1° × 0.1°. When multiple OCO-2 observations exist within the same grid, the average value is used as the grid representative value.

[0086] TCCON matches OCO-2:

[0087] Spatial matching: Taking the latitude and longitude of the TCCON station as the center, select OCO-2 observation points within 5 grids in front, behind, left, and right (i.e., covering a range of ±0.5°);

[0088] Time matching: Based on the OCO-2 observation time, the observation values ​​of the TCCON station within ±30 minutes before and after the time are selected; if there are multiple matching points within the spatiotemporal window, the average value is taken as the matching result.

[0089] Unified processing of auxiliary data: To achieve spatial consistency with OCO-2 data, the following data were all resampled to a resolution of 0.1° × 0.1°:

[0090] ERA5 reanalyzes meteorological data: extracts meteorological elements such as surface temperature, 2-meter air temperature, 10-meter u / v wind field, boundary layer height, specific humidity, surface air pressure, mean sea level air pressure, total precipitation, total column water vapor, and 100-meter wind speed, and performs spatial interpolation to a 0.1° grid.

[0091] Aerosol optical thickness (AOD) data (MODIS MYD04_L2): Spatial resampling of the original 10 km data to generate 0.1° gridded AOD data;

[0092] Cloud mask data (MODIS MYD35_L2): Generates binary cloud / clear sky masks based on cloud detection results and resamples them to a 0.1° grid.

[0093] Nighttime light data (VIIRS / DNB VNP46A2): Daily or monthly average light intensity values ​​are taken and resampled to 0.1° resolution to characterize the intensity of human activity;

[0094] Digital Elevation Data (SRTMGL1): Extracts terrain elevation information and resamples it to 0.1° resolution;

[0095] Land cover type data (MODIS MCD12Q1): Based on the main category, resampled to a 0.1° grid for land surface type identification.

[0096] The OCO-2 gridded data and TCCON matching results were integrated, and corresponding ERA5, AOD, cloud mask, nighttime light, DEM, and land cover data were spatially matched. Only records meeting the spatiotemporal matching criteria were retained and stored as a CSV file by date. Each row contains the following fields: date, latitude and longitude, OCO-2 XCO2, TCCON XCO2, ERA5 meteorological variable, AOD, cloud mask value, nighttime light intensity, terrain elevation, and land cover type.

[0097] II. Deviation Sequence Decomposition and Region Feature Construction

[0098] After obtaining matching samples from OCO-2 satellite and TCCON ground-based observations, the first step is to calculate the deviation sequence between the two:

[0099] This sequence reflects the systematic and random differences between satellite products and benchmark observations. To identify the main patterns of deviation and potential influencing factors, the state space decomposition method was used to decompose the deviation sequence.

[0100] Specifically, the bias time series is decomposed into three parts: the long-term trend term reflects the long-term drift of satellite sensors and the systematic error of the inversion algorithm; the seasonal component characterizes the periodic biases related to vegetation activity and changes in meteorological conditions; and the random residual term is obtained after subtracting the trend and seasonal components, representing short-term disturbances and random noise.

[0101] Building upon this foundation, to enhance the model's adaptability to different geographical and environmental variations, a regional tag is introduced.

[0102] Regional labels comprehensively reflect the geographical, ecological, and human activity characteristics of observation points, and are mainly generated based on the following variables:

[0103] Geographic features: latitude and longitude, elevation (DEM); Ecological environment features: NDVI, land cover type; Human activity features: nighttime light intensity.

[0104] The regional division adopts a combination of two strategies:

[0105] Automatic zoning based on cluster analysis: Input the above standardized variables into the K-means or Gaussian Mixture Model (GMM) clustering algorithm to automatically identify regional types with similar geographical and ecological characteristics;

[0106] Rule-based modified classification: The clustering results are manually constrained and optimized, such as merging highly similar regions or distinguishing typical urban / farmland / forest areas, to ensure that the classification has geographical significance and interpretability.

[0107] Ultimately, each matched sample is accompanied by a region label (such as "urban area", "farmland area", "forest area", etc.), which together with the bias decomposition results constitute the model input feature set, providing a basis for subsequent machine learning modeling.

[0108] III. XGBoost Model Construction and Training

[0109] (1) Input features and target variables:

[0110] Feature variables: The bias sequence components (trend term, seasonal term and random residual term) obtained from state space decomposition, multi-source auxiliary factors (meteorology, AOD, cloud mask, NDVI, night light, DEM, land cover, etc.) and regional labels are used as input features.

[0111] Target variable: the deviation between satellite and ground-based observations.

[0112] (2) XGBoost model configuration:

[0113] Algorithm selection: The XGBoost regression model is adopted, with the optimization objectives of minimizing the root mean square error (RMSE) and maximizing the coefficient of determination (R²) to achieve the modeling and correction of XCO2 bias.

[0114] Hyperparameter optimization: To obtain optimal model performance, Bayesian optimization (Optuna framework) combined with cross-validation is adopted. The optimization objective is to minimize the RMSE of the validation set, the maximum number of iterations is set to 200 rounds, and an early stopping mechanism (early_stopping_rounds=50) is enabled to prevent overfitting.

[0115] Training strategy: A five-fold cross-validation scheme is used during training. The training set is randomly divided into five subsets, with four subsets used for training and one for validation. This process is repeated five times, and the average result is taken.

[0116] (3) Adaptive regional weighting mechanism:

[0117] In the initial stage, all samples are assigned the same weight; after the first round of training, the RMSE of each region is calculated, and the weights are adjusted according to their relative error levels.

[0118]

[0119] Here, α is the adjustment coefficient (usually set to 1.0), and the weight range is limited to [0.5, 5]. After updating the weights, the model is retrained, and this process is repeated 2–3 times until the regional error converges. XGBoost supports sample-level weights (sample_weight parameter), which can directly implement this strategy.

[0120] (4) Model evaluation and interpretation

[0121] Model performance was evaluated using cross-validation (CV) to ensure robustness and generalization ability. This study employed a group K-Fold CV strategy, dividing all samples into five subsets. One subset was selected as the validation set at each iteration, while the remaining four subsets were used for training. Model evaluation metrics included root mean square error (RMSE) and coefficient of determination (R²). The model's optimization objectives were minimizing RMSE and maximizing R².

[0122] To quantitatively analyze the role of each input factor in bias correction, the SHAP method was used for interpretive analysis of the model. Based on the SHAP results, factors with significant influence or high-error regions were identified; feature sets or region weights were adjusted; the model was retrained and the improvement effect was evaluated.

[0123] (5) Output and Product Generation

[0124] After model training is complete, the model with the optimal parameter configuration is applied to all OCO-2 gridded observation data, the bias prediction values ​​are calculated and corrected, and a high-precision XCO2 product is obtained. The output results are saved in NC file format by date.

Claims

1. A method for correcting biases in XCO2 satellite products based on multi-source ground-based data, comprising the following steps: Step S1: Obtain the XCO2 inversion product of the target satellite, and at the same time collect multi-source auxiliary data corresponding to the spatiotemporal region of the observation area, including meteorological reanalysis data, vegetation index NDVI, aerosol optical thickness AOD, topographic features and anthropogenic emission indicators; perform temporal and spatial registration on the multi-source data to construct a sample set that pairs satellite observations with ground reference observations; Step S2: Calculate the deviation sequence ΔXCO2 between the satellite inversion XCO2 and the baseline observation, and decompose the deviation into a long-term trend term, a seasonal periodic term, and a random residual term using state-space decomposition; Step S3: Generate regional label information based on the geographical location, climate zone, ecosystem type, and intensity of human activities of the observation points; In particular, the regional label is constructed based on multi-source auxiliary data : Divided by latitude and longitude grids or administrative divisions; differentiated between high and low vegetation areas based on mean NDVI value; differentiated between high and low human activity areas based on nighttime light intensity; sample representation: , in Indicates multi-source auxiliary factors, For region labels; Step S4: Using the trend term, periodic term, residual term, multi-source auxiliary factors and region labels obtained from time series decomposition as input features, an initial bias correction model is established using an ensemble learning algorithm; Step S5: During model training, dynamically adjust sample weights based on the error performance of each region; assign higher training weights to regions with larger deviations so that the model can achieve stronger fitting ability in high error regions. Specifically, during initial training, the weights of all samples are set to... ; After completing the first round of training, calculate the mean absolute error for each region: , in: Let represent the mean absolute error of region L, where L represents a specific region within L. This represents the total number of samples in the region. The true value of the i-th sample, and the model prediction value of the i-th sample; In the next round of training, update the weights: , Where: α is the adjustment coefficient; yes ; Step S6: Use the optimized model to correct the deviation of the satellite XCO2 product and output the corrected result; Step S7: Use the SHAP method to quantify the contribution of different factors to the bias, forming a closed-loop iterative mechanism of "diagnosis-optimization-retraining".

2. The XCO2 satellite product deviation correction method based on multi-source ground-based data according to claim 1, characterized in that: The multi-source auxiliary data are as follows: OCO-2 Level 2 XCO2 product, with a spatial resolution of approximately 1.3 km × 2.25 km; TCCON ground-based observation data; ERA5 reanalysis meteorological data: spatial resolution 0.25° × 0.25°, temporal resolution 1 h, including surface temperature, 2-meter air temperature, 10-meter u / v wind field, boundary layer height, specific humidity, surface air pressure, mean sea level air pressure, total precipitation, total column water vapor, and 100-meter wind speed; MODIS aerosol optical thickness, cloud mask, and land cover type data; and SRTM digital elevation data. Nighttime light data for VIIRS.

3. The XCO2 satellite product deviation correction method based on multi-source ground-based data according to claim 1, characterized in that: In step S2, let the deviation sequence be: , Constructing a state-space model: sh Equations of state: , Observation equation: , in: This represents the XCO2 value observed by the satellite at time t. This represents the XCO2 reference value obtained from ground-based observations at the same time t. This refers to the deviation between satellite observations and ground-based reference values. Let F be the state vector, representing the trend term, seasonal term, and random term, respectively; F and H are the state transition matrix and observation matrix, respectively. , These are process noise and observation noise, respectively. Recursive estimation is performed using Kalman filtering and smoothing algorithms. This enables the dynamic decomposition of deviations.

4. The XCO2 satellite product deviation correction method based on multi-source ground-based data according to claim 1, characterized in that: In step S4, the bias sequence components, multi-source auxiliary factors, and region labels obtained from state space decomposition are used as input features; The output feature variable is the deviation between satellite and ground-based observations; Algorithm selection: The XGBoost regression model is adopted, with the optimization objectives of minimizing the root mean square error (RMSE) and maximizing the coefficient of determination (R²) to achieve the modeling and correction of XCO2 bias.

5. The XCO2 satellite product deviation correction method based on multi-source ground-based data according to claim 4, characterized in that: In step S7, SHAP is used to interpret the model output: , in This represents the contribution of the j-th feature to the prediction result; by comparing the SHAP distribution under different region labels, the source of bias is diagnosed and the model feature system is optimized.

6. An XCO2 satellite product bias correction system based on multi-source ground-based data, characterized in that: The method for implementing the XCO2 bias correction based on multi-source remote sensing data according to any one of claims 1 to 5 specifically includes the following modules: Data acquisition module: used to automatically retrieve satellite XCO2 products and multi-source auxiliary data from NASA GES DISC, Copernicus, and GEE platforms at regular intervals. The auxiliary data includes meteorological parameters, topographic elevation, vegetation index, and surface albedo. Preprocessing module: used to perform spatial resampling, temporal registration and outlier removal on the acquired data to ensure that data from different sources are consistent under the same spatial grid and temporal scale; Time series decomposition module: Based on the state space decomposition method, it decomposes the deviation sequence between satellite and benchmark observation into long-term trend term, seasonal periodic term and random residual term to characterize the systematic and random deviation features. Model prediction module: Used to generate high-precision XCO2 products with bias correction based on a trained machine learning model, which integrates time series decomposition results, multi-source auxiliary factors and regional labels; Adaptive weighting module: used to dynamically adjust the sample training weights according to the deviation characteristics of each region, so as to achieve key optimization and adaptive enhancement of the model in high error regions; Visualization and Output Module: Used to output analysis and display products including the following: distribution map of XCO2 bias correction results at the regional scale; heat map of the contribution of characteristic factors in each region and SHAP interpretation results.