Urban underlying surface carbon flux prediction and inversion method based on machine learning

By integrating multi-source data and machine learning methods, combined with outlier screening and nonlinear modeling, the problem of low accuracy in urban underlying carbon flux inversion is solved, achieving high-precision and interpretable carbon flux prediction and inversion, which is applicable to urban environmental management and carbon emission control.

CN122047646APending Publication Date: 2026-05-15NANJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING NORMAL UNIVERSITY
Filing Date
2026-03-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In complex urban underlying surface environments, existing technologies suffer from low carbon flux inversion accuracy, high data noise, weak interpretability, and insufficient data coverage, especially at the urban scale where it is difficult to achieve continuous high-precision estimation.

Method used

A machine learning-based approach was adopted, which integrates meteorological observation data, WRF model output data and remote sensing vegetation index. Outliers were screened by combining the friction velocity threshold method and IQR statistical method. The XGBoost gradient boosting tree model was used for carbon flux prediction and inversion, and the SHAP algorithm was used for model interpretability analysis.

Benefits of technology

It improves the accuracy and reliability of urban underlying carbon flux prediction and inversion, enhances the adaptability and interpretability of the model, can more accurately identify key driving factors of carbon flux changes, and is applicable to different urban environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047646A_ABST
    Figure CN122047646A_ABST
Patent Text Reader

Abstract

According to the urban underlying surface carbon flux prediction and inversion method based on machine learning, meteorological observation data, WRF mode output data and remote sensing vegetation index data are fused, a friction speed abnormal value screening mechanism is introduced, and high-precision prediction and inversion of carbon flux are realized in combination with an XGBoost gradient boosting tree model. According to the method, a multi-dimensional time sequence feature matrix including time lag features, rolling statistical features, periodic time features and static features is constructed, time sequence cross validation is carried out, the time sequence sensitivity and generalization ability of the model are improved, an SHAP algorithm is used for carrying out explanatory analysis on the model, quantitative evaluation of feature contributions is achieved, and the accuracy of the model is improved. Therefore, the key driving factor of the carbon flux change is effectively revealed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of urban carbon cycle and greenhouse gas monitoring, specifically to a method for predicting and inverting urban underlying surface carbon flux based on machine learning. Background Technology

[0002] Urban areas account for more than 80% of global carbon emissions, and their carbon fluxes are complex in their spatiotemporal distribution, influenced by multiple factors including meteorological factors, underlying surface characteristics, and human activities. Traditional Eddy Covariance methods are limited by representative area and turbulent conditions, making it difficult to achieve continuous high-precision estimation at the urban scale. At the same time, empirical models or linear regression methods struggle to capture the nonlinear relationship between carbon fluxes and meteorological factors.

[0003] On the one hand, eddy covariance (EC) is currently the most commonly used carbon flux observation technique, estimating carbon flux by measuring gas concentration and wind speed fluctuations in turbulence. However, in complex urban environments, the diverse topography, buildings, roads, and other underlying surface features lead to complex turbulent structures, posing significant challenges to carbon flux estimation. First, EC typically assumes a large representative area around the observation point, but this assumption often fails in urban environments due to the limited representative area. For example, tall buildings in cities alter the wind field structure, causing significant differences in turbulence characteristics near the observation point compared to the surrounding area, making the observation results difficult to represent the carbon flux of the entire urban area. Second, EC relies on well-developed turbulent conditions to ensure the statistical representativeness of gas concentration and wind speed fluctuations; however, in urban environments, especially at night or under low wind speed conditions, turbulence intensity may be insufficient, increasing the uncertainty of the observation data. For instance, the boundary layer stabilizes at night, reducing turbulent activity and significantly decreasing the accuracy of carbon flux observations. Furthermore, due to limitations imposed by turbulent conditions, eddy covariance methods may fail to provide reliable carbon flux data during certain periods (such as nighttime or periods of low wind speed), leading to data gaps. This data discontinuity severely impacts the spatiotemporal continuous monitoring and long-term analysis of urban carbon flux. Additionally, the urban environment contains numerous sources of interference, such as traffic noise and construction activity. These sources introduce additional noise, affecting the quality of the observational data. Moreover, complex urban terrain and building obstructions can cause abnormal fluctuations in wind speed and gas concentrations, further increasing data noise.

[0004] On the other hand, existing empirical models and linear regression methods are widely used in carbon flux estimation, but these methods also face many problems in complex urban environments. First, there is a complex nonlinear relationship between carbon flux and meteorological factors (such as temperature, humidity, and wind speed). Empirical models and linear regression methods usually assume that these relationships are linear, which cannot accurately capture the actual nonlinear characteristics. For example, the response of vegetation photosynthesis and respiration to temperature is nonlinear, and linear models cannot accurately describe this relationship, leading to large estimation errors. Second, empirical models and linear regression methods are usually based on fixed empirical formulas or linear equations, lacking adaptability and flexibility to different environmental conditions. In urban environments, the characteristics of the underlying surface and the diversity of human activities make the driving factors of carbon flux more complex, and fixed empirical models are difficult to adapt to this diversity. For example, the vegetation cover, building density, and traffic flow vary greatly in different cities, and fixed empirical models cannot accurately reflect the impact of these differences on carbon flux. Furthermore, while empirical models and linear regression methods can provide some interpretability, they cannot quantitatively explain the contribution of each feature to carbon flux. In urban environments, carbon flux changes are influenced by a combination of factors, and the lack of quantitative analysis of the contributions of these factors casts doubt on the scientific validity and credibility of the model results. For example, it is impossible to accurately identify which meteorological factors or underlying surface features are the key drivers of carbon flux changes. Additionally, empirical models and linear regression methods typically rely on limited observational data, making it difficult to cover the complex spatiotemporal distribution of urban areas. Carbon fluxes in cities vary significantly across different regions and at different times, and limited observational data cannot fully reflect this complexity. For instance, different functional zones in a city (such as commercial areas, residential areas, and parks) exhibit different carbon flux characteristics, but empirical models and linear regression methods struggle to distinguish these differences.

[0005] In summary, existing technologies suffer from problems such as low accuracy in carbon flux inversion, high data noise, weak interpretability, and insufficient data coverage in complex urban underlying environments. There is an urgent need for a new method for carbon flux prediction and inversion that integrates multi-source data, has nonlinear modeling capabilities, and is interpretable. Summary of the Invention

[0006] To address the technical challenges of low accuracy, high data noise, weak interpretability, and insufficient data coverage in current urban complex underlying surface environments, this application provides a machine learning-based method for predicting and inverting carbon fluxes in urban underlying surfaces. This method integrates multi-source data, including meteorological observation data, WRF model output data, and remote sensing vegetation index (NDVI), and incorporates a friction velocity threshold method combined with IQR statistics as an outlier screening mechanism. The XGBoost gradient boosting tree model is then used to achieve high-precision carbon flux prediction and inversion. Furthermore, the construction of time lag, rolling window, and periodic features improves the model's temporal sensitivity and generalization ability. Additionally, the SHAP algorithm is used for model interpretability analysis, enabling a quantitative assessment of feature contributions and effectively revealing the key driving factors of carbon flux changes. This method improves the accuracy of carbon flux prediction and inversion, and enhances the accuracy of carbon budget estimation for complex urban underlying surfaces, through multi-source data fusion, nonlinear model construction, and interpretable analysis algorithms.

[0007] This invention provides a method for predicting and inverting urban underlying carbon flux based on machine learning, the specific steps of which are as follows: S1: Data acquisition, collecting and storing meteorological observation data, WRF model output data and remote sensing vegetation index data; The data collection steps also include: S11: Collect meteorological observation data, including meteorological parameters such as temperature, humidity, wind speed, radiation, and surface flux, obtained through meteorological instruments.

[0008] S2: Outlier screening step, which automatically removes outliers using friction speed threshold detection and IQR statistical methods; S21: Use the friction speed threshold to detect outliers and remove data that is below the friction speed threshold; Among them, the friction speed threshold The value can be between 0.06 and 0.09; S22: Use the IQR statistical method to remove outliers from the data obtained in S21.

[0009] S3: Feature construction step, extracting relevant features and generating a multi-dimensional temporal feature matrix; The relevant features include time lag features, rolling statistical features, periodic time features, and static features, and the above multi-dimensional time series feature matrix is ​​used to construct a training feature matrix; Among them, static features include NDVI static features, which reflect the differences in vegetation coverage on different underlying surfaces, and rolling statistical features include rolling window features.

[0010] S4: Model training steps, build and train the carbon flux inversion and prediction model based on the XGBoost gradient boosting tree algorithm, and optimize the parameters through 5-fold time series cross-validation; Specifically, to achieve the inversion and prediction of urban underlying carbon flux, the XGBoost gradient boosting tree algorithm is adopted. Its input features are the training feature matrix. After time series cross-validation and hyperparameter optimization, the model can establish a nonlinear mapping relationship: ; in, For the input feature matrix, This is the set of model parameters.

[0011] The specific model building and training steps are as follows: S41: Data preparation, using the feature matrix as the input feature matrix, and dividing the training set and test set in chronological order to ensure consistency with time series cross-validation.

[0012] S42: Model initialization, initialize the XGBoost model, and set the initial hyperparameters; S43: Time series cross-validation, using the time series cross-validation method to evaluate model performance; Furthermore, Bayesian optimization algorithms can be used to further optimize the model and its hyperparameters. S44: Obtain and store the optimization hyperparameters, and obtain the corresponding optimization hyperparameters through time series cross-validation and / or Bayesian optimization; The optimized hyperparameters include one or more of the following: learning rate, maximum depth of a single tree, number of trees (n_estimators), subsample ratio, colsample_bytree ratio, minimum leaf node weight, minimum loss reduction for node splits (gamma), and L2 / L1 regularization coefficients (lambda / alpha). S45: Select the optimal hyperparameters from the optimized hyperparameters and retrain the model; The relevant loss function and overall objective are set as follows: Since carbon flux prediction is a continuous variable regression problem, the squared error is used as the loss function: ; The overall objective is to minimize the weighted squared error and the model complexity. ; in, Represents the actual observed value. This represents the model's predicted value. Regularization term representing model complexity S46: Evaluate model performance on the test set.

[0013] S5: Carbon flux inversion step, inputting real-time features into the trained carbon flux prediction and inversion model, and outputting the predicted carbon flux; Real-time features may include real-time meteorological and environmental features.

[0014] S6: Model interpretation step, using the SHAP algorithm to quantitatively interpret the contribution of each input feature to carbon flux prediction and identify key driving factors of carbon flux.

[0015] This invention also provides a system applicable to the machine learning-based urban underlying carbon flux prediction and inversion method described above. The system includes six core modules: a data acquisition module, an outlier screening module, a feature construction module, a model training module, a carbon flux inversion module, and a model interpretation module. By employing time lag, rolling window, and periodic feature construction, the model's temporal sensitivity and generalization ability are improved.

[0016] The data acquisition module collects and stores meteorological observation data, WRF model output data, and remote sensing vegetation index (NDVI) data. The outlier filtering module automatically removes outliers using friction speed threshold detection and IQR statistical methods. The feature construction module extracts features and generates a multi-dimensional time series feature matrix, wherein the features include time lag features, rolling statistical features, periodic time features and static features, and constructs the above multi-dimensional time series feature matrix into a feature matrix; The model training module is based on the XGBoost gradient boosting tree algorithm to build and train a carbon flux inversion and prediction model, and optimizes the parameters through 5-fold time series cross-validation. The carbon flux inversion module takes real-time features as input to the trained carbon flux prediction and inversion model and outputs the predicted carbon flux. The model interpretation module uses the SHAP algorithm to quantitatively interpret the contribution of each input feature to carbon flux prediction and identify key driving factors of carbon flux.

[0017] Compared with the prior art, the advantages of the method provided by the present invention include: 1. This application proposes a machine learning-based method for predicting and inverting carbon fluxes on urban underlying surfaces. This method significantly improves the accuracy and reliability of predicting and inverting carbon fluxes on urban underlying surfaces by integrating multi-source data fusion, constructing nonlinear models and interpretable analysis algorithms, and comprehensively utilizing technologies such as multi-source data fusion, outlier screening, time series analysis, high-precision model construction and interpretability analysis. It has important application value for urban environmental management and carbon emission control.

[0018] 2. The method in this application integrates multi-source data such as meteorological observation data, WRF model output data, and remote sensing vegetation index (NDVI). These data cover multiple aspects of the urban environment, such as meteorological conditions, vegetation cover, and surface features, thereby providing more comprehensive input information, enhancing the model's predictive ability, reducing the limitation of the area around the observation point in the urban environment, improving the predictive model's adaptability to the intensity of turbulence in the urban environment, and improving the accuracy of carbon budget estimation for complex urban underlying surfaces, thus significantly improving the accuracy of carbon flux prediction and inversion.

[0019] 3. This application establishes an outlier screening mechanism, combining the friction velocity threshold method and the IQR statistical method for a dual-mechanism outlier screening. This dual mechanism, combining physical meaning and statistical characteristics, can more comprehensively identify and remove outliers. It can not only eliminate observational biases caused by insufficient turbulence but also identify and remove outliers in the data, thereby significantly improving the overall quality and reliability of the data.

[0020] By introducing a friction velocity outlier screening mechanism, this method can automatically remove abnormal observation data caused by factors such as the urban heat island effect and building obstruction. It is particularly effective in screening observation samples under urban nighttime turbulent conditions, effectively eliminating observational biases caused by insufficient turbulence, thereby improving data quality, consistency, and reliability. Furthermore, the IQR statistical method is employed to further eliminate potential outliers from the perspective of data distribution.

[0021] 4. This application employs the XGBoost gradient boosting tree model to achieve nonlinear modeling. It constructs a feature matrix using features such as time lag characteristics, rolling statistical characteristics, periodic time characteristics, and static characteristics. The application also improves the sensitivity of the model to time series data through time series cross-validation. This allows for the evaluation of the model's accuracy and reliability in time series prediction tasks from multiple perspectives, enhancing the model's generalization ability and time sensitivity. Furthermore, the application can combine Bayesian algorithms to optimize the model's hyperparameters, achieving high-precision and rapid prediction and inversion of carbon flux. This model is suitable for processing large-scale datasets and helps to more accurately estimate the carbon budget of complex urban underlying surfaces, which is of great significance for urban environmental management and carbon emission control.

[0022] 5. In view of the characteristics of the XGBoost tree model, this application adopts the TreeSHAP algorithm, which calculates the expected contribution of each feature at the split node by traversing the path probability of each tree, thereby improving the computational efficiency of the tree model, reducing the computation time and hardware requirements, improving the reliability and interpretability, and providing visualization tools to display the global feature importance graph, feature dependency graph and local interpretation graph, making the interpretation results of the model more intuitive and easy to understand, and facilitating user understanding and analysis.

[0023] 6. The method of this application takes into account a variety of complex factors in the urban environment, such as buildings, roads, vegetation and water bodies. Through the fusion of multi-source data and outlier screening, it can be applied to different climate zones and underlying surface types, which enhances the adaptability of the model to different urban environments and improves the coverage and representativeness of the data, enabling the model to be applied in a wider range of environments.

[0024] 7. It offers a modular design that can be integrated with existing urban carbon monitoring systems to achieve automated inversion calculations, enhancing the scalability of model applications. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 This is the system flowchart of this application; Figure 2 This is a comparison chart of model predictions and actual values ​​of carbon flux in embodiments of this application; Figure 3 This is a visualization data graph of the importance (global) of SHAP features in an embodiment of this application; Figures 4 to 7 These are multiple visualized data graphs of the SHAP dependency graph in the embodiments of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] This invention provides a method for predicting and inverting urban underlying carbon flux based on machine learning, the specific steps of which are as follows: S1: Data Acquisition. Acquire and store meteorological observation data, WRF model output data, and remote sensing vegetation index (NDVI) data.

[0028] Specifically, meteorological observation data, WRF model output data, and remote sensing vegetation data are fused from multiple sources to construct a multi-source data set and achieve data multidimensionality.

[0029] The data collection steps also include: S11: Collect meteorological observation data, including meteorological parameters such as temperature, humidity, wind speed, radiation, and surface flux, obtained through meteorological instruments.

[0030] Temperature reflects the thermal state of the environment and is a key factor affecting vegetation physiological processes and soil microbial activity. For example, temperature changes affect vegetation photosynthesis and respiration. Higher temperatures usually promote photosynthesis, but excessively high temperatures may lead to a decrease in photosynthesis. At the same time, rising temperatures also accelerate the respiration of vegetation and soil, increasing carbon emissions. Furthermore, in urban environments, underlying surface features such as buildings and roads can lead to local temperature increases, forming the urban heat island effect. Therefore, temperature is of great significance for the prediction and inversion of carbon flux on urban underlying surfaces.

[0031] Humidity, including relative humidity and specific humidity, reflects the amount of water vapor in the atmosphere and affects vegetation transpiration and soil moisture retention. In urban environments, humidity is influenced not only by weather and plant factors but also by the obstruction of buildings and roads, as well as human activities such as irrigation and fountains.

[0032] Wind speed reflects the speed at which gases are transported in the atmosphere. Wind speed affects the transport and diffusion of gases. Higher wind speeds help with the mixing and transport of gases, but excessively high wind speeds may lead to excessive diffusion of gases, reducing the accuracy of carbon flux observation. Wind speed has significant spatiotemporal variations. In urban environments, buildings and terrain can alter the wind field structure, forming complex turbulence.

[0033] Radiation reflects the input of solar energy. Higher radiation intensity is conducive to photosynthesis, thereby increasing carbon absorption. In addition to having obvious diurnal and seasonal variations, the reflection and absorption effects of buildings and roads in urban environments affect the distribution of radiation.

[0034] Surface fluxes reflect the energy and momentum exchange between the Earth's surface and the atmosphere. They influence the thermal and moisture states of the surface, thereby affecting vegetation physiological processes and soil microbial activity. Surface fluxes include sensible heat flux, latent heat flux, and momentum flux, reflecting the energy and momentum exchange between the Earth's surface and the atmosphere. In urban environments, the thermal capacity and roughness of buildings and roads affect the distribution of surface fluxes.

[0035] Specifically, data can be collected periodically through temperature sensors, humidity sensors, wind speed sensors, radiation sensors, eddy covariance systems, or surface energy balance models at weather stations, for example, every 10 or 30 minutes.

[0036] The Weather Research and Forecasting (WRF) model is a mesoscale numerical weather prediction system. Its output data can be stored in NetCDF format and includes data on multiple temporal and spatial meteorological elements, such as temperature, humidity, wind speed, and precipitation. This data can be downloaded from the official websites or data sharing platforms of relevant research institutions. Based on research needs, appropriate input data (such as initial and boundary conditions) are obtained, and the WRF model software is used to perform simulations, thereby generating the required output data.

[0037] NDVI (Normalized Difference Vegetation Index) is a commonly used vegetation index that reflects the growth status of vegetation by calculating the difference in reflectance between the red and near-infrared bands in satellite images. The corresponding NDVI data can be downloaded from satellite remote sensing platforms, such as MODIS NDVI data, which can be downloaded from data sources such as the NASA MODISWeb website.

[0038] To acquire satellite remote sensing imagery data containing both red and near-infrared bands, such as Landsat and Sentinel-2 satellite data, the relevant formula is as follows: ; in, , These are the red and near-infrared band reflection parameters from Sentinel-2 satellite data, respectively.

[0039] S2: Outlier screening step, which automatically removes outliers using friction speed threshold detection and IQR statistical methods; In urban environments, the combined influence and interaction of various factors, including the urban heat island effect, the obstruction and friction of buildings, changes in the nighttime boundary layer structure, the impact of human activities, the street canyon effect, the influence of nighttime urban lighting, and the effects of urban vegetation and water bodies, lead to enhanced and persistent nighttime turbulence. Therefore, in observational data under nighttime turbulent conditions, turbulence intensity directly affects the reliability and representativeness of the data. Friction velocity can characterize turbulence intensity; a low friction velocity indicates insufficient turbulent exchange, potentially leading to inaccurate carbon flux observations. To improve data accuracy, this invention proposes setting a friction velocity threshold to directly eliminate observational samples with insufficient turbulence (which may not accurately reflect the turbulent exchange process of the atmospheric boundary layer in a physical sense). This screening based on nighttime turbulent conditions effectively eliminates observational biases caused by insufficient turbulence, ensuring better consistency and reliability of the remaining data in terms of physical processes.

[0040] Based on this, the present invention also employs the IQR (interquartile range) statistical method to further filter out outliers in the data. First, a friction velocity threshold is set according to nighttime turbulence conditions. The process involves first removing observations with insufficient turbulence, and then using the interquartile range (IQR) method to identify high-value outliers. An automatic outlier removal mechanism is achieved by combining the friction velocity threshold method with the IQR statistical method. The friction velocity threshold method, based on the physical process, ensures the reliability of the data under turbulent conditions; the IQR statistical method, from the perspective of data distribution, further eliminates potential outliers. This dual mechanism, combining physical meaning and statistical characteristics, can more comprehensively identify and remove outliers. It not only eliminates observational biases caused by insufficient turbulence but also identifies and removes outliers in the data, thereby significantly improving the overall quality and reliability of the data.

[0041] Step S21: Use the friction speed threshold to detect outliers and remove data that are below the friction speed threshold; Among them, the friction speed threshold The value can be between 0.06 and 0.09.

[0042] Step S22: Use the IQR statistical method to remove outliers from the data obtained in step S21.

[0043] Specifically, the data is sorted from smallest to largest, and the following calculations are performed: ; ; ; in, This represents the first quartile, the value at the 25th percentile, which divides the dataset into the smaller 25% and the larger 75%. The third quartile represents the value at the 75th percentile, dividing the dataset into the smaller 75% and the larger 25%. Based on this, samples exceeding the upper and lower bounds are removed. The upper bound is set to a multiple of 3.0. Any data point below the lower bound or above the upper bound is removed as a potential outlier, which can balance outlier identification and sample retention.

[0044] Table 1 shows the friction velocity thresholds at Nanjing Normal University Station and Yangtze River Station during June–September 2024. The results of two-step outlier removal were presented. The results showed that the nighttime turbulence thresholds at the two stations were 0.073 m / s and 0.087 m / s, respectively, with an overall outlier removal rate within the range of 5%–7%, indicating good data quality. Table 1 .

[0045] S3: Feature construction step, extracting features and generating a multi-dimensional time series feature matrix, wherein the features include time lag features, rolling statistical features, periodic time features and static features, and constructing the above multi-dimensional time series feature matrix into a training feature matrix; Specifically, after cleaning the observation data, the WRF model output data is upsampled to 30 minutes and aligned with the observation data by time to construct a training feature matrix. The features include periodic time features, time lag features, rolling statistical features, and static features. These features reflect the multi-dimensional features of multiple data sources, such as meteorological observation data features, WRF model output data features, NDVI data features, and time features.

[0046] Among them, static features include NDVI static features, which reflect the differences in vegetation coverage on different underlying surfaces, and rolling statistical features include rolling window features.

[0047] S4: Model training steps, build and train the carbon flux inversion and prediction model based on the XGBoost gradient boosting tree algorithm, and optimize the parameters through 5-fold time series cross-validation; Specifically, to achieve the inversion and prediction of urban underlying carbon flux, this invention employs the XGBoost gradient boosting tree algorithm. Its input features are the training feature matrix. After time-series cross-validation and hyperparameter optimization, the model can establish a nonlinear mapping relationship: ; in, For the input feature matrix, This is the set of model parameters.

[0048] The specific model building and training steps are as follows: S41: Data preparation: Use the feature matrix as the input feature matrix, and divide the training set and test set in chronological order to ensure consistency with time series cross-validation.

[0049] Furthermore, the above operations can also prevent information leakage.

[0050] S42: Model initialization, initialize the XGBoost model, and set the initial hyperparameters; S43: Time series cross-validation, using the time series cross-validation method to evaluate model performance; Specifically, to evaluate the model's robustness on out-of-time samples, a 5-fold time-series cross-validation was used to optimize the parameters, i.e., 5-fold cross-validation based on time windows (TimeSeriesSplit, n_splits=5). In each fold, the model is trained using an earlier time period and validated using a later time period to maintain temporal causality. After completing the cross-validation, the model was retrained using all samples to obtain the final model for subsequent inversion and prediction.

[0051] The evaluation indicators include the coefficient of determination. Root mean square error (RMSE) and mean absolute error (MAE): ; ; ; in, Represents the actual observed value. This represents the model's predicted value. Represents all actual observations The average value.

[0052] Coefficient of determination The first metric represents the goodness of fit between the model's predicted values ​​and the actual observed values; a value closer to 1 indicates a better fit. The second metric, Root Mean Square Error (RMSE), represents the square root of the average error between the model's predicted and actual observed values; a smaller value indicates a more accurate prediction. The third metric, Mean Absolute Error (MAE), represents the average absolute error between the model's predicted and actual observed values; a smaller value indicates a more accurate prediction. These three metrics work together to evaluate the model's accuracy and reliability in time series forecasting tasks from multiple perspectives.

[0053] Furthermore, Bayesian optimization algorithms can be used to further optimize the model and its hyperparameters. During model training, time-series cross-validation can be used first to evaluate the model's performance under different parameter settings. Then, Bayesian optimization algorithms can be employed to further optimize hyperparameters, such as the learning rate and the maximum depth of the tree. Bayesian optimization predicts model performance under different hyperparameter combinations by constructing surrogate models (such as Gaussian process regression) and uses functions such as the desired improvement in EI as a retrieval function to find the optimal model configuration.

[0054] The specific optimization process can be as follows: Step 1: Select the objective function as the optimization objective based on the characteristics of the hyperparameters.

[0055] Furthermore, the objective function can be chosen as the root mean square error (RMSE). By adjusting the hyperparameters, the objective function can be minimized to find the optimal model configuration.

[0056] Step 2: Based on the characteristics of the hyperparameters, select a prior distribution for each hyperparameter.

[0057] Furthermore, when the hyperparameter is the learning rate, a uniform distribution can be chosen as the prior distribution.

[0058] Step 3: Construct a surrogate model. The surrogate model can be constructed using the previous historical evaluation results to predict the objective function under different combinations of hyperparameters.

[0059] Furthermore, Gaussian process regression (GPR) can be used to construct the proxy model.

[0060] Step 4: Select the acquisition function to balance the relevant indicators.

[0061] Furthermore, the acquisition function may include indicators such as expected improvement in index (EI), probability improvement in index (PI), and upper confidence bound (UCB).

[0062] Step 5: Iterate and optimize until the termination condition is met; Step 51, Initialization: Randomly select a set of hyperparameters from the prior distribution for evaluation; Step 52: Evaluate the model performance through the objective function, store and analyze the objective function values, update the surrogate model based on the evaluation results, and select the next hyperparameter combination based on the relevant indicators of the obtained function. Step 53: Repeat step 52 until the termination condition is met.

[0063] Furthermore, the termination condition could be a predetermined number of attempts or a cessation of performance improvement.

[0064] S44. Obtain and store the optimized hyperparameters, and obtain the corresponding optimized hyperparameters through time series cross-validation and Bayesian optimization. The optimized hyperparameters include one or more of the following: learning rate, maximum depth of a single tree, number of trees, subsample, colsample bytree, minimum leaf node weight, minimum loss reduction for node splits (gamma), and L2 / L1 regularization coefficients (lambda / alpha).

[0065] S45. Model training: Retrain the model using the optimal hyperparameters.

[0066] S45. Select the optimal hyperparameters from the optimized hyperparameters and retrain the model. The relevant loss function and overall objective are set as follows: Since carbon flux prediction is a continuous variable regression problem, the squared error is used as the loss function: ; The overall objective is to minimize the weighted squared error and the model complexity. ; in, Represents the actual observed value. This represents the model's predicted value. Represents all actual observations The average value, Regularization term representing model complexity.

[0067] Preferably, the model can be composed of 100 decision trees, with a maximum depth of 6 for each tree, a learning rate of 0.1, a squared error as the loss function, and an L2 regularization term to prevent overfitting.

[0068] S46. Evaluate the model performance on the test set.

[0069] Figure 2 The paper presents a comparison between the model predictions and actual observation data in the final stage of cross-validation between Nanjing Normal University Station and Yangtze River Station, with the data covering the period from August 21 to September 5, 2024.

[0070] from Figure 2As can be seen, the predicted results of Nanjing Normal University Station and Yangtze River Station are highly consistent with the overall trend of the measured data. The model proposed in this application can accurately reflect the diurnal variation and fluctuation characteristics of carbon flux. Among them, the flux change of Yangtze River Station is relatively gentle, and the model prediction almost coincides with the measurement with small error. Local deviations mainly occur under extreme weather conditions or during low turbulence at night, which is consistent with the characteristics of the uncertainty of eddy covariance observation. Although the data fluctuation of Nanjing Normal University Station is more obvious, the model proposed in this application can still maintain a high degree of fit during the night and morning periods.

[0071] S5, Carbon flux inversion step: Input the real-time features into the trained carbon flux prediction and inversion model, and output the predicted carbon flux. Real-time features may include real-time meteorological and environmental features.

[0072] S6. Model interpretation step: The SHAP algorithm is used to quantitatively interpret the contribution of each input feature to carbon flux prediction and identify the key driving factors of carbon flux.

[0073] To improve the interpretability of the model results, this invention introduces the SHAP algorithm after the model training is completed to quantitatively interpret the contribution of each input feature to carbon flux prediction.

[0074] Specifically, since the model used in this invention for carbon flux inversion and prediction is based on the XGBoost gradient boosting tree algorithm and is essentially composed of multiple decision trees, using the commonly used traditional Shapley algorithm for interpretation would lead to high complexity, placing high demands on computation time, hardware requirements, and processing speed, making it very difficult to implement in practice. Considering the above problems, this application proposes using the SHAP algorithm to interpret the model. The SHAP algorithm interprets the model prediction by calculating the marginal contribution of each feature to the model output. For the XGBoost-based model, we use the TreeSHAP algorithm, which calculates the expected contribution of each feature at the split node by traversing the path probability of each tree. This allows us to obtain the quantitative impact of each feature on the model prediction, thereby identifying the key driving factors of carbon flux changes and improving the computational efficiency of the tree model.

[0075] When interpreting the SHAP algorithm, the model output is set to... Let the input feature set be... For any input sample ,feature The Shapley value is defined as its value across all possible subsets of features. Average marginal contribution in: ; in, Indicates using only a subset of features The expected output of the model at that time; Indicates when adding features The expected output is given by the coefficients, which are used to assign weights to ensure that the contributions of all features are distributed equally.

[0076] The above average marginal contribution reflects the characteristics The average marginal contribution to the predicted output, the sum of the SHAP values ​​of all features equals the model prediction bias: ; The expected contribution of each feature at the split node is calculated by traversing the path probability of each tree, and its time complexity is reduced by... Down to ,in: Indicates the number of trees; Indicates the number of leaf nodes; This indicates the maximum depth of the tree.

[0077] The predicted value for each tree can be broken down as follows: ; in As the global baseline value, Representation of features The marginal contribution.

[0078] Furthermore, settings can be configured for model implementation and visualization.

[0079] The SHAP value for each sample and each feature can be calculated using the SHAP library in Python, and a global feature importance plot (summary_plot) can be drawn to show the average absolute contribution. The feature dependency plot reveals the nonlinear response relationship between features and predicted values; and the force plot shows the positive and negative contributions of each feature to the predicted values ​​in a single sample.

[0080] Preferably, the feature importance matrix can be calculated as follows: ; in The total number of samples, Representation of features The average absolute influence.

[0081] Figure 3 The paper presents a global feature importance map of SHAP for carbon flux prediction at Yangtze River stations. Figures 4-7 The corresponding SHAP dependency graphs for each feature are given, from Figures 3-7As can be seen, the SHAP algorithm can well explain the positive and negative contributions of each feature to the predicted value, and can visualize them for easy observation.

[0082] This invention also provides a system applicable to the machine learning-based urban underlying carbon flux prediction and inversion method described above. The system comprises six core modules: a data acquisition module, an outlier screening module, a feature construction module, a model training module, a carbon flux inversion module, and a model interpretation module. By employing time lag, rolling windows, and periodic feature construction, the model's temporal sensitivity and generalization ability are improved.

[0083] The data acquisition module collects and stores meteorological observation data, WRF model output data, and remote sensing vegetation index (NDVI) data. The outlier filtering module automatically removes outliers using friction speed threshold detection and IQR statistical methods. The feature construction module extracts features and generates a multi-dimensional time series feature matrix, wherein the features include time lag features, rolling statistical features, periodic time features and static features, and constructs the above multi-dimensional time series feature matrix into a feature matrix; The model training module is based on the XGBoost gradient boosting tree algorithm to build and train a carbon flux inversion and prediction model, and optimizes the parameters through 5-fold time series cross-validation. The carbon flux inversion module takes real-time features as input to the trained carbon flux prediction and inversion model and outputs the predicted carbon flux. The model interpretation module uses the SHAP algorithm to quantitatively interpret the contribution of each input feature to carbon flux prediction and identify key driving factors of carbon flux.

[0084] This invention utilizes the aforementioned machine learning-based method for predicting and inverting urban underlying carbon flux. It integrates meteorological observation data, WRF model output data, and remote sensing vegetation indices, and introduces a friction velocity outlier screening mechanism. Combined with the XGBoost gradient boosting tree model, it achieves high-precision prediction and inversion of carbon flux. The method comprises six core steps: data acquisition, outlier screening, feature construction, model training, carbon flux inversion, and model interpretation. By constructing a multi-dimensional time-series feature matrix including time lag features, rolling statistical features, periodic time features, and static features, and performing time-series cross-validation, the model's time-series sensitivity and generalization ability are improved. Furthermore, the SHAP algorithm is used for interpretive analysis of the model, enabling a quantitative assessment of feature contributions, thereby effectively revealing the key driving factors of carbon flux changes.

[0085] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0086] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0087] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

Claims

1. A method for predicting and inverting urban underlying carbon flux based on machine learning, characterized in that, Includes the following steps: S1. Data Acquisition: Acquire and store meteorological observation data, WRF model output data, and remote sensing vegetation index data. S2, Outlier screening step: Automatically remove outliers using friction speed threshold detection and IQR statistical methods; S3. Feature construction step: Extract relevant features and generate a multi-dimensional temporal feature matrix; S4. Model training steps: Construct and train a carbon flux inversion and prediction model based on the XGBoost gradient boosting tree algorithm. S5, Carbon flux inversion step: Input the real-time features into the trained model and output the predicted carbon flux. S6. Model interpretation step: The SHAP algorithm is used to quantitatively interpret the contribution of each input feature to carbon flux prediction and identify the key driving factors of carbon flux.

2. The method for predicting and inverting urban underlying carbon flux based on machine learning according to claim 1, characterized in that, The model training process also includes the following steps: S41. Data preparation: Use the feature matrix as the input feature matrix, and divide the training set and test set according to time order to ensure consistency with time series cross-validation. S42. Model initialization: Initialize the XGBoost model and set the initial hyperparameters. S43. Time series cross-validation: Use the time series cross-validation method to evaluate model performance. S44. Obtain and store the optimization hyperparameters, and obtain the corresponding optimization hyperparameters through time series cross-validation and / or Bayesian optimization. S45. Select the optimal hyperparameter from the optimized hyperparameters and retrain the model; S46. Evaluate the model performance on the test set.

3. The method for predicting and inverting urban underlying carbon flux based on machine learning according to claim 2, characterized in that, The squared error is used as the loss function: ; in, Represents the actual observed value. This represents the model's predicted value.

4. The method for predicting and inverting urban underlying carbon flux based on machine learning according to claim 2, characterized in that, The overall objective is to minimize the weighted squared error and model complexity. ; in, Represents the actual observed value. This represents the model's predicted value. Regularization term representing model complexity.

5. The method for predicting and inverting urban underlying carbon flux based on machine learning according to claim 1, characterized in that, The relevant features include time lag features, rolling statistical features, periodic time features, and static features.

6. The method for predicting and inverting urban underlying carbon flux based on machine learning according to claim 1, characterized in that, The real-time features may include real-time meteorological and environmental features.

7. The method for predicting and inverting urban underlying carbon flux based on machine learning according to claim 1, characterized in that, The outlier screening process also includes the following steps: S21. Use the friction speed threshold to detect outliers and remove data that are below the friction speed threshold. S22. Use the IQR statistical method to remove outliers from the data obtained in S21.

8. The method for predicting and inverting urban underlying carbon flux based on machine learning according to claim 1, characterized in that, Parameters were optimized using 5-fold time series cross-validation.

9. The method for predicting and inverting urban underlying carbon flux based on machine learning according to claim 5, characterized in that, Static features include NDVI static features, which reflect differences in vegetation cover on different underlying surfaces.