NEP estimation method based on feature selection and machine learning

By using multi-source data fusion and feature selection methods, combined with XGBoost Regression and SHAP analysis, the accuracy and interpretability issues in NEP estimation were resolved, and high-precision estimation and mechanism analysis of complex mountain ecosystems were achieved.

CN120673215APending Publication Date: 2025-09-19SOUTHWEST FORESTRY UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510959215.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies in NEP estimation face problems such as high cost, limited spatial coverage, insufficient model stability and insufficient machine learning interpretability, making it difficult to achieve high-precision estimation and mechanism explanation in complex mountain ecosystems.

Method used

By fusing multi-source remote sensing data with environmental factors, using feature selection and the Extreme Gradient Boosting Regression (XGBoost Regression, XGBR) algorithm for training, and combining it with the Shapley Sum Interpretation Method (SHAP) for analytical verification, a high-precision NEP estimation model was constructed.

Benefits of technology

High-precision estimation of NEP and analysis of driving mechanisms in complex mountain ecosystems have been achieved, which has improved the estimation accuracy and mechanism interpretation ability and overcome the limitations of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673215A_ABST
    Figure CN120673215A_ABST
Patent Text Reader

Abstract

The invention discloses an NEP estimation method based on feature selection and machine learning, and belongs to the technical field of ecological environment monitoring and carbon cycle research, and the method comprises the steps: obtaining remote sensing image data and environment factors of a target region, and carrying out the preprocessing of the obtained remote sensing image data, and obtaining remote sensing factors; based on the remote sensing factor and the environmental factor, obtaining a net ecosystem productivity NEP feature influencing multi-source data fusion of the target area, performing screening, obtaining a preset number of key factors, and performing training through an extreme gradient lifting regression algorithm; analyzing the training result of the extreme gradient lifting regression algorithm by adopting a Shapril addition interpretation method; the extreme gradient lifting regression algorithm is verified through the evaluation indexes, and target area NEP estimation is completed; according to the method, high-precision estimation and driving mechanism analysis of the NEP of the complex mountain ecosystem are realized by fusing multi-source data and a machine learning technology, and clear mechanism explanation is provided for ecological research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of ecological environment monitoring and carbon cycle research, and specifically relates to a NEP estimation method based on feature selection and machine learning. Background Art

[0002] Currently, three main technical means are relied upon: flux observation methods obtain data through direct measurements at the site level, but are costly and have limited spatial coverage; process models simulate carbon cycle processes based on ecological mechanisms but face challenges of complex parameters and large uncertainties; and although machine learning methods can integrate multi-source data to improve prediction efficiency, they are difficult to reveal the driving mechanism.

[0003] Existing methods face significant bottlenecks in NEP estimation: flux observations are difficult to meet the needs of continuous monitoring at the regional scale, process models are overly reliant on prior knowledge and lack stability in simulation results, and machine learning methods, while offering predictive advantages, are unable to effectively quantify the contribution of various environmental factors and their interactions. This is primarily due to feature selection methods' inability to capture nonlinear relationships, leading to redundant input data, and the lack of an explanatory framework, which makes it difficult to distinguish between global and local driving mechanisms. These issues severely limit the in-depth application of machine learning in carbon cycle mechanism research, and there is an urgent need to develop new technical solutions that combine high precision and strong interpretability. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a NEP estimation method based on feature selection and machine learning.

[0005] To implement the above technology, the specific steps include: S1. Acquire remote sensing image data and environmental factors of the target area, perform preprocessing operations on the acquired remote sensing image data, and obtain remote sensing factors; Remote sensing image data include: interannual global vegetation net primary productivity data products, 8-day surface reflectance composite data products, and annual global land cover type datasets; Environmental factors in the target area include: climate factors, soil factors, topographic factors, human factors and vegetation factors; climate factors include: temperature, precipitation and potential evapotranspiration; soil factors include: soil pH, soil carbon content, soil clay content and soil sand content; topographic factors include: altitude, slope and aspect; human factors include: night light intensity and population density; vegetation factors refer to vegetation factors in the target area; The methods for performing preprocessing operations on acquired remote sensing image data include: Spatial resampling: Use bilinear interpolation for spatial resampling; Time scale synthesis processing: Perform annual time scale synthesis processing on monthly data of remote sensing image data; The remote sensing factors obtained include: Difference Vegetation Index (DVI), Enhanced Vegetation Index (EVI), Green Chlorophyll Vegetation Index (GCVI), Greenness Index (GI), Green Normalized Difference Vegetation Index (GNDVI), Kernel Normalized Difference Vegetation Index (kNDVI), Modified Soil-Adjusted Vegetation Index (MSAVI), Modified Simple Ratio Index (MSR), Normalized Difference Vegetation Index (NDVI), Normalized Green-Red Difference Index (NGRDI), Non-Linear The following wavelengths are used for the analysis: soil vegetation index (SVI), soil-adjusted vegetation index (NLI), optimized soil-adjusted vegetation index (OSAVI), renormalized difference vegetation index (RDVI), ratio vegetation index (RVI), soil-adjusted vegetation index (SAVI), triangular vegetation index (TVI), visible atmospheric resistive index (VARI), red band, green band, blue band, and near infrared band (NIR Band).

[0006] S2. Based on remote sensing factors and environmental factors, a net ecosystem productivity (NEP) estimation feature system is constructed by integrating multi-source data of the target area to obtain the NEP characteristics affecting the target area; The method to obtain the NEP characteristics affecting the target area is: Based on the pre-processed 8-day synthetic surface reflectance data product, the red band (Red Band), green band (Green Band), blue band (Blue Band), and near-infrared band (NIR Band) of the target area are extracted, and a preset number of NEP features are calculated through software; at the same time, the NEP features of environmental factors are sorted out.

[0007] S3. Based on the Net Ecosystem Productivity (NEP) characteristics of the multi-source data fusion affecting the target area, a preset number of key factors are obtained and trained using the Extreme Gradient Boosting Regression (XGBoost Regression, XGBR) algorithm; The method to obtain the preset number of key factors is: In the R_Studio environment, the VSURF package (Version 1.1.0) using the variable selection method using random forests (VSURF) was used for feature selection. The feature selection method was a two-stage optimization method, namely: The first stage loads the package and reads the Excel data file containing NEP observations and NEP features of the target area. The VSURF(*) function is called with the NEP observations as the target value and the feature matrix formed by the Excel data file of the NEP features of the target area as the input parameter. The random forest parameters are set to the preset values ​​and the number of variable trials is set to the default value. The expression of the NEP observation value is as follows: NEP Observed value = NPP-R h Where NPP represents the net primary productivity of vegetation, which is obtained through the interannual global vegetation net primary productivity data product MOD17A3HGF; Represents pixels exist Soil heterotrophic respiration in the month; Represents pixels exist Monthly average temperature (℃); Represents pixels exist Monthly total precipitation (mm); if >0, indicating that the carbon fixed by the ecosystem is greater than the carbon released, which is a carbon sink, otherwise it is a carbon source; The second stage adopts a forward stepwise optimization strategy, starting from the empty model, iteratively adding the feature variables that can most significantly reduce the OOB error, and terminates the iteration when the error change rate is less than 1%; Finally, a preset number of key factors are retained; The training method for the extreme gradient boosting regression algorithm (XGBoost Regression, XGBR) is as follows: the XGBoost library in Python is used to build the XGBR regression model and the training is performed through k-fold cross validation; During the specific implementation process, a preset number of key factors and NEP observations are retained as the training set, and the training set and test set are divided into 10 mutually exclusive subsets using a preset ratio. Nine subsets are used as training data in turn, and the remaining subset is used as validation data. The full sample validation is completed after k cycles.

[0008] S4. Analytically validate the XGBoost Regression (XGBR) training results using Shapley Additive Explanations (SHAP). The analytical verification method is to call the SHAP library in the Python environment, calculate the Shapley value of the characteristic variable, and systematically quantify the contribution of different driving factors to NEP and their interactions.

[0009] S5. Evaluate and verify the analytically verified extreme gradient boosting regression algorithm (XGBoost Regression, XGBR); The validation indicators include: coefficient of determination R², root mean square error RMSE and mean absolute error MAE.

[0010] Beneficial effects of the present invention This study, by integrating multi-source data with machine learning techniques, combined with VSURF feature selection and SHAP interpretability analysis, achieves high-precision estimation of NEP in complex mountain ecosystems and explains the driving mechanisms. Compared with traditional methods, this method significantly improves both estimation accuracy and mechanism interpretation.

[0011] The present invention integrates multi-source data and adopts the VSURF algorithm for feature optimization to construct an XGBR regression model, which significantly improves the estimation accuracy of NEP and solves the problem of insufficient accuracy of traditional methods in complex mountain ecosystems.

[0012] This paper combines the SHAP interpretability analysis framework to quantitatively analyze the contribution of each driving factor to NEP and the interaction mechanism, overcoming the "black box" limitations of machine learning models and providing a clear mechanism explanation for ecological research. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a principle framework diagram of the present invention; Figure 3 is a scatter plot of the XGBR model fitting according to an embodiment of the present invention; Figure 4 is a histogram of the global importance of features according to an embodiment of the present invention; Figure 5 It is a SHAP summary diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0014] The present invention is further described in detail below with reference to specific embodiments.

[0015] like Figure 1 and Figure 2 As shown in FIG, a NEP estimation method based on feature selection and machine learning includes the following steps: S1. Acquire remote sensing image data and environmental factors of the target area, perform preprocessing operations on the acquired remote sensing image data, and obtain remote sensing factors; In this embodiment, the target area is Yunnan Province; The remote sensing image data of the target area are obtained from the MODIS products provided by the National Aeronautics and Space Administration (NASA); Remote sensing imagery data includes: an interannual global net primary productivity product (MOD17A3HGF, 500-meter spatial resolution), an 8-day composite surface reflectance product (MOD09A1, 500-meter spatial resolution), and an annual global land cover data set (MCD12Q1, 500-meter spatial resolution). The 8-day composite surface reflectance product (MOD09A1, 500-meter spatial resolution) is based on the MODIS surface reflectance level 2 gridded product (MOD09GHK, 500-meter resolution) and is synthesized using optimal observation selection criteria, including the highest observation quality score, minimum observation zenith angle, and the absence of aerosol interference flag, as well as alignment of pixels identified using bitmask technology to exclude cloud cover. The environmental factors of the target area include: climate factors, soil factors, terrain factors, human factors and vegetation factors; among them, climate factors include: temperature, precipitation and potential evapotranspiration, and the spatial resolution is set to 1 km when collected; soil factors include: soil pH value, soil carbon content, soil clay content and soil sand content, and the spatial resolution is set to 250 meters when collected; terrain factors include: altitude, slope and slope aspect. In this invention, terrain factors are collected based on SRTM data, and the spatial resolution is set to 90 meters when collected; human factors include: night light intensity and population density. Night light intensity is collected based on VIIRS data, and population density is collected based on WorldPop data. The spatial resolution of night light intensity and population density is set to 1 km when collected; vegetation factors are vegetation factors of the target area, and the spatial resolution is set to 500 meters when collected; The methods for performing preprocessing operations on acquired remote sensing image data include: Spatial resampling: Bilinear interpolation is used for spatial resampling to reduce the spatial resolution of remote sensing image data to 1000 meters; Time scale synthesis processing: Perform annual time scale synthesis processing on monthly data of remote sensing image data; The remote sensing factors obtained include: Difference Vegetation Index (DVI), Enhanced Vegetation Index (EVI), Green Chlorophyll Vegetation Index (GCVI), Greenness Index (GI), Green Normalized Difference Vegetation Index (GNDVI), Kernel Normalized Difference Vegetation Index (kNDVI), Modified Soil-Adjusted Vegetation Index (MSAVI), Modified Simple Ratio Index (MSR), Normalized Difference Vegetation Index (NDVI), Normalized Green-Red Difference Index (NGRDI), Non-Linear The following wavelengths are used for the analysis: soil vegetation index (SVI), soil-adjusted vegetation index (NLI), optimized soil-adjusted vegetation index (OSAVI), renormalized difference vegetation index (RDVI), ratio vegetation index (RVI), soil-adjusted vegetation index (SAVI), triangular vegetation index (TVI), visible atmospheric resistive index (VARI), red band, green band, blue band, and near infrared band (NIR Band).

[0016] S2. Based on remote sensing factors and environmental factors, a net ecosystem productivity (NEP) estimation feature system is constructed by integrating multi-source data of the target area to obtain the NEP characteristics affecting the target area; The method to obtain the NEP characteristics affecting the target area is: Based on the preprocessed 8-day synthetic surface reflectance data product (MOD09A1, spatial resolution 1000 meters), the red band, green band, blue band, and near-infrared band (NIR band) of the target area were extracted. A preset number of NEP features were calculated using Python 3.9. In this paper, 16 vegetation indices were set, corresponding to the remote sensing factors after removing the red band, green band, blue band, and near-infrared band (NIR band), as shown in Table 1. At the same time, the NEP features of environmental factors were obtained, as shown in Table 2. Table 1: 16 vegetation indices obtained using Python 3.9 In Table 1, G, R, B, RE, and NIR represent the reflectance values ​​of the green band, red band, blue band, red edge band, and near infrared band, respectively; Table 2: Environmental factors S3. Based on the Net Ecosystem Productivity (NEP) characteristics of the multi-source data fusion affecting the target area, a preset number of key factors are obtained and trained using the Extreme Gradient Boosting Regression (XGBoost Regression, XGBR) algorithm; The method to obtain the preset number of key factors is: In the R_Studio environment, the VSURF package (Version 1.1.0) using the variable selection method using random forests (VSURF) was used for feature selection. The feature selection method was a two-stage optimization method, namely: The first stage loads the package and reads the Excel data file containing NEP observations and NEP features of the target area. The VSURF(*) function is called with the NEP observations as the target value and the feature matrix formed by the Excel data file of the NEP features of the target area as the input parameter. The random forest parameter n_tree is set to 100 and the number of variable attempts m_try is set to the default value. The VSURF(*) function algorithm first performs a preliminary screening based on the Gini importance score, that is, the importance random noise level is automatically determined through out-of-bag error (OOB) analysis, and features with importance significantly higher than the random noise level are retained. The expression of the NEP observation value is as follows: NEP Observed value = NPP-R h Where NPP represents the net primary productivity of vegetation, which is obtained through the interannual global vegetation net primary productivity data product MOD17A3HGF; Represents pixels exist Soil heterotrophic respiration in the month; Represents pixels exist Monthly average temperature (℃); Represents pixels exist Monthly total precipitation (mm); if >0, indicating that the carbon fixed by the ecosystem is greater than the carbon released, which is a carbon sink, otherwise it is a carbon source; The second stage adopts a forward stepwise optimization strategy, starting from the empty model, iteratively adding the feature variables that can most significantly reduce the OOB error, and terminates the iteration when the error change rate is less than 1%; Finally, 11 key factors were retained, including: Enhanced Vegetation Index (EVI), vegetation type (VegType), soil pH (pH), soil clay content (SC), temperature (Temp), soil sand content (SS), solar radiation (SolRad), precipitation (Precip), greenness index (GI), potential evapotranspiration (PET), and difference vegetation index (DVI); The training method using the extreme gradient boosting regression algorithm (XGBoost Regression, XGBR) is as follows: using the XGBoost library in Python to build an XGBR regression model, and training through k-fold cross validation. In this embodiment, k is 10; During the specific implementation process, the 11 retained key factors and NEP observations were used as the training set. For the total sample size of 182,040 data, the training set (127,428 samples) and the test set (54,612 samples) were divided into a ratio of 7:3. In this embodiment, the training set has 127,428 samples, which are randomly divided into 10 mutually exclusive subsets. 9 subsets (114,685 samples) are used as training data, and the remaining subset (12,743 samples) is used as validation data. The full sample validation is completed in 10 cycles. The model parameters are set as follows: 1000 iterations (n_estimators=1000), a learning rate of 0.1 (learning_rate=0.1), a maximum tree depth of 10 (max_depth=10), and an early stopping mechanism (early_stopping_rounds=50) is set to prevent overfitting. like Figure 3 The scatter plot of the obtained results and the measured values ​​shows that the data points are closely distributed on both sides of the 1:1 line, verifying the high-precision prediction ability of the model in complex mountain ecosystems.

[0017] S4. Analytically validate the XGBoost Regression (XGBR) training results using Shapley Additive Explanations (SHAP). The analytical verification method is to call the SHAP library (version 0.41.0) in the Python environment to calculate the Shapley value of the characteristic variable. The system quantifies the contribution of different driving factors to NEP and their interactions, such as Figure 4 and Figure 5 As shown in the figure, 11 factors are the most important factors affecting NEP.

[0018] S5. Evaluate and verify the analytically verified extreme gradient boosting regression algorithm (XGBoost Regression, XGBR); The validation indicators include: coefficient of determination R², root mean square error RMSE and mean absolute error MAE; The verification results are shown in Table 3; Table 3: XGBR model evaluation As shown in Table 3, the coefficient of determination (R²) reaches 0.94, the root mean square error (RMSE) is 76.82 gC / (m²·a), and the mean absolute error (MAE) is 55.11 gC / (m²·a). The model performs well on the test set (54,612 samples).

[0019] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely illustrative of the principles and applications of the invention. It should be understood that many modifications may be made to the illustrative embodiments, and that other arrangements may be devised, without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in ways other than those described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be employed in conjunction with other described embodiments.

Claims

1. A NEP estimation method based on feature selection and machine learning, characterized in that: The following steps are involved: S1. Acquire remote sensing image data and environmental factors of the target area, perform preprocessing operations on the acquired remote sensing image data, and obtain remote sensing factors; S2. Based on remote sensing factors and environmental factors, a multi-source data fusion net ecosystem productivity (NEP) estimation feature system for the target area is constructed to obtain the multi-source data fusion net ecosystem productivity (NEP) characteristics that affect the target area. S3. Screen the NEP characteristics obtained from the fusion of multi-source data affecting the target area to obtain a preset number of key factors, and train them using the extreme gradient boosting regression algorithm XGBR; The screening method for screening the net ecosystem productivity (NEP) characteristics obtained by fusion of multi-source data affecting the target area is: a variable selection method based on random forest; The method of training by the extreme gradient boosting regression algorithm XGBR is as follows: using the XGBoost library in Python to build an XGBR regression model and training it by k-fold cross validation; S4. Use Shapley’s sum interpretation method SHAP to analyze the training results of the extreme gradient boosting regression algorithm XGBR; S5. Verify the extreme gradient boosting regression algorithm through evaluation indicators and complete the NEP estimation of the target area; The validation indicators include: determination coefficient R², root mean square error RMSE and mean absolute error MAE.

2. The NEP estimation method based on feature selection and machine learning according to claim 1, characterized in that: The remote sensing image data and environmental factors of the target area are obtained, and a preprocessing operation is performed on the obtained remote sensing image data to obtain the remote sensing factors, wherein the remote sensing image data includes: an interannual global vegetation net primary productivity data product, an 8-day synthetic data product of surface reflectance, and an annual global data set of land cover types; the 8-day synthetic data product of surface reflectance is an 8-day data synthesized based on the MODIS surface reflectance secondary gridded product and using the best observation selection criteria; the selection criteria include: the highest observation quality score, the minimum observation zenith angle, and the no aerosol interference flag, and alignment of pixels identified using a bit masking technique to exclude cloud cover; The preprocessing operation includes: spatial resampling and time scale synthesis processing.

3. The NEP estimation method based on feature selection and machine learning according to claim 1, characterized in that: The method of obtaining the net ecosystem productivity NEP characteristics of the multi-source data fusion affecting the target area by constructing a net ecosystem productivity NEP estimation feature system based on remote sensing factors and environmental factors and the multi-source data fusion affecting the target area is as follows: extracting the red light band, green light band, blue light band and near-infrared band of the target area based on the preprocessed 8-day synthetic data product of surface reflectance, and calculating a preset number of NEP characteristics through software, wherein the preset number of NEP characteristics obtained is the same as the number of remote sensing factors after removing the red light band, green light band, blue light band and near-infrared band.

4. The NEP estimation method based on feature selection and machine learning according to claim 1, characterized in that: The net ecosystem productivity (NEP) characteristics obtained by fusion of multi-source data affecting the target area are screened to obtain a preset number of key factors, and are trained through the extreme gradient boosting regression algorithm XGBR. The method of obtaining the preset number of key factors is: calling the random forest-based variable selection method VSURF in the R_Studio environment to perform feature selection; wherein the feature selection method is a two-stage optimization method.

5. The NEP estimation method based on feature selection and machine learning according to claim 4, characterized in that: The feature selection method is a two-stage optimization method including: The first stage loads the package and reads the Excel data file containing NEP observations and NEP features of the target area. The VSURF(*) function is called with the NEP observations as the target value and the feature matrix formed by the Excel data file of the NEP features of the target area as the input parameter. The random forest parameters and the number of variable attempts with default values ​​are set. The VSURF(*) function algorithm automatically determines the importance of random noise level through out-of-bag error (OOB) analysis and retains features whose importance is significantly higher than the random noise level. The expression of the NEP observation is as follows: NEP Observed value = NPP-R h Where NPP represents the net primary productivity of vegetation, which is obtained through the interannual global vegetation net primary productivity data product MOD17A3HGF; Represents pixels exist Soil heterotrophic respiration in the month; Represents pixels exist Average monthly temperature; Represents pixels exist The total precipitation in a month; >0, indicating that the carbon fixed by the ecosystem is greater than the carbon released, which is a carbon sink, otherwise it is a carbon source; The second stage adopts a forward stepwise optimization strategy: starting from the empty model, iteratively add NEP features that can most significantly reduce the out-of-bag error (OOB) error in the target area, and terminate the iteration when the error change rate is less than 1%.

Citation Information

Cited By

  • Vegetation carbon sink space optimization method and system based on interpretable machine learning

    CN121390476A