Soil organic carbon remote sensing estimation method based on habitat plaque division and machine learning

By combining habitat patch delineation and machine learning methods with remote sensing data and environmental covariates, the problem of insufficient accuracy and interpretability in remote sensing estimation of soil organic carbon in complex topographic areas was solved, and high-precision and transparent prediction of soil organic carbon content was achieved.

CN120992512APending Publication Date: 2025-11-21SOUTHWEST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511101342.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy and insufficient interpretability in remote sensing estimation of soil organic carbon in complex terrain areas. The "black box" nature of machine learning models limits their transparency and interpretability.

Method used

This study employed a habitat patch division and machine learning approach, combining remote sensing data and environmental covariates. The PAM clustering algorithm was used to divide the study area into habitat patches, and the SHAP model was used to explain the contributions of feature variables. Support vector machine, random forest, and extreme gradient boosting decision tree models were used to predict soil organic carbon content.

Benefits of technology

This study enabled precise remote sensing estimation of soil organic carbon content in complex terrain areas, improving the model's prediction accuracy and interpretability, and revealing the relationship between characteristic variables and soil organic carbon content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120992512A_ABST
    Figure CN120992512A_ABST
Patent Text Reader

Abstract

The invention discloses a soil organic carbon remote sensing estimation method based on habitat plaque division and machine learning, and belongs to the technical field of ecological environment evaluation, and the method comprises the following steps: S1, obtaining remote sensing data and basic geographic data of a research area, and carrying out the preprocessing of the remote sensing data; s2, according to the preprocessed remote sensing data and basic geographic data, calculating a spectral index and an environment covariable of a research area; s3, performing sub-region division on the research region; and S4, simulating and generating a spatial distribution diagram of the organic carbon content of the soil by utilizing machine learning based on the sub-region division result of the research region, the spectral index and the environment covariable. According to the method, the problem of relatively low simulation precision possibly caused by using a single model to simulate the existing complex terrain region and the limitation of a machine learning model in the explanatory aspect are made up.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of ecological environment evaluation, and particularly relates to a soil organic carbon remote sensing estimation method based on habitat patch division and machine learning. BACKGROUND

[0002] It is of great significance to obtain high-resolution and large-scale spatial information of soil organic carbon quickly, estimate its spatial distribution accurately and understand its spatial variation for mitigating global climate change, developing precision agriculture, realizing soil carbon sequestration potential evaluation and improving environmental quality. However, it is challenging to carry out high-precision digital soil mapping in complex terrain areas (high spatial heterogeneity). In the aspect of soil organic carbon variability analysis, soil sampling and field investigation are traditionally the main data acquisition methods. With the emergence of complex and high spatiotemporal resolution remote sensing data, satellite remote sensing data are increasingly widely used in soil organic carbon research. By using climate factors, terrain factors, soil types, optical-microwave remote sensing information, land use change and population density and other human activity information, combined with geostatistics and spatial analysis methods, the spatial distribution pattern of soil organic carbon in complex terrain areas can be obtained. However, most current methods are based on modeling and estimating soil organic carbon in the whole study area, and there are still some deficiencies in remote sensing estimation of soil organic carbon content in complex terrain areas, and there is a lack of a special remote sensing fine estimation method for soil organic carbon in complex terrain areas. In addition, although machine learning models have achieved remarkable results in improving the estimation accuracy of SOC content, they cannot explain the relationship between the model and the input parameters, and the inherent "black box" characteristics limit the improvement of model transparency and interpretability, and cannot be visualized. Traditional methods, such as machine learning algorithms based on tree models, can provide certain clues by sorting the importance of features, but it is difficult to quantify the specific influence and direction of each feature on the target output. To solve this problem, the SHAP method is based on the principle of game theory, which compares the performance difference of the model containing and not containing a specific variable to accurately evaluate the individual contribution of each variable. By calculating the variable contribution for each specific instance, the local and global importance of the input variable is comprehensively considered, the relationship between the feature variable and the SOC content is revealed, and the limitations of machine learning models in interpretability are effectively made up. SUMMARY

[0003] The application is proposed to solve the above problems, and provides a soil organic carbon remote sensing estimation method based on habitat patch division and machine learning.

[0004] The technical scheme of the application is as follows: a soil organic carbon remote sensing estimation method based on habitat patch division and machine learning comprises the following steps:

[0005] S1, acquire remote sensing data and basic geographic data of a study area, and pre-process the remote sensing data;

[0006] S2, calculate spectral indices and environmental covariates of the study area according to the pre-processed remote sensing data and the basic geographic data;

[0007] S3, divide the study area into sub-regions;

[0008] S4, based on the sub-region division result, the spectral indices and the environmental covariates of the study area, simulate and generate a spatial distribution map of soil organic carbon content by using machine learning.

[0009] Further, in S1, the specific method for pre-processing the remote sensing data is to perform radiation calibration, atmospheric correction, image stitching and cropping processing on the image data.

[0010] Further, in S2, according to the wave bands of the pre-processed remote sensing data, calculate a plurality of spectral indices of the study area;

[0011] The spectral indices of the study area include a first brightness index, a second brightness index, a color index, a clay index, a ratio vegetation index, a green-red vegetation index, a normalized difference vegetation index, a surface water index, a second modified soil corrected vegetation index, a normalized water index, a modified normalized water index, a water stress index, a surface water index, a redness index, a total vegetation index adjusted by soil, a vegetation index corrected by soil, a transformed vegetation index, a leaf area index, a normalized vegetation index, a greenness normalized vegetation index and an enhanced vegetation index.

[0012] Further, in S3, the PAM clustering algorithm is used to divide the study area into a plurality of habitat patches, thereby completing the sub-region division.

[0013] Further, in S4, based on the plurality of spectral indices and the environmental covariates, the best prediction model of the soil organic carbon content of each sub-region and the best prediction model of the soil organic carbon content of the study area are determined by using each machine learning model, and according to the prediction accuracy of each machine learning model, the spatial distribution map of the soil organic carbon content is generated.

[0014] Further, in S4, the machine learning models include support vector machines, random forests, extreme gradient boosting decision trees and light gradient boosting machines.

[0015] Further, in S4, the prediction accuracy of the machine learning model includes a determination coefficient, a root mean square error, a mean absolute error and a relative analysis error.

[0016] The beneficial effects of this invention are as follows: Based on a full consideration of environmental variables, remote sensing optical indices, and microwave remote sensing data, and considering the strong heterogeneity of the surface in the study area, this invention introduces a clustering analysis method to rationally divide the study area, constructs a remote sensing estimation model for soil organic carbon in complex terrain areas of Southwest China, realizes a precise remote sensing estimation of soil organic carbon content in complex terrain areas of Southwest China, and introduces the SHAP model to interpret the selected digital soil mapping model, thus overcoming the problem that the use of a single model to simulate complex terrain areas may lead to low simulation accuracy and the limitations of machine learning models in terms of interpretability. Attached Figure Description

[0017] Figure 1 This is a flowchart of a remote sensing estimation method for soil organic carbon based on habitat patch delineation and machine learning.

[0018] Figure 2 This is a schematic diagram of a region selected for studying complex terrain in a specific embodiment of the present invention. Detailed Implementation

[0019] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0020] like Figure 1 As shown, this invention provides a remote sensing estimation method for soil organic carbon based on habitat patch delineation and machine learning, comprising the following steps:

[0021] S1. Acquire remote sensing data and basic geographic data of the study area, and preprocess the remote sensing data;

[0022] S2. Calculate the spectral index and environmental covariates of the study area based on the preprocessed remote sensing data and basic geographic data;

[0023] S3. Divide the study area into sub-regions;

[0024] S4. Based on the sub-regional division results, spectral indices, and environmental covariates of the study area, machine learning is used to simulate and generate a spatial distribution map of soil organic carbon content.

[0025] The present application acquires and pre-processes multi-source environment-remote sensing data in the southwest region; selects environmental covariates including climate, terrain, soil and human activities, and also selects Sentinel-1 polarization data, Sentinel-2A band data and 15 spectral indices calculated from these bands as remote sensing variables; the clustering algorithm around the central partition PAM is used to divide the study area into different habitat patch types; the effects of optimizing the auxiliary variable combination of each region by using filter, wrapper and embedded feature selection methods are compared; a variety of machine learning models are used to establish soil organic carbon content prediction models in independent sub-regions and the whole region respectively; the spatial distribution map of soil organic carbon content in the southwest region is generated; the Shapley Additive ExPlanations (SHAP) model is introduced to explain the selected digital soil mapping model to verify its reliability. The present application realizes the remote sensing fine estimation of soil organic carbon content, and overcomes the problem of low simulation accuracy.

[0026] The present embodiment selects the southwest region as shown in Figure 2 as the research complex terrain region of the present application, first collects Landsat8 image data in the southwest region, and pre-processes the image data. VV and VH polarization data of Sentinel-1, annual mean temperature, annual precipitation, land surface temperature LST, population density, land use, DEM and its extracted terrain indicators such as terrain relief TU, slope, aspect and terrain humidity index TWI are collected.

[0027] In S1, SOC measured data, annual mean temperature, annual precipitation, land surface temperature LST, population density, land use, DEM and its extracted related terrain indicators are collected.

[0028] In the embodiment of the present application, in S1, the specific method for pre-processing the remote sensing data is: performing radiation calibration, atmospheric correction, image stitching and cropping processing on the image data.

[0029] In the embodiment of the present application, in S2, according to the bands of the pre-processed remote sensing data, a plurality of spectral indices of the study area are calculated.

[0030] The spectral indices of the study area include first brightness index, second brightness index, color index, clay index, ratio vegetation index, green-red vegetation index, normalized difference vegetation index, surface water index, quadratic modified soil corrected vegetation index, normalized water index, modified normalized water index, water stress index, surface water index, redness index, soil-adjusted total vegetation index, soil-corrected vegetation index, transformed vegetation index, leaf area index, normalized vegetation index, greenness normalized vegetation index and enhanced vegetation index.

[0031] Table 1

[0032]

[0033] Wherein, B2, B3, B4, B5, B6 and B7 respectively represent the reflectivity of the blue band (450-515nm), green band (525-600nm), red band (630-680nm), near-infrared band (845-885nm), short-wave infrared band (1560-1660nm) and short-wave infrared band 2 (2100-2300nm) corresponding to Landsat8 image.

[0034] In the embodiment of the present application, in S3, the study area is divided into several habitat patches by using PAM clustering algorithm, and the sub-area division is completed.

[0035] The division method of habitat patch surface unit is adopted; the concept of habitat patch is: a specific geographical space of ecological environment formed under the action of similar natural and human factors, and the patch has relative uniformity of environmental factors inside, including the state and time variation characteristics of environmental factors; the specific division method includes clustering analysis: the clustering algorithm of PAM around the center partition is used to divide the study area into four different habitat patch types; after division, the effects of optimizing the combination of auxiliary variables of each region by using filter, wrapper and embedded feature selection method are compared.

[0036] In the embodiment of the present application, in S4, based on several spectral indices and environmental covariates, the best prediction model of soil organic carbon content of each sub-area and the best prediction model of soil organic carbon content of the study area are determined by using each machine learning model, and the spatial distribution map of soil organic carbon content is generated according to the prediction accuracy of each machine learning model.

[0037] In the embodiment of the present application, in S4, the machine learning model includes support vector machine, random forest, extreme gradient boosting decision tree and light gradient boosting machine.

[0038] In the embodiment of the present application, in S4, the prediction accuracy of the machine learning model includes determination coefficient, root mean square error, mean absolute error and relative analysis error.

[0039] The accuracy of the simulation output result of the machine learning model is respectively evaluated by determination coefficient R 2 , root mean square error RMSE, mean absolute error MAE and relative analysis error RPD, and the specific evaluation formula is as follows:

[0040]

[0041]

[0042] P = Predicted soil organic carbon content i and O i are the predicted and observed soil organic carbon content, respectively, n is the number of sample points, and are the mean of predicted and observed soil organic carbon content, respectively, SD is the standard deviation of observed soil organic carbon content.

[0043] In the embodiments of the present application, the SHAP model is introduced to interpret the selected digital soil mapping model to verify its reliability. Although the machine learning model has achieved remarkable results in improving the estimation accuracy of SOC content, it cannot explain the relationship between the model and the input parameters. Its inherent "black box" characteristics limit the improvement of model transparency and interpretability, and cannot be visualized. Traditional methods, such as machine learning algorithms based on tree models, can provide some clues by sorting the importance of features, but it is difficult to quantify the specific influence and direction of each feature on the target output. To solve this problem, the SHAP method emerged as the times require. It cleverly draws on the principles of game theory, and compares the performance difference between models with and without a specific variable to accurately assess the individual contribution of each variable. In addition, SHAP can calculate the variable contribution for each specific instance, thereby achieving a comprehensive consideration of the local and global importance of input variables.

[0044] Based on the optimal feature combination and the combination of machine learning model architecture determined in step S4, the SHAP interpretation method is used to analyze the interpretability of the SOC content estimation model constructed according to the environmental similarity division strategy. The influence of each input feature on the model estimation result is analyzed from the global and local perspectives. This analysis aims to comprehensively evaluate and quantify the contribution of each factor in the process of estimating SOC content, reveal the influence mechanism of variable driving model output, and thereby improve the understanding of the prediction mechanism of the model.

[0045] Those skilled in the art will appreciate that the embodiments described herein are intended to aid the reader in understanding the principles of the present application and should not be construed as limiting the scope of protection of the present application to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the spirit of the present application, and these modifications and combinations are still within the scope of protection of the present application.

Claims

1. A method for estimating soil organic carbon based on habitat patch division and machine learning, characterized in that, The method comprises the following steps: S1, acquiring remote sensing data and basic geographic data of a research area, and preprocessing the remote sensing data; S2, calculating spectral indices and environmental covariates of the research area according to the preprocessed remote sensing data and the basic geographic data; S3, dividing the research area into sub-regions; S4, based on the sub-region division result of the research area, the spectral indices and the environmental covariates, simulating and generating a spatial distribution map of soil organic carbon content by using machine learning.

2. The habitat patch-based and machine learning-based method for soil organic carbon remote sensing estimation according to claim 1, characterized in that, In the S1, the specific method for preprocessing the remote sensing data is: performing radiation calibration, atmospheric correction, image stitching and cutting processing on the image data.

3. The habitat patch-based and machine learning-based method for soil organic carbon remote sensing estimation according to claim 1, characterized in that, In the S2, a plurality of spectral indices of the research area are calculated according to the wave bands of the preprocessed remote sensing data; The spectral indices of the research area include a first brightness index, a second brightness index, a color index, a clay index, a ratio vegetation index, a green-red vegetation index, a normalized difference vegetation index, a surface water index, a second modified soil corrected vegetation index, a normalized water index, a modified normalized water index, a water stress index, a surface water index, a redness index, a soil adjusted total vegetation index, a soil corrected vegetation index, a transformed vegetation index, a leaf area index, a normalized vegetation index, a greenness normalized vegetation index and an enhanced vegetation index.

4. The habitat patch-based and machine learning-based method for soil organic carbon remote sensing estimation according to claim 1, characterized in that, In the S3, the PAM clustering algorithm is used to divide the research area into a plurality of habitat patches, thereby completing the sub-region division.

5. The habitat patch-based and machine learning-based soil organic carbon remote sensing estimation method according to claim 1, characterized in that, In the S4, based on the plurality of spectral indices and the environmental covariates, the best prediction model of the soil organic carbon content of each sub-region and the best prediction model of the soil organic carbon content of the research area are determined by using each machine learning model, and the spatial distribution map of the soil organic carbon content is generated according to the prediction accuracy of each machine learning model.

6. The habitat patch-based and machine learning-based method for soil organic carbon remote sensing estimation according to claim 5, characterized in that, In the S4, the machine learning models include a support vector machine, a random forest, an extreme gradient boosting decision tree and a light gradient boosting machine.

7. The habitat patch-based and machine learning-based method for soil organic carbon remote sensing estimation according to claim 5, characterized in that, In the S4, the prediction accuracy of the machine learning model includes a determination coefficient, a root mean square error, a mean absolute error and a relative analysis error.

Citation Information

Cited By

  • Natural ecosystem soil carbon reserve calculation method based on machine learning

    CN121901843A