Soil organic carbon density space non-stationary dominant factor identification and high-resolution mapping method

The random forest model is trained through raster data of feature variables and combined with SHAP algorithm to identify the dominant factors of soil organic carbon density, solving the problem of insufficient SOCD prediction accuracy and interpretability in the existing technology, realizing high-resolution and spatial interpretation analysis, supporting precise agricultural management and soil protection.

CN120298535APending Publication Date: 2025-07-11JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510275276.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has complex interactions with limited accuracy and difficulty in comprehensively considering multiple environmental factors in the spatial prediction mapping of soil organic carbon density (SOCD). The random forest model has poor interpretability and cannot provide spatial-level explanatory analysis, which limits the guidance of precise agricultural management and soil protection.

Method used

The random forest model is trained using raster data of feature variables, combined with SHAP interpretable machine learning algorithm, and superimposes the spatial distribution map of the SHAP value of feature variables, identify the dominant factors of each cell, and realize high-resolution prediction and spatial-level explanatory analysis.

Benefits of technology

The high-resolution prediction of SOCD is achieved, and the explanatory analysis is provided at the spatial level, which can quantify the specific contribution and impact direction of each influencing factor to the predicted results, and provides a scientific basis for farmland management policies and soil protection measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298535A_ABST
    Figure CN120298535A_ABST
Patent Text Reader

Abstract

The invention discloses a soil organic carbon density spatial non-stationary dominant factor identification and high-resolution mapping method, which uses raster data of characteristic variables as input to train an RF model, so that the trained RF model can correspondingly output and predict spatial distribution of soil organic carbon density, and high-resolution prediction of SOCD in space is realized. Meanwhile, the SHAP value spatial distribution diagrams of all the feature variables are superposed, the SHAP value of each feature variable on each pixel can be obtained, and then the dominant factor of each pixel is obtained, so that the specific influence of the feature variables of different regions on SOCD spatial distribution is obtained, and interpretive analysis of the spatial level is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of digital soil mapping (DSM) technology, and particularly relates to a method for identifying dominant factors of spatial non-stationarity of soil organic carbon density and high-resolution mapping. Background Art

[0002] Soil organic carbon density (SOCD) is a key indicator for measuring soil carbon storage, and is of great significance for the assessment of the carbon cycle in farmland ecosystems, the formulation of soil protection policies, and the response to climate change. However, there are still many challenges in the spatial prediction mapping of SOCD.

[0003] Literatures (McBratney, A. B., Santos, M. M., & Minasny, B. (2003). On digital soil mapping. Geoderma, 117(1-2), 3-52; Liang, Z., Chen, S., Yang, Y., Zhao, R., Shi, Z., & Rossel, R. A. V. (2019). National digital soil map of organic matter in topsoil and its associated uncertainty in 1980’s China. Geoderma, 335, 47-56; Hu, B., Xie, M., Zhou, Y., Chen, S., Zhou, Y., Ni, H.,... & Shi, Z. (2024). A high-resolution map of soil organic carbon in cropland of Southern China. Catena, 237, 107813) disclose traditional prediction methods, such as predictions based on interpolation or simple regression models, which often have limited accuracy and are difficult to comprehensively consider the effects of multiple environmental factors and their complex interactions.

[0004] Literature (Breiman, L. (2001). Random forests. Machine learning, 45, 5-32; Were, K., Bui, D. T., Dick, B., & Singh, B. R. (2015). A comparative assessment of support vector regression, artificial neural networks, and random forests for predicting and mapping soil organic carbon stocks across an Afromontane landscape. Ecological Indicators, 52, 394 - 403; Hu, B., Xie, M., Zhou, Y., Chen, S., Zhou, Y., Ni, H.,... & Shi, Z. (2024). A high - resolution map of soil organic carbon in cropland of Southern China. Catena, 237, 107813) have disclosed that in recent years, the random forest (RF) model, as an advanced machine learning algorithm, has been widely used in SOCD prediction. This model improves the accuracy and robustness of prediction by constructing multiple decision trees and integrating their prediction results. However, the random forest model still has obvious limitations in the spatial prediction mapping of SOCD.

[0005] Literature (Minasny, B., Setiawan, B. I., Saptomo, S. K., & McBratney, A. B. (2018). Open digital mapping as a cost - effective method for mapping peat thickness and assessing the carbon stock of tropical peatlands. Geoderma, 313, 25 - 40; Hengl, T., Nussbaum, M., Wright, M. N., Heuvelink, G. B., & B. (2018). Random forest as a generic framework for predictive modeling of spatial and spatio-temporal variables. PeerJ, 6, e5518) discloses that, first, the interpretability of the RF model is poor and it is called a "black box" model. Although this model can provide relatively accurate SOCD prediction results, it is difficult to quantitatively analyze the specific contributions and influence directions of various influencing factors on the prediction results. This makes it lack a scientific basis for decision-makers to formulate farmland management policies or implement soil protection measures and unable to clarify the action mechanisms of various factors. Second, the RF model cannot analyze the influence of influencing factors on the prediction results at the spatial level. This model usually can only provide global predictions and feature importance rankings, but cannot reveal the specific influence of different regional characteristic variables on the spatial distribution of SOCD. This limits the understanding and analysis of the spatial variation characteristics of SOCD and cannot provide targeted guidance for precision agriculture management and soil protection.

[0006] Therefore, there is an urgent need to develop a new method that can overcome the above limitations to achieve high-resolution prediction of SOCD and provide explanatory analysis at the spatial level. Summary of the Invention

[0007] The present invention provides a method for identifying dominant factors of spatial non-stationarity and high-resolution mapping of soil organic carbon density, which can achieve high-resolution prediction of SOCD and provide explanatory analysis at the spatial level.

[0008] A specific embodiment of the present invention provides a method for identifying dominant factors of spatial non-stationarity and high-resolution mapping of soil organic carbon density, including:

[0009] Converting the obtained soil organic matter content into soil organic carbon density, converting the obtained covariates into raster format data through spatial preprocessing, where the covariates include soil attribute data and environmental covariate data, and performing feature screening on the covariates converted into raster format through the RFE algorithm to obtain feature variables and raster data of the feature variables;

[0010] Training an RF model based on the feature variables, raster data of the feature variables, and soil organic carbon density, and obtaining the spatial distribution of the predicted soil organic carbon density through the trained RF model;

[0011] The SHAP values of the screened feature variables are obtained through the SHAP interpretable machine learning algorithm. Based on the raster data of the feature variables, the spatial distribution maps of the SHAP values of each feature variable are generated at the pixel level. The spatial distribution maps of the SHAP values of all the screened feature variables are superimposed, and the feature variable corresponding to the maximum absolute value of the SHAP value at each pixel is used as the dominant factor of the pixel, thus completing the identification of the dominant factor.

[0012] Preferably, generating the spatial distribution maps of the SHAP values of each feature variable at the pixel level based on the raster data of the feature variables includes:

[0013] Based on the raster data of the feature variables, the SHAP values of the feature variables are mapped to each pixel through the OK interpolation method to form the spatial distribution map of the SHAP values.

[0014] Preferably, the absolute values of the SHAP values of the screened feature variables are sorted from large to small to obtain the importance ranking of the screened feature variables for predicting soil organic carbon density.

[0015] Preferably, multiple SHAP values within the value range of the feature variables participating in the spatial prediction modeling are obtained through the SHAP interpretable machine learning algorithm. Based on the changes and positive and negative properties of the multiple SHAP values within the value range, the influence degree and influence direction of the screened feature variables on soil organic carbon density are judged.

[0016] Preferably, the obtained covariates are converted into raster format data through spatial preprocessing, including:

[0017] Using ArcGIS 10.8 software, the original raster data of the covariates is drawn by the OK interpolation method, the original raster data of the covariates is projected onto the Universal Transverse Mercator reference system, and the original raster data is resampled to the set resolution by the nearest neighbor resampling method, and then the ranges of each covariate data are unified to the same spatial range by extraction by mask.

[0018] Preferably, before performing feature screening on the covariates converted into raster format through the RFE algorithm, the covariates data converted into raster format and the soil organic carbon density are divided into a training sample set and a validation sample set, and the raster data of the feature variables is obtained by performing feature screening on the training sample set through the RFE algorithm.

[0019] Preferably, verifying the accuracy of the trained RF model through the validation sample set includes:

[0020] The spatial distribution of the predicted soil organic carbon density is exported in raster form to the ArcGIS 10.8 software. The predicted soil organic carbon density corresponding to each point in the validation set is extracted from the spatial distribution of the predicted soil organic carbon density through the Extract to Points tool to obtain the model prediction values, and the accuracy of the observed values of the actually measured soil organic carbon density in the validation sample set and the model prediction values is verified.

[0021] Preferably, the raster data of the characteristic variables and the training RF model with soil organic carbon density include:

[0022] Taking the characteristic variables, the raster data of the characteristic variables, and the soil organic carbon density as inputs, through training the RF model, the mapping relationship between the characteristic variables and the soil organic carbon density is obtained, and the spatial distribution of the predicted soil organic carbon density is obtained through the trained RF model.

[0023] Preferably, the obtained soil organic matter content is converted into soil organic carbon density. Among them, the soil organic carbon density SOCD of the i-th sampling point i is:

[0024]

[0025] Among them, SOCD i is the soil organic carbon density (kg m -2 ), 0.58 is the conversion coefficient between the soil organic matter content and the soil organic carbon content, S i is the soil organic matter content (g kg -1 ), B i is the soil bulk density (g cm -3 ), H i is the sampling depth (cm), n is the number of samplings. The collected soil samples are air-dried, ground, and passed through a sieve, and then the soil organic matter content of the soil samples is measured.

[0026] Compared with the prior art, the beneficial effects of the present invention are:

[0027] The present invention uses the raster data of the characteristic variables as input to train the RF model, enabling the trained RF model to correspondingly output the spatial distribution of the predicted soil organic carbon density, realizing high-resolution prediction of SOCD in space. At the same time, the present invention also superimposes the spatial distribution maps of the SHAP values of all characteristic variables, and can obtain the SHAP values of each characteristic variable on each pixel, and then obtain the dominant factor of each pixel, so as to obtain the specific influence of the characteristic variables in different regions on the spatial distribution of SOCD, providing explanatory analysis at the spatial level. Description of the Drawings

[0028] Figure 1Logic block diagram of a method for identifying dominant factors and high-resolution mapping of soil organic carbon density spatial non-stationarity provided by a specific embodiment of the present invention;

[0029] Figure 2 Result graph of variable screening by the RFE algorithm provided by a specific embodiment of the present invention;

[0030] Figure 3 Spatial prediction result graph of 30m resolution SOCD by the RF model provided by a specific embodiment of the present invention;

[0031] Figure 4 Prediction accuracy verification graph of 30m resolution SOCD by the RF model provided by a specific embodiment of the present invention;

[0032] Figure 5 Global importance ranking graph of characteristic variables provided by a specific embodiment of the present invention;

[0033] Figure 6 Partial dependence graph of characteristic variables provided by a specific embodiment of the present invention;

[0034] Figure 7 Spatial distribution graph of SHAP values of characteristic variables provided by a specific embodiment of the present invention;

[0035] Figure 8 Spatial distribution graph of dominant factors provided by a specific embodiment of the present invention. Detailed implementation manners

[0036] The present invention will be further described below in conjunction with specific operation examples.

[0037] The method of the present invention can make full use of multi-source data, consider the comprehensive influence of various environmental factors, and adopt advanced machine learning algorithms and interpretability technologies to reveal the spatial variation characteristics, the role of influencing factors, and identify dominant factors of SOCD, which can provide a scientific basis and technical support for the formulation of farmland management policies and the implementation of soil protection measures.

[0038] A specific embodiment of the present invention provides a method for identifying dominant factors and high-resolution mapping of soil organic carbon density spatial non-stationarity, as Figure 1 shown, including:

[0039] S1. Convert the obtained soil organic matter content into soil organic carbon density, and convert the obtained covariates into raster format data through spatial preprocessing, where the covariates include soil attribute data and environmental covariate data.

[0040] In the specific embodiments of the present invention, the obtained soil organic matter content is converted into soil organic carbon density, and the obtained covariates are converted into raster format data through spatial preprocessing. The covariates include soil attribute data and environmental covariate data.

[0041] In a specific embodiment, the method for obtaining soil organic matter content and covariate data provided in this embodiment includes:

[0042] Measure the relevant attribute contents of all soil samples using traditional chemical determination methods. Among them, pH is measured using a 1:1 suspension of soil sample and water, the soil organic matter (SOM) content of the soil sample is analyzed using the potassium dichromate volumetric method - external heating method, the TN content is measured using the semi-micro Kjeldahl method. Among the covariate data: soil type (SG), clay, sand, and silt data are downloaded from the Data Center for Resources and Environmental Sciences, Chinese Academy of Sciences (https: / / www.resdc.cn), with a spatial resolution of 1 km. The climate indicators and biological indicators of MAP, MAT, NPP, and NDVI of the soil formation factors are also downloaded from RESDC, with a resolution of 1 km (https: / / www.resdc.cn). Record PLD (cm) during the on-site soil sampling process. Based on the SRTM DEM data of CIAT (https: / / srtm.csi.cgiar.org), use SAGA GIS software (https: / / saga-gis) to calculate the relevant topographic factors with a spatial resolution of 30 m in the example area, including TWI, TPI, VD, multi-resolution valley flatness (MrVBF), Slope, Aspect, SL, and CI.

[0043] In a specific embodiment, the conversion of the obtained soil organic matter content into soil organic carbon density and the conversion of the obtained covariates into raster format data through spatial preprocessing provided in this embodiment include:

[0044] Convert the obtained soil organic matter content into soil organic carbon density data, perform spatial preprocessing on the covariate data using ArcGIS 10.8 software, and process it into raster format data. Divide the soil organic carbon density and the raster format covariate data into a training sample set and a validation sample set. The raster format covariate data is the covariate data under spatial coordinates.

[0045] Among them, the soil organic carbon density (SOCD) content provided in the specific embodiments of the present invention is:

[0046]

[0047] Among them, SOCD i is the soil organic carbon density (kg m -2), 0.58 is the conversion coefficient between soil organic matter content and soil organic carbon content, S i is the soil organic matter content (g kg -1 ), B i is the soil bulk density (g cm -3 ), H i is the sampling depth (cm), n is the number of samples. The collected soil samples are air-dried, ground and passed through a sieve, and then the soil organic matter content of the soil samples is measured.

[0048] In a specific embodiment, the obtained covariates are converted into raster format data through spatial preprocessing, including:

[0049] In this embodiment, the original raster data of the soil attribute data in the covariates is drawn by using the OK interpolation method with ArcGIS 10.8 software. The original raster data of the soil attribute data is projected onto the Universal Transverse Mercator reference system, and the raster data is resampled to the set resolution by using the nearest neighbor resampling method. Then, the range of each covariate data is unified to the same spatial range by using extract by mask.

[0050] In this embodiment, the original raster data of the environmental covariate data in the covariates is projected onto the Universal Transverse Mercator reference system, and the raster data is resampled to the set resolution by using the nearest neighbor resampling method. Then, the range of each covariate data is unified to the same spatial range by using extract by mask. The original raster data of the environmental covariate data provided by the specific embodiment of the present invention can be directly obtained from public channels.

[0051] In an embodiment, in this embodiment, ArcGIS 10.8 software is used to draw BD, CEC, pH, TN, TK, TP, SEI, EMg, ECa, ASi, AS, and PLD of the entire sample area by using the OK interpolation method. Then, the original raster data of all environmental covariates is re-projected onto the Universal Transverse Mercator reference system (WGS1984 UTM Zone 49N), and the raster data is resampled to a 30 m resolution by using the nearest neighbor resampling method. Finally, the range of each environmental covariate data is unified to the same spatial range by using extract by mask. The specific soil attributes and environmental covariate data are shown in Table 1.

[0052] S2. In the specific embodiment of the present invention, the covariates converted into raster format are subjected to feature screening by the RFE algorithm to obtain raster data of feature variables: In the specific embodiment of the present invention, the RFE algorithm is combined with 10-fold cross-validation for feature selection, and the minimum RMSE is used as the standard for establishing the optimal feature variable set.

[0053] In a specific embodiment, before performing feature screening on the covariates converted into a raster format through the RFE algorithm, the covariate data converted into a raster format is divided into a training sample set and a validation sample set. In one embodiment, the ratio of the training sample set to the validation sample set is 7:3. Feature screening is performed on the training sample set through the RFE algorithm to obtain raster data of feature variables, and an optimal covariate data set is established. The specific results are as Figure 2 shown in Table 1 and Table 2. Figure 2 In the figure, the abscissa is the number of feature variables included in each model during the feature variable screening process, and the ordinate is the RMSE of cross-validation of each model. It can be seen from the figure that when the number of feature variables included in the modeling is 15, the RMSE is the smallest (0.76 kg m-2). Therefore, the optimal covariate data set screened by the RFE algorithm contains 15 feature variables, as specifically shown in Table 2.

[0054] S3. In a specific embodiment of the present invention, an RF model is trained based on feature variables, raster data of feature variables, and soil organic carbon density. The spatial distribution of predicted soil organic carbon density is obtained through the trained RF model: taking feature variables, raster data of feature variables, and soil organic carbon density as inputs, an RF model is trained to obtain the mapping relationship between feature variables and soil organic carbon density. Through the raster data of feature variables, the spatial distribution of predicted soil organic carbon density can be obtained. Based on this mapping relationship, the spatial distribution of soil organic carbon density corresponding to the raster data of feature variables can be obtained. Therefore, in a specific embodiment of the present invention, the spatial distribution of predicted soil organic carbon density can be accurately obtained through the trained RF model.

[0055] In a specific embodiment, through R v4.3.3 software, using the raster data of feature variables and soil organic carbon density as inputs of the RF model, and taking the spatial distribution data in the raster form of SOCD as the output of the model, a high-resolution spatial distribution model of SOCD is constructed (the specific parameters set for RF used are: ntree is 500; mtry is 4), and mapping work is carried out. The specific spatial prediction results of SOCD are shown in Figure 3 , generally speaking, the SOCD in the example area is between 2.00 - 6.58 kg m -2 -2. The SOCD spatial distribution map generated by the RF model shows strong spatial differentiation. Among them, the high-value areas of SOCD are mainly distributed in the southwestern, southern, and northeastern parts of the example area, while the low-value areas of SOCD are mainly distributed in the northern part.

[0056] In a specific implementation, the accuracy of the trained RF model is verified through a verification sample set, including: exporting the predicted spatial distribution of soil organic carbon density in a grid format to ArcGIS 10.8 software, extracting the predicted soil organic carbon density corresponding to each point in the verification set from the predicted spatial distribution of soil organic carbon density through the Extract Values to Points tool to obtain the model prediction values, and performing accuracy verification on the observed values of the actually measured soil organic carbon density and the model prediction values in the verification sample set.

[0057] In one embodiment, the accuracy verification of the spatial prediction results is carried out: R is selected 2 , RMSE, and LCCC are used as evaluation indicators for accuracy verification. The SOCD results predicted by the spatial model are exported to ArcGIS 10.8 software in a grid format, and then the predicted results are extracted to the verification set through the Extract Values to Points tool to obtain the model prediction values. Finally, accuracy verification is performed on the observed values and the predicted values in the verification set.

[0058] As Figure 4 shown, the abscissa represents the measured values of soil organic carbon density, the ordinate represents the predicted values of soil organic carbon density, and the dashed line represents the 1:1 fitting line. The results of the accuracy verification show that Figure 4 the spatial prediction of soil organic carbon density has good prediction accuracy and effect. Through RFE variable screening and the R of the spatial prediction of the RF model 2 reaches 0.46, the RMSE is as low as 0.64 kg m -2 , and the LCCC is 0.60.

[0059] S4. In the specific embodiment of the present invention, the SHAP values of the screened feature variables are obtained through the SHAP interpretable machine learning algorithm, and the SHAP value spatial distribution maps of each feature variable are generated at the pixel level based on the raster data of the feature variables.

[0060] In a specific embodiment, generating the SHAP value spatial distribution maps of each feature variable at the pixel level based on the raster data of the feature variables includes: mapping the SHAP values of the feature variables to each pixel through the OK interpolation method based on the raster data of the feature variables to form the SHAP value spatial distribution maps.

[0061] In a specific embodiment, the contributions of different characteristic variables to SOCD prediction were quantitatively analyzed: the SHAP values of the characteristic variables in all soil samples were calculated using the "kernelshap" package in R v4.3.3 software. These SHAP values represent the contribution or influence degree of each characteristic variable in predicting SOCD, and their positive and negative values respectively indicate the positive or negative influence of the characteristic variable on the predicted value of SOCD. By processing and sorting these SHAP values, a global importance ranking of the characteristic variables can be obtained, that is, which characteristic variables have the most significant influence on the change of SOCD. The specific importance ranking results are shown in Figure 5 .

[0062] As shown in Figure 5 a, the vertical axis represents different characteristic variables, the horizontal axis represents the SHAP values of the characteristic variables, and the depth of the dot color maps the level of the corresponding characteristic variable values. The pie chart represents the contribution ratio of each variable category to the change of SOCD. Figure 5 In b, the vertical axis represents different characteristic variables, and the horizontal axis represents the average value of the absolute values of the SHAP values of different characteristic variables. The order of the variables from top to bottom represents a gradual decrease in importance to the change of SOCD. Figure 5 The results in Figure 5 show that TN has the greatest influence on the change of SOCD, followed by BD, MAP, ASi, etc. Among the characteristic variable categories, soil properties contribute the most to SOCD mapping (87.9%) ( Figure 5 a), followed by climate (4.8%), topography (4.2%), biology (1.8%) and parent material (1.2%). According to the global contribution analysis and importance ranking diagram of SHAP ( Figure 5 b), the importance ranking of different environmental covariate categories for the change of SOCD in the RF model is: soil properties (1.45) > climate (0.08) > topography (0.07) > biology (0.03) > parent material (0.02).

[0063] S5. In the specific embodiment of the present invention, the spatial distribution maps of the SHAP values of all the screened characteristic variables are superimposed, and the characteristic variable corresponding to the maximum value of the absolute value of the SHAP value on each pixel is used as the dominant factor of the pixel, thereby completing the identification of the dominant factor.

[0064] In a specific embodiment, the spatial distribution maps of the SHAP values of all the screened characteristic variables are superimposed, and the characteristic variable corresponding to the maximum value of the absolute value of the SHAP value on each pixel is used as the dominant factor of the pixel, thereby completing the identification of the dominant factor.

[0065] In a specific embodiment, in this embodiment, multiple SHAP values of the selected feature variables within the value range are obtained through the SHAP interpretable machine learning algorithm. Based on the changes and positive / negative nature of the multiple SHAP values within the value range, the influence degree and influence direction of the selected feature variables on soil organic carbon density are judged.

[0066] In one embodiment, analysis of the influence degree and influence direction of feature variables: Based on the SHAP values of each calculated feature variable, further quantitatively analyze the influence degree and influence direction of each feature variable on the change of SOCD. By generating the partial dependence plot of each feature variable, it is possible to intuitively observe how each feature variable affects the change of SOCD when it changes within its value range. The specific analysis results are shown in Figure 6 .

[0067] As Figure 6 shown, each point represents the marginal contribution of a feature variable to the local prediction relative to the global average prediction (the vertical axis of the dependence plot is the SHAP value, the horizontal axis is the corresponding feature variable value, and the depth of the point color represents the magnitude of the feature variable value). The black curve on the partial dependence plot is the fitted line of the SHAP value, which is only for visualization purposes. Among the top six environmental covariates in terms of importance ranking, TN, BD, and ECa show a positive influence on SOCD prediction, while ASi and EMg show a negative influence on SOCD. MAP also has a positive influence on SOCD. This indicates that soil properties and climate are most closely related to the change of SOCD.

[0068] In one embodiment, this embodiment provides spatial analysis of feature variable influence and dominant factors: Based on the SHAP values of each feature variable, through the OK interpolation method in ArcGIS 10.8 software, the SHAP values of each feature variable are mapped to each pixel to reveal the influence of these feature variables on the change of SOCD in space. Then, the generated raster data of the spatial distribution of the SHAP values of the feature variables is input into R v4.3.3 software for overlay processing of the raster data. Based on the principle that the largest absolute SHAP value on each pixel is used as the dominant factor, the corresponding feature variable is identified as the dominant factor of that pixel. The specific results are shown in Figure 7 and Figure 8 .

[0069] As Figure 7 shown, a positive SHAP value indicates a positive influence of the feature variable on soil organic carbon density, and a negative value indicates a negative influence on soil organic carbon density. As Figure 8 shown, the pie chart and ring chart represent the proportion of different dominant factors in this area. From Figure 8It can be seen that in this area, TN accounts for 75.78%, followed by BD (20.49%), MAP (1.85%), and ECa (1.14%). The proportions of other variables are less than 1%, indicating that soil properties are the most important factors affecting the spatial distribution and change of SOCD.

[0070] The climate covariates provided in the specific embodiments of the present invention are the second most important factors for SOCD change, but their importance is much lower than that of soil property covariates ( Figure 6 ). In addition, for MAP, the spatial distribution pattern of the SHAP values of MAP in this area reflects the complex influence of precipitation on the spatial distribution of SOCD ( Figure 7 c). Specifically, except for a small part in the northwest where precipitation has a certain inhibitory effect on SOCD accumulation, most areas in the study area show positive areas, indicating that precipitation is one of the main positive factors promoting SOCD accumulation in the study area. In addition, topography, biology, and parent material also have a significant impact on the spatial variation of SOCD ( Figure 6 ), and they also play an indispensable role in the soil formation process, although their respective influence weights are relatively small, accounting for only 4.2%, 1.8%, and 1.2% respectively.

[0071] By plotting the dominant factors of SOCD in the specific embodiments of the present invention, the dominant factors that have the greatest impact on the change of SOCD in different regions can be clarified at the pixel level. For example, the SHAP value of TN on SOCD is negative in the north and positive in the south and east ( Figure 7 a). The SHAP value of TN increases with the increase of TN content ( Figure 6 a), which highlights the importance of increasing nitrogen content to improve carbon storage in this example area. Especially in the north of the example area, TN has a negative impact on SOCD and is an obstructive factor affecting carbon storage ( Figure 7 a, Figure 8 ). In addition, MAP is the primary dominant factor affecting the spatial distribution of SOCD in the northwest of the example area and has an obvious negative impact on SOCD ( Figure 7 c, Figure 8 ). Therefore, precipitation may be the main obstacle to improving carbon storage in this area. As shown in Figure 6 c, in the area where MAP is less than 1500 mm, the SHAP value of MAP on SOCD is negative, while in the area where MAP is greater than 1500 mm, the SHAP value of MAP on SOCD is positive.

[0072] In the specific implementation example of the present invention, a method for identifying the dominant factors of soil organic carbon density spatial non-stationarity and high-resolution mapping is established. This method not only provides a high-precision SOCD high-resolution spatial mapping product by integrating variable selection and machine learning methods, but also combines the interpretable machine learning (SHAP) method to quantitatively analyze the influencing factors and identify the dominant factors at the global, local, and pixel space levels, emphasizing the feasibility and reliability of using the SHAP method to map the spatial variation of SOCD. The framework developed in the specific embodiment of the present invention also provides a new tool for us to interpret the DSM results at the global and pixel scales.

[0073] Table 1 Soil property and environmental covariate data

[0074]

[0075] Table 2 Selection of optimal covariates by the RFE algorithm

[0076]

Claims

1. A method for identifying the dominant factors of spatial non-stationarity of soil organic carbon density and high-resolution mapping, characterized in that Including: Converting the obtained soil organic matter content into soil organic carbon density, converting the obtained covariates into raster format data through spatial preprocessing, where the covariates include soil attribute data and environmental covariate data, and performing feature screening on the covariates converted into raster format through the RFE algorithm to obtain feature variables and raster data of the feature variables; Training an RF model based on the feature variables, raster data of the feature variables, and soil organic carbon density, and obtaining the spatial distribution of predicted soil organic carbon density through the trained RF model; Obtaining the SHAP values of the selected feature variables through the SHAP interpretable machine learning algorithm, generating a spatial distribution map of the SHAP values of each feature variable at the pixel level based on the raster data of the feature variables, superimposing the spatial distribution maps of the SHAP values of all the selected feature variables, and taking the feature variable corresponding to the maximum absolute value of the SHAP value at each pixel as the dominant factor of the pixel, thereby completing the identification of the dominant factor.

2. The method for identifying the dominant factors of the spatial non-stationarity of soil organic carbon density and high-resolution mapping according to claim 1, wherein Generating a spatial distribution map of the SHAP values of each feature variable at the pixel level based on the raster data of the feature variables, including: Mapping the SHAP values of the feature variables to each pixel through the OK interpolation method based on the raster data of the feature variables to form a spatial distribution map of the SHAP values.

3. The method for identifying the dominant factors of soil organic carbon density spatial non-stationarity and high-resolution mapping according to claim 1, characterized in that, Sorting the absolute values of the SHAP values of the selected feature variables from large to small to obtain the importance ranking of the selected feature variables for predicting soil organic carbon density.

4. The method for identifying the dominant factors of soil organic carbon density spatial non-stationarity and high-resolution mapping according to claim 1, characterized in that Obtaining multiple SHAP values within the value range of the feature variables participating in the spatial prediction modeling through the SHAP interpretable machine learning algorithm, and judging the influence degree and influence direction of the selected feature variables on soil organic carbon density based on the change and positive and negative nature of the multiple SHAP values within the value range.

5. The method for identifying dominant factors of soil organic carbon density spatial non-stationarity and high-resolution mapping according to claim 1, characterized in that Converting the obtained covariates into raster format data through spatial preprocessing, including: Using ArcGIS 10.8 software to draw the original raster data of the covariates by the OK interpolation method, projecting the original raster data of the covariates to the Universal Transverse Mercator reference system, resampling the original raster data to the set resolution by the nearest neighbor resampling method, and then unifying the ranges of each covariate data to the same spatial range by extraction by mask.

6. The method for identifying dominant factors of soil organic carbon density spatial non-stationarity and high-resolution mapping according to claim 1, characterized in that Before performing feature screening on the covariates converted into raster format through the RFE algorithm, dividing the covariates data and soil organic carbon density converted into raster format into a training sample set and a validation sample set, and performing feature screening on the training sample set through the RFE algorithm to obtain raster data of the feature variables.

7. The method for identifying dominant factors of spatial non-stationarity of soil organic carbon density and high-resolution mapping according to claim 6, wherein Verifying the accuracy of the trained RF model through the validation sample set, including: Exporting the spatial distribution of the predicted soil organic carbon density in raster form to ArcGIS 10.8 software, extracting the predicted soil organic carbon density corresponding to each point in the validation set from the spatial distribution of the predicted soil organic carbon density through the extract to point tool to obtain the model prediction value, and performing accuracy verification on the observed value of the actually measured soil organic carbon density and the model prediction value in the validation sample set.

8. The method for identifying the dominant factors of spatial non-stationarity and high-resolution mapping of soil organic carbon density according to claim 1, characterized in that Training the RF model based on the raster data of the feature variables and the soil organic carbon density, including: Taking the characteristic variables, the raster data of the characteristic variables, and the soil organic carbon density as inputs, by training the RF model, the mapping relationship between the characteristic variables and the soil organic carbon density is obtained, and the spatial distribution of the predicted soil organic carbon density is obtained through the trained RF model.

9. The method for identifying dominant factors of spatial non-stationarity of soil organic carbon density and high-resolution mapping according to claim 1, wherein Convert the obtained soil organic matter content into soil organic carbon density. Among them, the soil organic carbon density SOCD at the i-th sampling point i is as follows: Among them, SOCD i is the soil organic carbon density, 0.58 is the conversion coefficient between soil organic matter content and soil organic carbon content, S i is the soil organic matter content, B i is the soil bulk density, H i is the sampling depth, n is the number of samples. The collected soil samples are air-dried, ground and passed through a sieve, and then the soil organic matter content of the soil samples is measured.

Citation Information

Cited By

  • Vegetation carbon sink space optimization method and system based on interpretable machine learning

    CN121390476A

  • Vegetation carbon sink space optimization method and system based on interpretable machine learning

    CN121390476B