Digital mapping method and device for soil organic carbon based on forward iterative variable screening
Through the method based on forward iterative variable screening, environmental variable sets are screened and modeled, which solves the problems of model complexity and low efficiency in digital soil mapping technology, and realizes efficient and accurate soil organic carbon distribution data characterization.
Patent Information
- Application Number
- CN202211503075.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-11-28
AI Technical Summary
The existing digital soil mapping technology has problems such as complex model structure, poor interpretability and large data calculation volume, which is difficult to meet the needs of efficient and accurate soil organic carbon distribution data.
The method based on forward iterative variable screening is adopted to filter the set of environmental variables, and important environmental variables are gradually screened out through the random forest model, and quantile regression forest model is constructed to improve the accuracy and efficiency of the model.
It realizes efficient and accurate characterization of soil organic carbon content and its uncertain spatial distribution, simplifies the model structure, and improves the mapping efficiency and accuracy.
Smart Images

Figure CN116106472B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to a soil organic carbon component prediction method, and in particular relates to a soil organic carbon digital mapping method and device based on forward iterative variable screening. Background Art
[0002] Due to the high temporal and spatial heterogeneity of soil, accurately grasping the fine distribution of soil information is the core of achieving sustainable development and utilization of soil resources. Traditional soil mapping is often completed by soil experts based on their personal experience and knowledge combined with the landscape information of soil sampling points. It has shortcomings such as long cycle, high cost, difficult updating, low accuracy, and lack of uncertainty analysis. With the continuous advancement of the digitalization process, the demand for real-time and accurate soil information distribution data in agricultural resource management and ecological environment modeling is increasing. Traditional soil mapping products can no longer meet today's urgent requirements, so it is urgent to develop a new technology to solve this problem.
[0003] Rooted in the theory of soil forming factors, the digital soil mapping technology proposed at the beginning of this century has great application potential in the field of fine characterization of soil information. Based on the correlation between soil sample data and environmental factors, this technology uses a quantitative soil environmental model to perform soil inference with the assistance of a computer to generate a soil type or property distribution map. Digital soil mapping has been gradually applied by countries around the world to the development of mapping products for various soil properties (such as organic matter, organic carbon, color, mechanical composition and pH) at different scales, and has made a series of research progress (Chen, S., Arrouays, D., Mulder, VL, Poggio, L., Minasny, B., Roudier, P., Libohova, Z., Lagacherie, P., Shi, Z., Hannam, J., Meersmans, J., Walter, C., 2022. Digital mapping of soil properties at a broadscale: A review. Geoderma, 409, 115567.)
[0004] With the continuous development of remote sensing technology, more and more remote sensing big data are being applied to digital soil mapping research. Poggio et al. (2021) (Poggio, L., de Sousa, LM, Batjes, NH, Heuvelink, GBM, Kempen, B., Ribeiro, E., and Rossiter, D.: SoilGrids 2.0: producing soil information for the globe with quantified spatial uncertainty, SOIL, 7, 217–240) used more than 400 environmental variables including vegetation, topography, and climate combined with a random forest model to draw a global 250-meter resolution distribution map of soil properties including soil organic carbon, total nitrogen, gravel, pH, etc. Liu et al. (2022) (Liu, F., Wu, H., Zhao, Y., Li, D., Yang, J.-L., Song, X., Shi, Z., Zhu, AX., Zhang, G.-L., 2022. Mapping high resolution National Soil Information Grids of China. Science Bulletin, 67(3), 328-340.) used more than 50 environmental variables including vegetation, topography, climate, and parent material combined with a random forest model to construct my country's first version of a 90-meter high-resolution national soil information grid.
[0005] However, while remote sensing big data improves quantitative soil environmental prediction models, it also brings disadvantages such as complex model structure, poor interpretability and large amount of data calculation. Therefore, it is urgent to develop a new digital soil mapping environmental variable screening algorithm to simplify the model structure and improve efficiency while ensuring the accuracy of the model. Summary of the invention
[0006] The present invention provides a soil organic carbon digital mapping method based on forward iterative variable screening. The method can efficiently and accurately characterize the soil organic carbon content and its uncertainty spatial distribution, thereby improving mapping efficiency.
[0007] A soil organic carbon digital mapping method based on forward iterative variable screening, comprising:
[0008] (1) obtaining a soil sample from the surface of the test area and determining the soil organic carbon content of the soil sample; deriving an initial environmental variable set based on the longitude and latitude of the soil sample through satellite remote sensing images, and resampling the initial environmental variable set through bilinear interpolation to obtain an environmental variable set with consistent spatial resolution;
[0009] (2) Using the forward iterative variable screening method to screen the environmental variable set, including:
[0010] (2.1) Remove different environmental variables from the environmental variable set in turn to obtain multiple groups of environmental variables. The difference between the accuracy of the random forest fitted by each group of environmental variables and the accuracy of the random forest fitted by the environmental variable set is used as the importance of the removed environmental variable. The environmental variable with the highest importance is marked as the selected environmental variable.
[0011] (2.2) Fitting the initial random forest model with the selected environmental variables and obtaining the accuracy of the initial random forest model through the ten-fold crossover method;
[0012] (2.3) Combining the selected environmental variables with each of the remaining environmental variables in the environmental variable set, fitting multiple sets of environmental variable combinations to obtain multiple random forest models, and comparing the accuracy of multiple random forest models to obtain the random forest model with the highest accuracy;
[0013] (2.4) marking the environmental variables that fit the random forest model with the highest accuracy as selected environmental variables;
[0014] (2.5) Iterate steps (2.3)-(2.4) until all environment variables in the environment variable set are marked as selected environment variables, then stop iterating to obtain the highest accuracy random forest model in each iteration;
[0015] (2.6) Compare the accuracy of the initial random forest model with the accuracy of the random forest model with the highest accuracy in each iteration, and take the environmental variables corresponding to the random forest model with the highest accuracy in the comparison results as the important environmental variables;
[0016] (3) The soil organic carbon content and the corresponding important environmental variables are used as sample data, and a sample data set is constructed based on multiple sample data. The quantile regression forest model is trained based on the sample data set to obtain a prediction model;
[0017] (4) When applied, the set of important environmental variables is input into the prediction model to obtain the large-scale soil organic carbon content and the uncertainty distribution of the large-scale soil organic carbon content.
[0018] The step of obtaining a soil sample from the surface of the area to be tested comprises:
[0019] Collect 0-30cm surface soil samples from the test area, determine the spatial position of soil sampling through random stratification based on comprehensive land cover, utilization and terrain, determine the sampling point based on the spatial position, collect multiple sub-samples through the diagonal method, and form a soil sample with the latitude and longitude of the point recorded as the latitude and longitude of the soil sample.
[0020] The method for determining the organic carbon content of the soil sample comprises: air-drying, grinding and sieving the soil sample, and then determining the organic carbon content of the soil by a combustion method.
[0021] The initial environmental variable set is derived from the longitude and latitude of the soil sample through satellite remote sensing images, and the initial environmental variable set includes climate variables, terrain variables and vegetation variables.
[0022] Said climate variables include mean annual precipitation, mean annual temperature, wind speed, vapor pressure, solar radiation and day / night surface temperature;
[0023] The terrain variables include elevation, slope, aspect, plan curvature, profile curvature and terrain moisture factor;
[0024] The vegetation variables include normalized difference vegetation index, enhanced vegetation index, gross primary productivity, net primary productivity and leaf area index.
[0025] The evaluation index of the accuracy of the random forest model is the root mean square error RMSE, and the root mean square error RMSE is:
[0026]
[0027] Among them, n is the number of cross-validation sample data, y i is the soil organic carbon content of the i-th sample data, is the predicted soil organic carbon content of the i-th sample generated by the random forest model, and each sample data corresponds to a set of environmental variables.
[0028] The quantile regression forest model generates multiple data sets with the same size as the training sample set through the bootstrap method for constructing multiple decision trees, randomly divides the environmental variable set into multiple environmental variable subsets, and in each decision tree, constructs the node branches of each decision tree by randomly dividing the environmental variable subsets; and optimizes the parameters used to divide the environmental variable subsets into the decision tree nodes through the grid search method combined with ten-fold cross validation to obtain a prediction model.
[0029] A soil organic carbon digital mapping device based on forward iterative variable screening, comprising a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor, wherein the computer memory is provided with a prediction model constructed by using the soil organic carbon digital mapping method based on forward iterative variable screening;
[0030] When the computer processor executes the computer program, the following steps are implemented:
[0031] The set of important environmental variables was input into the prediction model to obtain the large-scale soil organic carbon content and the uncertainty distribution of the large-scale soil organic carbon content.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] The present invention uses a forward iterative variable screening method to first obtain the most important environmental variable in the environmental variable set, and then sequentially combine the most important environmental variable with each environmental variable in the remaining environmental variable set, and obtain the most accurate random forest model in the random forest model fitted by each environmental variable combination, and mark the environmental variable combination corresponding to the most accurate random forest model as a selected environmental variable, complete one iteration, combine the selected environmental variable with the remaining environmental variables again in sequence, repeat the above steps, obtain new selected environmental variables and the most accurate random forest model of this iteration, complete the next iteration, and complete the iteration until all variables in the environmental variable set are marked as selected environmental variables, compare the accuracy of the most accurate random forest model obtained in each iteration, and use the most accurate environmental variable corresponding to the most accurate random forest model as an important environmental variable. The above method can take into account the accuracy of the random forest model fitted under different importance environmental variable combinations, thereby avoiding the removal of environmental variables that are less important but can still fit the random forest model with higher accuracy after being combined with other environmental variables, thereby improving the model prediction accuracy, and the method provided by the present invention can screen a large number of environmental variables to obtain important environmental variables with a small number of environmental variables, thereby improving the prediction accuracy and the prediction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A flow chart of a soil organic carbon digital mapping method based on forward iterative variable screening provided by an embodiment of the present invention;
[0035] Figure 2 The scatter plot of the measured value and the predicted value of the soil organic carbon content in the validation set provided by the embodiment of the present invention represents the accuracy of the soil organic carbon spatial prediction model estimated by the present invention;
[0036] Figure 3 This is the distribution map of organic carbon content in the surface soil of Northeast China prepared in Example 1 of the present invention. DETAILED DESCRIPTION
[0037] The present invention will be further described and illustrated below in conjunction with the accompanying drawings and specific implementation methods.
[0038] The present invention provides a soil organic carbon digital mapping method based on forward iterative variable screening, such as Figure 1 As shown, including:
[0039] (1) Obtain soil organic carbon content and environmental variable set of soil samples: Collect 0-30cm surface soil samples in the test area. Generally, the spatial location of soil sampling is determined by random stratified sampling based on comprehensive land cover / use and terrain. Find the sampling point in the spatial location, generally collect 5 sub-samples by the diagonal method to form the final soil sample, and record the longitude and latitude of the center point of the sampling point. After the soil sample is air-dried, ground, and sieved with a 2mm hole, the soil organic carbon content is determined by the combustion method to form a soil data set.
[0040] Climate, topography and vegetation environmental variables derived from satellite remote sensing images were collected based on the latitude and longitude of soil samples. Climate variables included average annual precipitation, average annual temperature, wind speed, vapor pressure, solar radiation and day / night surface temperature. Topographic variables included elevation, slope, aspect, plan curvature, profile curvature and topographic moisture factor. Vegetation variables included normalized difference vegetation index, enhanced vegetation index, gross primary productivity, net primary productivity and leaf area index.
[0041] The average annual precipitation, average annual temperature, wind speed, vapor pressure and solar radiation in the climate variables provided by the present invention are derived from the WorldClim Version2 global 1 km climate data product (http: / / www.worldclim.com / version2), and the day / night surface temperature is derived from the MODIS surface temperature data (https: / / modis.gsfc.nasa.gov / data / dataprod / mod11.php), with a spatial resolution of 1 km. The elevation data in the terrain variables are derived from the SRTM 90-meter digital elevation data (https: / / cgiarcsi.community / data / srtm-90m-digital-elevation-database-v4-1 / ), and the slope, aspect, plan curvature, profile curvature and terrain humidity factor are calculated through the elevation data. The normalized difference vegetation index and enhanced vegetation index in vegetation variables are calculated by 30-meter resolution Landsat 8 surface reflectance image, the total primary productivity and net primary productivity are derived from 1-kilometer resolution MODIS total primary productivity and net primary productivity (https: / / modis.gsfc.nasa.gov / data / dataprod / mod17.php), and the leaf area index is derived from MODIS leaf area index (https: / / modis.gsfc.nasa.gov / data / dataprod / mod15.php), with a spatial resolution of 1 km. Each vegetation variable calculates both the annual and seasonal averages.
[0042] Set the spatial resolution of the digital soil mapping product required for the area to be tested, resample the environmental variables with different spatial resolutions, and obtain raster data with consistent spatial resolution. In order to meet the needs of soil survey and management, the spatial resolution of the digital soil mapping product can be set to 90 meters or 30 meters, and the environmental variable resampling method is set to bilinear interpolation.
[0043] (2) Using the forward iterative variable screening method to screen the environmental variable set, including:
[0044] (2.1) Different environmental variables are removed from the environmental variable set in turn, that is, different single environmental variables are removed in turn to obtain new environmental variable combinations. After removing the single environmental variables in the environmental variable set in turn, multiple groups of environmental variables are obtained. The difference between the accuracy of the random forest fitted by each group of environmental variables and the accuracy of the random forest fitted by the environmental variable set is used as the importance of the removed environmental variables, and the environmental variable with the highest importance is marked as the selected environmental variable.
[0045] (2.2) Fitting the initial random forest model with the selected environmental variables and obtaining the accuracy of the initial random forest model through the ten-fold crossover method;
[0046] (2.3) Combining the selected environmental variables with each remaining environmental variable in the environmental variable set, fitting a random forest model through each environmental variable combination to obtain multiple random forest models, and obtaining the accuracy of multiple random forest models through a ten-fold crossover method; comparing the accuracy of multiple random forest models to obtain the random forest model with the highest accuracy;
[0047] (2.4) marking the environmental variables that fit the random forest model with the highest accuracy as selected environmental variables;
[0048] (2.5) Iterate steps (2.3)-(2.4) until all variables in the environment variable set are marked as selected environment variables and the iteration stops, and the highest accuracy random forest model is obtained in each iteration;
[0049] (2.6) Compare the accuracy of the initial random forest model with the accuracy of the random forest model with the highest accuracy in each iteration, and take the environmental variables corresponding to the random forest model with the highest accuracy in the comparison results as the important environmental variables;
[0050] The accuracy evaluation index of the random forest model provided by the present invention is the root mean square error RMSE, and the root mean square error RMSE is:
[0051]
[0052] Among them, n is the number of cross-validation sample data, y i is the soil organic carbon content of the i-th sample data, is the predicted soil organic carbon content generated by the random forest model for the i-th sample data, and each sample data corresponds to a set of environmental variables.
[0053] (3) Constructing a sample data set, which includes a modeling set and an independent validation set: soil organic carbon content and the corresponding important environmental variables are used as sample data, and multiple sample data are used to construct a sample data set, which is divided into a modeling set and an independent validation set. Through random sampling, 70% of the data in the soil data set are assigned to the modeling set, and the remaining 30% of the data are assigned to the independent validation set. These sample data all contain soil organic carbon content and important environmental variables screened out by the forward iterative variable screening method.
[0054] (4) The prediction model was obtained by training the quantile regression forest model based on the sample data set: the soil organic carbon content in the modeling set and the corresponding important environmental variables were used as training data to train the quantile regression forest model. The prediction model was obtained by optimizing the quantile regression forest model parameters through ten-fold cross validation.
[0055] The quantile regression forest model provided by the present invention generates multiple data sets with the same size as the training sample set through the bootstrap method for constructing multiple decision trees, randomly divides the environmental variable set into multiple environmental variable subsets, and in each decision tree, constructs the node branches of each decision tree by randomly dividing the environmental variable subsets; and optimizes the parameters used to divide the environmental variable subsets into the decision tree nodes through the grid search method combined with ten-fold cross validation to obtain a prediction model.
[0056] The verification random forest prediction model provided by the present invention: the established prediction model is used to predict an independent verification set, the soil organic carbon content in the independent verification set is compared with the predicted soil organic carbon content, and the prediction accuracy of the random forest conversion function is evaluated; when the prediction accuracy meets the standard, the established prediction model can be used to predict the soil organic carbon content.
[0057] (5) The fitted prediction model is applied to the corresponding important environmental variables to obtain a large-scale soil organic carbon content and its uncertainty distribution map. The prediction model ultimately outputs three soil organic carbon digital mapping products: average, 5th percentile, and 95th percentile.
[0058] The present invention also provides a soil organic carbon digital mapping device based on forward iterative variable screening, comprising a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor, wherein the computer memory adopts a prediction model constructed by the soil organic carbon digital mapping method based on forward iterative variable screening;
[0059] When the computer processor executes the computer program, the following steps are implemented:
[0060] The set of important environmental variables was input into the prediction model to obtain the large-scale soil organic carbon content and the uncertainty distribution of the large-scale soil organic carbon content.
[0061] The forward iterative variable screening algorithm combined with digital soil mapping technology proposed in the present invention can efficiently and accurately characterize the spatial distribution of soil organic carbon content and its uncertainty, which not only improves the mapping efficiency and simplifies the model structure, but also improves the mapping efficiency while improving the model accuracy. It provides a new idea for the precise spatiotemporal guarantee of large-scale and high-precision soil information, is beneficial to the sustainable utilization and management of soil resources, and has certain theoretical and practical significance and promotion and application value.
[0062] Example 1
[0063] In this embodiment, Northeast my country is selected as the research area, and 400 surface soil sample data collected in 2017 are used. The vegetation, topography and climate environmental variables obtained by remote sensing technology are optimized using a forward iterative variable screening algorithm, and a quantile random forest model is constructed to finally obtain the surface soil organic carbon content and its uncertainty distribution with a spatial resolution of 90 meters. The basic steps of this method are as described in steps (1) to (6) of the above embodiment, and will not be fully repeated. The following mainly shows the specific data and implementation details:
[0064] Step (1): Surface soil samples were collected during the off-season; during the grinding and sieving process of mixed samples, straw, plant roots and other exogenous substances that may affect the determination of soil organic carbon should be carefully removed.
[0065] Step (2): WorldClim Version2 in the climate and environmental variables uses data from 60,000 climate stations around the world, combined with elevation, distance from the coast, maximum / minimum surface temperature and cloud cover data, and uses thin plate spline interpolation to obtain global 1 km resolution annual average precipitation, annual average temperature, wind speed, vapor pressure and solar radiation during 1970-2000. The cross-validation correlation of these products is between 0.76 and 0.99. The official download link for this data is http: / / www.worldclim.com / version2. The Moderate Resolution Imaging Spectroradiometer (MODIS) surface temperature, gross primary productivity and net primary productivity, and leaf area index products used in the climate and vegetation environmental variables were all collected in 2017. This data can be downloaded from Google Earth Engine (GEE). MODIS is a medium-resolution imaging spectrometer carried on the Terra and Aqua satellites. It is an important instrument in the U.S. Earth Observing System (EOS) program for observing global biological and physical processes. It has an orbital altitude of 705 kilometers, a scan width of 2,330 kilometers, a revisit period of 16 days, 36 medium-resolution spectral bands (0.4-14.4 microns), and a spatial resolution of 250-1000 meters. The Landsat 8 reflectance data in the vegetation environmental variables were collected in 2017. After removing images with cloud cover greater than 20%, the annual and quarterly averages of vegetation variables were calculated. The data can be downloaded in GEE. Landsat 8 is a satellite launched by the National Aeronautics and Space Administration (NASA). It has an orbital altitude of 705 kilometers, a scan width of 185 kilometers, a revisit period of 16 days, and carries an OLI land imager with 9 bands (0.43-11.19 microns) and a spatial resolution of 30 meters.
[0066] Step (3): All environmental variables were resampled to 90 meters. This function was implemented in R language using the resample function of the raster package, and the resampling method was set to bilinear interpolation.
[0067] Step (4): Extraction of environmental variables at soil sampling points was implemented in R language using the extract function of the raster package. The importance of factors in step (a) of the forward iterative variable screening algorithm was evaluated by %IncMSE, which was obtained by calculating the mean square error (MSE) of the fitted model after removing each variable and the percentage increase in the mean square error of the fitted model of all variables. The larger the %IncMSE, the more important the variable. This step was implemented in R using the importance function of the randomForest package, and the method selected was mean decrease in accuracy. The ten-fold cross validation in steps (b) and (c) of the forward iterative variable screening algorithm was implemented in R language using the train function of the caret package.
[0068] Step (5): The quantile random forest conversion function is implemented in R language using the quantregForest function of the quantregForest package. The number of trees is set to the default value of 500, and the number of predictor variables for decision tree node partitioning is optimized to 5.
[0069] Step (6): The large-scale soil organic carbon content and its uncertainty distribution map are drawn using the predict function of the quantregForest package in R language, and the final TIFF format file is output using the writeRaster function of the raster package.
[0070] In this example, the data is divided into a modeling set (300) and an independent validation set (100). The modeling set data is used to fit the prediction model, and the independent validation set is used to objectively evaluate the accuracy of the model. 2 ) and root mean square error (RMSE) are used to evaluate the prediction accuracy of the independent validation set. The number of environmental variables after the forward iteration variable screening in this embodiment is 9, which greatly simplifies the model compared to the 38 environmental variables before screening. The prediction model evaluation accuracy after the forward iteration variable screening in this embodiment is as follows Figure 2 shown by Figure 2 It can be seen that the R of the independent validation set 2 is 0.60, and the RMSE is 5.52 g kg -1 , which has a good prediction effect and is better than the prediction model fitted by the first 38 environmental variables (R 2 is 0.56, and the RMSE is 6.02 g kg -1 ). In step (6) of mapping the soil organic carbon content and its uncertainty distribution map, the present embodiment takes a total of 34 minutes, which is much better than the original 120 minutes for 38 environmental variables, greatly improving the mapping efficiency.
[0071] Northeast my country was selected as the study area. The surface soil organic carbon content (average soil organic carbon digital mapping) and its uncertainty spatial distribution (5th and 95th percentile soil organic carbon digital mapping) predicted by this embodiment in 2017 are shown in the figure. Figure 3 As shown, the spatial resolution is 90 meters.
[0072] The above-described embodiment is only a preferred solution of the present invention, but it is not intended to limit the present invention. A person skilled in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.
Claims
1. A soil organic carbon digital mapping method based on forward iterative variable screening, characterized in that: include: (1) obtaining a soil sample from the surface of the test area and determining the soil organic carbon content of the soil sample; deriving an initial environmental variable set based on the longitude and latitude of the soil sample through satellite remote sensing images, and resampling the initial environmental variable set through bilinear interpolation to obtain an environmental variable set with consistent spatial resolution; (2) Using the forward iterative variable screening method to screen the environmental variable set, including: (2.1) Remove different environmental variables from the environmental variable set in turn to obtain multiple groups of environmental variables. The difference between the accuracy of the random forest fitted by each group of environmental variables and the accuracy of the random forest fitted by the environmental variable set is used as the importance of the removed environmental variable. The environmental variable with the highest importance is marked as the selected environmental variable. (2.2) Fitting the initial random forest model with the selected environmental variables and obtaining the accuracy of the initial random forest model through the ten-fold crossover method; (2.3) Combining the selected environmental variables with each of the remaining environmental variables in the environmental variable set, fitting multiple sets of environmental variable combinations to obtain multiple random forest models, and comparing the accuracy of multiple random forest models to obtain the random forest model with the highest accuracy; (2.4) marking the environmental variables that fit the random forest model with the highest accuracy as selected environmental variables; (2.5) Iterate steps (2.3)-(2.4) until all environment variables in the environment variable set are marked as selected environment variables, then stop iterating to obtain the highest accuracy random forest model in each iteration; (2.6) Compare the accuracy of the initial random forest model with the accuracy of the random forest model with the highest accuracy in each iteration, and take the environmental variables corresponding to the random forest model with the highest accuracy in the comparison results as the important environmental variables; (3) The soil organic carbon content and the corresponding important environmental variables are used as sample data, and a sample data set is constructed based on multiple sample data. The quantile regression forest model is trained based on the sample data set to obtain a prediction model; (4) When applied, the set of important environmental variables is input into the prediction model to obtain the large-scale soil organic carbon content and the uncertainty distribution of the large-scale soil organic carbon content.
2. The soil organic carbon digital mapping method based on forward iterative variable screening according to claim 1 is characterized in that: The step of obtaining a soil sample from the surface of the area to be tested comprises: Collect 0-30cm surface soil samples from the test area, determine the spatial position of soil sampling through random stratification based on comprehensive land cover, utilization and terrain, determine the sampling point based on the spatial position, collect multiple sub-samples through the diagonal method, and form a soil sample with the latitude and longitude of the point recorded as the latitude and longitude of the soil sample.
3. The soil organic carbon digital mapping method based on forward iterative variable screening according to claim 1 is characterized in that: The method for determining the soil organic carbon content of the soil sample comprises: air-drying, grinding and sieving the soil sample, and then determining the soil organic carbon content by a combustion method.
4. The soil organic carbon digital mapping method based on forward iterative variable screening according to claim 1 is characterized in that: The initial environmental variable set is derived from the longitude and latitude of the soil sample through satellite remote sensing images, and the initial environmental variable set includes climate variables, terrain variables and vegetation variables.
5. The soil organic carbon digital mapping method based on forward iterative variable screening according to claim 4 is characterized in that: Said climate variables include mean annual precipitation, mean annual temperature, wind speed, vapor pressure, solar radiation and day / night surface temperature; The terrain variables include elevation, slope, aspect, plan curvature, profile curvature and terrain moisture factor; The vegetation variables include normalized difference vegetation index, enhanced vegetation index, gross primary productivity, net primary productivity and leaf area index.
6. The soil organic carbon digital mapping method based on forward iterative variable screening according to claim 1, characterized in that: The evaluation index of the accuracy of the random forest model is the root mean square error RMSE, and the root mean square error RMSE is: Among them, n is the number of cross-validation sample data, y i is the soil organic carbon content of the i-th sample data, is the predicted soil organic carbon content of the i-th sample generated by the random forest model, and each sample data corresponds to a set of environmental variables.
7. The soil organic carbon digital mapping method based on forward iterative variable screening according to claim 1, characterized in that: The quantile regression forest model generates multiple data sets with the same size as the training sample set through the bootstrap method for constructing multiple decision trees, randomly divides the environmental variable set into multiple environmental variable subsets, and in each decision tree, constructs the node branches of each decision tree by randomly dividing the environmental variable subsets; and optimizes the parameters used to divide the environmental variable subsets into the decision tree nodes through the grid search method combined with ten-fold cross validation to obtain a prediction model.
8. A soil organic carbon digital mapping device based on forward iterative variable screening, comprising a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor, characterized in that: The computer memory contains a prediction model constructed by using a soil organic carbon digital mapping method based on forward iterative variable screening as described in any one of claims 1 to 7; When the computer processor executes the computer program, the following steps are implemented: The set of important environmental variables was input into the prediction model to obtain the large-scale soil organic carbon content and the uncertainty distribution of the large-scale soil organic carbon content.
Citation Information
Patent Citations
Soil organic carbon content prediction method based on random forest-ordinary Kriging method
CN109342697A
Soil nutrient prediction and comprehensive evaluation method based on machine learning algorithm
CN109374860A