A digital soil mapping method and device based on soil spatial neighborhood information

By introducing soil spatial neighborhood information into digital soil mapping, using the European-style distance and inverse distance weighting method to obtain spatial neighborhood soil attribute data, and combining environmental variables to train a random forest model, the problem of insufficient prediction accuracy of soil attribute data in the existing technology is solved, and a higher precision digital soil map is achieved.

CN118537441BActive Publication Date: 2025-06-13ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410661656.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-06-13
Estimated Expiration
2044-05-27

AI Technical Summary

Technical Problem

The existing digital soil mapping method still has room for improvement in the prediction accuracy of soil attribute data.

Method used

By using soil spatial neighborhood information, the Euclidean distance between the soil sample to be tested and other soil samples is calculated, the set of spatial neighborhood soil samples near the front k is obtained, and the spatial neighborhood soil attribute data is obtained through the inverse distance weighting method. Combining environmental variables and spatial neighborhood soil attribute data, a random forest model is trained to predict soil attribute data.

Benefits of technology

It significantly improves the prediction accuracy of soil attribute data and achieves more accurate digital soil mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118537441B_ABST
    Figure CN118537441B_ABST
Patent Text Reader

Abstract

The present invention discloses a digital soil mapping method and device based on soil spatial neighborhood information. The method includes obtaining a set of spatially neighboring soil samples that are the top k closest in spatial distance to the soil sample to be measured by comparing the Euclidean distances between the soil sample to be measured and other soil samples, and averaging the soil attribute data of each spatially neighboring soil sample after inverse distance weighting to obtain spatially neighboring soil attribute data; training a random forest model based on the environmental variable set of each training sample in the training sample set and the corresponding spatially neighboring soil attribute data to obtain a soil attribute data prediction model, inputting the environmental variable set of the area to be measured into the soil attribute data prediction model to obtain predicted soil attribute data, and realizing digital soil mapping of the area to be measured based on the predicted soil attribute data. This method can accurately predict soil attribute data, and thus can draw a more accurate digital soil map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of soil mapping, and particularly relates to a digital soil mapping method and device based on soil spatial neighborhood information. Background Art

[0002] Soil, as the cornerstone of all terrestrial ecosystems, is facing significant threats from anthropogenic disturbances including climate change, land use change, pollution, and erosion. Therefore, timely updating of fine-resolution soil information is crucial for protecting limited soil resources and making decisions on their green and sustainable development.

[0003] Digital soil mapping has become an important tool for providing high-quality cross-scale spatio-temporal soil information and quantifying its uncertainty. Improving the prediction accuracy of digital soil maps is crucial for promoting the practical application of digital soil mapping in decision-making. Therefore, a large number of studies have been devoted to this research topic, reflecting its importance in this field.

[0004] Deeply exploring environmental variables related to target soil properties is one of the feasible solutions to improve the accuracy of digital soil mapping products. In the past decade, a series of progress has been made in characterizing new environmental variables, including the description of agricultural management, geomorphic variables, and three-dimensional spatio-temporal covariates.

[0005] The patent application of invention with the publication number CN103955953A discloses a method for selecting terrain co-variables for digital soil mapping, which adopts a selection method of multiple terrain factors and multiple algorithms, uses a functional test strategy to preprocess and select different terrain factor variables, realizes the rapid and accurate selection of complex terrain factor variables through combining its correlation mechanism with soil properties, and adopts the technology of "evaluation analysis as the main, correlation analysis as the auxiliary", realizing a quantitative digital soil mapping terrain factor variable selection system of "general selection mechanism for different terrain factor variables; dynamic factor screening for different dependency relationships; evaluation control strategy, taking into account algorithm performance", and has broad industrial application prospects.

[0006] The invention patent application with the publication number CN116206011A discloses a digital soil mapping method and system based on multi-source data. The method includes obtaining the soil use type and satellite remote sensing image raster data of the area to be measured, dividing the area to be measured into multiple soil areas, where each soil use type includes at least one soil area, identifying all soil areas to obtain the topographic features of each soil area and the distribution characteristics of soil areas of the same soil use type according to the identification results; determining whether the soil areas of the same soil use type are adjacent; if not, obtaining the sample points of each soil unit, and obtaining the sample data of each sample point according to the sample points for soil mapping. By dividing the area to be measured into multiple soil areas and then respectively determining the sample points of each soil area and the sample data corresponding to the sample points, the obtained sample data is more representative, and the prediction accuracy of the prediction model is improved.

[0007] However, there is still room for further improvement in the prediction accuracy of the soil attribute data provided by the above patent's digital soil mapping method. Summary of the Invention

[0008] The present invention provides a digital soil mapping method based on soil spatial neighborhood information, which can accurately predict soil attribute data and thus can more accurately realize digital soil mapping.

[0009] A specific embodiment of the present invention provides a digital soil mapping method based on soil spatial neighborhood information, including:

[0010] By comparing the Euclidean distance between the soil sample to be measured and other soil samples, a set of spatially neighboring soil samples that are the top k closest in spatial distance to the soil sample to be measured is obtained, and the soil attribute data of each spatially neighboring soil sample is inverse distance weighted and then averaged to obtain spatially neighboring soil attribute data;

[0011] Based on the environmental variable set of each training sample in the training sample set and the corresponding spatially neighboring soil attribute data, a random forest model is trained to obtain a soil attribute data prediction model, and the training sample set comes from the set of soil samples to be measured constructed from the soil samples to be measured;

[0012] During application, the environmental variable set of the area to be measured is input into the soil attribute data prediction model to obtain predicted soil attribute data, and digital soil mapping of the area to be measured is realized based on the predicted soil attribute data.

[0013] Preferably, obtaining a set of spatially neighboring soil samples that are the top k closest in spatial distance to the soil sample to be measured by comparing the Euclidean distance between the soil sample to be measured and other soil samples includes:

[0014] Calculate the Euclidean distance between the soil sample to be measured and other soil samples based on the longitude and latitude of the soil sample to be measured and other soil samples;

[0015] Take the k other soil samples with the closest Euclidean distance as the spatial neighborhood soil sample set.

[0016] Preferably, the soil sample set to be measured further includes a verification sample set and the target soil attribute data corresponding to each verification sample;

[0017] Input the environmental variable set of the verification sample into the soil attribute data prediction model to obtain the predicted soil attribute data, and obtain the prediction accuracy based on the predicted soil attribute data and the target soil attribute data using the coefficient of determination, root mean square error, and concordance correlation coefficient;

[0018] When the prediction accuracy is lower than the set accuracy threshold, the corresponding soil attribute data prediction model is used as the final soil attribute data prediction model.

[0019] Preferably, the method for obtaining the soil sample to be measured includes the following steps:

[0020] Collect the soil surface layer samples of 0-20 cm in the area to be measured, and use the random stratified method to determine the spatial positions of the soil sampling points;

[0021] Collect the soil samples to be measured at the spatial positions: collect multiple sub-samples at each soil sampling point using the diagonal method, and mix these multiple sub-samples into a soil sample to be measured, and record the longitude and latitude information of the soil sampling points at the same time.

[0022] Preferably, the soil attribute data includes one or more of soil organic carbon density data, soil heavy metal content data, soil texture data, and soil gravel content data.

[0023] Preferably, the environmental variables include one or more of climate environmental variables, terrain environmental variables, spatial position environmental variables, and biological environmental variables.

[0024] Preferably, training a random forest model based on the environmental variable set of each training sample in the training sample set and the corresponding spatial neighborhood soil attribute data to obtain a soil attribute data prediction model includes:

[0025] Based on the environmental variable set of each training sample in the training sample set and the corresponding spatial neighborhood soil attribute data, use ten-fold cross-validation to determine the optimal number of trees and the number of branching variables of the random forest, and fit the corresponding relationship between the environmental variables and the predicted soil attribute data based on the determined optimal number of trees and the number of branching variables, so as to obtain the soil attribute data prediction model.

[0026] On the other hand, a digital soil mapping device based on soil spatial neighborhood information is also provided in a specific embodiment of the present invention, including:

[0027] A data processing module; configured to obtain a set of spatial neighborhood soil samples that are the top k closest in spatial distance to the soil sample to be measured by comparing the Euclidean distances between the soil sample to be measured and other soil samples, take the average of the soil attribute data of each spatial neighborhood soil sample after inverse distance weighting to obtain spatial neighborhood soil attribute data; train a random forest model based on the environmental variable set and the corresponding spatial neighborhood soil attribute data of each training sample in the training sample set to obtain a soil attribute data prediction model, where the training sample set comes from a set of soil samples to be measured constructed from the soil sample to be measured;

[0028] An input module: input the environmental variable set of the area to be measured into the soil attribute data prediction model to obtain predicted soil attribute data, and implement digital soil mapping of the area to be measured based on the predicted soil attribute data.

[0029] On the other hand, a digital soil mapping device based on soil spatial neighborhood information is also provided in a specific embodiment of the present invention, characterized by including a memory and one or more processors, where executable code is stored in the memory, and when the one or more processors execute the executable code, it is used to implement the digital soil mapping method based on soil spatial neighborhood information as described above.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] Different from directly fitting the corresponding relationship between environmental variables and soil attribute data in the prior art, the present invention simultaneously uses the spatial neighborhood soil attribute data and environmental variables of the soil adjacent to the soil to be measured for training the random forest model, thereby considering the influence of adjacent soil samples on the soil sample to be measured when training the random forest model, and then being able to accurately obtain the corresponding soil attribute data, realizing the accurate drawing of digital soil mapping. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flowchart of the digital soil mapping method based on soil spatial neighborhood information provided by a specific embodiment of the present invention;

[0033] Figure 2 It is a prediction accuracy graph of the organic carbon density of the soil to be measured by the digital soil mapping method based on soil spatial neighborhood information provided in Embodiment 1 of the present invention;

[0034] Figure 3 It is a prediction accuracy graph of the heavy metal lead in the soil to be measured by the digital soil mapping method based on soil spatial neighborhood information provided in Embodiment 2 of the present invention;

[0035] Figure 4The prediction accuracy map of heavy metal zinc in the soil to be measured by the digital soil mapping method based on soil spatial neighborhood information provided in Embodiment 2 of the present invention. Detailed implementation manners

[0036] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0037] In the prior art, there has been no research attempt to incorporate soil spatial neighborhood information into digital soil mapping. Except in the field of environmental variables, the soil property data at specific sampling points is inevitably correlated with the soil samples at neighboring sampling points. In order to further improve the accuracy and stability of digital soil mapping, the specific embodiments of the present invention attempt to introduce soil spatial neighborhood information enhancement technology into digital soil mapping. This technology can enhance the accuracy of the soil property data prediction model by capturing the information of neighboring soil samples and integrating it with the existing environmental variables into digital soil mapping.

[0038] Specific embodiments of the present invention propose a digital soil mapping method based on soil spatial neighborhood information, as Figure 1 shown, including the following steps:

[0039] (1) Obtain a set of soil samples to be measured, and divide the set of soil samples to be measured into a training set and a validation set:

[0040] Specific embodiments of the present invention provide a method for obtaining soil samples to be measured, including the following steps: Collect soil surface layer samples of 0-20 cm in the area to be measured, and determine the spatial positions of soil sampling points by using the random stratified method by comprehensively considering land cover and terrain factors; Collect soil samples to be measured at these spatial positions, and collect 5 sub-samples at each sampling point by using the diagonal method, and mix these 5 sub-samples into a soil sample to be measured, and the soil sample to be measured includes soil property data, an environmental variable set, and longitude and latitude; At the same time, record the longitude and latitude information of the sampling points as the geographical location identifier of the soil samples.

[0041] Specific embodiments of the present invention provide a method for the soil property data of the soil sample to be measured, including the following steps: Air-dry, grind, and sieve each collected soil sample to be measured, and then determine the soil property data by using standard laboratory physical and chemical analysis methods. The soil property data includes one or more of soil organic carbon density data, soil heavy metal content data, soil texture data, and soil gravel content data.

[0042] The environmental variable method for the soil sample to be measured provided by the specific embodiment of the present invention includes extracting satellite remote sensing images and products through the longitude and latitude of the soil sample to obtain an environmental variable set, and the environmental variables include one or more of climate environmental variables, terrain environmental variables, spatial position environmental variables, and biological environmental variables.

[0043] (2) Obtain a spatial neighborhood soil sample set by comparing the Euclidean distance between the soil sample to be measured and other soil samples, and take the average of the soil attribute data of each spatial neighborhood soil sample after inverse distance weighting. The specific description is as follows:

[0044] In the specific embodiment of the present invention, the Euclidean distance between the longitude and latitude of the soil sample to be measured and other soil samples is calculated. It can be understood that other soil samples are also collected in the area to be measured. By comparing the magnitudes of the Euclidean distances of the longitude and latitude, the top k soil samples that are closest to the soil sample to be measured in terms of spatial distance are obtained. These top k soil samples are used as the spatial neighborhood soil sample set, so as to obtain the soil samples adjacent to the soil sample to be measured. The inverse distance weighting is used to assign weights to the soil attribute data of each spatial neighborhood soil sample. The closer to the soil sample to be measured, the higher the weight, so as to realize the idea that the spatial neighborhood soil samples closer to the soil sample to be measured have a greater impact on the soil sample to be measured. The weighted soil attribute data is averaged to obtain the spatial neighborhood soil attribute data.

[0045] (3) Train a random forest model based on the environmental variable set of each training sample in the training sample set and the corresponding spatial neighborhood soil attribute data to obtain a soil attribute data prediction model. The specific description is as follows:

[0046] In the specific embodiment of the present invention, based on the environmental variable set of each training sample in the training sample set and the corresponding spatial neighborhood soil attribute data, ten-fold cross-validation is used to determine the optimal number of trees and the number of branching variables of the random forest. Based on the determined optimal number of trees and the number of branching variables, the corresponding relationship between the environmental variables and the predicted soil attribute data is fitted, so as to obtain a soil attribute data prediction model.

[0047] In the specific embodiment of the present invention, the spatial neighborhood soil attribute data is used as a parameter together with the environmental variables for training the random forest, so as to take into account the influence of the soil samples adjacent to the soil sample to be measured on the soil sample to be measured. Furthermore, the mapping relationship between the environmental variables and the soil attribute data is enhanced, so that the obtained soil attribute data prediction model can more accurately predict the soil attribute data.

[0048] (4) Verify the soil attribute data prediction model based on the validation sample set. If the verification result is less than the accuracy threshold, the final soil attribute data prediction model is obtained.

[0049] In a specific embodiment, the environmental variable set of the verification sample is input into the soil property data prediction model to obtain the predicted soil property data, and the prediction accuracy is obtained based on the predicted soil property data and the target soil property data by using the coefficient of determination, the root mean square error, and the concordance correlation coefficient; when the prediction accuracy is lower than the set accuracy threshold, the corresponding soil property data prediction model is used as the final soil property data prediction model.

[0050] The accuracy of the model provided by the present invention uses the coefficient of determination, the root mean square error, and the concordance correlation coefficient to evaluate the prediction accuracy of the prediction model. The coefficient of determination R 2 is:

[0051]

[0052] The root mean square error RMSE is:

[0053]

[0054] The concordance correlation coefficient LCCC is:

[0055]

[0056] where n is the number of verification samples, y i is the target soil property data of the i-th verification sample, is the predicted soil property data generated by the prediction model for the i-th sample, represents the covariance between and, var(y i ) and respectively represent the variances of y i and , μy i and respectively represent the means of and.

[0057] (5) A digital soil map is prepared based on the final soil property data prediction model. The environmental variable set of the area to be measured is input into the final soil property data prediction model to obtain the predicted soil property data, and the digital soil map of the area to be measured is realized based on the predicted soil property data.

[0058] On the other hand, a specific embodiment of the present invention also provides a digital soil mapping device based on soil spatial neighborhood information, including:

[0059] A data processing module; it is used to obtain a set of spatially neighboring soil samples that are the top k closest in spatial distance to the soil sample to be measured by comparing the Euclidean distances between the soil sample to be measured and other soil samples, and take the average of the soil property data of each spatially neighboring soil sample after inverse distance weighting to obtain spatially neighboring soil property data; based on the environmental variable sets of each training sample in the training sample set and the corresponding spatially neighboring soil property data, train a random forest model to obtain a soil property data prediction model, and the training sample set comes from the set of soil samples to be measured constructed from the soil samples to be measured.

[0060] An input module: input the environmental variable set of the area to be measured into the soil property data prediction model to obtain predicted soil property data, and implement digital soil mapping of the area to be measured based on the predicted soil property data.

[0061] On the other hand, a specific embodiment of the present invention further provides a digital soil mapping device based on soil spatial neighborhood information, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the digital soil mapping method based on soil spatial neighborhood information as claimed.

[0062] Embodiment 1

[0063] In this embodiment, France is selected as the research area. Using 916 surface soil sample data (0 - 20 cm) of the LUCAS Soil 2018 dataset, using other soil properties and environmental variables obtained through remote sensing technology, soil spatial neighborhood information is constructed through the inverse distance weighting method, and finally the prediction accuracy of the target soil property is improved. The specific description is as follows:

[0064] Step (1): Collect soil surface samples of 0 - 20 cm in the area to be measured. Determine the spatial positions of soil sampling points by random stratified sampling that comprehensively considers land cover and terrain. Based on the spatial positions, determine the sampling points. Collect 5 sub-samples by the diagonal method. The 5 sub-samples form 1 mixed soil sample, that is, the soil sample to be measured, and record the longitude and latitude of the sampling point of the soil sample to be measured as the longitude and latitude of the soil sample to be measured.

[0065] The soil property data of the soil sample to be measured provided in this embodiment is soil organic carbon density data. The calculation of this soil organic carbon density data includes soil bulk density and soil organic carbon content (soil organic carbon density = soil bulk density × soil organic carbon content × soil layer thickness / 100 × (1 - gravel content)).

[0066] The soil bulk density provided in this embodiment is obtained by using the core sampling method. In this method, a ring knife of known volume is used to obtain samples by vertically or horizontally knocking it into undisturbed soil, and then the samples are taken back to the laboratory for soil quality determination. The soil bulk density can be calculated by the following formula: Soil bulk density = soil sample mass / ring knife volume.

[0067] The organic carbon content of the soil samples provided in this embodiment includes: after air-drying, grinding, and sieving the soil samples, the organic carbon content of the soil is determined by the combustion method.

[0068] The gravel content of the soil provided in this embodiment includes: collecting soil samples and preparing wet sieving tools; separating the gravel in the soil samples using the wet sieving method; weighing the gravel and calculating its percentage content in the samples:

[0069] Step (2): The environmental variables provided in this embodiment include climate environmental variables, topographic factors, clay, and USDA soil texture classification

[0070] In the climate environmental variables provided in this embodiment, WorldClim Version 2 uses data from 60,000 climate stations worldwide (http: / / www.worldclim.com / version2), combined with elevation, distance from the coast, maximum / minimum surface temperature, and cloud cover data, to obtain annual average precipitation, annual average temperature, wind speed, vapor pressure, and solar radiation at a 1-kilometer resolution globally through spline interpolation.

[0071] The topographic factors provided in this embodiment obtain elevation data from the 30-meter spatial resolution ASTER GDEM V2 (Earth Spatial Data Cloud, https: / / www.gscloud.cn / ), and use SAGA (Automated Geoscientific Analysis System) to calculate topographic factors, including filling depressions and available water capacity.

[0072] The clay and USDA soil texture classification provided in this embodiment are extracted from the LUCAS Soil2009 database of the European Soil Data Center. Soil organic carbon in the topsoil (0-20 cm) of Europe in 2009 is extracted from the 1-km resolution product (https: / / esdac.jrc.ec.europa.eu / projects / lucas). Soil types are extracted from the French national soil type map at a scale of 1:5,000,000.

[0073] The French land use map provided in this embodiment is extracted from the CORINE Land Cover in 2018 ( https: / / land.copernicus.eu / ).

[0074] All environmental variables provided in this embodiment are resampled to a 30-meter spatial resolution using bilinear interpolation (soil type and land use use majority voting) and reprojected to the Lambert-93 system.

[0075] Step (3): The extraction of environmental variables at the soil sampling points provided in this embodiment is implemented through the extract function of the raster package in R language, so as to represent the latitude information and environmental variable data of the sampling points of each soil sample to be measured through a data matrix.

[0076] Step (4): First, in R language, the createDataPartition function of the caret package is used to divide the data set into a training set and a validation set. Then, the longitude and latitude information of each sample is extracted from the data set, and the Euclidean distance matrix between samples is calculated using the dist function. For each sample in the training set, the Euclidean distance between it and other training set samples is calculated in a loop, and the k nearest neighbor samples with the shortest distance are selected. Then, using the inverse distance weighting method, according to the organic carbon density values and their distance weights of these neighboring samples, the weighted average organic carbon density value is calculated and added to the training set as a new spatial neighborhood information variable. For each sample in the validation set, the Euclidean distance between it and the training set samples is also calculated, and k nearest training set samples are selected. Subsequently, the weighted average organic carbon density value is calculated using the inverse distance weighting method and added to the validation set as a new spatial neighborhood information variable. In this study, the value range of k is set from 1 to 10, and the optimal model parameters can be selected by comparing different k values. As Figure 2 shown, as k increases, the prediction accuracy (R 2 ) of the model gradually improves. When k = 7, the prediction accuracy of the model is the highest, and R 2 is 0.37.

[0077] Step (5): Build a random forest (RF) model. The RF model is built using the randomForest package. This model improves the prediction accuracy by integrating multiple decision trees. When building, the number of trees (ntree) is set to the default value of 500 to ensure the model stability and generalization ability. At the same time, the number of predictor variables (mtry) for splitting the decision tree nodes is optimized to 3 to make full use of the data feature information.

[0078] Step (6): Input the environmental variable set into the final prediction model to enhance the prediction accuracy of the digital mapping model.

[0079] In this embodiment, the data is divided into a training sample set (687) and a validation sample set (229), which are used to fit the prediction model and objectively evaluate the model accuracy respectively. By evaluating the model performance of the random forest model without using soil spatial neighborhood information and the random forest model combined with soil spatial neighborhood information on the same validation data, the potential of soil spatial neighborhood information in enhancing the spatial modeling accuracy is objectively evaluated. The research results show that in the modeling of soil organic carbon density in France, the random forest model combined with soil spatial neighborhood information is superior to the random forest model without using soil spatial neighborhood information. The R 2 increased by 0.05, the LCCC increased by 0.04, and the RMSE decreased by 0.9 t ha -1 , which proves the effectiveness and superiority of the method of this embodiment.

[0080] Embodiment 2

[0081] In this embodiment, the Meuse River floodplain near Stein in the Netherlands is selected as the research area. 152 topsoil samples (0 - 20 cm) collected nearby are used as soil samples. Using other soil property data and environmental variables obtained through remote sensing technology, soil spatial neighborhood information is constructed by the inverse distance weighting method, ultimately improving the prediction accuracy of the target soil property. The following mainly shows the specific data and implementation details:

[0082] Step (1): The research dataset of this study includes the data of soil heavy metals (zinc and lead) content, organic matter, altitude, etc. observed at 152 sampling points, which are specifically described as follows:

[0083] Soil zinc and lead concentration data. The Meuse dataset included in the sp software package is used in the research. This dataset provides the location and many soil and landscape variables of the observation sites, and the soil zinc and lead concentrations are from mixed samples with an area of approximately 15 m × 15 m.

[0084] The environmental variables provided in this embodiment mainly include distance to the river, altitude, land use, flood frequency, soil type, and lime application frequency.

[0085] In R language, the createDataPartition function of the caret software package is used to divide the dataset into a training set and a validation set. Then, the longitude and latitude information of each sample is extracted from the dataset, and the Euclidean distance matrix between samples is calculated using the dist function.

[0086] The research on soil heavy metal zinc provided in this embodiment:

[0087] For each sample in the training set in this embodiment, the Euclidean distance between it and other training set samples is calculated iteratively, and the nearest k neighboring samples are selected. Subsequently, the weighted average of the zinc contents of these k neighboring samples is calculated, and this value is added to the training set as a new spatial neighborhood information variable. For the samples in the validation set, the same operation is performed, i.e., the weighted average zinc content of the nearest k neighboring samples in the surrounding training set is calculated and added to the validation set as a new spatial neighborhood information variable. In this study, for the prediction of zinc, this embodiment attempts to use different k values (ranging from 1 to 10) to explore the influence of the k value on the prediction performance of the model, as Figure 3 shown. As k increases, the prediction accuracy (R 2 ) of the model gradually improves. When k = 9, the prediction accuracy of the model is the highest, and R 2 is 0.64.

[0088] Regarding the study on soil heavy metal lead provided by this embodiment:

[0089] Similar to the study process of zinc, this embodiment calculates the Euclidean distance between each sample in the training set and validation set and other samples, and selects the nearest k neighboring samples. Then, the weighted average of the lead contents of these k neighboring samples is calculated and added to the training set and validation set as a new spatial neighborhood information variable. Similarly, this embodiment attempts to use different k values (ranging from 1 to 10) to evaluate its influence on the performance of the lead content prediction model, as Figure 4 shown. As k increases, the prediction accuracy (R 2 ) of the model gradually improves. When k = 5, the prediction accuracy of the model is the highest, and R 2 is 0.73.

[0090] Through the above steps, this embodiment constructs training sets and validation sets containing spatial information for soil heavy metals zinc and lead respectively, providing a basis for the subsequent construction of the random forest model.

[0091] Step (3): Construct a random forest (RF) model

[0092] Using the randomForest software package, this embodiment constructs random forest models for zinc and lead respectively for prediction. This model improves the prediction accuracy by integrating multiple decision trees. During the model construction process, this embodiment sets the number of decision trees (ntree) to the default value of 500 to ensure the stability and generalization ability of the model. At the same time, in order to make full use of the data feature information, this embodiment optimizes the number of predictor variables (mtry) for splitting the decision tree nodes and sets it to an appropriate value.

[0093] Step (4): Enhance the performance of the prediction model

[0094] When enhancing the performance of the prediction model, this embodiment also operates separately for zinc and lead. This embodiment uses the dataset containing environmental variables and the calculated k-value weighted average soil heavy metal values (for zinc and lead respectively) as input features to train their respective prediction models. By introducing this spatial neighborhood information, the spatial prediction performance in the digital mapping model can be enhanced. These new features not only consider the spatial relationship between samples but also incorporate the influence of environmental variables, thus helping to improve the prediction accuracy and reliability of the model.

[0095] This embodiment divides the data into a training sample set (114) and a validation sample set (38), which are used to fit the prediction model and objectively evaluate the model accuracy respectively. By evaluating the model performance of the random forest model without using soil spatial neighborhood information and the random forest model incorporating soil spatial neighborhood information on the same validation data, the potential of soil spatial neighborhood information in enhancing the spatial modeling accuracy is objectively evaluated. The research results show that in the spatial modeling of soil heavy metals in the Meuse River floodplain near Stein in the Netherlands, the random forest model incorporating soil spatial neighborhood information is superior to the random forest model without using soil spatial neighborhood information, where for lead, the R 2 increased by 0.17, the LCCC increased by 0.12, and the RMSE decreased by 14.8 ppm; for zinc, the R 2 increased by 0.06, the LCCC increased by 0.03, and the RMSE decreased by 20.18 ppm, proving the effectiveness and superiority of the method of this embodiment.

[0096] The embodiments described above are only preferred solutions of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting equivalent replacement or equivalent transformation methods fall within the protection scope of the present invention.

Claims

1. A digital soil mapping method based on soil spatial neighborhood information, characterized in that: include: By comparing the Euclidean distance between the soil sample to be tested and other soil samples, a set of spatial neighborhood soil samples with the closest spatial distance to the soil sample to be tested is obtained, and the soil attribute data of each spatial neighborhood soil sample is inversely distance weighted and averaged to obtain the spatial neighborhood soil attribute data; A soil attribute data prediction model is obtained by training a random forest model based on an environmental variable set of each training sample in a training sample set and corresponding spatial neighborhood soil attribute data, wherein the training sample set comes from a soil sample set to be tested constructed by the soil sample to be tested; When applied, the environmental variable set of the area to be tested is input into the soil attribute data prediction model to obtain predicted soil attribute data, and digital soil mapping of the area to be tested is realized based on the predicted soil attribute data; By comparing the Euclidean distance between the soil sample to be tested and other soil samples, a set of spatial neighborhood soil samples with the closest spatial distance to the soil sample to be tested is obtained, including: Calculate the Euclidean distance between the soil sample in the tested area and other soil samples based on the latitude and longitude of the soil sample and other soil samples; The k other soil samples with the closest Euclidean distance are taken as the spatial neighborhood soil sample set; The soil attribute data includes one or more of soil organic carbon density data, soil heavy metal content data, soil texture data, and soil gravel content data.

2. The digital soil mapping method based on soil spatial neighborhood information according to claim 1 is characterized in that: The soil sample set to be tested also includes a verification sample set and target soil property data corresponding to each verification sample; The environmental variable set of the validation sample is input into the soil property data prediction model to obtain the predicted soil property data, and the prediction accuracy is obtained by using the determination coefficient, root mean square error and consistency correlation coefficient based on the predicted soil property data and the target soil property data; When the prediction accuracy is lower than the set accuracy threshold, the corresponding soil attribute data prediction model is used as the final soil attribute data prediction model.

3. The digital soil mapping method based on soil spatial neighborhood information according to claim 1 is characterized in that: The method for obtaining a soil sample to be tested comprises the following steps: Collect 0-20 cm soil surface samples from the test area and use the random stratification method to determine the spatial location of the soil sampling points; The soil sample to be tested is collected at the spatial position: a plurality of sub-samples are collected at each soil sampling point by using the diagonal method, and the plurality of sub-samples are mixed into a soil sample to be tested, and the latitude and longitude information of the soil sampling point is recorded at the same time.

4. The digital soil mapping method based on soil spatial neighborhood information according to claim 1 is characterized in that: Environmental variables include one or more of climate environmental variables, topographic environmental variables, spatial location environmental variables, and biological environmental variables.

5. The digital soil mapping method based on soil spatial neighborhood information according to claim 1 is characterized in that: The soil attribute data prediction model is obtained by training the random forest model based on the environmental variable set of each training sample in the training sample set and the corresponding spatial neighborhood soil attribute data, including: Based on the environmental variable set of each training sample in the training sample set and the corresponding spatial neighborhood soil attribute data, ten-fold cross validation was used to determine the optimal number of trees and branch variables of the random forest. The corresponding relationship between the environmental variables and the predicted soil attribute data was fitted based on the determined optimal number of trees and branch variables, thus obtaining the soil attribute data prediction model.

6. A digital soil mapping device based on soil spatial neighborhood information, characterized in that: include: Data processing module; Used to obtain a spatial neighborhood soil sample set with the closest spatial distance to the soil sample to be tested by comparing the Euclidean distance between the soil sample to be tested and other soil samples, and to obtain the spatial neighborhood soil property data by averaging the soil property data of each spatial neighborhood soil sample after inverse distance weighting; and to train a random forest model based on the environmental variable set of each training sample in the training sample set and the corresponding spatial neighborhood soil property data to obtain a soil property data prediction model, wherein the training sample set comes from the soil sample set to be tested constructed by the soil sample to be tested; By comparing the Euclidean distance between the soil sample to be tested and other soil samples, a set of spatial neighborhood soil samples with the closest spatial distance to the soil sample to be tested is obtained, including: Calculate the Euclidean distance between the soil sample in the tested area and other soil samples based on the latitude and longitude of the soil sample and other soil samples; The k other soil samples with the closest Euclidean distance are taken as the spatial neighborhood soil sample set; The soil attribute data includes one or more of soil organic carbon density data, soil heavy metal content data, soil texture data, and soil gravel content data; Input module: Input the environmental variable set of the area to be tested into the soil property data prediction model to obtain predicted soil property data, and realize digital soil mapping of the area to be tested based on the predicted soil property data.

7. A digital soil mapping device based on soil spatial neighborhood information, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable codes, and when the one or more processors execute the executable codes, they are used to implement the digital soil mapping method based on soil spatial neighborhood information as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Terrain collaborative variable selection method for digital soil cartography

    CN103955953A

  • Digital soil mapping method and system based on multi-source data

    CN116206011A

  • Southern hilly area cultivated land soil available phosphorus drawing method based on high-resolution environment variables

    CN115266612A

  • Soil organic matter spectrum prediction method and device based on nonlinear memory learning

    CN117057464A