Three-dimensional soil attribute distribution prediction method based on adaptive soil depth function
By optimizing the soil property distribution model through adaptive parameter correction and inverse distance weighting algorithm, the problem of insufficient adaptability of the soil property depth distribution function in different geographical locations is solved, and higher accuracy of soil property prediction is achieved.
Patent Information
- Application Number
- CN202511068367.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-14
AI Technical Summary
The existing soil property depth distribution function is not adaptable to different geographical locations, exhibiting spatial heterogeneity and data inhomogeneity, resulting in low prediction accuracy.
Based on the adaptive parameter correction method, the optimal prediction model is selected by fitting soil property data with multiple functions, and the model parameters are iteratively optimized using the inverse distance weighted algorithm to construct a three-dimensional soil property distribution prediction model.
It improves the accuracy and reliability of soil property distribution prediction, and can more accurately simulate the distribution patterns of soil properties in different soil-forming environments and depth ranges, reduce the impact of missing data, and improve the adaptability and accuracy of the model.
Smart Images

Figure CN120951168A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of soil environmental monitoring technology, and more specifically, relates to a three-dimensional soil property distribution prediction method based on an adaptive soil depth function. Background Technology
[0002] The study of the depth distribution of soil properties aims to reveal the variation patterns and driving mechanisms of soil physicochemical and biological characteristics (such as organic matter, moisture, nutrients, and microorganisms) with soil depth. By constructing mathematical functions (such as exponential decay and piecewise models) or physical models (such as convection-diffusion equations), the vertical heterogeneity of soil properties can be quantified, and the processes of material migration, energy transformation, and biogeochemical cycles can be analyzed.
[0003] This research is of crucial significance for optimizing agricultural productivity (such as nutrient management in the root zone), ecological restoration (such as vertical control of pollutants), climate change response (such as assessment of vertical carbon pool distribution), and groundwater protection (such as prediction of solute leaching). Its findings provide a scientific basis for precision agriculture, environmental risk assessment, and sustainable development strategies, and represent an important area of interdisciplinary research in soil science, ecology, and earth system science.
[0004] Currently, research on the depth distribution of soil properties typically relies on depth distribution functions constructed from sampled data. However, due to variations in soil-forming factors and processes across different geographical locations, soil properties exhibit strong spatial heterogeneity and uncertainty at different depths. Furthermore, much of the sampled data suffers from uneven distribution, abrupt changes, and missing data. Therefore, existing depth distribution functions based on sampled data lack adaptability. To address this, this application attempts to construct an adaptive soil depth distribution function suitable for different soil-forming environments, thereby better fitting the distribution patterns and characteristics of soil properties at different depths. Summary of the Invention
[0005] In view of the above-mentioned defects or improvement needs of the existing technology, this application provides a soil property depth distribution prediction method based on adaptive parameter correction method, which aims to solve the technical problem of low prediction accuracy of current soil depth property models.
[0006] To achieve the above objectives, in a first aspect, this application provides a method for predicting the depth distribution of soil properties based on an adaptive parameter correction method, comprising: S1. Obtain the attribute depth distribution data of multiple sampling points in the target soil area; S2. Taking the attribute depth distribution data of each sampling point as the object, select multiple functions to fit the attribute depth distribution data of each sampling point. Multiple candidate prediction models are obtained for each sampling point. Based on the fitting accuracy, the optimal candidate prediction model is selected as the basic prediction model. S3. Set a standard depth range for the soil, use the basic prediction model to predict the attribute data of the first standard depth range to obtain the predicted value, and optimize the parameters of the basic prediction model by minimizing the difference between the attribute data of the first depth range of the sampling point and the predicted value. S4. The attribute data of each standard depth interval is predicted layer by layer from shallow to deep using the basic prediction model. The obtained normal prediction values are then used to correct the erroneous prediction values by reverse distance weighting. The parameters of the basic prediction model are then optimized using the corrected prediction values. This process is iterated until the prediction values of each standard depth interval are normal. S5. Collect multi-source environmental variables that affect soil property data, combine the predicted values of each sampling point in each standard depth range as a dataset, and construct a three-dimensional soil property distribution prediction model based on machine learning algorithms.
[0007] Preferably, step S1 specifically comprises: Acquire the attribute depth distribution data of multiple soil sampling points in the target area; Remove sampling points with a depth range of less than or equal to 2 layers; In each sampling point, the attribute data of the depth interval within the preset topsoil layer are averaged, and then the attribute data of the first depth interval of the sampling point is updated to the average value. The attribute data of each depth interval of the sampling points are smoothed by using a k-th order polynomial.
[0008] Preferably, in step S2, the optimal candidate prediction model is selected as the basic prediction model based on the fitting accuracy, specifically as follows: Each candidate prediction model predicts attribute data for a preset surface interval; The fitting accuracy of the candidate prediction model is comprehensively evaluated by the difference between the obtained predicted value and the attribute data of the first depth interval of the sampling point, the determination coefficient of the candidate prediction model, and the root mean square error of the candidate prediction model. The candidate prediction model with the best fitting accuracy is selected as the base prediction model.
[0009] Preferably, in step S2, the optimal candidate prediction model is selected as the basic prediction model based on the fitting accuracy, specifically as follows: Each candidate prediction model predicts attribute data for a preset surface interval. The difference between the predicted value and the attribute data for the first depth interval of the sampling point is taken as the surface prediction error. The candidate prediction model is scored based on the surface prediction error, with a higher score for a smaller error. This results in surface prediction error scores for multiple candidate prediction models. ; Calculate the coefficient of determination for each candidate prediction model, and assign a score to each candidate model based on the coefficient of determination. The higher the coefficient of determination, the higher the score. This process yields the coefficient of determination scores for multiple candidate prediction models. ; Calculate the root mean square error (RMSE) of each candidate prediction model, and assign a score to each candidate model based on the RMSE. The smaller the RMSE, the higher the score. This yields the RMSE scores for multiple candidate prediction models. ; Calculate the total score for each candidate prediction model:
[0010] in, For the first The total score of each candidate prediction model , and The weights for the coefficient of determination score, root mean square error score, and surface prediction error score are respectively, satisfying:
[0011]
[0012] The total score The highest candidate prediction model As a basic prediction model.
[0013] Preferably, in step S3, the parameters of the basic prediction model are optimized by minimizing the difference between the attribute data and the predicted value of the first depth interval of the sampling point, specifically as follows: The parameters of the basic prediction model can be obtained using the following formula. :
[0014] in, Indicates minimization. This represents the attribute data of the first depth interval of the sampling point. Indicates the first standard depth range. This represents the predicted value from the basic prediction model.
[0015] Preferably, step S4 specifically includes the following sub-steps: S41. Use the basic prediction model to predict the attribute data of the Kth standard depth interval. If the obtained prediction value is found to be incorrect, proceed to the next step; otherwise, proceed to step S44. S42. The predicted values of the Kth standard depth interval are corrected by using the reverse distance weighting of the predicted values of each standard depth interval before the Kth standard depth interval. S43. Correct the parameters of the basic prediction model using the predicted values of the first K standard depth intervals; S44. Determine whether the Kth standard depth interval is the last standard depth interval. If yes, the iterative optimization ends; otherwise, K=K+1 and return to step S41.
[0016] Preferably, the criteria for determining incorrect predictions are as follows: If the predicted value of the Kth standard depth interval is obtained and the dispersion of the predicted values of the first K standard depth intervals is greater than the preset dispersion, then the predicted value of the Kth standard depth interval is considered to be incorrect. If the predicted value is negative or zero, the predicted value is considered incorrect.
[0017] Preferably, the dispersion of the predicted values for the first K standard depth intervals is obtained by evaluating the standard deviation or variance of the predicted values for the first K standard depth intervals.
[0018] Preferably, the predicted values for the Kth standard depth interval are corrected using a reverse distance weighted method, specifically as follows:
[0019]
[0020] in, Indicates the first Predicted values for the standard depth range, For the first Weights of standard depth intervals Represents the Kth standard depth interval and the... Median distance of the standard depth range This represents the predicted value for the corrected Kth standard depth interval.
[0021] In a second aspect, this application provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation thereof.
[0022] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art: (1) Given the complexity and variability of soil-forming factors and processes, in order to more accurately characterize the vertical distribution patterns of soil properties, gain a deeper understanding of soil formation and evolution mechanisms, and improve the adaptability and prediction accuracy of models to different soil conditions, this application determines multiple candidate fitting functions based on the property depth distribution data of existing sampling points, and then uses a preset evaluation method to initially adaptively select a suitable basic prediction model from the multiple fitting functions. Subsequently, a standard depth range of soil is set to predict the property data, and an inverse distance weighted algorithm is used to iteratively correct the predicted values to obtain the optimal property depth distribution data. Finally, a three-dimensional soil property distribution prediction model is constructed based on the property depth distribution data. This method effectively improves the accuracy and reliability of soil property distribution prediction.
[0023] (2) In order to solve the problem of missing deep sampling data, this application determines the basic prediction model based on the range of existing sampling data through various candidate fitting functions, and then uses the deep prediction data predicted by the basic prediction model. At the same time, it monitors whether the deep prediction data is reasonable, and uses the reverse distance weighted interpolation method to replace unreasonable deep prediction data. Based on the replacement data, the basic prediction model is corrected in reverse, and finally the basic prediction model can obtain the best-fit sampling value and the prediction value that conforms to the attribute change trend.
[0024] (3) This application introduces a variety of novel functions to fit the depth distribution of soil properties, aiming to select the optimal prediction model to more accurately simulate and predict the distribution patterns and characteristics of soil properties in different soil-forming environments and at various depth intervals. These functions can effectively adaptively capture the variations and characteristics of soil properties in different soil-forming environments and at various depth levels.
[0025] (4) In the sampling data, the surface data is easily affected by human factors. In order to reduce this influence, this application performs preliminary preprocessing on the attribute data of the topsoil layer at the sampling point, thereby ensuring that the sampling data can accurately reflect the influence of environmental factors on the depth distribution of attributes.
[0026] (5) The applicant found that the surface attribute data of soil determines the range of variation of deep attribute data. Therefore, when screening the basic prediction model from multiple candidate functions, this application regards the surface prediction error as the most important evaluation item of fitting accuracy. The basic prediction model obtained by this correction is more in line with the actual distribution of soil attributes. Attached Figure Description
[0027] Figure 1 This is a flowchart of the three-dimensional soil property distribution prediction method provided in the embodiments of this application.
[0028] Figure 2 This is a flowchart of the final prediction model construction method provided in the embodiments of this application.
[0029] Figure 3 This is a schematic diagram showing the distribution of attribute data obtained from the modified basic prediction model provided in this application as a function of depth.
[0030] Figure 4 This is a schematic diagram showing the distribution of attribute data obtained from the final prediction model provided in this application as a function of depth.
[0031] Figure 5 These are the various soil property depth distribution prediction models provided in the embodiments of this application. Mean comparison chart.
[0032] Figure 6 This is a comparison chart of the RMSE mean values of various soil property depth distribution prediction models provided in the embodiments of this application.
[0033] Figure 7 This is a distribution map showing the proportion of various soil property depth distribution prediction models under different soil orders provided in the embodiments of this application.
[0034] Figure 8 This is a distribution map showing the proportion of various soil attribute depth distribution prediction models under different land uses provided in the embodiments of this application.
[0035] Figure 9 This is a heat map showing the average values of the k-parameter under different land uses, provided in the embodiments of this application.
[0036] Figure 10 This is a ranking chart of the importance of various environmental variables to attribute data at a resolution of 30m, provided in an embodiment of this application.
[0037] Figure 11 This is a schematic diagram showing the frequency of soil organic carbon values at various depths provided in the embodiments of this application.
[0038] Figure 12 This is a schematic diagram illustrating the prediction accuracy of the three-dimensional soil property distribution prediction model at various depths in various standard depth ranges provided in the embodiments of this application.
[0039] Figure 13 This is a schematic diagram showing the distribution of soil organic carbon content in various standard depth ranges within the target area provided in this application embodiment.
[0040] Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0042] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. For example, "first model" and "second model," etc., are used to distinguish different models, not to describe a specific order of models.
[0043] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0044] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple sampling points means two or more sampling points, multiple function models means two or more function models, etc.
[0045] First, let's introduce the technical terms or English abbreviations that appear in this application: SOC (Soil Organic Carbon) refers to soil organic carbon.
[0046] (R-Squared) coefficient of determination.
[0047] RMSE (Root Mean Squared Error).
[0048] The embodiments of this application are described below with reference to the accompanying drawings. Figure 1 The diagram shows a flowchart of a three-dimensional soil property distribution prediction method, which includes the following steps: 1. Determine the source of soil data In this embodiment, the depth distribution of soil organic carbon content in the target area is taken as the research object.
[0049] The target region exhibits diverse soil types, strong spatial and temporal heterogeneity, multiple layers of soil barriers, and varied degradation processes, significantly constraining agricultural production. Current land conservation and utilization models lack technical specificity; furthermore, related technologies primarily focus on topsoil fertility, neglecting the characteristics of subsurface soils. Overall, the lack of high-precision spatial and temporal variability information on soil properties leads to mismatches between conservation and utilization technologies and soil characteristics, resulting in insufficient accuracy in their application. Therefore, implementing high-precision soil attribute mapping provides a crucial data foundation for supporting the efficient and sustainable utilization of soil resources.
[0050] Using existing sampling data in the target area, the organic carbon content of several depth ranges was extracted from each sampling point, and these data were collected for subsequent steps.
[0051] 2. Data Processing 2.1 Data Removal Sampling points with a depth range of less than or equal to 2 layers were removed.
[0052] 2.2 Topsoil Data Preprocessing Considering that the topsoil layer in different regions is subject to significant human disturbance, and that topsoil sampling data may not reflect the patterns and characteristics of soil property changes with depth, preprocessing is performed on the property data of the topsoil layer at the sampling points. This preprocessing specifically involves: The attribute data of the depth range within the topsoil layer of the sampling point are averaged and used as the attribute data of the first depth range of the sampling point, thereby reducing the human influence on the attribute data of the first depth range.
[0053] Let's illustrate this with an example: In this embodiment, the 0-40cm layer is defined as the topsoil layer. The original sampling data for some sampling points are shown in Table 1. Table 1
[0054] First, data removal is performed. Sampling points numbered 03-004 in Table 1 only have attribute data for one depth interval, so sampling points 03-004 are removed.
[0055] In Table 1, the first depth interval (0-13) and the second depth interval (13-40) of sampling point 03-001 are located within the topsoil layer (0-40cm). Therefore, the attribute data of the first depth interval and the second depth interval are averaged:
[0056] Therefore The attribute data for the first depth interval of sampling points 03-001 has been updated, see Table 2: Table 2
[0057] 2.3 Smoothing of Attribute Data A k-th order polynomial is used to fit the attribute data for all depth intervals at each sampling point, and the attribute data is expanded based on the fitted k-th order polynomial. For example, sampling point 03-001 has attribute data in three depth spaces, which are now expanded using a... order polynomial:
[0058] Fit these three attribute data points, and then take two data points before and after each attribute data point from the curve of the fitted k-th order polynomial. This expands one attribute data point into five attribute data points, and three attribute data points into fifteen attribute data points. The attribute data points 03-001 are smoothed and then expanded into fifteen attribute data points.
[0059] 3. Construct a basic prediction model 3.1 Preliminary Fitting of Candidate Prediction Model Because the soil-forming environment and depth range vary at each sampling point, in order to better fit the pattern and characteristics of soil property changes with depth, multiple candidate prediction models are fitted to the attribute data of each sampling point so that the model with the highest prediction accuracy can be selected subsequently. In this embodiment, the candidate prediction models include: (1) Power Function (First Model)
[0060] (2) Logarithmic Function (Second Model)
[0061] (3) Negative Exponential Function (Third Model)
[0062] (4) Trinomial Exponential Function (Fourth Model)
[0063] (5) Fully Inverted Polynomial Function (Fifth Model)
[0064] (6) Modified Exponential Function (Sixth Model)
[0065] (7) Exponential Decay Functions (Seventh Model)
[0066] (8) Gaussian Function (Eighth Model)
[0067] (9) Spherical Function (Ninth Model)
[0068] in, This indicates the degree of variation in soil properties at vertical depth intervals of ℎ; The nugget value primarily reflects shallow microscale variations or measurement errors. Represents structural variance, representing the strength of the explainable variation caused by deep structure; It is the sill value, which is the total variation value of the variogram that eventually stabilizes as the depth increases; and To adjust the constant of the function shape, and control the rate of variation with depth and the smoothness of the curve; The rate of change coefficient reflects the intensity of variation of soil properties with increasing depth in the profile; Residuals represent the portion of variation that the model fails to explain, including the effects of measurement errors, sampling noise, or other non-structural disturbances on variations in soil profile properties.
[0069] 3.2 Parameter Optimization The best-fit parameters for each candidate prediction model at each sampling point are found using the least squares method, as follows: The parameter is calculated using the following formula. Sum of squared errors of the next candidate prediction model :
[0070] in, For the first A depth range, For the first Actual attribute data for each depth range For candidate prediction models in parameters Next, the Predicted values of attribute data for each depth interval; Adjust parameters iteratively gradually decrease until reached The minimum value or the maximum number of iterations is reached; The model parameters at this time As the best-fit parameters for the model.
[0071] 3.3 Evaluation and Screening of Candidate Prediction Models After fitting the candidate prediction models, nine candidate prediction models are obtained for each sampling point.
[0072] Calculate the coefficient of determination for each candidate prediction model ,according to Assign a score to each candidate prediction model. The larger the value, the higher the score, thus yielding the coefficient of determination scores for the nine candidate prediction models. .
[0073] Calculate the root mean square error (RMSE) of each candidate prediction model, and assign a score to each candidate prediction model based on the RMSE. The smaller the RMSE, the higher the score. Thus, the root mean square error scores RMSE(1)-RMSE(9) of the 9 candidate prediction models are obtained.
[0074] Each candidate prediction model predicts the attribute data of a preset surface interval (the preset surface depth is 0-5cm in this embodiment). The difference between the obtained attribute data and the attribute data of the first depth interval is used as the surface prediction error. The candidate prediction model is scored according to the surface prediction error. The smaller the surface prediction error, the higher the score. Thus, the surface prediction error scores F(1)-F(9) of the 9 candidate prediction models are obtained.
[0075] Calculate the total score for each candidate prediction model:
[0076] in, For the first The total score of each candidate prediction model , and Let the importance weights of the coefficient of determination score, root mean square error score, and surface prediction error score be respectively, satisfying the following:
[0077]
[0078] , and The first The coefficient of determination score, root mean square error score, and surface prediction error score of each candidate prediction model.
[0079] The candidate prediction model with the highest total score is the best-performing candidate prediction model. During the selection process, the surface prediction error has the highest weight, meaning it is the most important. Therefore, the candidate prediction model with the highest total score is selected as the basic prediction model. Table 3 shows the basic prediction models and their parameters for the three sampling points in this embodiment: Table 3
[0080] 4. Optimize the basic prediction model to obtain the final prediction model. After evaluating and selecting candidate prediction models for each sampling point, a basic prediction model is obtained. However, due to the low quality of sampling data at some points—for example, the first sampling point 03-001 lacks attribute data for the deep depth range, and the third sampling point 03-003 lacks attribute data for both the shallow and deep depth ranges—the basic prediction models for the first and third sampling points suffer from insufficient accuracy in predicting soil properties across different soil-forming environments and depth ranges. Therefore, to more accurately fit the attribute patterns and characteristics of each depth range, adaptive parameter correction of the basic prediction model is necessary. For details, please refer to [link to documentation]. Figure 2 This includes the following steps: 4.1. Correcting the basic prediction model based on surface prediction data First, the standard depth range is defined. See Table 4 for the standard depth range defined in this embodiment: Table 4
[0081] The attribute data for each standard depth interval is predicted using the basic prediction model and is shown in Table 5. Table 5
[0082] The first depth interval closest to the ground surface is identified from the sampling data of the sampling points, and the basic prediction model is corrected by using the difference in attribute data between the first depth interval and the first standard depth interval.
[0083] Referring to Table 2, the first depth interval closest to the surface for sampling points 03-001 is 0-13, corresponding to an attribute value of 17.61. Therefore, by minimizing... This will allow for further adjustments to the parameters of the corresponding basic prediction model.
[0084] 4.2 Further iterative optimization of the basic prediction model based on the prediction data of each layer. The modified base prediction model is used to predict attribute data for each standard depth interval. The distribution of the predicted attribute data is as follows: Figure 3 As shown, some data in the sixth standard depth range are negative. This is because the sampling data from sampling points 03-001 lacks attribute data for depths greater than 70cm. Therefore, the basic prediction model built upon this attribute data may have significant errors in predicting deep soil layers. Thus, this basic prediction model still has problems and needs further optimization, specifically including the following steps: 4.2.1. Use the modified basic prediction model to predict the attribute data of each standard depth interval layer by layer. If the attribute data is found to be incorrect, output the standard depth space K where the incorrect attribute data is located and proceed to the next step; otherwise, the iterative optimization ends. If the attribute data obtained from the current prediction is negative or zero, then the attribute data is incorrect. If the difference between the currently predicted attribute data and the mean of all attribute data is greater than the preset difference, then the attribute data is incorrect. 4.2.2. The inverse distance weighted algorithm is used to correct the attribute data of the Kth standard depth interval using the attribute data of each standard depth interval above the Kth standard depth interval. :
[0085]
[0086] in, Indicates the predicted result of the first Attribute data for standard depth ranges, For the first Weights of standard depth intervals Represents the Kth standard depth interval and the... The median distance of the standard depth range.
[0087] For example, if the attribute data for the fourth standard depth interval is negative, then the attribute data for the fourth standard depth interval can be derived using the attribute data predicted from the first, second, and third standard depth intervals.
[0088] 4.2.3. Use the attribute data of the first K standard depth intervals to correct the parameters of the basic prediction model.
[0089] 4.2.4 Determine whether the Kth standard depth interval is the last standard depth interval. If yes, the iterative optimization ends; otherwise, return to step 4.2.1.
[0090] After iterative optimization of the basic prediction model, the final prediction model for some sampling points is obtained, as shown in Table 6: Table 6
[0091] The attribute data for each standard depth range of each sampling point were predicted using the final prediction model of some sampling points. The results are shown in Table 7. Table 7
[0092] The data was obtained by plotting the attribute data of sampling points 03-001 in each standard depth range. Figure 4The distribution of attribute data with depth shows that, compared to Figure 3 The trend of organic carbon values at sampling points 03-001 with depth is more accurate.
[0093] In this embodiment, the final prediction model is obtained through adaptive matching based on the sampling data from different sampling points.
[0094] Next, we compare the fitting accuracy of the adaptive final prediction model with other prediction models. See the results below. Figure 5 and Figure 6 In the x-axis, PF represents the power function model, LF the logarithmic function model, NEF the negative exponential function model, TEF the three-type exponential function model, FIPF the univariate polynomial inverse function model, EDF the exponential decay function model, GF the Gaussian function model, SF the spherical function model, and AF the adaptive final prediction model.
[0095] The results show that the adaptive final prediction model of this application performs best among all prediction models, with an average fit mean. It achieves the highest value and the lowest RMSE. This model form, through its flexible expressive power, better captures the nonlinear variation trend of SOC along the soil profile, possesses broad adaptability and promotion potential, and is suitable as the preferred method for vertical modeling of soil properties.
[0096] Soil organic carbon (SOC), as a core element reflecting soil quality and ecological function, plays a crucial role in elucidating soil carbon cycle processes and fertility maintenance mechanisms through its functional model fitting characteristics across different soil orders. This application, based on the distribution frequency data of various functional model types across soil orders such as semi-hydrogenic soils, leached soils, and calcareous soils, explores the differences in soil order distribution patterns among SOC prediction model types through statistical analysis and visualization. Figure 7As shown, the distribution of SOC prediction models varies significantly across different soil orders. Gaussian function models dominate in most soil orders, including semi-hydrated soils (44 times), leached soils (26 times), and calcareous soils (20 times), reflecting their universality in fitting SOC variation patterns. Meanwhile, each soil order exhibits unique distributions due to different soil-forming conditions. For example, exponential decay function models (17 times) and modified exponential function models (14 times) are prevalent in semi-hydrated soils, while logarithmic function models (13 times) are somewhat present in leached soils. In calcareous soils, spherical function models (7 times) and exponential decay function models (7 times) are relatively evenly distributed. Although Gaussian function models still dominate in semi-leached soils and primary turquoise soils, their frequency is lower, and other function model types are more dispersed. Influenced by special conditions such as human intervention and saline-alkali stress, the patterns are more complex. In summary, the distribution of SOC prediction model types is deeply correlated with soil order characteristics. The differences in the frequency distribution of SOC prediction model types under different soil orders reflect the spatial heterogeneity of profile carbon migration and stabilization mechanisms, and also provide a basis for soil classification for selecting appropriate prediction models.
[0097] The distribution of soil organic carbon (SOC) adaptive function types varies significantly across different land use types. Different land use patterns alter soil physicochemical properties and vegetation-microbe interactions, thereby affecting the SOC function fitting patterns. Figure 8 As shown, this application analyzes the soil science mechanisms underlying the differences in the distribution of four dominant function models (logarithmic, exponential decay, modified exponential, and Gaussian) for three typical land use types: forest, farmland, and grassland. In farmland, human cultivation reshapes SOC dynamics, and the Gaussian function model reflects the fluctuation characteristics of SOC content after homogenization. The modified exponential and exponential decay models reflect the human intervention imprint on organic matter turnover. In forest, natural vegetation ensures that SOC accumulation and distribution conform to natural laws, and the Gaussian function model reflects the distribution trend under conditions of minimal disturbance. The logarithmic and exponential decay models are related to the natural succession characteristics of litter decomposition. In grassland, herbaceous vegetation and environmental stress influence SOC distribution. The Gaussian function model stems from the relatively concentrated spatial distribution of SOC, while the exponential decay model and others are related to the stress response mechanism of herbaceous carbon turnover. Different land uses shape the distribution pattern of SOC function models by altering soil structure, organic matter input, and microbial activity, providing theoretical support for accurately simulating carbon processes and optimizing land use strategies, thus helping to enhance the value of soil carbon sinks and ecosystem services.
[0098] Building upon this, and further based on the correlation between soil organic carbon (SOC) function model distribution and land use, a normalized thermogram of the rate of change parameter k was constructed. (See [link to relevant documentation]). Figure 9 This study aims to reveal the k-value differentiation characteristics and their soil scientific significance for different land use types (forest, farmland, grassland) under the main function models (logarithmic function model, exponential decay function model, modified exponential function model, and Gaussian function model), as detailed below: (1) From the perspective of the function model response within the land use type, in the forest system, the k value of the exponential decay function model (0.36) is close to and higher than the k value of the logarithmic function model (0.25) and the k value of the Gaussian function model (0.20). This characteristic aligns with the phased nature of forest soil carbon turnover: after forest litter input, the initial organic matter such as lignin and cellulose is rapidly decomposed by microorganisms (corresponding to the exponential decay function, reflecting the rapid loss of carbon pool); as humification progresses, the accumulation of difficult-to-decompose components slows down the carbon decay rate, and the modified exponential function, which can characterize the transition from "rapid decay to slow stabilization" (fitting the transition of carbon pool from dynamic loss to steady-state accumulation), becomes dominant in this stage; the low k value of the logarithmic function reflects the long-term slow turnover of forest soil carbon (such as the inert accumulation of deep humus); the low k value of the Gaussian function corresponds to the normal distribution pattern of forest soil carbon under the bioclimatic gradient (such as the spatial heterogeneity of litter mass and microbial activity leading to a concentrated distribution of carbon density), presenting an overall natural carbon cycle sequence of "dynamic decomposition - steady-state accumulation - slow turnover - spatial homogenization".
[0099] (2) In the farmland system, the k value of the exponential decay function model (0.32) is the highest, the k value of the modified exponential function model (0.12) is the lowest, and the k value of the logarithmic function model (0.15) and the k value of the Gaussian function model (0.17) are in between. This is directly related to the reshaping of carbon processes in farmland by human interference: tillage disturbances (such as plowing) destroy soil aggregates, exposing organic matter and accelerating decomposition (significant initial exponential decay); however, artificial carbon supplementation such as fertilizer input and crop residue return makes it difficult for the soil carbon pool to form the "accumulation-stability" pattern of the natural ecosystem (the low k value of the modified exponential function reflects that the carbon pool accumulation rhythm is interrupted by tillage and harvesting); the low k value of the logarithmic function reflects the slow linear response of farmland carbon turnover after regulation by irrigation, fertilization, etc. (such as the gradual accumulation of soil carbon under long-term fertilization); the k value of the Gaussian function corresponds to the more uniform distribution of soil carbon density under the homogenization effect of tillage (such as the reduction of spatial variation of carbon content at the field scale), showing the characteristics of artificial carbon cycle of "accelerated decomposition by human intervention - accumulation by supplementation disturbance - homogenized distribution".
[0100] (3) In the grassland system, the k value of the exponential decay function model (0.44) is significantly higher than that of other functions, while the k value of the logarithmic function model (0.27), the k value of the Gaussian function model (0.24), and the k value of the modified exponential function model (0.19) decrease in that order. This is due to the stress-induced carbon turnover of the grassland ecosystem: herbaceous roots are shallow and turnover is fast (such as the seasonal death of annual herbs), and grazing trampling accelerates the breaking and decomposition of litter, so that the soil carbon pool is in a dynamic state of "input-rapid loss" for a long time (dominated by exponential decay); the k value of the logarithmic function reflects the stage response of grassland carbon input (such as the slow decomposition after the surge in biomass during the rainy season); the k value of the Gaussian function corresponds to the concentration trend of carbon distribution under grassland stress (such as the peak carbon density in the root distribution area of dominant species); the low k value of the modified exponential function is due to the difficulty of achieving steady-state accumulation of carbon pool caused by drought and grazing (continuous disturbance inhibits the humification process), showing the natural-disturbance coupled carbon cycle characteristics of "stress-induced rapid decomposition-stage slow response-concentrated distribution-accumulation inhibition".
[0101] In summary, the differentiation of the k-parameter essentially reflects how land use alters carbon input patterns (such as forest litter, grassland roots, and farmland stubble), decomposition intensity (microbial activity, disturbance frequency), and spatial distribution patterns (aggregate structure, biogradient), thereby reshaping the rate and pattern of soil carbon turnover. This provides crucial evidence for elucidating the soil carbon cycle mechanism driven by land use and optimizing carbon model parameters (such as function responses that distinguish between natural and anthropogenic disturbances), contributing to a deeper understanding of soil carbon sink potential assessment and regulation from a process model perspective.
[0102] Based on the results obtained from the above prediction models, we can significantly enhance our ability to predict dynamic changes in soil quality. This means that in future land use planning and management, we can formulate more scientific strategies to effectively maintain soil health and productivity. Healthy soil is not only the foundation of agricultural production, but also plays a crucial role in carbon storage, soil and water conservation, and biodiversity maintenance.
[0103] 5. Collect multi-source environmental variables that affect the spatial heterogeneity and uncertainty of soil properties, and combine them with the attribute data of each sampling point in each standard depth range predicted by the final prediction model (see Table 7) as the dataset. Based on the machine learning algorithm, construct a soil property distribution prediction model at different depths. Thus, based on the known multi-source environmental variables in various parts of the target area, the depth distribution prediction of local soil properties can be realized, forming a three-dimensional soil property distribution map.
[0104] In this embodiment, a 30-meter resolution soil attribute environmental variable dataset was constructed by analyzing soil survey data. This dataset, combined with attribute data from each sampling point at various standard depth intervals, served as the training set. Machine learning algorithms (including Random Forest, LightGBM, XGBoost, and Bagging models) were introduced to build soil attribute prediction models to predict the content of soil attributes at different depths. After validation and optimization, the optimal prediction results were selected, and a ranking of the importance of environmental variables on attribute data was output. See [link to relevant documentation]. Figure 10 The darker the color, the more important the influence of environmental variables on the attribute data.
[0105] Figure 11 The figure shows the frequency of various predicted soil organic carbon (SOC) content values across different standard depth ranges in the entire target area. As can be seen from the figure, the SOC content in the soil gradually decreases with increasing depth throughout the target area.
[0106] The method described in this application successfully predicted the soil organic carbon (SOC) content in the target area across six standard depth intervals (0–5cm, 5–15cm, 15–30cm, 30–60cm, 60–100cm, and 100–200cm). Modeling results show that the adaptive optimization-based variable selection strategy significantly improved the model's stability and accuracy, with random forest performing best at each depth level. Values range from 0.22 to 0.46, such as... Figure 12 As shown. A high-resolution three-dimensional distribution map of soil properties was ultimately constructed, from... Figure 13 It is evident that the surface SOC content in the target area exhibits a significant spatial distribution characteristic. The SOC content is relatively concentrated in the central and northeastern parts of the target area, with the highest content in the northern part and relatively low content in the northwestern part. As the soil layer deepens, the SOC content decreases layer by layer.
[0107] Therefore, by constructing a soil property distribution prediction model and combining it with scientific soil management strategies, technical support can be provided for the efficient utilization of soil resources in this major grain-producing area of the country. This not only helps ensure national food security but also promotes the sustainable development of the regional ecological environment, achieving a synergistic improvement in economic, social, and ecological benefits. Furthermore, the feasibility of this method in maintaining stable modeling capabilities and wide application in highly heterogeneous regional environments has been verified, demonstrating promising prospects for technological promotion.
[0108] Based on the methods in the above embodiments, this application provides an electronic device, such as... Figure 14As shown, it includes a processor, a communications interface, a memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions stored in the memory to execute the methods described in the above embodiments.
[0109] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0110] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0111] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0112] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.
[0113] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.
[0114] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0115] It is understood that the various numerical designations used in the embodiments of this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application.
[0116] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for predicting the distribution of three-dimensional soil properties based on an adaptive soil depth function, characterized in that, include: S1. Obtain the attribute depth distribution data of multiple sampling points in the target soil area; S2. Taking the attribute depth distribution data of each sampling point as the object, select multiple functions to fit the attribute depth distribution data of each sampling point. Multiple candidate prediction models are obtained for each sampling point. Based on the fitting accuracy, the optimal candidate prediction model is selected as the basic prediction model. S3. Set a standard depth range for the soil, use the basic prediction model to predict the attribute data of the first standard depth range to obtain the predicted value, and optimize the parameters of the basic prediction model by minimizing the difference between the attribute data of the first depth range of the sampling point and the predicted value. S4. The attribute data of each standard depth interval is predicted layer by layer from shallow to deep using the basic prediction model. The obtained normal prediction values are then used to correct the erroneous prediction values by reverse distance weighting. The parameters of the basic prediction model are then optimized using the corrected prediction values. This process is iterated until the prediction values of each standard depth interval are normal. S5. Collect multi-source environmental variables that affect soil property data, combine the predicted values of each sampling point in each standard depth range as a dataset, and construct a three-dimensional soil property distribution prediction model based on machine learning algorithms.
2. The three-dimensional soil property distribution prediction method according to claim 1, characterized in that, Step S1 specifically involves: Acquire the attribute depth distribution data of multiple soil sampling points in the target area; Remove sampling points with a depth range of less than or equal to 2 layers; In each sampling point, the attribute data of the depth interval within the preset topsoil layer are averaged, and then the attribute data of the first depth interval of the sampling point is updated to the average value. The attribute data of each depth interval of the sampling points are smoothed by using a k-th order polynomial.
3. The three-dimensional soil property distribution prediction method according to claim 1, characterized in that, In step S2, the optimal candidate prediction model is selected as the basic prediction model based on the fitting accuracy, specifically as follows: Each candidate prediction model predicts attribute data for a preset surface interval; The fitting accuracy of the candidate prediction model is comprehensively evaluated by the difference between the obtained predicted value and the attribute data of the first depth interval of the sampling point, the determination coefficient of the candidate prediction model, and the root mean square error of the candidate prediction model. The candidate prediction model with the best fitting accuracy is selected as the base prediction model.
4. The three-dimensional soil property distribution prediction method according to claim 1, characterized in that, In step S2, the optimal candidate prediction model is selected as the basic prediction model based on the fitting accuracy, specifically as follows: Each candidate prediction model predicts attribute data for a preset surface interval. The difference between the predicted value and the attribute data for the first depth interval of the sampling point is taken as the surface prediction error. The candidate prediction model is scored based on the surface prediction error, with a higher score for a smaller error. This results in surface prediction error scores for multiple candidate prediction models. ; Calculate the coefficient of determination for each candidate prediction model, and assign a score to each candidate model based on the coefficient of determination. The higher the coefficient of determination, the higher the score. This process yields the coefficient of determination scores for multiple candidate prediction models. ; Calculate the root mean square error (RMSE) of each candidate prediction model, and assign a score to each candidate model based on the RMSE. The smaller the RMSE, the higher the score. This yields the RMSE scores for multiple candidate prediction models. ; Calculate the total score for each candidate prediction model: in, For the first The total score of each candidate prediction model , and The weights for the coefficient of determination score, root mean square error score, and surface prediction error score are respectively, satisfying: The total score The highest candidate prediction model As a basic prediction model.
5. The three-dimensional soil property distribution prediction method according to claim 1, characterized in that, In step S3, the parameters of the basic prediction model are optimized by minimizing the difference between the attribute data and the predicted value of the first depth interval of the sampling point. Specifically: The parameters of the basic prediction model can be obtained using the following formula. : in, Indicates minimization. This represents the attribute data of the first depth interval of the sampling point. Indicates the first standard depth range. This represents the predicted value from the basic prediction model.
6. The three-dimensional soil property distribution prediction method according to claim 1, characterized in that, Step S4 specifically includes the following sub-steps: S41. Use the basic prediction model to predict the attribute data of the Kth standard depth interval. If the obtained prediction value is found to be incorrect, proceed to the next step; otherwise, proceed to step S44. S42. The predicted values of the Kth standard depth interval are corrected by using the reverse distance weighting of the predicted values of each standard depth interval before the Kth standard depth interval. S43. Correct the parameters of the basic prediction model using the predicted values of the first K standard depth intervals; S44. Determine whether the Kth standard depth interval is the last standard depth interval. If yes, the iterative optimization ends; otherwise, K=K+1 and return to step S41.
7. The three-dimensional soil property distribution prediction method according to claim 6, characterized in that, The specific criteria for determining incorrect predictions are as follows: If the predicted value of the Kth standard depth interval is obtained and the dispersion of the predicted values of the first K standard depth intervals is greater than the preset dispersion, then the predicted value of the Kth standard depth interval is considered to be incorrect. If the predicted value is negative or zero, the predicted value is considered incorrect.
8. The three-dimensional soil property distribution prediction method according to claim 7, characterized in that, The dispersion of the predicted values for the first K standard depth intervals is obtained by evaluating the standard deviation or variance of the predicted values for the first K standard depth intervals.
9. The three-dimensional soil property distribution prediction method according to claim 6, characterized in that, The predicted values for the Kth standard depth interval are corrected using a reverse distance weighting method, specifically as follows: in, Indicates the first Predicted values for the standard depth range, For the first Weights of standard depth intervals Represents the Kth standard depth interval and the... Median distance of the standard depth range This represents the predicted value for the corrected Kth standard depth interval.
10. An electronic device, characterized in that, include: At least one memory for storing computer programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-9.
Citation Information
Patent Citations
Three-dimensional soil field water-holding capacity prediction method and system
CN110909467A
Soil organic matter prediction mapping method, device and equipment based on deep learning and medium
CN120147468A