Multi-factor soil quality evaluation method based on quantile random forest model
By using a quantile random forest model combined with coefficient of variation and principal component analysis to screen core indicators, and by utilizing remote sensing data to enhance the model's adaptability, the problems of indicator redundancy and uncertainty in soil quality assessment were solved. This resulted in a stable and highly adaptable multi-level soil quality assessment, providing reliable decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN UNIV OF TECH
- Filing Date
- 2026-03-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing soil quality assessment methods suffer from problems such as subjective redundancy in indicator selection, insufficient model generalization ability, and lack of uncertainty measurement in assessment results. These issues result in insufficient hierarchical differentiation of evaluation results and decreased prediction accuracy and stability of models when applied to different land use types or across geographical regions, making it difficult to support standardized operational assessments.
A multi-factor soil quality assessment method based on the quantile random forest model was adopted. Core indicators were screened through coefficient of variation analysis and principal component analysis. Soil quality index was constructed by combining vegetation index and topographic data obtained by remote sensing. The predicted value and uncertainty interval were output by the quantile random forest model to achieve multi-level soil quality assessment.
It effectively reduces data collection and model complexity, enhances model stability and adaptability, provides dynamic decision support including reliability information, and improves the reliability and regional applicability of evaluation results.
Smart Images

Figure CN121899384A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of soil quality assessment, and more specifically to a multi-factor soil quality assessment method based on a quantile random forest model. Background Technology
[0002] Soil quality assessment is an important foundation for sustainable agricultural management, land planning, and ecological restoration. Its core task is to quantify the soil’s ability to maintain ecological functions and productivity through a comprehensive analysis of the soil’s physical, chemical, and biological properties.
[0003] In terms of constructing the indicator system, existing methods often rely on expert experience or statistical analysis based on all available indicators, which often results in a large number of indicators and collinearity among them. This redundancy not only increases the cost and workload of field sampling and laboratory analysis, but also introduces noise in subsequent modeling, reducing the stability and interpretability of the model. More importantly, the dominant factors and variation characteristics of different soil layers are different, and most methods do not screen indicators hierarchically, resulting in the selected indicators not being able to fully represent the true quality status of each soil layer, and the hierarchical differentiation of the evaluation results is insufficient.
[0004] In terms of model construction and adaptability, most existing assessment models only use soil properties as input and fail to effectively integrate key environmental factors that affect soil formation and evolution. This single data source leads to a decline in the prediction accuracy and stability of the model when dealing with different land use types or cross-geographical applications due to the heterogeneity of the environmental background. The generalization ability of the model is limited, making it difficult to support standardized operational assessments.
[0005] In terms of presenting assessment results and supporting decision-making, most existing technologies only output a specific soil quality index or grade. This output format cannot reflect the uncertainty of the model prediction itself. For example, in areas with sparse data or for samples with attributes at the boundary, the reliability of the prediction results cannot be quantified. As a result, decision-makers cannot distinguish which results have high confidence and which have a large risk of error. Therefore, when making land management or ecological restoration decisions based on assessment results, potential uncertainties are easily overlooked, affecting the accuracy of the implementation of measures.
[0006] Therefore, how to design a multi-factor soil quality assessment method based on the quantile random forest model that can screen representative indicators, quantify prediction uncertainties, and integrate multi-source environmental information to improve the reliability and regional applicability of the assessment results is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a multi-factor soil quality assessment method based on a quantile random forest model, which aims to solve the problems of subjective redundancy in the selection of indicators, insufficient generalization ability of the model, and lack of uncertainty measurement in the assessment results in existing soil quality assessment methods. It can achieve a more comprehensive, stable, and adaptable evaluation of soil profile quality, thereby providing a more reliable basis for land management and ecological decision-making.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] A multi-factor soil quality assessment method based on a quantile random forest model includes the following steps: S1. Collect multi-layer soil samples of different land use types in the target area and obtain the normalized vegetation index (NDVI) and slope data at each sampling point. S2. Measure the physical, chemical and biological indicators of the multilayer soil samples to obtain raw indicator data; S3. Standardize the original indicator data to obtain standardized indicator data; S4. Based on the standardized index data, use coefficient of variation analysis and principal component analysis to screen out the core indicators that characterize the soil quality of each soil layer and obtain the minimum dataset for each soil layer. S5. Based on the minimum dataset of each soil layer, construct the corresponding soil quality index (SQI); S6. Using the index data in the minimum dataset, as well as the corresponding Normalized Vegetation Index (NDVI) and slope data, as input features, and the Soil Quality Index (SQI) as the output label, train the quantile random forest model. S7. Input the corresponding features of the sample to be evaluated into the trained quantile random forest model to obtain the point prediction value and prediction interval of the soil quality index SQI for each soil layer, and determine the confidence level based on the width of the prediction interval. S8. Based on the point prediction value of the soil quality index (SQI), output the soil quality level of each soil layer; and perform weighted fusion of the soil quality index (SQI) of each soil layer to obtain the comprehensive soil profile quality index.
[0010] Preferably, in step S1, multi-layer soil samples are collected at depths of 0-20cm, 20-40cm, 40-60cm, 60-80cm, and 80-100cm; the Normalized Difference Vegetation Index (NDVI) is calculated from satellite remote sensing images of the same region and period as the soil sampling; and the slope data is calculated from the Digital Elevation Model (DEM) of the target area.
[0011] Preferably, in step S3, the standardization process includes: For positive indicators, standardized values Represented as: For negative indicators, their standardized values Represented as:
[0012] Where X is the original measured value of the indicator. and These are the maximum and minimum values of the indicator across all samples, respectively.
[0013] Preferably, S4 includes: For the standardized index data of each soil layer, the coefficient of variation (CV) of each index is calculated, and the indexes are divided into three categories: high sensitivity, medium sensitivity, and low sensitivity. Highly sensitive and some moderately sensitive indicators were selected to form a candidate indicator set for each soil layer; Principal component analysis was performed on the candidate index set to extract principal components with eigenvalues greater than 1. In each extracted principal component, indicators with an absolute factor loading of not less than 0.5 are selected as candidate representative indicators for that principal component; and when the absolute factor loading of the same indicator in multiple principal components is not less than 0.5, it is assigned to the principal component with the largest absolute factor loading. Calculate the Norm value of each index within each principal component, where the Norm value is the L2 norm of the factor loadings of that index on all principal components; The indicator with the largest Norm value is selected as the representative indicator of the principal component; if the absolute value of the correlation coefficient between the two indicators with the highest Norm values is less than 0.7, then both indicators are selected. By merging the representative indices selected from all principal components, the minimum dataset for each soil layer is obtained.
[0014] Preferably, in step S5, the soil quality index (SQI) is expressed as:
[0015] Where m is the number of indicators concentrated in the minimum data set of the corresponding soil layer. For the i-th metric in the smallest dataset, Let be the weight of the i-th indicator.
[0016] Preferably, the weights are determined. include: For each metric in the minimum dataset, determine the principal component with the largest absolute value of its loading, and use it as the representative principal component of that metric. Obtain the absolute values of the loadings of this index on its representative principal components. And obtain the variance contribution rate of the representative principal component. ; Calculate the initial weight of this indicator. ; The initial weights of all indicators are normalized to obtain the final weights of each indicator. .
[0017] Preferably, in step S6, the quantile random forest model outputs the point prediction value of the soil quality index (SQI) at the 0.5 quantile. It outputs the 0.05 quantile prediction value. and 0.95 quantile predicted value As the upper and lower bounds of the prediction interval; The predicted point value Represented as:
[0018] Where N is the total number of decision trees in the quantile random forest model. Let t be the 0.5 quantile prediction of the t-th decision tree for the input feature vector x.
[0019] Preferably, in step S7, determining the confidence level based on the prediction interval width includes: Calculate the prediction interval width ; Compare W with a preset threshold , Comparison, among which ;like If it is, then it is judged as high confidence; if If it is, then it is determined to be of medium confidence level; if If the confidence level is low, then it is determined to be low confidence.
[0020] Preferably, in step S8, the soil quality grades for each soil layer are output as follows: Map the point prediction values of the Soil Quality Index (SQI) to the [0,1] interval:
[0021] in, This is the predicted value of the SQI point. , These represent the maximum and minimum predicted values of SQI points for all samples within the evaluation area, respectively. Will They are divided into different quality levels according to preset numerical ranges.
[0022] Preferably, in step S8, the comprehensive quality index of the soil profile is... Represented as:
[0023] Where L represents the total number of soil layers. The predicted value of the Soil Quality Index (SQI) for the j-th soil layer is given. Let be the weight of the j-th soil layer.
[0024] As can be seen from the above technical solution, compared with the prior art, the technical solution of the present invention has the following beneficial effects: 1. This method constructs a minimal dataset by dividing the soil layers, effectively overcoming the problems of index redundancy and collinearity in soil quality assessment. By combining coefficient of variation analysis and principal component analysis, it can screen out a few core and complementary representative indicators from physical, chemical and biological indicators. This not only simplifies the input feature dimensions required for subsequent models and reduces the cost of data collection and measurement, but also improves the stability of the assessment model by eliminating redundant information, ensuring the scientific nature of the evaluation process.
[0025] 2. By using the normalized vegetation index and topographic slope obtained from remote sensing as environmental covariates and inputting them into the model along with core soil indicators, the evaluation model can inherently reflect the potential modulating effect of vegetation cover and topographic features on soil quality, thereby enhancing the model's adaptability and generalization ability to different land use types and complex geographical environments.
[0026] 3. By utilizing the quantile random forest model, this method achieves a leap from providing a single predicted value to providing a complete information set including predicted value, uncertainty interval, and confidence level. The prediction interval output by this method quantitatively characterizes the uncertainty of the assessment results, while the confidence level based on this provides decision-makers with intuitive risk warnings. This upgrades the soil quality assessment results from static level judgments to dynamic decision support tools that include reliability information, which helps guide the optimal allocation of subsequent monitoring resources and risk avoidance of land management measures. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0028] Figure 1 A flowchart of a multi-factor soil quality assessment method based on a quantile random forest model is provided for an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the process of obtaining the minimum dataset for each soil layer, as provided in an embodiment of the present invention. Figure 3 This is a graph showing the linear regression analysis results between the SQI-TDS full dataset and the SQI-MDS minimum dataset provided in this embodiment of the invention. Figure 4 A graph showing the results of linear regression analysis of observed and predicted soil quality index values provided in this embodiment of the invention; Figure 5 Box plots of SQI prediction residuals for different soil layers provided in embodiments of the present invention; Figure 6 Box plots of SQI prediction residuals for different land use types provided in embodiments of the present invention; Figure 7 This is a distribution map of soil quality index levels for different soil layers provided in an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] like Figure 1 As shown, this embodiment provides a multi-factor soil quality assessment method based on a quantile random forest model, including the following steps: S1. Collect multi-layer soil samples of different land use types in the target area and obtain the normalized vegetation index (NDVI) and slope data at each sampling point. S2. Measure the physical, chemical and biological indicators of the multilayer soil samples to obtain raw indicator data; S3. Standardize the original indicator data to obtain standardized indicator data; S4. Based on the standardized index data, use coefficient of variation analysis and principal component analysis to screen out the core indicators that characterize the soil quality of each soil layer and obtain the minimum dataset for each soil layer. S5. Based on the minimum dataset of each soil layer, construct the corresponding soil quality index (SQI); S6. Using the index data in the minimum dataset, as well as the corresponding Normalized Vegetation Index (NDVI) and slope data, as input features, and the Soil Quality Index (SQI) as the output label, train the quantile random forest model. S7. Input the corresponding features of the sample to be evaluated into the trained quantile random forest model to obtain the point prediction value and prediction interval of the soil quality index SQI for each soil layer, and determine the confidence level based on the width of the prediction interval. S8. Based on the point prediction value of the soil quality index (SQI), output the soil quality level of each soil layer; and perform weighted fusion of the soil quality index (SQI) of each soil layer to obtain the comprehensive soil profile quality index.
[0031] This method selects core evaluation indicators through an objective data-driven approach, reducing subjective human interference; it uses a quantile random forest model to synchronously output predicted values and their uncertainty ranges, enabling the evaluation results to combine quality judgment and reliability measurement; and by integrating multi-source data and hierarchical modeling, it achieves a comprehensive evaluation of soil profile quality that is systematic, stable, and more adaptable.
[0032] The following provides a further explanation of each step and related features in the above method; In this embodiment, S1, multi-layer soil samples of different land use types in the target area are collected, and the normalized vegetation index (NDVI) and slope data of each sampling point are obtained. In practice, sampling points were set up within the study area for different land types, including woodland, sloping farmland, dammed land, grassland, and terraced fields, based on the principles of representativeness and uniformity. A nine-point mixed sampling method was used, with soil samples collected at each sampling point at five depth intervals: 0-20cm, 20-40cm, 40-60cm, 60-80cm, and 80-100cm. GPS coordinates and other information were recorded. Simultaneously, satellite remote sensing images from the same period and region as the soil sampling were selected, and the Normalized Difference Vegetation Index (NDVI) was obtained by calculating the ratio of near-infrared to red light band reflectance. At the same time, slope raster data was calculated based on the digital elevation model (DEM) of the area, and the corresponding NDVI and slope values were extracted based on the sampling point coordinates. This correlation between remote sensing and topographic environmental factors and soil samples provides data support for enhancing the spatial generalization ability of the subsequent model.
[0033] In this embodiment, S2, the physical, chemical, and biological indicators of the multilayer soil sample are measured to obtain raw indicator data; To achieve scientific selection and optimization of evaluation indicators while ensuring the comprehensiveness of the indicator system, the selection of evaluation indicators in this embodiment gradually extends from traditional physicochemical indicators to indicators including biological characteristics and ecological functions, fully guaranteeing the comprehensiveness and scientific nature of the soil quality evaluation indicator selection. Among them, physical indicators include bulk density, water content, and mechanical composition; chemical indicators include organic matter, total nitrogen, total phosphorus, available potassium, available phosphorus, pH, ammonia nitrogen, and nitrate nitrogen; biological indicators include microbial biomass carbon, urease, alkaline phosphatase, and sucrase. Details are shown in Table 1 below. Table 1
[0034] In this embodiment, S3, the original index data is standardized to obtain standardized index data. The standardization process includes: For positive indicators (such as organic matter and microbial biomass carbon), the standardized value Represented as: For negative indicators (such as bulk density), their standardized values Represented as:
[0035] Where X is the original measured value of the indicator. and These are the maximum and minimum values of the indicator across all samples, respectively.
[0036] In this embodiment S4, based on the standardized index data, the core indicators characterizing the soil quality of each soil layer are screened out using coefficient of variation analysis and principal component analysis, and the minimum dataset of each soil layer is obtained. like Figure 2 The specific steps include: S41. For the standardized index data of each soil layer, calculate the coefficient of variation (CV) for each index:
[0037] in, The standard deviation of this indicator. This is the mean of the indicator; The indicators are divided into three categories: high sensitivity, medium sensitivity, and low sensitivity. The thresholds for high sensitivity, medium sensitivity, and low sensitivity are as follows: CV ≥ 35% is high sensitivity, 15% ≤ CV < 35% is medium sensitivity, and CV < 15% is low sensitivity. Taking the 0–20 cm soil layer as an example, the coefficient of variation at a depth of 0–20 cm ranges from 0.7% to 54.36%, with a mean of 20.51. The sensitivity classification results for all indicators are shown in Table 2 below: Table 2
[0038] S42. Select highly sensitive and some moderately sensitive indicators to form a candidate indicator set for each soil layer; S43. Perform principal component analysis on the candidate index set and extract principal components with eigenvalues greater than 1 (cumulative explained variance reaches a preset threshold of 80%). In this embodiment, seven principal components with eigenvalues greater than 1 (PC1-PC7) are extracted, with a cumulative variance contribution rate of 85.40%, which is sufficient to represent most of the original information. S44. In each extracted principal component, select an index with an absolute factor loading of not less than 0.5 as a candidate representative index of that principal component; and when the absolute factor loading of the same index in multiple principal components is not less than 0.5, assign it to the principal component with the largest absolute factor loading. S45. Calculate the Norm value of each index within each principal component. The Norm value is the L2 norm of the factor loading of the index on all principal components. S46. Select the indicator with the largest Norm value as the representative indicator of the principal component; if the absolute value of the correlation coefficient between the two indicators with the highest Norm values is less than 0.7, then select both indicators at the same time. S47. Merge the representative indices selected from all principal components to obtain the minimum dataset for each soil layer.
[0039] The final minimum dataset MDS for the 0-20cm soil layer was determined to be: Si (physical), BD (physical), SOM (chemical), SU (biological), Zn (chemical), pH (chemical), and MBC (biological). This MDS contains only 7 indicators, which is a significant reduction from the original more than 20 indicators. It covers the three major categories of physical, chemical, and biological, and characterizes various functions such as soil structure, fertility, and biological activity.
[0040] The minimum dataset is shown in Table 3 below: Table 3
[0041] like Figure 3 As shown, the minimum dataset MDS is significantly linearly correlated with the full dataset TDS (R²≈0.74), verifying the representativeness and substitutability of the minimum dataset. It can be seen that it can effectively eliminate redundant indicators, reduce the impact of collinearity, and form a core indicator system with a small number of indicators, strong representativeness, and complementary information. This significantly reduces the cost of subsequent data collection and model complexity, while ensuring the scientific nature and stability of the evaluation.
[0042] In this embodiment, S5, based on the minimum dataset of each soil layer, the corresponding soil quality index (SQI) is constructed. The Soil Quality Index (SQI) is expressed as follows:
[0043] Where m is the number of indicators concentrated in the minimum data set of the corresponding soil layer. For the i-th metric in the smallest dataset, Let be the weight of the i-th indicator.
[0044] Furthermore, determine the weights. include: For each metric in the minimum dataset, determine the principal component with the largest absolute value of its loading, and use it as the representative principal component of that metric. Obtain the absolute values of the loadings of this index on its representative principal components. And obtain the variance contribution rate of the representative principal component. ; Calculate the initial weight of this indicator. ; The initial weights of all indicators are normalized to obtain the final weights of each indicator. .
[0045] In this embodiment S6, the index data in the minimum dataset, as well as the corresponding Normalized Vegetation Index (NDVI) and slope data, are used as input features, and the Soil Quality Index (SQI) is used as the output label to train the quantile random forest model. Among them, the quantile random forest model outputs point predictions of the soil quality index (SQI) at the 0.5 quantile. It outputs the 0.05 quantile prediction value. and 0.95 quantile predicted value As the upper and lower bounds of the prediction interval; The predicted point value Represented as:
[0046] Where N is the total number of decision trees in the quantile random forest model. Let t be the 0.5 quantile prediction of the t-th decision tree for the input feature vector x.
[0047] Specifically, taking the training of the 0-20cm soil layer model as an example: the input feature is a 9-dimensional vector for each sample, including the standardized values of the 7 indicators in its MDS, as well as the standardized values of NDVI and slope corresponding to the sampling point. The output label is the SQI value of the sample in the 0-20cm soil layer calculated in the above steps. The quantile random forest algorithm is used, the number of decision trees is set to 500, the random seed is fixed to ensure that the results are repeatable, and stratified sampling is adopted to divide the training set and the validation set in a ratio of 70%:30% to ensure that different land use types are distributed in both groups. Furthermore, to evaluate the predictive performance of the constructed quantile random forest model, this embodiment uses the coefficient of determination R², root mean square error (RMSE), mean absolute error (MAE), and mean bias to compare and analyze the observed and predicted values of the soil quality index (SQI). like Figure 4 As shown, there is a positive correlation between the SQI predicted values and the observed values. The scatter plots of each sample point are mainly distributed near the 1:1 reference dashed line, and the linear fitting solid line is close to the 1:1 line. The statistical results based on all sample points are: R²=0.77, RMSE=0.037, MAE=0.028, Bias= 0.001 indicates that the method in this embodiment can explain the variation of SQI well, with a small overall prediction error and no obvious systematic overestimation or underestimation; To further examine the predicted stability across different soil layers and land use types, this embodiment plots residual box plots grouped by midpoint of soil layer depth and land use type, as shown below. Figure 5 As shown, the median residuals for each soil layer (0–20, 20–40, 40–60, 60–80, 80–100 cm) are all close to 0, with values approximately [missing value]. The box height is relatively small, ranging from 0.01 to 0.00, indicating that the overall fluctuation of the prediction error for each soil layer is low; for example... Figure 6 As shown, the median residuals for different land types (BD = dammed land, CD = grassland, LD = forest land, PGD = sloping farmland, TT = terraced fields) are also close to 0, mainly distributed in The residuals are between 0.01 and 0.01, and most of the sample residuals fall within a narrow quartile range, indicating that the method in this embodiment has good predictive consistency and robustness under different land use types.
[0048] In this embodiment, S7, the corresponding features of the sample to be evaluated are input into the trained quantile random forest model to obtain the point prediction value and prediction interval of the soil quality index SQI for each soil layer, and the confidence level is determined according to the width of the prediction interval. Determining the confidence level based on the prediction interval width includes: Calculate the prediction interval width ; Compare W with a preset threshold , Comparison, among which ;like If it is, then it is judged as high confidence; if If it is, then it is determined to be of medium confidence level; if If the threshold is not met, it is considered a low confidence level; the threshold at this point is... and Determined by the quantiles of the predicted interval width from the validation set samples: Take the 33rd percentile. Take the 66th percentile.
[0049] Specifically, the nine standardized features of a new, untrained sloping farmland sampling point (0-20cm soil layer) are input into a pre-trained QRF model for the 0-20cm soil layer, and the model outputs the predicted SQI value for that point. =0.52, the prediction interval is [0.43, 0.65], then the interval width W=0.22, based on the threshold obtained from the statistical analysis of the validation set (for this soil layer). , =0.27), because Therefore, the confidence level of the SQI prediction result for this sample point is determined to be medium. This step provides three layers of information: soil quality level, predicted possible fluctuation range, qualitative assessment of the reliability of the prediction, and the determined confidence level. This provides a reliability measure for the soil quality level and profile comprehensive quality index output in subsequent steps. It transforms the uncertainty of model prediction into an intuitive confidence level classification, so that the final evaluation result not only includes quality level information, but also adds the credibility of the results for each level or region. In soil management or ecological restoration decisions, it can distinguish between reliable conclusions with high confidence and areas with low confidence that require further verification, thus improving the risk warning capability of the assessment results.
[0050] In this embodiment, S8, based on the point prediction value of the Soil Quality Index (SQI), the soil quality level of each soil layer is output; and the soil quality index (SQI) of each soil layer is weighted and fused to obtain the comprehensive soil profile quality index.
[0051] The output soil quality grades for each soil layer include: Map the point prediction values of the Soil Quality Index (SQI) to the [0,1] interval:
[0052] in, This is the predicted value of the SQI point. , These represent the maximum and minimum predicted values of SQI points for all samples within the evaluation area, respectively. Will They are divided into different quality levels according to preset numerical ranges; like Figure 7 As shown, based on the soil quality classification standards (Level I: 0-0.2, Level II: 0.2-0.4, Level III: 0.4-0.6, Level IV: 0.6-0.8, Level V: 0.8-1.0), the proportion of each level in different soil layers can be statistically analyzed to obtain the soil quality grade distribution results shown in the figure. The classification demonstrates that it can distinguish the differences in soil quality across different soil layers, providing a basis for vertical profile evaluation. Furthermore, the comprehensive quality index of soil profiles Represented as:
[0053] Where L represents the total number of soil layers. The predicted value of the Soil Quality Index (SQI) for the j-th soil layer is given. Let be the weight of the j-th soil layer.
[0054] By combining the distribution maps of each soil layer, the vertical distribution and overall condition of soil quality at that location can be systematically assessed, providing a basis for decisions on deep soil improvement or utilization.
[0055] The multi-factor soil quality assessment method based on the quantile random forest model in this embodiment achieves a comprehensive diagnosis of soil quality at multiple levels, including reliability assessment, through data collection, indicator screening, index construction, intelligent modeling, and result output. This method not only improves the efficiency and accuracy of assessment, but also provides a more reliable decision-making basis for refined land management, identification of priority areas for ecological restoration, and sustainable agricultural practices through its output of quality level and confidence level information.
[0056] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0057] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-factor soil quality assessment method based on a quantile random forest model, characterized in that, Includes the following steps: S1. Collect multi-layer soil samples of different land use types in the target area and obtain the normalized vegetation index (NDVI) and slope data at each sampling point. S2. Measure the physical, chemical and biological indicators of the multilayer soil samples to obtain raw indicator data; S3. Standardize the original indicator data to obtain standardized indicator data; S4. Based on the standardized index data, use coefficient of variation analysis and principal component analysis to screen out the core indicators that characterize the soil quality of each soil layer and obtain the minimum dataset for each soil layer. S5. Based on the minimum dataset of each soil layer, construct the corresponding soil quality index (SQI); S6. Using the index data in the minimum dataset, as well as the corresponding Normalized Vegetation Index (NDVI) and slope data, as input features, and the Soil Quality Index (SQI) as the output label, train the quantile random forest model. S7. Input the corresponding features of the sample to be evaluated into the trained quantile random forest model to obtain the point prediction value and prediction interval of the soil quality index SQI for each soil layer, and determine the confidence level based on the width of the prediction interval. S8. Based on the point prediction value of the soil quality index (SQI), output the soil quality level of each soil layer; and perform weighted fusion of the soil quality index (SQI) of each soil layer to obtain the comprehensive soil profile quality index.
2. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 1, characterized in that, In S1, multi-layer soil samples are collected at depths of 0-20cm, 20-40cm, 40-60cm, 60-80cm, and 80-100cm; the Normalized Difference Vegetation Index (NDVI) is calculated from satellite remote sensing images of the same region and period as the soil sampling; and the slope data is calculated from the Digital Elevation Model (DEM) of the target area.
3. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 1, characterized in that, In S3, the standardization process includes: For positive indicators, standardized values Represented as: For negative indicators, their standardized values Represented as: Where X is the original measured value of the indicator. and These are the maximum and minimum values of the indicator across all samples, respectively.
4. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 1, characterized in that, S4 includes: For the standardized index data of each soil layer, the coefficient of variation (CV) of each index is calculated, and the indexes are divided into three categories: high sensitivity, medium sensitivity, and low sensitivity. Highly sensitive and some moderately sensitive indicators were selected to form a candidate indicator set for each soil layer; Principal component analysis was performed on the candidate index set to extract principal components with eigenvalues greater than 1. In each extracted principal component, indicators with an absolute factor loading of not less than 0.5 are selected as candidate representative indicators for that principal component; and when the absolute factor loading of the same indicator in multiple principal components is not less than 0.5, it is assigned to the principal component with the largest absolute factor loading. Calculate the Norm value of each index within each principal component, where the Norm value is the L2 norm of the factor loadings of that index on all principal components; The indicator with the largest Norm value is selected as the representative indicator of the principal component; if the absolute value of the correlation coefficient between the two indicators with the highest Norm values is less than 0.7, then both indicators are selected. By merging the representative indices selected from all principal components, the minimum dataset for each soil layer is obtained.
5. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 1, characterized in that, In S5, the soil quality index (SQI) is expressed as: Where m is the number of indicators concentrated in the minimum data set of the corresponding soil layer. For the i-th metric in the smallest dataset, Let be the weight of the i-th indicator.
6. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 5, characterized in that, Determine the weights include: For each metric in the minimum dataset, the principal component with the largest absolute loading is identified as the representative principal component of that metric. Obtain the absolute values of the loadings of this index on its representative principal components. And obtain the variance contribution rate of the representative principal component. ; Calculate the initial weight of this indicator. ; The initial weights of all indicators are normalized to obtain the final weights of each indicator. .
7. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 1, characterized in that, In step S6, the quantile random forest model outputs point predictions of the soil quality index (SQI) at the 0.5 quantile. It outputs the 0.05 quantile prediction value. and 0.95 quantile predicted value As the upper and lower bounds of the prediction interval; The predicted point value Represented as: Where N is the total number of decision trees in the quantile random forest model. Let t be the 0.5 quantile prediction of the t-th decision tree for the input feature vector x.
8. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 1, characterized in that, In step S7, determining the confidence level based on the prediction interval width includes: Calculate the prediction interval width ; Compare W with a preset threshold , Comparison, among which ;like If it is, then it is judged as high confidence; if If it is, then it is determined to be of medium confidence level; if If the confidence level is low, then it is determined to be low confidence.
9. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 1, characterized in that, In step S8, the soil quality grades for each soil layer are output as follows: Map the point prediction values of the Soil Quality Index (SQI) to the [0,1] interval: in, This is the predicted value of the SQI point. , These represent the maximum and minimum predicted values of SQI points for all samples within the evaluation area, respectively. Will They are divided into different quality levels according to preset numerical ranges.
10. The multi-factor soil quality assessment method based on a quantile random forest model according to claim 1, characterized in that, In S8, the comprehensive quality index of the soil profile Represented as: Where L represents the total number of soil layers. The predicted value of the Soil Quality Index (SQI) for the j-th soil layer is given. Let be the weight of the j-th soil layer.
Citation Information
Patent Citations
Screening and weight determining method for wetland ecosystem health evaluation index system
CN106548015A
Method for evaluating soil quality of larch forest
CN116341984A