Soil typical sampling point sampling method based on machine learning
Through machine learning technology and redundant feature elimination strategy, the soil type prediction model is optimized, and the problems of poor representativeness and low efficiency in soil sample sampling are solved, achieving high-precision soil type prediction.
Patent Information
- Application Number
- CN202510449624.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
The existing soil sample point sampling methods have the problems of model overfitting caused by poor sample representativeness, low sampling efficiency and feature redundancy, which affects the prediction accuracy of soil type.
The soil type prediction model is optimized through environmental covariate resampling, random sampling, redundant feature elimination and weighted sampling strategies, and the soil type prediction model is used to predict soil type using a random forest classification model.
It improves the representativeness and prediction accuracy of soil sample points, provides a spatial prediction scheme for soil type in large-scale areas, and ensures the reliability and accuracy of prediction results.
Smart Images

Figure CN120337020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soil type prediction, and particularly to a method for sampling typical soil points based on machine learning. Background Art
[0002] The distribution of soil types has important impacts on agricultural production, environmental protection, land use planning, etc. Traditional soil sampling methods rely on manual selection of sample locations, usually suffering from problems such as poor sample representativeness, low sampling efficiency, and uneven spatial distribution. With the continuous development of remote sensing technology and environmental data, soil type spatial prediction methods based on environmental covariates (such as climate, terrain, etc.) have gradually attracted attention. However, the existing sampling methods still have the following problems: Insufficient sample representativeness. Traditional methods of random sampling fail to fully consider the spatial representativeness of soil sampling points, which may lead to insufficiently representative samples and affect the prediction accuracy.
[0003] Low sampling efficiency. Traditional methods rely on manual experience for sampling, with low efficiency and possible sample bias.
[0004] Problems of feature redundancy and overfitting: Environmental covariates usually have redundant features, which may lead to model overfitting and affect the accurate prediction of soil types.
[0005] Therefore, how to improve the sampling efficiency of soil sampling points, enhance the representativeness of samples, and optimize the prediction ability of the model through machine learning methods has become an important research topic in the current field of soil mapping. Summary of the Invention
[0006] Aiming at the problems of few soil sampling points, poor sample representativeness, improper processing of environmental covariate data, and model overfitting existing in the existing soil sampling points sampling and soil type prediction methods mentioned in the above background art, the present invention proposes a method for sampling typical soil points based on machine learning. By introducing machine learning technology, redundant feature elimination strategies, and considering the distribution of feature environmental covariates for weighted sampling strategies, the soil type prediction model is optimized. The present invention can improve the representativeness of soil sampling points, optimize the prediction accuracy of soil types, and provide a soil type spatial prediction scheme in a large-scale area.
[0007] To achieve the above object, the present invention adopts the following technical solutions: The present invention provides a method for sampling typical soil points based on machine learning, including the following steps: S1: Resample the spatial resolution of the selected environmental covariate data to make it consistent with the spatial resolution of the national second soil census soil map; S2: Conduct a preliminary random sampling on the target area to form a preliminary soil sample dataset, and combine the soil type of each soil sample with the corresponding environmental covariates to construct a regression matrix; S3: Through autocorrelation analysis of the regression matrix, identify and remove the highly correlated features among the environmental covariates, and use the recursive feature elimination algorithm to screen out the features that have the most influence on the prediction of the soil type, obtaining an optimized sub-regression matrix; S4: Use the optimized sub-regression matrix to train a random forest classification model to predict the soil species type. After model training, extract the feature importance scores, sort them according to the importance scores, and identify the environmental covariates ranked at the top; S5: For each soil species type, cyclically select the environmental covariates ranked at the top, divide the value range of each environmental covariate into multiple value range intervals, perform a union operation based on the value range intervals to obtain an intersection area, and select typical soil samples by assigning a high sampling weight to the intersection area; S6: Overlay the typical soil samples with the corresponding environmental covariates, reconstruct the regression matrix, and after performing step S3, obtain a new sub-regression matrix. Use the new sub-regression matrix to train a random forest classification model to establish a spatial prediction model of the soil type. Input the important environmental covariates, and obtain the soil type distribution map and its uncertainty distribution map of the large-scale area through the spatial prediction model, thereby realizing the accurate prediction of the soil type in a large range.
[0008] As a further description of the present invention, the data sources of the environmental covariates include climate data, terrain data, land cover data, and geographic information system data.
[0009] As a further description of the present invention, resample the spatial resolution of the selected environmental covariate data, which is achieved by using nearest neighbor interpolation, bilinear interpolation, or Kriging interpolation methods.
[0010] As a further description of the present invention, during the preliminary random sampling of the target area, select a sample size that is 20 times the number of soil species types and limit the minimum sampling number to form a preliminary soil sample dataset.
[0011] As a further description of the present invention, during the preliminary random sampling process, select soil samples by the random sampling method, and ensure that the number of samples of each soil type conforms to the regional distribution of the soil type.
[0012] As a further description of the present invention, in the regression matrix constructed in step S2, the independent variable is the environmental covariate, and the dependent variable is the soil species type.
[0013] As a further illustration of the present invention, the identification and removal of highly correlated features among the environmental covariates by performing autocorrelation analysis on the regression matrix specifically includes: Screen out the covariates with an absolute value of correlation greater than 0.85 between two environmental covariates, and delete the environmental covariate with higher correlation with other environmental covariates. Check all environmental covariates in turn until the absolute value of the correlation between all environmental covariates is not greater than 0.85.
[0014] As a further illustration of the present invention, the use of the recursive feature elimination algorithm to screen out the features most influential on the soil type prediction to obtain an optimized sub-regression matrix specifically includes: Use the recursive feature elimination algorithm to perform multiple iterations on the regression matrix, combine the model performance evaluation metrics to optimize the feature set, gradually eliminate irrelevant or less contributing feature variables to the model, and finally select the optimal feature subset with the greatest contribution to the soil type prediction through cross-validation, thereby obtaining the optimized sub-regression matrix.
[0015] As a further illustration of the present invention, the random forest classification model is based on multiple decision trees, predicts the soil type through the way of ensemble learning, can handle high-dimensional environmental covariate data, and obtains the final prediction result through a voting mechanism, and the training process of the random forest classification model adopts multiple cross-validation techniques.
[0016] As a further illustration of the present invention, in step S4, specifically identify the top 5 most important environmental covariates, and in step S5, specifically divide the value range of each environmental covariate into 3 value range intervals.
[0017] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention optimizes the soil type prediction model by introducing machine learning technology, redundant feature elimination strategy and considering the distribution of feature environmental covariates to perform weighted sampling strategy. The present invention can improve the representativeness of soil sample points, optimize the prediction accuracy of soil types, and provide a soil type spatial prediction scheme in a large-scale area.
[0018] Other features and advantages of this technical solution will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing this technical solution. The purpose and other advantages of this technical solution can be achieved and obtained through the structures specifically pointed out in the written specification and the drawings.
[0019] The technical solution of this technical solution will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0020] The accompanying drawings are used to provide a further understanding of the technical solution, and constitute a part of the specification. Together with the embodiments of the technical solution, they are used to explain the technical solution and do not constitute a limitation to the technical solution. In the accompanying drawings: Figure 1 It is a flowchart of the method for sampling typical soil points based on machine learning provided by the present invention.
[0021] Figure 2 It is a flowchart of the role of the method for sampling typical soil points based on machine learning provided by the present invention in large-scale soil type mapping. Detailed implementation manners
[0022] The following describes the preferred embodiments of the technical solution with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the technical solution and are not used to limit the technical solution.
[0023] As Figure 1 shown, the present invention provides a method for sampling typical soil points based on machine learning, including the following steps: S1: Spatial resolution resampling of environmental covariates: Resample the spatial resolution of the selected environmental covariate data to make it consistent with the spatial resolution of the national second soil census soil map.
[0024] In order to make the environmental covariate data consistent with the national second soil census (abbreviated as "second census") soil map in terms of spatial resolution, the present invention first resamples the spatial resolution of the environmental covariates. The purpose of resampling is to ensure the complete spatial consistency between the environmental covariates and the second census soil map, so that the subsequent soil samples and environmental covariates can be accurately matched, avoiding errors caused by inconsistent spatial resolutions.
[0025] Further, the methods used by the present invention for resampling include: nearest neighbor interpolation, bilinear interpolation, and Kriging interpolation, etc. By selecting a suitable resampling method, the geographical information consistency of the environmental data can be maintained, and the accuracy of subsequent model analysis can be improved.
[0026] Further, the above-mentioned environmental covariate data sources include climate data, terrain data, land cover data, and geographical information system (GIS) data.
[0027] Specifically, in the resampling process, set the spatial resolution of the target soil map as R_target, and the original spatial resolution of the environmental covariates as R_original. Adjust the spatial resolution of the environmental covariate data through interpolation operations to make the resolution between the environmental covariates and the second census soil map completely consistent, so as to ensure the consistency and accurate matching of the data.
[0028] S2: Preliminary random sampling and regression matrix construction: Conduct preliminary random sampling on the target area to form a preliminary soil sample dataset, and combine the soil type of each soil sample with the corresponding environmental covariates to construct a regression matrix.
[0029] The present invention uses a preliminary random sampling method to select samples from the target area. The selected soil sample size is 20 times the number of target soil species types, and the minimum sampling quantity is limited to 10. Each soil sample combines its soil type with the corresponding environmental covariates to construct a regression matrix. The independent variables in the regression matrix are environmental covariate data, while the dependent variable is soil type information. This regression matrix is used for the training of subsequent machine learning models to achieve spatial prediction of soil types.
[0030] During the preliminary random sampling process, the present invention selects soil samples through a random sampling method to ensure that the number of samples of each soil type conforms to the regional distribution of that soil type, thereby ensuring that the regression matrix has a certain representativeness, avoiding the problem of too many or too few samples of a certain soil type, and further improving the effectiveness of model training.
[0031] S3: Redundant feature elimination: Through autocorrelation analysis of the regression matrix, identify and remove highly correlated features among environmental covariates, and use a recursive feature elimination algorithm to screen out the features that have the greatest impact on soil type prediction, obtaining an optimized sub-regression matrix.
[0032] After the regression matrix is constructed, the present invention performs redundant feature elimination on the regression matrix to ensure that the model training process is not interfered by irrelevant or redundant features. Specifically, covariates with an absolute value of correlation greater than 0.85 between two environmental covariates are screened out, and the environmental covariate with a higher correlation with other environmental covariates is deleted. All environmental covariates are checked in turn until the absolute value of the correlation between all environmental covariates is not greater than 0.85. This process identifies highly correlated features by calculating the correlation between environmental covariates, and covariates with an absolute value of correlation greater than 0.85 will be eliminated. This step effectively reduces the complexity and training time of the model by reducing redundant features.
[0033] Then, the Recursive Feature Elimination (RFE) algorithm is used to further screen out the features that are most influential in predicting soil types. The RFE algorithm recursively removes unimportant or less contributing features, evaluates the impact of features on model performance metrics (such as accuracy, F1 score, etc.) in each iteration, gradually eliminates irrelevant or less contributing feature variables, and finally selects the features that contribute the most to soil type prediction through cross-validation to form an optimal feature subset, thereby obtaining an optimized sub-regression matrix. Through this process, the regression matrix is optimized, and the training efficiency and accuracy of the model are significantly improved.
[0034] S4: Train a random forest classification model and extract feature-important feature variables: Use the optimized sub-regression matrix to train a random forest classification model to predict soil species types. After model training, extract the feature importance scores, sort them according to the importance scores, and identify the top-ranked environmental covariates.
[0035] Based on the optimized regression matrix, the present invention uses a random forest classification model to predict soil types. Random forest is an ensemble learning method that constructs multiple decision trees and obtains the final prediction result through a voting mechanism. Since random forest can handle high-dimensional data and its training process uses multiple cross-validation techniques to prevent overfitting and improve the generalization ability of the model, thus ensuring the consistency and reliability of the prediction results under different regions and different environmental conditions, it is very suitable for the prediction task of soil types.
[0036] After the model training is completed, by analyzing the feature importance scores of the random forest, the environmental covariates that are most important for soil type prediction can be identified. The feature importance scores are accumulated through the training results of each tree in the random forest to evaluate the contribution of each feature in the prediction process. These important features will provide an important basis for subsequent typical sample point selection.
[0037] S5: Selection of typical soil samples and weighted sampling: For each soil species type, cyclically select the top-ranked environmental covariates, divide the value range of each environmental covariate into multiple value range intervals, perform a union operation based on the value range intervals to obtain an intersection region, and select typical soil samples by assigning a relatively high sampling weight to the intersection region.
[0038] For each soil type, the present invention further selects the top-ranked environmental covariates (specifically, the top 5 most important environmental covariates), and divides the value ranges of these characteristic variables into multiple intervals (specifically, 3 intervals can be divided). By performing a "union" operation on these intervals, the sampleable areas of the soil samples are determined. Then, according to the area size of the intersection areas, larger sampling weights are assigned to these areas, and typical soil samples are preferentially selected. These typical samples are highly representative spatially and can better reflect the characteristics of the soil type.
[0039] Through the weighted sampling strategy, the probability of selecting soil samples with higher weights will increase, thereby enhancing the representativeness of the sample data. Finally, the number of soil samples obtained through this strategy is 20 times the number of target soil types.
[0040] S6: Generate a regression matrix and perform spatial prediction: Overlay the typical soil samples with the corresponding environmental covariates, reconstruct the regression matrix, and after performing step S3, obtain a new sub-regression matrix. Use the new sub-regression matrix to train a random forest classification model to establish a spatial prediction model of the soil type. Input the important environmental covariates, and through the spatial prediction model, obtain the soil type distribution map and its uncertainty distribution map of the large-scale area, so as to achieve accurate prediction of the soil type on a large scale.
[0041] After the selection of typical samples is completed, these samples are combined with the environmental covariate data again to construct a new regression matrix. According to the above-mentioned step S3 redundancy feature elimination step, an optimized regression matrix is obtained, and this matrix is used to train a random forest classification model. Through this model, spatial prediction of the soil type in the large-scale area can be achieved.
[0042] This process generates a spatial distribution map of the soil type by inputting the most important environmental variables. These prediction results not only include the distribution of the soil type, but also can combine uncertainty analysis to output the confidence interval or uncertainty distribution map of the soil type distribution. Through the uncertainty distribution map, typical sampling is performed again in places with high uncertainty, and spatial prediction is performed again. Thus, a prediction result with lower uncertainty is achieved. This method can provide the decision-maker with the credibility of the soil type prediction result, thereby providing accurate data support for further soil mapping.
[0043] In order to further improve the accuracy of soil type prediction results, the present invention combines an uncertainty analysis method to further verify the soil type prediction results. By analyzing the uncertainty of the prediction results, confidence information for the prediction results of each region can be obtained, which is of great significance for applications such as soil mapping, agricultural planning, and environmental protection. At the same time, through uncertainty analysis, typical sampling can be recycled to further optimize sample selection and model training, ultimately achieving high accuracy and high reliability in soil type prediction. Through these steps, the present invention provides a systematic and scientific method for sampling typical soil points and spatially predicting soil types, which can be widely applied to the field of large-scale soil type prediction and soil resource management.
[0044] Figure 2 The following shows the operation process of the soil typical point sampling method based on machine learning provided by the present invention in large-scale soil type mapping. First, obtain the actual sample training set samples (including data such as virtual points of typical sampling, second soil survey points, and newly added points) and the graphic sample training set samples (including environmental covariate data such as elevation, slope, climate, and land use). Subsequently, perform autocorrelation analysis on the regression matrix composed of soil type and environmental covariates to screen out important environmental covariates. Based on the random forest classification model, use the optimized regression matrix as the training model, and use the trained random forest for spatial prediction to generate a raster spatial distribution map and an uncertainty map of soil types. According to the uncertainty of the prediction results, carry out work such as indoor verification of points, extraction of changed areas, field verification, and correction of second soil survey patches, so as to obtain an updated soil mapping result. Through uncertainty cycle optimization, finally generate a high-precision soil thematic map with low uncertainty.
[0045] Obviously, those skilled in the art can make various changes and modifications to the present technical solution without departing from the spirit and scope of the present technical solution. Thus, if these modifications and variations of the present technical solution fall within the scope of the claims of the present technical solution and its equivalent technologies, the present technical solution is also intended to include these changes and modifications.
Claims
1. A method for sampling typical soil points based on machine learning, characterized in that, It includes the following steps: S1: Resample the spatial resolution of the selected environmental covariate data to make it consistent with the spatial resolution of the second national soil census soil map; S2: Conduct preliminary random sampling on the target area to form a preliminary soil sample dataset, and combine the soil type of each soil sample with the corresponding environmental covariates to construct a regression matrix; S3: Through autocorrelation analysis of the regression matrix, identify and remove the highly correlated features among the environmental covariates, and use the recursive feature elimination algorithm to screen out the features that have the greatest impact on the prediction of the soil type, obtaining an optimized sub-regression matrix; S4: Use the optimized sub-regression matrix to train a random forest classification model to predict the soil species type. After model training, extract the feature importance scores, sort them according to the importance scores, and identify the environmental covariates with the highest rankings; S5: For each soil species type, cyclically select the environmental covariates with the highest rankings, divide the value range of each environmental covariate into multiple value range intervals, perform a union operation based on the value range intervals to obtain an intersection area, and select typical soil samples by assigning a high sampling weight to the intersection area; S6: Overlay the typical soil samples with the corresponding environmental covariates, reconstruct the regression matrix, and after performing step S3, obtain a new sub-regression matrix. Use the new sub-regression matrix to train a random forest classification model to establish a spatial prediction model of the soil type. Input the important environmental covariates, and obtain the soil type distribution map and its uncertainty distribution map of the large-scale area through the spatial prediction model, thereby realizing the accurate prediction of the soil type in a large range.
2. The method for sampling typical soil points based on machine learning according to claim 1, wherein The data sources of the environmental covariates include climate data, terrain data, land cover data, and geographic information system data.
3. The soil typical sampling point sampling method based on machine learning according to claim 1, characterized in that The resampling of the spatial resolution of the selected environmental covariate data is achieved by using nearest neighbor interpolation, bilinear interpolation, or Kriging interpolation methods.
4. The soil typical sampling point sampling method based on machine learning according to claim 1, wherein During the preliminary random sampling of the target area, select a sample size that is 20 times the number of soil species types and limit the minimum sampling number to form a preliminary soil sample dataset.
5. The method for sampling typical soil points based on machine learning according to claim 1, wherein During the preliminary random sampling process, select soil samples by the random sampling method and ensure that the number of samples of each soil type is consistent with the regional distribution of the soil type.
6. The method for sampling typical soil points based on machine learning according to claim 1, wherein In the regression matrix constructed in step S2, the independent variable is the environmental covariate, and the dependent variable is the soil species type.
7. The soil typical sampling point sampling method based on machine learning according to claim 1, wherein The identification and removal of the highly correlated features among the environmental covariates through autocorrelation analysis of the regression matrix specifically include: Screen out the covariates with an absolute value of correlation greater than 0.85 between two environmental covariates, and delete the environmental covariate with a higher correlation with other environmental covariates. Cyclically check all environmental covariates in turn until the absolute value of the correlation among all environmental covariates is not greater than 0.
85.
8. The soil typical sampling point sampling method based on machine learning as described in claim 1, wherein The use of the recursive feature elimination algorithm to screen out the features that have the greatest impact on the prediction of the soil type and obtain an optimized sub-regression matrix specifically includes: The regression matrix is iterated multiple times using the recursive feature elimination algorithm, and combined with model performance evaluation metrics to optimize the feature set, gradually eliminating irrelevant or less contributing feature variables to the model. Finally, the optimal feature subset that makes the greatest contribution to the soil type prediction is selected through cross-validation, thereby obtaining the optimized sub-regression matrix.
9. The method for sampling typical soil points based on machine learning according to claim 1, wherein The random forest classification model is based on multiple decision trees and predicts the soil type through the method of ensemble learning. It can handle high-dimensional environmental covariate data and obtain the final prediction result through a voting mechanism. The training process of the random forest classification model adopts the multiple cross-validation technique.
10. The soil typical sampling point sampling method based on machine learning according to claim 1, characterized in that In step S4, specifically identify the top 5 most important environmental covariates, and in step S5, specifically divide the value range of each environmental covariate into 3 value range intervals.