A soil thickness type prediction method based on feature ensemble learning

By constructing a soil thickness ensemble prediction model based on feature ensemble learning, the problems of high uncertainty and low prediction accuracy in soil thickness spatial prediction are solved, and higher prediction accuracy and robustness are achieved, reducing investigation costs.

CN114970934BActive Publication Date: 2025-05-13LIANGSHAN BRANCH OF SICHUAN TOBACCO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210185325.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-05-13
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

The prior art has problems of high uncertainty and low prediction accuracy in the spatial prediction of soil thickness, especially in mountainous areas and remote areas, which are difficult to obtain sufficient soil thickness observation data.

Method used

Using a method based on feature ensemble learning, a continuous soil thickness ensemble prediction model and a soil depth interval ensemble prediction model are constructed, and integrated prediction is combined with environmental variables to obtain the spatial distribution of soil thickness intervals.

Benefits of technology

It improves the accuracy and robustness of soil thickness prediction, reduces the number of soil thickness observation points, saves field survey costs, and improves the accuracy of calculation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114970934B_ABST
    Figure CN114970934B_ABST
Patent Text Reader

Abstract

The present invention relates to a soil thickness type prediction method based on feature ensemble learning, which effectively utilizes environmental variables that affect soil thickness changes to construct a prediction model, and constructs a continuous soil thickness integrated prediction model and a soil depth interval integrated prediction model by screening an optimal set of environmental variables, and finally obtains the spatial distribution of soil thickness intervals covering a target area. The method has higher prediction accuracy than traditional spatial interpolation technology, and in the future, when there is a lack of sufficient soil thickness observation data in mountainous areas or remote areas, the present invention can reduce the number of soil thickness observation points required, thereby saving field survey costs while ensuring the prediction accuracy of soil thickness spatial distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a soil thickness type prediction method based on feature ensemble learning, and belongs to the technical field of soil thickness in soil hydrology and quantitative soil science. Background Art

[0002] Soil thickness is a physical property of soil, usually used to indicate the depth of soil material from the surface to a specific interface. Therefore, soil thickness can effectively indicate the storage depth of nutrients or available water in soil material, and is of great significance for accurate simulation of soil erosion, crop growth, biodiversity, soil carbon storage estimation and soil water turnover. Therefore, obtaining a high-precision spatial distribution map of regional soil thickness plays an important practical guiding role in ecosystem assessment, agricultural production, soil and water conservation, vegetation restoration, and desertification control.

[0003] Researchers with different disciplinary backgrounds and technicians from different application departments have different definitions of soil thickness. There are different standards for the classification of soil thickness. Some scholars define soil thickness as the vertical depth from the soil surface to the soil parent material layer; some scholars define soil thickness as the maximum depth that soil nutrients or plant roots can reach; technicians from relevant departments also define soil thickness as the thickness of the tillage layer. Regardless of the different definitions of soil thickness in specific applications, the prediction of the spatial distribution of soil thickness is often studied and applied as the same type of technology.

[0004] Soil is a material that is continuous in both time and space, and has complex spatial heterogeneity. Even at the same time, soil thickness shows complex variation characteristics at the field scale, watershed scale or national scale, and it is difficult to use a constant to represent the change in soil thickness in the region. Unlike conventional surface vegetation coverage surveys, soil thickness is difficult to observe directly. Traditional soil surveys use profile excavation, drilling or natural exposed bedrock surveys to record the thickness of soil at specific locations in the field. Soil thickness surveys in large areas, especially in mountainous areas where road accessibility is low, are difficult to conduct soil surveys, require a lot of manpower, material resources, time and funds, are inefficient, and can only obtain soil thickness sample data at discrete locations. Therefore, traditional soil surveys often use the average soil thickness value observed at one or more sample points to represent the soil thickness of the area.

[0005] Since soil materials and rocks have different densities, conductivity and magnetism, some technicians try to use geophysical exploration technology to detect soil thickness. The more commonly used technologies include high-density resistivity method, geological radar method, seismic exploration method, magnetotelluric, gamma ray detection method, etc. Geophysical exploration technology has the advantage of non-destructive detection. It does not need to destroy the original physical structure of the soil. The soil physical information is collected quickly, efficiently and accurately. For example, relevant studies have shown that the soil thickness inverted using the EM38 electromagnetic induction instrument has high accuracy, and its determination coefficient can reach 0.7-0.8. However, since geophysical exploration technology cannot obtain the true value of soil thickness, it often needs to be combined with traditional soil surveys such as field drilling to calibrate and verify the geophysical exploration model based on the soil thickness values ​​observed in the field. Due to the technical characteristics of different geophysical exploration equipment, different technical methods often have specific operating environments or application scope requirements. For example, in areas with high soil moisture content and shallow groundwater levels, the amplitude of ground penetrating radar is easily affected by changes in moisture content. In addition, geophysical exploration is often conducted in the form of field survey lines, which can only obtain soil thickness distribution values ​​in the area covered by the survey line.

[0006] In order to obtain a spatial distribution map of soil thickness covering a larger area, technicians often use geographic information system technology to perform spatial interpolation on discrete soil thickness sample data. This method assumes that the spatial distribution of soil thickness has a certain regularity, that is, to quantify the spatial variation characteristics of soil thickness. Commonly used geostatistical methods include ordinary kriging, simple kriging, pan-kriging and co-kriging. However, the case results conducted by researchers show that the prediction accuracy of spatial interpolation methods varies in different regions, and no unified conclusion has been reached on the optimal prediction method.

[0007] In recent years, with the rapid development of digital soil mapping, domestic and foreign scholars and technicians prefer to use the "soil landscape model" to predict soil thickness. This model assumes that soil properties are affected by soil-forming factors such as topography, parent material, vegetation, and climate. The prediction model constructed using these environmental variables as covariates can more accurately obtain the spatial distribution map of soil thickness. More mature spatial prediction technologies include support vector machines, random forests, geographically weighted regression, deep learning, fuzzy C-means clustering, etc. Related studies have also shown that the overall prediction accuracy of this prediction model is significantly higher than that of geostatistical or statistical models.

[0008] Due to the high spatial variation of soil thickness and the combined influence of environmental factors such as topography and climate, the spatial prediction accuracy of soil thickness is low, and different prediction technologies are difficult to be directly applied to the prediction of soil thickness in the target area. In summary, the existing technologies mainly have the following technical problems:

[0009] (1) Whether it is traditional soil profile excavation or geological drilling, field soil surveys can only obtain information about the depth of the surveyed soil. In some areas (such as the Loess Plateau), the soil thickness is as high as hundreds of meters. If the survey depth is less than the actual soil thickness, the staff often only record the actual observation depth, such as marking the soil thickness as ">2m" or "2m". In the actual spatial prediction process, using such observation data that is lower than the actual soil thickness to build a prediction model can easily underestimate the soil thickness in some areas.

[0010] (2) Traditional soil survey methods are inefficient and cannot obtain a large amount of observation data. Although geophysical exploration technology has a high operating efficiency, it can only obtain the numerical value of soil thickness under the survey line sequence. The amount of "point" and "line" soil thickness information data obtained by these two operating methods is very limited, and it is difficult to operate in remote areas and mountainous areas.

[0011] (3) Due to the complex topography in mountainous areas, the spatial distribution heterogeneity of soil thickness is very high. Since the physical formation mechanism of soil thickness is very complex, involving soil redistribution in the soil weathering rate and soil erosion process, the correlation between soil thickness and environmental variables (covariates) is often low, which directly leads to the low prediction accuracy of geostatistics and digital soil mapping technology based on soil landscape models. Some easily accessible environmental variables, such as remote sensing factors, land use, and vegetation coverage, are difficult to effectively characterize the spatial variation characteristics of soil thickness.

[0012] (4) Currently available spatial prediction technologies have significant deficiencies in robustness. Different spatial prediction technologies are often based on specific model assumptions. Although there are significant differences in the overall accuracy of the prediction models, different technologies may achieve different prediction accuracies in different local areas of the operation area. How to accurately identify and effectively integrate these accurate prediction results is a significant shortcoming of existing technologies.

[0013] In summary, the technical deficiencies in the above analysis also appear in the spatial prediction of some soil physical and chemical properties. Summary of the invention

[0014] The technical problem to be solved by the present invention is to provide a soil thickness type prediction method based on feature ensemble learning, which covers two key technical links: continuous soil thickness ensemble learning and different types of soil thickness data ensemble learning, and can effectively improve the problems of high uncertainty and low prediction accuracy in existing soil thickness prediction.

[0015] In order to solve the above technical problems, the present invention adopts the following technical solutions: the present invention designs a soil thickness type prediction method based on feature ensemble learning, and according to the following steps A to F, obtains the soil thickness ensemble prediction model corresponding to the target area and the corresponding accuracy, and obtains the soil depth interval ensemble prediction model corresponding to the target area and the corresponding accuracy; then according to steps i to iii, obtains the soil thickness interval spatial distribution corresponding to the target area;

[0016] Step A. Based on the preset sampling points corresponding to different land types in the target area, soil thickness values ​​at the sampling points are used to form soil thickness data feature vectors at the sampling points, and then the soil thickness data feature vectors at the sampling points are used to form a continuous soil thickness data set ConDep;

[0017] At the same time, based on the preset ordering of each soil thickness threshold from small to large or from large to small, and each soil depth interval divided, the soil depth interval corresponding to the soil thickness value of each sample point position is obtained to form a soil depth interval feature vector of each sample point position, and then the soil depth interval feature vector of each sample point position is used to form a discrete soil depth interval data set DisDep; then enter step B;

[0018] Step B. Obtain the data of preset environmental variables covering the target area, and adopt a resampling method to obtain the data of each environmental variable corresponding to the resolution Res grid division of the target area at each sample point location distribution, and then enter step C;

[0019] Step C. Based on the data of each environmental variable corresponding to each grid in the target area, that is, obtaining the data of each environmental variable corresponding to each sample point position, according to the correlation between each environmental variable and the soil thickness value, determining each optimal environmental variable related to the soil thickness value in each environmental variable, forming a target environmental variable group, and then entering step D;

[0020] Step D. Based on the data of each environmental variable corresponding to each grid in the target area, the data of each optimal environmental variable in the target environmental variable group corresponding to each sample point position are added to the soil thickness data feature vector corresponding to the sample point position for updating, thereby updating the continuous soil thickness data set ConDep;

[0021] At the same time, the data of each optimal environmental variable in the target environmental variable group corresponding to each sampling point position are added to the soil depth interval feature vector of the corresponding sampling point position for updating, thereby updating the discrete soil depth interval data set DisDep, and then entering step E;

[0022] Step E. For each type of pre-set model to be trained, based on the continuous soil thickness data set ConDep, the data of each optimal environmental variable in the target environmental variable group corresponding to the sample point position is used as input, and the soil thickness value corresponding to the sample point position is used as output, and the model to be trained is trained to obtain a continuous soil thickness prediction model PreTch_i(dep), and obtain the determination coefficient R2_i of the soil thickness prediction model; wherein 1≤i≤I, I represents the number of each type of model to be trained, PreTch_i(dep) represents the i-th continuous soil thickness prediction model, and R2_i represents the determination coefficient of the i-th soil thickness prediction model;

[0023] At the same time, for each type of preset model to be trained, based on the discrete soil depth interval data set DisDep, the data of each optimal environmental variable in the target environmental variable group corresponding to the sample point position is used as input, and the soil depth interval corresponding to the sample point position is used as output. The model to be trained is trained to obtain the soil depth interval prediction model ClaTch_i(dep), and the accuracy Accu_i of the soil depth interval prediction model is obtained; ClaTch_i(dep) represents the i-th soil depth interval prediction model, and Accu_i represents the accuracy of the i-th soil depth interval prediction model;

[0024] Then proceed to step F;

[0025] Step F. Based on each continuous soil thickness prediction model PreTch_i(dep) and the corresponding determination coefficient R2_i, a continuous soil thickness integrated prediction model is constructed as follows:

[0026]

[0027] Among them, f con (dep) represents the soil thickness value, and according to the continuous soil thickness dataset ConDep, the accuracy Con_R2 of the soil thickness integrated prediction model is obtained;

[0028] At the same time, based on each soil depth interval prediction model ClaTch_i(dep) and the corresponding accuracy Accu_i, the soil depth interval integrated prediction model is constructed as follows:

[0029]

[0030] Among them, f dis (dep) represents the soil depth interval, and according to the discrete soil depth interval dataset DisDep, the accuracy DisAccu of the soil depth interval integrated prediction model is obtained;

[0031] Step i. Obtain the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, and then proceed to step ii;

[0032] Step ii. According to the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, the soil depth interval integrated prediction model is applied to obtain the first spatial distribution Map_Dis1 of the soil thickness interval covering the target area;

[0033] At the same time, according to the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, the continuous soil thickness integrated prediction model is applied to obtain the numerical spatial distribution of soil thickness covering the target area; and combined with each soil depth interval divided based on each soil thickness threshold in step A, the second spatial distribution Map_Dis2 of the soil thickness interval corresponding to the numerical spatial distribution of soil thickness covering the target area is obtained; then proceed to step iii;

[0034] Step iii. According to the following formula:

[0035]

[0036] Obtain the spatial distribution of soil thickness intervals covering the target area f cd (dep).

[0037] As a preferred technical solution of the present invention, step A includes steps A1 to A3 as follows:

[0038] Step A1. For each preset sample point location corresponding to different land types in the target area, obtain the soil thickness value of the sample point location, form the soil thickness data feature vector of the sample point location, and then form a continuous soil thickness data set ConDep = {cp_1, ..., cp_m, ..., cp_M} from the soil thickness data feature vectors of each sample point location, and then enter step A2; wherein 1≤m≤M, M represents the number of sample point locations, and cp_m represents the soil thickness data feature vector of the mth sample point location;

[0039] Step A2. Process the continuous soil thickness dataset ConDep so that the continuous soil thickness dataset ConDep conforms to a normal distribution, and then proceed to step A3;

[0040] Step A3. Based on the preset soil thickness thresholds sorted in order from small to large or from large to small, and the divided soil depth intervals, the soil depth interval corresponding to the soil thickness value of each sample point position is obtained to form a soil depth interval feature vector of each sample point position, and then the soil depth interval feature vector of each sample point position is used to form a discrete soil depth interval data set DisDep = {dp_1, ..., dp_m, ..., dp_M}, and then enter step D; wherein dp_m represents the soil depth interval feature vector of the mth sample point position.

[0041] As a preferred technical solution of the present invention: in the step A1, for each preset sample point location corresponding to different land types in the target area, the soil thickness value, coordinate longitude information, coordinate latitude information, and land use type of the sample point location are obtained to form a soil thickness data feature vector of the sample point location;

[0042] In step A3, based on the preset soil thickness thresholds and the soil depth intervals divided in order from small to large or from large to small, the soil depth interval corresponding to the soil thickness value of each sample point position is obtained, and the coordinate longitude information, coordinate latitude information, and land use type of each sample point position are combined to form a soil depth interval feature vector of each sample point position.

[0043] As a preferred technical solution of the present invention: in the step A2, a natural logarithm function is applied to process the continuous soil thickness dataset ConDep so that the continuous soil thickness dataset ConDep conforms to a normal distribution.

[0044] As a preferred technical solution of the present invention: the step B includes the following steps B1 to B2;

[0045] Step B1. Obtain data of preset environmental variables covering the target area, and apply the Z-Score standardization method to each environmental variable, perform data standardization processing, update the data of each environmental variable, and then enter step B2;

[0046] Step B2. Unify the geographic coordinate system and data format of the data of each environmental variable, update the data of each environmental variable, and use the resampling method to obtain the data of the target area under the resolution Res grid division, where each grid corresponds to each environmental variable, and then enter step C.

[0047] As a preferred technical solution of the present invention: in the step i, the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area is obtained, and the geographic coordinate system and data format of the data of each environmental variable are unified according to the operation of step B1 and step B2, the data distribution of each optimal environmental variable is updated, and then step ii is entered.

[0048] As a preferred technical solution of the present invention: the preset environmental variables include terrain factors, remote sensing factors, climate variables, biological factors, and geological factors;

[0049] Among them, terrain factors: elevation, slope, slope aspect, terrain moisture index, plane curvature, slope surface curvature, slope position, slope shape, slope length, terrain undulation, surface roughness, and surface cutting depth;

[0050] Remote sensing factors: each band of remote sensing images, leaf area index, ratio vegetation index, difference environmental vegetation index, green vegetation index, vertical vegetation index, and normalized difference vegetation index;

[0051] Climate variables: average annual rainfall, average annual temperature, average sunshine hours, average wind speed;

[0052] Biological factors: vegetation type, land use;

[0053] Geological factors: parent material, hydrogeological map.

[0054] As a preferred technical solution of the present invention: the resampling method is any one of a bilinear interpolation method, a nearest neighbor assignment method, a cubic convolution interpolation method, and a majority method.

[0055] As a preferred technical solution of the present invention: the various types of models to be trained preset in step E include geographically weighted regression, random forest, support vector machine, deep learning, decision tree, k-nearest neighbor algorithm, Bayesian classification, and classification tree.

[0056] As a preferred technical solution of the present invention: the calculation of the accuracy of each soil thickness prediction model and the calculation of the accuracy of each soil depth interval prediction model in step E, and the calculation of the accuracy of the soil thickness integrated prediction model and the calculation of the accuracy of the soil depth interval integrated prediction model in step F, the accuracy calculation method adopted is any one of ten-fold cross validation, five-fold cross validation, three-fold cross validation, and leave-one-out validation, and any one of the root mean square error, performance deviation ratio, sum of squares of error, and correction determination coefficient is adopted as the accuracy standard.

[0057] The soil thickness type prediction method based on feature ensemble learning described in the present invention adopts the above technical solution and has the following technical effects compared with the prior art:

[0058] (1) The soil thickness type prediction method based on feature ensemble learning designed by the present invention effectively utilizes environmental variables that affect soil thickness changes to construct a prediction model. By screening the optimal set of environmental variables, a continuous soil thickness integrated prediction model and a soil depth interval integrated prediction model are constructed, and finally the spatial distribution of soil thickness intervals covering the target area is obtained. Compared with the traditional spatial interpolation technology, it has higher prediction accuracy. In the future, when there is a lack of sufficient soil thickness observation data in mountainous areas or remote areas, the present invention can reduce the number of soil thickness observation points required, thereby ensuring the accuracy of soil thickness spatial distribution prediction and saving field survey costs;

[0059] (2) The soil thickness type prediction method based on feature ensemble learning designed by the present invention is different from the traditional continuous soil thickness prediction. It proposes a soil thickness type prediction oriented to specific business needs, that is, discrete soil thickness data prediction, which can fully tap the prediction advantages of the classification model and avoid the shortcomings of the traditional prediction technology that only focuses on the generalization ability of the continuous prediction model. It can maximize the prediction ability of the classification model in soil thickness.

[0060] (3) The soil thickness type prediction method based on feature ensemble learning designed by the present invention can effectively integrate weak learners, is very flexible, and can effectively avoid the overfitting problem of each sub-model. The constructed multi-type ensemble learning model has a low generalization error rate and high precision. Users do not need to adjust too many model parameters during actual use, which maximizes the accuracy of the calculation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is the main flow chart of the soil thickness type prediction method based on feature ensemble learning designed by the present invention;

[0062] Figure 2 is the spatial distribution of soil thickness sampling points of the implementation case;

[0063] Figure 3 It is the spatial distribution map of the environmental variable elevation in the implementation case area;

[0064] Figure 4 is the spatial distribution map of the slope of environmental variables in the implementation case area;

[0065] Figure 5 It is the spatial distribution map of the topographic humidity index of the environmental variable in the implementation case area;

[0066] Figure 6 It is the spatial distribution map of the average annual rainfall of the environmental variable in the implementation case area;

[0067] Figure 7It is the spatial distribution map of the annual average temperature of the environmental variable in the implementation case area;

[0068] Figure 8 It is the spatial distribution map of the normalized difference vegetation index of environmental variables in the implementation case area;

[0069] Fig. 9 It is the spatial distribution map of soil thickness intervals predicted in the implementation case area. DETAILED DESCRIPTION

[0070] The specific implementation modes of the present invention will be further described in detail below in conjunction with the accompanying drawings.

[0071] The present invention designs a soil thickness type prediction method based on feature ensemble learning. In practical applications, Figure 1 As shown, according to the following steps A to F, the soil thickness integrated prediction model corresponding to the target area and the corresponding accuracy are obtained, and the soil depth interval integrated prediction model corresponding to the target area and the corresponding accuracy are obtained.

[0072] Step A. Based on the preset sample point locations corresponding to different land types in the target area, the soil thickness values ​​at each sample point location are used to form a soil thickness data feature vector at each sample point location, and then the soil thickness at each sample point location and the data feature vector are used to form a continuous soil thickness data set ConDep.

[0073] At the same time, based on the preset soil thickness thresholds and the soil depth intervals that are sorted in order from small to large or from large to small, the soil depth intervals corresponding to the soil thickness values ​​at each sample point are obtained to form the soil depth interval feature vectors at each sample point, and then the discrete soil depth interval data set DisDep is formed by the soil depth interval feature vectors at each sample point; then the process proceeds to step B. In practical applications, the soil depth intervals divided here are, for example, four intervals of 0-20cm, 20-50cm, 50-100cm and >100cm, where the soil thickness thresholds are 20cm, 50cm and 100cm.

[0074] In actual application, the above step A specifically executes the following steps A1 to A3.

[0075] Step A1. For each preset sample point location corresponding to different land types in the target area, obtain the soil thickness value, coordinate longitude information, coordinate latitude information, and land use type of the sample point location to form a soil thickness data feature vector of the sample point location, and then form a continuous soil thickness data set ConDep = {cp_1, ..., cp_m, ..., cp_M} from the soil thickness data feature vectors of each sample point location, and then enter step A2; wherein 1≤m≤M, M represents the number of sample point locations, and cp_m represents the soil thickness data feature vector of the mth sample point location.

[0076] In practical applications, for the soil thickness data feature vector of each sample point location, if the land use type is water body, the soil thickness value corresponding to the water body is 0. In practical applications, for sample point locations that do not have soil thickness value, coordinate longitude information, coordinate latitude information, and land use type information, they will be deleted and will not participate in further analysis.

[0077] Step A2. Apply the natural logarithm function to process the continuous soil thickness dataset ConDep so that the continuous soil thickness dataset ConDep conforms to the normal distribution, and then proceed to step A3.

[0078] Step A3. Based on the preset soil thickness thresholds sorted in order from small to large or from large to small, and the divided soil depth intervals, the soil depth interval corresponding to the soil thickness value of each sample point position is obtained, and the coordinate longitude information, coordinate latitude information, and land use type of each sample point position are combined to form a soil depth interval feature vector of each sample point position, and then the soil depth interval feature vector of each sample point position is used to form a discrete soil depth interval data set DisDep = {dp_1, ..., dp_m, ..., dp_M}, and then enter step D; wherein dp_m represents the soil depth interval feature vector of the mth sample point position.

[0079] Step B. Obtain the data of preset environmental variables covering the target area, and use the resampling method to obtain the data of each environmental variable corresponding to the resolution Res grid division of the target area corresponding to the distribution of each sample point position, and then enter step C.

[0080] In practical application, the above step B specifically performs the following steps B1 to B2.

[0081] Step B1. Obtain the data of each preset environmental variable covering the target area, and apply the Z-Score standardization method to each environmental variable, perform data standardization processing, update the data of each environmental variable, so that its data conforms to the standard normal distribution, that is, the mean is 0 and the standard deviation is 1, and then enter step B2.

[0082] Step B2. Unify the geographic coordinate system and data format of the data of each environmental variable, update the data of each environmental variable, and use any resampling method including bilinear interpolation method, nearest neighbor assignment method, cubic convolution interpolation method, and majority method to obtain the data of the target area under the resolution Res grid division, where each grid corresponds to each environmental variable, and then enter step C.

[0083] In actual application, based on the geographic environment data collection and shared data download of the business department, environmental variable data of different resolutions that affect soil thickness in the coverage study area are obtained. The environmental variables include terrain factors, remote sensing factors, climate variables, biological factors, and geological factors.

[0084] Among them, terrain factors include: elevation, slope, slope aspect, terrain moisture index, plane curvature, slope surface curvature, slope position, slope shape, slope length, terrain undulation, surface roughness, and surface cutting depth.

[0085] Remote sensing factors: various bands of remote sensing images, leaf area index, ratio vegetation index, difference environmental vegetation index, greenness vegetation index, vertical vegetation index, and normalized difference vegetation index.

[0086] Climate variables: average annual rainfall, average annual temperature, average sunshine hours, average wind speed.

[0087] Biological factors: vegetation type, land use.

[0088] Geological factors: parent material, hydrogeological map.

[0089] The format of raster data can be Esri Grid, GIF, IMG, JPEG, TIFF, MRF, CRF, etc.

[0090] Step C. Based on the data of each environmental variable corresponding to each grid in the target area, that is, obtaining the data of each environmental variable corresponding to each sample point position, according to the correlation between each environmental variable and the soil thickness value, determine the optimal environmental variables related to the soil thickness value in each environmental variable, form a target environmental variable group, and then enter step D.

[0091] Step D. Based on the data of each environmental variable corresponding to each grid in the target area, the data of each optimal environmental variable in the target environmental variable group corresponding to each sample point position are added to the soil thickness data feature vector of the corresponding sample point position for updating, thereby updating the continuous soil thickness dataset ConDep.

[0092] At the same time, the data of each optimal environmental variable in the target environmental variable group corresponding to each sample point position are added to the soil depth interval feature vector of the corresponding sample point position for updating, thereby updating the discrete soil depth interval data set DisDep, and then entering step E.

[0093] Step E. For each type of pre-set model to be trained, based on the continuous soil thickness data set ConDep, the data of each optimal environmental variable in the target environmental variable group corresponding to the sample point position is used as input, and the soil thickness value corresponding to the sample point position is used as output. The model to be trained is trained to obtain a continuous soil thickness prediction model PreTch_i(dep), and the determination coefficient R2_i of the soil thickness prediction model is obtained; wherein 1≤i≤I, I represents the number of each type of model to be trained, PreTch_i(dep) represents the i-th continuous soil thickness prediction model, and R2_i represents the determination coefficient of the i-th soil thickness prediction model.

[0094] In actual applications, the preset types of models to be trained include geographically weighted regression, random forest, support vector machine, deep learning, decision tree, k-nearest neighbor algorithm, Bayesian classification, and classification tree.

[0095] At the same time, for each type of preset model to be trained, based on the discrete soil depth interval data set DisDep, the data of each optimal environmental variable in the target environmental variable group corresponding to the sample point position is used as input, and the soil depth interval corresponding to the sample point position is used as output. The model to be trained is trained to obtain the soil depth interval prediction model ClaTch_i(dep), and the accuracy Accu_i of the soil depth interval prediction model is obtained; ClaTch_i(dep) represents the i-th soil depth interval prediction model, and Accu_i represents the accuracy of the i-th soil depth interval prediction model; then enter step F.

[0096] Step F. Based on each continuous soil thickness prediction model PreTch_i(dep) and the corresponding determination coefficient R2_i, a continuous soil thickness integrated prediction model is constructed as follows:

[0097]

[0098] Among them, f con (dep) represents the soil thickness value, and based on the continuous soil thickness dataset ConDep, the accuracy Con_R2 of the soil thickness integrated prediction model is obtained.

[0099] At the same time, based on each soil depth interval prediction model ClaTch_i(dep) and the corresponding accuracy Accu_i, the soil depth interval integrated prediction model is constructed as follows:

[0100]

[0101] Among them, f dis (dep) represents the soil depth interval, and based on the discrete soil depth interval dataset DisDep, the accuracy DisAccu of the soil depth interval integrated prediction model is obtained.

[0102] In the actual application of the above design, the calculation of the accuracy of each soil thickness prediction model and the calculation of the accuracy of each soil depth interval prediction model in step E, as well as the calculation of the accuracy of the soil thickness integrated prediction model and the calculation of the accuracy of the soil depth interval integrated prediction model in step F, the accuracy calculation method adopted is any one of ten-fold cross validation, five-fold cross validation, three-fold cross validation, and leave-one-out validation, and any one of the root mean square error, performance deviation ratio, sum of squares of error, and correction determination coefficient is used as the accuracy standard.

[0103] Based on the above, the soil thickness integrated prediction model and the corresponding accuracy corresponding to the target area are obtained, and the soil depth interval integrated prediction model and the corresponding accuracy corresponding to the target area are obtained; further according to steps i to iii, the spatial distribution of soil thickness intervals corresponding to the target area is obtained.

[0104] Step i. Obtain the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, and unify the geographic coordinate system and data format of the data of each environmental variable according to the operation of step B1 and step B2, update the data distribution of each optimal environmental variable, and then enter step ii.

[0105] Step ii. According to the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, the soil depth interval integrated prediction model is applied to obtain the first spatial distribution Map_Dis1 of the soil thickness interval covering the target area.

[0106] At the same time, according to the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, the continuous soil thickness integrated prediction model is applied to obtain the numerical spatial distribution of soil thickness covering the target area; and combined with the various soil depth intervals divided based on each soil thickness threshold in step A, the second spatial distribution Map_Dis2 of the soil thickness interval corresponding to the numerical spatial distribution of soil thickness covering the target area is obtained; then enter step iii.

[0107] Step iii. According to the following formula:

[0108]

[0109] Obtain the spatial distribution of soil thickness intervals covering the target area f cd(dep).

[0110] The soil thickness type prediction method based on feature ensemble learning designed by the present invention is applied in practice. The soil thickness prediction in the southern part of Sichuan Province is taken as an example. In this operating area, the soil thickness is defined as the depth from the surface to the weakly weathered layer or fresh bedrock. The soil thickness prediction in Liangshan Prefecture and Panzhihua area in southern Sichuan Province is taken as an example.

[0111] Specifically, in the design and implementation of the present invention, the spatial distribution of the positions of the sample points is as follows: Figure 2 As shown, in step A, based on the soil thickness thresholds of 60cm and 100cm, the soil thickness is divided into three intervals, that is, the first soil thickness interval is a 0cm-60cm depth interval, the second soil thickness interval is a 60cm-100cm depth interval, and the third soil thickness interval is a depth interval greater than 100cm, and then the soil depth interval corresponding to the soil thickness value of each sample point position is obtained to form a soil depth interval feature vector of each sample point position, and then the soil depth interval feature vector of each sample point position is used to form a discrete soil depth interval data set DisDep.

[0112] Execute step B to obtain data of preset environmental variables covering the target area, including elevation, slope, aspect, terrain moisture index, surface roughness, normalized vegetation index, difference environmental vegetation index, land use, average annual rainfall, average annual temperature, average sunshine hours, and soil parent material, and the format of the raster data is Esri Grid or TIFF; and use the Z-score standardization method to standardize each environmental variable, and use the geographic information system software to set all environmental variable data to a unified geographic coordinate system (WGS_1984_Albers); then use the resampling method to obtain the data of the target area under the 1km grid division corresponding to the distribution of each sample point location, and each grid corresponds to each environmental variable, in TIFF format.

[0113] Based on the execution of step C, the optimal environmental variables related to the soil thickness value are determined, including elevation, slope, terrain moisture index, annual average rainfall, annual average temperature, and normalized vegetation index, respectively. Figure 3-8 As shown, a target environment variable group is formed, and steps D to E are further executed.

[0114] Continue to execute step F, build a continuous soil thickness integrated prediction model and a soil depth interval integrated prediction model, and finally in the application of the embodiment, execute steps i to iii to obtain the soil thickness interval spatial distribution f covering the target area cd (dep), such as Fig. 9 shown.

[0115] Through the actual implementation of the scheme designed by the present invention, the continuous soil thickness prediction sub-model and the discrete soil thickness prediction sub-model can be effectively integrated to maximize the robustness of soil thickness prediction. The feature-based ensemble learning method has high universality and can be applied not only to the spatial prediction of soil thickness types, but also to the calculation of the spatial distribution of similar geographic entities, such as glacier thickness, bedrock depth, etc.

[0116] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A soil thickness type prediction method based on feature ensemble learning, characterized by: According to the following steps A to F, the soil thickness integrated prediction model and the corresponding accuracy corresponding to the target area are obtained, and the soil depth interval integrated prediction model and the corresponding accuracy corresponding to the target area are obtained; then according to steps i to iii, the spatial distribution of the soil thickness interval corresponding to the target area is obtained; Step A. Based on the preset sampling points corresponding to different land types in the target area, soil thickness values ​​at the sampling points are used to form soil thickness data feature vectors at the sampling points, and then the soil thickness data feature vectors at the sampling points are used to form a continuous soil thickness data set ConDep; At the same time, based on the preset ordering of each soil thickness threshold from small to large or from large to small, and each soil depth interval divided, the soil depth interval corresponding to the soil thickness value of each sample point position is obtained to form a soil depth interval feature vector of each sample point position, and then the soil depth interval feature vector of each sample point position is used to form a discrete soil depth interval data set DisDep; then enter step B; Step B. Obtain the data of preset environmental variables covering the target area, and adopt a resampling method to obtain the data of each environmental variable corresponding to the resolution Res grid division of the target area at each sample point location distribution, and then enter step C; Step C. Based on the data of each environmental variable corresponding to each grid in the target area, that is, obtaining the data of each environmental variable corresponding to each sample point position, according to the correlation between each environmental variable and the soil thickness value, determining each optimal environmental variable related to the soil thickness value in each environmental variable, forming a target environmental variable group, and then entering step D; Step D. Based on the data of each environmental variable corresponding to each grid in the target area, the data of each optimal environmental variable in the target environmental variable group corresponding to each sample point position are added to the soil thickness data feature vector corresponding to the sample point position for updating, thereby updating the continuous soil thickness data set ConDep; At the same time, the data of each optimal environmental variable in the target environmental variable group corresponding to each sampling point position are added to the soil depth interval feature vector of the corresponding sampling point position for updating, thereby updating the discrete soil depth interval data set DisDep, and then entering step E; Step E. For each type of pre-set model to be trained, based on the continuous soil thickness data set ConDep, the data of each optimal environmental variable in the target environmental variable group corresponding to the sample point position is used as input, and the soil thickness value corresponding to the sample point position is used as output, and the model to be trained is trained to obtain a continuous soil thickness prediction model PreTch_i(dep), and obtain the determination coefficient R2_i of the soil thickness prediction model; wherein 1≤i≤I, I represents the number of each type of model to be trained, PreTch_i(dep) represents the i-th continuous soil thickness prediction model, and R2_i represents the determination coefficient of the i-th soil thickness prediction model; At the same time, for each type of preset model to be trained, based on the discrete soil depth interval data set DisDep, the data of each optimal environmental variable in the target environmental variable group corresponding to the sample point position is used as input, and the soil depth interval corresponding to the sample point position is used as output. The model to be trained is trained to obtain the soil depth interval prediction model ClaTch_i(dep), and the accuracy Accu_i of the soil depth interval prediction model is obtained; ClaTch_i(dep) represents the i-th soil depth interval prediction model, and Accu_i represents the accuracy of the i-th soil depth interval prediction model; Then proceed to step F; Step F. Based on each continuous soil thickness prediction model PreTch_i(dep) and the corresponding determination coefficient R2_i, a continuous soil thickness integrated prediction model is constructed as follows: Among them, f con (dep) represents the soil thickness value, and according to the continuous soil thickness dataset ConDep, the accuracy Con_R2 of the soil thickness integrated prediction model is obtained; At the same time, based on each soil depth interval prediction model ClaTch_i(dep) and the corresponding accuracy Accu_i, the soil depth interval integrated prediction model is constructed as follows: Among them, f dis (dep) represents the soil depth interval, and according to the discrete soil depth interval dataset DisDep, the accuracy DisAccu of the soil depth interval integrated prediction model is obtained; Step i. Obtain the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, and then proceed to step ii; Step ii. According to the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, the soil depth interval integrated prediction model is applied to obtain the first spatial distribution Map_Dis1 of the soil thickness interval covering the target area; At the same time, according to the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area, the continuous soil thickness integrated prediction model is applied to obtain the numerical spatial distribution of soil thickness covering the target area; and combined with each soil depth interval divided based on each soil thickness threshold in step A, the second spatial distribution Map_Dis2 of the soil thickness interval corresponding to the numerical spatial distribution of soil thickness covering the target area is obtained; then proceed to step iii; Step iii. According to the following formula: Obtain the spatial distribution of soil thickness intervals covering the target area f cd (dep).

2. A soil thickness type prediction method based on feature ensemble learning according to claim 1, characterized in that: The step A includes steps A1 to A3 as follows: Step A1. For each preset sample point location corresponding to different land types in the target area, obtain the soil thickness value of the sample point location, form the soil thickness data feature vector of the sample point location, and then form a continuous soil thickness data set ConDep = {cp_1, ..., cp_m, ..., cp_M} from the soil thickness data feature vectors of each sample point location, and then enter step A2; wherein 1≤m≤M, M represents the number of sample point locations, and cp_m represents the soil thickness data feature vector of the mth sample point location; Step A2. Process the continuous soil thickness dataset ConDep so that the continuous soil thickness dataset ConDep conforms to a normal distribution, and then proceed to step A3; Step A3. Based on the preset soil thickness thresholds sorted in order from small to large or from large to small, and the divided soil depth intervals, the soil depth interval corresponding to the soil thickness value of each sample point position is obtained to form a soil depth interval feature vector of each sample point position, and then the soil depth interval feature vector of each sample point position is used to form a discrete soil depth interval data set DisDep = {dp_1, ..., dp_m, ..., dp_M}, and then enter step D; wherein dp_m represents the soil depth interval feature vector of the mth sample point position.

3. A soil thickness type prediction method based on feature ensemble learning according to claim 2, characterized in that: In the step A1, for each preset sample point location corresponding to different land types in the target area, the soil thickness value, coordinate longitude information, coordinate latitude information, and land use type of the sample point location are obtained to form a soil thickness data feature vector of the sample point location; In step A3, based on the preset soil thickness thresholds and the soil depth intervals divided in order from small to large or from large to small, the soil depth interval corresponding to the soil thickness value of each sample point position is obtained, and the coordinate longitude information, coordinate latitude information, and land use type of each sample point position are combined to form a soil depth interval feature vector of each sample point position.

4. The soil thickness type prediction method based on feature ensemble learning according to claim 2 is characterized in that: In the step A2, a natural logarithm function is applied to process the continuous soil thickness dataset ConDep so that the continuous soil thickness dataset ConDep conforms to a normal distribution.

5. The soil thickness type prediction method based on feature ensemble learning according to claim 1 is characterized in that: The step B includes the following steps B1 to B2; Step B1. Obtain data of preset environmental variables covering the target area, and apply the Z-Score standardization method to each environmental variable, perform data standardization processing, update the data of each environmental variable, and then enter step B2; Step B2. Unify the geographic coordinate system and data format of the data of each environmental variable, update the data of each environmental variable, and use the resampling method to obtain the data of the target area under the resolution Res grid division, where each grid corresponds to each environmental variable, and then enter step C.

6. The soil thickness type prediction method based on feature ensemble learning according to claim 5 is characterized in that: In step i, the data distribution of each optimal environmental variable in the target environmental variable group corresponding to the target area is obtained, and the geographic coordinate system and data format of the data of each environmental variable are unified according to the operation of step B1 and step B2, and the data distribution of each optimal environmental variable is updated, and then step ii is entered.

7. A soil thickness type prediction method based on feature ensemble learning according to claim 1 or 5, characterized in that: The preset environmental variables include terrain factors, remote sensing factors, climate variables, biological factors, and geological factors; Among them, terrain factors: elevation, slope, slope aspect, terrain moisture index, plane curvature, slope surface curvature, slope position, slope shape, slope length, terrain undulation, surface roughness, and surface cutting depth; Remote sensing factors: various bands of remote sensing images, leaf area index, ratio vegetation index, difference environmental vegetation index, green vegetation index, vertical vegetation index, and normalized difference vegetation index; Climate variables: average annual rainfall, average annual temperature, average sunshine hours, average wind speed; Biological factors: vegetation type, land use; Geological factors: parent material, hydrogeological map.

8. A soil thickness type prediction method based on feature ensemble learning according to claim 1 or 5, characterized in that: The resampling method is any one of a bilinear interpolation method, a nearest neighbor assignment method, a cubic convolution interpolation method, and a majority method.

9. The soil thickness type prediction method based on feature ensemble learning according to claim 1, characterized in that: The various types of models to be trained that are preset in step E include geographically weighted regression, random forest, support vector machine, deep learning, decision tree, k-nearest neighbor algorithm, Bayesian classification, and classification tree.

10. The soil thickness type prediction method based on feature ensemble learning according to claim 1, characterized in that: The calculation of the accuracy of each soil thickness prediction model and the calculation of the accuracy of each soil depth interval prediction model in step E, as well as the calculation of the accuracy of the soil thickness integrated prediction model and the calculation of the accuracy of the soil depth interval integrated prediction model in step F, adopt any one of ten-fold cross validation, five-fold cross validation, three-fold cross validation, and leave-one-out validation as the accuracy calculation method, and adopt any one of the root mean square error, performance deviation ratio, sum of squares of error, and correction determination coefficient as the accuracy standard.

Citation Information

Patent Citations

  • Large-scale soil organic carbon spatial distribution simulation method involving environmental factors

    CN104408258A

  • Canonical correspondence analysis-based soil nitrogen reserve estimation method

    CN106980750A