Soil type prediction method based on small sample number

By adopting generalized regression neural networks and multiple environmental factors in soil type prediction, the problem of insufficient prediction accuracy and generalization capabilities in the prior art is solved, especially in the case of small samples, higher prediction accuracy and lower sample number requirements are achieved.

CN120070938APending Publication Date: 2025-05-30INST OF LAND ENG & TECH SHAANXI PROVINCIAL LAND ENG CONSTR GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411963142.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing soil type prediction technology still needs to improve in terms of prediction accuracy and generalization ability, especially when there is little sample data, it is difficult to obtain better prediction results.

Method used

A generalized regression neural network is used to combine traditional soil maps, digital elevation models and remote sensing image data to extract a variety of environmental factors for soil type prediction, including parent material information, topographic information, principal component information, texture feature information, land use type and Euclidean distance information.

Benefits of technology

It improves the accuracy and generalization ability of soil type prediction, can ensure high prediction accuracy under small samples, and reduces the requirements for sample size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070938A_ABST
    Figure CN120070938A_ABST
Patent Text Reader

Abstract

The invention discloses a soil type prediction method based on a small sample number, and relates to the technical field of soil type prediction, and the method comprises the following steps: collecting soil basic data; performing indoor processing on the soil basic data, and constructing an environment factor data set; performing field sampling investigation to obtain sampled soil type data; and performing soil type prediction according to the environment factor data set and the sampled soil type data by adopting a generalized regression neural network to obtain a prediction result. According to the method, the generalized regression neural network is adopted to predict the soil type, so that the method has relatively high nonlinear mapping capability and learning speed, the prediction precision and generalization capability are improved, and meanwhile, the requirement on the number of samples is relatively low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of soil type prediction, and particularly to a soil type prediction method based on a small number of samples. Background Art

[0002] The distribution of soil types is the basis for research in multiple fields such as agriculture, environmental science, and geology. The distribution of soil types not only affects the growth of crops and the formulation of agricultural production strategies, but also is directly related to the inference of geological structures, the prediction of natural disasters, and the sustainable utilization of land resources. Therefore, accurately predicting the distribution of soil types is of great significance for improving agricultural production efficiency, protecting the ecological environment, guiding urban planning and construction, etc. In recent years, with the rapid development of artificial intelligence and big data technologies, using neural networks for soil type prediction has become a cutting-edge and effective method. Neural networks can automatically extract the features of soil data and predict the distribution of soil types in unknown areas by learning the patterns in known data, which provides strong support for the scientific management and rational utilization of soil resources.

[0003] In the prior art, Chinese Patent CN103529189A discloses a method for predicting the spatial distribution of soil organic matter based on qualitative and quantitative auxiliary variables. Using an artificial neural network model, on the basis of integrating qualitative and quantitative auxiliary environmental variables such as soil types, topographic factors, and vegetation indices, the spatial distribution prediction of soil organic matter content is carried out, in order to provide a method reference for the spatial distribution prediction of regional high-precision soil properties.

[0004] However, although the above prior art uses an artificial neural network model to carry out the spatial distribution prediction of soil organic matter content, the prediction accuracy and generalization ability still need to be improved; in addition, when the sample data is small, it is difficult for the above prior art to obtain good prediction results. Summary of the Invention

[0005] This application provides a soil type prediction method based on a small number of samples to solve the problems that the prediction accuracy and generalization ability of the existing soil type prediction technology still need to be improved, and it is difficult to obtain good prediction results when the sample data is small.

[0006] On the one hand, this application provides a soil type prediction method based on a small number of samples, including the following steps:

[0007] Step 1, collect soil basic data.

[0008] Step 2, perform in-office processing on the soil basic data to construct an environmental factor dataset.

[0009] Step 3, conduct field sampling surveys to obtain sampled soil type data.

[0010] Step 4: Use a generalized regression neural network to perform soil type prediction based on the environmental factor dataset and the sampled soil type data to obtain a prediction result.

[0011] In a possible implementation, in Step 1, the soil basic data includes: a traditional soil map, a digital elevation model, and remote sensing image data.

[0012] In a possible implementation, in Step 2, the in-office processing of the soil basic data includes:

[0013] Extract the parent material information from the traditional soil map as the parent material-related factor.

[0014] Extract the terrain information from the digital elevation model as the terrain-related factor.

[0015] Extract the principal component information and texture feature information from the remote sensing image data as the remote sensing-related factors.

[0016] Extract the land use type and Euclidean distance information from the remote sensing image data as the human activity factors.

[0017] In a possible implementation, in Step 2, the parent material information includes: the parent material type.

[0018] In a possible implementation, in Step 2, use spatial analysis to extract the terrain information from the digital elevation model.

[0019] The terrain information includes: slope, aspect, plane curvature, profile curvature, terrain wetness index.

[0020] In a possible implementation, in Step 2, use band synthesis and principal component transformation to extract the principal component information from the remote sensing image data, and use the gray-level co-occurrence matrix to extract the texture feature information from the remote sensing image data.

[0021] The principal component information uses the transformed first principal component.

[0022] The texture feature information includes: mean, variance, information entropy.

[0023] In a possible implementation, in Step 2, use the Euclidean distance analysis tool to extract the Euclidean distance information from the remote sensing image data.

[0024] The Euclidean distance information includes: the Euclidean distance from the river and the Euclidean distance from the road.

[0025] In a possible implementation, in step four, the environmental factor data set is first grouped according to geographical location to obtain a sample environmental factor data set and a to-be-predicted environmental factor data set, and the sample environmental factor data set corresponds to the sampled soil type data in terms of geographical location.

[0026] The sample environmental factor data set and the sampled soil type data are used as the training sample data set of the generalized regression neural network, and the to-be-predicted environmental factor data set and the corresponding coordinates are used as the input variable data set of the generalized regression neural network to perform soil type prediction to obtain a prediction result.

[0027] In a possible implementation, after step four, it further includes:

[0028] Step five, perform accuracy evaluation on the prediction result.

[0029] In a possible implementation, the accuracy evaluation uses one or more of the correlation coefficient, root mean square error, and mean absolute error.

[0030] A soil type prediction method based on a small sample size in this application has the following advantages:

[0031] By using the generalized regression neural network for soil type prediction, it has strong non-linear mapping ability and learning speed, can accurately capture the non-linear relationship between input data and output data, improves the prediction accuracy, can learn effective features from limited data in a short time, and improves the generalization ability; at the same time, the generalized regression neural network can converge to the optimized regression surface with a relatively large sample size and can also ensure a high prediction accuracy under small samples, with a low requirement for the sample size.

[0032] It is proposed to extract the parent material information from the traditional soil map as the parent material-related factor, extract the terrain information from the digital elevation model as the terrain-related factor, extract the principal component information and texture feature information from the remote sensing image data as the remote sensing-related factor, and extract the land use type and Euclidean distance information from the remote sensing image data as the human activity factor. By extracting multiple factors to construct the environmental factor data set, the comprehensiveness and reliability of the data are improved, and thus the prediction accuracy is improved.

[0033] It is proposed that the accuracy evaluation uses one or more of the correlation coefficient, root mean square error, and mean absolute error. By performing accuracy evaluation on the prediction result, the performance of the generalized regression neural network can be evaluated and verified. Description of the Drawings

[0034] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0035] Figure 1 It is a schematic flowchart of a soil type prediction method based on a small sample size provided by an embodiment of the present application. Detailed implementation manners

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0037] As Figure 1 shown, an embodiment of the present application provides a soil type prediction method based on a small sample size, including the following steps:

[0038] Step 1: Collect soil basic data.

[0039] Step 2: Perform in-office processing on the soil basic data to construct an environmental factor dataset.

[0040] Step 3: Conduct field sampling surveys to obtain sampled soil type data.

[0041] Step 4: Use a generalized regression neural network to perform soil type prediction based on the environmental factor dataset and the sampled soil type data to obtain a prediction result.

[0042] Exemplarily, in Step 1, the soil basic data includes: traditional soil maps, digital elevation models, and remote sensing image data.

[0043] Specifically, in this embodiment, the traditional soil map is obtained from an existing database. The digital elevation model uses the digital elevation model with a resolution of 30 m in the Geospatial Data Cloud Platform of the Computer Network Information Center, Chinese Academy of Sciences (http: / / www.gscloud.cn). The remote sensing image data uses the Sentinel-2 satellite remote sensing image data in the Copernicus Data Space Ecosystem (https: / / dataspace.copemicus.eu / ). The Sentinel-2 satellite has a revisit period of 10 days and is equipped with a multi-spectral sensor, supporting 13 spectral bands in the visible light, near-infrared (NIR), and short-wave infrared (SWIR) ranges, which has unique advantages for monitoring vegetation information. The image information can fully meet the research needs in fields such as crop yield estimation, forest fires, and marine disasters. The growth and development of plants on the soil will block the surface soil, affecting the acquisition of soil information from remote sensing images. Therefore, the selected remote sensing image data is for spring and winter months.

[0044] Exemplarily, in step two, the in-office processing of the soil basic data includes:

[0045] Extract the parent material information from the traditional soil map as the parent material-related factor.

[0046] Extract the terrain information from the digital elevation model as the terrain-related factor.

[0047] Extract the principal component information and texture feature information from the remote sensing image data as the remote sensing-related factors.

[0048] Extract the land use type and Euclidean distance information from the remote sensing image data as the human activity factors.

[0049] Exemplarily, in step two, the parent material information includes: parent material type.

[0050] Exemplarily, in step two, spatial analysis is used to extract the terrain information from the digital elevation model.

[0051] The terrain information includes: slope, aspect, plane curvature, profile curvature, terrain wetness index.

[0052] Specifically, in this embodiment, the spatial analysis uses Arcgis analysis.

[0053] Exemplarily, in step two, band synthesis and principal component transformation are used to extract the principal component information from the remote sensing image data, and the gray-level co-occurrence matrix is used to extract the texture feature information from the remote sensing image data.

[0054] The principal component information uses the transformed first principal component.

[0055] The texture feature information includes: mean, variance, and information entropy.

[0056] Specifically, in this embodiment, the Layer Stacking tool of ENVI software is used to perform band synthesis on the remote sensing image data, and then the PCA Rotation tool is used for principal component transformation. The first principal component after transformation is selected as the principal component information.

[0057] Specifically, the gray-level co-occurrence matrix is a calculation method based on second-order probability statistics, with good adaptability, which improves the accuracy of image detection and classification. The texture feature information used in this embodiment is also made in ENVI software. Taking the first principal component extracted from the principal component analysis as the basic data, the gray-level co-occurrence matrix in the Filter toolbox is used to extract the texture feature information of the remote sensing image data.

[0058] Specifically, in this embodiment, the land use type is obtained through visual interpretation of the remote sensing image data. In other possible embodiments, it can also be obtained using existing publicly available datasets.

[0059] Exemplarily, in step two, the Euclidean distance analysis tool is used to extract the Euclidean distance information from the remote sensing image data.

[0060] The Euclidean distance information includes: the Euclidean distance from the river and the Euclidean distance from the road.

[0061] Specifically, the Euclidean distance refers to the straight-line distance from a certain pixel in the figure to the nearest source. The larger the value at a certain position, the farther the distance from that position to a certain land feature. In this embodiment, the rivers and roads in the land use type are extracted, and the spatial analysis - distance analysis - Euclidean distance analysis tool in the ArcGIS10.2 toolbox is used to calculate the Euclidean distance from the river and the Euclidean distance from the road. In this embodiment, the Euclidean distance from the river represents the degree of influence of the river on land development and utilization in the study area, and the Euclidean distance from the road represents the density of the traffic road network in the study area. Both can be used as indicators to reflect the degree of human activities from the side.

[0062] Specifically, in this embodiment, in the field sampling survey, the sampled soil type data is obtained by selecting typical sample points based on the collected traditional soil maps, combined with remote sensing image data and land use types, and conducting soil type surveys.

[0063] Exemplarily, in step four, the environmental factor dataset is first grouped according to geographical location to obtain a sample environmental factor dataset and a to-be-predicted environmental factor dataset. The sample environmental factor dataset corresponds to the sampled soil type data in terms of geographical location.

[0064] Taking the sample environmental factor dataset and the sampled soil type data as the training sample dataset of the generalized regression neural network, and taking the environmental factor dataset to be predicted and the corresponding coordinates as the input variable dataset of the generalized regression neural network, soil type prediction is performed to obtain the prediction result.

[0065] Specifically, the generalized regression neural network is composed of four layers in structure, namely the input layer, the pattern layer, the summation layer, and the output layer.

[0066] The input layer is used to receive the input variables and transfer the input variables to the pattern layer. The number of neurons in the input layer is equal to the dimension of the input vector. In this embodiment, the input feature X = (x, y, ef), where x and y represent the horizontal and vertical coordinates, and ef represents the environmental factor data corresponding to the coordinates in the environmental factor dataset to be predicted.

[0067] The number of neurons in the pattern layer is equal to the number of training samples. Each neuron corresponds to a training sample. The transfer function used in the pattern layer (i.e., the function for processing the input variables) generally uses the Gaussian function, as shown in the following formula:

[0068]

[0069] where, P i represents the output of the i-th neuron in the pattern layer, D represents the Euclidean distance between the input variable and the training sample, and σ represents the smoothing parameter. A smaller smoothing parameter will make the network fit the training data more closely, but may lead to overfitting; a larger smoothing parameter will make the network fit the training data more loosely and may lead to underfitting.

[0070] The Euclidean distance D between the input variable and the training sample is shown in the following formula:

[0071]

[0072] where, n represents the dimension of the input variable, x i represents the i-th input vector, and y i represents the i-th training sample.

[0073] The summation layer uses two types of neurons for summation. One type performs arithmetic summation on the outputs of all neurons in the pattern layer, and the other type performs weighted summation on all neurons in the pattern layer. The weight of the weighted summation is the correlation between the output of the pattern layer neuron and the target variable. The number of nodes in the summation layer is equal to the output sample dimension plus 1, and is divided into two parts, including arithmetic summation neurons and weighted summation neurons.

[0074] Among them, the main function of the arithmetic summation neuron is to perform simple summation on the output of the pattern layer, as shown in the following formula:

[0075]

[0076] Among them, S D represents the arithmetic sum value, N represents the number of training samples, and P i represents the output of the i-th neuron in the pattern layer, that is, the similarity value between the input variable and the i-th training sample. By summing all the similarity values, the overall similarity value between the input variable and the entire training sample data set is obtained.

[0077] The weighted sum neuron performs a weighted sum on the output of the pattern layer, as shown in the following formula:

[0078]

[0079] Among them, S N represents the weighted sum value, and y i represents the i-th training sample, that is, by multiplying each training sample by its corresponding similarity value and then summing, a value that comprehensively considers the similarity and the training sample is obtained.

[0080] The number of nodes in the output layer is equal to the number of output dimensions, which is 1 in this embodiment. The output of the node is equal to the output of the summation layer divided, as shown in the following formula:

[0081]

[0082] Among them, SOM represents the node output of the output layer, that is, the prediction result.

[0083] Exemplarily, after step four, it further includes:

[0084] Step five, evaluating the accuracy of the prediction result.

[0085] Exemplarily, the accuracy evaluation uses one or more of the correlation coefficient, root mean square error, and mean absolute error.

[0086] Specifically, the correlation coefficient, root mean square error, and mean absolute error are all common accuracy evaluation indicators, and will not be elaborated here.

[0087] In the embodiment of the present application, by using a generalized regression neural network for soil type prediction, it has strong non-linear mapping ability and learning speed, can accurately capture the non-linear relationship between input data and output data, improves the prediction accuracy, can learn effective features from limited data in a short time, and improves the generalization ability; at the same time, the generalized regression neural network can converge to the optimized regression surface with a large number of samples, and can also ensure a high prediction accuracy with a small number of samples, and has a low requirement for the number of samples.

[0088] The proposed method extracts parent material information from traditional soil maps as parent material-related factors, extracts topographic information from digital elevation models as topographic-related factors, extracts principal component information and texture feature information from remote sensing image data as remote sensing-related factors, and extracts land use types and Euclidean distance information from remote sensing image data as human activity factors. By extracting multiple factors to construct an environmental factor dataset, the comprehensiveness and reliability of the data are improved, thereby improving the prediction accuracy.

[0089] The proposed accuracy evaluation uses one or more of the correlation coefficient, root mean square error, and mean absolute error. By evaluating the accuracy of the prediction results, the performance of the generalized regression neural network can be evaluated and verified.

[0090] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0091] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A soil type prediction method based on a small sample size, characterized in that: The following steps are involved: Step 1: Collect basic soil data; Step 2: Perform internal processing on the soil basic data to construct an environmental factor data set; Step 3: Conduct field sampling survey to obtain sampling soil type data; Step 4: Use a generalized regression neural network to predict the soil type based on the environmental factor data set and the sampled soil type data to obtain a prediction result.

2. A soil type prediction method based on small sample quantity according to claim 1, characterized in that: In step 1, the soil basic data includes: traditional soil map, digital elevation model and remote sensing image data.

3. A soil type prediction method based on small sample quantity according to claim 2, characterized in that: In step 2, the internal processing of the soil basic data includes: Extracting parent material information from the traditional soil map as a parent material related factor; Extracting terrain information from the digital elevation model as a terrain-related factor; Extracting principal component information and texture feature information from the remote sensing image data as remote sensing related factors; The land use type and Euclidean distance information are extracted from the remote sensing image data as human activity factors.

4. A soil type prediction method based on small sample quantity according to claim 3, characterized in that: In step 2, the parent material information includes: parent material type.

5. The soil type prediction method based on small sample quantity according to claim 3, characterized in that: In step 2, spatial analysis is used to extract terrain information from the digital elevation model; The terrain information includes: slope, slope direction, plane curvature, profile curvature, and terrain humidity index.

6. A soil type prediction method based on small sample quantity according to claim 3, characterized in that: In step 2, principal component information is extracted from the remote sensing image data by using band synthesis and principal component transformation, and texture feature information is extracted from the remote sensing image data by using gray level co-occurrence matrix; The principal component information adopts the transformed first principal component; The texture feature information includes: mean, variance, and information entropy.

7. A soil type prediction method based on small sample quantity according to claim 3, characterized in that: In step 2, using a Euclidean distance analysis tool to extract Euclidean distance information from the remote sensing image data; The Euclidean distance information includes: the Euclidean distance to the river and the Euclidean distance to the road.

8. The soil type prediction method based on small sample quantity according to claim 1, characterized in that: In step 4, the environmental factor data set is first grouped according to geographical location to obtain a sample environmental factor data set and a to-be-predicted environmental factor data set, wherein the sample environmental factor data set corresponds to the sampled soil type data in terms of geographical location; The sample environmental factor data set and the sampled soil type data are used as training sample data sets of a generalized regression neural network, and the environmental factor data set to be predicted and the corresponding coordinates are used as input variable data sets of the generalized regression neural network to perform soil type prediction and obtain prediction results.

9. The soil type prediction method based on small sample quantity according to claim 1, characterized in that: Step 4 and beyond also include: Step five: evaluating the accuracy of the prediction results.

10. A soil type prediction method based on small sample quantity according to claim 9, characterized in that: The accuracy evaluation adopts one or more of correlation coefficient, root mean square error, and mean absolute error.

Citation Information

Patent Citations

  • Soil organic matter space distribution predication method based on qualitative and quantitative auxiliary variables

    CN103529189A