Farmland soil sampling space optimization method based on local heterogeneity

By calculating the soil variation coefficient and dividing spatial continuous sub-regions, and optimizing soil sampling with fuzzy C-mean classifier and environmental covariates, the problem of local heterogeneity of soil is solved, the accuracy and robustness of soil mapping are improved, and refined agricultural management is supported.

CN120333893APending Publication Date: 2025-07-18ZHENGZHOU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510516908.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing soil sampling methods fail to effectively consider the impact of local soil heterogeneity level on sample size allocation and spatial layout optimization, resulting in insufficient accuracy of soil mapping results.

Method used

The local heterogeneity level was evaluated by calculating the coefficient of variation of the soil, and the spatial continuous sub-regions were divided using the heterogeneity change algorithm, and the key sampling locations were identified in combination with the fuzzy C-mean classifier and environmental covariates, and the soil sampling space was optimized.

Benefits of technology

It has achieved improvement of the accuracy and robustness of farmland soil mapping with a small sample size, ensured sufficient coverage of geospatial and characteristic space, and supported refined agricultural management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120333893A_ABST
    Figure CN120333893A_ABST
Patent Text Reader

Abstract

The invention discloses a farmland soil sampling space optimization method based on local heterogeneity, and belongs to the field of soil sampling space optimizing.The method comprises the steps that the heterogeneity level of soil in a research area is calculated through a variable coefficient; according to the soil heterogeneity level of the research area, dividing the research area into spatial continuous sub-areas with similar soil heterogeneity level by using a heterogeneity change algorithm; and on the basis of the environment covariables, a method of combining a fuzzy c-mean classifier and a heterogeneity change algorithm is adopted, geographic and feature spaces are optimized, key sampling positions in continuous sub-regions of each space are identified, and farmland soil sampling space optimization is completed. According to the method, the problem that the influence of the soil local heterogeneity level on sample size distribution and spatial layout optimization is not considered in an existing method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of soil sampling spatial optimization, and particularly relates to a method for optimizing the spatial distribution of farmland soil sampling based on local heterogeneity. Background Art

[0002] Soil, as an important part of the earth's surface system, has spatial variability. A full understanding of the spatial variability of soil nutrients in farmland will contribute to the development goal of precision agriculture. Currently, the main method for obtaining the soil conditions in a certain area is still to collect soil samples in the field and then analyze the samples to obtain the soil condition results. However, if soil samples are collected on a large scale and intensively, it will consume a large amount of manpower, material resources and financial resources. Soil sampling is to obtain a small number of key soil samples and then use a certain prediction model to estimate the soil conditions, which can achieve a good balance between the sampling cost and the expected soil prediction accuracy. Accurate soil sample data is the basis for correctly inferring the soil conditions. Therefore, soil sampling design is crucial for the effective implementation of digital soil mapping. If the sample quality is not high, it is difficult to obtain reliable digital soil mapping results even using the most advanced prediction models. Therefore, it is necessary to develop a reasonable and effective soil sampling method to ensure the reliability of digital soil mapping results.

[0003] To achieve this goal, a variety of sampling design methods have been developed. Currently, soil sampling methods are mainly divided into design-based sampling methods and model-based sampling methods. Design-based methods determine the number and location of soil samples through probability sampling or environmental variable-assisted sampling. Probability sampling methods include grid sampling, stratified random sampling, spatial balance sampling and spatial coverage sampling; while environmental variable-assisted sampling methods include fuzzy C-means sampling, feature space coverage sampling, multi-level representative sampling and conditional Latin hypercube sampling. Model-based methods determine the number and location of soil samples by minimizing the estimation variance of the geostatistical model, such as sampling based on the Kriging model. To improve the efficiency of the model, a simulated annealing algorithm is usually used to optimize the distribution of samples. However, this method requires prior knowledge of soil spatial variability, which is usually obtained through preliminary soil sampling.

[0004] Regardless of the sampling design method adopted, existing studies mainly focus on optimizing the quantity and location of samples in the geographical space or feature space. Achieving a good sample distribution in the geographical space is crucial for obtaining reliable interpolation results. Especially when using prediction models such as machine learning to infer soil distribution, good sample coverage in the feature space is particularly critical for calibration. However, in practical applications, different soil regions may exhibit different characteristics, which may lead to the inability to establish an effective prediction model during the sample design stage. Research shows that when samples are evenly distributed in the geographical space, the estimation accuracy of the spatial mean of environmental variables is higher. Therefore, achieving an appropriate coverage balance in the geographical space and feature space is of great significance for improving the robustness of sampling design and resisting model prediction hypothesis biases.

[0005] When designing a soil sampling plan, efforts are made to ensure effective coverage in the geographical space and feature space. For example, existing plans have proposed an unbiased soil sampling optimization method based on categorical auxiliary variable information. These sampling design methods have achieved remarkable results in their respective application fields. However, these methods usually start from a global perspective and optimize the quantity and distribution of samples. In agricultural areas, the soil often exhibits significant local variability. Existing research has shown that in areas with large soil variability, more soil samples need to be collected. If insufficient sampling is carried out in these areas with large variability, it may lead to insufficient representativeness of soil samples, thus affecting the accuracy of soil mapping. Therefore, it is crucial to analyze the spatial structure of the soil to determine key sampling locations. Integrating the local heterogeneity of the soil is an effective strategy to ensure comprehensive coverage of the geographical space and feature space, thereby formulating a more objective sampling plan. However, in many study areas, there is a lack of prior knowledge of local soil heterogeneity. Given the strong correlation between soil properties and certain environmental covariates, the local soil heterogeneity level can be effectively inferred by analyzing these environmental factors. In addition, these covariates also help to determine the key sampling locations in the sampling plan. Summary of the Invention

[0006] Aiming at the above deficiencies in the prior art, a method for optimizing the spatial layout of farmland soil sampling based on local heterogeneity provided by the present invention solves the problem that the existing methods do not consider the influence of the local soil heterogeneity level on the sample quantity allocation and spatial layout optimization.

[0007] To achieve the above invention objective, the technical solution adopted by the present invention is: A method for optimizing the spatial layout of farmland soil sampling based on local heterogeneity, comprising:

[0008] Calculating the heterogeneity level of the soil in the study area using the coefficient of variation;

[0009] According to the heterogeneity level of the soil in the study area, the study area is divided into several spatially continuous sub-regions with similar soil heterogeneity levels by using the heterogeneity change algorithm;

[0010] Based on environmental covariates, the fuzzy C-means classifier and the heterogeneity change algorithm are used to optimize the geographical and feature spaces, identify the key sampling locations within each spatially continuous sub-region, and complete the spatial optimization of farmland soil sampling.

[0011] Furthermore, the heterogeneity level of the soil in the study area is calculated using the coefficient of variation, specifically as follows:

[0012] Collect soil sample data;

[0013] Determine the window size;

[0014] According to the window size, the coefficient of variation within the window is calculated using the digital elevation model as the heterogeneity level of the soil in the study area:

[0015]

[0016] where CV is the heterogeneity level of the current window of the soil in the study area; X i is the DEM value of the i-th pixel within the window; is the average value of the DEM digital elevation model values within the window; m is the number of pixels of the DEM digital elevation model within the window.

[0017] Furthermore, the specific method for dividing the study area into several spatially continuous sub-regions with similar soil heterogeneity levels by using the heterogeneity change algorithm according to the heterogeneity level of the soil in the study area is as follows:

[0018] Determine the adjacency relationship between soils, and calculate the similarity of local heterogeneity between adjacent soils based on the heterogeneity level of the soil in the study area, the area size of adjacent soil regions, and the length of the common boundary;

[0019] According to the similarity of local heterogeneity between adjacent soils, gradually merge adjacent soils until the preset number of candidate sub-regions is reached;

[0020] Continuously change the value of the number of candidate sub-regions to obtain the division results of the study area corresponding to different numbers of candidate sub-regions;

[0021] Use the overall goodness index to evaluate the division results of the study area corresponding to different numbers of candidate sub-regions, generate an overall goodness curve that changes with the number of candidate sub-regions, and take the number of candidate sub-regions corresponding to the data point with the maximum value of the overall goodness curve as the optimal number of candidate sub-regions;

[0022] Determine the final research area division result according to the optimal number of candidate sub-regions, and obtain several spatially continuous sub-regions with similar soil heterogeneity levels.

[0023] Furthermore, the expression for the similarity of local heterogeneity between adjacent soils is:

[0024]

[0025] Where HC() is the similarity calculation function of local heterogeneity; p1 is the soil unit; p2 is the adjacent soil unit of p1; s1 is the area of p1; s2 is the area of p2; hc() is the local heterogeneity parameter calculation function; l is the common boundary length of p1 and p2; is the variance of the heterogeneity level after the merger of p1 and p2; is the variance of the heterogeneity level of p1; is the variance of the heterogeneity level of p2.

[0026] Furthermore, the expression for the overall goodness is:

[0027]

[0028]

[0029] Where OG is the overall goodness; WV norm is the normalized WV value; MI norm is the normalized MI value; WV is the area-weighted variance; MI is the Moran index; s j is the area of the jth spatially continuous sub-region; v j is the variance of the heterogeneity level of the jth spatially continuous sub-region; n is the number of candidate sub-regions; w jk is the proportion of l jk in the jth spatially continuous sub-region; y j is the mean value of the heterogeneity level of the jth spatially continuous sub-region; is the mean value of the heterogeneity level of the research area; y k is the mean value of the heterogeneity level of the kth spatially continuous sub-region; l jk is the common boundary length between the jth and kth spatially continuous sub-regions; l j is the boundary length of the jth spatially continuous sub-region.

[0030] Furthermore, based on the environmental covariates, the fuzzy C-means classifier and the heterogeneity change algorithm are used to optimize the geographical and feature spaces, and the key sampling positions within each spatially continuous sub-region are identified, specifically as follows:

[0031] For the continuous sub-regions of the current space, use the fuzzy C-means classifier to perform fuzzy classification based on different numbers of clusters, and obtain the fuzzy classification results corresponding to each number of clusters;

[0032] Based on the environmental covariates, according to the fuzzy classification results corresponding to each number of clusters in the continuous sub-regions of the current space, generate a Q-index curve that varies with the number of clusters, and use the number of clusters at the data point where the change value of the tangent slope is greater than the preset Q-index change threshold as the optimal number of clusters for the continuous sub-regions of the current space;

[0033] Obtain the fuzzy classification result corresponding to the optimal number of clusters in the continuous sub-regions of the current space as the final classification result of the continuous sub-regions of the current space;

[0034] Use the heterogeneity change algorithm to divide the continuous sub-regions of the current space into several soil regions with similar soil heterogeneity levels, and obtain the heterogeneity change result of the continuous sub-regions of the current space; the number of the soil regions is the same as the optimal number of clusters;

[0035] Overlay the final classification result and the heterogeneity change result of the continuous sub-regions of the current space, and based on the overlaid result, use the fuzzy category with the largest area in each soil region divided from the continuous sub-regions of the current space as the main category of the corresponding soil region;

[0036] Extract the fuzzy membership degrees of the main categories in each soil region respectively, and identify the soil position with the highest fuzzy membership degree as the key sampling position;

[0037] Obtain the key sampling positions of all continuous sub-regions of the space.

[0038] Further, the expression of the Q-index is:

[0039]

[0040] where Q is the value of the Q-index; is the Q value of the i1-th environmental covariate in the continuous sub-regions of the current space; n1 is the number of environmental covariates; N h is the area of the h-th cluster in the continuous sub-regions of the current space; N is the area of the continuous sub-regions of the current space; is the variance of the values of the i1-th environmental covariate in the h-th cluster of the continuous sub-regions of the current space; L is the number of clusters; is the variance of the values of the i1-th environmental covariate in the continuous sub-regions of the current space.

[0041] The beneficial effects of the present invention are as follows: The present invention ensures full coverage in the geographical space and the feature space; by considering the differential level of soil local heterogeneity, sub-regions are divided, the number of samples in different sub-regions is reasonably allocated, and the fuzzy C-means and heterogeneity change (HC) algorithms are used to optimize the spatial positions of the samples, so that good mapping accuracy of multiple soil properties in farmland can be achieved with a smaller sample size, providing technical support and data support for precision agricultural soil resource management such as variable fertilization and farmland quality assessment; the local spatial structure information is effectively utilized, so that soil samples that can accurately reflect the spatial variation of different soil characteristics can be collected; by using the inferred local heterogeneity level, the soil is divided into spatially continuous sub-regions, thus realizing objective sample allocation, because regions with higher local heterogeneity usually require more samples for accurate soil mapping; effective coverage of the geographical space and the feature space is achieved within these sub-regions, promoting the identification of key sampling positions and further enhancing the robustness of the method. Description of the Drawings

[0042] Figure 1 It is a flowchart of the method of the present invention.

[0043] Figure 2 It is a schematic diagram of the CV results when the window size is different in the embodiment of the present invention.

[0044] Figure 3 It is a schematic diagram of the results with different numbers of sub-regions in the embodiment of the present invention.

[0045] Figure 4 It is a schematic diagram of the Q-index results when the number of clusters changes in each sub-region in the embodiment of the present invention.

[0046] Figure 5 It is a schematic diagram of obtaining key sampling positions in the embodiment of the present invention.

[0047] Figure 6 It is a schematic diagram of the sampling results of different methods in the embodiment of the present invention.

[0048] Figure 7 It is a schematic diagram of the kernel density analysis results generated by four sampling methods in the embodiment of the present invention.

[0049] Figure 8 It is a schematic diagram of the box plot results of five soil properties of four sampling methods in the embodiment of the present invention.

[0050] Figure 9 It is a schematic diagram of the OK mapping accuracy of four sampling methods under different soil properties in the embodiment of the present invention.

[0051] Figure 10 It is a schematic diagram of the OK mapping results of multiple soil properties in the embodiment of the present invention. Detailed Embodiments

[0052] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0053] As Figure 1 shown, in an embodiment of the present invention, a method for optimizing the spatial sampling of farmland soil based on local heterogeneity includes:

[0054] Calculating the heterogeneity level of the soil in the study area using the coefficient of variation;

[0055] According to the heterogeneity level of the soil in the study area, using the heterogeneity change algorithm to divide the study area into several spatially continuous sub-regions with similar soil heterogeneity levels;

[0056] Based on environmental covariates, using the fuzzy C-means classifier and the heterogeneity change algorithm to optimize the geographical and feature spaces, identifying the key sampling positions within each spatially continuous sub-region, and completing the optimization of the spatial sampling of farmland soil.

[0057] The specific method of calculating the heterogeneity level of the soil in the study area using the coefficient of variation is as follows:

[0058] Collecting soil sample data;

[0059] Determining the window size;

[0060] According to the window size, calculating the coefficient of variation within the window using the digital elevation model as the heterogeneity level of the soil in the study area:

[0061]

[0062] where CV is the heterogeneity level of the current window of the soil in the study area; X i is the DEM value of the i-th pixel within the window; is the average value of the DEM digital elevation model values within the window; m is the number of pixels of the DEM digital elevation model within the window.

[0063] In this embodiment, inferring the heterogeneity level of the local soil using the coefficient of variation (CV) includes:

[0064] Collecting soil sample data and soil map products, including their spatial distribution, attribute characteristics, and related environmental covariates;

[0065] Use a digital elevation model (DEM) to calculate the coefficient of variation (CV) within windows of different sizes. It should be noted that a larger CV value represents a stronger level of local soil heterogeneity.

[0066] Reasons for using DEM to calculate the CV coefficient to characterize the level of local soil heterogeneity: Topographic attributes are crucial for understanding water flow and the migration of soil substances and particles in the terrain. In areas where topographic changes are obvious within a small range, the local heterogeneity of soil composition is often very high. Therefore, use DEM to calculate the coefficient of variation within different windows. The larger the coefficient of variation, the more obvious the topographic changes within that range, and thus the higher the level of soil heterogeneity.

[0067] Determine an appropriate window for CV calculation. Different window sizes will also produce different CV results. To objectively evaluate the CV value, analyze the mean and variance curves of the CV value as the window size increases, and select an appropriate window size.

[0068] In this embodiment, the process of selecting key soil sample locations using this method:

[0069] First, generate curves showing how the mean and variance of the coefficient of variation (CV) change as the window size increases, as shown in Figure 2 (a). The window sizes range from 3×3 pixels to 37×37 pixels, with an interval of 2 pixels. The mean of the coefficient of variation shows an upward trend when the window size initially increases, but then decreases after the window continues to increase. This indicates that larger windows are helpful for evaluating local soil heterogeneity. However, after exceeding a certain critical value, the impact of further increasing the window size on the evaluation of local soil heterogeneity tends to weaken. The variance of the coefficient of variation (CV) increases as the window size increases, but its upward trend is relatively slow. This indicates that larger windows make the differences in local soil heterogeneity more stable. To more objectively evaluate the local soil heterogeneity level, a window size of 21×21 pixels (i.e., 630×630 meters) was selected to calculate the coefficient of variation. Figure 2 (b) shows the corresponding CV data, reflecting the degree of difference in local soil heterogeneity within different soil regions.

[0070] According to the heterogeneity level of the soil in the study area, use the heterogeneity change algorithm to divide the study area into several spatially continuous sub-regions with similar soil heterogeneity levels. Specifically:

[0071] Determine the adjacency relationship between soils, and calculate the similarity of local heterogeneity between adjacent soils based on the heterogeneity level of the soil in the study area, the area size of adjacent soil regions, and the length of the common boundary.

[0072] Gradually merge adjacent soils according to the similarity of local heterogeneity between adjacent soils until the preset number of candidate sub-regions is reached.

[0073] Continuously change the value of the number of candidate sub-regions to obtain the research area division results corresponding to different numbers of candidate sub-regions;

[0074] Use the overall goodness index to evaluate the research area division results corresponding to different numbers of candidate sub-regions, generate an overall goodness curve that changes with the number of candidate sub-regions, and take the number of candidate sub-regions corresponding to the data point with the maximum value of the overall goodness curve as the optimal number of candidate sub-regions;

[0075] According to the optimal number of candidate sub-regions, determine the final research area division result to obtain several spatially continuous sub-regions with similar soil heterogeneity levels.

[0076] The expression for the similarity of local heterogeneity between adjacent soils is:

[0077]

[0078] Among them, HC() is a function for calculating the similarity of local heterogeneity; p1 is a soil unit; p2 is an adjacent soil unit of p1; s1 is the area of p1; s2 is the area of p2; hc() is a function for calculating local heterogeneity parameters; l is the common boundary length of p1 and p2; is the variance of the heterogeneity level after p1 and p2 are merged; is the variance of the heterogeneity level of p1; is the variance of the heterogeneity level of p2.

[0079] In this embodiment, the contribution of adjacent regions is weighted based on the area and common boundary of adjacent regions. Adjacent regions with a smaller area or a larger common boundary are more likely to belong to the same partition because regions with a smaller area or a larger common boundary are more likely to be part of adjacent objects. This formula believes that the smaller the area of adjacent soil regions and the larger the adjacent common boundary, the higher the probability that they belong to the same partition. When the areas of adjacent regions are equal and the adjacent common boundary is small, the probability that they belong to the same partition is low; when the areas of adjacent soil regions are all small or the difference in area is large, the probability that they belong to the same partition is high.

[0080] The expression for the overall goodness is:

[0081]

[0082]

[0083] Among them, OG is the overall goodness; WV norm is the normalized WV value; MI norm is the normalized MI value; WV is the area-weighted variance; MI is the Moran index; sj is the area of the j-th spatially continuous sub-region; v j is the variance of the heterogeneity level of the j-th spatially continuous sub-region; n is the number of candidate sub-regions; w jk is the proportion of l jk in the j-th spatially continuous sub-region; y j is the mean of the heterogeneity level of the j-th spatially continuous sub-region; is the mean of the heterogeneity level of the study area; y k is the mean of the heterogeneity level of the k-th spatially continuous sub-region; l jk is the common boundary length between the j-th spatially continuous sub-region and the k-th spatially continuous sub-region; l j is the boundary length of the j-th spatially continuous sub-region.

[0084] In this embodiment, w of the present invention jk is the ratio of the length of the common boundary of adjacent regions to the length of their regional boundaries. Since only the relationship between adjacent regions is considered here, the original distance-based weight is meaningless, so this weight coefficient is adjusted.

[0085] In this embodiment, an OG curve is generated to evaluate the optimal number of candidate sub-regions, as shown in Figure 3 (a). The normalized WV value increases with the increase in the number of sub-regions, indicating that the homogeneity within the sub-regions is improved. On the contrary, with the increase in the number of sub-regions, the normalized MI value shows a downward trend, indicating that the heterogeneity between the sub-regions decreases. In order to maximize the homogeneity within the regions and the heterogeneity between the regions, the OG curve obtained by F-measure is further analyzed. As the number of sub-regions increases, the OG curve initially shows an upward trend and then a downward trend. The highest OG value appears when the number of sub-regions is 12, so this number of sub-regions is determined as the optimal value. Figure 3 (b) shows the segmentation result when the number of sub-regions is 12.

[0086] Based on the environmental covariates, the fuzzy C-means classifier and the heterogeneity change algorithm are used to optimize the geographical and feature spaces, and the key sampling positions within each spatially continuous sub-region are identified, specifically:

[0087] For the current spatially continuous sub-region, the fuzzy C-means classifier is used to perform fuzzy classification based on different numbers of clusters, and the fuzzy classification results corresponding to each number of clusters are obtained;

[0088] Based on the environmental covariates, according to the fuzzy classification results corresponding to each number of clusters of the current spatially continuous sub-region, a Q-index curve that changes with the number of clusters is generated, and the number of clusters at the data point where the slope change value of the tangent line is greater than the preset Q-index change threshold is used as the optimal number of clusters of the current spatially continuous sub-region;

[0089] Obtain the fuzzy classification result corresponding to the optimal number of clusters for the current spatially continuous sub-region as the final classification result of the current spatially continuous sub-region;

[0090] Use the heterogeneity change algorithm to divide the current spatially continuous sub-region into several soil regions with similar soil heterogeneity levels to obtain the heterogeneity change result of the current spatially continuous sub-region; the number of the soil regions is the same as the optimal number of clusters;

[0091] Overlay the final classification result and the heterogeneity change result of the current spatially continuous sub-region. Based on the overlaid result, take the fuzzy category with the largest area in each soil region divided from the current spatially continuous sub-region as the main category of the corresponding soil region;

[0092] Extract the fuzzy membership degrees of the main categories in each soil region respectively, and identify the soil positions with the highest fuzzy membership degrees as the key sampling positions;

[0093] Obtain the key sampling positions for all spatially continuous sub-regions.

[0094] The expression of the Q index is:

[0095]

[0096] where Q is the Q index value; is the Q value of the i1-th environmental covariate in the current spatially continuous sub-region; n1 is the number of environmental covariates; N h is the area of the h-th cluster in the current spatially continuous sub-region; N is the area of the current spatially continuous sub-region; is the variance of the values of the i1-th environmental covariate in the h-th cluster in the current spatially continuous sub-region; L is the number of clusters; is the variance of the values of the i1-th environmental covariate in the current spatially continuous sub-region.

[0097] In this embodiment, the Q index takes into account the variability of multiple variables.

[0098] To analyze the appropriate number of samples after applying FCM to six environmental covariates, the Q value curves of each sub-region are generated, as Figure 4As shown. The curve shows that when the cluster number varies from 1 to 10, the Q values of most subregions show a trend of first rising rapidly and then leveling off. The increase in the Q value indicates an improvement in the ability to classify environmental covariates and divide them into different groups. However, subsequently, as the number of clusters increases, the change in the Q value is small, indicating that the spatial heterogeneity has tended to be stable. To maintain a high level of heterogeneity in the FCM results, the inflection points of each curve were further analyzed to determine the appropriate number of clusters within each subregion. For subregions 2, 3, 4, 5, 6, 7, 10, 11, and 12, their inflection points are relatively obvious, so the appropriate number of clusters can be determined. In contrast, for subregions 1, 8, and 9, the appropriate number of clusters was determined through a comprehensive analysis of the existing data.

[0099] In this embodiment, six environmental covariates are used, and the HC algorithm is applied to each subregion, with the number of segments matching the number of clusters;

[0100] Overlay analysis is performed on the fuzzy C-means classification results and HC results of each subregion, the fuzzy membership degrees of the main categories in each subregion are extracted, and the soil positions with the highest fuzzy membership degree are identified as key sampling points;

[0101] Based on the optimal number of clusters in each subregion and six environmental covariates, overlay analysis of the FCM and HC results is performed to determine the key sample positions. Figure 5 The corresponding experimental results are shown. Figure 5 (a) and Figure 5 (b) show that the FCM results show spatial separability, while the HC results show spatial continuity. To ensure comprehensive coverage of the geographical space and feature space, the fuzzy membership degrees of the main categories in each segment of each subregion are calculated ( Figure 5 (c)). Subsequently, the position with the highest membership degree within each segment is selected as the key sampling position ( Figure 5 (d)).

[0102] To verify the effectiveness of the proposed method, the present invention is compared with stratified random sampling (SRS), k-means sampling (KMS), and conditional Latin hypercube sampling (cLHS). It should be particularly noted that the SRS method only uses the digital elevation model (DEM) to ensure spatial continuity between layers, while the KMS and cLHS methods respectively use six environmental covariates.

[0103] To evaluate the mapping accuracy of five soil properties (soil organic matter content (SOM), pH value, total nitrogen (TN), available phosphorus (AP), and available potassium (AK)), two indices were adopted: mean absolute error (MAE) and root mean square error (RMSE).

[0104] Figure 6 The sampling plans for different methods are shown, Figure 6 (a) corresponding to the SRS method, Figure 6 (b) corresponding to the KMS method, Figure 6 (c) corresponding to the cLHS method. The results show that different sampling plans led to differences in the distribution of soil samples. Figure 7 and Figure 8 The results of kernel density analysis and box plots are shown, Figure 7 (a) corresponding to the proposed method, Figure 7 (b) corresponding to the SRS method, Figure 7 (c) corresponding to the KMS method, Figure 7 (d) corresponding to the cLHS method. By comparing the sampling plans generated by the proposed method with other methods, the differential characteristics of the soil sample distribution were analyzed. A higher kernel density value means a higher degree of aggregation of sampling points. The kernel density rankings are as follows: KMS (41.84) > cLHS (37.96) > SRS (33.11) > the proposed method (30.78). This indicates that the proposed method performs well in the spatial distribution of soil samples, while the KMS method has the worst coverage effect of the sample distribution. As Figure 8 shown, the box plots of the five soil properties generated by the four sampling methods further show that the proposed method performs well in obtaining more stable and wider ranges of soil property values. For the SOM property, the ranges of soil sample values generated by the KMS and cLHS methods are relatively limited. However, for the AP and AK properties, the performances of the KMS and SRS methods are exactly the opposite. These observations indicate that the proposed method has strong stability in ensuring the effective spatial distribution of soil characteristics. Therefore, the proposed method is superior to other competing methods in terms of the coverage potential in the geospatial and feature spaces.

[0105] The ordinary kriging (OK) mapping accuracy of the five soil properties under the four sampling methods was quantitatively evaluated. Figure 9The corresponding OK mapping accuracy is presented. For the two soil properties of SOM and AK, the proposed method performs best among the four sampling methods, specifically showing the lowest MAE and RMSE values. On the contrary, the soil maps generated by the cLHS method have the lowest accuracy. In the soil mapping of the pH property, the proposed method is significantly superior to the SRS method compared with KMS and cLHS. In addition, the proposed method and cLHS are also superior to SRS and KMS in the mapping accuracy of the TN property. However, in the soil mapping of the AP property, the performances of the four methods are comparable. Generally speaking, the SRS and cLHS methods vary in the mapping accuracy of different soil properties, while the proposed method shows higher stability in the mapping of multiple soil properties and is superior to KMS in most cases. In summary, the proposed method demonstrates stronger stability than other competing sampling methods in obtaining satisfactory OK mapping accuracy.

[0106] The proposed method shows good potential in obtaining soil samples that can provide accurate mapping results. Therefore, these soil samples are used for ordinary kriging (OK) mapping of multiple soil attributes. The results of the multiple soil attribute mapping are as Figure 10 shown. It can be seen from the figure that the SOM level gradually increases from the southwest direction to the northeast direction ( Figure 10 (a)). To improve and maintain the soil health of farmland, it is recommended to implement straw return to the field in areas with relatively low SOM levels in the next few years. In addition, the pH value of the farmland is higher than 7 ( Figure 10 (b)), indicating that the soil is alkaline. Different crops have different tolerances to alkaline soil. Therefore, the types of crops to be planted need to be adjusted according to the specific pH value. The TN content in the central farmland area is relatively high, while the western and eastern farmland areas show relatively low TN content. The spatial distributions of the AP and AK levels are inversely proportional to the TN content. Therefore, the fertilization plan, crop types, and rotation systems need to be adjusted accordingly to achieve a balance in soil nutrient levels.

Claims

1. A method for optimizing the spatial sampling of farmland soil based on local heterogeneity, characterized in that Including: Calculating the heterogeneity level of the soil in the study area using the coefficient of variation; Dividing the study area into several spatially continuous sub-regions with similar soil heterogeneity levels using the heterogeneity change algorithm according to the heterogeneity level of the soil in the study area; Based on environmental covariates, optimizing the geographical and feature spaces using the fuzzy C-means classifier and the heterogeneity change algorithm, and identifying the key sampling locations within each spatially continuous sub-region to complete the spatial optimization of farmland soil sampling.

2. The method for optimizing the spatial sampling of farmland soil based on local heterogeneity according to claim 1, characterized in that, The specific method for calculating the heterogeneity level of the soil in the study area using the coefficient of variation is as follows: Collecting soil sample data; Determining the window size; Calculating the coefficient of variation within the window using the digital elevation model according to the window size as the heterogeneity level of the soil in the study area: Among them, CV is the heterogeneity level of the current window of the soil in the study area; X i is the DEM value of the i-th pixel within the window; is the average value of the DEM digital elevation model values within the window; m is the number of pixels of the DEM digital elevation model within the window.

3. The method for optimizing the spatial sampling of farmland soil based on local heterogeneity according to claim 1, wherein The specific method for dividing the study area into several spatially continuous sub-regions with similar soil heterogeneity levels using the heterogeneity change algorithm according to the heterogeneity level of the soil in the study area is as follows: Determining the adjacency relationship between soils, and calculating the similarity of local heterogeneity between adjacent soils based on the heterogeneity level of the soil in the study area, the area size of adjacent soil regions, and the length of the common boundary; Gradually merging adjacent soils according to the similarity of local heterogeneity between adjacent soils until the preset number of candidate sub-regions is reached; Continuously changing the value of the number of candidate sub-regions to obtain the division results of the study area corresponding to different numbers of candidate sub-regions; Evaluating the division results of the study area corresponding to different numbers of candidate sub-regions using the overall goodness index, generating an overall goodness curve that changes with the number of candidate sub-regions, and taking the number of candidate sub-regions at the data point corresponding to the maximum value of the overall goodness curve as the optimal number of candidate sub-regions; Determining the final division result of the study area according to the optimal number of candidate sub-regions to obtain several spatially continuous sub-regions with similar soil heterogeneity levels.

4. The method for optimizing the spatial sampling of farmland soil based on local heterogeneity according to claim 3, characterized in that The expression for the similarity of local heterogeneity between adjacent soils is: Among them, HC() is the similarity calculation function for local heterogeneity; p1 is the soil unit; p2 is the adjacent soil unit of p1; s1 is the area of p1; s2 is the area of p2; hc() is the local heterogeneity parameter calculation function; l is the common boundary length of p1 and p2; is the variance of the heterogeneity level after the merger of p1 and p2; is the variance of the heterogeneity level of p1; is the variance of the heterogeneity level of p2.

5. The method for optimizing the spatial sampling of farmland soil based on local heterogeneity according to claim 3, wherein The expression for the overall goodness is: Among them, OG is the overall goodness; WV norm is the normalized WV value; MI norm is the normalized MI value; WV is the area-weighted variance; MI is the Moran index; s j is the area of the j-th spatially continuous sub-region; v j is the variance of the heterogeneity level of the j-th spatially continuous sub-region; n is the number of candidate sub-regions; w jk is the proportion of l jk in the j-th spatially continuous sub-region; y j is the mean of the heterogeneity level of the j-th spatially continuous sub-region; is the mean of the heterogeneity level of the study area; y k is the mean of the heterogeneity level of the k-th spatially continuous sub-region; l jk is the common boundary length between the j-th spatially continuous sub-region and the k-th spatially continuous sub-region; l j is the boundary length of the j-th spatially continuous sub-region.

6. The optimized method for the spatial sampling of farmland soil based on local heterogeneity according to claim 1, characterized in that The specific method for optimizing the geographical and feature spaces using the fuzzy C-means classifier and the heterogeneity change algorithm based on environmental covariates and identifying the key sampling locations within each spatially continuous sub-region is as follows: For the current spatially continuous sub-region, performing fuzzy classification using the fuzzy C-means classifier based on different numbers of clusters to obtain the fuzzy classification results corresponding to each number of clusters; Based on environmental covariates, generating a Q-index curve that changes with the number of clusters according to the fuzzy classification results corresponding to each number of clusters in the current spatially continuous sub-region, and taking the number of clusters at the data point where the change value of the slope of the tangent line is greater than the preset Q-index change threshold as the optimal number of clusters in the current spatially continuous sub-region; Obtaining the fuzzy classification result corresponding to the optimal number of clusters in the current spatially continuous sub-region as the final classification result of the current spatially continuous sub-region; Dividing the current spatially continuous sub-region into several soil regions with similar soil heterogeneity levels using the heterogeneity change algorithm to obtain the heterogeneity change result of the current spatially continuous sub-region; the number of the soil regions is the same as the optimal number of clusters; Overlay the final classification results and heterogeneity change results of the current spatially continuous sub-regions. Based on the overlay results, identify the fuzzy class with the largest area in each soil region divided from the current spatially continuous sub-regions as the main class of the corresponding soil region; Extract the fuzzy membership degrees of the main classes in each soil region respectively, and identify the soil positions with the highest fuzzy membership degrees as the key sampling positions; Obtain the key sampling positions for all spatially continuous sub-regions.

7. The method for optimizing the spatial sampling of farmland soil based on local heterogeneity according to claim 6 is characterized in that The expression of the Q index is as follows: Among them, Q is the Q-index value; is the Q value of the i1-th environmental covariate in the current spatially continuous sub-region; n1 is the number of environmental covariates; N h is the area of the h-th cluster in the current spatially continuous sub-region; N is the area of the current spatially continuous sub-region; is the variance of the values of the i1-th environmental covariate in the h-th cluster in the current spatially continuous sub-region; L is the number of clusters; is the variance of the values of the i1-th environmental covariate in the current spatially continuous sub-region.

Citation Information

Cited By

  • Rapid generation method of soil nutrient distribution diagram

    CN120953424A

  • A method for rapid generation of soil nutrient distribution maps

    CN120953424B

  • Agricultural planting soil parameter analysis system

    CN121303605A