Soil organic carbon three-dimensional space drawing method fused with ground visible near infrared spectrum
By integrating the visible near-infrared spectroscopy and INLA-SPDE model on the ground, the sampling point and triangular grid parameters are optimized, and the high cost and low efficiency problems of traditional soil organic carbon prediction are solved, achieving high-precision three-dimensional spatial prediction and uncertain mapping of soil organic carbon.
Patent Information
- Application Number
- CN202510630038.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The traditional soil organic carbon prediction method is costly and time-consuming and labor-intensive, and cannot meet the high-density and high-efficiency soil information acquisition needs. The existing three-dimensional mapping technology has insufficient accuracy and efficiency, especially the INLA-SPDE model lacks limitations in the selection of environmental covariates.
The sampling points are optimized by ground visible near-infrared spectroscopy combined with principal component analysis and Kennard-stone algorithm, and the INLA-SPDE model is constructed. The Gauss Markov random field is solved by the finite element method, and the three-dimensional spatial prediction of soil organic carbon is carried out. Combined with the Matérn covariance function and triangular grid parameter optimization, the posterior distribution estimation and uncertainty mapping of soil organic carbon are realized.
It improves the accuracy and efficiency of three-dimensional spatial prediction of soil organic carbon, reduces the number of sampling points, saves manpower and material resources, improves the accuracy and operability of model prediction, can reflect soil depth information and quantify prediction uncertainty.
Smart Images

Figure CN120543791A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of soil property prediction, and in particular to a three-dimensional spatial mapping method of soil organic carbon that integrates ground visible and near-infrared spectra. Background Art
[0002] As an important component of the Earth's land surface, soil is an important guarantee for crop growth and human activities. Among them, soil properties determine the functional value of soil. Rational assessment of the spatial distribution of soil properties plays an important role in the implementation of precision agriculture. Soil organic carbon (SOC) usually exists in the soil in the form of organic matter and is an important component of soil. SOC can improve soil quality by increasing the aggregate structure in the soil, thereby ensuring the ventilation of soil pores and retaining water. It is one of the key factors affecting the physical and chemical properties of soil. However, traditional SOC prediction methods are costly, time-consuming and labor-intensive, and can no longer meet the current demand for high-density and timely soil information acquisition. With the advancement of technology, digital soil mapping (DSM) has been widely used as a high-precision, high-quality mapping technology. Conventional digital soil mapping is based on the collection of soil surface information and is drawn on a two-dimensional plane. It cannot consider the changes in SOC at different depths. Therefore, how to quickly achieve three-dimensional SOC mapping is the key to current digital soil mapping.
[0003] In recent years, commonly used three-dimensional mapping technologies include 3D-IDW and INLA-SPDE. As a commonly used three-dimensional mapping model, 3D-IDW can achieve rapid visualization of three-dimensional mapping, but because it cannot add environmental covariates, its mapping accuracy is usually low; and the INLA-SPDE model is a Bayesian spatial model. Compared with non-Bayesian spatial models, it can fully consider the uncertainty of prior information of model parameters by embedding horizontal spatial geographic location, vertical depth and environmental covariate variation trend, thereby improving the model's spatial prediction accuracy for soil organic carbon; however, in the existing technology for the application of the INLA-SPDE model, there is no restriction on the selection of environmental covariates. Although the existing technology can overcome the defects of the 3D-IDW three-dimensional mapping model to a certain extent, there is still room for improvement. Summary of the Invention
[0004] The purpose of this application is to provide a three-dimensional spatial mapping method of soil organic carbon that integrates ground visible and near-infrared spectra, which can improve the accuracy and efficiency of three-dimensional spatial prediction of soil organic carbon.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] This application provides a three-dimensional spatial mapping method for soil organic carbon by integrating ground visible and near-infrared spectra, comprising the following steps:
[0007] According to the scope of the study area, several training set sampling points and several validation set sampling points were set up, and surface disturbed soil samples were collected at each training set sampling point to obtain the original field ground visible-near infrared spectra of the surface disturbed soil samples.
[0008] The collected original field ground visible and near-infrared spectra were preprocessed respectively, and the preprocessed field ground visible and near-infrared spectra were reduced in dimension using principal component analysis to extract the principal component data of the field ground visible and near-infrared spectra of each surface disturbed soil sample.
[0009] Based on the principal component data of visible and near-infrared spectra of field surface disturbed soil samples of each surface layer, the Kennard-stone algorithm was used to optimize and screen the training set sampling points to obtain several screened training set sampling points; the Kennard-stone algorithm was also used to reduce the number of training set sampling points to one-fourth of the original number.
[0010] For several screened training set sampling points and several validation set sampling points, disturbed soil sample data at different profile depths are collected respectively to obtain original field ground visible near-infrared spectra and soil organic carbon data at different profile depths of each sampling point. The sampling points are screened training set sampling points or validation set sampling points.
[0011] The original field visible and near-infrared spectra collected at different depths of each sampling point profile were preprocessed, and principal component analysis was used to reduce the dimension of the preprocessed field visible and near-infrared spectra at different depths of each sampling point profile, and the principal component data of the field visible and near-infrared spectra at different depths of each sampling point profile were extracted.
[0012] The optimal triangular grid parameters are determined based on the latitude and longitude of the screened sampling points of the training set and the soil organic carbon data.
[0013] Based on the optimal triangular grid parameters, the Matérn covariance function is set to construct an SPDE model, and the SPDE model is solved by using the finite element method to obtain a discretized Gaussian Markov random field.
[0014] The Gaussian Markov random field is used as a spatial random effect, the principal component data of the field ground visible near-infrared spectrum at different depths of the profile of each sampling point is used as an environmental covariate, and the soil organic carbon data at different depths of the profile of each sampling point is used as a model input. The INLA model is applied to perform three-dimensional spatial prediction of soil organic carbon to obtain a posterior distribution estimate of soil organic carbon; the posterior distribution estimate includes a posterior distribution mean and a posterior distribution standard deviation.
[0015] The spatial distribution of soil organic carbon content is mapped based on the mean of the posterior distribution, and the uncertainty results of soil organic carbon content are mapped based on the standard deviation of the posterior distribution.
[0016] Optionally, several training set sampling points and validation set sampling points are set according to the scope of the study area, and surface disturbed soil samples are collected at each training set sampling point, specifically including the following steps:
[0017] The scope of the study area is the field scale. A grid point method is used to set up several training set sampling points in the study area, and a random sampling method is used to set up several validation set sampling points in the study area.
[0018] Optionally, obtaining the original field ground visible-near infrared spectrum of each surface disturbed soil sample specifically comprises the following steps:
[0019] For any surface disturbed soil sample at a sampling point in the training set, the collected surface disturbed soil sample is placed in a culture dish and compacted to half its volume to obtain a compacted soil sample; the surface disturbed soil sample is compacted until it is ensured that the surface disturbed soil sample has no obvious macropores and no water seepage;
[0020] Field ground visible and near-infrared spectra are collected at several selected positions on the surface of the compacted soil sample, and the field ground visible and near-infrared spectra of each position are averaged to obtain the original field ground visible and near-infrared spectrum of the surface disturbed soil sample.
[0021] Optionally, the collected original field ground visible and near-infrared spectra are preprocessed respectively, including: for any collected original field ground visible and near-infrared spectrum, absorbance conversion, multivariate scattering correction and Savitzky-Golay convolution smoothing are performed in sequence to obtain a preprocessed field ground visible and near-infrared spectrum.
[0022] Optionally, principal component analysis is used to reduce the dimension of the pre-processed field ground visible near-infrared spectrum, and principal component data of the field ground visible near-infrared spectrum of each surface disturbed soil sample are extracted, including:
[0023] For the field visible near-infrared spectrum of any surface disturbed soil sample, when performing principal component analysis dimensionality reduction, the characteristic spectrum with a cumulative contribution rate greater than 85% is extracted from the field visible near-infrared spectrum of each surface disturbed soil sample as the principal component data of the field visible near-infrared spectrum of the surface disturbed soil sample.
[0024] Optionally, for the selected training set sampling points and the validation set sampling points, disturbed soil sample data at different profile depths are collected respectively to obtain original field ground visible near-infrared spectra and soil organic carbon data at different profile depths of each sampling point, specifically including the following steps:
[0025] For any sampling point, disturbed soil samples at different depths of the profile are collected to obtain disturbed soil samples at different depths of the profile of the sampling point; the sampling point is a training set sampling point or a validation set sampling point;
[0026] For the disturbed soil sample at any depth of the sampling point profile, the disturbed soil sample is divided into two sub-samples; the two sub-samples of the disturbed soil sample are used for determination of soil organic carbon data and collection of original field ground visible near-infrared spectra, respectively;
[0027] For a subsample of the disturbed soil sample, measuring soil organic carbon data according to the subsample to obtain soil organic carbon data of the disturbed soil sample;
[0028] For another subsample of the disturbed soil sample, the subsample is placed in a culture dish and compacted to half its volume to obtain a compacted soil sample; after compaction to ensure that there are no obvious large pores and no water seepage, field ground visible and near-infrared spectra are collected at several optional positions on the surface of the compacted soil sample, and the field ground visible and near-infrared spectra of each position are averaged to obtain the original field ground visible and near-infrared spectrum of the disturbed soil sample.
[0029] Optionally, the original field ground visible and near-infrared spectra collected at different depths of the profile of each sampling point are preprocessed, specifically including: performing absorbance conversion, multivariate scattering correction and Savitzky-Golay convolution smoothing processing on the original field ground visible and near-infrared spectra of the disturbed soil samples collected at different depths of any profile in sequence to obtain the preprocessed field ground visible and near-infrared spectra.
[0030] Optionally, principal component analysis is used to reduce the dimension of the field visible and near-infrared spectra of the field ground at different depths of each sampling point profile after preprocessing, and principal component data of the field visible and near-infrared spectra of the field ground at different depths of each sampling point profile are extracted, specifically including: when performing principal component analysis dimensionality reduction, characteristic spectra with a cumulative contribution rate greater than 85% are extracted from the field visible and near-infrared spectra of the field ground at different depths of each profile, as the principal component data of the field visible and near-infrared spectra of the field ground at different depths of each sampling point profile.
[0031] Optionally, the optimal triangular grid parameters are determined based on the latitude and longitude of the selected training set sampling points and the soil organic carbon data, specifically including:
[0032] According to the latitude and longitude of the selected training set sampling points, several sets of different triangular grid parameters are set; the triangular grid parameters include the minimum distance and maximum distance between each training set sampling point and the distance between the training set sampling point and the boundary of the study area;
[0033] According to the soil organic carbon data, several groups of different triangular grid parameters are screened according to the general information criterion to obtain the optimal triangular grid parameters.
[0034] Alternatively, the general information criterion is as follows:
[0035]
[0036] Among them, Y i is the true value of soil organic carbon at the i-th sample point; θ i is the predicted value of soil organic carbon at the i-th sample point; n is the total number of sample points; S is the number of sample points in the triangular grid; P(|) represents the conditional probability of the true value of soil organic carbon at each triangular grid sample point; Var() represents the variance function.
[0037] Alternatively, solve the SPDE model according to the following formula:
[0038]
[0039] Where ξ(s) is the discretized Gaussian Markov random field, n=1,…N, N is the total number of vertices of the triangular grid; is the piecewise linear basis function of the triangular grid; ξ n is the weight of the n-th vertex of the triangular grid under Gaussian distribution.
[0040] Alternatively, three-dimensional spatial prediction of soil organic carbon is performed according to the following formula:
[0041] y(s i )=β0+β1×VNIR(s i )+ξ(si )+ε i .
[0042] Among them, y(s i ) is the soil organic carbon content at the i-th sample point; β0 is the intercept; VNIR(s i ) is the principal component data of the field ground visible near infrared spectrum at the i-th sample point; β1 is the coefficient of the VNIR environmental covariate; ξ(s i ) is the spatial random effect, i.e., the discretized Gaussian Markov random field; ε i It is Gaussian white noise that is independent of spatial position.
[0043] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0044] The present application provides a three-dimensional spatial mapping method for soil organic carbon by integrating ground visible and near-infrared spectra. In this method, according to the scope of the study area, sampling points of the training set and the validation set are set, the surface disturbed soil samples of the training set of the study area are obtained, the original field ground visible and near-infrared spectra are collected for preprocessing, and then the principal component analysis is applied to the preprocessed field ground visible and near-infrared spectra for principal component analysis and dimensionality reduction; based on the extracted principal component data of the field ground visible and near-infrared spectra, the KS algorithm is used to optimize and screen the training set, and the number of sampling points of the optimized and screened training set is reduced to one-fourth of the original training set; based on the screened data set, the original field ground visible and near-infrared spectra and soil organic carbon at different depths of the profile of each sampling point are obtained. Carbon data, the original field ground visible near-infrared spectra at different depths of each sampling point profile were preprocessed and the principal component data were extracted, and the optimal triangular grid parameters were determined at the same time; then, based on the optimal triangular grid parameters, the Matérn covariance function was set to construct the SPDE model, and the finite element method was used to solve it to obtain the discretized Gaussian Markov random field. Subsequently, the Gaussian Markov random field was used as the spatial random effect, and the principal component data of the field ground visible near-infrared spectra were used as the environmental covariate. The INLA model was applied to perform three-dimensional spatial prediction of soil organic carbon, and the posterior distribution estimate of soil organic carbon was obtained. Finally, the spatial distribution of soil organic carbon content and the uncertainty results of soil organic carbon content were mapped based on this. Compared with traditional methods of obtaining soil organic carbon through excavation profiles and laboratory testing, this application combines near-ground sensing technology. On the one hand, by collecting field ground visible near-infrared spectral data of surface disturbed soil samples, the KS algorithm is used to optimize the screening training set to reduce it to one-fourth of the original, thereby improving data acquisition efficiency. On the other hand, by integrating field ground visible near-infrared spectra that can reflect the spectral characteristics related to soil organic carbon, it not only ensures the improvement of accuracy, but also provides a guarantee for the operability of subsequent technology promotion. On the other hand, by using the INLA-SPDE model in this application, it can achieve flexible addition of environmental covariates. At the same time, by establishing and screening the triangular grid of the sampling points, it can take into account the correlation between the points in the study area, thereby further improving the model prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0046] Figure 1 A flowchart of a method for three-dimensional spatial mapping of soil organic carbon by integrating terrestrial visible and near-infrared spectra is provided in accordance with an embodiment of the present application.
[0047] Figure 2 A schematic diagram of the SPDE spatial effect in a three-dimensional spatial mapping method of soil organic carbon that integrates terrestrial visible and near-infrared spectra provided in one embodiment of the present application.
[0048] Figure 3 A schematic diagram of the predicted distribution of soil organic carbon in a three-dimensional spatial mapping method of soil organic carbon that integrates terrestrial visible and near-infrared spectra provided in one embodiment of the present application.
[0049] Figure 4 A schematic diagram of the distribution of predicted standard deviation values in a three-dimensional spatial mapping method for soil organic carbon that integrates ground visible and near-infrared spectra provided in one embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0051] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0052] The present application provides a method for three-dimensional spatial mapping of soil organic carbon by integrating ground visible and near-infrared spectra. In an exemplary embodiment, Figure 1 As shown, the following steps are included:
[0053] A1. Based on the scope of the study area, several training set sampling points and several validation set sampling points are set, and surface disturbed soil samples are collected at each training set sampling point to obtain the original field ground visible and near-infrared spectra of each surface disturbed soil sample. In this embodiment, the means of "setting several training set sampling points and several validation set sampling points based on the scope of the study area" in step A1 specifically includes the following steps:
[0054] A11. The study area is plot-sized. A grid-based sampling method is used to set several training set sampling points within the study area. A random sampling method is used to set several validation set sampling points within the study area. In this example, a grid-based sampling method is used to set 69 training set sampling points at intervals of 100 m × 100 m. A random sampling method is used to set 10 validation set sampling points.
[0055] A12. Collect disturbed surface soil samples from the sampling points of the training set. In this embodiment, soil samples with a thickness of 0-10 cm are collected as disturbed surface soil samples.
[0056] The method of "obtaining the original field ground visible-near infrared spectrum of each surface disturbed soil sample" in step A1 specifically includes the following steps:
[0057] A13. For any surface disturbed soil sample from a sampling point in the training set, place the collected surface disturbed soil sample into a culture dish and compact it to half its volume to obtain a compacted soil sample; compact the surface disturbed soil sample until it has no obvious macropores and no water seepage;
[0058] A14. Collect field ground visible and near-infrared spectra at several selected locations on the surface of the compacted soil sample, and average the field ground visible and near-infrared spectra at each location to obtain the original field ground visible and near-infrared spectrum of the surface disturbed soil sample.
[0059] Specifically, the soil sample should be placed in a petri dish and compacted to half its volume, ensuring that the soil sample has no obvious large pores and no water seepage. A location on the soil sample surface without obvious gaps and debris should be selected for visible-near-infrared (VNIR) spectrum collection. The VNIR spectrum is measured using a Quality Spec Trek portable spectrometer with a spectral range of 350-2500nm. Three locations are selected for each soil sample during measurement, and three spectra are measured repeatedly at each location. The nine spectra obtained are averaged and used as the field ground visible-near-infrared spectral data of the disturbed soil sample at that point.
[0060] A2. Preprocess the collected original field ground visible and near-infrared spectra respectively, use principal component analysis to reduce the dimension of the preprocessed field ground visible and near-infrared spectra, and extract the principal component data of the field ground visible and near-infrared spectra of each surface disturbed soil sample.
[0061] In this embodiment, step A2 of "preprocessing each collected original field visible and near-infrared spectrum" includes sequentially performing absorbance conversion, multivariate scattering correction, and Savitzky-Golay convolution smoothing on each collected original field visible and near-infrared spectrum to obtain a preprocessed field visible and near-infrared spectrum. Before preprocessing the spectral data, noise in the 350-450nm and 2400-2500nm bands is removed.
[0062] In this embodiment, step A2 of "using principal component analysis to reduce the dimensionality of the pre-processed field ground visible near-infrared spectrum and extracting the principal component data of the field ground visible near-infrared spectrum of each surface disturbed soil sample" includes: for the field ground visible near-infrared spectrum of any surface disturbed soil sample, when performing principal component analysis dimensionality reduction, extracting the characteristic spectrum with a cumulative contribution rate greater than 85% of the field ground visible near-infrared spectrum of the surface disturbed soil sample as the principal component data of the field ground visible near-infrared spectrum of the surface disturbed soil sample.
[0063] A3. Based on the principal component data of the field visible-near infrared spectra of each surface disturbed soil sample, the Kennard-stone algorithm was used to optimize and screen the training set sampling points to obtain several screened training set sampling points; the Kennard-stone algorithm was also used to reduce the number of training set sampling points to one-fourth of the original number.
[0064] In this embodiment, in step A3, "using the Kennard-stone algorithm to optimize and screen the training set sampling points to obtain a number of screened training set sampling points" specifically selects the principal component data of the field ground visible-near infrared spectrum of the surface (0-10 cm) disturbed soil sample after principal component analysis, and applies the Kennard-stone algorithm to optimize and screen the existing 69 training set sampling points, and a total of 17 sampling points are screened as new training set sampling points.
[0065] A4. For the selected training set sampling points and the validation set sampling points, disturbed soil sample data at different profile depths are collected to obtain original field ground visible near-infrared spectra and soil organic carbon data at different profile depths of each sampling point.
[0066] In this embodiment, the means of "collecting disturbed soil sample data at different profile depths for the selected training set sampling points and the selected validation set sampling points, respectively, to obtain original field ground visible-near infrared spectra and soil organic carbon data at different profile depths of each sampling point" in step A4 specifically includes the following steps:
[0067] A41. For any sampling point, disturbed soil samples are collected at different depths along the profile. In this embodiment, soil samples are collected at different depths from the surface downward at each sampling point, with each 10 cm layer representing a depth of 1 m. The first layer is 0-10 cm, the second layer is 10-20 cm, the third layer is 20-30 cm, the fourth layer is 30-40 cm, the fifth layer is 40-50 cm, the sixth layer is 50-60 cm, the seventh layer is 60-70 cm, the eighth layer is 70-80 cm, the ninth layer is 80-90 cm, and the tenth layer is 90-100 cm. A total of 10 layers of disturbed soil samples are collected.
[0068] A42. The sampling points are training set sampling points or validation set sampling points. In this embodiment, there are 17 training set sampling points and 10 validation set sampling points.
[0069] A43. For any depth of the disturbed soil sample at the sampling point, divide the disturbed soil sample into two subsamples. These two subsamples are used for soil organic carbon data determination and collection of the original field surface visible and near-infrared spectrum, respectively. Specifically, divide each layer of the disturbed soil sample at each sampling point into two subsamples: one for soil organic carbon data determination; the other, sealed in a sealed bag to maintain a constant moisture content, is used for collection of surface visible and near-infrared spectrum data.
[0070] A44. For a subsample of the disturbed soil sample, soil organic carbon data is measured based on the subsample to obtain soil organic carbon data of disturbed soil samples at different depths in the profile.
[0071] A45. For another subsample of the disturbed soil sample, the subsample is placed in a culture dish and compacted to half its volume to obtain a compacted soil sample; after compaction to ensure that there are no obvious large pores and no water seepage, field ground visible and near-infrared spectra are collected at several optional positions on the surface of the compacted soil sample, and the field ground visible and near-infrared spectra of each position are averaged to obtain the original field ground visible and near-infrared spectrum of the disturbed soil sample.
[0072] Specifically, the soil sample should be placed in a petri dish and compacted to half its volume to ensure that the soil sample has no obvious large pores and no water seepage. The visible near-infrared spectrum is collected at a location on the soil sample surface without obvious gaps and debris. The visible near-infrared spectrum is measured using a Quality Spec Trek portable spectrometer with a spectral range of 350-2500nm. Three locations are selected for each soil sample during the measurement, and three spectra are measured repeatedly at each location. The nine spectra obtained are averaged and used as the raw data of the ground visible near-infrared spectrum of the soil sample.
[0073] A5. Preprocess the original field visible and near-infrared spectra collected at different depths of the sampling point profiles, use principal component analysis to reduce the dimension of the preprocessed field visible and near-infrared spectra at different depths of the sampling point profiles, and extract the principal component data of the field visible and near-infrared spectra at different depths of the sampling point profiles.
[0074] In this embodiment, step A5 of "preprocessing the raw field visible and near-infrared spectra collected at different depths of each sampling point profile" includes sequentially performing absorbance conversion, multivariate scattering correction, and Savitzky-Golay convolution smoothing on the raw field visible and near-infrared spectra of the disturbed soil samples collected at different depths of any profile to obtain preprocessed field visible and near-infrared spectra. Before preprocessing the spectral data, noise in the 350-450 nm and 2400-2500 nm bands is removed.
[0075] In this embodiment, step A5 of "using principal component analysis to reduce the dimensionality of the field visible and near-infrared spectra of the field ground at different depths of each sampling point profile after preprocessing, and extracting principal component data of the field visible and near-infrared spectra of the field ground at different depths of each sampling point profile" includes: when performing principal component analysis dimensionality reduction, extracting characteristic spectra with a cumulative contribution rate greater than 85% of the field visible and near-infrared spectra of the field ground at different depths of each profile as the principal component data of the field visible and near-infrared spectra of the field ground at different depths of each sampling point profile.
[0076] A6. Determine optimal triangular grid parameters based on the latitude and longitude of the selected training set sampling points and the soil organic carbon data. In this embodiment, step A6 specifically includes the following steps:
[0077] A61. According to the latitude and longitude of the selected training set sampling points, set several different sets of triangulated grid parameters; the triangulated grid parameters include the minimum distance and maximum distance between each training set sampling point and the distance between the training set sampling point and the boundary of the study area.
[0078] A62. Based on the soil organic carbon data, several groups of different triangular grid parameters are screened according to a general information criterion to obtain optimal triangular grid parameters.
[0079] Specifically, based on the training set results after the Kennard-Stone algorithm was used to divide the data, the latitude and longitude coordinates of the 17 sampling points and the spatial range data of the study area were obtained. The minimum and maximum distances between the sampling points and the distance between the sampling points and the boundary of the study area were calculated. The parameters of the constrained refined Delaunay triangulation mesh (MESH) of the model were set based on the above results, and the triangulation mesh was screened according to the Widely Available Information Criterion (WAIC). The smaller the WAIC value, the higher the degree of model fit and the optimal model triangulation mesh parameters. The specific formula of the Widely Available Information Criterion (WAIC) is as follows:
[0080]
[0081] Among them, Y i is the true value of soil organic carbon at the i-th sample point; θ i is the predicted value of soil organic carbon at the i-th sample point; n is the total number of sample points; S is the number of sample points in the triangular grid; P(|) represents the conditional probability of the true value of soil organic carbon at each triangular grid sample point; Var() represents the variance function.
[0082] By comparing the WAIC of triangular grid parameters under different combinations, as shown in Table 1.
[0083] Table 1 Triangular grid parameters and WAIC values under different combinations
[0084] max.edge offset cutoff triangular grid Total number of vertices WAIC 100,150 40 25,150 MESH1 327 -983.17 70,150 40 25,150 MESH2 456 -1037.92 70,150 10 25,150 MESH3 569 -958.09 70,150 10 50,200 MESH4 602 -969.17
[0085] Triangulation parameters primarily include max.edge, cutoff, and offset. Max.edge represents the maximum and minimum edge lengths of a triangle; cutoff sets the correlation distance between two points. When the distance between two points is greater than the set value, they do not affect each other; and offset controls the placement of triangles outside the study area. In Table 1, MESH2 has the lowest WAIC value, and this value increases as the total number of triangle vertices increases. Therefore, MESH2 was selected as the optimal triangulation for the study area, ensuring that the model parameters better meet the suitability of the study area.
[0086] A7. Based on the optimal triangular grid parameters, set the Matérn covariance function to construct the SPDE model, and use the finite element method to solve the SPDE model to obtain a discretized Gaussian Markov random field. Specifically, Figure 2The SPDE model links the spatially continuous Matérn GF with the discretized Gaussian Markov random field and solves the SPDE model according to the following formula:
[0087]
[0088] Where ξ(s) is the discretized Gaussian Markov random field, n=1,…N, N is the total number of vertices of the triangular grid; is the piecewise linear basis function of the triangular grid; ξ n is the weight of the n-th vertex of the triangular grid under Gaussian distribution.
[0089] A8. Using the Gaussian Markov random field as a spatial random effect, the principal component data of the ground visible near-infrared spectrum at different depths of each sampling point profile as an environmental covariate, and the soil organic carbon data at different depths of each sampling point profile as a model input, the INLA model is applied to perform a three-dimensional spatial prediction of soil organic carbon to obtain a posterior distribution estimate of soil organic carbon; the posterior distribution estimate includes a posterior distribution mean and a posterior distribution standard deviation. In this embodiment, the three-dimensional spatial prediction of soil organic carbon is performed according to the following formula:
[0090] y(s i )=β0+β1×VNIR(s i )+ξ(s i )+ε i .
[0091] Among them, y(s i ) is the soil organic carbon content at the i-th sample point; β0 is the intercept; VNIR(s i ) is the principal component data of the field ground visible near infrared spectrum at the i-th sample point; β1 is the coefficient of the VNIR environmental covariate; ξ(s i ) is the spatial random effect, i.e., the discretized Gaussian Markov random field; ε i It is Gaussian white noise that is independent of spatial position.
[0092] For the above predictions, the accuracy evaluation index is selected to comprehensively evaluate the prediction accuracy of the model. The accuracy evaluation index includes the determination coefficient (R 2 ), root mean square error (RMSE) and relative analytical error (RPD). 2 The larger the RPD and the smaller the RMSE, the better the prediction effect of the model and the stronger the prediction ability.
[0093] The proposed INLA-SPDE model that integrates ground visible and near-infrared spectral data is compared with the conventional three-dimensional mapping method 3D-IDW. The results are shown in Table 2.
[0094] Table 2 Comparison of prediction accuracy between INLA-SPDE model and 3D-IDW model
[0095]
[0096] 3D-IDW mapping is performed after using the conventional grid division method, and the mapping accuracy is R 2 =0.62, RMSE=1.98, RPD=1.63, which has a good simulation result; when using the 17 training sets screened after KS partitioning, the 3D-IDW model is affected by the spatial heterogeneity of the number of sampling points, and the mapping accuracy cannot be guaranteed (R 2 =0.07, RMSE=3.10, RPD=1.04). However, this application can still ensure higher mapping accuracy after KS partitioning compared with the 3D-IDW result by applying the INLA-SPDE model after integrating the ground visible near-infrared spectrum. Among them, compared with the same KS partitioning result, the accuracy R 2 The result was improved by 0.63, RMSE decreased by 1.32, and RPD increased by 0.78. Compared with the result of retaining 69 training sets, the accuracy R 2 The improvement was 0.08, the RMSE decreased by 0.20, and the RPD increased by 0.19. Therefore, the application of the INLA-SPDE model based on ground visible near-infrared spectroscopy provided in this application can ensure good mapping accuracy and reflect the information of current soil environmental factors; it can also reduce the number of sampling points in the current mapping, thereby saving manpower and material resources wasted in data acquisition.
[0097] A9. Map the spatial distribution of soil organic carbon content based on the mean of the posterior distribution, and map the uncertainty results of soil organic carbon content based on the standard deviation of the posterior distribution. Specifically, the posterior distribution estimate obtained based on the INLA-SPDE model was used with ArcGIS 10.8 software to map the spatial distribution of soil organic carbon content based on the mean of the posterior distribution, as shown in the following example: Figure 3 As shown; the uncertainty results of soil organic carbon content are mapped based on the standard deviation of the posterior distribution, as shown Figure 4 shown.
[0098] The solutions provided by the above embodiments of this application have the following characteristics compared to conventional technologies:
[0099] First of all, the traditional research method for obtaining the training set is usually a grid sampling method with a fixed length as the sampling interval. This method is based on the fact that the spatial scope of the study area is easier to lay out at the field scale, but when the study area is too large, the number of sampling points is often too many, which can easily cause a loss of manpower and material resources. However, this application uses the KS algorithm to obtain the principal component data of the ground visible near-infrared spectrum of the surface (0-10cm) disturbed soil for screening. The number of training sets is reduced from 69 sampling points to 17 sampling points, and the sample screening rate reaches 75%. The soil sampling points can be determined only by collecting surface disturbed soil samples, saving a lot of manpower and material resources. In addition, the sampling point attributes after KS division ensure that the differences between the sample points are significant, and the feasibility of predicting soil organic carbon can be achieved when the sampling point data is sharply reduced. When the accuracy of the INLA-SPDE model is compared with 3D-IDW, it can still maintain the best, further verifying the advantages of the INLA-SPDE model in predicting soil organic carbon content.
[0100] Secondly, this application integrates the principal component data of the visible and near-infrared spectra of the field ground as environmental covariates, effectively considering the complex relationship between the random effects of soil space and environmental variables, and improving the prediction accuracy of the model. Compared with the current data selection methods for digital soil mapping, many studies have combined remote sensing data and terrain data as sources of mapping covariates, but such data can only reflect surface information and are often interfered by information such as image spatiotemporal resolution and cloud cover, making it difficult to achieve three-dimensional modeling prediction of soil depth. This application integrates ground visible and near-infrared spectral data that can reflect the relevant spectral characteristics of soil organic carbon. On the one hand, it improves the model's prediction of soil depth information. On the other hand, due to the ease of operation of the technology, it also provides a guarantee for the operability of subsequent technology promotion.
[0101] Furthermore, by using the INLA-SPDE model, this application can, on the one hand, achieve the flexible addition of environmental covariates, and on the other hand, by establishing and screening the MESH network of sampling points, it can consider the correlation between the points in the study area, thereby further improving the model prediction accuracy.
[0102] Finally, this application combines the INLA-SPDE model for three-dimensional spatial prediction of soil organic carbon, that is, obtaining soil organic carbon results through posterior distribution probability, enriching the current research progress on soil organic carbon reserves at different scales and providing relevant references, and can also quantify the uncertainty of the prediction and provide support for in-depth analysis of mapping accuracy.
Claims
1. A method for three-dimensional spatial mapping of soil organic carbon by integrating ground visible and near-infrared spectra, characterized in that: include: According to the scope of the study area, several training set sampling points and several validation set sampling points were set up, and surface disturbed soil samples were collected at each training set sampling point to obtain the original field ground visible near-infrared spectrum of each surface disturbed soil sample; The collected original field visible and near-infrared spectra were preprocessed respectively, and the principal component analysis was used to reduce the dimension of the preprocessed field visible and near-infrared spectra, and the principal component data of the field visible and near-infrared spectra of each surface disturbed soil sample were extracted; Based on the principal component data of the field visible-near infrared spectra of each surface disturbed soil sample, the Kennard-stone algorithm was used to optimize and screen the training set sampling points to obtain several screened training set sampling points; the Kennard-stone algorithm was also used to reduce the number of training set sampling points to one-fourth of the original number; For the selected training set sampling points and the selected validation set sampling points, disturbed soil sample data at different profile depths are collected to obtain original field ground visible near-infrared spectra and soil organic carbon data at different profile depths of each sampling point; the sampling points are the selected training set sampling points or validation set sampling points; The original field visible and near-infrared spectra collected at different depths of each sampling point section were preprocessed, and the principal component analysis was used to reduce the dimension of the preprocessed field visible and near-infrared spectra at different depths of each sampling point section, and the principal component data of the field visible and near-infrared spectra at different depths of each sampling point section were extracted; Determining optimal triangular grid parameters based on the latitude and longitude of the selected training set sampling points and the soil organic carbon data; Based on the optimal triangular grid parameters, a Matérn covariance function is set to construct an SPDE model, and the SPDE model is solved using a finite element method to obtain a discretized Gaussian Markov random field; The Gaussian Markov random field is used as a spatial random effect, the principal component data of the field ground visible near-infrared spectrum at different depths of each sampling point profile is used as an environmental covariate, and the soil organic carbon data at different depths of each sampling point profile is used as a model input. The INLA model is applied to perform a three-dimensional spatial prediction of soil organic carbon to obtain a posterior distribution estimate of soil organic carbon; the posterior distribution estimate includes a posterior distribution mean and a posterior distribution standard deviation. The spatial distribution of soil organic carbon content is mapped based on the mean of the posterior distribution, and the uncertainty results of soil organic carbon content are mapped based on the standard deviation of the posterior distribution.
2. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 1, characterized in that: According to the scope of the study area, several training set sampling points and validation set sampling points were set, and surface disturbed soil samples were collected at each training set sampling point, including: The scope of the study area is the field scale. A grid point method is used to set up several training set sampling points in the study area, and a random sampling method is used to set up several validation set sampling points in the study area.
3. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 1, characterized in that: Obtain original field ground visible-near infrared spectra of each surface disturbed soil sample, including: For any surface disturbed soil sample at a sampling point in the training set, the surface disturbed soil sample is placed in a culture dish and compacted to half its volume to obtain a compacted soil sample; the surface disturbed soil sample is compacted until it is ensured that the surface disturbed soil sample has no obvious macropores and no water seepage; Field ground visible and near-infrared spectra are collected at several selected positions on the surface of the compacted soil sample, and the field ground visible and near-infrared spectra of each position are averaged to obtain the original field ground visible and near-infrared spectrum of the surface disturbed soil sample.
4. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 1, characterized in that: The collected original field ground visible and near-infrared spectra are preprocessed separately, including: For any original field ground visible near-infrared spectrum collected, absorbance conversion, multivariate scattering correction and Savitzky-Golay convolution smoothing are performed in sequence to obtain the preprocessed field ground visible near-infrared spectrum.
5. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 1, characterized in that: The principal component analysis was used to reduce the dimension of the pre-processed field ground visible near-infrared spectra and extract the principal component data of the field ground visible near-infrared spectra of each surface disturbed soil sample, including: For the field visible near-infrared spectrum of any surface disturbed soil sample, when performing principal component analysis dimensionality reduction, the characteristic spectrum with a cumulative contribution rate greater than 85% is extracted from the field visible near-infrared spectrum of the surface disturbed soil sample as the principal component data of the field visible near-infrared spectrum of the surface disturbed soil sample.
6. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 1, characterized in that: For the selected training set sampling points and the validation set sampling points, data collection of disturbed soil samples at different depths of the profile is performed to obtain original field ground visible near-infrared spectra and soil organic carbon data at different depths of the profile of each sampling point, specifically including: For any sampling point, disturbed soil samples at different depths of the profile are collected to obtain disturbed soil samples at different depths of the profile of the sampling point; The sampling points are training set sampling points or validation set sampling points; For the disturbed soil sample at any depth of the sampling point profile, the disturbed soil sample is divided into two sub-samples; the two sub-samples of the disturbed soil sample are used for determination of soil organic carbon data and collection of original field ground visible near-infrared spectra, respectively; For a subsample of the disturbed soil sample, measuring soil organic carbon data according to the subsample to obtain soil organic carbon data of the disturbed soil sample; For another subsample of the disturbed soil sample, the subsample was placed in a culture dish and compacted to half its volume to obtain a compacted soil sample; after compaction to ensure that there were no obvious macropores and no water seepage, Field ground visible near-infrared spectra are collected at several selected positions on the surface of the compacted soil sample, and the field ground visible near-infrared spectra of each position are averaged to obtain the original field ground visible near-infrared spectrum of the disturbed soil sample.
7. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 1, characterized in that: Based on the latitude and longitude of the selected training set sampling points and the soil organic carbon data, the optimal triangular grid parameters are determined, including: According to the latitude and longitude of the selected training set sampling points, several sets of different triangular grid parameters are set; the triangular grid parameters include the minimum distance and maximum distance between each training set sampling point and the distance between the training set sampling point and the boundary of the study area; According to the soil organic carbon data, several groups of different triangular grid parameters are screened according to the general information criterion to obtain the optimal triangular grid parameters.
8. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 7, characterized in that: The general information criterion is as follows: Among them, Y i is the true value of soil organic carbon at the i-th sample point; θ i is the predicted value of soil organic carbon at the i-th sample point; n is the total number of sample points; S is the number of sample points in the triangular grid; P(|) represents the conditional probability of the true value of soil organic carbon at each triangular grid sample point; Var() represents the variance function.
9. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 1, characterized in that: Solve the SPDE model according to the following formula: Where ξ(s) is the discretized Gaussian Markov random field, n=1,…N, N is the total number of vertices of the triangular grid; is the piecewise linear basis function of the triangular grid; ξ n is the weight of the n-th vertex of the triangular grid under Gaussian distribution.
10. The method for three-dimensional mapping of soil organic carbon by integrating ground visible and near-infrared spectra according to claim 1, characterized in that: The three-dimensional spatial prediction of soil organic carbon is performed according to the following formula: y(s i )=β0+β1×VNIR(s i )+ξ(s i )+ε i ; Among them, y(s i ) is the soil organic carbon content at the i-th sample point; β0 is the intercept; VNIR(s i ) is the principal component data of the field ground visible near-infrared spectrum at the i-th sample point; β1 is the coefficient of the VNIR environmental covariate; ξ(s i ) is the spatial random effect, i.e., the discretized Gaussian Markov random field; ε i It is Gaussian white noise that is independent of spatial position.
Citation Information
Patent Citations
Plain area soil organic carbon three-dimensional space distribution simulation method
CN110276160A
Soil organic carbon density mapping method and system based on deep learning algorithm, computer equipment and storage medium
CN114782584A
Method for acquiring spatial distribution diagram of organic carbon in farmland soil
CN115561432A
Fine three-dimensional soil mapping method fusing finite profile and surface soil sampling points
CN115908735A
Soil condition analysis system and process
US20180188225A1