A method for mapping soil organic carbon in three-dimensional space by fusing ground visible near-infrared spectrum

By fusing ground-based visible and near-infrared spectroscopy with the INLA-SPDE model and optimizing sampling points and triangular grid parameters, the high cost and low efficiency of traditional soil organic carbon prediction were solved, achieving high-precision three-dimensional spatial prediction and uncertainty analysis of soil organic carbon.

CN120543791BActive Publication Date: 2025-11-18CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510630038.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-11-18
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Traditional methods for predicting soil organic carbon are costly and time-consuming, failing to meet the demand for high-density and timely soil information acquisition. Existing three-dimensional mapping technologies are insufficient in terms of accuracy and efficiency, especially the INLA-SPDE model, which lacks constraints in the selection of environmental covariates.

Method used

By combining ground-based visible and near-infrared spectroscopy with principal component analysis and Kennard-Stone algorithm to optimize sampling points, an INLA-SPDE model was constructed. The Gaussian Markov random field was solved using the finite element method to predict soil organic carbon in three dimensions. By combining Matérn covariance function and triangular grid parameter optimization, the posterior distribution estimation and uncertainty mapping of soil organic carbon were achieved.

Benefits of technology

It improves the accuracy and efficiency of three-dimensional spatial prediction of soil organic carbon, reduces data acquisition costs and manpower and material resources, enhances the model's ability to predict soil depth information, and provides operability and uncertainty analysis of soil organic carbon content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543791B_ABST
    Figure CN120543791B_ABST
Patent Text Reader

Abstract

The application discloses a kind of soil organic carbon three-dimensional space mapping methods of fusing ground visible near infrared spectrum, it is related to soil property prediction technical field, in the method: based on the range of study area setting sampling point, extracts training set sampling point surface field ground visible near infrared spectrum principal component data, application Kennard-stone algorithm optimization reduces training set sampling point to one fourth, then obtain the field ground visible near infrared spectrum and soil organic carbon data of different depth of sampling point profile, and determine the best triangular grid parameter;Then construct SPDE model and solve using finite element method, make Gaussian Markov random field as spatial random effect, the principal component data of field ground visible near infrared spectrum of different depth of each sampling point profile as environmental covariate, application INLA model is carried out soil organic carbon three-dimensional space prediction, obtain the posterior distribution estimate, and soil organic carbon three-dimensional space mapping is carried out accordingly.The application can quickly obtain field ground visible near infrared spectrum, improve data acquisition efficiency, while fusing field ground visible near infrared spectrum principal component data, improve model prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of soil property prediction, in particular to a soil organic carbon three-dimensional space mapping method fusing ground visible near-infrared spectrum. BACKGROUND

[0002] Soil is an important component of the earth's land surface and an important guarantee for crop growth and human activities. Among them, the soil property determines the soil function value, and reasonable evaluation of the spatial distribution of the soil property plays an important role in the implementation of precision agriculture. Soil organic carbon (SOC) usually exists in the form of organic matter in soil and is an important component of soil. SOC can improve soil quality by increasing soil aggregate structure, thereby ensuring the aeration of soil pores and the role of water retention, and is one of the key elements affecting soil physical and chemical properties. However, the traditional SOC prediction method is high in cost and time-consuming and laborious, and has been unable to meet the needs of high-density and high-time-efficiency soil information acquisition. With the advancement of technology, digital soil mapping (DSM) has been widely used as a high-precision and high-quality mapping technology. Conventional digital soil mapping is drawn on a two-dimensional plane based on the collection of soil surface information, and cannot consider the SOC changes at different depths, so how to quickly realize the three-dimensional mapping of SOC is the key to current digital soil mapping.

[0003] In recent years, common three-dimensional mapping techniques include 3D-IDW and INLA-SPDE. 3D-IDW, as a commonly used three-dimensional mapping model, can realize the rapid visualization of three-dimensional mapping, but it is generally low in mapping accuracy because it cannot increase environmental covariates. The INLA-SPDE model, as a Bayesian spatial model, can fully consider the uncertainty of the prior information of the model parameters compared with non-Bayesian spatial models by embedding the horizontal spatial location, the vertical depth and the variation trend of the environmental covariates, thereby improving the spatial prediction accuracy of the soil organic carbon. However, in the application of the INLA-SPDE model in the prior art, the selection of the environmental covariates is not limited, so that the prior art can overcome the defects of the 3D-IDW three-dimensional mapping model to a certain extent, but still has certain improvement space. SUMMARY

[0004] The application aims to provide a soil organic carbon three-dimensional space mapping method fusing ground visible near-infrared spectrum, which can improve the accuracy and efficiency of the three-dimensional space prediction of soil organic carbon.

[0005] To achieve the above object, the application provides the following scheme.

[0006] The application provides a soil organic carbon three-dimensional space mapping method fusing ground visible near-infrared spectrum, comprising the following steps:

[0007] According to the range of the research area, a plurality of training set sampling points and a plurality of validation set sampling points are set, and surface disturbed soil samples are collected at each training set sampling point to obtain original field ground visible near-infrared spectrum of the surface disturbed soil samples.

[0008] Each collected original field ground visible near-infrared spectrum is pretreated, and principal component analysis is used to reduce the dimension of the pretreated field ground visible near-infrared spectrum, so as to extract principal component data of the field ground visible near-infrared spectrum of each surface disturbed soil sample.

[0009] According to the principal component data of the field ground visible near-infrared spectrum of each surface disturbed soil sample, Kennard-stone algorithm is used to optimize and screen the training set sampling points, so as to obtain a plurality of screened training set sampling points; and Kennard-stone algorithm is used to reduce the number of training set sampling points to one fourth of the original number.

[0010] For the plurality of screened training set sampling points and the plurality of validation set sampling points, profile disturbed soil sample data at different depths is collected to obtain original field ground visible near-infrared spectrum and soil organic carbon data at different depths of each sampling point, and the sampling point is the screened training set sampling point or the validation set sampling point.

[0011] Each collected original field ground visible near-infrared spectrum at different depths of each sampling point is pretreated, and principal component analysis is used to reduce the dimension of the pretreated field ground visible near-infrared spectrum at different depths of each sampling point, so as to extract principal component data of the field ground visible near-infrared spectrum at different depths of each sampling point.

[0012] According to the longitude and latitude of the plurality of screened training set sampling points and the soil organic carbon data, the best triangular grid parameter is determined.

[0013] Based on the best triangular grid parameter, a Matérn covariance function is set to construct an SPDE model, and a finite element method is used to solve the SPDE model to obtain a discretized Gaussian Markov random field.

[0014] The Gaussian Markov random field is taken as a spatial random effect, principal component data of field ground visible near-infrared spectrum of each sampling point profile at different depths is taken as an environmental covariate, and soil organic carbon data of each sampling point profile at different depths is taken as a model input, and INLA model is applied to three-dimensional spatial prediction of soil organic carbon to obtain posterior distribution estimation value of soil organic carbon.

[0015] Based on the posterior distribution mean value, soil organic carbon content spatial distribution mapping is carried out, and based on the posterior distribution standard deviation value, soil organic carbon content uncertainty result mapping is carried out.

[0016] Optionally, according to the range of the study area, a plurality of training set sampling points and validation set sampling points are set, and surface disturbed soil samples are collected at each training set sampling point, specifically including the following steps:

[0017] The size of the study area is a field scale, a plurality of training set sampling points are set in the study area by using a grid point method, and a plurality of validation set sampling points are set in the study area by using a random sampling method.

[0018] Optionally, the original field ground visible near-infrared spectrum of each surface disturbed soil sample is obtained, specifically including the following steps:

[0019] For the surface disturbed soil sample of any training set sampling point, the collected surface disturbed soil sample is loaded into a culture dish for compaction to half the volume to obtain a compacted soil sample; the surface disturbed soil sample is compacted to ensure that the surface disturbed soil sample has no obvious large pores and no water seepage;

[0020] The field ground visible near-infrared spectrum is collected at optional positions on the surface of the compacted soil sample, and the field ground visible near-infrared spectrum at each position is averaged to obtain the original field ground visible near-infrared spectrum of the surface disturbed soil sample.

[0021] Optionally, the collected original field ground visible near-infrared spectrum is pretreated, including: for any original field ground visible near-infrared spectrum collected, absorbance conversion, multivariate scatter correction and Savitzky-Golay convolution smoothing processing are sequentially performed to obtain pretreated field ground visible near-infrared spectrum.

[0022] Optionally, principal component analysis is used to reduce the dimension of the pretreated field ground visible near-infrared spectrum, and principal component data of the field ground visible near-infrared spectrum of each surface disturbed soil sample is extracted, including:

[0023] For the field ground visible near-infrared spectrum of each surface disturbed soil sample, when performing principal component analysis dimension reduction, the characteristic spectrum with a cumulative contribution rate greater than 85% of the field ground visible near-infrared spectrum of each surface disturbed soil sample is extracted as the principal component data of the field ground visible near-infrared spectrum of the surface disturbed soil sample.

[0024] Optionally, for the screened training set sampling points and the screened verification set sampling points, the profile different depth disturbed soil sample data of each sampling point is collected to obtain the original field ground visible near-infrared spectrum and the soil organic carbon data of the profile different depth of each sampling point, and the method comprises the following steps:

[0025] For any sampling point, the profile different depth disturbed soil sample of the sampling point is collected to obtain the profile different depth disturbed soil sample of the sampling point; the sampling point is a training set sampling point or a verification set sampling point;

[0026] For the profile any depth disturbed soil sample of the sampling point, the profile any depth disturbed soil sample is divided into two sub-samples; the two sub-samples of the profile any depth disturbed soil sample are used for the determination of the soil organic carbon data and the collection of the original field ground visible near-infrared spectrum, respectively;

[0027] For one sub-sample of the profile any depth disturbed soil sample, the soil organic carbon data of the sub-sample is determined to obtain the soil organic carbon data of the profile any depth disturbed soil sample;

[0028] For another sub-sample of the profile any depth disturbed soil sample, the sub-sample is loaded into a culture dish and compacted to half the volume to obtain a compacted soil sample; after the compacted soil sample is compacted to ensure that there are no obvious large pores and no water seepage, the field ground visible near-infrared spectrum is collected at optional positions on the surface of the compacted soil sample, respectively, and the field ground visible near-infrared spectrum at each position is averaged to obtain the original field ground visible near-infrared spectrum of the profile any depth disturbed soil sample.

[0029] Optionally, the collected original field ground visible near-infrared spectrum of the profile different depth of each sampling point is pretreated, and the pretreatment specifically comprises: for the original field ground visible near-infrared spectrum of any profile different depth disturbed soil sample collected, the absorbance conversion, the multivariate scatter correction and the Savitzky-Golay convolution smoothing processing are sequentially performed to obtain the pretreated field ground visible near-infrared spectrum.

[0030] Optionally, the field ground visible near-infrared spectrum of each sampling point profile at different depths is reduced in dimension by principal component analysis after pretreatment, and the principal component data of the field ground visible near-infrared spectrum of each sampling point profile at different depths is extracted, specifically including: when performing principal component analysis dimension reduction, the characteristic spectrum with a cumulative contribution rate greater than 85% is extracted from the field ground visible near-infrared spectrum of each profile at different depths as the principal component data of the field ground visible near-infrared spectrum of each sampling point profile at different depths.

[0031] Optionally, the optimal triangular grid parameters are determined according to the longitude and latitude of the screened several training set sampling points and the soil organic carbon data, specifically including:

[0032] According to the longitude and latitude of the screened several training set sampling points, a plurality of different triangular grid parameters are set; the triangular grid parameters include the minimum distance between each training set sampling point, the maximum distance and the distance between the training set sampling point and the boundary of the study area;

[0033] According to the soil organic carbon data, the plurality of different triangular grid parameters are screened according to the general information amount criterion to obtain the optimal triangular grid parameters.

[0034] Optionally, the general information amount criterion is as follows:

[0035]

[0036] Wherein, Y i is the true value of the soil organic carbon of the i th sample point; θ i is the predicted value of the soil organic carbon of the i th sample point; n is the total number of sample points; S is the number of sample points of the triangular grid; P (|) represents the conditional probability of the true value of the soil organic carbon at each triangular grid sample point; Var () represents the variance value function.

[0037] Optionally, the SPDE model is solved according to the following formula:

[0038]

[0039] Wherein, ξ(s) is a discretized Gaussian Markov random field, n = 1,…N, N is the total number of triangular grid vertices; is a triangular grid piecewise linear basis function; ξ n is the weight of the n th vertex of the triangular grid under the Gaussian distribution.

[0040] Optionally, the soil organic carbon three-dimensional space prediction is performed according to the following formula:

[0041] y(s i )=β0+β1×VNIR(s i )+ξ(si )+ epsilon i .

[0042] where y(s i ) is the soil organic carbon content at the i-th sample point; beta0 is the intercept; VNIR(s i ) is the principal component data of the field ground visible near-infrared spectrum at the i-th sample point; beta1 is the coefficient of the VNIR environmental covariate; xi(s i ) is the spatial random effect, i.e. the discretized Gaussian Markov random field; epsilon i is the Gaussian white noise irrelevant to the spatial position.

[0043] According to the specific embodiments provided in the application, the following technical effects are disclosed:

[0044] The application provides a soil organic carbon three-dimensional space mapping method fusing ground visible near-infrared spectrum, in the method, according to the range of the study area, the training set and the validation set sampling points are set, the surface disturbed soil samples of the training set of the study area are obtained, the original field ground visible near-infrared spectrum is collected and pretreated, and then the principal component analysis is applied to the principal component analysis dimension reduction of the pretreated field ground visible near-infrared spectrum; based on the extracted field ground visible near-infrared spectrum principal component data, the KS algorithm is used to optimize and screen the training set, and the number of training set sampling points is reduced to one fourth of the original training set; based on the screened data set, the original field ground visible near-infrared spectrum and the soil organic carbon data of each sampling point profile at different depths are obtained, the original field ground visible near-infrared spectrum of each sampling point profile at different depths is pretreated and the principal component data is extracted, and the best triangular grid parameters are determined; then based on the best triangular grid parameters, the SPDE model is constructed by setting the Matern covariance function, and the discretized Gaussian Markov random field is obtained by using the finite element method, then the Gaussian Markov random field is used as a spatial random effect, the principal component data of the field ground visible near-infrared spectrum is used as an environmental covariate, and the INLA model is applied to the three-dimensional space prediction of the soil organic carbon to obtain the posterior distribution estimation value of the soil organic carbon, and finally the soil organic carbon content spatial distribution mapping and the soil organic carbon content uncertainty result mapping are carried out. Compared with the traditional way of obtaining soil organic carbon by digging profiles and laboratory testing, the application combines near-surface sensing technology, on the one hand, the field ground visible near-infrared spectrum data of the surface disturbed soil sample is collected, the KS algorithm is used to optimize and screen the training set to one fourth of the original, the data acquisition efficiency is improved, on the other hand, the field ground visible near-infrared spectrum which can reflect the related spectral characteristics of soil organic carbon is fused, which not only ensures the improvement of precision, but also provides guarantee for the operability of subsequent technology popularization; on the other hand, in the application, the INLA-SPDE model is used, the environmental covariate can be flexibly added, and the correlation between each point in the study area can be considered by establishing and screening the triangular grid of the sampling points, so as to further improve the prediction accuracy of the model. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0046] Figure 1 A flow chart of a soil organic carbon three-dimensional space mapping method fusing ground visible near-infrared spectrum provided by an embodiment of the application.

[0047] Figure 2 A schematic diagram of SPDE spatial effect in a method for mapping soil organic carbon in three-dimensional space by fusing ground visible near-infrared spectrum according to an embodiment of the present application.

[0048] Figure 3 A schematic diagram of predicted distribution of soil organic carbon in a method for mapping soil organic carbon in three-dimensional space by fusing ground visible near-infrared spectrum according to an embodiment of the present application.

[0049] Figure 4 A schematic diagram of predicted standard deviation value distribution in a method for mapping soil organic carbon in three-dimensional space by fusing ground visible near-infrared spectrum according to an embodiment of the present application. DETAILED DESCRIPTION

[0050] The technical solutions in the embodiments of the present application will be apparently and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without any creative work fall within the scope of protection of the present application.

[0051] The above purposes, features and advantages of the present application will be more apparent and easy to understand. The present application will be further described in detail below with reference to the drawings and specific embodiments.

[0052] A method for mapping soil organic carbon in three-dimensional space by fusing ground visible near-infrared spectrum according to an embodiment of the present application includes the following steps in an exemplary embodiment, as shown in Figure 1

[0053] A1, according to the range of the study area, a plurality of training set sampling points and a plurality of validation set sampling points are set, and surface disturbed soil samples are collected at each training set sampling point to obtain the original field ground visible near-infrared spectrum of each surface disturbed soil sample. In the present embodiment, the means of "setting a plurality of training set sampling points and a plurality of validation set sampling points according to the range of the study area" in step A1 specifically includes the following steps:

[0054] A11, the size of the study area is field scale, a plurality of training set sampling points are set in the study area by using grid point method, and a plurality of validation set sampling points are set in the study area by using random sampling method. In the present embodiment, the training set is arranged at intervals of 100m x 100m by using grid point method, and a total of 69 sampling points are set; the validation set is arranged by using random sampling method, and a total of 10 sampling points are set.

[0055] ​A12, collect the surface disturbed soil sample of the training set sampling point, in this embodiment, collect the soil sample of 0-10 cm as the surface disturbed soil sample.

[0056] The means of "obtaining the original field ground visible near-infrared spectrum of each surface disturbed soil sample" in step A1 specifically includes the following steps:

[0057] A13, for the surface disturbed soil sample of any training set sampling point, put the collected surface disturbed soil sample into a culture dish for compaction to half the volume, to obtain a compacted soil sample; compact the surface disturbed soil sample to ensure that the surface disturbed soil sample has no obvious large pores and no water seepage;

[0058] A14, collect the field ground visible near-infrared spectrum at optional positions on the surface of the compacted soil sample, and obtain the original field ground visible near-infrared spectrum of the surface disturbed soil sample by averaging the field ground visible near-infrared spectrum at each position.

[0059] Specifically, the soil sample should be put into a culture dish and compacted to half the volume to ensure that the soil sample has no obvious large pores and no water seepage. The visible near-infrared (Visible-Near Infrared, VNIR) spectrum is collected at a position on the surface of the soil sample without obvious gaps and impurities. The Quality Spec Trek portable spectrometer is used to measure the VNIR spectrum, and the spectral range is 350-2500 nm. When measuring, three positions are selected for each soil sample, and three spectra are measured at each position. The average of the nine spectra obtained is taken as the field ground visible near-infrared spectrum data of the disturbed soil sample at the point.

[0060] A2, respectively, pretreat each original field ground visible near-infrared spectrum, and use principal component analysis to reduce the dimension of the pretreated field ground visible near-infrared spectrum, and extract the principal component data of the field ground visible near-infrared spectrum of each surface disturbed soil sample.

[0061] In this embodiment, "respectively, pretreat each original field ground visible near-infrared spectrum" in step A2 includes: for any original field ground visible near-infrared spectrum collected, sequentially perform absorbance conversion, multivariate scatter correction and Savitzky-Golay convolution smoothing processing to obtain the pretreated field ground visible near-infrared spectrum. Before pretreating the spectrum data, remove the noise of 350-450 nm and 2400-2500 nm.

[0062] In the present embodiment, the step A2 of "using principal component analysis to reduce the dimension of the pre-processed field ground visible near-infrared spectrum, and extracting the principal component data of the field ground visible near-infrared spectrum of each surface disturbed soil sample" includes: for the field ground visible near-infrared spectrum of any surface disturbed soil sample, when performing principal component analysis dimension reduction, extracting the characteristic spectrum with a cumulative contribution rate of more than 85% from the field ground visible near-infrared spectrum of the surface disturbed soil sample as the principal component data of the field ground visible near-infrared spectrum of the surface disturbed soil sample.

[0063] A3, according to the principal component data of the field ground visible near-infrared spectrum of each surface disturbed soil sample, using Kennard-stone algorithm to optimize and screen the training set sampling points, obtaining a plurality of screened training set sampling points; using Kennard-stone algorithm, reducing the number of training set sampling points to one fourth of the original number.

[0064] In the present embodiment, the step A3 of "using Kennard-stone algorithm to optimize and screen the training set sampling points, and obtaining a plurality of screened training set sampling points" specifically selects the principal component data of the field ground visible near-infrared spectrum of the surface (0-10cm) disturbed soil sample after principal component analysis, applies Kennard-stone algorithm to optimize and screen the existing 69 training set sampling points, and a total of 17 sampling points are screened as new training set sampling points.

[0065] A4, for the plurality of screened training set sampling points and a plurality of said validation set sampling points, respectively collecting the data of the disturbed soil samples at different depths of the profile, and obtaining the original field ground visible near-infrared spectrum and soil organic carbon data at different depths of the profile of each sampling point.

[0066] In the present embodiment, the step A4 of "for the plurality of screened training set sampling points and a plurality of said validation set sampling points, respectively collecting the data of the disturbed soil samples at different depths of the profile, and obtaining the original field ground visible near-infrared spectrum and soil organic carbon data at different depths of the profile of each sampling point" specifically includes the following steps:

[0067] A41, for any sampling point, collecting the disturbed soil samples of different depths of the profile. In this embodiment, at each sampling point, the soil samples of different depths are collected from the ground surface, with each 10 cm as a layer, and the depth reaching 1 m. The first layer is 0-10 cm, the second layer is 10-20 cm, the third layer is 20-30 cm, the fourth layer is 30-40 cm, the fifth layer is 40-50 cm, the sixth layer is 50-60 cm, the seventh layer is 60-70 cm, the eighth layer is 70-80 cm, the ninth layer is 80-90 cm, and the tenth layer is 90-100 cm. A total of 10 layers of disturbed soil samples are collected.

[0068] A42, the sampling points are training set sampling points or validation set sampling points. In this embodiment, there are 17 training set sampling points and 10 validation set sampling points.

[0069] A43, for the disturbed soil sample of any depth of the profile of the sampling point, the disturbed soil sample is divided into two sub-samples; the two sub-samples of the disturbed soil sample are respectively used for the determination of soil organic carbon data and the collection of original field ground visible near-infrared spectrum. Specifically, each layer of the disturbed soil sample of each sampling point is evenly divided into two parts, one part is used for the determination of soil organic carbon data; and the other part is sealed in a sealed bag to keep the moisture content unchanged, and is used for the collection of ground visible near-infrared spectrum data.

[0070] A44, for one sub-sample of the disturbed soil sample, the determination of soil organic carbon data is performed according to the sub-sample, and the soil organic carbon data of the disturbed soil sample of different depths of the profile is obtained.

[0071] A45, for the other sub-sample of the disturbed soil sample, the sub-sample is loaded into a culture dish for compaction to half the volume, and a compacted soil sample is obtained; after compaction to ensure that there are no obvious large pores and no water seepage, the surface of the compacted soil sample is optionally selected at several positions for the collection of field ground visible near-infrared spectrum, and the field ground visible near-infrared spectrum of each position is averaged to obtain the original field ground visible near-infrared spectrum of the disturbed soil sample.

[0072] Specifically, the soil sample should be loaded into a culture dish for compaction to half the volume, and the compaction should ensure that there are no obvious large pores and no water seepage. The positions on the surface of the soil sample without obvious gaps and impurities are selected for the collection of visible near-infrared spectrum. The Quality Spec Trek portable spectrometer is used to measure the visible near-infrared spectrum, and the spectral range is 350-2500 nm. When measuring, three positions are selected for each soil sample, and three spectra are repeatedly measured at each position. The average of the 9 spectra obtained is taken as the original data of the ground visible near-infrared spectrum of the soil sample.

[0073] A5, preprocessing the original field visible near-infrared spectrum of each sampling point profile at different depths, using principal component analysis to reduce the dimension of the preprocessed field visible near-infrared spectrum of each sampling point profile at different depths, and extracting the principal component data of the field visible near-infrared spectrum of each sampling point profile at different depths.

[0074] In the embodiment, the "preprocessing the original field visible near-infrared spectrum of each sampling point profile at different depths" in step A5 includes: for the original field visible near-infrared spectrum of the disturbed soil sample at different depths of any profile collected, sequentially performing absorbance conversion, multivariate scatter correction and Savitzky-Golay convolution smoothing processing to obtain the preprocessed field visible near-infrared spectrum. Before preprocessing the spectrum data, remove the noise at 350-450 nm and 2400-2500 nm.

[0075] In the embodiment, the "using principal component analysis to reduce the dimension of the preprocessed field visible near-infrared spectrum of each sampling point profile at different depths, and extracting the principal component data of the field visible near-infrared spectrum of each sampling point profile at different depths" in step A5 includes: when performing principal component analysis dimension reduction, extracting the characteristic spectrum with a cumulative contribution rate greater than 85% from the field visible near-infrared spectrum at different depths of each profile as the principal component data of the field visible near-infrared spectrum at different depths of each sampling point profile.

[0076] A6, determining the best triangular grid parameters according to the longitude and latitude of the screened several training set sampling points and the soil organic carbon data. In the embodiment, step A6 specifically includes the following steps:

[0077] A61, setting several groups of different triangular grid parameters according to the longitude and latitude of the screened several training set sampling points; the triangular grid parameters include the minimum distance between the training set sampling points, the maximum distance and the distance between the training set sampling points and the boundary of the study area.

[0078] A62, screening several groups of different triangular grid parameters according to the soil organic carbon data according to the general information content criterion to obtain the best triangular grid parameters.

[0079] Specifically, based on the Kennard-stone algorithm, the longitude and latitude coordinates of 17 sampling points and the spatial range data of the study area are obtained, the minimum distance, maximum distance and distance between the sampling points and the boundary of the study area are calculated, the constraint subdivision Delaunay triangular mesh (MESH) parameters of the model are set according to the above results, and the triangular mesh is screened according to the widely available information criterion (WAIC). The smaller the WAIC value, the higher the fitting degree of the model, and the triangular mesh parameters of the model are optimal. The specific formula of the widely available information criterion is as follows:

[0080]

[0081] Y i is the true value of soil organic carbon of the i th sample point; θ i is the predicted value of soil organic carbon of the i th sample point; n is the total number of sample points; S is the number of triangular mesh sample points; P (|) represents the conditional probability of the true value of soil organic carbon at each triangular mesh sample point; Var () represents the variance value function.

[0082] By comparing the WAIC of triangular mesh parameters under a plurality of different combinations, as shown in Table 1.

[0083] Table 1 Triangular mesh parameters and WAIC values under different combinations

[0084] max.edge offset cutoff triangulation total number of vertices WAIC 100,150 40 25,150 MESH1 327 -983.17 70,150 40 25,150 MESH2 456 -1037.92 70,150 10 25,150 MESH3 569 -958.09 70,150 10 50,200 MESH4 602 -969.17

[0085] The triangular mesh parameters mainly include max.edge, cutoff and offset. Max.edge is the maximum and minimum edge length of the triangular mesh; cutoff is used to set the correlation distance between two points. When the distance between two points of the triangular mesh is greater than the set value, the two points do not affect each other; offset is used to control the triangular mesh layout outside the study area. In Table 1, the WAIC value of MESH2 is the lowest, and the WAIC value increases after the total number of triangular vertices increases, so MESH2 is selected as the optimal triangular mesh of the study area, which can ensure that the model parameters meet the suitability of the study area.

[0086] A7, based on the optimal triangular mesh parameters, a Matérn covariance function is set to construct a SPDE model, and a finite element method is used to solve the SPDE model to obtain a discretized Gaussian Markov random field. Specifically, as Figure 2The diagram shown illustrates the spatial effects of SPDE. The SPDE model links the spatially continuous Matérn GF with a discretized Gaussian Markov random field. The SPDE model is solved according to the following equation:

[0087]

[0088] Where ξ(s) is a discretized Gaussian Markov random field, n = 1, ..., N, where N is the total number of vertices of the triangular grid; It is a piecewise linear basis function of a triangular grid; ξ n It is the weight of the nth vertex of the triangular grid under a Gaussian distribution.

[0089] A8. Using the Gaussian Markov random field as a spatial random effect, the principal component data of the visible and near-infrared spectra of the ground at different depths of the sampling point profile are used as environmental covariates, and the soil organic carbon data at different depths of the sampling point profile are used as model inputs. The INLA model is applied to predict the three-dimensional spatial distribution of soil organic carbon, obtaining the posterior distribution estimate. The posterior distribution estimate includes the posterior distribution mean and the posterior distribution standard deviation. In this embodiment, the three-dimensional spatial prediction of soil organic carbon is performed according to the following formula:

[0090] y(s i )=β0+β1×VNIR(s i )+ξ(s i )+ε i .

[0091] Wherein, y(s) i ) represents the soil organic carbon content at the i-th sample point; β0 is the intercept; VNIR(s) i ) represents the principal component data of the visible and near-infrared spectra of the ground at the i-th sample point; β1 represents the coefficients of the VNIR environmental covariates; ξ(s) i ) represents spatial random effects, i.e., discretized Gaussian Markov random fields; ε i It is Gaussian white noise that is independent of spatial location.

[0092] For the above predictions, accuracy evaluation indicators are selected to comprehensively assess the model's prediction accuracy. These accuracy evaluation indicators include the coefficient of determination (R²). 2 ), root mean square error (RMSE), and relative analysis error (RPD). Among them, R... 2 The larger the RPD and the smaller the RMSE, the better the model's predictive performance and the stronger its predictive ability.

[0093] The proposed INLA-SPDE model, which integrates terrestrial visible and near-infrared spectral data, was compared with the conventional 3D mapping method 3D-IDW. The results are shown in Table 2.

[0094] Table 2 Comparison of prediction accuracy between INLA-SPDE model and 3D-IDW model

[0095]

[0096] 3D-IDW mapping was performed using a conventional mesh generation method, with a mapping accuracy of R0. 2 =0.62, RMSE=1.98, RPD=1.63, indicating good simulation results; however, when using the 17 training sets selected after KS partitioning, the 3D-IDW model cannot guarantee mapping accuracy due to the spatial heterogeneity reflected by the number of sampling points. 2 =0.07, RMSE=3.10, RPD=1.04). This application, by applying the INLA-SPDE model after fusing terrestrial visible and near-infrared spectra, maintains high mapping accuracy after KS partitioning, compared to the 3D-IDW results. Specifically, compared to the results using the same KS partitioning, the accuracy R... 2 The accuracy improved by 0.63, RMSE decreased by 1.32, and RPD improved by 0.78; compared to the results of retaining 69 training sets, the accuracy R... 2 The accuracy improved by 0.08, RMSE decreased by 0.20, and RPD improved by 0.19. Therefore, applying the INLA-SPDE model based on the ground visible and near-infrared spectrum provided in this application can ensure good mapping accuracy and reflect the information in the current soil environmental elements; it can also reduce the number of sampling points in the current mapping, thereby saving the manpower and material resources wasted on data acquisition.

[0097] A9. Spatial distribution mapping of soil organic carbon content is performed based on the posterior distribution mean, and the uncertainty of soil organic carbon content is mapped based on the posterior distribution standard deviation. Specifically, the posterior distribution estimates obtained based on the INLA-SPDE model are used in ArcGIS 10.8 software to map the spatial distribution of soil organic carbon content based on the posterior distribution mean, such as... Figure 3 As shown; a graph is plotted to illustrate the uncertainty of soil organic carbon content based on the posterior distribution standard deviation, as shown below. Figure 4 As shown.

[0098] The solutions provided in the above embodiments of this application have the following characteristics compared to traditional technologies:

[0099] Firstly, the traditional research method usually uses grid sampling method with fixed length as sampling interval to obtain the training set. This method is easier to arrange when the spatial range of the study area is at the field scale. However, when the area of the study area is too large, the number of sampling points is often too large, which is easy to cause the loss of manpower and material resources. In this application, the principal component data of the ground visible near-infrared spectrum of the disturbed soil in the surface layer (0-10 cm) is screened by using the KS algorithm, and the number of training set is reduced from 69 sampling points to 17 sampling points, and the sample screening rate reaches 75%. Only by collecting the surface layer disturbed soil sample, the soil sampling point can be determined, and a large amount of manpower and material resources can be saved. In addition, the application of KS divided sampling point attribute ensures the significant difference between the sampling points, and realizes the feasibility of predicting soil organic carbon under the condition of sharp reduction of sampling point data. When the precision of INLA-SPDE model is compared with that of 3D-IDW, the best can still be maintained, which further verifies the advantage of INLA-SPDE model in predicting soil organic carbon content.

[0100] Secondly, this application integrates the principal component data of the ground visible near-infrared spectrum as the environmental covariate, effectively considers the complex relationship between soil spatial random effect and environmental variables, and improves the prediction accuracy of the model. Compared with the current data selection method of digital soil mapping, many studies combine remote sensing data and terrain data as mapping covariates, but such data can only reflect the surface information, and are often disturbed by image spatial resolution, cloud and other information, which is difficult to realize three-dimensional modeling and prediction of soil depth. This application integrates the ground visible near-infrared spectrum data which can reflect the spectral characteristics of soil organic carbon, on the one hand, improves the prediction of soil depth information by the model, and on the other hand, due to the convenience of technology operation, it also provides guarantee for the operability of subsequent technology popularization.

[0101] Further, by using INLA-SPDE model, on the one hand, the environmental covariates can be flexibly added, and on the other hand, by screening the MESH network of sampling points, the correlation between each point in the study area can be considered, so as to further improve the prediction accuracy of the model.

[0102] Finally, this application combines INLA-SPDE model for three-dimensional spatial prediction of soil organic carbon, that is, the posterior distribution probability is used to obtain the result of soil organic carbon, which enriches the current research progress of soil organic carbon storage at different scales and provides relevant reference, and can also quantify the uncertainty of prediction, which provides support for in-depth analysis of mapping accuracy.

Claims

1. A method for three-dimensional spatial mapping of soil organic carbon by integrating visible and near-infrared spectroscopy, characterized in that, include: Based on the study area, several training set sampling points and several validation set sampling points were set up, and surface disturbed soil samples were collected at each training set sampling point to obtain the original field visible and near-infrared spectra of each surface disturbed soil sample. The original field visible and near-infrared spectra were preprocessed, and principal component analysis was used to reduce the dimensionality of the preprocessed field visible and near-infrared spectra to extract the principal component data of the field visible and near-infrared spectra of each disturbed surface soil sample. Based on the principal component data of the visible and near-infrared spectra of the surface disturbed soil samples, the Kennard-Stone algorithm was used to optimize and screen the training set sampling points, resulting in a number of screened training set sampling points. The Kennard-Stone algorithm was then used to reduce the number of training set sampling points to one-quarter of the original number. For several selected training set sampling points and several selected validation set sampling points, soil sample data with disturbance at different profile depths were collected to obtain the original field visible and near-infrared spectra and soil organic carbon data at different depths of the profile at each sampling point; the sampling points are either selected training set sampling points or validation set sampling points. The raw visible and near-infrared spectra of the field ground at different depths of the sampling points were preprocessed. Principal component analysis was used to reduce the dimensionality of the preprocessed visible and near-infrared spectra of the field ground at different depths of the sampling points, and the principal component data of the visible and near-infrared spectra of the field ground at different depths of the sampling points were extracted. Based on the latitude and longitude of the selected training set sampling points and the soil organic carbon data, the optimal triangular grid parameters are determined. Based on the optimal triangular grid parameters, the SPDE model is constructed by setting the Matérn covariance function, and the SPDE model is solved by the finite element method to obtain the discretized Gaussian Markov random field. Using the Gaussian Markov random field as a spatial random effect, the principal component data of the visible and near-infrared spectra of the field ground at different depths of the sampling points are used as environmental covariates, and the soil organic carbon data at different depths of the sampling points are used as model inputs. The INLA model is applied to predict the three-dimensional spatial distribution of soil organic carbon, and the posterior distribution estimate of soil organic carbon is obtained. The posterior distribution estimate includes the posterior distribution mean and the posterior distribution standard deviation. The spatial distribution of soil organic carbon content is mapped based on the mean of the posterior distribution, and the uncertainty of soil organic carbon content is mapped based on the standard deviation of the posterior distribution.

2. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 1, characterized in that, Based on the study area, several training set sampling points and validation set sampling points were set up, and surface disturbed soil samples were collected at each training set sampling point, specifically including: The study area is at the field scale. A grid-based sampling method was used to set up several training set sampling points in the study area, and a random sampling method was used to set up several validation set sampling points in the study area.

3. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 1, characterized in that, Obtain the original field visible and near-infrared spectra of each disturbed surface soil sample, including: For any surface disturbed soil sample from any training set sampling point, the surface disturbed soil sample is placed in a petri dish and compacted to half its volume to obtain a compacted soil sample; the surface disturbed soil sample is compacted until it is ensured that the surface disturbed soil sample has no obvious large pores and no water seepage. The visible and near-infrared spectra of the surface of the compacted soil sample were collected at several selected locations. The visible and near-infrared spectra of the surface of the sample were then averaged to obtain the original visible and near-infrared spectra of the surface disturbed soil sample.

4. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 1, characterized in that, The raw visible and near-infrared spectra collected from the field were preprocessed, specifically including: For any raw field visible and near-infrared spectrum collected, absorbance conversion, multivariate scattering correction, and Savitzky-Golay convolution smoothing are performed sequentially to obtain the preprocessed field visible and near-infrared spectrum.

5. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 1, characterized in that, Principal component analysis was used to reduce the dimensionality of the preprocessed visible and near-infrared spectra of the field surface, and principal component data of the visible and near-infrared spectra of each disturbed surface soil sample were extracted, specifically including: For any surface disturbed soil sample's visible and near-infrared spectrum, during principal component analysis dimensionality reduction, the characteristic spectra with a cumulative contribution rate greater than 85% are extracted from the surface visible and near-infrared spectrum of the surface disturbed soil sample and used as the principal component data of the surface visible and near-infrared spectrum of the surface disturbed soil sample.

6. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 1, characterized in that, For several selected training set sampling points and several validation set sampling points, data were collected from disturbed soil samples at different depths of the profile. This yielded raw field visible and near-infrared spectra and soil organic carbon data at different depths of the profile at each sampling point. Specifically, this included: For any given sampling point, disturbed soil samples at different depths of the profile are collected to obtain disturbed soil samples at different depths of the profile at the sampling point. The sampling points are either training set sampling points or validation set sampling points; For disturbed soil samples at any depth of the sampling point profile, the disturbed soil samples are divided into two sub-samples; the two sub-samples are used for the determination of soil organic carbon data and the acquisition of raw field visible and near-infrared spectra, respectively. For a subsample of the disturbed soil sample, soil organic carbon data is measured based on the subsample to obtain the soil organic carbon data of the disturbed soil sample. For another subsample of the disturbed soil sample, the subsample was placed in a petri dish and compacted to half its volume to obtain a compacted soil sample; compaction was continued until no obvious large pores were observed and no water seepage occurred. Several locations were randomly selected on the surface of the compacted soil sample to collect visible and near-infrared spectra of the ground. The visible and near-infrared spectra of the ground at each location were averaged to obtain the original visible and near-infrared spectra of the disturbed soil sample.

7. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 1, characterized in that, Based on the latitude and longitude of the selected training set sampling points and the soil organic carbon data, the optimal triangular grid parameters are determined, specifically including: Based on the latitude and longitude of several selected training set sampling points, several different sets of triangular grid parameters are set; the triangular grid parameters include the minimum distance, maximum distance and distance between each training set sampling point and the boundary of the study area; Based on the soil organic carbon data, several different sets of triangular grid parameters were screened according to the general information content criterion to obtain the optimal triangular grid parameters.

8. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 7, characterized in that, The general information content criterion is shown in the following formula: Among them, Y i Let θ be the true value of soil organic carbon at the i-th sample point; i Let be the predicted value of soil organic carbon at the i-th sample point; n is the total number of sample points; S is the number of sample points in the triangular grid; P(|) represents the conditional probability of the true value of soil organic carbon at each triangular grid sample point; Var() represents the variance function.

9. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 1, characterized in that, Solve the SPDE model using the following formula: Where ξ(s) is a discretized Gaussian Markov random field, n = 1, ..., N, where N is the total number of vertices of the triangular grid; It is a piecewise linear basis function of a triangular grid; ξ n It is the weight of the nth vertex of the triangular grid under a Gaussian distribution.

10. The method for three-dimensional spatial mapping of soil organic carbon based on terrestrial visible and near-infrared spectroscopy according to claim 1, characterized in that, The following formula is used to predict the three-dimensional spatial distribution of soil organic carbon: y(s i )=β0+β1×VNIR(s i )+ξ(s i )+ε i ; Wherein, y(s) i ) represents the soil organic carbon content at the i-th sample point; β0 is the intercept; VNIR(s) i ) represents the principal component data of the visible and near-infrared spectra of the ground at the i-th sample point; β1 represents the coefficients of the VNIR environmental covariates; ξ(s) i ) represents spatial random effects, i.e., discretized Gaussian Markov random fields; ε i It is Gaussian white noise that is independent of spatial location.

Citation Information

Patent Citations

  • Plain area soil organic carbon three-dimensional space distribution simulation method

    CN110276160A

  • Soil organic carbon density mapping method and system based on deep learning algorithm, computer equipment and storage medium

    CN114782584A