A soil organic matter spectral prediction method based on texture perception residual superposition

By using the method of texture perception residual superposition, a global main effect and residual model is constructed using texture classification and spectral information. This solves the problem of texture influence in soil organic matter prediction, achieves efficient and accurate prediction on soils with different textures, and improves prediction performance and robustness.

CN121167680BActive Publication Date: 2026-04-10豫章师范学院
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods for predicting soil organic matter by spectral means are susceptible to systematic errors in soils of different textures, resulting in low prediction accuracy and difficulty in rapid and accurate measurement under in-situ field conditions.

Method used

By classifying soil texture into three categories—coarse, medium, and fine—and converting them into dummy variables, and combining them with spectral information to construct a global main effect model and a residual model, predictions can be made directly using texture category and spectral information, avoiding the need for soil texture testing and correcting systematic errors.

Benefits of technology

It enables efficient and accurate prediction of soil organic matter content in soils with different textures, improves the robustness and generalization ability of prediction performance, and maintains the advantages of in-situ soil spectroscopy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167680B_ABST
    Figure CN121167680B_ABST
Patent Text Reader

Abstract

The application discloses a soil organic matter spectrum prediction method based on texture perception residual superposition, which converts the classified texture into a dummy variable and uses the dummy variable for model training, so that when the method is applied, the soil organic matter content can be obtained by directly inputting the texture category and spectrum information into the trained model, without the need for testing the texture, without weakening the advantages of soil in-situ spectrum, and being able to meet the needs of simply and efficiently predicting the soil organic matter content. The method also trains a second model to establish a mapping relationship between different textures and residual prediction values, so that the residual model provided by the application can accurately capture the difference between the baseline prediction value of different textures and the true soil organic carbon content, thereby compared with the prior art, the application can display the systematic error correction, the model prediction performance is more stable, and the soil organic carbon content of more large-scale and highly heterogeneous soil can be accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of soil spectral quantitative analysis and soil quality monitoring, and particularly relates to a soil organic matter spectral prediction method based on texture perception residual superposition. BACKGROUND

[0002] The soil spectral library is an existing resource that has been established. How to use the existing spectral library information to improve the prediction accuracy of regional soil in-situ spectrum is a scientific problem that needs to be solved at present. The existing in-situ soil organic matter quantitative methods based on the spectral library (such as PLS, SVR, RF, etc.) are easily affected by soil texture (sand / silt / clay ratio) changes: under the same spectral form, due to the differences in mineralogy and structure, the model bias has systematic errors in different texture soils. Moreover, the interaction of in-situ soil moisture and soil texture affects the soil reflectance spectrum, resulting in more complex differences in in-situ soil spectrum. The greater the systematic error is, the lower the prediction accuracy is.

[0003] The patent application with the publication number CN118169056A discloses a hyperspectral soil organic matter content prediction method based on DBO-SVR, which includes the following steps: 1) soil sample collection and processing; 2) determining the available potassium content of the soil sample in step 1) and the hyperspectral data of the soil sample; 3) pre-processing the hyperspectral data in step 2) and extracting effective variables by PCA; 4) constructing a support vector machine regression model and a partial least squares regression model; 5) building a DBO-SVR model to train the spectral data; 6) establishing a hyperspectral soil organic matter content prediction model; and 7) model accuracy evaluation. The method can realize real-time, rapid and accurate indoor detection of soil organic matter content, and has sufficient practical application significance. However, the prediction method disclosed in the above patent application is for indoor chemical analysis classification prediction, and the influence of texture on soil organic carbon content is not fully considered.

[0004] CN113436153A discloses a method for predicting carbon components of undisturbed soil profile based on hyperspectral imaging and support vector machine technology. The hyperspectral image of the soil profile sample at a predetermined depth of each sample position is obtained. The characteristic spectral bands of the soil carbon component type corresponding to the target sample spectral region are used as input, and the soil carbon component data corresponding to the soil carbon component type corresponding to the target sample spectral region are used as output. Through training, the soil carbon component prediction model corresponding to the soil carbon component type is obtained, and the prediction of the target region soil profile carbon component is realized. The entire design scheme can quickly and accurately predict the component content of undisturbed soil profile organic carbon, soluble carbon, easily oxidizable carbon and soil microbial biomass carbon, and realize the fine drawing of their spatial distribution in the soil profile. It makes up for the shortcomings of traditional laboratory chemical analysis method. However, the prediction accuracy of the method disclosed in the patent application needs to be further improved, and the soil carbon component classification disclosed in the above patent application still needs indoor chemical analysis, which weakens the advantage of spectrum.

[0005] The traditional scheme either uses indoor drying and grinding spectrum or needs chemical analysis classification to improve precision, which is difficult to measure quickly in the field in situ condition. Either ignore the texture or consider the texture information but ignore the residual error. Moreover, even if the texture information is used, the existing method still needs to test the soil texture (clay, sand and silt content), which weakens the advantage of soil in-situ spectrum. Therefore, there is an urgent need for a new method that can use the auxiliary information of texture correction quickly acquired in the field in-situ condition to improve the generalization precision and robustness in different texture scenes without the need for testing the soil texture. SUMMARY

[0006] The present application provides a soil organic matter spectrum prediction method based on texture-aware residual superposition, which can simply, efficiently and accurately predict the soil organic matter content.

[0007] The present application provides a soil organic matter spectrum prediction method based on texture-aware residual superposition, which can simply, efficiently and accurately predict the soil organic matter content.

[0008] According to the sand, powder and clay content of the soil sample in the soil spectrum library, the texture is classified, the classified texture is converted into a dummy variable, and the training sample set and the verification sample set are constructed based on the spectral information of each soil sample in the soil spectrum library, the dummy variable of the corresponding texture and the true soil organic carbon content.

[0009] A first model is obtained, the training sample set is used, the spectral information and the dummy variable are used as independent variables, and the true soil organic carbon content is used as the dependent variable. The first model is trained to obtain a global main effect model, and the within-fold prediction value is obtained through the global main effect model.

[0010] obtaining a second model, using the training sample set, taking the spectral information and the dummy variable as independent variables, and taking the difference between the fold prediction value and the true soil organic carbon content as the dependent variable, training the second model to obtain a residual model, and obtaining a residual prediction value through the residual model;

[0011] In application, the soil spectral information and the category of texture of the current soil sample are respectively input into the global main effect model to obtain a baseline prediction value, the soil spectral information and the category of texture of the current soil sample and the baseline prediction value are input into the residual model to obtain a residual prediction value, and the final soil organic carbon content prediction value is obtained based on the difference between the baseline prediction value and the residual prediction value.

[0012] Preferably, according to the sand, silt and clay content of the soil sample in the soil spectral library, the corresponding texture is divided into three categories of coarse, medium and fine according to the rules, and the texture of each sample is converted into a three-column binary digital combination dummy variable based on the classification result.

[0013] Preferably, the corresponding texture is divided into three categories of coarse, medium and fine according to the rules, including:

[0014] When the clay content in the soil sample is ≥ 35%, the texture of the soil sample is fine;

[0015] When the sand content in the soil sample is ≥ 65% and the clay content is ≤ 18%, the texture of the soil sample is coarse;

[0016] The rest is medium.

[0017] Preferably, the spectral information of each soil sample and the dummy variable of the corresponding texture are used to construct a feature matrix, and the feature matrix is used as the independent variable for training the first model and the second model.

[0018] Preferably, the step of training the first model comprises:

[0019] Taking the spectral information and the dummy variable as independent variables and the true soil organic carbon content as the dependent variable, the first model is trained by using a partial least squares regression model, the latent variable is selected through multi-fold cross-validation to obtain a global main effect model, the training sample is input into the global main effect model to obtain a fold prediction value, and the verification sample is input into the global main effect model to obtain a baseline prediction value.

[0020] Preferably, the first model comprises supervised learning-elastic net regression, least squares support vector regression or kernel partial least squares regression.

[0021] Preferably, the second model comprises random forest, Cubsit, XGBoost, LightGBM or CatBoost.

[0022] Preferably, before training the first model, the spectral reflectance R is subjected to log(1 / R) and then Savitzky-Golay processing, and the processed spectrum and the dummy variable are fused, and the fused features are used as the dependent variable for training the first model.

[0023] Compared with the prior art, the present application has the following advantages:

[0024] In the present application, the classified texture is converted into a dummy variable for model training, so that when applied, the soil organic matter content can be obtained by directly inputting the texture category and spectral information into the trained model, without the need for testing the texture, without weakening the advantages of in-situ soil spectrum, and meeting the needs of simple and efficient prediction of soil organic matter content.

[0025] In the present application, the second model is trained to establish a mapping relationship between different textures and residual prediction values, so that the residual model provided by the present application can accurately capture the difference between the baseline prediction value of different textures and the true soil organic carbon content, thereby compared with the prior art, the present application can display systematic errors, the model prediction performance is more stable, and the soil organic carbon content can be accurately predicted for a more large-scale and highly heterogeneous soil. BRIEF DESCRIPTION OF DRAWINGS

[0026] Figure 1 The flowchart of the soil organic matter spectral prediction method based on texture perception residual superposition provided by the specific embodiment of the present application is provided;

[0027] Figure 2 The overall flowchart of the soil organic matter spectral prediction method based on texture perception residual superposition provided by the specific embodiment of the present application is provided;

[0028] Figure 3 The prediction effect comparison chart provided by the specific embodiment of the present application is provided. DETAILED DESCRIPTION

[0029] The specific embodiment of the present application proposes a texture perception residual superposition correction scheme: first, a model is established based on spectral library data, the model is used for prediction to obtain the baseline prediction value of the to-be-measured region, then the residual is modeled again in combination with the spectral library spectrum and the classified texture to obtain a residual model based on the texture and the spectrum, the texture (and the spectrum) is used as "domain information" to learn and offset the systematic deviation of the global baseline model in different texture spaces. The model is used in combination with the spectrum of the to-be-measured region to obtain the residual prediction value (pre_res) of the to-be-measured region, and the final prediction value is the difference between the baseline prediction value and the residual prediction value (pre_val-pre_res). This scheme does not sacrifice the global sample size while making directional correction to the systematic error caused by different textures.

[0030] The embodiment of the present application provides a soil organic matter spectrum prediction method based on texture perception residual superposition, as shown in Figure 1 and Figure 2 includes:

[0031] S1, constructing a training sample set and a verification sample set: according to the texture content of the soil sample in the soil spectrum library, that is, the sand, silt and clay content, the texture is classified, the classified texture is converted into a dummy variable, and the training sample set and the verification sample are constructed based on the spectrum information (Spe) of each soil sample in the soil spectrum library, the dummy variable of the corresponding texture and the real soil organic carbon content (SOM_cal).

[0032] In an embodiment, the step of performing texture coding to obtain the dummy variable of the texture provided by the embodiment includes:

[0033] According to the sand / silt / clay content of the soil sample in the spectrum, the corresponding texture is divided into three categories of coarse texture, medium texture and fine texture according to the Chinese three-part texture (coarse / fine) rule, that is, Texture∈ {coarse, medium, fine}, and is converted into a dummy variable (such as three columns), that is, based on the classification result, the texture of each sample is converted into a dummy variable of a three-column binary digital combination.

[0034] Specifically, the step of dividing the corresponding texture into three categories of coarse texture, medium texture and fine texture according to the rule provided by the embodiment includes:

[0035] When the clay content in the soil sample is ≥35%, the texture of the soil sample is fine;

[0036] When the sand content in the soil sample is ≥65% and the clay content is ≤18%, the texture of the soil sample is coarse;

[0037] The rest is medium.

[0038] By using the above method, after obtaining the soil to be tested, the soil does not need to be tested to accurately judge the texture content, but the coarse, medium and fine categories of the texture can be directly input to obtain the prediction result. In actual application, the coarse, medium and fine categories can be accurately judged by hand grinding, so that the method is simple and efficient, and the advantages of soil in-situ spectrum are not weakened.

[0039] After converting the texture content of each sample into a dummy variable, the spectrum variable of the spectrum information is spliced to form a feature matrix, and the feature matrix is taken as the input of the first model and the second model.

[0040] S2, training the first model to obtain a global main effect model (PLSR) based on the training sample set: using the training sample set, taking the spectral information and the dummy variable of the texture as the independent variable, and taking the true soil organic carbon content as the dependent variable, training the first model to obtain the global main effect model, and obtaining the fold-in prediction value through the global main effect model (PLSR).

[0041] In a specific embodiment, the step of training the first model comprises:

[0042] Taking the spectral information and the dummy variable of the texture as the independent variable, taking the true soil organic carbon content (SOC) as the dependent variable, training the first model by using the partial least squares regression model, selecting the latent variable through K-fold cross-validation, in an embodiment, K is 10, and the minimum cross-validation root mean square error RMSECV" or "one-sigma" criterion can be used for verification, to obtain the global main effect model, inputting the training sample into the global main effect model to obtain the fold-in prediction value, and inputting the verification sample into the global main effect model to obtain the baseline prediction value.

[0043] Specifically, the first model provided in the embodiment comprises a supervised learning-elastic network regression (Elastic Net), a least squares support vector regression (LSVR) or a kernel partial least squares regression (Kernel PLS).

[0044] S3, training the second model to obtain a residual model, i.e., a residual learning model (RF), based on the training sample set and the fold-in prediction value: using the training sample set, taking the spectral information (Pre) and the dummy variable (Tex) as the independent variable X, and taking the difference between the fold-in prediction value and the true soil organic carbon content as the dependent variable Y, training the second model to obtain the residual learning model (RF), and obtaining the residual estimation value (res_train) through the residual learning model, in an embodiment, the number of trees of the second model is 800-1200, and mtry≈ , p is the number of independent variables used by the secondary "residual model", wherein mtry is the number of feature variables considered by each tree at a split node.

[0045] In a specific embodiment, the second model provided in the embodiment comprises a random forest, Cubsit, XGBoost, LightGBM or CatBoost.

[0046] In a specific embodiment, before training the first model, the spectral information is preprocessed by absorbance (log(l / R) and Savitzky-Golay smoothing (w = 11, p = 2, m = 0), w is the window width, p is the polynomial order, and m is the differential order. The preprocessed light and dummy variables are fused, and the fused features are used as the dependent variable for training the first model. In this embodiment, the absorbance and Savitzky-Golay smoothing are introduced into the input, which can more effectively remove spectral noise and moisture information, thereby reducing spectral noise as much as possible and improving spectral prediction accuracy.

[0047] Before training the first model and the second model, the field judgment method of the texture needs to use hands to touch the soil to determine the texture, and the texture determination level is the coarse, medium and fine three categories of Chinese texture classification, and the three categories are coded, so that the texture information can be used to eliminate the heterogeneity error caused by the texture.

[0048] S4, in application, the soil spectral information of the current soil sample and the category of the texture are input into the global main effect model to obtain a baseline prediction value, the soil spectral information of the current soil sample and the category of the texture and the baseline prediction value are input into the residual model to obtain a residual prediction value, and the difference between the baseline prediction value and the residual prediction value is used to obtain a final soil organic carbon content prediction value.

[0049] The verification process of the verification set provided in this embodiment on the model is:

[0050] The soil spectrum (Spe_val) of the verification sample set (verification set) and the empirically perceived texture classification are input into the PLSR model to obtain a baseline prediction of the organic matter content of the verification set (SOM_val_pre).

[0051] The baseline prediction of the organic matter content of the verification set and the soil spectrum (Spe_val) of the verification sample set (verification set) and the empirically perceived texture classification are input into the RF model to obtain a residual prediction of the organic matter content of the verification set (Ses_val).

[0052] The difference between the baseline prediction value and the residual prediction value is used to obtain a final soil organic carbon content prediction value, i.e., a final prediction (SOM_val_preSes_val).

[0053] Then, the RMSE, R 2 , RPIQ and LCCC are used to evaluate the output results of the model. Compared with the common clay, silt and sand obtained by chemical analysis for soil texture classification (texture classification in the figure, blue points), texture perception and incorporation of these texture information (texture incorporation, red points), the results are as follows Figure 3The texture classification is first divided into three classes, and prediction is made for each class. The prediction results of each class are combined and the prediction results are calculated. The texture incorporation is to standardize the clay, silt and sand values and incorporate them into the spectral data to model and obtain the prediction results.

Claims

1. A soil organic matter spectral prediction method based on texture perception residual superposition, characterized in that, The application relates to a method for predicting soil organic carbon content based on soil spectrum information. According to the sand, silt and clay content of the soil samples in the soil spectrum library, the corresponding soil texture is classified into three categories, coarse texture, medium texture and fine texture, and the soil texture of each sample is converted into a dummy variable of a three-column binary digital combination based on the classification result. The second model comprises a random forest, Cubsit, XGBoost, LightGBM or CatBoost. The step of training the first model comprises the following steps: The spectrum information and the dummy variable are used as independent variables, and the real soil organic carbon content is used as a dependent variable, and a partial least squares regression model is used to train the first model, a latent variable is selected through multi-fold cross-validation, a global main effect model is obtained, the training sample is input into the global main effect model to obtain an intra-fold prediction value, and the verification sample is input into the global main effect model to obtain a baseline prediction value. Before training the first model, the spectrum reflectivity R is subjected to log (1 / R), and then subjected to Savitzky-Golay processing; the processed spectrum and the dummy variable are fused, and the fused features are used as the dependent variable for training the second model. According to the rules, the corresponding soil texture is classified into three categories, coarse texture, medium texture and fine texture, which comprises the following steps: When the clay content in the soil sample is greater than or equal to 35%, the soil texture of the soil sample is fine texture; When the sand content in the soil sample is greater than or equal to 65% and the clay content is less than or equal to 18%, the soil texture of the soil sample is coarse texture; The rest is medium texture.

2. The soil organic matter spectral prediction method based on texture perception residual superposition according to claim 1, characterized in that, The spectrum information of each soil sample and the dummy variable of the corresponding soil texture are used to construct a feature matrix, and the feature matrix is used as the independent variable for training the first model and the second model. The first model comprises supervised learning-elastic network regression, least squares support vector regression or kernel partial least squares regression. ​ ​ 3. The soil organic matter spectral prediction method based on texture perception residual superposition according to claim 1, characterized in that, ​ 4. The soil organic matter spectral prediction method based on texture perception residual superposition according to claim 1, characterized in that, ​

Citation Information

Patent Citations

  • Undisturbed soil profile carbon component prediction method based on hyperspectral imaging and support vector machine technology

    CN113436153A

  • Hyperspectral soil organic matter content prediction method based on DBO-SVR

    CN118169056A

  • Soil organic carbon spectrum prediction method and device based on spectrum guided ensemble learning

    CN116818687A

  • LIBS-oriented residual learning lightweight convolutional neural network quantification method

    CN117235512A

  • A method for constructing a prediction model for non-grain soil acidification

    CN119782942A