Method and device for predicting soil organic matter based on in-situ wet soil spectrum availability discrimination

By using a decision tree model to screen wet soil samples and construct selection paths, the problems of accuracy and efficiency in spectral prediction under wet soil conditions were solved, and adaptive soil organic matter prediction was achieved, improving prediction accuracy and stability.

CN121808523BActive Publication Date: 2026-05-29豫章师范学院
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
豫章师范学院
Filing Date
2026-03-10
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing spectral prediction models fail to effectively distinguish the differences in water effects among different textures, strata, and parent materials under wet soil conditions, resulting in large fluctuations in prediction results, difficulty in guaranteeing accuracy, and increased computational complexity and field operation time.

Method used

Wet soil samples were screened using a decision tree model. Based on the prediction accuracy, the samples were divided into three categories: direct prediction, uncertain, and dry soil priority prediction. Selection paths were constructed using spectral feature vectors and background parameters to determine whether water impact compensation treatment was needed or different prediction models were selected, thus achieving adaptive decision-making.

Benefits of technology

It improves the accuracy and efficiency of soil organic matter prediction in wet soils, reduces unnecessary correction steps, and enhances the stability and accuracy of in-situ prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808523B_ABST
    Figure CN121808523B_ABST
Patent Text Reader

Abstract

The application discloses a soil organic matter prediction method and device based on in-situ wet soil spectrum availability discrimination, the wet soil samples with high prediction accuracy of soil organic matter screened by the method can be analyzed by a second model to obtain a selection path for the wet soil samples, and based on the selection path, it can be judged which wet soil samples directly predict soil organic matter and which wet soil samples need to be supplemented with water influence, so that the prediction accuracy is improved, unnecessary correction steps are reduced, and in-situ prediction efficiency and stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of soil spectral quantitative analysis and soil quality monitoring, specifically involving a method and device for predicting soil organic matter based on the availability of in-situ wet soil spectral analysis. Background Technology

[0002] Soil organic matter is an important indicator reflecting soil fertility and ecological function. Traditional measurement methods rely on chemical analysis, which are complex, time-consuming, and costly. Visible-near-infrared spectroscopy, due to its advantages of speed, non-destructive testing, and on-site application, has been widely studied and applied in the field of soil organic matter prediction.

[0003] Existing spectral prediction models are mostly based on laboratory air-dried soil samples. However, in actual field surveys and remote sensing applications, soil is usually in a natural water content state. Moisture has significant absorption characteristics in the approximately 1400 nm and 1900 nm wavelength bands, which can easily mask the spectral information related to organic matter, thus leading to a significant decrease in prediction accuracy.

[0004] Patent application CN117929325A discloses a method for estimating soil organic matter content. The method includes: collecting soil samples, performing spectral measurements to obtain reflectance spectral data in the 500–2400 nm band; combining the reflectance spectral data pairwise to calculate the soil ratio index, soil difference index, and soil normalization index to obtain a two-band mixed spectrum of soil organic matter; selecting two-band spectral indices with an absolute correlation coefficient greater than 0.85 from the two-band mixed spectrum to form a two-band spectral index matrix; using an external parameter orthogonalization algorithm, processing the two-band spectral index matrix as input; and constructing an estimation model using partial least squares prediction to obtain the estimated soil organic matter content. The advantages are: selecting only characteristic bands, eliminating the influence of soil moisture on the accuracy of soil organic matter spectral monitoring, and improving detection accuracy.

[0005] Patent application CN118212522A discloses a method for predicting soil organic matter (SOM) content based on spectral coupling effects. This method improves the spatiotemporal transferability of the SOM prediction model by mitigating the coupling effect of soil physical properties on the spectrum. Based on satellite hyperspectral imagery and soil physical variables such as soil moisture, surface roughness, and bulk density, a soil spectral correction strategy based on information decomposition is established. Results show that soil spectral correction based on fourth-order polynomials and the XG-Boost algorithm has good accuracy and generalization ability. Furthermore, when the soil spectral correction strategy is adopted, the accuracy of the SOM prediction model and its generalization ability after model transfer are significantly improved. Compared with direct transfer prediction, the RMSE of the SOM prediction results is reduced by 57.90% and 60.27%, respectively, using the soil spectral correction strategy based on fourth-order polynomials and XG-Boost. This work provides a new research paradigm for the prediction of soil property parameters in other regions.

[0006] The aforementioned patent applications typically perform uniform moisture correction on all wet soil samples, but they do not consider the differences in moisture effects under different textures, strata, parent materials, and mineral compositions. They also do not determine whether the wet soil spectrum has reliable predictive capabilities under different soil-forming backgrounds. This can easily lead to over-correction or under-correction under certain conditions, resulting in large fluctuations in prediction results and difficulty in ensuring accuracy. It also increases computational complexity and field operation time. Summary of the Invention

[0007] This invention provides a method for predicting soil organic matter based on the availability of in-situ wet soil spectra. This method can make adaptive decisions on whether the wet soil spectrum needs to be compensated for by water influence or different prediction models should be selected, thereby improving prediction accuracy and reducing unnecessary correction steps, thus improving the efficiency and stability of in-situ prediction.

[0008] This invention provides a method for predicting soil organic matter based on in-situ wet soil spectral availability, comprising:

[0009] The first model is used to predict soil organic matter multiple times for each wet soil sample and the prediction accuracy is statistically analyzed. Based on the prediction accuracy, each wet soil sample is divided into three categories: wet soil direct prediction, uncertain, and dry soil priority prediction. The category is used as the dependent variable, and soil organic matter content, background parameters, and spectral feature vector are used as independent variables. Based on the dependent and independent variables, the second model can obtain the selection path pointing to the selected wet soil direct prediction category. Each node in the selection path includes background parameters and / or spectral feature vector.

[0010] Visible-near-infrared reflectance spectra of the soil under natural moisture conditions are collected to obtain raw wet soil spectral data. The raw wet soil spectral data is preprocessed to obtain characteristic parameters that can characterize the water response and spectral morphology, thereby forming a spectral feature vector and obtaining the background parameters of the soil under test. If the spectral feature vector and background parameters of the soil under test meet the requirements of each node on the selected path, the soil organic matter content of the soil under test is directly predicted; otherwise, water influence compensation processing is performed first before predicting the soil organic matter content of the soil under test.

[0011] Preferably, based on the dependent and independent variables, a selection path pointing to the screened category of directly predicted wet soil can be obtained through the second model, including:

[0012] The independent variables include soil organic matter content, moisture index, soil organic matter, soil strata, soil texture, parent material, and moisture content.

[0013] During training, the independence of each variable from the three dependent variables is first tested by permutation. If the most significant variable satisfies the splitting criterion, the variable is selected and its optimal splitting threshold is automatically searched for to divide the nodes; otherwise, the splitting is terminated and leaf nodes are formed.

[0014] By recursively splitting until the maximum tree depth or minimum sample size constraint is met, the trained decision tree structure is obtained.

[0015] Subsequently, the splitting variables and corresponding numerical ranges of each layer are automatically derived based on the decision tree structure to form selection paths under different scenarios. Among them, samples with high moisture index and high water content are given priority to enter the wet soil direct prediction path. The wet soil direct prediction path is the selection path that is selected and points to the wet soil direct prediction category.

[0016] Preferably, the soil organic matter content of each obtained wet soil sample is predicted multiple times using the first model, and the prediction accuracy is statistically analyzed, including:

[0017] The first model is used to predict soil organic matter multiple times for each wet soil sample. Each prediction result is compared with the true value. If the difference is less than the set difference threshold, the prediction is accurate; otherwise, it is inaccurate. Thus, the prediction accuracy of soil organic matter for each wet soil sample is calculated.

[0018] The first model is partial least squares regression, random forest, or decision tree.

[0019] Preferably, based on prediction accuracy, each wet soil sample is divided into three categories: direct prediction for wet soil, uncertain prediction for wet soil, and priority prediction for dry soil, including:

[0020] Pre-set upper and lower accuracy limits. When the prediction accuracy is greater than or equal to the upper accuracy limit, the corresponding wet soil sample is classified as wet soil with direct prediction. When the prediction accuracy is less than or equal to the lower accuracy limit, the corresponding wet soil sample is classified as dry soil with priority prediction. When the prediction accuracy is between the upper and lower accuracy limits, the corresponding wet soil sample is classified as uncertain.

[0021] Preferably, the background parameters include soil property parameters and / or environmental factor parameters.

[0022] Preferably, the original wet soil spectral data are preprocessed with absorbance and SG first derivative smoothing to obtain characteristic parameters that can characterize the moisture response and spectral morphology.

[0023] Preferably, moisture impact compensation treatment is performed, including:

[0024] A moisture change feature subspace is constructed based on paired samples of wet and dry soil. The components of the spectral feature vector of the soil to be tested in the moisture change feature subspace are orthogonally projected and eliminated to obtain a corrected spectrum. The corrected spectrum is then input into the corrected prediction model to obtain the predicted value of soil organic matter.

[0025] The corrected prediction model is obtained by multiplying the dependent variable of an established model for directly predicting the soil organic matter content of the soil to be tested by a correction matrix.

[0026] Preferably, moisture impact compensation treatment is performed, including:

[0027] Preliminary predicted values ​​of soil organic matter were obtained using a wet soil prediction model.

[0028] The background parameters of the soil to be tested are added to the dependent variable of the wet soil prediction model, and the model is re-established. Based on the difference between the predicted value output by the re-established model and the predicted value output by the wet soil prediction model, and the true parameters, a loss function is constructed to train and re-establish the model to obtain a residual regression model. The preliminary soil organic matter prediction value is compensated and corrected by the residual value output by the residual regression model to obtain the soil organic matter prediction value.

[0029] The wet soil prediction model is an established model used to directly predict the soil organic matter content of the soil to be tested.

[0030] Preferably, the wet soil prediction model includes partial least squares regression, random forest, or decision tree.

[0031] The present invention also provides a soil organic matter prediction device based on in-situ wet soil spectral availability discrimination, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the soil organic matter prediction method based on in-situ wet soil spectral availability discrimination.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] The present invention selects wet soil samples with high accuracy in predicting soil organic matter. The second model can analyze the selection path to the wet soil sample. Based on the selection path, it can determine which wet soil samples can be directly predicted for soil organic matter and which need to be supplemented by water influence. This achieves both improved prediction accuracy and reduced unnecessary correction steps, thereby improving the efficiency and stability of in-situ prediction. Attached Figure Description

[0034] Figure 1 A flowchart of a method for predicting soil organic matter based on in-situ wet soil spectral availability is provided in a specific embodiment of the present invention.

[0035] Figure 2 This is a selection path diagram provided in Embodiment 1 of the present invention;

[0036] Figure 3 The comparison diagram shows the prediction effect comparison between the comparison method provided in a specific embodiment of the present invention and the method provided by the present invention. Detailed Implementation

[0037] In a specific embodiment of the present invention, wet soil samples with good prediction results when directly predicted are first screened out, and then the selection path that can point to the wet soil sample is analyzed by the model. In practical applications, the background parameters and spectral feature vector of the wet soil to be tested are compared with each node of the selection path, so as to determine whether the organic matter content of the soil to be tested can be directly predicted.

[0038] The specific embodiments of this invention provide a method for predicting soil organic matter based on in-situ wet soil spectral availability, such as... Figure 1 As shown, it includes:

[0039] S1. The soil organic matter content of each wet soil sample is predicted multiple times using the first model, and the prediction accuracy is statistically analyzed. Based on the prediction accuracy, each wet soil sample is divided into three categories: wet soil direct prediction, uncertain, and dry soil priority prediction. The category is used as the dependent variable, and the soil organic matter content, background parameters, and spectral feature vector are used as independent variables. Based on the dependent and independent variables, the second model can obtain the selection path pointing to the wet soil direct prediction category. Each node in the selection path is the background parameter and / or spectral feature vector, thereby obtaining prediction path selection control.

[0040] In a specific embodiment of the present invention, based on the dependent and independent variables, a second model can be used to obtain a selection path pointing to the screened wet soil samples, including:

[0041] After obtaining the dependent and independent variables, a training dataset is constructed and input into the second model. The dependent variables are categorized into three types: direct prediction for wet soil, uncertainty, and priority prediction for dry soil. The independent variables include the moisture index (MDI_1900), soil organic matter (SOM), soil layers, soil texture, parent material, and moisture content. The second model employs a conditional inference tree (CTree) classification model. During the training phase, a (weighted) permutation test is first performed on the independence of each independent variable from the three dependent variables. If the most significant variable meets the splitting criterion, that variable is selected, and its optimal splitting threshold is automatically searched for for node partitioning; otherwise, splitting terminates and leaf nodes are formed. Recursive splitting continues until the maximum tree depth or minimum sample size constraint is met, resulting in a trained decision tree structure. Subsequently, the splitting variables and corresponding numerical ranges for each layer are automatically derived from this structure, forming selection paths under different scenarios. Samples with high moisture index and high moisture content are preferentially entered into the "direct prediction for wet soil" path.

[0042] In one embodiment, the conditional inference tree (CTree) from the partykit package is used to classify and model the class (class ~ .). At each node, a (weighted) permutation test is first performed on the independence between all candidate features and the response variable. If the test result of the most significant feature satisfies the splitting criterion (mincriterion = 0.90, corresponding to a significance level of approximately 0.10), then that feature is selected and its optimal split point is searched for to split the node; otherwise, splitting stops and leaf nodes are generated. To control model complexity, the minimum split sample size minsplit = 30 and the maximum tree depth maxdepth = 3 are set. The leaf node class is determined by the weighted class ratio of the samples within the node (weighted majority vote / highest probability class).

[0043] The second model provided in the specific embodiment of the present invention is a machine learning classification model, which can be any one of a decision tree model, a conditional inference tree model, or a random forest model.

[0044] In specific embodiments of the present invention, this method can be used to analyze the parameter values ​​or categories of background parameters and spectral feature vectors of wet soil samples with high prediction accuracy, providing a judgment standard for subsequent determination of whether the wet soil sample to be tested can be directly predicted.

[0045] In one specific embodiment, this embodiment uses a first model to perform multiple soil organic matter predictions on each obtained wet soil sample and statistically analyzes the prediction accuracy, including:

[0046] The first model is used to predict soil organic matter multiple times for each wet soil sample. Each prediction result is compared with the true value. If the difference is less than the set threshold, the prediction is accurate; otherwise, it is inaccurate. The prediction accuracy of soil organic matter for each wet soil sample is then calculated. The first model includes partial least squares regression, random forest, and decision tree.

[0047] In one specific embodiment, each wet soil sample is divided into three categories based on prediction accuracy: wet soil direct prediction, uncertain prediction, and dry soil priority prediction, including:

[0048] Pre-set upper and lower accuracy limits. When the prediction accuracy is greater than or equal to the upper accuracy limit, the corresponding wet soil sample is classified as wet soil with direct prediction. When the prediction accuracy is less than or equal to the lower accuracy limit, the corresponding wet soil sample is classified as dry soil with priority prediction. When the prediction accuracy is between the upper and lower accuracy limits, the corresponding wet soil sample is classified as uncertain.

[0049] The present invention can obtain high-quality wet soil samples for organic matter identification by using the above-mentioned method for judging prediction accuracy.

[0050] S2. In a specific embodiment of the present invention, in-situ field samples are first collected. The visible-near-infrared reflectance spectra of the soil to be tested are collected under natural water content to obtain the original wet soil spectral data. Then, the original wet soil spectral data is subjected to spectral preprocessing and feature extraction to obtain feature parameters that can characterize water response and spectral morphology, thereby forming a spectral feature vector and obtaining the background parameters of the soil to be tested.

[0051] In one specific embodiment, the raw wet soil spectral data is preprocessed with absorbance and SG first derivative smoothing to obtain characteristic parameters that can characterize the moisture response and spectral morphology. The SG first derivative is obtained by smoothing the signal with a Savitzky-Golay filter and calculating the first derivative, which is used to highlight signal changes and peak characteristics.

[0052] In one specific embodiment, the background parameters include soil property parameters and / or environmental factor parameters.

[0053] S3. If the spectral feature vector and background parameters of the soil to be tested meet the requirements of each node on the selected path, the soil organic matter content of the soil to be tested is directly predicted using the wet soil prediction model; otherwise, the moisture influence compensation processing is performed on the wet soil spectrum, and a calibrated prediction model is used for prediction; or a prediction model matching the moisture conditions is selected for prediction. The moisture influence compensation processing refers to the process of weakening or offsetting the adverse effects of moisture on the spectral prediction relationship through spectral transformation, prediction result correction, or model selection.

[0054] The wet soil prediction model provided in this specific embodiment of the invention can be a data-driven model such as partial least squares regression, random forest, extreme gradient boosting, support vector regression, or artificial neural network. This embodiment uses partial least squares regression (PLSR) as an example. First, a wet soil training dataset is constructed, with wet soil spectral reflectance as the input independent variable and measured soil organic matter content as the output dependent variable. Then, a PLSR model is established based on preprocessed spectra, and the optimal number of latent variables is determined through cross-validation. Finally, the R-squared value of the model is evaluated on an independent validation set. 2 After the RMSE and RPD values ​​meet the preset accuracy requirements, the model parameters are saved as the final wet soil prediction model for subsequent direct prediction of new wet soil samples.

[0055] In one specific embodiment, within the adaptive prediction framework of this invention, moisture impact compensation is used as one of the prediction paths. The control results are selectively invoked based on the prediction path selection. Specific implementation methods include, but are not limited to, the following:

[0056] Method 1: Orthogonal projection correction method. A moisture change feature subspace is constructed based on paired samples of wet and dry soil. The components of the wet soil spectrum in the direction of this subspace are orthogonally projected to eliminate the moisture, and the corrected spectrum is obtained before prediction.

[0057] Specifically, a moisture change feature subspace is constructed based on paired samples of wet and dry soil. The components of the spectral feature vector of the soil to be tested on the moisture change feature subspace are orthogonally projected and eliminated to obtain a corrected spectrum. The corrected spectrum is then input into the correction prediction model to obtain the predicted value of soil organic matter.

[0058] The corrected prediction model is obtained by multiplying the dependent variable of an established model for directly predicting the soil organic matter content of the soil to be tested by a correction matrix.

[0059] Method 2: Residual Compensation Method. First, a preliminary prediction value is obtained using a wet soil prediction model. Then, a residual regression model is constructed based on spectral characteristic parameters and background parameters to compensate and correct the preliminary prediction results, thus obtaining the final prediction result.

[0060] Specifically, a preliminary soil organic matter prediction value is obtained through a wet soil prediction model; background parameters of the soil to be tested are added to the dependent variable of the wet soil prediction model, and the model is re-established; a loss function is constructed based on the difference between the predicted value output by the re-established model and the predicted value output by the wet soil prediction model, and the true parameters to train and re-establish the model to obtain a residual regression model; the preliminary soil organic matter prediction value is compensated and corrected by the residual value output by the residual regression model to obtain the soil organic matter prediction value; the wet soil prediction model is an established model used to directly predict the soil organic matter content of the soil to be tested.

[0061] Example 1: S1. The visible-near infrared reflectance spectrum of the soil to be tested was collected under natural water content to obtain the original wet soil spectral data.

[0062] S2. The original wet soil spectral data is preprocessed with absorbance and SG first derivative smoothing, and characteristic parameters characterizing moisture response and spectral morphology are extracted to form a spectral feature vector.

[0063] S3. Obtain the background parameters corresponding to the soil to be tested. The background parameters include at least soil property parameters, environmental factor parameters, or a combination thereof.

[0064] S4. Generate a predicted path selection control result based on the spectral feature vector and the background parameters. Specific steps include:

[0065] (1) Repeatedly sample wet soil samples and use the PLSR model to predict soil organic matter multiple times for each wet soil sample. The results of multiple predictions are compared with the true value. If the difference between the prediction result and the true value is less than the set difference threshold, the prediction is accurate; otherwise, the prediction is inaccurate. Calculate the prediction accuracy of each wet soil sample and set the lower and upper limits of the accuracy probability. When the prediction accuracy is greater than or equal to the upper limit of the accuracy probability, the wet soil sample is a wet soil direct prediction category. When the prediction accuracy is less than or equal to the lower limit of the accuracy probability, the wet soil sample is a wet soil poor prediction category, i.e., a dry soil priority prediction category, or a dry soil better prediction category. When the prediction accuracy is between the upper and lower limits of the accuracy probability, the wet soil sample is a fuzzy category, i.e., an uncertain category.

[0066] (2) Obtain soil samples corresponding to the better predicted category of wet soil. The categories (wet soil direct prediction category, dry soil priority prediction category, and uncertain category) are used as dependent variables. Soil organic matter content, background parameters, and spectral feature vectors are used as independent variables. Based on these dependent and independent variables, a second model can be used to obtain a selection path pointing to the selected wet soil samples. Each node in the selection path includes background parameters and / or spectral feature vectors. Using the dependent and independent variables as inputs, an existing machine learning classification model is used to obtain the selection path pointing to the selected wet soil samples. Each node in the selection path is a background parameter and / or spectral feature vector. The spectral reflectance vector obtained for the same sample under wet soil conditions is used as the main independent variable, and background parameters such as moisture content, soil strata, texture, and parent material are introduced as auxiliary independent variables for modeling. The reason is that, on the one hand, the wet soil spectrum contains comprehensive spectral information of soil organic matter and moisture, which is the main information carrier for predicting SOM; on the other hand, introducing background parameters helps to distinguish different soil environmental scenarios, reduce the interference of moisture and mineral background on the spectral-SOM relationship, thereby improving the robustness and generalization ability of the model. To ensure model stability, we prioritize training with representative and reliable wet soil samples.

[0067] (3) When applying, the visible-near infrared reflectance spectrum of the soil to be tested is collected under natural water content to obtain the original wet soil spectral data. The original wet soil spectral data is preprocessed to obtain characteristic parameters that can characterize the water response and spectral morphology, thereby forming a spectral feature vector and obtaining the background parameters of the soil to be tested. If the spectral feature vector and background parameters of the soil to be tested meet the requirements of each node on the selected path, the soil organic matter content of the soil to be tested is directly predicted. Otherwise, water influence compensation processing is performed first and then the soil organic matter content of the soil to be tested is predicted.

[0068] Specifically, such as Figure 2 As shown, Figure 2The background parameters include organic matter-related spectral characteristic parameters, soil type parameters, and soil layer parameters (stratum type). Organic matter-related spectral characteristic parameters are obtained from in-situ wet soil spectral data and background parameters. Based on the prediction path selection control results generated in step (2), when the organic matter spectral characteristic parameter is greater than or equal to the threshold T1, the soil type parameter judgment is entered; otherwise, the prediction path after water influence supplementation is directly executed. When the soil type parameter meets the requirements, such as dry land or paddy soil, the water response spectral characteristic parameter judgment is performed. When the water response spectral characteristic parameter is greater than or equal to the threshold T2, i.e. the first threshold, the wet soil direct prediction path is directly executed; otherwise, the prediction path after water influence supplementation is executed. When the soil type parameter does not meet the requirements, such as forest land, the soil layer parameter judgment is performed. When the soil layer parameter meets the requirements, such as the first and third soil layers, the wet soil direct prediction path is directly executed; otherwise, the prediction path after water influence supplementation is executed, thus completing the execution of the adaptive prediction path.

[0069] Example of path selection prediction based on multi-feature joint discrimination. For example... Figure 2 As shown, the samples are classified step by step based on the initial estimated level of soil organic matter, soil type, soil layer information and 1900 nm water absorption characteristics, and the corresponding prediction path is selected based on the final classification results.

[0070] Example 2: As Figure 3 As shown, the adaptive soil organic matter prediction method based on in-situ wet soil spectral availability discrimination (the method of this invention) proposed in Example 1 is compared with two comparative models established based on experimental data of this invention:

[0071] (1) Dry-PLSR (Dry Soil Model): A dry soil prediction model is established using partial least squares regression (PLSR) with the air-dried soil spectrum collected in this invention as input and the measured SOM as output.

[0072] (2) Wet-PLSR (Direct Measurement of Wet Soil): A direct prediction model for wet soil is established using partial least squares regression (PLSR) with the in-situ wet soil spectrum collected by this invention as input and the measured SOM as output.

[0073] The results show that, under different soil moisture conditions, the adaptive prediction method proposed in this invention has a higher consistency with the measured values ​​(R0). 2(Higher accuracy, lower RMSE, larger RPD) significantly reduces prediction error, outperforming single dry soil models or direct wet soil prediction models. A specific embodiment of this invention also provides a soil organic matter prediction device based on in-situ wet soil spectral availability discrimination, characterized by including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the soil organic matter prediction method based on in-situ wet soil spectral availability discrimination.

[0074] On the other hand, the present invention also provides a soil organic matter prediction device based on in-situ wet soil spectral availability discrimination, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the soil organic matter prediction method based on in-situ wet soil spectral availability discrimination.

[0075] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for predicting soil organic matter based on in-situ wet soil spectral availability discrimination, characterized in that, include: The first model is used to predict soil organic matter multiple times for each wet soil sample and the prediction accuracy is statistically analyzed. Based on the prediction accuracy, each wet soil sample is divided into three categories: wet soil direct prediction, uncertain, and dry soil priority prediction. The category is used as the dependent variable, and soil organic matter content, background parameters, and spectral feature vector are used as independent variables. Based on the dependent and independent variables, the second model can obtain the selection path pointing to the selected wet soil direct prediction category. Each node in the selection path includes background parameters and spectral feature vector. Visible-near-infrared reflectance spectra of the soil under natural moisture conditions are collected to obtain raw wet soil spectral data. The raw wet soil spectral data is preprocessed to obtain characteristic parameters that can characterize water response and spectral morphology, thereby forming a spectral feature vector and obtaining background parameters of the soil under test. If the spectral feature vector and background parameters of the soil under test meet the requirements of each node on the selected path, the soil organic matter content of the soil under test is directly predicted; otherwise, water influence compensation processing is performed first before predicting the soil organic matter content of the soil under test. Based on the dependent and independent variables, the second model can obtain a selection path pointing to the screened category of direct prediction of wet soil, including: The independent variables include soil organic matter content, moisture index, soil organic matter, soil strata, soil texture, parent material, and moisture content. During training, the independence of each variable from the three dependent variables is first tested by permutation. If the most significant variable satisfies the splitting criterion, the variable is selected and its optimal splitting threshold is automatically searched for to divide the nodes; otherwise, the splitting is terminated and leaf nodes are formed. By recursively splitting until the maximum tree depth or minimum sample size constraint is met, the trained decision tree structure is obtained. Subsequently, based on the decision tree structure, the splitting variables and corresponding numerical ranges of each layer are automatically derived to form selection paths under different scenarios. Among them, samples with high moisture index and high water content are given priority to enter the wet soil direct prediction path. The wet soil direct prediction path is the selection path that is selected to point to the wet soil direct prediction category. Based on prediction accuracy, each wet soil sample is divided into three categories: direct prediction for wet soil, uncertain prediction, and priority prediction for dry soil. Pre-set upper and lower accuracy limits. When the prediction accuracy is greater than or equal to the upper accuracy limit, the corresponding wet soil sample is classified as wet soil with direct prediction. When the prediction accuracy is less than or equal to the lower accuracy limit, the corresponding wet soil sample is classified as dry soil with priority prediction. When the prediction accuracy is between the upper and lower accuracy limits, the corresponding wet soil sample is classified as uncertain.

2. The method for predicting soil organic matter based on in-situ wet soil spectral availability discrimination according to claim 1, characterized in that, The soil organic matter content of each obtained wet soil sample was predicted multiple times using the first model, and the prediction accuracy was statistically analyzed, including: The first model is used to predict soil organic matter multiple times for each wet soil sample. Each prediction result is compared with the true value. If the difference is less than the set difference threshold, the prediction is accurate; otherwise, it is inaccurate. Thus, the prediction accuracy of soil organic matter for each wet soil sample is calculated. The first model is partial least squares regression, random forest, or decision tree.

3. The method for predicting soil organic matter based on in-situ wet soil spectral availability discrimination according to claim 1, characterized in that, The background parameters include soil property parameters and environmental factor parameters.

4. The method for predicting soil organic matter based on in-situ wet soil spectral availability discrimination according to claim 1, characterized in that, The original wet soil spectral data were preprocessed with absorbance and SG first derivative smoothing to obtain characteristic parameters that can characterize the moisture response and spectral morphology.

5. The method for predicting soil organic matter based on in-situ wet soil spectral availability discrimination according to claim 1, characterized in that, Moisture impact compensation treatment includes: A moisture change feature subspace is constructed based on paired samples of wet and dry soil. The components of the spectral feature vector of the soil to be tested in the moisture change feature subspace are orthogonally projected and eliminated to obtain the corrected spectrum. The corrected spectrum is then input into the correction prediction model to obtain the predicted value of soil organic matter. The corrected prediction model is obtained by multiplying the dependent variable of an established model for directly predicting the soil organic matter content of the soil to be tested by a correction matrix.

6. The method for predicting soil organic matter based on in-situ wet soil spectral availability discrimination according to claim 1, characterized in that, Moisture impact compensation treatment includes: Preliminary predicted values ​​of soil organic matter were obtained using a wet soil prediction model. The background parameters of the soil to be tested are added to the dependent variable of the wet soil prediction model, and the model is re-established. Based on the difference between the predicted value output by the re-established model and the predicted value output by the wet soil prediction model, and the true parameters, a loss function is constructed to train and re-establish the model to obtain a residual regression model. The preliminary soil organic matter prediction value is compensated and corrected by the residual value output by the residual regression model to obtain the soil organic matter prediction value. The wet soil prediction model is an established model used to directly predict the soil organic matter content of the soil to be tested.

7. The method for predicting soil organic matter based on in-situ wet soil spectral availability discrimination according to claim 6, characterized in that, The wet soil prediction model includes partial least squares regression, random forest, or decision tree.

8. A soil organic matter prediction device based on in-situ wet soil spectral availability discrimination, characterized in that, The method includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the soil organic matter prediction method based on in-situ wet soil spectral availability discrimination as described in any one of claims 1-7.