Soil organic matter determination method and system

Through near-infrared spectroscopy technology and deep learning models, the complexity and accuracy problems of traditional soil organic matter measurement methods have been solved, and fast and accurate soil organic matter determination has been achieved, which is suitable for agricultural production.

CN120668606APending Publication Date: 2025-09-19KUNMING COMPREHENSIVE NATURAL RESOURCES SURVEY CENT OF CHINA GEOLOGICAL SURVEY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511034635.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional soil organic matter measurement methods are complex, time-consuming, and costly, and are not suitable for rapid on-site agricultural testing. Existing spectral technology is not sophisticated enough in spectral data preprocessing in soil organic matter determination, and the model prediction accuracy needs to be improved.

Method used

Near-infrared spectroscopy technology is used to establish an organic matter prediction model through sample preparation, spectral acquisition, data preprocessing, feature extraction, model building and humidity and temperature correction steps, combined with principal component analysis, multivariate scattering correction, standard normal transformation, detrending processing and deep learning model, and the organic matter is measured using a near-infrared spectrometer and humidity/temperature correction model.

Benefits of technology

It achieves rapid and accurate measurement of soil organic matter content, reduces measurement costs, improves measurement efficiency and accuracy, is suitable for various soil types and agricultural production environments, and supports precise fertilization and scientific management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120668606A_ABST
    Figure CN120668606A_ABST
Patent Text Reader

Abstract

The invention discloses a soil organic matter determination method and system, and relates to the technical field of soil detection. The method comprises the steps of sample preparation, spectrum acquisition, data preprocessing, feature extraction, model establishment, humidity and temperature correction and sample determination. The method comprises the following steps: firstly, collecting a soil sample from a to-be-detected area and treating to ensure uniformity; then, a near infrared spectrometer is used for collecting spectral data, and preprocessing is carried out to extract feature information; thirdly, training and verifying an organic matter prediction model by using a Lucas soil public data set; meanwhile, collecting soil humidity and temperature information, and establishing a humidity / temperature correction model to correct a measurement result; and finally, inputting spectral data of a to-be-detected sample into the model to obtain the organic matter content. The system comprises a sample preparation device, a spectrum acquisition device, a data processing unit, a humidity and temperature measurement device and an output device, and is used for realizing the measurement method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of soil detection, and in particular relates to a soil organic matter determination method and system. Background Art

[0002] In modern agricultural production, scientific management of fertilization and increasing grain yields are core issues. Organic matter, as a crucial component of soil, is a key indicator of soil fertility and crucial for crop growth. Traditionally, soil organic matter measurement has relied on chemical methods such as potassium dichromate volumetric analysis, dry burning, and calcination. While accurate, these methods are complex, time-consuming, and costly, requiring expensive testing equipment, making them unsuitable for rapid field testing in agriculture.

[0003] In recent years, with the rapid development of spectral technology, near-infrared, remote sensing, and hyperspectral techniques have been widely used for rapid soil nutrient estimation. These techniques offer advantages such as non-destructiveness, rapidity, and high efficiency, providing new approaches for the rapid determination of soil organic matter content. However, existing spectral techniques for soil organic matter determination still face some challenges, such as insufficient preprocessing of spectral data and the need to improve model prediction accuracy. Summary of the Invention

[0004] To solve the above problems, the present invention proposes a method and system for rapid determination of soil organic matter. The method includes the steps of sample preparation, spectral acquisition, data preprocessing, feature extraction, model establishment, humidity and temperature correction, and sample determination, and the system includes corresponding devices and equipment.

[0005] To solve the above technical problems, the present invention is achieved through the following technical solutions: The present invention provides a soil organic matter determination method and system, comprising: S1. Sample preparation: Collect soil samples from the area to be tested, mix, dry, grind and sieve to ensure sample homogeneity; S2. Spectral acquisition: Use a near-infrared spectrometer to collect spectra of the prepared soil samples to obtain spectral data of the samples; S3. Data preprocessing: The collected spectral data are subjected to outlier removal, multivariate scattering correction, standard normal transformation, detrending, and feature band extraction to obtain preprocessed spectral data; S4, feature extraction: extracting feature information related to soil organic matter content from the preprocessed spectral data; S5. Model building: Build an organic matter prediction model and use the known Lucas soil public dataset to train and validate the organic matter prediction model. S6. Humidity and temperature correction: Collect humidity and temperature information of soil samples and establish a humidity / temperature correction model through mathematical modeling methods to correct the measurement results; S7. Sample determination: Input the spectral data of the soil sample to be tested into the established digital organic matter prediction model and humidity / temperature correction model to obtain the organic matter content of the sample.

[0006] As a preferred technical solution of the present invention, the S1 sample preparation step specifically includes the following sub-steps: S1.1 Selection of representative sampling points: Based on the research objectives or monitoring requirements, select representative sampling points in the area to be tested that can fully reflect the soil characteristics; S1.2 Sampling operation specifications: Use clean, non-contaminated sampling tools for sampling, remove the surface soil to avoid interference from impurities, determine the sampling depth based on the research purpose, and collect soil samples at a depth of 0-20 cm; S1.3 Sample mixing: In each sampling unit, the soil samples collected at each point are thoroughly mixed to form a mixed sample; S1.3 Drying and grinding: Spread the mixed soil sample on a tray and dry it naturally in a cool and ventilated environment. After drying, use stainless steel or ceramic tools to finely grind it. S1.4 Particle size screening: Based on the analysis requirements, the ground sample is screened to obtain soil particles of different particle size ranges; S1.5 Container selection and labeling information: Use clean, well-sealed containers to store soil samples. Avoid using materials that can chemically react with the soil. Mark the container with key information such as sample number, collection location, and date for subsequent analysis and traceability.

[0007] As a preferred technical solution of the present invention, the S2 spectrum acquisition step specifically includes the following sub-steps: S2.1 Instrument preparation: Before collecting spectral data, preheat and calibrate the near-infrared spectrometer to ensure the accuracy and stability of the measurement results; S2.2 Sample measurement: Place the prepared soil sample in the measuring container of the infrared spectrometer and collect the spectral data of the sample; S2.3 Data recording: The collected spectral data are recorded in the computer to provide a basis for subsequent data preprocessing and sample determination.

[0008] As a preferred technical solution of the present invention, in order to ensure the accuracy and stability of the spectral data in the S2 spectral collection step, spectral data of each soil sample is collected at least three times, and the average value is taken as the final spectral data.

[0009] As a preferred technical solution of the present invention, the S3 data preprocessing step specifically includes the following sub-steps: S3.1 Spectral principal component calculation: Use principal component analysis (PCA) to reduce the dimensionality of spectral data and extract the main spectral characteristic components; S3.2 Mahalanobis distance calculation: Calculate the Mahalanobis distance between each sample spectral data and the principal component space; S3.3 Plot the Mahalanobis distance of samples and calculate the threshold: Plot the Mahalanobis distance of all samples, observe their distribution, and set a threshold based on the distribution characteristics to separate abnormal data; S3.4 Use Matlab to perform multivariate scatter correction, standard normal transformation and detrending: S3.41 Multivariate Scattering Correction: To address the uneven particle size, shape, and distribution in soil samples, a multivariate scattering correction method is used to reduce the impact of scattering effects on spectral data. S3.42 Standard Normal Transformation: Perform a standard normal transformation on the spectral data to make it conform to the normal distribution characteristics; S3.43 Detrending: Remove the trend component in the spectral data, that is, the trend that gradually increases or decreases with time or wavelength; S3.5 Extract characteristic bands using partial least squares regression: The spectral data are further processed using partial least squares regression (PLSR). By calculating the regression coefficients, the characteristic bands that are most correlated with the soil organic matter content are identified. These characteristic bands will serve as the basis for subsequent modeling and prediction.

[0010] As a preferred technical solution of the present invention, the S5 model establishment process specifically includes the following sub-steps: S5.1 Build a deep learning model: S5.11 Design a multilayer perceptron (MLP) neural network model consisting of an input layer, at least one hidden layer, and an output layer. The input layer of S5.12 receives the spectral data of the soil, and the output layer outputs the predicted organic matter content; S5.13 Adjust the number of hidden layers and the number of neurons in each layer based on the specific problem and dataset characteristics; S5.14 Select the activation function, ReLU, sigmoid, or tanh, to increase the nonlinearity of the network; S5.15 Define the loss function, mean squared error (MSE) or cross entropy loss, which measures the difference between the network's predictions and the true values. S5.2 Model training: S5.21 Use the training set data to train the neural network model; S5.22 During the training process, forward propagation calculations are performed to calculate the activation values ​​and output values ​​of each layer; S5.23 calculates the difference between the network prediction value and the true value based on the loss function, that is, the loss value; S5.24 uses the backpropagation algorithm to calculate the gradient of the loss function with respect to the weights and biases of each layer; S5.25 Update the network weights and biases using the gradient descent method based on the calculated gradient; S5.26 Repeat the forward propagation, loss calculation, and weight update steps until a predetermined number of training rounds is reached or the loss value converges; S5.3 Model Validation: S5.31 Use the validation set data to validate the trained model and evaluate the model's predictive performance and accuracy; S5.32 optimizes the model based on the verification results, including adjusting the network structure, activation function, loss function and optimizing algorithm parameters.

[0011] As a preferred technical solution of the present invention, the use of Type transfer function is used as the transfer function of the hidden layer, using The transfer function outputs a non-negative number and transmits it in the hidden layer. The type transfer function is transferred in the output layer; By finding the total error Adjust the weights between different layers of the error signal back propagation; the specific correlation function formula is as follows; Type transfer function: Type transfer function: Type transfer function: Total error: .

[0012] As a preferred technical solution of the present invention, the humidity and temperature correction step S6 specifically includes the following steps: S6.1 Data collection: Obtain relevant data of the LUCAS soil dataset as a training set; S6.2 Correction model construction: Using the LUCAS soil dataset containing moisture and temperature information and corresponding organic matter content data, a moisture and temperature correction model was established using multiple linear regression or neural network mathematical modeling methods; S6.3 Calibration model validation: Use an independent validation set to validate the established moisture and temperature calibration model to evaluate the calibration effect and accuracy of the model. The validation set should contain soil samples different from the training set to ensure that the calibration model has good generalization ability.

[0013] A soil organic matter determination system, used to cooperate with the above-mentioned soil organic matter determination method, comprises: Sample preparation device: used for collecting, mixing, drying, grinding and sieving soil samples; Spectral acquisition device: including a near-infrared spectrometer, used to collect spectra of the prepared soil samples; Data processing unit: used for preprocessing, feature extraction, model building and sample measurement of the collected spectral data; Humidity and temperature measuring device: used to collect humidity and temperature information of soil samples; Output device: used to display and output the measurement results, including organic matter content and the results after humidity and temperature correction.

[0014] The present invention has the following beneficial effects: The present invention adopts near-infrared spectroscopy technology to achieve rapid measurement of soil organic matter content. Compared with traditional chemical methods, the measurement cycle is greatly shortened and the measurement efficiency is improved. Compared with traditional detection equipment, the near-infrared spectrometer is lighter and easier to operate, and does not require complex sample pretreatment process, which reduces the measurement cost. Through data preprocessing, feature extraction and the establishment of a deep learning model, the present invention can more accurately predict the soil organic matter content, and through humidity and temperature correction, the accuracy of the measurement results is further improved. The method and system of the present invention are applicable to various soil types and agricultural production environments, and can be easily promoted and applied in agricultural production, providing strong support for precision fertilization and scientific management.

[0015] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 Schematic diagram of a process for determining soil organic matter in the present invention; Figure 2 Schematic diagram of the specific process of the S3 data preprocessing step in the present invention; Figure 3 The figure is a schematic diagram of the specific process of establishing the S5 model in the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0019] Example 1 like Figure 1 The present invention provides a method for determining soil organic matter, comprising: S1. Sample preparation: Collect soil samples from the area to be tested, mix, dry, grind and sieve to ensure sample homogeneity; S2. Spectral acquisition: Use a near-infrared spectrometer to collect spectra of the prepared soil samples to obtain spectral data of the samples; S3. Data preprocessing: The collected spectral data are subjected to outlier removal, multivariate scattering correction, standard normal transformation, detrending, and feature band extraction to obtain preprocessed spectral data; S4, feature extraction: extracting feature information related to soil organic matter content from the preprocessed spectral data; S5. Model building: Build an organic matter prediction model and use the known Lucas soil public dataset to train and validate the organic matter prediction model. S6. Humidity and temperature correction: Collect humidity and temperature information of soil samples and establish a humidity / temperature correction model through mathematical modeling methods to correct the measurement results; S7. Sample determination: Input the spectral data of the soil sample to be tested into the established digital organic matter prediction model and humidity / temperature correction model to obtain the organic matter content of the sample.

[0020] like Figure 1 As shown, the soil organic matter determination method provided by the present invention follows the following detailed steps when it is implemented: S1. Sample preparation: Description of steps: First, collect a certain number of soil samples from the farmland to be tested or the designated area. Pay attention to randomness and representativeness during the collection process to ensure that the samples can truly reflect the soil characteristics of the area to be tested. After collection, the soil samples are thoroughly mixed to eliminate local differences. Next, the mixed soil samples are placed in a drying oven for drying to remove excess water. After drying, the soil samples are ground using a grinder to achieve a certain fineness. Finally, large particles of impurities in the ground soil are removed through sieving to ensure that the samples are uniform and fine. Through this series of pretreatment steps, the uniformity and representativeness of the soil samples are ensured, providing a reliable basis for subsequent spectral acquisition and analysis.

[0021] S2. Spectral acquisition: Step Description: Use a high-precision near-infrared spectrometer to collect spectra from the prepared soil samples. Near-infrared spectrometers can capture the reflected or transmitted spectral information of soil samples within a specific wavelength range. This information is closely related to the chemical composition of the soil (including organic matter). Spectral data of the soil samples is obtained through spectral acquisition, providing data support for subsequent data processing and organic matter content prediction.

[0022] S3. Data preprocessing: Step description: A series of preprocessing operations are performed on the collected spectral data, including outlier removal (removing abnormal data points caused by instrument errors or improper operation), multivariate scattering correction (reducing the scattering effect caused by differences in soil particle size, shape and distribution), standard normal transformation (making the data conform to the normal distribution to facilitate subsequent analysis), detrending (eliminating long-term trends or baseline drift in spectral data) and feature band extraction (screening out the bands most relevant to organic matter content from the full spectral data). Through data preprocessing, the quality and reliability of spectral data are improved, laying a solid foundation for subsequent feature extraction and model establishment.

[0023] S4. Feature extraction: Step description: Advanced algorithms are used to extract characteristic information closely related to soil organic matter content from preprocessed spectral data. These characteristic information are usually expressed as specific spectral bands or combinations of spectral parameters. By extracting these characteristics, the complexity of spectral data is simplified, and the accuracy and efficiency of organic matter content prediction are improved.

[0024] S5. Model establishment: Step Description: Using the publicly available Lucas soil dataset as a training set and combining it with the extracted feature information, we developed an organic matter prediction model. This model predicts soil organic matter content based on input spectral data. During the model development process, we validated and optimized the model to ensure its stability and accuracy. This model allows for rapid prediction of soil organic matter content based on spectral data, providing a scientific basis for soil management and agricultural production.

[0025] S6. Humidity and temperature correction: Step Description: During the measurement process, soil sample humidity and temperature information is collected. Mathematical modeling is then used to establish a humidity / temperature correction model. This model accounts for the effects of humidity and temperature on spectral data, thereby correcting the measurement results. This humidity and temperature correction eliminates the influence of environmental factors on spectral data and organic matter content predictions, improving the accuracy and reliability of the measurement.

[0026] S7. Sample determination: Step description: Input the spectral data of the soil sample to be tested into the established organic matter prediction model and humidity / temperature correction model. The model will calculate the organic matter content of the sample based on the input spectral data and correction parameters. Through sample measurement, the organic matter content information of the soil sample to be tested can be quickly and accurately obtained, providing guidance for agricultural production activities such as soil improvement and fertilizer application.

[0027] In summary, the soil organic matter determination method of this embodiment achieves rapid and accurate determination of soil organic matter content through a series of scientific and systematic steps. This method not only has the advantages of simple operation, high efficiency, and low cost, but also has wide application in fields such as farmland soil management and agricultural production guidance, providing strong support for sustainable agricultural development.

[0028] Example 2 Based on the first embodiment, the difference of this embodiment is that: The S1 sample preparation step specifically includes the following sub-steps: S1.1 Selection of representative sampling points: Based on the research objectives or monitoring requirements, select representative sampling points in the area to be tested that can fully reflect the soil characteristics; S1.2 Sampling operation specifications: Use clean, non-contaminated sampling tools for sampling, remove the surface soil to avoid interference from impurities, determine the sampling depth based on the research purpose, and collect soil samples at a depth of 0-20 cm; S1.3 Sample mixing: In each sampling unit, the soil samples collected at each point are thoroughly mixed to form a mixed sample; S1.3 Drying and grinding: Spread the mixed soil sample on a tray and dry it naturally in a cool and ventilated environment. After drying, use stainless steel or ceramic tools to finely grind it. S1.4 Particle size screening: Based on the analysis requirements, the ground sample is screened to obtain soil particles of different particle size ranges; S1.5 Container selection and labeling information: Use clean, well-sealed containers to store soil samples. Avoid using materials that can chemically react with the soil. Mark the container with key information such as sample number, collection location, and date for subsequent analysis and traceability.

[0029] The S2 spectrum acquisition step specifically includes the following sub-steps: S2.1 Instrument preparation: Before collecting spectral data, preheat and calibrate the near-infrared spectrometer to ensure the accuracy and stability of the measurement results; S2.2 Sample measurement: Place the prepared soil sample in the measuring container of the infrared spectrometer and collect the spectral data of the sample; S2.3 Data recording: The collected spectral data are recorded in the computer to provide a basis for subsequent data preprocessing and sample determination.

[0030] In the S2 spectrum collection step, in order to ensure the accuracy and stability of the spectral data, spectral data of each soil sample was collected at least three times, and the average value was taken as the final spectral data.

[0031] Based on Example 1, this example further refines and optimizes the sample preparation (S1) and spectrum acquisition (S2) steps, as follows: S1. Sample preparation: S1.1 Selection of representative sampling points: Based on research objectives or monitoring needs, a Geographic Information System (GIS) should be used to scientifically select representative sampling points within the study area that fully reflect soil characteristics, taking into account factors such as soil type, topography, and vegetation cover. The number and distribution of sampling points should ensure coverage of the entire study area to avoid missing key areas.

[0032] S1.2 Sampling operation specifications: Use clean and disinfected sampling tools (such as stainless steel shovels and soil augers) to avoid cross-contamination. Remove the topsoil (approximately 5 cm) to minimize contamination from human activities, plant roots, and other impurities. Determine the appropriate sampling depth (0-20 cm in this example) based on the research objective, such as assessing the vertical distribution of soil organic matter.

[0033] S1.3 Sample mixing: In each sampling unit (such as a sampling point or multiple points within a certain area), the soil samples collected from each point are fully mixed in a certain proportion (such as equal weight) to form a mixed sample. During the mixing process, care is taken to avoid external contamination to ensure the representativeness of the sample.

[0034] S1.4 Drying and grinding: Spread the evenly mixed soil sample on a clean tray and place it in a cool and ventilated place to dry naturally to constant weight. Avoid direct sunlight and high temperature during the drying process to prevent the sample from deteriorating. After drying, use stainless steel or ceramic tools to finely grind the sample to ensure that the sample particle size meets the analysis requirements.

[0035] S1.5 Particle size screening: According to the analysis requirements, use a particle size sieve (such as 2mm, 0.5mm, 0.25mm, etc.) to screen the ground sample to obtain soil particles in different particle size ranges. During the screening process, pay attention to maintaining the purity of the sample and avoid cross contamination.

[0036] S1.6 Container selection and labeling information: Use clean, well-sealed containers (such as glass bottles, plastic bottles, etc., avoid using materials that can chemically react with the soil) to store soil samples. Mark the container with key information such as sample number, collection location, date, sampling depth, sampler, etc. for subsequent analysis and traceability.

[0037] S2. Spectral acquisition: S2.1 Instrument Preparation: Before collecting spectral data, the near-infrared spectrometer must be fully warmed up and calibrated. The warm-up time is determined according to the instrument's manual to ensure optimal operating conditions. The calibration process includes wavelength and sensitivity calibration to ensure the accuracy and stability of the measurement results.

[0038] S2.2 Sample Measurement: Spread the prepared soil sample evenly in the infrared spectrometer's measurement container, ensuring the sample is of moderate thickness and evenly distributed. Set the measurement parameters (such as scan count and resolution) according to the instrument's manual, start the measurement program, and collect the sample's spectral data.

[0039] S2.3 Data Recording: Record the collected spectral data in real time on a computer and save it in a standard file format (e.g., CSV, TXT, etc.) for subsequent data preprocessing and sample measurement. To ensure the accuracy and stability of the spectral data, collect spectral data at least three times for each soil sample and take the average value as the final spectral data. When recording data, pay attention to check the integrity and consistency of the data to ensure accuracy.

[0040] This example further improves the accuracy and reliability of the soil organic matter determination method by refining the sample preparation and spectrum acquisition steps. At the same time, through detailed operating specifications and record requirements, the traceability and scientific nature of the entire determination process are ensured. Example 3 Based on the first embodiment, the difference of this embodiment is that: like Figure 2 As shown: The S3 data preprocessing step specifically includes the following sub-steps: S3.1 Spectral principal component calculation: Use principal component analysis (PCA) to reduce the dimensionality of spectral data and extract the main spectral characteristic components; S3.2 Mahalanobis distance calculation: Calculate the Mahalanobis distance between each sample spectral data and the principal component space; S3.3 Plot the Mahalanobis distance of samples and calculate the threshold: Plot the Mahalanobis distance of all samples, observe their distribution, and set a threshold based on the distribution characteristics to separate abnormal data; S3.4 Use Matlab to perform multivariate scatter correction, standard normal transformation and detrending: S3.41 Multivariate Scattering Correction: To address the uneven particle size, shape, and distribution in soil samples, a multivariate scattering correction method is used to reduce the impact of scattering effects on spectral data. S3.42 Standard Normal Transformation: Perform a standard normal transformation on the spectral data to make it conform to the normal distribution characteristics; S3.43 Detrending: Remove the trend component in the spectral data, that is, the trend that gradually increases or decreases with time or wavelength; S3.5 Extract characteristic bands using partial least squares regression: The spectral data are further processed using partial least squares regression (PLSR). By calculating the regression coefficients, the characteristic bands that are most correlated with the soil organic matter content are identified. These characteristic bands will serve as the basis for subsequent modeling and prediction.

[0041] Based on the first embodiment, this embodiment optimizes the data preprocessing step (S3) in more detail and in depth, specifically including the following sub-steps: S3. Data preprocessing: S3.1 Spectral principal component calculation: Principal component analysis (PCA) was used to reduce the dimensionality of the collected spectral data. PCA calculates the correlations between variables in the dataset and extracts the most representative principal components. These principal components maximize the information of the original data while reducing its dimensionality and complexity. In spectral data, principal components typically correspond to different chemical compositions or physical properties. Therefore, PCA can be used to extract spectral signatures related to soil organic matter content.

[0042] S3.2 Mahalanobis distance calculation: Calculate the Mahalanobis distance between each sample spectral data and the principal component space. Mahalanobis distance is a distance measurement method that takes into account the data distribution characteristics. It can effectively identify abnormal data points that differ greatly from the overall data distribution. By calculating the Mahalanobis distance, the degree of abnormality of each sample spectral data can be evaluated, providing a basis for subsequent data cleaning and outlier processing.

[0043] S3.3 Plot the sample Mahalanobis distance and calculate the threshold: The Mahalanobis distances of all samples are plotted to observe their distribution. Based on this distribution, a reasonable threshold is set to isolate abnormal data. The threshold setting requires comprehensive consideration of factors such as the overall distribution of the data, the proportion of abnormal data, and the accuracy requirements of subsequent analysis. By setting a threshold, outliers in spectral data can be effectively removed, improving data reliability and accuracy.

[0044] S3.4 Multivariate scatter correction, standard normal transformation, and detrending were performed using Matlab: S3.41 Multivariate Scattering Correction: To address the problems of uneven particle size, shape and distribution in soil samples, the multivariate scattering correction method is used. The multivariate scattering correction corrects the spectral data by establishing a scattering model to reduce the influence of the scattering effect on the spectral data. The corrected spectral data more accurately reflects the chemical composition and physical properties of the soil samples.

[0045] S3.42 Standard normal transformation: The spectral data is subjected to a standard normal transformation to make it conform to the normal distribution characteristics. The standard normal transformation converts the data into a standard normal distribution form by calculating the mean and standard deviation of the data, thereby eliminating the dimensional differences and uneven distribution problems between the data. The transformed data is more suitable for subsequent statistical analysis and modeling.

[0046] S3.43 Detrending: Remove trend components from spectral data. Trend components refer to trends that gradually increase or decrease over time or wavelength. These components may be caused by factors such as instrument errors, environmental changes, or inhomogeneities in sample preparation. Detrending can eliminate the effects of these trend components on spectral data, improving data stability and accuracy.

[0047] S3.5 Extract characteristic bands using least partial squares regression: The spectral data were further processed using partial least squares regression (PLSR). PLSR is a multiple regression analysis method that can identify the characteristic bands most correlated with soil organic matter content by calculating regression coefficients in the presence of multicollinearity. These characteristic bands will serve as the basis for subsequent modeling and prediction, providing strong support for the rapid and accurate determination of soil organic matter content.

[0048] This example further improves the accuracy and reliability of the soil organic matter determination method by refining the data preprocessing steps and introducing advanced data processing methods such as principal component analysis, Mahalanobis distance calculation, multivariate scatter correction, standard normal transformation, detrending, and least partial squares regression. These methods also provide a more robust and reliable data foundation for subsequent modeling and prediction.

[0049] Example 4 Based on the first embodiment, the difference of this embodiment is that: like Figure 3 As shown in the figure: The S5 model establishment process specifically includes the following sub-steps: S5.1 Build a deep learning model: S5.11 Design a multilayer perceptron (MLP) neural network model consisting of an input layer, at least one hidden layer, and an output layer. The input layer of S5.12 receives the spectral data of the soil, and the output layer outputs the predicted organic matter content; S5.13 Adjust the number of hidden layers and the number of neurons in each layer based on the specific problem and dataset characteristics; S5.14 Select the activation function, ReLU, sigmoid, or tanh, to increase the nonlinearity of the network; S5.15 Define the loss function, mean squared error (MSE) or cross entropy loss, which measures the difference between the network's predictions and the true values. S5.2 Model training: S5.21 Use the training set data to train the neural network model; S5.22 During the training process, forward propagation calculations are performed to calculate the activation values ​​and output values ​​of each layer; S5.23 calculates the difference between the network prediction value and the true value based on the loss function, that is, the loss value; S5.24 uses the backpropagation algorithm to calculate the gradient of the loss function with respect to the weights and biases of each layer; S5.25 Update the network weights and biases using the gradient descent method based on the calculated gradient; S5.26 Repeat the forward propagation, loss calculation, and weight update steps until a predetermined number of training rounds is reached or the loss value converges; S5.3 Model Validation: S5.31 Use the validation set data to validate the trained model and evaluate the model's predictive performance and accuracy; S5.32 optimizes the model based on the verification results, including adjusting the network structure, activation function, loss function and optimizing algorithm parameters.

[0050] Used in S5.24 Type transfer function is used as the transfer function of the hidden layer, using The transfer function outputs a non-negative number and transmits it in the hidden layer. The type transfer function is transferred in the output layer; By finding the total error Adjust the weights between different layers of the error signal back propagation; the specific correlation function formula is as follows; Type transfer function: Type transfer function: Type transfer function: Total error: .

[0051] The S6 humidity and temperature calibration step specifically includes the following steps: S6.1 Data collection: Obtain relevant data of the LUCAS soil dataset as a training set; S6.2 Correction model construction: Using the LUCAS soil dataset containing moisture and temperature information and corresponding organic matter content data, a moisture and temperature correction model was established using multiple linear regression or neural network mathematical modeling methods; S6.3 Calibration model validation: Use an independent validation set to validate the established moisture and temperature calibration model to evaluate the calibration effect and accuracy of the model. The validation set should contain soil samples different from the training set to ensure that the calibration model has good generalization ability.

[0052] This example describes the process of establishing a soil organic matter content prediction model based on deep learning, and details the steps of humidity and temperature correction.

[0053] S5 model building process S5.1 Building a Deep Learning Model Design a multi-layer perceptron (MLP) neural network model: The model consists of an input layer, at least one hidden layer, and an output layer.

[0054] The input layer receives the spectral data of the soil, which is the basis for the model prediction.

[0055] The output layer outputs the predicted organic matter content, which is the target output of the model.

[0056] Adjust the number of hidden layers and the number of neurons in each layer: It depends on the characteristics of the specific problem and dataset. Generally, more complex tasks may require more hidden layers and neurons.

[0057] Select activation function: Activation functions such as ReLU, sigmoid, or tanh are used to increase the nonlinear ability of the network, enabling the model to learn more complex features.

[0058] Define the loss function: Loss functions such as mean squared error (MSE) or cross entropy loss are used to measure the difference between the network's predicted value and the true value, which is the goal of model optimization.

[0059] S5.2 Model Training Use the training set data to train the neural network model: The activation value and output value of each layer are calculated through forward propagation.

[0060] The loss function calculates the difference between the network prediction value and the true value, that is, the loss value.

[0061] Calculate the gradient using the backpropagation algorithm: The backpropagation algorithm is used to calculate the gradient of the loss function with respect to the weights and biases of each layer.

[0062] Update the weights and biases of the network: Based on the calculated gradient, use gradient descent or other optimization algorithms to update the weights and biases of the network.

[0063] The training process is repeated until a predetermined number of training rounds is reached or the loss converges, meaning the model's performance on the training set no longer improves significantly.

[0064] S5.3 Model Validation Validate the model using the validation set data: Evaluate the model's predictive performance and accuracy to ensure that the model performs well on unseen data.

[0065] Optimize the model based on the verification results: including adjusting the network structure, activation function, loss function and optimizing algorithm parameters, etc., to improve the generalization ability of the model.

[0066] S6 Humidity and Temperature Calibration Procedure Data collection: Obtain relevant data from the LUCAS soil dataset as a training set. This data should include information on humidity, temperature, and corresponding organic matter content.

[0067] Correction model construction: Use mathematical modeling methods such as multiple linear regression or neural networks to establish humidity and temperature correction models to eliminate the effects of humidity and temperature on the prediction of organic matter content.

[0068] Calibration model validation: Use an independent validation set to validate the established calibration model to ensure good generalization ability. The validation set should contain soil samples that are different from the training set.

[0069] This example introduces the steps of humidity and temperature correction. Through reasonable model design, training and verification, as well as the construction and verification of the correction model, an accurate and reliable soil organic matter content prediction model can be established. A soil organic matter determination system, used to cooperate with the above-mentioned soil organic matter determination method, includes: Sample preparation device: used for collecting, mixing, drying, grinding and sieving soil samples; Spectral acquisition device: including a near-infrared spectrometer, used to collect spectra of the prepared soil samples; Data processing unit: used for preprocessing, feature extraction, model building and sample measurement of the collected spectral data; Humidity and temperature measuring device: used to collect humidity and temperature information of soil samples; Output device: used to display and output the measurement results, including organic matter content and the results after humidity and temperature correction.

[0070] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0071] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A soil organic matter determination method and system, characterized in that: include: S1. Sample preparation: Collect soil samples from the area to be tested, mix, dry, grind and sieve to ensure sample homogeneity; S2. Spectral acquisition: Use a near-infrared spectrometer to collect spectra of the prepared soil samples to obtain spectral data of the samples; S3. Data preprocessing: The collected spectral data are subjected to outlier removal, multivariate scattering correction, standard normal transformation, detrending, and feature band extraction to obtain preprocessed spectral data; S4, feature extraction: extracting feature information related to soil organic matter content from the preprocessed spectral data; S5. Model building: Build an organic matter prediction model and use the known Lucas soil public dataset to train and validate the organic matter prediction model. S6. Humidity and temperature correction: Collect humidity and temperature information of soil samples and establish a humidity / temperature correction model through mathematical modeling methods to correct the measurement results; S7. Sample determination: Input the spectral data of the soil sample to be tested into the established digital organic matter prediction model and humidity / temperature correction model to obtain the organic matter content of the sample.

2. A soil organic matter determination method according to claim 1, characterized in that, The S1 sample preparation step specifically includes the following sub-steps: S1.1 Selection of representative sampling points: Based on the research objectives or monitoring requirements, select representative sampling points in the area to be tested that can fully reflect the soil characteristics; S1.2 Sampling operation specifications: Use clean, non-contaminated sampling tools for sampling, remove the surface soil to avoid interference from impurities, determine the sampling depth based on the research purpose, and collect soil samples at a depth of 0-20 cm; S1.3 Sample mixing: In each sampling unit, the soil samples collected at each point are thoroughly mixed to form a mixed sample; S1.3 Drying and grinding: Spread the mixed soil sample on a tray and dry it naturally in a cool and ventilated environment. After drying, use stainless steel or ceramic tools to finely grind it. S1.4 Particle size screening: Based on the analysis requirements, the ground sample is screened to obtain soil particles of different particle size ranges; S1.5 Container selection and labeling information: Use clean, well-sealed containers to store soil samples. Avoid using materials that can chemically react with the soil. Mark the container with key information such as sample number, collection location, and date for subsequent analysis and traceability.

3. A soil organic matter determination method according to claim 1, characterized in that, The S2 spectrum acquisition step specifically includes the following sub-steps: S2.1 Instrument preparation: Before collecting spectral data, preheat and calibrate the near-infrared spectrometer to ensure the accuracy and stability of the measurement results; S2.2 Sample measurement: Place the prepared soil sample in the measuring container of the infrared spectrometer and collect the spectral data of the sample; S2.3 Data recording: The collected spectral data are recorded in the computer to provide a basis for subsequent data preprocessing and sample determination.

4. A soil organic matter determination method according to claim 3, characterized in that, In the S2 spectrum collection step, in order to ensure the accuracy and stability of the spectrum data, spectrum data of each soil sample is collected at least three times, and the average value is taken as the final spectrum data.

5. A soil organic matter determination method according to claim 1, characterized in that, The S3 data preprocessing step specifically includes the following sub-steps: S3.1 Spectral principal component calculation: Use principal component analysis (PCA) to reduce the dimensionality of spectral data and extract the main spectral characteristic components; S3.2 Mahalanobis distance calculation: Calculate the Mahalanobis distance between each sample spectral data and the principal component space; S3.3 Plot the Mahalanobis distance of samples and calculate the threshold: Plot the Mahalanobis distance of all samples, observe their distribution, and set a threshold based on the distribution characteristics to separate abnormal data; S3.4 Use Matlab to perform multivariate scatter correction, standard normal transformation and detrending: S3.41 Multivariate Scattering Correction: To address the uneven particle size, shape, and distribution in soil samples, a multivariate scattering correction method is used to reduce the impact of scattering effects on spectral data. S3.42 Standard Normal Transformation: Perform a standard normal transformation on the spectral data to make it conform to the normal distribution characteristics; S3.43 Detrending: Remove the trend component in the spectral data, that is, the trend that gradually increases or decreases with time or wavelength; S3.5 Extract characteristic bands using partial least squares regression: The spectral data are further processed using partial least squares regression (PLSR). By calculating the regression coefficients, the characteristic bands that are most correlated with the soil organic matter content are identified. These characteristic bands will serve as the basis for subsequent modeling and prediction.

6. A soil organic matter determination method according to claim 1, characterized in that, The S5 model establishment process specifically includes the following sub-steps: S5.1 Build a deep learning model: S5.11 Design a multilayer perceptron (MLP) neural network model consisting of an input layer, at least one hidden layer, and an output layer. The input layer of S5.12 receives the spectral data of the soil, and the output layer outputs the predicted organic matter content; S5.13 Adjust the number of hidden layers and the number of neurons in each layer based on the specific problem and dataset characteristics; S5.14 Select the activation function, ReLU, sigmoid, or tanh, to increase the nonlinearity of the network; S5.15 Define the loss function, mean squared error (MSE) or cross entropy loss, which measures the difference between the network's predictions and the true values. S5.2 Model training: S5.21 Use the training set data to train the neural network model; S5.22 During the training process, forward propagation calculations are performed to calculate the activation values ​​and output values ​​of each layer; S5.23 calculates the difference between the network prediction value and the true value based on the loss function, that is, the loss value; S5.24 uses the backpropagation algorithm to calculate the gradient of the loss function with respect to the weights and biases of each layer; S5.25 Update the network weights and biases using the gradient descent method based on the calculated gradient; S5.26 Repeat the forward propagation, loss calculation, and weight update steps until a predetermined number of training rounds is reached or the loss value converges; S5.3 Model Validation: S5.31 Use the validation set data to validate the trained model and evaluate the model's predictive performance and accuracy; S5.32 optimizes the model based on the verification results, including adjusting the network structure, activation function, loss function and optimizing algorithm parameters.

7. A soil organic matter determination method according to claim 6, characterized in that, The S5.24 is used Type transfer function is used as the transfer function of the hidden layer, using The transfer function outputs a non-negative number and transmits it in the hidden layer. The type transfer function is transferred in the output layer; By finding the total error Adjust the weights between different layers of the error signal back propagation; the specific correlation function formula is as follows; Type transfer function: Type transfer function: Type transfer function: Total error: .

8. A soil organic matter determination method according to claim 1, characterized in that: The humidity and temperature correction step S6 specifically includes the following steps: S6.1 Data collection: Obtain relevant data of the LUCAS soil dataset as a training set; S6.2 Correction model construction: Using the LUCAS soil dataset containing moisture and temperature information and corresponding organic matter content data, a moisture and temperature correction model was established using multiple linear regression or neural network mathematical modeling methods; S6.3 Calibration model validation: Use an independent validation set to validate the established moisture and temperature calibration model to evaluate the calibration effect and accuracy of the model. The validation set should contain soil samples different from the training set to ensure that the calibration model has good generalization ability.

9. A soil organic matter determination system, characterized in that: Used in conjunction with the soil organic matter determination method according to claims 1 to 8 above, the system comprises: Sample preparation device: used for collecting, mixing, drying, grinding and sieving soil samples; Spectral acquisition device: including a near-infrared spectrometer, used to collect spectra of the prepared soil samples; Data processing unit: used for preprocessing, feature extraction, model building and sample measurement of the collected spectral data; Humidity and temperature measuring device: used to collect humidity and temperature information of soil samples; Output device: used to display and output the measurement results, including organic matter content and the results after humidity and temperature correction.

Citation Information

Cited By

  • Penthorum chinense pursh seed inspection data correction method based on transfer learning

    CN121278494A

  • A method for correcting seed inspection data of *Gynura divaricata* based on transfer learning

    CN121278494B