A soil organic matter content prediction method based on multi-source data

By combining multi-source data with remote sensing, topographic, climate, and hydrological data, the problem of poor prediction accuracy of soil organic matter in black soil areas has been solved, achieving efficient and accurate prediction of soil organic matter content and supporting sustainable agricultural development.

CN118522366BActive Publication Date: 2026-07-31NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S
Filing Date
2024-05-11
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies have poor accuracy in predicting soil organic matter content in black soil regions, especially in areas with complex farmland types where topographic factors have a significant impact, resulting in low prediction accuracy.

Method used

This study employs multi-source data, combining remote sensing, topographic, climate, and hydrological data. Data is downloaded from the Google Earth Engine platform and preprocessed in ENVI and SAGA GIS. Topographic and hydrological parameters are calculated using ArcGIS. A soil organic matter prediction model is established using the random forest algorithm, and 10-fold cross-validation is performed. MATLAB is used for inversion and mapping to achieve accurate predictions.

Benefits of technology

It improves the accuracy of soil organic matter content prediction, reflects the importance of topographic parameters in prediction, adapts to different types of arable land, and enables efficient and accurate mapping of soil organic matter spatial distribution, supporting land management and sustainable agricultural development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118522366B_ABST
    Figure CN118522366B_ABST
Patent Text Reader

Abstract

This invention relates to a method for predicting soil organic matter content based on multi-source data. This invention addresses the problem of poor accuracy in existing soil organic matter prediction methods. The method includes: 1. Remote sensing image data and meteorological data of the study area; 2. Data preprocessing; 3. Calculation of topographic parameters; 4. Calculation of hydrological parameters; 5. Accuracy assessment; 6. Saving of soil organic matter prediction results; 7. Obtaining a spatial distribution map of soil organic matter. This invention constructs models for environmental variables and spectral characteristic indices in different cultivated land types for comparison, achieving regional inversion and mapping based on differences in cultivated land types, thereby improving the accuracy of soil organic matter prediction. This invention produces accurate spatial distribution maps of soil organic matter, enabling accurate and efficient prediction of soil organic matter content, which is of great significance for land management and agriculture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for predicting organic matter content. Background Technology

[0002] Soil organic matter (SOM) is a crucial component of the soil carbon pool, playing a vital role in global carbon cycling, soil health, and food security. Black soil, as a valuable soil resource, has its SOM content serving as an important indicator of soil fertility quality. However, in recent decades, excessive human exploitation has led to significant losses of organic matter and nutrients in black soil regions, resulting in severe soil degradation (Hu, 2020). Therefore, accurate prediction of SOM content in black soil regions is fundamental to promoting sustainable agricultural development, enhancing soil carbon sequestration potential, and mitigating global climate change.

[0003] Traditional soil organic matter monitoring mainly relies on field sampling and indoor chemical analysis for inversion, a method that suffers from long cycles and low efficiency (Dong, 2021). Currently, with the development of remote sensing (RS) technology, multispectral remote sensing has gradually become an effective means of obtaining SOM content. Multispectral satellite remote sensing has advantages such as low cost and wide coverage, enabling real-time continuous observation of the Earth's surface and is widely used in soil monitoring. Zhou et al. (2021) combined optical and radar remote sensing data to predict and spatially map soil organic carbon (SOC).

[0004] Traditional soil organic matter (SOM) mapping primarily relies on spectral data modeling, which involves combining different spectral bands from sensors and establishing an SOM content estimation model using linear or nonlinear modeling methods. Besides remote sensing imagery, the spatial distribution of SOM is influenced by various environmental factors such as climate, topography, soil properties, and vegetation. With the advent of Geographic Information Systems (GIS) and Remote Sensing (RS) technologies, high-resolution climate, topography, and soil type data are readily available and used to estimate SOM content (Kumar et al., 2012). Many researchers combine environmental variables with remote sensing imagery to construct SOM estimation models, thereby improving model accuracy. Lotfollahi et al. (2023) effectively improved the predictive ability of SOM content in semi-arid mountainous areas by using topographic and spectral variables for modeling. Katebikord et al. (2022) studied the correlation between soil organic matter (SOC) and environmental variables and combined remote sensing imagery with environmental variables to map the spatial distribution of SOC content. However, in areas with complex arable land types, various arable land type factors have different effects on the spatial distribution of SOM in different topographic regions, thus affecting the accuracy of soil organic matter prediction. Summary of the Invention

[0005] To address the problem of poor accuracy in existing soil organic matter prediction methods, this invention provides a method for predicting soil organic matter content based on multi-source data, which is a soil organic matter mapping method based on remote sensing, topography, climate and hydrological data.

[0006] The method for predicting soil organic matter content based on multi-source data of this invention is carried out according to the following steps:

[0007] Step 1: Download remote sensing image data and digital elevation model data of the study area from the Google Earth Engine platform, and download meteorological data of the study area from the National Meteorological Information Center.

[0008] Step 2: In ENVI, preprocess the remote sensing image data, digital elevation model data, and meteorological data from Step 1, including radiometric correction, atmospheric correction, image cropping, and stitching.

[0009] Step 3: Calculate a set of terrain parameters using the terrain analysis function in SAGA GIS and the preprocessed data from Step 2;

[0010] Step 4: Calculate a set of hydrological parameters using the line density tool in ArcGIS 10.6 and the preprocessed data from Step 2;

[0011] Step 5: Divide the study area into paddy field area and dry field area. In both dry field and paddy field areas, use different combinations of variables as input variables and use the ten-fold cross-validation method to evaluate the accuracy.

[0012] Step 6: Combine the Random Forest (RF) method to establish a soil organic matter prediction model based on remote sensing image data, topographic parameters, meteorological data, and hydrological parameters. Adjust and determine the optimal number of decision trees, and use the 10-fold cross-validation method to verify the stability of the model. After confirming the model's stability, save its output soil organic matter prediction results.

[0013] Step 7: Based on the prediction results of Step 6, use MATLAB to perform inversion and obtain the spatial distribution map of soil organic matter; that is, the prediction of soil organic matter content is realized.

[0014] This invention utilizes data sources from the Google Earth Engine cloud platform (https: / / earthengine.google.com / ), selects ground feature points, and then constructs an input feature set based on remote sensing, topography, climate, and hydrological data. It integrates optical image features and combines the Random Forest (RF) algorithm from machine learning to establish a soil organic matter (SOM) prediction model. This invention predicts SOM based on remote sensing, topography, climate, and hydrological data, taking into account the influence of hydrology on soil SOM content, enriching the types of environmental variables, and thus improving the prediction accuracy. Simultaneously, it demonstrates the importance of topographic parameters for SOM prediction; combining topographic-related variables with other environmental covariates helps improve the accuracy of SOM content prediction, thereby enhancing the prediction precision of soil organic matter content.

[0015] The method of this invention constructs models for environmental variables and spectral characteristic indices in different cultivated land types and compares them to achieve regional inversion and mapping based on differences in cultivated land types, thereby improving the accuracy of soil organic matter prediction.

[0016] This invention produces a precise spatial distribution map of soil organic matter, which can accurately and efficiently predict soil organic matter content, and is of great significance for land management and agriculture. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the framework of the method in Example 1;

[0018] Figure 2 This is a spatial distribution map of soil organic matter content in the region between 131°27′50″ and 132°15′ east longitude and 46°28′14″ and 46°59′38″ north latitude. Detailed Implementation

[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0021] Specific Implementation Method 1: This implementation method for predicting soil organic matter content based on multi-source data follows these steps:

[0022] Step 1: Download remote sensing image data and digital elevation model data of the study area from the Google Earth Engine platform, and download meteorological data of the study area from the National Meteorological Information Center.

[0023] Step 2: In ENVI, preprocess the remote sensing image data, digital elevation model data, and meteorological data from Step 1, including radiometric correction, atmospheric correction, image cropping, and stitching.

[0024] Step 3: Calculate a set of terrain parameters using the terrain analysis function in SAGA GIS and the preprocessed data from Step 2;

[0025] Step 4: Calculate a set of hydrological parameters using the line density tool in ArcGIS 10.6 and the preprocessed data from Step 2;

[0026] Step 5: Divide the study area into paddy field area and dry field area. In both dry field and paddy field areas, use different combinations of variables as input variables and use the ten-fold cross-validation method to evaluate the accuracy.

[0027] Step 6: Combine the Random Forest (RF) method to establish a soil organic matter prediction model based on remote sensing image data, topographic parameters, meteorological data, and hydrological parameters. Adjust and determine the optimal number of decision trees, and use the 10-fold cross-validation method to verify the stability of the model. After confirming the model's stability, save its output soil organic matter prediction results.

[0028] Step 7: Based on the prediction results of Step 6, use MATLAB to perform inversion and obtain the spatial distribution map of soil organic matter; that is, the prediction of soil organic matter content is realized.

[0029] In step five of this implementation method, in order to evaluate the performance of environmental covariates and remote sensing data in predicting SOM content in different cultivated land types, the sample is partitioned and inverted, and the study area is divided into paddy field area and dry field area. Different combinations of variables are used as input variables in both dry field and paddy field areas, and the accuracy is evaluated by using the ten-fold cross-validation method.

[0030] The remote sensing image data, topographic parameters, meteorological data, and hydrological parameters in this embodiment specifically include: Band1_4, Band2_4, Band3_4, Band4_4, Band5_4, Band6_4, Band1_5, Band2_5, Band3_5, Band4_5, Band5_5, Band6_5, catchment area (CA), profile curvature (PC), general curvature (GC), slope (SL), terrain ruggedness index (TRI), terrain moisture index (TWI), valley depth (VD), vector terrain roughness (VTR), drainage density (DD), drainage frequency (DF), annual precipitation, and annual average temperature.

[0031] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that: in step one, two periods of Landsat-8 satellite images are extracted based on the Google Earth Engine platform, and the band information is extracted as band features using these images as the data source; other steps and parameters are the same as in Specific Implementation Method One.

[0032] Band information for this implementation: Band1_4, Band2_4, Band3_4, Band4_4, Band5_4, Band6_4, Band1_5, Band2_5, Band3_5, Band4_5, Band5_5, Band6_5.

[0033] Specific Implementation Method 3: This implementation method differs from Specific Implementation Method 1 in that the terrain parameters calculated using SAGAGIS in step 3 are: catchment area (CA), profile curvature (PC), general curvature (GC), slope (SL), terrain ruggedness index (TRI), terrain moisture index (TWI), valley depth (VD), and vector terrain roughness (VTR); the other steps and parameters are the same as in Specific Implementation Method 1.

[0034] Specific Implementation Method Four: This implementation method differs from Specific Implementation Method One in that the hydrological parameters calculated using ArcGIS in step four are: drainage density (DD) and drainage frequency (DF); the other steps and parameters are the same as in Specific Implementation Method One.

[0035] Specific Implementation Method 5: This implementation method differs from Specific Implementation Method 1 in that: in step six, during the RF model construction process, ntree is set to 500 and mtry is 1 / 3 of the number of input quantities; other steps and parameters are the same as in Specific Implementation Method 1.

[0036] The numerical values ​​in this embodiment are relatively small and the errors within the model are basically stable.

[0037] Example 1 This example utilizes the method of the present invention for predicting soil organic matter based on remote sensing, topography, climate, and hydrological data (the framework flowchart of this example is shown below). Figure 1 (As shown) Follow these steps:

[0038] Step 1: Download remote sensing image data and digital elevation model data of the study area from the Google Earth Engine platform, and download meteorological data of the study area from the National Meteorological Information Center.

[0039] Step 2: In ENVI, preprocess the remote sensing image data, digital elevation model data, and meteorological data from Step 1, including radiometric correction, atmospheric correction, image cropping, and stitching.

[0040] Step 3: Calculate a set of terrain parameters using the terrain analysis function in SAGA GIS and the preprocessed data from Step 2;

[0041] Step 4: Calculate a set of hydrological parameters using the line density tool in ArcGIS 10.6 and the preprocessed data from Step 2;

[0042] Step 5: Divide the study area into paddy field area and dry field area. In both dry field and paddy field areas, use different combinations of variables as input variables and use the ten-fold cross-validation method to evaluate the accuracy.

[0043] Step 6: Using the Random Forest (RF) method, establish a soil organic matter prediction model based on remote sensing image data, topographic parameters, meteorological data, and hydrological parameters (based on several feature variables: Band1_4, Band2_4, Band3_4, Band4_4, Band5_4, Band6_4, Band1_5, Band2_5, Band3_5, Band4_5, Band5_5, Band6_5, catchment area (CA), profile curvature (PC), general curvature (GC), slope (SL), terrain ruggedness index (TRI), terrain moisture index (TWI), valley depth (VD), vector terrain roughness (VTR), drainage density (DD), drainage frequency (DF), annual precipitation, and annual average temperature). Adjust and determine the optimal number of decision trees, and use the 10-fold cross-validation method to verify the stability of the model. After confirming the model's stability, save its output soil organic matter prediction results.

[0044] Step 7: Based on the prediction results of Step 6, use MATLAB to perform inversion and obtain the spatial distribution map of soil organic matter; that is, the prediction of soil organic matter content is realized.

[0045] Step S1 is based on Google Earth. The Engine platform extracted two phases of Landsat-8 satellite imagery and used these imagery as the data source to extract band information (Band1_4, Band2_4, Band3_4, Band4_4, Band5_4, Band6_4, Band1_5, Band2_5, Band3_5, Band4_5, Band5_5, Band6_5) as band features. Step S3 used SAGAGIS to calculate topographic parameters such as catchment area (CA), profile curvature (PC), general curvature (GC), slope (SL), terrain ruggedness index (TRI), terrain moisture index (TWI), valley depth (VD), and vector terrain roughness (VTR). Step S4 used ArcGIS to calculate hydrological parameters such as drainage density (DD) and drainage frequency (DF). Step S5, based on differences in cultivated land type, divided the study area into paddy field area and dryland area, assessed the impact of regional environmental covariates of different cultivated land types on soil organic matter prediction, and used the ten-fold cross-validation method to evaluate the accuracy.

[0046] Example 2: In May 2023, the method of Example 1 was used to predict the soil organic matter content in Study Area 1, Youyi County, Shuangyashan City, Heilongjiang Province.

[0047] Step S1: Download remote sensing image data and digital elevation model data of Youyi County, Shuangyashan City, Heilongjiang Province from the Google Earth Engine platform, and download meteorological data of Youyi County from the National Meteorological Information Center; that is, extract two periods of Landsat-8 satellite images based on the Google Earth Engine platform, and extract band information (Band1_4, Band2_4, Band3_4, Band4_4, Band5_4, Band6_4, Band1_5, Band2_5, Band3_5, Band4_5, Band5_5, Band6_5) as band features using these images as the data source;

[0048] Step S2: In ENVI, the downloaded remote sensing image data, digital elevation model data, and meteorological data are preprocessed, including radiometric correction, atmospheric correction, image cropping, and mosaicking. The remote sensing image data and meteorological data undergo radiometric correction, atmospheric correction, image cropping, and mosaicking. The digital elevation model data needs to be preprocessed using ESR ArcGIS 10.6 to fill depressions and mosaic, and then cropped according to the study area and cultivated land area.

[0049] Step S3: Use SAGA GIS to input digital elevation model data to calculate a set of terrain parameters, including catchment area (CA), profile curvature (PC), general curvature (GC), slope (SL), terrain ruggedness index (TRI), terrain moisture index (TWI), valley depth (VD), and vector terrain roughness (VTR).

[0050] Step S4: Calculate a set of hydrological parameters using ArcGIS, including drainage density (DD) and drainage frequency (DF);

[0051] Step S5: In order to evaluate the performance of environmental covariates and remote sensing data in predicting SOM content in different cultivated land types, we performed zonal inversion of the samples, using different combinations of variables as input variables in both dryland and paddy fields, and used the ten-fold cross-validation method to evaluate the accuracy.

[0052] Step S6: Using the Random Forest method, a soil organic matter prediction model based on remote sensing, topography, climate, and hydrological data is established. The optimal number of decision trees is determined, and the stability of the model is verified using a 10-fold cross-validation method. After confirming model stability, the output soil organic matter prediction results are saved. The Random Forest (RF) method can be implemented in the R programming environment using packages such as randomForest and caret. The randomForest() function has two crucial parameters that affect model accuracy: mtry and ntree. Generally, mtry is selected by trying different values ​​until a suitable value is found. The value of ntree can be roughly determined graphically when the model's error is stable. In this study, during the RF model construction process, ntree was set to 500, and mtry was set to 1 / 3 of the number of input variables. At this value, the values ​​are small, and the model's error is basically stable.

[0053] Step S7: Based on the prediction results of S6, MATLAB is used to make predictions to obtain the spatial distribution map of soil organic matter in Youyi County, Shuangyashan City, Heilongjiang Province.

[0054] In step S1, two phases of Landsat-8 satellite images are extracted based on the Google Earth Engine platform. Using these images as the data source, band information (Band1_4, Band2_4, Band3_4, Band4_4, Band5_4, Band6_4, Band1_5, Band2_5, Band3_5, Band4_5, Band5_5, Band6_5) is extracted as band features.

[0055] Accuracy Validation: To validate the predictive performance of the output results, we used K-fold cross-validation (CV) to train the model and evaluate its performance. CV is a practical method for splitting data samples into smaller subsets. This procedure divides the initial sample into K subsamples. One subsample is used to retain the data as validation data for the model, and the other K-1 samples are used for training. Cross-validation is repeated K times, validating each subsample once. The results from the K cross-validations are averaged to obtain a single estimate. In this study, we used 10-fold cross-validation and trained the soil organic matter content prediction model by region. For overall regression, each iteration generated 162 training samples and 18 validation samples; for regional regression, each iteration generated 95 training samples and 11 validation samples in mountainous areas and 67 training samples and 7 validation samples in plains areas. We used two metrics for model evaluation: root mean square error (RMSE) and coefficient of determination (R²). 2 R 2 The RMSE is used to evaluate the stability of the model, with higher values ​​indicating higher stability; the RMSE is used to evaluate the consistency between the model's predictions and the observed values, with smaller RMSE values ​​indicating higher model accuracy.

[0056] Figure 2 This is a spatial distribution map of soil organic matter content in the region between 131°27′50″ and 132°15′ east longitude and 46°28′14″ and 46°59′38″ north latitude. Figure 2 The figure shows the spatial distribution of SOM content predicted using 180 sampling points across the entire study area and by modeling the optimal variable combination across these 180 sampling points in different regions. The figure indicates that the predicted SOM content is higher in the northeastern part of the study area and lower in the central and southwestern parts, consistent with the results of my country's Second National Soil Survey.

[0057] The method of this invention can be widely applied in agricultural production, accurately and efficiently monitoring and analyzing soil organic matter content, and making land management and related decisions based on the analysis results. Relevant departments can adopt more sustainable agricultural practices and reduce soil erosion by accurately detecting the spatial distribution of soil organic matter content, thus promoting agricultural development. The mapping results can also provide a data foundation for agriculture and land management.

Claims

1. A method for predicting soil organic matter content based on multi-source data, characterized in that The method for predicting soil organic matter content based on multi-source data follows these steps: Step 1: Download remote sensing image data and digital elevation model data of the study area from the Google Earth Engine platform, and download meteorological data of the study area from the National Meteorological Information Center. Step 2: In ENVI, preprocess the remote sensing image data, digital elevation model data, and meteorological data from Step 1, including radiometric correction, atmospheric correction, image cropping, and stitching. Step 3: Calculate a set of terrain parameters using the terrain analysis function in SAGA GIS and the preprocessed data from Step 2; Step 4: Calculate a set of hydrological parameters using the line density tool in ArcGIS 10.6 and the preprocessed data from Step 2; Step 5: Divide the study area into paddy field area and dry field area. In both dry field and paddy field areas, use different combinations of variables as input variables and use the ten-fold cross-validation method to evaluate the accuracy. Step 6: Combine the Random Forest (RF) method to establish a soil organic matter prediction model based on remote sensing image data, topographic parameters, meteorological data, and hydrological parameters. Adjust and determine the optimal number of decision trees, and use the 10-fold cross-validation method to verify the stability of the model. After confirming the model's stability, save its output soil organic matter prediction results. Step 7: Based on the prediction results of Step 6, use MATLAB to perform inversion and obtain the spatial distribution map of soil organic matter; that is, the prediction of soil organic matter content is realized. 2.The method of claim 1, wherein In step one, two phases of Landsat-8 satellite imagery were extracted using the Google Earth Engine platform, and band information was extracted from these imagery as the data source to serve as band features. 3.The method of claim 1, wherein In step three, the terrain parameters calculated using SAGA GIS are: catchment area (CA), profile curvature (PC), general curvature (GC), slope (SL), terrain ruggedness index (TRI), terrain moisture index (TWI), valley depth (VD), and vector terrain roughness (VTR). 4.The method of claim 1, wherein In step four, the hydrological parameters calculated using ArcGIS are: drainage density (DD) and drainage frequency (DF). 5.The method of claim 1, wherein In step six, during the RF model building process, ntree is set to 500 and mtry is 1 / 3 of the number of input variables.