Soil total potassium content estimation method and device based on hyperspectral remote sensing

Through ZY1-02D ​​hyperspectral remote sensing and AdaBoost integrated learning method, combined with terrain, climate, and vegetation variables, the low-precision and high-cost problems of soil full potassium content prediction in traditional methods are solved, and a low-cost and high-precision large-scale soil full potassium spatial distribution mapping is achieved.

CN120526883APending Publication Date: 2025-08-22ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510516897.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve low-cost, high-precision, and large-scale spatial distribution prediction of soil full potassium content. The traditional method has a long cycle, high cost and pollutes the environment, and the accuracy of the machine learning model is low.

Method used

Using ZY1-02D ​​hyperspectral remote sensing data, combined with terrain, climate and vegetation environment variables, a soil full potassium content estimation model was constructed through partial least squares regression and AdaBoost integrated learning method, and overfitting was overcome by bootstrap and random feature selection mechanisms, enhancing model differences, and achieving high-precision prediction.

Benefits of technology

It improves the prediction accuracy of the machine learning model, realizes high-precision, large-scale spatial distribution mapping of soil full potassium, reduces costs, and is suitable for the needs of smart agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526883A_ABST
    Figure CN120526883A_ABST
Patent Text Reader

Abstract

The invention discloses a soil total potassium content estimation method and device based on hyperspectral remote sensing. Low-cost, high-precision and large-scale total potassium spatial distribution prediction can be realized by virtue of spaceflight hyperspectral data. The method comprises the following steps: constructing a model: adopting partial least square regression, and carrying out dimensionality reduction on an independent variable by searching a new orthogonal projection direction; input variables and soil characteristic variables are regressed through decision rules, bootstrap is adopted to overcome the overfitting problem, a random characteristic selection mechanism is introduced, and in the process of building each decision tree, only part of the characteristic variables are considered for division, and the difference between the decision trees is increased; combining a plurality of weak learners into a strong regression device through an AdaBoost ensemble learning method, calculating the weight of each weak regression device according to the prediction precision of each sample in each training set and the last total prediction precision, and updating the distribution weight of each sample at the same time; and finally, carrying out weighted summation on regression device results obtained by training each time and outputting results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and machine learning, and in particular to a method for estimating total potassium content in soil based on hyperspectral remote sensing, and a device for estimating total potassium content in soil based on hyperspectral remote sensing. Background Art

[0002] Total Potassium (TK) in soil is a crucial component of soil nutrients and a key indicator for evaluating farmland soil fertility and crop fertilization. In agricultural production, precise control of soil nutrient abundance is crucial for increasing grain yields. However, the distribution of soil TK is not spatially homogeneous, making it difficult to manage fertilization based on standardized empirical knowledge. Rapid and accurate monitoring of soil TK is a key area of ​​research and a key to implementing precision agriculture. To achieve this goal, numerous researchers have leveraged remote sensing technology to analyze the spatial distribution of soil TK, which holds important implications for precision fertilization, the rational and efficient use of agricultural resources, and the protection of farmland soil environments. Traditional methods for determining soil TK are not only time-consuming and costly, but also cause pollution and damage to the surface environment, making them inadequate for smart agriculture. Therefore, developing a high-precision, large-scale method for estimating soil TK is crucial.

[0003] In recent years, the emergence and application of hyperspectral remote sensing technology has provided new approaches and methods for rapidly and effectively estimating total soil potassium content. Compared to multispectral satellites, hyperspectral satellite remote sensing data offers the advantages of wide spectral coverage and high spectral resolution, enabling the detection of subtle differences in soil properties. The domestically produced hyperspectral Resource 1-02D ​​satellite (ZY1-02D) provides a powerful tool for global resource monitoring and environmental research, achieving high accuracy in estimating soil organic matter (SOM) and heavy metal content. Estimating soil properties using remote sensing data alone ignores the spatial autocorrelation of soil properties and the influence of environmental factors. Therefore, selecting appropriate environmental covariates is crucial for model fitting accuracy.

[0004] The complex physical and chemical composition of the soil increases the complexity of the spectrum, making it difficult to describe the relationship between soil properties and spectral reflectance linearly. Machine learning algorithms can solve these problems well. Machine learning has the advantage of extracting features from multidimensional data and is widely used in soil property prediction. Adaptive Boosting (AdaBoost) is an integrated learning method. The core is to combine multiple weak learners into a strong regressor with strong training capabilities. The weight of each weak regressor is calculated based on the prediction accuracy of each sample in each training set and the overall prediction accuracy of the last time, thereby improving the overall performance and stability. Researchers have conducted many estimates of total potassium content in soil with the help of laboratory spectra and multispectral images, but the accuracy of the constructed models is low and the results are subject to large uncertainty. Some people have analyzed the spatial distribution of exchangeable potassium in soil based on multispectral remote sensing images of different resolutions, and the fitting accuracy of the model is R 2 In order to reveal the influence of the time difference between soil sampling and remote sensing satellite images, some people used long-term Landsat images to predict soil total potassium content. The best model R constructed in the study was 2 is 0.52; some people used the Lateritic Red Soil Spectral Library (LRSSL) in Guangdong Province, China to build a soil property prediction model based on visible-near infrared spectroscopy (Vis-NIR), but the prediction performance of pH (Pondus Hydrogenii) and TK was low. Summary of the Invention

[0005] In order to overcome the defects of the existing technology, the technical problem to be solved by the present invention is to provide a method for estimating the total potassium content in soil based on hyperspectral remote sensing, which can achieve low-cost, high-precision, large-scale spatial distribution prediction of total potassium with the help of aerospace hyperspectral data.

[0006] The technical solution of the present invention is: this method for estimating total potassium content in soil based on hyperspectral remote sensing includes the following steps:

[0007] (1) Data acquisition and processing: Collect soil sample data, obtain ZY1-02D ​​remote sensing image data, and select three environmental variable data: topography, climate, and vegetation;

[0008] (2) Model construction: Partial least squares regression is used to reduce the dimensionality of independent variables by finding new orthogonal projection directions; input variables are regressed with soil characteristic variables through decision rules, bootstrapping is used to overcome the overfitting problem, and a random feature selection mechanism is introduced. In the process of building each decision tree, only some characteristic variables are considered for division to increase the differences between each decision tree; multiple weak learners are combined into a strong regressor through the AdaBoost ensemble learning method, and the weight of each weak regressor is calculated based on the prediction accuracy of each sample in each training set and the overall prediction accuracy of the previous training set. At the same time, the distribution weight of each sample is updated, and finally the weighted sum of the regressor results obtained from each training is output;

[0009] (3) Experimental setup: Experiment 1 is based on environmental variables, Experiment 2 is based on the original hyperspectral bands of ZY1-02D, Experiment 3 is based on the combination of environmental variables and spectral bands, Experiment 4 is based on the variables transformed by the first-order derivative of spectral reflectance, and Experiment 5 is based on the combination of environmental variables and the first-order derivative of spectral reflectance;

[0010] (4) Using the coefficient of determination R 2 , mean absolute error MAE and root mean square error RMSE to evaluate the accuracy of the prediction model;

[0011] (5) Predicting the fitting results of soil total potassium content based on the experimental setup of step (3) and the model of step (2);

[0012] (6) The optimal experimental setting in step (3) and AdaBoost were used to predict the spatial distribution of total potassium in cultivated soil in the study area.

[0013] The present invention uses ZY1-02D ​​hyperspectral remote sensing to detect subtle differences between soil properties, allowing machine learning models to pay more attention to the characteristics of multidimensional data. The constructed AdaBoost model has better accuracy and smaller errors in the training and test sets of soil total potassium content, and more accurate descriptions of nonlinear relationships in the data, which can achieve high-precision spatial distribution mapping. Therefore, the present invention can greatly improve the prediction accuracy of machine learning models, has great potential in the study of soil total potassium distribution, and provides the possibility for low-cost, high-precision, large-scale inversion of total potassium in cultivated land.

[0014] A device for estimating total potassium content in soil based on hyperspectral remote sensing is also provided, the device comprising:

[0015] The data acquisition and processing module collects soil sample data, obtains ZY1-02D ​​remote sensing image data, and selects three environmental variable data: terrain, climate, and vegetation;

[0016] The model construction module uses partial least squares regression to reduce the dimensionality of independent variables by finding new orthogonal projection directions; the input variables are regressed with soil characteristic variables through decision rules, bootstrapping is used to overcome the overfitting problem, and a random feature selection mechanism is introduced. In the process of building each decision tree, only some characteristic variables are considered for division to increase the difference between each decision tree; multiple weak learners are combined into a strong regressor through the AdaBoost ensemble learning method. The weight of each weak regressor is calculated based on the prediction accuracy of each sample in each training set and the overall prediction accuracy of the previous training set, and the distribution weight of each sample is updated at the same time. Finally, the weighted sum of the regressor results obtained from each training is output;

[0017] Experiment setting module, where Experiment 1 is based on environmental variables, Experiment 2 is based on the original hyperspectral bands of ZY1-02D, Experiment 3 is based on the combination of environmental variables and spectral bands, Experiment 4 is based on the variables transformed by the first-order derivative of spectral reflectance, and Experiment 5 is based on the combination of environmental variables and the first-order derivative of spectral reflectance;

[0018] Evaluation module, which uses the coefficient of determination R 2 , mean absolute error MAE and root mean square error RMSE to evaluate the accuracy of the prediction model;

[0019] A fitting prediction module, which predicts the total potassium content in soil based on the fitting results of the experiment setting module and the model building module;

[0020] The distribution prediction module uses the optimal experimental setting and AdaBoost combination to predict the spatial distribution of total potassium in cultivated soil in the study area. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 The geographical location of the study area and the spatial distribution of soil sampling points are shown. (a) Map of Henan Province; (b) Map of Zhengzhou City; (c) Location of soil sampling points in the study area.

[0022] Figure 2 The reflectance spectra of ZY1-02D ​​at different total potassium contents are shown. (a) Original spectral reflectance; (b) First-order inverse spectral reflectance.

[0023] Figure 3 The AdaBoost model framework is shown. For the input variables, the weight distribution of the training samples is first initialized, with each sample having the same weight. Then, weak classifiers are trained, their weights are increased or decreased, and the updated sample set is used to train the next classifier. Finally, all weak classifiers are combined into a strong classifier.

[0024] Figure 4Scatter plots of the total potassium content measured and predicted for the training and test sets using different model combinations are shown. (a)–(e) PLSR; (f)–(j) RF; (k)–(o) AdaBoost.

[0025] Figure 5 Figure 3 shows the spatial distribution of total potassium in cultivated soils in Xinzheng City based on the best model combination experiment 3 and AdaBoost. (a) 2021; (b) 2022.

[0026] Figure 6 The present invention is a flowchart of a method for estimating total potassium content in soil based on hyperspectral remote sensing. DETAILED DESCRIPTION

[0027] like Figure 6 As shown in Figure 2, this method for estimating soil total potassium content based on hyperspectral remote sensing includes the following steps:

[0028] (1) Data acquisition and processing: Collect soil sample data, obtain ZY1-02D ​​remote sensing image data, and select three environmental variable data: topography, climate, and vegetation;

[0029] (2) Model construction: Partial least squares regression is used to reduce the dimensionality of independent variables by finding new orthogonal projection directions; input variables are regressed with soil characteristic variables through decision rules, bootstrapping is used to overcome the overfitting problem, and a random feature selection mechanism is introduced. In the process of building each decision tree, only some characteristic variables are considered for division to increase the differences between each decision tree; multiple weak learners are combined into a strong regressor through the AdaBoost ensemble learning method, and the weight of each weak regressor is calculated based on the prediction accuracy of each sample in each training set and the overall prediction accuracy of the previous training set. At the same time, the distribution weight of each sample is updated, and finally the weighted sum of the regressor results obtained from each training is output;

[0030] (3) Experimental setup: Experiment 1 is based on environmental variables, Experiment 2 is based on the original hyperspectral bands of ZY1-02D, Experiment 3 is based on the combination of environmental variables and spectral bands, Experiment 4 is based on the variables transformed by the first-order derivative of spectral reflectance, and Experiment 5 is based on the combination of environmental variables and the first-order derivative of spectral reflectance;

[0031] (4) Using the coefficient of determination R 2 , mean absolute error MAE and root mean square error RMSE to evaluate the accuracy of the prediction model;

[0032] (5) Predicting the fitting results of soil total potassium content based on the experimental setup of step (3) and the model of step (2);

[0033] (6) The optimal experimental setting in step (3) and AdaBoost were used to predict the spatial distribution of total potassium in cultivated soil in the study area.

[0034] The present invention uses ZY1-02D ​​hyperspectral remote sensing to detect subtle differences between soil properties, allowing machine learning models to pay more attention to the characteristics of multidimensional data. The constructed AdaBoost model has better accuracy and smaller errors in the training and test sets of soil total potassium content, and more accurate descriptions of nonlinear relationships in the data, which can achieve high-precision spatial distribution mapping. Therefore, the present invention can greatly improve the prediction accuracy of machine learning models, has great potential in the study of soil total potassium distribution, and provides the possibility for low-cost, high-precision, large-scale inversion of total potassium in cultivated land.

[0035] Preferably, in step (1), a five-point sampling method is adopted to collect several surface soil samples from the study area at a sampling depth of 0 to 20 cm, and the geographical location of the sampling points is recorded; the soil collected at each sampling point is broken up, debris is picked out, and after thorough mixing, a soil sample is obtained by quartering; the soil sample is then naturally air-dried, and after drying, it is passed through a 10 mm sieve, and the sieved part is used as a soil sample, and each soil sample is placed in a sealed bag; and the total potassium content of the soil is determined by a flame photometer.

[0036] Preferably, in step (1), the remote sensing image data is subjected to radiometric calibration and atmospheric correction, and the digital number DN of the original image is converted into the real surface reflectance SR; the ground spectral reflectance of the soil sample point is extracted from the processed image in ENVI 5.3 according to the geographic coordinates of the sample point; and the original spectral reflectance is subjected to a reflectance first-order derivative transformation.

[0037] Preferably, in step (1), the terrain dataset is derived from the digital elevation dataset of the Space Shuttle Radar Topography Mission of the United States Geological Survey Earth Explorer, and the terrain attributes include altitude, slope and aspect; the climate dataset includes surface temperature, average annual temperature and average annual precipitation, with an original spatial resolution of 1 km and a resampled resolution of 30 m; the vegetation dataset is based on the Landsat 9OLI image obtained from the geospatial data cloud platform, and the normalized vegetation index, difference vegetation index, optimized soil adjusted vegetation index and greenness normalized vegetation index are calculated in ENVI.

[0038] Preferably, in step (2), the RF model in the scikit-learn library is called and the parameters are optimized using a random search method.

[0039] Preferably, in step (2), the AdaBoost model in the scikit-learn library is called, and the parameters are optimized using a grid search method.

[0040] Preferably, in step (4),

[0041]

[0042] Where n is the number of soil samples, y i is the observed value of TK content of sample i, is the predicted value of TK content of sample i.

[0043] Preferably, the optimal experimental setting is Experiment 3.

[0044] Those skilled in the art will appreciate that all or part of the steps in the above-described embodiment method can be accomplished by instructing the relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes the steps of the above-described embodiment method. The storage medium can be: ROM / RAM, a magnetic disk, an optical disk, a memory card, etc. Therefore, corresponding to the method of the present invention, the present invention also includes a soil total potassium content estimation device based on hyperspectral remote sensing. The device is generally represented in the form of functional modules corresponding to the steps of the method. The device includes:

[0045] The data acquisition and processing module collects soil sample data, obtains ZY1-02D ​​remote sensing image data, and selects three environmental variable data: terrain, climate, and vegetation;

[0046] The model construction module uses partial least squares regression to reduce the dimensionality of independent variables by finding new orthogonal projection directions; the input variables are regressed with soil characteristic variables through decision rules, bootstrapping is used to overcome the overfitting problem, and a random feature selection mechanism is introduced. In the process of building each decision tree, only some characteristic variables are considered for division to increase the difference between each decision tree; multiple weak learners are combined into a strong regressor through the AdaBoost ensemble learning method. The weight of each weak regressor is calculated based on the prediction accuracy of each sample in each training set and the overall prediction accuracy of the previous training set, and the distribution weight of each sample is updated at the same time. Finally, the weighted sum of the regressor results obtained from each training is output;

[0047] Experiment setting module, where Experiment 1 is based on environmental variables, Experiment 2 is based on the original hyperspectral bands of ZY1-02D, Experiment 3 is based on the combination of environmental variables and spectral bands, Experiment 4 is based on the variables transformed by the first-order derivative of spectral reflectance, and Experiment 5 is based on the combination of environmental variables and the first-order derivative of spectral reflectance;

[0048] Evaluation module, which uses the coefficient of determination R 2 , mean absolute error MAE and root mean square error RMSE to evaluate the accuracy of the prediction model;

[0049] A fitting prediction module, which predicts the total potassium content in soil based on the fitting results of the experiment setting module and the model building module;

[0050] The distribution prediction module uses the optimal experimental setting and AdaBoost combination to predict the spatial distribution of total potassium in cultivated soil in the study area.

[0051] Preferably, in the data acquisition and processing module, a five-point sampling method is used to collect several surface soil samples from the study area at a sampling depth of 0 to 20 cm, and the geographical location of the sampling points is recorded; the soil collected at each sampling point is broken up, debris is picked out, and after thorough mixing, a soil sample is obtained by quartering; the soil sample is then naturally air-dried, passed through a 10 mm sieve after drying, and the sieved portion is used as the soil sample. Each soil sample is placed in a sealed bag; the total potassium content of the soil is determined using a flame photometer;

[0052] The remote sensing image data was radiometrically calibrated and atmospherically corrected, and the digital values ​​(DN) of the original images were converted into true surface reflectance (SR). The ground spectral reflectance of the soil sample points was extracted from the processed images in ENVI 5.3 according to the geographic coordinates of the sample points. The original spectral reflectance was transformed into the first-order derivative of reflectance.

[0053] The terrain dataset is derived from the digital elevation dataset of the Shuttle Radar Topography Mission of the United States Geological Survey Earth Explorer. The terrain attributes include elevation, slope, and aspect. The climate dataset includes surface temperature, average annual temperature, and average annual precipitation. The original spatial resolution is 1 km and resampled to 30 m. The vegetation dataset is obtained from the Landsat 9

[0054] OLI imagery, calculate the Normalized Difference Vegetation Index, Difference Vegetation Index, Optimized Soil Adjusted Vegetation Index, and Greenness Normalized Difference Vegetation Index in ENVI.

[0055] The present invention is described in more detail below.

[0056] 1. Data acquisition and processing

[0057] ① Soil sample data: In March 2021 and March 2022, 120 and 120 surface soil samples were collected from the study area using the five-point sampling method (see Figure 1), with a sampling depth of 0 to 20 cm, and a portable global positioning system (GPS) was used to record the geographic location of the sampling points. The soil collected from each sub-sample point was broken up, and debris such as roots, straw, stones, and insect bodies were picked out. After thorough mixing, a single soil sample was obtained using the quartering method. The soil sample was then air-dried and passed through a 10 mm sieve. The total weight of the sieved portion was greater than 300 grams. Each soil sample was placed in a sealed bag and taken to the laboratory for further chemical analysis. The total potassium content of the soil was determined using a flame photometer.

[0058] ② Remote sensing image data: The two phases of ZY1-02D / AHSI remote sensing images used were imaged on January 28, 2021 and February 26, 2022, respectively. Since satellite remote sensing images may be affected by various factors during the imaging process, resulting in radiation distortion and geometric deformation. Therefore, ENVI 5.3 was used to perform radiometric calibration and atmospheric correction on the images, and the digital number (DN) of the original image was converted into the true surface reflectance (SR). The ground spectral reflectance of the soil sample point was extracted from the processed image in ENVI 5.3 according to the geographic coordinates of the sample point (see Figure 2 a). Perform reflectance first-order derivative transformation on the original spectral reflectance ( Figure 2 b), enhance the response band and reduce the noise in the spectral information.

[0059] ③ Environmental covariate data: Three environmental covariate data sets, topography, climate, and vegetation, were selected to predict soil total potassium content (see Table 1). The topography dataset was derived from the Shuttle Radar Topography Mission Digital Elevation Dataset (STRM DEM) from the U.S. Geological Survey Earth Explorer (USGS, 2018). Terrain attributes, including six variables such as elevation, slope, and aspect, were calculated in ArcGIS. The climate dataset included land surface temperature (LST), mean annual air temperature (MAT), and mean annual precipitation (MAP). The original spatial resolution was 1 km, and the data was resampled to 30 m. The vegetation dataset was obtained from Landsat 8OLI images on March 22, 2021, and Landsat 9OLI images on April 18, 2022, using the Geospatial Data Cloud Platform (https: / / www.gscloud.cn). The Normalized Difference Vegetation Index (NDVI), Difference Vegetation Index (DVI), Optimized Soil Adjusted Vegetation Index (OSAVI), and Green Normalized Difference Vegetation Index (GNDVI) were calculated in ENVI.

[0060] Table 1. Environmental covariate data

[0061]

[0062] 2. Model construction

[0063] 1. Partial Least Squares Regression: PLSR, first proposed by Wold et al., is a widely used multivariate statistical regression method and one of the methods used to predict soil properties. PLSR combines common multivariate regression analysis, principal component analysis, and correlation analysis, while retaining the advantages of all three regression methods. It is an optimization algorithm for conventional linear regression. This method reduces the dimensionality of the independent variables by finding new orthogonal projection directions (principal components) to improve the model's predictive performance.

[0064] ② Random Forest: RF is a tree-based method used to analyze the relationship between variables and influencing factors. RF regresses input variables against soil characteristic variables using decision rules, employing bootstrapping to overcome overfitting. Unlike traditional decision trees, random forests incorporate a random feature selection mechanism. This means that when constructing each decision tree, only a subset of feature variables is considered for partitioning. This increases the diversity between decision trees and effectively improves the model's generalization capabilities. The RF model from the scikit-learn library is used to optimize parameters using random search.

[0065] ③ Adaptive Boosting: AdaBoost is an ensemble learning method, the core of which is to combine multiple weak learners into a strong regressor with strong training ability, thereby improving the overall performance and stability. However, unlike RF, which has equal weights for each sample, the AdaBoost algorithm calculates the weight of each weak regressor based on the prediction accuracy of each sample in each training set and the overall prediction accuracy of the previous training set, and updates the distribution weight of each sample at the same time. Finally, the weighted sum of the regressor results obtained from each training is output (see Figure 3 ). Call the AdaBoost model in the scikit-learn library and use the grid search method to optimize the parameters.

[0066] 3. Experimental Setup

[0067] In order to evaluate and compare the ability of spectral variables, environmental variables and spectral transformation to predict soil total potassium, the following five experiments were set up: Experiment 1 was constructed based on environmental variables, Experiment 2 was based on the original hyperspectral bands of ZY1-02D, Experiment 3 was a combination of environmental variables and spectral bands, Experiment 4 was based on variables transformed from the first-order derivative of spectral reflectance, and Experiment 5 was a combination of environmental variables and the first-order derivative of spectral reflectance (see Table 2).

[0068] Table 2. Different combinations of input variables for the soil total potassium prediction model.

[0069]

[0070] 4. Model accuracy evaluation

[0071] The coefficient of determination (R2 ), Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) were used to evaluate the accuracy of the prediction model (Formulas (1) to (3)). 2 The value of R is between 0 and 1. 2 The closer the value is to 1, the better the model fit is. MAE and RMSE are used to measure the difference between the predicted value and the measured value. The smaller the value, the smaller the error between the model prediction value and the measured value.

[0072]

[0073] In formulas (1) to (3), n is the number of soil samples, y i is the observed value of TK content of sample i, is the predicted value of TK content of sample i.

[0074] 5. Model Accuracy Comparison

[0075] Based on five experimental settings, the fitting results of PLSR, RF and AdaBoost models for predicting total potassium are shown in Figure 2. Figure 4 The choice of model and experimental setup significantly affect the prediction accuracy of total potassium. Based on the validation set model, the machine learning model under Experiment 3 had higher prediction accuracy, and the AdaBoost model generally outperformed RF and PLSR.

[0076] From the perspective of experimental settings, when only environmental covariates are used, the validation set accuracy of the machine learning model is <0.6, while when only spectral variables are used, the validation set R of the AdaBoost model is 2 =0.74, indicating that ZY1-02D ​​hyperspectral remote sensing can effectively estimate soil total potassium content. Under the conditions of Experiment 3, that is, the synergistic spectral variables and environmental covariates, the prediction accuracy of the model is further improved. Comparing the three machine learning models of PLSR, RF, and AdaBoost, the calibration set and validation set R of the PLSR model are shown in Figure 2. 2 Between 0.26 and 0.58, the model estimation ability is poor. 2 The value is between 0.60 and 0.73, slightly lower than AdaBoost. The combination of Experiment 3 and AdaBoost is the best prediction model. 2 =0.75, MAE=0.37 g / kg, RMSE=0.49 g / kg, which can explain 75% of the total potassium content. Compared with the optimal PLSR and RF model combination, MAE was reduced by 42.79% and 2.63%, and RMSE was reduced by 28.99% and 2%, respectively.

[0077] Overall, regardless of the machine learning model used, predictive accuracy improved when both spectral variables and environmental covariates were used as model inputs. Models based on environmental covariates showed the lowest accuracy, outperforming models using only spectral variables as inputs. From a model perspective, the AdaBoost model performed best in estimating soil total potassium, followed by RF, and worst by PLSR.

[0078] 6. Spatial distribution of total potassium in soil

[0079] The model combination based on Experiment 3 and AdaBoost has a strong prediction ability. Therefore, this study uses the best model Experiment 3 and AdaBoost combination to predict the spatial distribution of total potassium in cultivated soil in the study area. Figure 5 shown. Figure 5 (a) shows the spatial distribution of total potassium in 2021, with a content range of 18.500-22.929 g / kg. The total potassium content in the northern and eastern parts of the study area is relatively high, while the total potassium content in the southwestern part is low. Figure 5 (b) shows the spatial distribution of total potassium in 2022, with a content range of 21.118 to 22.929 g / kg, and an overall trend of high in the west and low in the east.

[0080] In order to verify the effectiveness of the proposed scheme, measured soil sampling data from Xinzheng City, Henan Province in 2021 and 2022 were used, and ZY1-02D ​​hyperspectral images and environmental covariate data were obtained at the same time. The data were divided into training sets and test sets to verify the effectiveness of hyperspectral remote sensing in estimating soil total potassium content.

[0081] First, the ZY1-02D / AHSI imagery was processed for radiometric calibration and atmospheric correction, converting the original digital number (DN) into true surface reflectance (SR). The ground spectral reflectance of the soil sample points was then extracted based on the geographic coordinates of the measured sampling points. Finally, the data was transformed using the first-order derivative of the reflectance to enhance the spectral response band and reduce noise in the spectral information.

[0082] The processed spectral data and environmental covariate data were input into the machine learning model, and the parameters were tuned to test the model's ability to estimate soil total potassium content. The figure shows that the combination of Experiment 3 and the AdaBoost model provides the most accurate estimation of total potassium content, describing the nonlinear relationship between the input variables and soil total potassium content. These results confirm the effectiveness and necessity of the ZY1-02D ​​hyperspectral system for estimating soil properties and provide strong support for determining the spatial distribution of soil total potassium.

[0083] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of protection of the technical solution of the present invention.

Claims

1. A method for estimating total potassium content in soil based on hyperspectral remote sensing, characterized by: The method comprises the following steps: (1) Data acquisition and processing: Collect soil sample data, obtain ZY1-02D ​​remote sensing image data, and select three environmental variable data: topography, climate, and vegetation; (2) Model construction: Partial least squares regression is used to reduce the dimensionality of independent variables by finding new orthogonal projection directions; input variables are regressed with soil characteristic variables through decision rules, bootstrapping is used to overcome the overfitting problem, and a random feature selection mechanism is introduced. In the process of building each decision tree, only some characteristic variables are considered for division to increase the differences between each decision tree; multiple weak learners are combined into a strong regressor through the AdaBoost ensemble learning method, and the weight of each weak regressor is calculated based on the prediction accuracy of each sample in each training set and the overall prediction accuracy of the previous training set. At the same time, the distribution weight of each sample is updated, and finally the weighted sum of the regressor results obtained from each training is output; (3) Experimental setup: Experiment 1 is based on environmental variables, Experiment 2 is based on the original hyperspectral bands of ZY1-02D, Experiment 3 is based on the combination of environmental variables and spectral bands, Experiment 4 is based on the variables transformed by the first-order derivative of spectral reflectance, and Experiment 5 is based on the combination of environmental variables and the first-order derivative of spectral reflectance; (4) Using the coefficient of determination R 2 , mean absolute error MAE and root mean square error RMSE to evaluate the accuracy of the prediction model; (5) Predicting the fitting results of soil total potassium content based on the experimental setup of step (3) and the model of step (2); (6) The optimal experimental setting in step (3) and AdaBoost were used to predict the spatial distribution of total potassium in cultivated soil in the study area.

2. The method for estimating total potassium content in soil based on hyperspectral remote sensing according to claim 1, wherein: In the step (1), a five-point sampling method is used to collect several surface soil samples from the study area at a sampling depth of 0 to 20 cm, and the geographical location of the sampling points is recorded; the soil collected at each sampling point is broken up, debris is picked out, and after thorough mixing, a soil sample is obtained by quartering; the soil sample is then naturally air-dried, and after drying, it is passed through a 10 mm sieve, and the sieved portion is used as a soil sample, and each soil sample is placed in a sealed bag; The total potassium content in the soil was determined using a flame photometer.

3. The method for estimating total potassium content in soil based on hyperspectral remote sensing according to claim 2, wherein: In the step (1), the remote sensing image data is subjected to radiometric calibration and atmospheric correction, and the digital number DN of the original image is converted into the real surface reflectance SR; the ground spectral reflectance of the soil sample point is extracted from the processed image in ENVI 5.3 according to the geographic coordinates of the sample point; Perform reflectance first derivative transformation on the original spectral reflectance.

4. The method for estimating total potassium content in soil based on hyperspectral remote sensing according to claim 3, wherein: In step (1), the terrain dataset is derived from the digital elevation dataset of the Space Shuttle Radar Topography Mission of the United States Geological Survey Earth Explorer, and the terrain attributes include altitude, slope, and aspect; the climate dataset includes surface temperature, average annual temperature, and average annual precipitation, with an original spatial resolution of 1 km and a resampled resolution of 30 m; the vegetation dataset is based on the Landsat 9OLI image obtained from the geospatial data cloud platform, and the normalized vegetation index, difference vegetation index, optimized soil-adjusted vegetation index, and greenness normalized vegetation index are calculated in ENVI.

5. The method for estimating total potassium content in soil based on hyperspectral remote sensing according to claim 4, wherein: In step (2), the RF model in the scikit-learn library is called and the parameters are optimized using a random search method.

6. The method for estimating total potassium content in soil based on hyperspectral remote sensing according to claim 5, wherein: In step (2), the AdaBoost model in the scikit-learn library is called, and the parameters are optimized using the grid search method.

7. The method for estimating total potassium content in soil based on hyperspectral remote sensing according to claim 6, wherein: In the step (4), Where n is the number of soil samples, y i is the observed value of TK content of sample i, is the predicted value of TK content of sample i.

8. The method for estimating total potassium content in soil based on hyperspectral remote sensing according to claim 7, wherein: The best experimental setup is Experiment 3.

9. A device for estimating total potassium content in soil based on hyperspectral remote sensing, characterized by: The device includes: a data acquisition and processing module, which collects soil sample data, obtains ZY1-02D ​​remote sensing image data, and selects three environmental variable data: terrain, climate, and vegetation; The model construction module uses partial least squares regression to reduce the dimensionality of independent variables by finding new orthogonal projection directions; the input variables are regressed with soil characteristic variables through decision rules, bootstrapping is used to overcome the overfitting problem, and a random feature selection mechanism is introduced. In the process of building each decision tree, only some characteristic variables are considered for division to increase the difference between each decision tree; multiple weak learners are combined into a strong regressor through the AdaBoost ensemble learning method. The weight of each weak regressor is calculated based on the prediction accuracy of each sample in each training set and the overall prediction accuracy of the previous training set, and the distribution weight of each sample is updated at the same time. Finally, the weighted sum of the regressor results obtained from each training is output; Experiment setting module, where Experiment 1 is based on environmental variables, Experiment 2 is based on the original hyperspectral bands of ZY1-02D, Experiment 3 is based on the combination of environmental variables and spectral bands, Experiment 4 is based on the variables transformed by the first-order derivative of spectral reflectance, and Experiment 5 is based on the combination of environmental variables and the first-order derivative of spectral reflectance; Evaluation module, which uses the coefficient of determination R 2 , mean absolute error MAE and root mean square error RMSE to evaluate the accuracy of the prediction model; A fitting prediction module, which predicts the total potassium content in soil based on the fitting results of the experiment setting module and the model building module; The distribution prediction module uses the optimal experimental setting and AdaBoost combination to predict the spatial distribution of total potassium in cultivated soil in the study area.

10. The device for estimating total potassium content in soil based on hyperspectral remote sensing according to claim 9, characterized in that: In the data acquisition and processing module, a five-point sampling method is used to collect several surface soil samples from the study area at a sampling depth of 0 to 20 cm, and the geographical location of the sampling points is recorded. The soil collected at each sampling point is broken up, debris is picked out, and after thorough mixing, a soil sample is obtained using the quartering method. The soil sample is then naturally air-dried and passed through a 10 mm sieve after drying. The sieved portion is used as the soil sample, and each soil sample is placed in a sealed bag. The total potassium content of soil was determined using a flame photometer; The remote sensing image data were radiometrically calibrated and atmospherically corrected to convert the original image digital number (DN) into the true surface reflectance (SR). The ground spectral reflectance of the soil sample points was extracted from the processed image in ENVI 5.3 according to the geographic coordinates of the sample points. Perform reflectance first derivative transformation on the original spectral reflectance; The terrain dataset is derived from the digital elevation dataset of the U.S. Geological Survey's Earth Explorer Shuttle Radar Topography Mission. Terrain attributes include elevation, slope, and aspect. The climate dataset includes surface temperature, mean annual temperature, and mean annual precipitation. The original spatial resolution is 1 km, which was resampled to 30 m. The vegetation dataset is based on Landsat 9OLI images obtained from the geospatial data cloud platform. The Normalized Difference Vegetation Index, Difference Vegetation Index, Optimized Soil Adjusted Vegetation Index, and Greenness Normalized Vegetation Index were calculated in ENVI.