Multi-element type soil nutrient prediction method coupling multi-source covariable and in-situ spectrum
By combining multi-objective regression machine learning and multi-source covariates, the problems of low signal-to-noise ratio and information redundancy in large-scale multi-type soil nutrient prediction are solved, achieving efficient and accurate prediction of multi-type soil nutrients and meeting the application needs of precision agriculture.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SANYA INSTITUTE OF NANJING AGRICULTURAL UNIVERSITY
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to achieve rapid and accurate nutrient content prediction on large-scale, diverse soil types. They are affected by factors such as soil type differences, roughness, porosity, and color. Furthermore, collinearity and synergistic changes among various nutrients lead to low signal-to-noise ratios and information redundancy, making them unsuitable for field production applications.
Multi-objective regression machine learning algorithms such as multi-objective support vector regression (MO-SVR), multi-objective random forest (MO-RF), and multi-layer perceptron (MLP) are used, combined with meteorological and topographic data as multi-source covariates, to construct a model that shares hidden layer features and decision paths. By utilizing the strong correlation between nutrients and the environment-driven mechanism, the model is used to predict soil nutrients in multiple types of soil.
It significantly improves the overall accuracy and model robustness of simultaneous multi-nutrient prediction, and realizes efficient and non-destructive prediction of soil nutrients of various types at the provincial scale, meeting the needs of precision agriculture.
Smart Images

Figure CN121835987A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a multivariate soil nutrient prediction method that couples multiple source covariates and in-situ spectroscopy, belonging to the field of soil nutrient detection technology. Background Technology
[0002] Based on the concept of "precision agriculture," accurately and quickly obtaining the content and spatial distribution of the required soil fertility factors, and adjusting the amount and timing of fertilizer application according to local conditions to improve fertilizer utilization, is of great significance to the sustainable development of agriculture and environmental protection in my country.
[0003] Traditional methods for soil nutrient determination require the collection and pretreatment of soil samples followed by laboratory analysis and measurement. While offering high accuracy, these methods suffer from drawbacks such as high cost, long analysis cycles, and limitations imposed by sample size, significantly restricting the efficiency of soil nutrient monitoring and failing to characterize the spatial distribution and continuous variation of soil nutrients in the field. With the rapid development of computer science and chemometrics, agronomy is increasingly seeking practical applications in its intersection with spectroscopy. Infrared spectroscopy, a branch of wave optics that studies the structure and composition of matter and radiation, as well as spectral properties, encompasses the visible (vis), near-infrared (NIR), and short-wave infrared (SWIR) portions of the electromagnetic spectrum. Because different soil parameters correspond to different parts of the electromagnetic spectrum, soil spectra can be used to predict various soil properties, such as the content of soil organic matter or total nitrogen, soil acidity, and soil texture. Therefore, the use of visible-near-infrared (vis-NIR) spectral data and machine learning algorithms for modeling has become increasingly widespread in soil nutrient monitoring.
[0004] Currently, the application of spectroscopic techniques for predicting soil nutrients is mainly concentrated under controlled indoor conditions. Although vis-NIR laboratory spectroscopy has successfully predicted various soil properties, sample pretreatment (air-drying, grinding, and sieving) remains cumbersome and time-consuming. Compared with laboratory-based spectral prediction, in-situ spectroscopy can significantly reduce the need for soil sample pretreatment, offering greater efficiency and real-time performance. It can rapidly provide accurate quantitative information for soil monitoring and modeling in field environments, making it more suitable for rapid monitoring of soil nutrients over large areas. However, the low signal-to-noise ratio and high data redundancy caused by field environmental factors such as soil moisture, atmospheric water vapor, soil surface roughness, and light conditions significantly affect the performance of in-situ spectral measurements and subsequent modeling predictions. Therefore, reducing interference from external environmental information unrelated to soil properties, improving the accuracy of in-situ spectral nutrient prediction, and unlocking its enormous potential in precision agriculture and environmental protection have become key research areas in this field.
[0005] Furthermore, current research focuses only on single soil types within a single study area, such as paddy soil, cotton fields, and black soil, and most studies use only single spectral data for modeling. There are few reports on large-scale predictions of multiple nutrient contents in diverse soil types based on in-situ spectral data. Due to significant differences between soil types, the low signal-to-noise ratio caused by soil factors such as texture, roughness, porosity, and color can greatly affect in-situ spectral measurements. Moreover, because there are strong collinear and synergistic variations among various soil nutrients, independent modeling of a single target will lead to information redundancy, thus affecting the accuracy and efficiency of nutrient content prediction. Simultaneously, since soil nutrient content at different provincial scales is significantly regulated by meteorological and topographic factors, in-situ spectral measurements in the field are affected by a series of external meteorological factors, field topographic factors, and internal soil properties. How to remove the influence of these interfering factors to improve the accuracy of soil nutrient content prediction has always been a hot topic and a challenge in this field.
[0006] In summary, current mainstream models focus only on modeling using single spectral data or on predicting nutrient content in a single soil type at a small scale in a single region. These methods are insufficient for large-scale quantitative prediction of soil nutrients in the field and cannot meet the needs of field production applications. This invention addresses the aforementioned technical problems by developing a rapid and accurate method for predicting the total and available nutrient content of multiple soil types at the provincial scale. This enables efficient and non-destructive prediction of nutrient content in different soil types at the provincial scale, providing technical guidance and supplementary experience for the application of quantitative prediction of soil nutrient information. Summary of the Invention
[0007] The purpose of this invention is to provide a multivariate soil nutrient prediction method that couples multiple source covariates and in-situ spectra. It uses a combination of preprocessing methods and employs multi-objective regression machine learning algorithms such as multi-objective support vector regression (MO-SVR), multi-objective random forest (MO-RF), and multilayer perceptron (MLP) to construct the model. This allows multiple nutrient feature tasks to share hidden features and decision paths. Simultaneously, it couples meteorological and topographical data as multi-source covariates to share input features, fully utilizing the strong correlation between nutrients and the environment-driven mechanism. This significantly improves the overall accuracy and robustness of simultaneous multi-nutrient prediction, overcoming the shortcomings of existing technologies that suffer from large prediction bias and weak generalization ability when modeling multivariate soil types and their multi-objective nutrients.
[0008] To achieve the above objectives, the technical solution adopted by this invention is: a multivariate soil nutrient prediction method coupling multiple source covariates and in-situ spectroscopy, comprising the following steps:
[0009] Step 1: In-situ hyperspectral measurement and data acquisition of soil;
[0010] Step 2: Soil sample collection and preparation;
[0011] Step 3: Obtaining covariate data: Obtain meteorological and topographic data for the corresponding sampling area based on latitude and longitude coordinates;
[0012] Step 4: Obtaining soil nutrient data: Analyze the prepared soil samples and determine the soil nutrient content, which will serve as the target variable data for subsequent model construction.
[0013] Step 5: In-situ spectral data preprocessing: Perform preprocessing operations on the in-situ soil spectra;
[0014] Step 6: Spectral data outlier detection and removal: Outlier detection is performed on the preprocessed spectral data to identify and remove outliers, thus eliminating potential anomalies.
[0015] Step 7, Feature Dimension Reduction and Screening: Principal component analysis is used to perform feature dimensionality reduction and screening on the spectral data and multi-source covariates;
[0016] Step 8: Construct a multi-objective regression prediction model for soil nutrient content: A multi-objective regression machine learning algorithm is used to construct a soil nutrient prediction model. A global grid search method is used to find and adjust the optimal hyperparameters of each model. A 10-fold cross-validation method is used to test and evaluate the model.
[0017] Step 9: Model prediction performance evaluation and comparison: Randomly divide the test set from the modeling dataset, use the constructed model to predict the total and available nutrient content of soil samples, compare the obtained predicted values with the actual nutrient content values, and compare the prediction performance of each model through validation. Select the optimal prediction model based on the comparison results.
[0018] Step 10: Quantify the contribution of multi-source covariates in model prediction: The SHAP method is used to analyze and rank the importance of topographic and meteorological covariates in the prediction of soil nutrient content. The comparison object is the actual SHAP contribution value of each covariate and the average SHAP value of all spectral band features.
[0019] As a further preferred embodiment of this scheme, in step 1, sampling points are divided into equal-spaced grids according to the size of the plots in the area to be tested. A 25cm deep soil profile with a cross-section of 14cm×8cm is artificially created using a rectangular soil sampler based on the sampling points. Then, three locations are selected at different positions on the vertical profile of the soil sample, and in-situ spectral data of the soil are collected. During the measurement process, the reflectance probe is kept close to the soil surface for spectral acquisition. Ten spectra are collected at each location, and the arithmetic mean of the ten spectra is taken as the in-situ spectral data of the soil sample at that sampling point. The spectrometer has a measurement range of 350-2500nm.
[0020] As a further preferred option of this scheme, the grid spacing of the equidistant grid method is 20m×20m; a whiteboard calibration is performed before measuring the in-situ spectrum of each soil sample to ensure the quality of the spectral reflectance data; when measuring the in-situ spectrum of the soil, a built-in halogen lamp is used instead of natural sunlight as the measurement light source to ensure the stability of the measurement environment.
[0021] As a further preferred embodiment of this scheme, in step 2, after completing the in-situ spectral data collection of soil at each sampling point, a 2cm thick soil sample is taken from the same location on the soil profile where in-situ spectral measurements were performed. The three soil samples are mixed and placed in a sealed bag for transport back. The collected soil samples are then placed on a plastic sheet and placed indoors for ventilation and air drying. Large soil clods must be crushed to prevent them from hardening into hard lumps after drying, which would make them difficult to grind. After air drying, stones and plant and animal remains are removed from the soil samples. The samples are then divided into two parts for grinding and sieves with 1mm and 100-mesh apertures, respectively. The soil sample passing through the 1mm sieve is used to determine the available nutrients (alkaline nitrogen, available phosphorus, and available potassium) of the soil. The other part of the soil sample passing through the 100-mesh sieve is used to determine the organic matter and total nutrients (total nitrogen, total phosphorus, and total potassium).
[0022] As a further preferred embodiment of this scheme, in step 3, after the in-situ spectral data of the soil at each sampling point is collected, the coordinates of each sampling point are recorded using GPS. Based on the coordinates of latitude and longitude, meteorological data and topographic data of the corresponding sampling area are obtained. Specifically, the meteorological data includes information on average temperature, precipitation, and average wind speed related to regional variability. The topographic data is extracted by using Python to extract topographic information such as height, slope, aspect, and topographic humidity index of the corresponding sampling area from the DEM raster file.
[0023] As a further optimization of this scheme, during the meteorological data acquisition process, the accuracy of the meteorological information corresponding to the sampling date is selected as daily accuracy; during the topographic data extraction process, topographic information is extracted from raster files using the geopandas and rasterio libraries; during the topographic data acquisition process, latitude and longitude conditions are selected for search on the Earth Explorer website, and the raster data of the corresponding sampling area is located using the WGS-84 coordinate system; the downloaded file when acquiring topographic data is the raster file provided by the Space Shuttle Radar Topography 1-arsecond Digital Elevation Model Mission (SRTM1-DEM) on the Earth Explorer website; the formula used in the extraction of topographic humidity index information is described as follows:
[0024]
[0025] in, This represents the square root of the cumulative slope area flowing through each pixel. It is the surface slope of the corresponding pixel.
[0026] As a further preferred embodiment of this scheme, in step 4, the contents of organic matter, total nitrogen, total potassium, total phosphorus, as well as alkaline nitrogen, available phosphorus, and available potassium in the soil sample are determined, wherein:
[0027] The organic matter content of soil samples was determined using the potassium dichromate bulk density method.
[0028] The total nitrogen content in soil samples was determined using the semi-micro Kjeldahl method.
[0029] The contents of total phosphorus, total potassium, and available potassium in soil samples were determined by ionization-coupled plasma optical emission spectrometry.
[0030] The alkaline hydrolysis-diffusion method was used to determine the alkaline nitrogen content in soil samples.
[0031] The available phosphorus content in soil samples was determined using the antimony-molybdenum colorimetric method.
[0032] As a further preferred embodiment of this scheme, in step 5, the preprocessing method is combined with extended multivariate scattering correction and normalization processing, wherein the polynomial order of the extended multivariate scattering correction is set to 2; in step 6, the outlier detection method based on dynamic threshold isolation forest is used to perform outlier detection on the preprocessed spectral data, with parameters set to n_estimators=200 and random_state=42 to ensure the repeatability of the results.
[0033] As a further preferred option of this scheme, in step 9, after comparison, the XGBoost model is selected as the prediction model for the multi-type soil nutrient content. The XGBoost model is rapidly optimized by adding multiple regression trees in sequence and using second-order gradient information to minimize the overall loss function, thereby achieving accurate fitting of the complex nonlinear relationship between spectral features and nutrient content.
[0034] As a further optimization of this scheme, the key hyperparameters of the XGBoost model include tree-level sampling rate, number of decision trees, maximum tree depth, and L1 regularization coefficient. During the modeling process, a global grid search is used to traverse the hyperparameter combinations, and the optimal combination is randomly selected to maximize the model's predictive performance, balancing prediction accuracy and generalization. The formula is described as follows:
[0035]
[0036] The above formula represents the overall objective function of XGBoost in the t-th iteration, including the empirical loss term and the regularization term, where, This represents the measured nutrient content of the soil sample. These are the predicted values from the first t-1 iterations. Let t be the regression tree to be learned. The loss function;
[0037] Specifically, This is a structure regularization term for the tree model, used to control model complexity. Its specific form is:
[0038]
[0039] in, Let be the number of leaf nodes in the t-th tree. The weight of the j-th leaf node, and These are the first and second regularization coefficients, used to suppress overfitting;
[0040] After approximating the objective function using a second-order Taylor expansion, the optimization process is transformed into a closed-form solution for the weights of each leaf node. The optimal leaf node weights are represented as follows:
[0041]
[0042] in, Let j be the set of samples that fall into the j-th leaf node. For the first-order gradient, It is a second-order gradient;
[0043] After all iterations are completed, the final prediction model is represented as the sum of the regression trees from each round:
[0044]
[0045] Where K is the total number of iterations.
[0046] As a further preferred embodiment of this scheme, in step 9, the predictive performance of each model is compared using R2, RMSE, and RPD validation metrics, and the optimal predictive model is selected based on the comparison results; the expression is as follows:
[0047]
[0048]
[0049]
[0050] in, This represents the measured nutrient content of soil sample i. This indicates the predicted nutrient content of the soil sample. This represents the average nutrient content of the soil sample, where n represents the number of soil samples, and SD represents the standard deviation between the predicted nutrient content and the measured nutrient content.
[0051] As a further preferred option of this scheme, when R 2 When the RPD is greater than 0.85, RMSE is less than 10% of the mean of the target variable, and RPD is greater than 2.0, it indicates that the model has excellent predictive performance, can accurately predict the total amount and available nutrients in the soil, and can be applied to the quantitative prediction of soil nutrient information.
[0052] When R 2 When the RMSE is greater than 0.80, the RMSE is less than 20% of the average value of the target variable, and the RPD is greater than 1.5, it indicates that the model has good predictive performance and can basically achieve accurate prediction of the total amount and available nutrients in the soil. It can serve as an effective reference to assist in the non-destructive and efficient detection of field soil samples.
[0053] As a further optimization of this scheme, in addition to the evaluation metrics, the R-squared value is also used as the evaluation criterion for model prediction performance. 2 _train - R 2 The results of _test and RMSE_test / RMSE_train are used as validation criteria to evaluate the overfitting of machine learning models.
[0054] Compared with the prior art, the beneficial effects of the present invention are:
[0055] This invention provides a multivariate soil nutrient prediction method that couples multiple source covariates and in-situ spectra. Due to the strong collinearity and synergistic variation among various soil nutrients, independent modeling of a single objective leads to information redundancy and low prediction efficiency. Furthermore, because soil samples involve provincial scales and multiple soil types, and nutrient content is significantly regulated by meteorological and topographic factors, traditional single-output models struggle to fully utilize the intrinsic correlations between variables, easily resulting in overfitting or accuracy loss. Therefore, this invention uses a combination of preprocessing methods and employs multi-objective regression machine learning algorithms such as Multi-Objective Support Vector Regression (MO-SVR), Multi-Objective Random Forest (MO-RF), and Multilayer Perceptron (MLP) to construct the model. This allows multiple nutrient feature tasks to share hidden features and decision paths, while simultaneously coupling meteorological and topographic multi-source covariates as shared input features. This fully utilizes the strong correlations between nutrients and the environment-driven mechanism, significantly improving the overall accuracy and robustness of simultaneous multivariate prediction. It overcomes the shortcomings of existing technologies that model only a single nutrient, resulting in large prediction bias and weak generalization ability. This enables efficient and non-destructive prediction of nutrient content in diverse soil types at the provincial scale, meeting the need for accurate prediction of nutrient content in diverse soil types at the provincial scale, and providing more real-time, accurate, and diversified application solutions for large-scale soil nutrient monitoring. Attached Figure Description
[0056] Figure 1 This is the overall flowchart of the present invention.
[0057] Figure 2 This is a schematic diagram of in-situ spectra of various soil types and their pretreated spectra; among which: Figure 2 (a) Soil in-situ spectrum - original reflectance, Figure 2 (b) Soil in-situ spectral-standard normal transformation preprocessing, Figure 2 (c) Soil in-situ spectral-normalized pretreatment, Figure 2 (d) Soil in-situ spectra - extended multivariate scattering correction preprocessing, Figure 2 (e) Soil in-situ spectral-continuum removal pretreatment, Figure 2 (f) Soil in-situ spectra - first derivative preprocessing.
[0058] Figure 3 It is based on the dynamic threshold Iso Forest method for detecting outliers in soil samples;
[0059] Figure 4 This is a 3D visualization of the PCA principal components of soil nutrient variables; where: Figure 4 (a) Soil total nitrogen content, Figure 4 (b) Soil total phosphorus content, Figure 4 (c) Total potassium content in soil, Figure 4 (d) Soil available nitrogen content, Figure 4 (e) Soil available phosphorus content, Figure 4 (f) Soil available potassium content, Figure 4 (g) Soil organic matter content.
[0060] Figure 5 This is a schematic diagram of the Pearson correlation analysis between soil nutrient variables and multiple covariates;
[0061] Figure 6 This is a graph showing the linear and nonlinear correlation analysis of the content of various soil nutrient variables; where: Figure 6 (a) Total potassium and available potassium in soil, Figure 6 (b) Total phosphorus and available potassium in the soil, Figure 6 (c) Total potassium and available nitrogen in soil, Figure 6 (d) Soil organic matter and alkaline nitrogen, Figure 6 (e) Total phosphorus and total potassium in the soil, Figure 6 (f) Total phosphorus and available phosphorus in soil.
[0062] Figure 7 This is a bar chart comparing the performance of various machine learning models in predicting soil nutrient content; among them: Figure 7(a) A bar chart comparing the performance of predicting total nitrogen content in soil. Figure 7 (b) A bar chart comparing the performance of predicting total phosphorus content in soil. Figure 7 (c) Bar chart comparing the performance of predicting total potassium content in soil. Figure 7 (d) Bar chart comparing the performance of predicting soil available nitrogen content. Figure 7 (e) A bar chart comparing the performance of predicting available phosphorus content in soil. Figure 7 (f) Bar chart comparing the performance of predicting available potassium content in soil. Figure 7 (g) Performance comparison bar chart for predicting soil organic matter content.
[0063] Figure 8 This is a diagram comparing the RMSE (Recovery Mean Squared) of various machine learning models in predicting soil nutrient content on the training and test sets; where: Figure 8 (a) Line graph comparing the RMSE of predicted total nitrogen content in soil. Figure 8 (b) Line graph comparing RMSE for predicted total phosphorus content in soil. Figure 8 (c) Line graph comparing the RMSE of predicted total potassium content in soil. Figure 8 (d) Line graph comparing the RMSE of predicted soil available nitrogen content. Figure 8 (e) Line graph comparing RMSE for predicted available phosphorus content in soil. Figure 8 (f) Line graph comparing RMSE for predicting available potassium content in soil. Figure 8 (g) Line graph comparing the RMSE of predicted soil organic matter content.
[0064] Figure 9 This is a scatter plot of soil nutrient content predicted by the XGBoost model; where: Figure 9 (a) Scatter plot of predicted total nitrogen content in soil. Figure 9 (b) Scatter plot of predicted total phosphorus content in soil. Figure 9 (c) Scatter plot of predicted total potassium content in soil. Figure 9 (d) Scatter plot of predicted soil available nitrogen content. Figure 9 (e) Scatter plot of predicted available phosphorus content in soil. Figure 9 (f) Scatter plot of predicted available potassium content in soil. Figure 9 (g) Scatter plot of predicted soil organic matter content.
[0065] Figure 10 This is a graphical representation of the contributions of multiple covariates, including topography and meteorology, to soil nutrient prediction; among which: Figure 10 (a) Visual representation of covariate contributions to predict total nitrogen content in soil. Figure 10 (b) Visual chart of covariate contributions to predict total phosphorus content in soil. Figure 10 (c) Visual representation of covariate contributions to predict total potassium content in soil. Figure 10 (d) Visual diagram of covariate contributions to predict soil available nitrogen content. Figure 10 (e) Visual representation of covariate contributions to predict available phosphorus content in soil. Figure 10 (f) Visual representation of covariate contributions to predict soil available potassium content. Figure 10 (g) Visual representation of covariate contributions to predict soil organic matter content. Detailed Implementation
[0066] The specific embodiments or technical solutions of the present invention will be described more clearly and explicitly below with reference to the accompanying drawings.
[0067] Example 1: A multivariate soil nutrient prediction method coupling multiple source covariates and in-situ spectroscopy:
[0068] like Figure 1 As shown, this embodiment specifically includes the following steps:
[0069] Step 1: In-situ Hyperspectral Measurement and Data Acquisition of Soil: Sampling areas with different soil types in different provinces were determined. Sampling points were divided into equal-spaced grids according to the size of the plots in the area to be tested. A 25cm deep soil profile with a cross-section of 14cm×8cm was artificially created using a rectangular soil sampler. Three locations (corresponding to different depths) were selected at different positions on the vertical profile of the soil sample, and in-situ spectral data of the soil were collected. During the measurement process, the reflectance probe was kept close to the soil surface for spectral acquisition. Ten spectra were collected at each location. The arithmetic mean of the ten spectra was taken as the in-situ spectral data of the soil sample at that sampling point. The spectral reflectance data range was 350-2500nm.
[0070] Step 2, Soil Sample Collection and Preparation: After collecting in-situ spectral data of soil at each sampling point, take a 2cm thick soil sample from the same location on the soil profile where in-situ spectral measurements were performed. Mix the three soil samples and put them into a sealed bag and transport them to the laboratory. Place the collected soil samples on a plastic sheet and air dry them indoors. After drying, remove stones and plant and animal remains from the soil samples. Divide the soil samples into two parts, grind them, and then sieve them using 1mm and 100-mesh sieves, respectively.
[0071] Step 3: Acquisition of covariate data: After the in-situ spectral data of the soil at each sampling point is collected, the coordinates of each sampling point are recorded using GPS. Based on the coordinates of latitude and longitude, meteorological and topographic data of the corresponding sampling area are obtained. Specifically, the meteorological data includes the average temperature, precipitation, and average wind speed, which are most relevant to regional variability. The topographic data is extracted by using Python to extract topographic information such as height, slope, aspect, and topographic humidity index of the corresponding sampling area from the DEM raster file.
[0072] Step 4: Acquisition of soil nutrient data: The prepared soil samples are transported to the laboratory for indoor chemical analysis to determine the content of organic matter, total nitrogen, total potassium, total phosphorus and other soil nutrients, as well as the content of available soil nutrients such as alkaline nitrogen, available phosphorus and available potassium, which will be used as target variable data for subsequent model construction.
[0073] Step 5: In-situ spectral data preprocessing: Different methods were used to preprocess the in-situ soil spectra. The preprocessing methods used included Standard Normal Transform (SNV), Extended Multivariate Scattering Correction (EMSC), Derivative Transform (DT), Continuum Removal (CR), Normalization, and Savitzky-Golay smoothing. Through comparative analysis, the optimal preprocessing method for each machine learning model was selected. Finally, the optimal combination of preprocessing methods was determined to be Extended Multivariate Scattering Correction (EMSC) and Normalization, with the polynomial order of Extended Multivariate Scattering Correction set to 2.
[0074] Step 6: Outlier Detection and Removal in Spectral Data: An outlier detection method based on dynamic threshold isolation forest (IsolationForest) is used to detect outliers in the preprocessed spectral data. The parameters are set to n_estimators=200 and random_state parameter is set to 42 to ensure the repeatability of the results, so as to detect and remove outliers in the spectral data and eliminate potential anomalies.
[0075] Step 7, Feature Dimension Reduction and Screening: Principal component analysis was used to perform feature dimensionality reduction and screening on the spectral data and multi-source covariates. After repeated iterative attempts, the optimal number of principal components was set to 15.
[0076] Step 8: Construct a multi-objective regression prediction model for soil nutrient content: Use multi-objective regression machine learning algorithms such as K-Nearest Neighbors (KNN), Random Forest (RF), Multi-objective Support Vector Machine (MO-SVR), Multi-objective Linear Regression (MO-LR), Extreme Gradient Boosting (XGBoost), and Multilayer Perceptron (MLP) to construct a soil nutrient prediction model. Use a global grid search method to find and adjust the optimal hyperparameters of each model, and use a 10-fold cross-validation method to test and evaluate the model.
[0077] Step 9: Model Prediction Performance Evaluation and Comparison: During the model prediction process, 20% of the modeling dataset is randomly allocated as a test set. The constructed model is used to predict the total and available nutrient content of soil samples. The predicted values are compared with the actual nutrient content values, and the results are analyzed using R... 2 The predictive performance of each model is compared using validation metrics such as RMSE and RPD, and the optimal predictive model is selected based on the comparison results.
[0078] Step 10: Quantify the contribution of multi-source covariates in model prediction: The SHAP method is used to analyze and rank the importance of topographic and meteorological covariates in the prediction of soil nutrient content. The comparison object is the actual SHAP contribution value of each covariate and the average SHAP value of all spectral band features.
[0079] Table 1 presents information such as the number of soil samples collected, soil type, soil texture, and meteorological and topographic data of the sampling area. The grid spacing for the equidistant grid method is 20m x 20m.
[0080] Table 1: Soil Sample Attribute Information and Schematic Diagram of Meteorological and Topographic Data of Sampling Area
[0081]
[0082] Tables 2 to 4 present descriptive statistics of soil content for diverse soil types in the three sampling areas, including statistical results such as the standard deviation and coefficient of variation of seven nutrient variables and soil moisture content.
[0083] Table 2: Descriptive Statistical Results of Soil Variables in the Sampling Area of Rugao, Jiangsu Province
[0084]
[0085] Table 3: Descriptive Statistical Results of Soil Variables in the Sampling Area of Gao'an, Jiangxi Province
[0086]
[0087] Table 4: Descriptive Statistical Results of Soil Variables in the Sampling Area of Chongming, Shanghai
[0088]
[0089] Specifically, the steps for obtaining terrain data of the sampling area are as follows:
[0090] (1) Log in to the Earth Explorer website and select the Space Shuttle Radar Topography 1 Arc-Sec Digital Elevation Model (SRTM1-DEM) mission conducted by the United States Geological Survey (USGS).
[0091] (2) Select latitude and longitude search in the search criteria and enter the latitude and longitude information of the sampling area (select WGS-84 coordinate system) to download the corresponding raster file;
[0092] (3) Use Python and the corresponding geopandas and rasterio packages to extract topographic data such as height, slope, aspect and topographic humidity index of the corresponding sampling area contained in the raster file.
[0093] Specifically, the steps for obtaining meteorological data for the sampling area are as follows:
[0094] (1) Match the corresponding township or town-level meteorological stations based on the latitude and longitude of the sampling area;
[0095] (2) Select the corresponding meteorological information data according to the sampling date, and select daily precision;
[0096] (2) Collect the meteorological data such as average temperature, precipitation and average wind speed that are most relevant to the sampling point as meteorological data for the corresponding sampling area.
[0097] This invention utilizes a soil spectral measurement instrument to acquire in-situ soil spectral data. The specific spectrometer is the Fieldspec Pro 4, manufactured by ASD (Analytical Spectral Devices) in the United States, equipped with a high-density contact reflectance probe as a detection accessory. This probe has a built-in light source. This spectrometer is portable and provides stable measurements. Its measurement range is 350-2500 nm, with a spectral resolution of 3 nm in the 350-1000 nm band and 10 nm in the 1000-2500 nm band. Its resampling interval is 1 nm, thus providing high-precision spectral measurements. The spectrometer has a field of view of 25°.
[0098] The high-density contact reflective probe is 25.4 cm long and weighs 1.5 kg. It operates at 12-18V and has a power of 6.5W. The built-in light source is a halogen bulb with a tungsten lamp color temperature of 2901 ± 10°%K and a spot size of 10 mm. When measuring the in-situ spectrum of soil, the built-in halogen lamp in the probe is used as the light source to ensure that the spectrometer receives sufficient light intensity during measurement, thus guaranteeing the quality of the measurement data. At the same time, the use of the built-in light source has the advantages of stable illumination and wide application time. It can be used flexibly without being limited by weather or time, as it can avoid the influence of unstable or insufficient light intensity caused by weather or cloud cover.
[0099] Specifically, the implementation steps for in-situ soil spectral measurement are as follows:
[0100] (1) During the in-situ spectral measurement process, a reflectance probe is used to collect in-situ soil spectra in close contact with the soil surface;
[0101] (2) Ten spectra were collected at each location, and the arithmetic mean of the ten spectra was used as the in-situ spectral data of the soil sample at that sampling point.
[0102] (3) A whiteboard calibration is required before measuring the in-situ spectrum of each soil sample to ensure the quality of the spectral data;
[0103] Taking the average of 10 spectral data points can reduce random errors caused by environmental, human, or instrumental factors, thereby avoiding the problem of poor quality caused by fluctuations in spectral data.
[0104] The collected in-situ soil spectra are preprocessed because the spectral data obtained by the spectrometer, in addition to containing the soil spectral response signal, may also be mixed with noise generated during the work process and the influence of environmental factors. Spectral preprocessing is beneficial to improve the accuracy of spectral analysis, provide stability and effectiveness in the subsequent model construction process, and improve the prediction accuracy of soil nutrients. Therefore, this invention selects the most suitable preprocessing method and its combination according to the characteristics of different models.
[0105] like Figure 2 As shown, this invention employs different methods to preprocess the in-situ soil spectrum. The preprocessing methods used include: Standard Normal Transform (SNV), Extended Multivariate Scattering Correction (EMSC), Derivative Transform (DT), Continuum Removal (CR), Normalization, and Savitzky-Golay Smoothing. The optimal preprocessing method for each machine learning model is selected through comparative analysis.
[0106] like Figure 3As shown, outlier detection is performed on the spectral preprocessed data to screen and remove abnormal outliers, reducing prediction errors caused by sample bias, ensuring the reliability of multi-type soil prediction, and improving the model's prediction accuracy for subsequent soil nutrient content.
[0107] Specifically, this invention employs an outlier detection method based on dynamic threshold isolation forest to filter and remove outliers from the spectral data in the modeling dataset. The method first initializes the isolation forest model with parameters set to n_estimators=200 and random_state=42 to ensure reproducibility. Then, it calculates the decision function score by fitting the model to the dataset, and uses the 5th percentile of the score sequence as a dynamic threshold. Samples with scores below this threshold are identified as outliers to exclude potential anomalies.
[0108] like Figure 4 As shown, the spectral data, after spectral preprocessing and outlier detection, is coupled with topographic data and meteorological data, among other multi-source covariates, to construct a modeling dataset. Before modeling, to further improve the stability and accuracy of the soil nutrient prediction model, especially to effectively reduce the redundancy and noise interference of high-dimensional spectral data when coupling multi-source data, this invention introduces Principal Component Analysis (PCA) algorithm for spectral dimensionality reduction. The Principal Component Analysis algorithm is based on the principle of linear transformation. By extracting the main variable components of the data, the high-dimensional spectral variables are projected into a low-dimensional orthogonal space. The principal component fraction matrix is obtained through eigenvalue decomposition. The first K principal components retain the main information of the cumulative variance contribution rate (95%). By retaining the first K principal components, the dimensionality reduction effect of high-dimensional data is achieved. In this invention, the bias introduced by band collinearity and noise is significantly reduced, and the generalization and accuracy of the machine learning model for predicting multi-type soil nutrients are improved.
[0109] Specifically, through repeated iterative attempts, this invention sought to balance the model's prediction accuracy and speed with preserving as much nutrient-related spectral information as possible, and ultimately selected 15 principal components as the optimal number.
[0110] To reveal the intrinsic relationships among nutrient contents in various soil types and between nutrients and meteorological and topographic covariates, and to provide a theoretical basis for subsequent model design, this invention conducted Pearson correlation analysis on all variables.
[0111] like Figure 5 As shown, this invention conducted Pearson correlation analysis on major soil nutrients (TN, SOM, TP, TK, AN, AP, AK) with meteorological and topographic multi-source covariates and presented the results in the form of a heatmap. The results show that:
[0112] (1) Nutrients showed significant covariance: available potassium (AK) had a correlation coefficient of 0.81 with total phosphorus (TP) and 0.83 with total potassium (TK); total phosphorus (TP) had a correlation coefficient of up to 0.94 with total potassium (TK); organic matter (SOM) also showed a moderate to strong correlation with other nutrients (0.60~0.90).
[0113] (2) Nutrients and environmental covariates also showed significant correlations: available potassium (AK), total potassium (TK), and total phosphorus (TP) showed significant negative correlations with slope, precipitation and elevation (r < -0.81), and significant positive correlations with topographic humidity index (twi) and wind speed (r > 0.8).
[0114] like Figure 6 As shown, this invention further selects the nutrient pairs with the strongest correlation (such as TP and TK, r=0.94) for linear or exponential regression fitting, with a coefficient of determination R0 2 =0.89, the fitted equation is y=53.62e 0.46x – 60.61 indicates a strong correlation between the two.
[0115] The above results indicate that there are strong collinear and synergistic variation patterns among various soil nutrients. Independent modeling for a single objective leads to information redundancy and low prediction efficiency. Furthermore, because the soil samples involve provincial scales and multiple soil types, and nutrient content is significantly regulated by meteorological and topographic factors, traditional single-output models cannot fully utilize the intrinsic correlations between variables, easily resulting in model overfitting or accuracy loss. Therefore, this invention employs multi-objective regression machine learning algorithms such as Multi-Objective Support Vector Regression (MO-SVR), Multi-Objective Random Forest (MO-RF), and Multilayer Perceptron (MLP) to construct the model. This allows multiple nutrient feature tasks to share hidden features and decision paths, while simultaneously coupling meteorological and topographic multi-source covariates as shared input features. This fully utilizes the strong correlations between nutrients and the environment-driven mechanism, significantly improving the overall accuracy and robustness of simultaneous prediction of multiple nutrients. This overcomes the shortcomings of existing technologies that only model for a single nutrient, resulting in large prediction bias and weak generalization ability.
[0116] This invention specifically employs the coefficient of determination (R²). 2 The root mean square error (RMSE) and relative prediction bias are used as evaluation metrics to assess the performance of the model's predictions. The formulas are described below:
[0117]
[0118]
[0119]
[0120] in, This represents the measured nutrient content of soil sample i. This indicates the predicted nutrient content of the soil sample. This represents the average nutrient content of the soil sample, where n represents the number of soil samples, and SD represents the standard deviation between the predicted nutrient content and the measured nutrient content.
[0121] This invention synthesizes the rating criteria of various scholars, and the referenced model accuracy verification criteria are as follows:
[0122] R 2 A value >0.85, RMSE <10% of the target variable mean, and RPD ≥2.0 indicate that the model has excellent predictive performance, can accurately predict the total amount and available nutrients in the soil, and can be applied to the quantitative prediction of soil nutrient information.
[0123] R 2 When the RMSE is greater than 0.80, the RPD is less than 20% of the average value of the target variable, and the RPD is greater than 1.5, it indicates that the model has good predictive performance and can basically achieve accurate prediction of total phosphorus and available nutrients in the soil. It can serve as an effective reference to assist in the non-destructive and efficient detection of field soil samples.
[0124] Ten-fold cross-validation was adopted as the model validation method to reduce the randomness caused by a single dataset split, and to make full use of multi-source datasets including multi-type soil samples and multi-source covariates. By performing multiple random splits, the generalization ability of the model prediction was improved.
[0125] like Figures 7 to 8 As shown, this invention employs machine learning algorithms such as K-Nearest Neighbors (KNN), Random Forest (RF), Multi-Objective Support Vector Machine (MO-SVR), Multi-Objective Linear Regression (MO-LR), Extreme Gradient Boosting (XGBoost), and Multilayer Perceptron (MLP) to construct models and compares the R-values of each model. 2 The RMSE validation index was used to compare the prediction performance of each model. Through comparison, the optimal multivariate soil nutrient content prediction model coupled with multi-source covariates and in-situ spectroscopy was found to be the XGBoost model. It achieved the highest or second-best performance in predicting both total soil nutrients and available soil nutrients, taking into account the prediction accuracy, prediction speed and prediction generalization of multi-objective nutrients.
[0126] XGBoost (Extreme Gradient Boosting) is an ensemble learning method based on gradient boosting decision trees. By sequentially adding multiple regression trees and utilizing second-order gradient information, it aims to minimize the overall loss function for rapid optimization, achieving accurate fitting of the complex nonlinear relationship between spectral features and nutrient content.
[0127] Key hyperparameters of XGBoost include tree-level sampling rate (colsample_bytree), number of decision trees (n_estimators), maximum tree depth (max_depth), and L1 regularization coefficient (reg_alpha). During the modeling process using machine learning algorithms, a global grid search is used to traverse the combination of hyperparameters and randomly select the optimal combination to maximize the model's prediction performance, balancing the model's prediction accuracy and generalization.
[0128] Its formula is described as follows:
[0129]
[0130] The above represents the overall objective function of XGBoost in the t-th iteration, mainly including the empirical loss term and the regularization term, where, This represents the measured nutrient content of the soil sample. This is the predicted value from the previous t-1 iterations. Let t be the regression tree to be learned. The loss function is preferably the squared error loss.
[0131] Specifically, This is a structure regularization term for the tree model, used to control model complexity. Its specific form is:
[0132] in, Let be the number of leaf nodes in the t-th tree. The weight of the j-th leaf node, and These are the first and second regularization coefficients, used to suppress overfitting.
[0133] After approximating the objective function using a second-order Taylor expansion, the optimization process can be transformed into a closed-form solution for the weights of each leaf node. The optimal leaf node weights are represented as follows:
[0134]
[0135] in, Let j be the set of samples that fall into the j-th leaf node. For the first-order gradient, It is a second-order gradient.
[0136] After all iterations are completed, the final prediction model can be represented as the sum of the regression trees from each round:
[0137]
[0138] Where K is the total number of iterations.
[0139] Through the XGBoost algorithm described above, this implementation method can significantly improve the prediction accuracy, generalization and robustness of soil nutrient content while controlling model complexity, and is suitable for multi-source data scenarios that couple covariates and in-situ spectra.
[0140] like Figure 9 As shown, the image is a scatter plot of the multi-objective prediction of total and available nutrients in soil using the XGBoost machine learning method coupled with multi-source covariates and in-situ spectral data. It indicates that the constructed XGBoost model achieves excellent predictions for total phosphorus, total potassium, available nitrogen, available phosphorus, readily available potassium, and organic matter content in soil, and can be applied to the quantitative prediction of soil nutrient information. In the model's prediction results for total nitrogen content, both RPD and RMSE indicators reached excellent levels. 2 The results were rated as good. Overall, the XGBoost model constructed in this invention can accurately predict total phosphorus and available nutrients in the soil.
[0141] This invention synthesizes the rating criteria of various scholars, and the reference machine learning model overfitting verification criteria are as follows:
[0142] R 2 _train - R 2 When _test < 0.05 and RMSE_test / RMSE_train < 1.1 (less than 10%), it indicates that the model prediction results have no risk of overfitting and can guide the model to be applied to actual production.
[0143] 0.10 > R 2 _train - R 2 When _test > 0.05, 1.1 > RMSE_test / RMSE_train > 1.2 (10% ~ 20%), it indicates that the model prediction results have a slight overfitting phenomenon, and there will be a slight deviation when applied to actual production.
[0144] 0.20 > R 2 _train - R 2When _test > 0.10, 1.5 > RMSE_test / RMSE_train > 1.2 (20% ~ 50%), it indicates that the model prediction results have obvious overfitting, and there will be obvious deviations when applied to actual production, making it unsuitable for quantitative prediction of soil nutrients.
[0145] Table 5 below shows a comparison of the training and test set performance of the XGBoost model coupled with multi-source covariates and in-situ spectral analysis for predicting soil nutrient content. The R-values of the XGBoost model in predicting various soil nutrient contents are shown in the table. 2 Relative deviation (R) 2 _train -R 2 The RMSE values for all samples (_test) are less than 0.05, indicating no risk of overfitting. The relative RMSE deviation for predicting total potassium is 1.1169, showing only slight overfitting. Furthermore, the relative RMSE deviations for predicting other nutrient contents are all less than 10%, indicating no risk of overfitting. Overall, the XGBoost model constructed in this invention can achieve accurate predictions of total phosphorus and available nutrients in soil.
[0146] Table 5: Performance Comparison of XGBoost Model in Predicting Soil Nutrient Content on Training and Test Sets
[0147]
[0148] To further verify the effect of introducing multi-source covariates such as topographic factors and meteorological data on improving the model's predictive ability, in the embodiments, the model is subjected to SHAP (Shapley Additive explanations) interpretability analysis to quantify the contribution of each input feature to the prediction results of total and available soil nutrients.
[0149] like Figure 10 As shown, this invention performs SHAP analysis and importance ranking on topographic and meteorological covariates in the prediction of soil nutrient content. The vertical axis represents multi-source covariates, including meteorological factors (Precip, windspeed, Temp_mean, dpt) and topographic factors (twi, aspect, slope, elevation). The horizontal axis represents the average SHAP value of each feature. The red dashed line in the figure represents the average SHAP value of all spectral band features, and the blue bars represent the actual contribution value of each covariate.
[0150] The results show that coupling multiple source covariates as model inputs has a high contribution to model prediction, basically at or above the average contribution level of the spectrum.
[0151] Table 6 further quantifies the contribution of multi-source covariates such as topographic and meteorological factors in the prediction of total and available soil nutrients. Among them, the sum of the SHAP values of wind speed is as high as 0.0649, which is about 15.83 times the spectral average (0.0649 / 0.0041≈15.83). The aspect, which has the lowest contribution among the covariates, has a contribution value of 0.0084, which is also 2.05 times the spectral average (0.0084 / 0.0041≈2.048). The above results fully demonstrate that by coupling topographic and meteorological factors as multi-source covariate input data, this invention significantly improves the model's explanatory power and prediction accuracy for soil nutrient content under diverse soil types at the provincial scale. It overcomes the problem of insufficient prediction stability caused by existing technologies that rely solely on spectral data and are sensitive to topographic relief and meteorological differences. It realizes in-situ prediction under different topographic and meteorological conditions at the provincial scale, significantly improving the robustness and accuracy of the model prediction.
[0152] Table 6: Contribution of Multi-Source Covariates such as Topography and Meteorology to Soil Nutrient Prediction
[0153]
[0154] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the scope of protection of the present invention in any way, and all technical solutions obtained by equivalent substitution or other means fall within the scope of protection of the present invention. Parts not covered in this invention are the same as or can be implemented using existing technology.
Claims
1. A multivariate type soil nutrient prediction method coupling multi-source covariates and in-situ spectra, characterized in that, The method comprises the following steps: Step 1, in-situ hyperspectral measurement of soil and data collection; Step 2, soil sample collection and preparation; Step 3, acquisition of covariate data: obtain meteorological data and topographic data of the corresponding sampling area according to the coordinate latitude and longitude; Step 4, acquisition of soil nutrient data: analyze the prepared soil samples to determine the soil nutrient content as the target variable data for subsequent model construction; Step 5, in-situ spectral data preprocessing: perform preprocessing operations on the in-situ soil spectrum; Step 6, spectral data outlier detection and elimination: detect outliers in the preprocessed spectral data and eliminate them to exclude potential anomalies; Step 7, feature dimension reduction and selection: use principal component analysis to reduce and select the features of the spectral data and multi-source covariates; Step 8, construction of multi-target regression prediction model for soil nutrient content: use multi-target regression machine learning algorithm to construct the soil nutrient prediction model, use global grid search method to find and adjust the optimal hyperparameters of each model, and use 10-fold cross-validation method to test and evaluate the model; Step 9, model prediction performance evaluation and comparison: randomly divide the test set from the modeling data set, use the constructed model to predict the total and available nutrient content of the soil samples, compare the predicted values with the actual nutrient content values, and compare the prediction performance of each model through verification, and select the optimal prediction model according to the comparison result; Step 10, quantification of the contribution of multi-source covariates in model prediction: use SHAP method to analyze and sort the importance of topographic and meteorological covariates in the prediction of soil nutrient content, and the comparison object is the actual SHAP contribution value of each covariate and the average SHAP value of all spectral band features.
2. The multivariate type soil nutrient prediction method of coupling multi-source covariates and in-situ spectra according to claim 1, characterized in that, In step 1, the sampling points are divided according to the equal interval grid method in the area to be measured according to the size of the field, and a 25cm deep soil profile with a cross section of 14cm×8cm is manufactured using a rectangular soil sampler according to the sampling points. Then, three positions are selected on the vertical profile of the soil sample, and the in-situ spectral data of the soil are collected. During the measurement process, the reflection probe is kept close to the soil surface for spectral collection. Ten spectra are collected at each position, and the arithmetic mean of the ten spectra is taken as the in-situ spectral data of the soil sample at that sampling point. The spectral range measured by the spectrometer is 350-2500nm.
3. The multivariate type soil nutrient prediction method of coupling multi-source covariates and in-situ spectra according to claim 1, characterized in that, In step 2, after the in-situ spectral data of each sampling point is collected, 2cm thick soil samples are taken from the same positions where the in-situ spectral measurement is performed on the soil profile using an aluminum box. The three soil samples are mixed and placed in a sealed bag for transportation. Then, the transported soil samples are placed on a plastic cloth and dried in a well-ventilated indoor environment. After drying, the stones and plant and animal residues in the soil samples are removed, and the soil samples are ground and sieved through a 1mm sieve and a 100 mesh sieve. The soil samples sieved through the 1mm sieve are used to measure the available nutrients of the soil, and the soil samples sieved through the 100 mesh sieve are used to measure the organic matter and total nutrients.
4. The multivariate type soil nutrient prediction method of coupling multi-source covariates and in-situ spectra according to claim 1, characterized in that, In step 3, after the soil in-situ spectral data of each sampling point is collected, the coordinate position of each sampling point is recorded using GPS, and the meteorological data and terrain data of the corresponding sampling area are obtained according to the coordinate latitude and longitude, wherein the meteorological data specifically collects average temperature, precipitation and average wind speed index information related to regional variability, and the terrain data is extracted by extracting the height, slope, slope direction and terrain humidity index terrain information of the corresponding sampling area in the DEM grid file through python.
5. The multivariate type soil nutrient prediction method of coupling multi-source covariates and in-situ spectra according to claim 1, characterized in that, In step 4, the organic matter, total nitrogen, total potassium, total phosphorus, alkali-hydrolyzable nitrogen, available phosphorus and available potassium contents of the soil sample are determined, wherein: The organic matter content in the soil sample is determined by potassium dichromate container weight method; The total nitrogen content in the soil sample is determined by semi-micro Kjeldahl nitrogen determination method; The total phosphorus, total potassium and available potassium contents in the soil sample are determined by ion-coupled plasma optical emission spectrometry method; The alkali-hydrolyzable nitrogen content in the soil sample is determined by alkali-hydrolyzation diffusion method; The available phosphorus content in the soil sample is determined by antimony molybdenum anti-colorimetric method.
6. The multivariate type soil nutrient prediction method of coupling multi-source covariates and in-situ spectra according to claim 1, characterized in that, In step 5, the preprocessing method is combined as extended multiplicative scatter correction and normalization processing, wherein the polynomial order of the extended multiplicative scatter correction is set to 2; in step 6, the preprocessed spectral data are subjected to outlier detection based on dynamic threshold isolation forest method, and the parameter is set to n_estimators=200 and random_state=42 to ensure the repeatability of the results.
7. The multivariate type soil nutrient prediction method of coupling multi-source covariates and in-situ spectra according to claim 1, characterized in that, In step 9, the XGBoost model is selected as the multivariate type soil nutrient content prediction model by comparison, the XGBoost model adds multiple regression trees in sequence and uses second-order gradient information to minimize the overall loss function for fast optimization, thereby realizing accurate fitting of the complex nonlinear relationship between spectral features and nutrient content.
8. The multivariate soil nutrient prediction method of coupling multi-source covariates and in-situ spectra according to claim 7, characterized in that, The key hyperparameters of the XGBoost model include tree level sampling rate, number of decision trees, maximum tree depth and L1 regularization coefficient, and in the modeling process, the optimal combination is selected to give the strongest prediction performance of the model by using global grid search to traverse the hyperparameter combination, taking into account the prediction accuracy and generalization of the model; the formula is described as follows: The above formula represents the overall objective function of XGBoost in the tth iteration, including an empirical loss term and a regularization term, where, is the measured nutrient content of the soil sample, is the predicted value of the previous t-1 iterations, is the tth regression tree to be learned, is the loss function; In particular, The structure regularization term for tree model is used to control the model complexity, and its specific form is: wherein, is the number of leaf nodes in the tth tree, is the weight of the jth leaf node, and are first and second regularization coefficients, respectively, for suppressing overfitting; After the target function is approximated by second-order Taylor expansion, the optimization process is converted into closed-form solution for each leaf node weight, and the optimal leaf node weight is represented as follows: wherein, is the set of samples falling into the jth leaf node, is the first order gradient, is the second order gradient; When all iterations are completed, the final prediction model is represented as the sum of each round of regression tree: Wherein K is the total number of iterations.
9. The multivariate type soil nutrient prediction method of coupling multi-source covariates and in-situ spectra according to claim 1, characterized in that, In step 9, the prediction performance of each model is compared by R2, RMSE and RPD verification indexes, and the optimal prediction model is selected according to the comparison result; the expression is as follows: wherein, represents the measured nutrient content of soil sample i, represents the predicted nutrient content of soil sample i, represents the average of the nutrient content of soil samples, n represents the number of soil samples, and SD represents the standard deviation of the predicted nutrient content from the measured nutrient content.
10. The multivariate type soil nutrient prediction method coupled with multiple source covariates and in-situ spectrum according to claim 9, characterized in that, When R 2 > 0.85, RMSE < 10% of the mean of the target variable, and RPD ≥ 2.0, it indicates that the model has excellent prediction performance and can achieve accurate prediction of the total and available nutrients in the soil, and can be applied to quantitative prediction of soil nutrient information. When R 2 >0.80, RMSE < 20% of the mean of the target variable, and 2) RPD ≥ 1.5, it indicates that the model has good prediction performance and can basically achieve accurate prediction of total and available nutrients in soil, which can be used as an effective reference to assist in non-destructive and efficient detection of soil samples in the field.