Soil organic matter content prediction method based on hyperspectrum

Through the improved spotted hyena optimization ISHO and IRIV algorithms, and combined with pretreatment methods and XGBoost model, the problems of noise and information redundancy in soil organic matter content prediction are solved, achieving high-precision soil organic matter content prediction.

CN120048382AInactive Publication Date: 2025-05-27HUZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510118023.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has noise and information redundancy in the prediction of soil organic matter content, resulting in poor loss of some effective bands and optimization results, and has a great impact on the external environment, resulting in strong collinearity between the selected bands and cannot better reflect the spectral information related to soil organic matter.

Method used

The improved spotted hyena optimization ISHO algorithm and iterative retained information variable IRIV algorithm were used to screen the soil spectral reflectivity data in characteristic bands, combining Savitzky-Golay smoothing, multivariate scattering correction and first-order differential pretreatment methods to enhance data correlation, and an extreme gradient enhancement XGBoost method was used to establish a soil organic matter content prediction model.

Benefits of technology

Effectively remove random noise and baseline drift, reduce multiple scattering effects, reduce data redundancy and reduce interband collinearity, and improve the accuracy and stability of soil organic matter content prediction models, providing important reference and guidance for soil carbon cycle research and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048382A_ABST
    Figure CN120048382A_ABST
Patent Text Reader

Abstract

The invention provides a soil organic matter content prediction method based on hyperspectrum, and relates to the technical field of hyperspectrum prediction. The method comprises the following steps: firstly, collecting and treating a soil sample, and carrying out SOM content measurement and soil spectral reflectivity measurement; the soil spectral reflectivity data are preprocessed, and spectral noise and redundant information are eliminated; secondly, preliminarily screening characteristic wave bands by adopting an improved spotted serow optimization ISHO algorithm, and then carrying out secondary screening on the characteristic wave bands by utilizing an iterative reserved information variable IRIV algorithm; and finally, establishing a hyperspectral prediction model to predict the SOM content, and evaluating the precision of the model. According to the method, the ISHO algorithm and the IRIV algorithm are adopted to perform characteristic wave band selection on the spectral reflectivity data, and the XGBoost method is adopted to establish the SOM content prediction model, so that the precision of the prediction model is improved, and theoretical and technical support is provided for SOM content monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hyperspectral prediction, and in particular to a method for predicting soil organic matter content based on hyperspectral. Background Art

[0002] Soil organic matter (SOM) is a general term for various carbon-containing organic compounds in the soil. It is a key parameter in many applications such as soil resource evaluation, precision agricultural management, and ecological and environmental protection. Therefore, achieving efficient and accurate estimation of SOM content is of great significance for global climate regulation and carbon cycle. Traditional SOM content detection is mainly achieved through chemical analysis methods such as Rockwell method, calcination method, and oxidation method. Although its detection accuracy is high, it has problems such as high cost, time-consuming operation, and environmental damage, and is not suitable for large-scale estimation of SOM content. In recent years, hyperspectral technology has gradually become the most promising means to replace traditional methods with its pollution-free, efficient, accurate, economical, and non-destructive characteristics. Using this technology, the SOM content can be accurately estimated based on the spectral reflectance of a small amount of soil samples.

[0003] In order to study the SOM content prediction method based on hyperspectral technology, many scholars have used extreme gradient boosting XGBoost, partial least squares regression, random forest, convolutional neural network and back propagation neural network to build models. However, there is a lot of noise in hyperspectral data, and the information redundancy between bands is high. Therefore, preprocessing and feature band selection methods are needed for denoising and dimensionality reduction. Common preprocessing methods include Savitzky-Golay (SG), standard normal variate (SNV), multiplicative scatter correction (MSC), first derivative (FD), logarithmic reciprocal (1 / logR), etc. Common feature band selection methods include competitive adaptive reweighted sampling (CARS) algorithm, particle swarm optimization (PSO) algorithm, ant colony optimization algorithm and simulated annealing algorithm.

[0004] Although the above-mentioned characteristic band selection algorithm can effectively eliminate irrelevant and invalid characteristic bands, there are still problems such as the loss of some effective bands and poor optimization effect. In addition, the use of a single band selection method is easily affected by the external environment, and the collinearity between the selected bands is strong, which cannot well reflect the spectral information related to SOM. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide a soil organic matter content prediction method based on hyperspectral to achieve SOM content prediction in view of the above-mentioned deficiencies in the prior art.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: a method for predicting soil organic matter content based on hyperspectral, comprising the following steps:

[0007] Step 1: Collect and process soil samples to determine SOM content and soil spectral reflectance;

[0008] The processed soil samples were divided into two parts, which were used for soil spectral reflectance data measurement and SOM content determination respectively; the SOM content was determined by potassium dichromate external heating method;

[0009] Step 2: Determine the soil spectral reflectance data; use a ground object spectrometer to collect the spectral reflectance data of the soil sample in a dark room, the wavelength range of the spectral reflectance data is 350 to 2500nm, and the sampling interval is 1nm;

[0010] Step 3: Preprocess the soil spectral reflectance data to eliminate spectral noise and redundant information;

[0011] Firstly, the Savitzky-Golay smoothing method is used to denoise the soil spectral reflectance data to obtain the original spectral reflectance data R. In order to further enhance the correlation between the soil spectral reflectance data and the SOM content, the original spectral reflectance data R is further corrected for multivariate scattering and then preprocessed for the first-order differential.

[0012] Step 4: Use the band selection algorithm to screen the characteristic bands of the preprocessed spectral reflectance data; first, use the improved spotted hyena optimized ISHO algorithm to preliminarily screen the characteristic bands, and then use the iterative information retention variable IRIV algorithm to perform secondary screening on the characteristic bands;

[0013] The ISHO algorithm uses Tent chaos mapping and binary discretization to optimize the initial individual position of spotted hyenas, and uses nonlinear decreasing inertia weight and dynamic Laplace crossover strategy to optimize the position update process of spotted hyenas;

[0014] The process of preliminary screening of characteristic bands using the ISHO algorithm is as follows:

[0015] (1) Set the parameters of the ISHO algorithm; set the number of spotted hyenas SHN, the spectral reflectance data matrix I and the maximum number of iterations MM Itn ; Among them, the spectral reflectance data matrix I contains s samples and m bands. Represents a set of m bands; construct an m-dimensional vector space, then set a population consisting of q spotted hyena individuals, where q < m, and search for prey in the m-dimensional space. Using the fitness function as a criterion, continuously update the positions of the spotted hyenas until the maximum number of iterations is reached, and finally select a characteristic band combination from the band set Select the characteristic band combination 1 ≤ i ≤ m;

[0016] (2) Set the spotted hyena individuals to represent a group of selected characteristic bands, and the position is represented by a vector with the same number as the total number of bands; regard the spotted hyena individual with the best fitness value in the current search space as the target prey, and the remaining spotted hyenas update their positions according to the position of the target prey; the calculation process is as follows:

[0017]

[0018] In the formula: is the distance between the prey and the spotted hyena; x is the current number of iterations; is the position vector of the current spotted hyena, y n represents the spectral reflectance data corresponding to the nth band; is the position vector of the current prey; and are the wobbling factor and the convergence factor respectively, as shown in the following formula:

[0019]

[0020] In the formula: and are random vectors in the interval [0, 1]; is the control factor, linearly decreases from 5 to 0; Itn = 1, 2,..., MM Itn ;

[0021] Initialize the spotted hyena population using the Tent chaotic map, as shown in the following formula:

[0022]

[0023] In the formula, y i ' is the chaotic variable, that is, the new spectral reflectance value;

[0024] After the Bernoulli displacement transformation, we can get:

[0025] y i ' = (2y i ) mod 1 (7)

[0026] (3) Discretize the individuals of the spotted hyena population, as shown in the following formula:

[0027]

[0028] Where: It is the result of binary discretization; is the result after chaos initialization; Thd is the binary discretization threshold, set Thd = 0.5; i is the index of the spotted hyena individual; j is the dimension index of the spotted hyena individual;

[0029] (4) Design the fitness function shown in the following formula to calculate the fitness value of each spotted hyena individual:

[0030]

[0031] Where: F is the fitness function; α is the performance weight parameter of the regression model; μ is the weight parameter for selecting the number of feature bands; α∈[0,1], μ=1-α; t is the number of selected feature bands; T is the total number of bands; J is the determination coefficient;

[0032] (5) The position of the individual spotted hyena with the smallest fitness value is taken as the prey position

[0033] (6) Determine whether the current number of iterations x has reached the maximum number of iterations. If so, output the prey position. Output 1 represents selection, 0 represents unselected, and the selected feature band combination is obtained Otherwise, continue to step (7);

[0034] (7) Determine whether the currently processed spotted hyena index i reaches the spotted hyena number SHN. If so, increase the current iteration number x by one and re-execute step (5); otherwise, continue to execute step (8);

[0035] (8) Calculate the nonlinear decreasing inertia weight under the current iteration number x. The inertia weight w is expressed as:

[0036]

[0037] Where: w st represents the weight at the beginning of the iteration, w st =0.9; w ed Represents the weight at the end of the iteration, w ed =0.4;

[0038] (9) Update the individual position of the spotted hyena according to the inertia weight w, as shown in the following formula:

[0039]

[0040] Where: It is the best location for spotted hyenas; is the location of the other spotted hyenas; is a set of SHN optimal solutions; is the updated optimal solution position, and the positions of other spotted hyenas are updated according to this position;

[0041] (10)Judgment Is it true? r is a random number between [0,1]. If it is true, the dynamic Laplace crossover strategy is used to update the position of the spotted hyena individual again, as shown in the following formula:

[0042]

[0043] Where: and are the positions of the offspring individuals produced after the Laplace operator crossover; and are the positions of the two parent individuals with the best fitness in the solution space; β is a random number generated in reverse according to the Laplace distribution function, as shown in the following formula:

[0044]

[0045] Where: c∈R is the location parameter; u is a random number in the interval [0,1] that follows a uniform distribution;

[0046] If it is not true, substitute the β value set in formula (17) into formula (14) (15) to update the position of the individual spotted hyena;

[0047]

[0048] (11) Calculate the fitness value of the spotted hyena individual after the position is updated, compare the fitness value of the current best spotted hyena individual with the fitness value of the prey, if the fitness value of the current best spotted hyena individual is less than the fitness value of the prey, take the current best spotted hyena individual position as the new prey position, and update the position of the best spotted hyena Otherwise keep the prey position and the best spotted hyena position constant;

[0049] (12) The spotted hyena index i is increased by 1, and step (7) is executed again;

[0050] The process of secondary screening of characteristic bands using the IRIV algorithm is as follows:

[0051] remember is the characteristic band set after ISHO algorithm screening, s is the number of soil samples, the spectral reflectance data matrix is ​​m' columns, and the soil samples are randomly divided into training set and validation set;

[0052] 1) For the matrix with s rows and m' columns in the training set, it contains only 1 and 0 1 and 0 respectively indicate whether the variable is used for modeling. The XGBoost model is established for each row of , and the RMSECV value obtained by 5-fold crossover is used as the evaluation standard. The vector of size s×1 is recorded as RMSECV 0 ; The matrix In the i'th column (), replace 1 with 0 and 0 with 1 to obtain the matrix B, i'=1,2,…,m'. Similarly, build an XGBoost model in each row of the matrix B to obtain a vector of size s×1, denoted as RMSECV i' ;

[0053] 2) Definition and To evaluate the importance of each feature band, the formula is as follows:

[0054]

[0055] Where: k th represents the kth row in the vector, then k th RMSECV 0 and k th RMSECV i' Represents vector RMSECV 0 and RMSECV i' The value of the kth row in ; and The mean value is denoted as M i',in and M i',out , subtract the two means to get DM i' ; If DM i' <0, the characteristic band is a strong information characteristic band or a weak information characteristic band; if DM i' >0, the characteristic band is a non-information characteristic band or an interference characteristic band; P = 0.05 is defined as the threshold for the Mann-Whitney U test, and the characteristic bands are finally divided into 4 categories;

[0056] 3) In each iteration, strong information feature bands and weak information feature bands are retained, and no information feature bands and interference feature bands are eliminated; return to step 1) for the next round of iteration until only strong information feature bands and weak information feature bands remain;

[0057] 4) Perform reverse elimination on the tt retained feature bands; First, establish an XGBoost model for the tt feature bands to obtain RMSECV t ; Then, by eliminating the jth feature band, the XGBoost model is established for the tt-1 feature bands to obtain RMSECV -j,j=1,2,…,tt;if RMSECV -j Less than RMSECV t Then the jth characteristic band is eliminated, otherwise it is retained; this process is repeated, and the remaining characteristic bands are the final selected characteristic bands D;

[0058] Step 5: Establish a hyperspectral prediction model to predict SOM content and evaluate the model accuracy;

[0059] The characteristic bands selected by the characteristic band selection algorithm were set as independent variables, and the SOM content was set as the dependent variable. The extreme gradient boosting XGBoost method was used to construct a hyperspectral prediction model to predict the SOM content.

[0060] The characteristic band combination D={(x v ,y v ):v=1,…,s,x v ∈R e ,y v ∈R}, that is, s soil samples with characteristic dimension e, x v represents the feature vector of the vth soil sample, whose dimension is e, and it contains all the feature information related to the vth soil sample; v represents the true value of the SOM content of the vth soil sample in the dataset;

[0061] definition The predicted value of SOM content of soil samples is:

[0062]

[0063] Among them, f t is the model function of the regression tree, F is the function space of the regression tree, which defines the form and parameter constraints of the regression tree; the goal of XGBoost is to iteratively select the optimal tree f from F t To minimize the objective function, thereby improving the prediction accuracy of the model, f t (x v ) represents the prediction score of the t-th tree for the v-th soil sample, f t By minimizing the loss function:

[0064]

[0065] Among them, Obj is the objective function, l represents the loss function, which is used to measure the gap between the predicted value and the true value; in order to avoid overfitting, Ω(f k ) is a regularization term that defines the complexity of the tree:

[0066]

[0067] Where K represents f t The number of leaf nodes, ww represents the weight of each leaf, γ and λ are regularization parameters;

[0068] The training of the superposition model adopts an iterative method, and the iterative form of the objective function is:

[0069]

[0070] In the formula, is the predicted value of the vth sample at the k-1th iteration, and constant is a constant term;

[0071] Perform a second-order Taylor expansion on the above formula (22), and the objective function is rewritten as:

[0072]

[0073] The candidate feature set of each tree split node is defined as I j ={v|q(x v )=j}, define G j is the sum of the gradients of all samples on leaf node j; H j is the sum of the second-order derivatives of all samples on leaf node j, g v is the first-order derivative of the loss function, h v is the second-order derivative of the loss function, then the final objective function is expressed as:

[0074]

[0075] For a tree with a fixed structure, the weight ww of leaf j is j * And the optimal solution of the objective function Obj * for:

[0076]

[0077] Since it is impossible to enumerate all tree structures, a greedy algorithm is introduced in XGBoost to split the existing leaf nodes in an iterative manner and calculate the maximum gain obtained. Let I L and I R are the sample sets of the left and right nodes after the node split, respectively, I=I L ∪I R , then the gain Gain is defined as:

[0078]

[0079] If Gain is greater than 0, it means that the loss after the split is reduced, that is, the split can be made, otherwise the split will be stopped.

[0080] The beneficial effects of the above technical solution are as follows: the method for predicting soil organic matter content based on hyperspectral provided by the present invention removes the influence of factors such as random noise, baseline drift and multiple scattering effects by pre-processing soil samples, adopting indoor spectral measurement means, taking the pre-processed soil spectral reflectance data as the research object, analyzing the spectral response band of SOM content, and mining spectral characteristic bands. In order to reduce data redundancy and reduce collinearity between bands, the ISHO algorithm and the IRIV algorithm are used to select characteristic bands of spectral reflectance data, and the XGBoost method is used to establish a SOM content prediction model, thereby improving the accuracy of the prediction model and providing theoretical and technical support for SOM content monitoring. It provides important reference and guidance for soil carbon cycle research and soil management. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 The study area and the spatial distribution map of soil sampling points provided by the embodiment of the present invention, wherein (a) is the study area, and (b) is the spatial distribution of soil sampling points;

[0082] Figure 2 A flow chart of a method for predicting SOM content based on hyperspectral provided in an embodiment of the present invention;

[0083] Figure 3 A schematic diagram of a process of a hyperspectral-based SOM content prediction method provided by an embodiment of the present invention;

[0084] Figure 4 Spectral reflectance curves at different SOM contents provided by the embodiment of the present invention;

[0085] Figure 5 The correlation coefficient curve graph between SOM content and spectral reflectance provided in the embodiment of the present invention, wherein (a) is a correlation coefficient graph between original spectral reflectance data and SOM content, (b) is a correlation coefficient graph between spectral data after 1 / logR preprocessing and SOM content, (c) is a correlation coefficient graph between spectral data after FD preprocessing and SOM content, (d) is a correlation coefficient graph between spectral data after MSC preprocessing and SOM content, (e) is a correlation coefficient graph between spectral data after 1 / logR-FD preprocessing and SOM content, and (f) is a correlation coefficient graph between spectral data after MSC-SD preprocessing and SOM content;

[0086] Figure 6 A flowchart of using the ISHO algorithm to perform initial screening of characteristic bands provided by an embodiment of the present invention;

[0087] Figure 7The spectral characteristic band diagram selected by the ISHO algorithm provided in an embodiment of the present invention, wherein (a) is the band selection probability determined by the ISHO algorithm, and (b) is the characteristic band distribution after the ISHO algorithm selection;

[0088] Figure 8 An iterative retained variable number graph for secondary screening using the IRIV algorithm provided in an embodiment of the present invention;

[0089] Fig. 9 A graph of SOM content results predicted by the XGBoost model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0090] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0091] In this embodiment, a soil sample collection area located in the northeast of Wuxing District, Huzhou City, Zhejiang Province, China is taken as an example, and the soil organic matter content prediction method based on hyperspectral of the present invention is used to predict the soil organic matter content in the area.

[0092] The longitude and latitude range of the soil sample collection area is 119°51′—120°22′ east longitude and 30°37′—30°57′ north latitude. It has high temperatures in summer, mild winters, distinct four seasons, sufficient sunlight, and abundant rainfall. It has a subtropical monsoon climate with an average annual temperature of about 16.3°C and an average annual rainfall of 1303.4 mm. The soil types in the study area are mainly paddy soil and tidal soil, and the SOM content in the area is rich. The collection of soil samples was completed in December 2023, and a total of 99 0-20 cm surface soil samples were collected, including farmland, highway and construction land samples. The spatial distribution of the study area and soil sampling points is as follows: Figure 1 shown.

[0093] In this embodiment, a soil organic matter content prediction method based on hyperspectral, such as Figure 2 , 3 As shown, the following steps are included:

[0094] Step 1: Collect and process soil samples to determine SOM content and soil spectral reflectance;

[0095] In this embodiment, 99 soil samples were collected, and the soil samples were naturally air-dried, ground and sieved; the processed soil samples were divided into two parts, which were used for soil spectral reflectance data measurement and SOM content determination respectively; the SOM content was determined by potassium dichromate external heating method;

[0096] In this embodiment, a handheld GPS locator is used to record the longitude and latitude coordinates of the sampling points during the soil sample collection process. When collecting soil samples, the five-point sampling method is used. About 1 kg of soil samples are collected at each sampling point. After the collected soil samples are sealed and marked, they are taken back to the laboratory for natural air drying, and stones, weed roots and other impurities are removed. After grinding, they are sieved with 2 mm. The soil samples are divided into two parts, one for indoor spectral measurement and the other for SOM content measurement. The SOM content is determined by potassium dichromate volumetric method-external heating method, and the SOM content measurement statistics are shown in Table 1. It can be seen from Table 1 that the SOM mass fraction in the total sample and training set varies from 3.904 g / kg to 58.827 g / kg, and the maximum SOM content in the validation set is 57.543 g / kg, and the minimum is 5.790 g / kg. The average values ​​of the total sample, training set and validation set were 24.853, 25.730 and 23.388 g / kg, the standard deviations were 12.436, 12.859 and 11.525, and the coefficients of variation were 49.8%, 49.6% and 48.5%, respectively, showing moderate variability, indicating that the spatial heterogeneity of SOM was significant. The skewness of the three sample sets was 0.811, 0.757 and 0.772, and the kurtosis was 0.357, 0.119 and 0.616, respectively, all of which were right-deviated from the normal distribution and had a flat kurtosis. According to the statistical results, the training set and validation set were close in statistical indicators, and the statistical characteristics of the total sample were slightly different, indicating that the training set and validation set were well representative of the total sample and could be used to establish and verify the SOM estimation model.

[0097] Table 1 Descriptive statistics of SOM content

[0098]

[0099] Step 2: Determine the soil spectral reflectance data; use a ground object spectrometer to collect the spectral reflectance data of the soil sample in a dark room, the wavelength range of the spectral reflectance data is 350nm to 2500nm, and the sampling interval is 1nm;

[0100] In this embodiment, the spectral data is measured using an ASDFieldSpec4 ground feature spectrometer, and its effective spectral range is 350nm to 2500nm. Before spectrum acquisition, preheat the instrument for 30 minutes and perform whiteboard calibration; the field of view angle is less than 25°, and the fiber optic probe is perpendicular to the surface of the soil sample at about 15cm. Ten spectral curves are collected for each soil sample. After removing abnormal spectral values, the remaining curves are averaged to obtain the final soil sample spectral reflectance data. All spectral curves uniformly remove the edge bands with low signal-to-noise ratio (350nm to 399nm and 2401nm to 2500nm), and finally select 2001 spectral data in the range of 400nm to 2400nm for each soil sample for subsequent processing and analysis. The soil samples were divided into five categories according to their SOM content: ≤10g / kg, 10g / kg-20g / kg, 20g / kg-30g / kg, 30g / kg-40g / kg, and ≥40g / kg. The average spectral reflectance of each category of samples was calculated, and the spectral reflectance curve was drawn as shown in the figure. Figure 4 As shown. It can be seen that the shapes of the five types of spectral curves change basically the same, and the SOM content is negatively correlated with the spectral reflectance. Due to the strong influence of atmospheric water vapor, the spectral reflectance fluctuates greatly between 1350nm~1450nm, 1800nm~2000nm and 2100nm~2250nm, and there are obvious moisture absorption valleys near 1400nm, 1900nm and 2200nm. In the range of 400nm~1000nm, the spectral reflectance shows an upward trend with the increase of wavelength; after 1000nm, except for the moisture absorption valley, the spectral reflectance tends to be flat as a whole.

[0101] Step 3: Preprocess the soil spectral reflectance data to eliminate spectral noise and redundant information;

[0102] Firstly, the Savitzky-Golay smoothing method is used to denoise the soil spectral reflectance data to obtain the original spectral reflectance data R. In order to further enhance the correlation between the soil spectral reflectance data and the SOM content, the original spectral reflectance data R is further subjected to multivariate scatter correction (MSC) and first-order differential (FD) preprocessing.

[0103] Since the spectral reflectance is inevitably affected by factors such as random noise, baseline drift and multiple scattering effects. This embodiment uses the SG filtering method to smooth and denoise the spectral curve to obtain R, and then performs 5 types of spectral pretreatments on R, including 1 / logR, FD, MSC, 1 / logR-FD and MSC-FD, which can effectively reduce the interference of external factors on R and make the spectral feature differences between samples more obvious, so as to further study the influence of different pretreatment methods on the correlation between spectral reflectance and SOM content. In this embodiment, Pearson correlation analysis was performed on the original spectral reflectance R and the spectral reflectance pretreated with 1 / logR, FD, MSC, 1 / logR-FD and MSC-FD and the SOM content, and the results are as follows. Figure 5 shown.

[0104] Depend on Figure 5 (a) It can be seen that SOM content is negatively correlated with R in the entire band, and the entire correlation coefficient curve is relatively smooth. The overall correlation is strongest in the 578nm~616nm band, and the absolute value of the correlation coefficient |r| reaches a maximum value of 0.567 at 587nm (p<0.01).

[0105] Depend on Figure 5 (b) It can be seen that after 1 / logR pretreatment, the SOM content was positively correlated with 1 / logR, and the correlation coefficient curve was similar to the symmetrical distribution of R. The absolute value of the correlation coefficient |r| reached a maximum value of 0.564 at 587nm (p<0.01).

[0106] Depend on Figure 5 (c) It can be seen that after FD preprocessing, the correlation is alternating between positive and negative, the correlation coefficient fluctuates greatly in the whole band, there are multiple absorption peaks, and the absolute value of the correlation coefficient reaches a maximum value of 0.539 at |r| at 443nm (p<0.01).

[0107] Depend on Figure 5 (d) It can be seen that after MSC pretreatment, the correlation coefficient curve has both positive and negative correlations, with large fluctuations. After 428nm, it is lower than the correlation between SOM content and R. The absolute value of the correlation coefficient |r| reaches a maximum value of 0.561 at 587nm (p<0.01).

[0108] Depend on Figure 5 (e) It can be seen that after 1 / logR-FD pretreatment, the correlation coefficient curve fluctuates violently and lacks regularity. The bands with higher correlation are mainly located between 431-440, 1419-1426 and 2153-2206nm. The absolute value of the correlation coefficient |r| reaches a maximum value of 0.537 at 443nm (p<0.01).

[0109] Depend on Figure 5(f) It is known that after MSC-FD pretreatment, the correlation coefficient curve fluctuates positively and negatively, and the peaks and valleys of the curve are easier to identify than R. The bands with higher correlation are mainly concentrated in the ranges of 488-528, 826-891 and 2160-2182nm. Bands with strong correlation with SOM content appear at wavelengths of 849, 858 and 862nm. The absolute value of the correlation coefficient |r| reaches a maximum value of 0.628 at 862nm (p<0.01).

[0110] In summary, after the spectral reflectance was preprocessed by five methods, the correlation between the spectral reflectance and the SOM content was improved to varying degrees in some bands. The correlation between the spectral reflectance and the SOM content after MSC-FD preprocessing was the most improved, and the absolute value of the maximum correlation coefficient between R and SOM content was increased by 0.061. Therefore, in this embodiment, the spectral data after MSC-FD preprocessing was selected for subsequent feature band screening and model construction.

[0111] Step 4: Use the band selection algorithm to screen the characteristic bands of the preprocessed spectral reflectance data; first, use the improved spotted hyena optimized ISHO algorithm to preliminarily screen the characteristic bands, and then use the iterative information retention variable IRIV algorithm to perform secondary screening on the characteristic bands;

[0112] Since hyperspectral images have many spectral bands, serious data redundancy, a large amount of useless and interfering information, and strong collinearity between bands, the stability and prediction ability of the model will be greatly affected. Therefore, the present invention uses a band selection algorithm to screen the characteristic bands of spectral reflectance data; constructs an ISHO-IRIV algorithm, first uses the ISHO algorithm to preliminarily screen the characteristic bands; and then uses the IRIV algorithm to perform a secondary screening of the characteristic bands;

[0113] The ISHO algorithm uses Tent chaos mapping and binary discretization to optimize the initial individual position of spotted hyenas, and uses nonlinear decreasing inertia weight and dynamic Laplace crossover strategy to optimize the position update process of spotted hyenas;

[0114] The process of using ISHO algorithm to preliminarily screen the characteristic bands is as follows: Figure 6 As shown, including:

[0115] (1) Set the parameters of the ISHO algorithm; set the number of spotted hyenas SHN, the spectral reflectance data matrix I and the maximum number of iterations MM Itn ; Among them, the spectral reflectance data matrix I contains s samples and m bands. Represents a set of m bands; construct an m-dimensional vector space, then set a population consisting of q spotted hyena individuals, where q < m, and search for prey in the m-dimensional space. Using the fitness function as a criterion, continuously update the positions of the spotted hyenas until the maximum number of iterations is reached. Finally, select the characteristic band combination from the band set Select the characteristic band combination 1 ≤ i ≤ m;

[0116] (2) Set the spotted hyena individuals to represent a group of selected characteristic bands, and the position is represented by a vector with the same number as the total number of bands; regard the spotted hyena individual with the best fitness value in the current search space as the target prey, and the remaining spotted hyenas update their positions according to the target prey position; the calculation process is as follows:

[0117]

[0118] In the formula: is the distance between the prey and the spotted hyena; x is the current number of iterations; is the position vector of the current spotted hyena, y n represents the spectral reflectance data corresponding to the nth band; is the position vector of the current prey; and are the wobbling factor and the convergence factor respectively, as shown in the following formula:

[0119]

[0120] In the formula: and are random vectors in the interval [0, 1]; is the control factor, linearly decreases from 5 to 0; Itn = 1, 2,..., MM Itn ;

[0121] Initialize the spotted hyena population using the Tent chaotic map, as shown in the following formula:

[0122]

[0123] In the formula, y i ' is the chaotic variable, that is, the new spectral reflectance value;

[0124] After the Bernoulli displacement transformation, it can be obtained:

[0125] y i ' = (2y i ) mod 1 (7)

[0126] (3) Discretize the individuals of the spotted hyena population, as shown in the following formula:

[0127]

[0128] Where: It is the result of binary discretization; is the result after chaos initialization; Thd is the binary discretization threshold, set Thd = 0.5; i is the index of the spotted hyena individual; j is the dimension index of the spotted hyena individual;

[0129] (4) Design the fitness function shown in the following formula to calculate the fitness value of each spotted hyena individual:

[0130]

[0131] Where: F is the fitness function; α is the performance weight parameter of the regression model (XGBoost); μ is the weight parameter for selecting the number of feature bands; α∈[0,1], μ=1-α; t is the number of selected feature bands; T is the total number of bands; J is the determination coefficient;

[0132] In the process of optimizing the fitness function, the impact of the regression model performance on the fitness function is given priority, followed by the number of characteristic bands. Therefore, it is necessary to increase the weight parameter α, taking α = 0.99. Thus, by minimizing the fitness function F, t can be minimized while maximizing J, so as to ensure that the performance of the regressor is optimal while the number of selected characteristic bands is the least and most representative.

[0133] (5) The position of the individual spotted hyena with the smallest fitness value is taken as the prey position

[0134] (6) Determine whether the current number of iterations x has reached the maximum number of iterations. If so, output the prey position. Output 1 represents selection, 0 represents unselected, and the selected feature band combination is obtained Otherwise, continue to step (7);

[0135] (7) Determine whether the currently processed spotted hyena index i reaches the spotted hyena number SHN. If so, increase the current iteration number x by one and re-execute step (5); otherwise, continue to execute step (8);

[0136] (8) Calculate the nonlinear decreasing inertia weight under the current iteration number x. The inertia weight w is expressed as:

[0137]

[0138] Where: w st represents the weight at the beginning of the iteration, w st =0.9; w edRepresents the weight at the end of the iteration, w ed =0.4;

[0139] (9) Update the individual position of the spotted hyena according to the inertia weight w, as shown in the following formula:

[0140]

[0141] Where: It is the best location for spotted hyenas; is the location of the other spotted hyenas; is a set of SHN optimal solutions; is the updated optimal solution position, and the positions of other spotted hyenas are updated according to this position;

[0142] (10)Judgment Is it true? r is a random number between [0,1]. If it is true, the dynamic Laplace crossover strategy is used to update the position of the spotted hyena individual again, as shown in the following formula:

[0143]

[0144] Where: and are the positions of the offspring individuals produced after the Laplace operator crossover; and are the positions of the two parent individuals with the best fitness in the solution space; β is a random number generated in reverse according to the Laplace distribution function, as shown in the following formula:

[0145]

[0146] If it is not true, substitute the β value set in formula (17) into formula (14) (15) to update the position of the individual spotted hyena;

[0147]

[0148] (11) Calculate the fitness value of the spotted hyena individual after the position is updated, compare the fitness value of the current best spotted hyena individual with the fitness value of the prey, if the fitness value of the current best spotted hyena individual is less than the fitness value of the prey, take the current best spotted hyena individual position as the new prey position, and update the position of the best spotted hyena Otherwise keep the prey position and the best spotted hyena position constant;

[0149] (12) The index i of the currently processed spotted hyena is increased by 1, and step (7) is executed again;

[0150] The process of secondary screening of characteristic bands using the IRIV algorithm is as follows:

[0151] remember is the characteristic band set after ISHO algorithm screening, s is the number of soil samples, the spectral reflectance data matrix is ​​m' columns, and the soil samples are randomly divided into training set and validation set;

[0152] 1) For the matrix with s rows and m' columns in the training set, it contains only 1 and 0 1 and 0 respectively indicate whether the variable is used for modeling. The XGBoost model is established for each row of , and the RMSECV value obtained by 5-fold crossover is used as the evaluation standard. The vector of size s×1 is recorded as RMSECV 0 ; The matrix In the i'th column (), replace 1 with 0 and 0 with 1 to obtain the matrix B, i'=1,2,…,m'. Similarly, build an XGBoost model in each row of the matrix B to obtain a vector of size s×1, denoted as RMSECV i' ;

[0153] 2) Definition and To evaluate the importance of each feature band, the formula is as follows:

[0154]

[0155] Where: k th represents the kth row in the vector, then k th RMSECV 0 and k th RMSECV i' Represents vector RMSECV 0 and RMSECV i' The value of the kth row in ; and The mean value is denoted as M i',in and M i',out , subtract the two means to get DM i' ; If DM i' <0, the characteristic band is a strong information characteristic band or a weak information characteristic band; if DM i' >0, the characteristic band is a non-information characteristic band or an interference characteristic band; P = 0.05 is defined as the threshold for the Mann-Whitney U test, and the characteristic bands are finally divided into 4 categories;

[0156] 3) In each iteration, strong information feature bands and weak information feature bands are retained, and no information feature bands and interference feature bands are eliminated; return to step 1) for the next round of iteration until only strong information feature bands and weak information feature bands remain;

[0157] 4) Perform reverse elimination on the tt retained feature bands; First, establish an XGBoost model for the tt feature bands to obtain RMSECV t ; Then, by eliminating the jth feature band, the XGBoost model is established for the tt-1 feature bands to obtain RMSECV -j ,j=1,2,…,tt;if RMSECV -j Less than RMSECV t Then the jth characteristic band is eliminated, otherwise it is retained; this process is repeated, and the remaining characteristic bands are the final selected characteristic bands D;

[0158] This embodiment also uses three characteristic band selection algorithms, SFLA, PSO and SHO, to compare with the ISHO algorithm. Since the characteristic band screening results of the above four algorithms are slightly different each time, they are random and may select some non-information variables or interference variables, resulting in low reproducibility of the algorithm results. Therefore, each algorithm is run independently 100 times, and then the probability of each band being selected during each operation is calculated, and the band with a selection probability exceeding the average value is taken as the characteristic band. First, the ISHO algorithm is used to screen the characteristic band, with a control factor of 0.5 and 1000 iterations. Figure 7 (a) is the probability of each band being selected after optimization using the ISHO algorithm. Figure 7 (b) is the distribution diagram of characteristic bands selected by ISHO algorithm. Figure 7 (a) It can be seen that the average probability of all bands being selected is 0.06. Bands with a probability greater than 0.06 are selected as the final feature bands, and 156 bands that meet the conditions are obtained, accounting for 7.79% of the total bands. Compared with the SHO algorithm, the number of feature bands selected by the proposed ISHO algorithm is significantly reduced. The screening results of the ISHO algorithm are better, the selection probability of most spectral bands is small, and the selection probability of a small number of spectral bands is large. Figure 7 (b) It can be seen that the characteristic bands are mainly distributed between 401nm~585nm and 683nm~920nm, and around 1220nm, 1441nm and 2161nm. The larger the selection probability, the more important the corresponding band is. The selection probabilities at 439nm, 490nm, 819nm, 854nm, 872nm, 876nm and 2168nm are all close to 1. This shows that the differences in the results of multiple runs of the ISHO algorithm are small, and the feature selection algorithm is relatively stable, and can more accurately select the characteristic bands related to modeling.

[0159] Through SPSS collinearity diagnosis analysis, it was found that there were collinearity problems among the bands selected by SFLA, PSO, SHO and ISHO algorithms. Therefore, the IRIV algorithm was used to perform secondary screening on the bands in order to reduce the collinearity between the bands, increase the probability of selecting the effective information bands related to SOM content, and improve the accuracy of the prediction model. The IRIV algorithm was used to perform secondary screening on the bands selected by ISHO, and a total of 7 rounds of iterations were performed. The results are as follows: Figure 8 As shown. Figure 8 It can be seen that before the fourth iteration, the number of retained variables dropped rapidly from 156 to 24. After the sixth iteration, the number of variables remained unchanged, and the useless information with SOM content was eliminated after reverse elimination. Finally, 16 characteristic bands were selected, accounting for 0.79% of the total bands. This shows that the ISHO-IRIV algorithm has a good dimensionality reduction effect on hyperspectral data.

[0160] Step 5: Establish a hyperspectral prediction model to predict SOM content and evaluate the model accuracy;

[0161] The characteristic bands selected by the characteristic band selection algorithm were set as independent variables, the SOM content was set as the dependent variable, and the extreme gradient boosting XGBoost method was used to construct a hyperspectral prediction model for SOM content.

[0162] XGBoost is an ensemble learning algorithm that improves the boosting algorithm based on the gradient boosting decision tree (GBDT). The basic idea is: first build multiple CART (classification and regression trees) tree models to predict the data set, then integrate these trees into a new tree model, and through continuous iterative improvement, the new tree model generated in each iteration will fit the residual of the previous tree until the best training effect is achieved. Set the learning rate of the model to 0.3 and the maximum depth of the tree to 6.

[0163] The characteristic band combination D={(x v ,y v ):v=1,…,s,x v ∈R e ,y v ∈R}, that is, s soil samples with characteristic dimension e, x v represents the feature vector of the vth sample, whose dimension is e, and it contains all the feature information related to the vth sample; y v Represents the true value of the SOM content of the vth sample in the data set;

[0164] definition The predicted value of SOM content of soil samples is:

[0165]

[0166] Among them, f t is the model function of the regression tree, F is the function space of the regression tree, which defines the form and parameter constraints of the regression tree; the goal of XGBoost is to iteratively select the optimal tree f from F t To minimize the objective function, thereby improving the prediction accuracy of the model, f t (x v ) represents the prediction score of the t-th tree for the v-th sample, f t By minimizing the loss function:

[0167]

[0168] Among them, Obj is the objective function, l represents the loss function, which is used to measure the gap between the predicted value and the true value; in order to avoid overfitting, Ω(f k ) is a regularization term that defines the complexity of the tree:

[0169]

[0170] Where K represents f t The number of leaf nodes, ww represents the weight of each leaf, γ and λ are regularization parameters;

[0171] The training of the superposition model adopts an iterative method, and the iterative form of the objective function is:

[0172]

[0173] In the formula, is the predicted value of the vth sample at the k-1th iteration, and constant is a constant term;

[0174] Perform a second-order Taylor expansion on the above formula, and the objective function is rewritten as:

[0175]

[0176] The candidate feature set of each tree split node is defined as I j ={v|q(x v )=j}, define G j is the sum of the gradients of all samples on leaf node j; H j is the sum of the second-order derivatives of all samples on leaf node j, g i is the first-order derivative (gradient) of the loss function, h vis the second-order derivative of the loss function, then the final objective function is expressed as:

[0177]

[0178] For a tree with a fixed structure, the weight ww of leaf j is j * And the optimal solution of the objective function Obj * for:

[0179]

[0180] Since it is impossible to enumerate all tree structures, a greedy algorithm is introduced in XGBoost to split the existing leaf nodes in an iterative manner and calculate the maximum gain obtained. Let I L and I R are the sample sets of the left and right nodes after the node split, respectively, I=I L ∪I R , then the gain obtained is defined as:

[0181]

[0182] If Gain is greater than 0, it means that the loss after the split is reduced, that is, the split can be made, otherwise the split will be stopped.

[0183] In order to evaluate the ability of different models to estimate SOM content, this example divided 70 samples into a training set and 29 into a validation set based on the joint XY distance method. 2 ), root mean square error (RMSE) and relative percent deviation (RPD) were used to analyze and evaluate the results. 2 It is used to evaluate the degree of fit of the model. The value range is between 0 and 1. The closer it is to 1, the better the degree of fit. RMSE is used to measure the deviation between the SOM predicted value and the measured value. The smaller its value, the better the prediction effect. RPD is used to describe the stability of the model. It is generally believed that when RPD<1.4, the model is unstable; when 1.4≤RPD≤2.0, the model is relatively stable; when RPD>2.0, the model is very stable. The calculation method of the evaluation index can be expressed as:

[0184]

[0185] In the formula, y i is the true value of the i-th sample, and are the mean value and model prediction value of the i-th sample respectively.

[0186] In this embodiment, the XGBoost model establishment effects based on the full-spectral dataset, SFLA dataset, PSO dataset, SHO dataset, ISHO dataset, SFLA-IRIV dataset, PSO-IRIV dataset, SHO-IRIV dataset, and ISHO-IRIV dataset are shown in Table 2.

[0187]

[0188] The ISHO-IRIV-XGBoost model has the best prediction effect, and the training set R 2 The validation set R 2 The prediction accuracy of the model is significantly improved compared with the Full-spectral-XGBoost model. The R 2 The model fits well and has strong generalization ability. The data points of the ISHO-IRIV-XGBoost model training set and validation set are basically concentrated near the 1:1 line, and the fitting effect is good. The 95% confidence intervals of the training set and validation set are narrow and overlap with each other, indicating that the prediction results of the model have high certainty and good consistency on the two data sets. The prediction results of the XGBoost model based on the ISHO-IRIV data set are shown in Figure 2. Fig. 9 shown.

[0189] In summary, the XGBoost model constructed based on the characteristic bands screened by the ISHO-IRIV algorithm can be used as the optimal estimation model for SOM content.

[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for predicting soil organic matter content based on hyperspectral, characterized in that: The following steps are involved: Step 1: Collect and process soil samples to determine SOM content and soil spectral reflectance; Step 2: Determine the soil spectral reflectance data; use a ground object spectrometer to collect the spectral reflectance data of the soil sample in a dark room; Step 3: Preprocess the soil spectral reflectance data to eliminate spectral noise and redundant information; Step 4: Use the band selection algorithm to screen the characteristic bands of the preprocessed spectral reflectance data; first, use the improved spotted hyena optimized ISHO algorithm to preliminarily screen the characteristic bands, and then use the iterative information retention variable IRIV algorithm to perform secondary screening on the characteristic bands; The ISHO algorithm uses Tent chaos mapping and binary discretization to optimize the initial individual position of spotted hyenas, and uses nonlinear decreasing inertia weight and dynamic Laplace crossover strategy to optimize the position update process of spotted hyenas; Step 5: Establish a hyperspectral prediction model to predict SOM content and evaluate the model accuracy; The characteristic bands selected by the characteristic band selection algorithm were set as independent variables, and the SOM content was set as the dependent variable. The extreme gradient boosting XGBoost method was used to construct a hyperspectral prediction model to predict the SOM content.

2. The method for predicting soil organic matter content based on hyperspectral according to claim 1, characterized in that: The step 1 divides the processed soil sample into two parts, which are used for soil spectral reflectance data measurement and SOM content determination respectively; The SOM content was determined by potassium dichromate external heating method.

3. The method for predicting soil organic matter content based on hyperspectral according to claim 2, characterized in that: The wavelength range of the spectral reflectance data in step 2 is 350 to 2500 nm, and the sampling interval is 1 nm.

4. The method for predicting soil organic matter content based on hyperspectral according to claim 1, characterized in that: The step 3 comprises: Firstly, the Savitzky-Golay smoothing method was used to denoise the soil spectral reflectance data to obtain the original spectral reflectance data R. In order to further enhance the correlation between the soil spectral reflectance data and the SOM content, the original spectral reflectance data R was further corrected for multivariate scattering and then preprocessed with the first-order differential.

5. The method for predicting soil organic matter content based on hyperspectral according to claim 1, characterized in that: The process of using the improved spotted hyena optimized ISHO algorithm to preliminarily screen the characteristic bands in step 4 is as follows: (1) Set the parameters of the ISHO algorithm; set the number of spotted hyenas SHN, the spectral reflectance data matrix I, and the maximum number of iterations MM Itn ; where the spectral reflectance data matrix I contains s samples and m bands, represents a set of m bands; construct an m-dimensional vector space, then set q spotted hyena individuals to form a population, q < m, search for prey in the m-dimensional space, and use the fitness function as a criterion to continuously update the positions of the spotted hyenas until the maximum number of iterations is reached. Finally, select the characteristic band combination from the band set 1 ≤ i ≤ m;​ (2) The spotted hyena individuals are set to represent a set of selected characteristic bands, and their positions are represented by a vector equal to the total number of bands. The spotted hyena individuals with the best fitness value in the current search space are regarded as target prey, and the remaining spotted hyenas update their positions according to the target prey positions. The calculation process is as follows: Where: is the distance between the prey and the spotted hyena; x is the current iteration number; is the current position vector of the spotted hyena, y n Represents the spectral reflectance data corresponding to the nth band; is the current prey position vector; and They are the swing factor and the convergence factor, as shown in the following formula: Where: and is a random vector in the interval [0, 1]; is the control factor, Linearly decreases from 5 to 0; Itn=1,2,…,MM Itn ; Use Tent chaos map to initialize the spotted hyena population as shown in the following formula: In the formula, y i ' is the chaotic variable, i.e. the new spectral reflectance value; After Bernoulli shift transformation, we can get: <h2 style=";text-align:left;direction:ltr">y<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> '=(2y<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> mod1 (7) (3) Binary discretization of spotted hyena population individuals is performed as shown in the following formula: Where: It is the result of binary discretization; is the result after chaos initialization; Thd is the binary discretization threshold, set Thd = 0.5; i is the index of the spotted hyena individual; j is the dimension index of the spotted hyena individual; (4) Design the fitness function shown in the following formula to calculate the fitness value of each spotted hyena individual: Where: F is the fitness function; α is the performance weight parameter of the regression model; μ is the weight parameter for selecting the number of feature bands; α∈[0,1], μ=1-α; t is the number of selected feature bands; T is the total number of bands; J is the determination coefficient; (5) The position of the individual spotted hyena with the smallest fitness value is taken as the prey position (6) Determine whether the current number of iterations x has reached the maximum number of iterations. If so, output the prey position. Output 1 represents selection, 0 represents unselected, and the selected feature band combination is obtained Otherwise, continue to step (7); (7) Determine whether the currently processed spotted hyena index i reaches the spotted hyena number SHN. If so, increase the current iteration number x by one and re-execute step (5); otherwise, continue to execute step (8); (8) Calculate the nonlinear decreasing inertia weight under the current iteration number x. The inertia weight w is expressed as: Where: w st represents the weight at the beginning of the iteration, w st =0.9; w ed Represents the weight at the end of the iteration, w ed =0.4; (9) Update the individual position of the spotted hyena according to the inertia weight w, as shown in the following formula: Where: It is the best location for spotted hyenas; is the location of the other spotted hyenas; is a set of SHN optimal solutions; is the updated optimal solution position, and the positions of other spotted hyenas are updated according to this position; (10)Judgment Is it true? r is a random number between [0,1]. If it is true, the dynamic Laplace crossover strategy is used to update the position of the spotted hyena individual again, as shown in the following formula: Where: and are the positions of the offspring individuals produced after the Laplace operator crossover; and are the positions of the two parent individuals with the best fitness in the solution space; β is a random number generated in reverse according to the Laplace distribution function, as shown in the following formula: Where: c∈R is the location parameter; u is a random number in the interval [0,1] that follows a uniform distribution; If it is not true, substitute the β value set in formula (17) into formula (14) (15) to update the position of the individual spotted hyena; (11) Calculate the fitness value of the spotted hyena individual after the position is updated, compare the fitness value of the current best spotted hyena individual with the fitness value of the prey, if the fitness value of the current best spotted hyena individual is less than the fitness value of the prey, take the current best spotted hyena individual position as the new prey position, and update the position of the best spotted hyena Otherwise keep the prey position and the best spotted hyena position constant; (12) The spotted hyena index i is increased by 1, and step (7) is executed again.

6. The method for predicting soil organic matter content based on hyperspectral according to claim 5, characterized in that: The process of performing secondary screening of the characteristic bands using the iterative information retention variable IRIV algorithm in step 4 is as follows: remember is the characteristic band set after ISHO algorithm screening, s is the number of soil samples, the spectral reflectance data matrix is ​​m' columns, and the soil samples are randomly divided into training set and validation set; 1) For the matrix with s rows and m' columns in the training set, it contains only 1 and 0 1 and 0 respectively indicate whether the variable is used for modeling. The XGBoost model is established for each row of , and the RMSECV value obtained by 5-fold crossover is used as the evaluation standard, and the vector of size s×1 is recorded as RMSECV0; the matrix In the i'th column (), replace 1 with 0 and 0 with 1 to obtain the matrix B, i'=1,2,…,m'. Similarly, build an XGBoost model in each row of the matrix B to obtain a vector of size s×1, denoted as RMSECV i' ; 2) Definition and To evaluate the importance of each feature band, the formula is as follows: Where: k th represents the kth row in the vector, then k th RMSECV0 and k th RMSECV i' Represents vectors RMSECV0 and RMSECV respectively i' The value of the kth row in ; and The mean value is denoted as M i',in and M i',out , subtract the two means to get DM i' ; If DM i' <0, the characteristic band is a strong information characteristic band or a weak information characteristic band; if DM i' >0, the characteristic band is a non-information characteristic band or an interference characteristic band; P = 0.05 is defined as the threshold for the Mann-Whitney U test, and the characteristic bands are finally divided into 4 categories; 3) In each iteration, strong information feature bands and weak information feature bands are retained, and no information feature bands and interference feature bands are eliminated; return to step 1) for the next round of iteration until only strong information feature bands and weak information feature bands remain; 4) Perform reverse elimination on the tt retained feature bands; First, establish an XGBoost model for the tt feature bands to obtain RMSECV t ; Then, by eliminating the jth feature band, the XGBoost model is established for the tt-1 feature bands to obtain RMSECV -j ,j=1,2,…,tt;if RMSECV -j Less than RMSECV t Then the jth characteristic band is eliminated, otherwise it is retained; this process is repeated, and the remaining characteristic bands are the final selected characteristic bands D.

7. The method for predicting soil organic matter content based on hyperspectral according to claim 6, characterized in that: The step 5 comprises: Set the characteristic band combination D selected by ISHO-IRIV = {(x v ,y v ):v=1,…,s,x v ∈R e ,y v ∈R}, that is, s soil samples with characteristic dimension e, x v represents the feature vector of the vth soil sample, whose dimension is e, and it contains all the feature information related to the vth soil sample; v represents the true value of the SOM content of the vth soil sample in the dataset; definition The predicted value of SOM content of soil samples is: Among them, f t is the model function of the regression tree, F is the function space of the regression tree, which defines the form and parameter constraints of the regression tree; the goal of XGBoost is to iteratively select the optimal tree f from F t To minimize the objective function, thereby improving the prediction accuracy of the model, f t (x v ) represents the prediction score of the t-th tree for the v-th soil sample, f t By minimizing the loss function: Among them, Obj is the objective function, l represents the loss function, which is used to measure the gap between the predicted value and the true value; Ω(f k ) is a regularization term that defines the complexity of the tree: Where K represents f t The number of leaf nodes, ww represents the weight of each leaf, γ and λ are regularization parameters; The training of the superposition model adopts an iterative method, and the iterative form of the objective function is: In the formula, is the predicted value of the vth sample at the k-1th iteration, and constant is a constant term; Perform a second-order Taylor expansion on the above formula, and the objective function is rewritten as: The candidate feature set of each tree split node is defined as I j ={v|q(x v )=j}, define G j is the sum of the gradients of all samples on leaf node j; H j is the sum of the second-order derivatives of all samples on leaf node j, g v is the first-order derivative of the loss function, h v is the second-order derivative of the loss function, then the final objective function is expressed as: For a tree with a fixed structure, the weight ww of leaf j is j * And the optimal solution of the objective function Obj * for: Since it is impossible to enumerate all tree structures, a greedy algorithm is introduced in XGBoost to split the existing leaf nodes in an iterative manner and calculate the maximum gain obtained. Let I L and I R are the sample sets of the left and right nodes after the node split, respectively, I=I L ∪I R , then the gain Gain is defined as: If Gain is greater than 0, it means that the loss after the split is reduced, that is, the split can be made, otherwise the split will be stopped.