Estimation Method of Vegetation Equivalent Water Thickness Based on Hyperspectral and Machine Learning
Through hyperspectral and machine learning methods, drones are used to obtain hyperspectral data, analyze spectral curves, and build machine learning models, which solves the accuracy and applicability of EWT estimation in complex terrain areas, and achieves more efficient vegetation moisture monitoring.
Patent Information
- Application Number
- CN202211482815.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-11-24
AI Technical Summary
In the area of terrain and complex terrain, the prior art is difficult to meet the practical application requirements of vegetation equivalent water thickness (EWT) estimation at local scale in terrain and terrain complex.
Using a method based on hyperspectral and machine learning, hyperspectral remote sensing data is obtained through drones, and spectral curves are analyzed by combining first-order derivatives, second-order derivatives and envelope removal methods, sensitive bands are identified, feature indexes are calculated, XGBoost and random forest machine learning models are constructed, and EWT estimation is performed.
It improves the accuracy and applicability of EWT estimation, is suitable for complex terrain areas, reduces data saturation, and improves the prediction ability of the model.
Smart Images

Figure CN115775357B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ecological environment research, and particularly to a method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning. Background Art
[0002] The change of leaf water content changes the nitrogen and carbon utilization rates of vegetation, affects the balance of vegetation carbon cycle and energy budget, and causes changes in leaf and canopy spectral reflectance. Therefore, the influence of vegetation water on the estimation accuracy of the model needs to be fully considered in the evaluation of ecological environment quality. As an important parameter characterizing the water content distribution of vegetation leaves, the equivalent water thickness has been widely used to characterize vegetation functions and ecosystem processes. Therefore, the rapid and accurate estimation of EWT has important guiding significance for the health of vegetation and the protection of the ecological environment.
[0003] Early EWT estimation mainly relied on field observations. This method has high accuracy, but has disadvantages such as strong destructiveness, high cost, and poor timeliness. The application of remote sensing technology makes up for the deficiencies of field observation methods and has advantages such as a wide monitoring range and fast data update. It is a large-scale, non-contact, dynamic observation method that can obtain information such as spectral radiation intensity, area change, and texture in vegetation-covered areas, providing an effective information source for extracting vegetation biochemical parameters. Among them, hyperspectral remote sensing technology has been widely used in vegetation monitoring due to its characteristics such as high efficiency, non-destructiveness, and multi-scale. For example, Ma Yanchuan et al. established a rapid and non-destructive monitoring and estimation model for cotton canopy EWT based on hyperspectral remote sensing data and achieved good results. However, when using hyperspectral data for vegetation water monitoring evaluation, due to limitations such as the large number of spectral bands and large amount of data in hyperspectral data, the data usability is reduced and the phenomenon of pixel oversaturation is likely to occur. Therefore, in the study of vegetation EWT, the limitations brought by data attributes need to be fully considered.
[0004] The EWT estimation method based on remote sensing data has made great progress. Commonly used methods include radiative transfer model inversion method, vegetation index method, etc. Among them, Li et al. proposed 3 new spectral absorption indices, which are suitable for the EWT estimation of various plant types; Pan Qingmei et al. analyzed the quantitative relationship between spectral indices of hyperspectral data and leaf water status, screened out the spectral indices most suitable for estimating the water content of walnut leaves, constructed an estimation model for the water content of walnut leaves, and realized the rapid and accurate monitoring of the water content of walnut leaves. However, in areas with complex terrain and topography, affected by factors such as landform, climate, and remote sensing model parameters, the previous estimation methods are difficult to meet the actual application needs of more regions. The application of hyperspectral technology effectively reduces the uncertainty of vegetation water estimation based on traditional remote sensing technology, but hyperspectral technology belongs to a near-ground observation method, and traditional radiative transfer models are not suitable for near-ground spectroscopy. Summary of the Invention
[0005] 1. Technical problems to be solved
[0006] The object of the present invention is to solve the problem that in complex terrain and topography areas, due to factors such as landform, climate, and remote sensing model parameters, existing inversion and estimation methods are difficult to meet the actual application requirements at the local scale, and a method for improving the accuracy of the vegetation equivalent water thickness estimation model based on different hyperspectral data resolutions and machine learning algorithms is proposed.
[0007] 2. Technical solutions
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A method for estimating the vegetation equivalent water thickness based on hyperspectral and machine learning, comprising the following steps:
[0010] Step 1: Acquisition of field observation data. First, in the study area, 151 unobstructed trees that can be directly photographed above the canopy are randomly selected for sampling, and 3 fresh leaves are collected from each tree. Secondly, an unmanned aerial vehicle is used to conduct an aerial photograph of the study area to obtain image data containing multiple spectral channels, and the data is preprocessed to obtain the unmanned aerial vehicle hyperspectral reflectance data of the study area;
[0011] Step 2: Calculation of the measured equivalent water thickness. First, when collecting fresh leaves, the surface area and fresh weight data of the leaves are statistically analyzed. Secondly, the obtained fresh leaves are dried in an indoor oven until a constant weight is reached and then weighed again to obtain the dry weight of the leaves. Finally, the measured equivalent water thickness value is calculated from the three parameters of the obtained fresh weight, dry weight, and leaf area. The equivalent water thickness calculation formula is:
[0012]
[0013] In the formula, EWT is the equivalent water thickness of the leaf, with the unit of g cm -2 ; FW is the fresh weight of the leaf, with the unit of g; DW is the dry weight of the leaf, with the unit of g; A is the leaf area of the leaf, with the unit of cm -2 ;
[0014] Step 3: Selection of sensitive bands. The equivalent water thickness in the field measurement data is divided into multiple range intervals, and then various methods are used to perform spectral feature analysis on the spectral curve to highlight the spectral features between different levels of equivalent water thickness and the spectral curve of the hyperspectral data, and obtain the bands sensitive to the change of the vegetation equivalent water thickness.
[0015] Step 4: Calculation of characteristic indices. According to the existing empirical formula, the empirical index is calculated, and at the same time, combined with the selected sensitive bands, the vegetation characteristic index based on the sensitive bands is calculated;
[0016] Step 5: Build a machine learning model. First, in this study, the feature index data and the original bands obtained above were resampled to 0.5 m, 1 m, and 2 m respectively, and the above data were extracted using the coordinate information of the measured data to establish a training dataset. Second, according to the variable importance ranking results of the XGBoost model, the input variables of the model were selected. At the same time, combining two machine learning algorithms, RF and XGBoost, six machine learning models for estimating EWT were constructed, and the ten-fold cross-validation method was used to verify the accuracy of all models. Finally, the differences and causes of the EWT estimation models with different resolutions and different models were compared and analyzed.
[0017] Preferably, in step 1, the drone is a DJI Matrice 600 multi-rotor drone, and a Nano-Hyperspec micro-hyperspectral imager is carried on the drone.
[0018] Preferably, the data preprocessing in step 1 includes reflectance calibration, radiometric calibration, atmospheric correction, and mosaicking.
[0019] Preferably, in step 2, the unobstructed trees that can be directly photographed above the canopy are used as the sampling criterion. During the field sampling process, the leaves are weighed on-site, and the leaf area is obtained by scanning with a Wanshen plant image analyzer. After the fresh leaves are dried to a constant weight indoors, the high-precision measured equivalent water thickness is calculated according to the empirical formula.
[0020] Preferably, the measured equivalent water thickness in step 3 is divided into four ranges: 0.001 - 0.01, 0.01 - 0.015, 0.015 - 0.02, and 0.02 g cm -2 And in ENVI 5.3, the first derivative, second derivative, and envelope removal methods are used to highlight the spectral characteristics between the spectral curves of different grades of equivalent water thickness and the hyperspectral data, and eliminate the influence of similar spectral curve situations.
[0021] Preferably, in step 4, 37 vegetation feature indices are selected, and 8 texture features including mean, variance, correlation, co-occurrence, dissimilarity, contrast, information entropy, and second moment are also selected as the input variable parameters of the model to avoid the possible saturation phenomenon of vegetation indices; the gray level of the texture features is set to 32, the window size is 5×5, and the direction is (0, 1). The indices used in the present invention mainly include: empirical indices, texture indices, and sensitive band indices, a total of 45 feature indices.
[0022] Preferably, the verification method in step 4 is to divide the data into 10 parts, each part having almost the same number of samples, leaving one part as the test data, and taking the other 9 parts in turn as the training set. A total of ten models are established, and the average prediction accuracy is taken as the final accuracy. The coefficient of determination R 2(Equation 2) and root mean square error RMSE (Equation 3) are used as the accuracy evaluation criteria for models with different resolutions;
[0023]
[0024]
[0025] Wherein, is the predicted value; is the mean value; y i is the measured value; n is the number of samples.
[0026] 3. Beneficial effects
[0027] Compared with the prior art, the advantages of the present invention are as follows:
[0028] (1) In the present invention, the key vegetation parameters (EWT) of the sample points are calculated based on the field survey data. At the same time, the spectral features of the spectral curve are analyzed by using the first derivative, second derivative, envelope removal method, etc., the sensitive bands are identified, and a variety of vegetation indices and texture indices are calculated, and then the environmental variable raster data with different resolutions are obtained. On this basis, data with different spatial resolutions are extracted to obtain a training data set containing environmental variable data such as original spectral information, vegetation indices, and texture indices, and then a machine learning model construction of data variables with different spatial resolutions is realized.
[0029] (2) In the present invention, the spectral map features of hyperspectral remote sensing data with different spatial resolutions are fully exploited, a multi-spatial resolution feature index suitable for the study area is constructed, and then the differences between models with different spatial resolutions are compared and analyzed, which proves that the conversion of spatial resolution has an important impact on the model estimation accuracy and has important reference significance for future similar research work. Description of the drawings
[0030] Figure 1 is the location map of the study area in Embodiment 1 of the present invention;
[0031] Figure 2 is the density distribution curve of the equivalent water thickness data observed on site in Embodiment 2 of the present invention;
[0032] Figure 3 is the spectral feature selection diagram in Embodiment 2 of the present invention. Specific embodiments
[0033] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0034] Embodiment 1:
[0035] Reference Figure 1 Figure 1 , the study area (103°37′29″-103°38′36″E, 31°1′25″-31°2′39″N) is located in the north of Dujiangyan City, Chengdu, on the east side of the Minjiang River estuary, and in the southeast of the Longmen Mountains in western Sichuan. The total area is about 1.88 km2. The terrain in the area is characterized by high in the southwest and low in the northeast, mainly mountainous and hilly, with an altitude of 812-1157 m and a relative height difference of 345 m. The average annual temperature is 15.2 °C, the frost-free period is 247-269 days, the rainfall is 528.7-1332.2 mm, and the sunshine is 1693.9-1042.2 h.
[0036] The area mainly includes various land use types such as natural forest land, artificial forest, cash crops, artificially planted landscapes, residential areas, roads, etc. The land types are complex and diverse and are significantly affected by human disturbances. Among them: the species in natural forest land include Acer pictum Thunb., Pinus, etc.; the species in artificial forest include Cinnamomum camphora (L.) Presl, bamboo, Ginkgo biloba L., etc.; the species in cash crops include Actinidia chinensis Planch., Camellia sinensis (L.) O.Ktze., etc.; the species in artificially planted landscapes include Osmanthus fragrans (Thunb.) Lour., Pinus Linn, etc.
[0037] The method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning includes the following steps:
[0038] Step 1: Acquisition of hyperspectral data. The hyperspectral data used is mainly obtained by observing with a Nano-Hyperspec micro-hyperspectral imager carried by a DJI Matrice 600 multi-rotor UAV. The UAV aerial photography lasted for 3 days (July 1, 2021 - July 3, 2021), with a total of 6 flights, and the true altitude was 300 m (including). Image data containing 270 spectral channels was obtained. After preprocessing such as reflectance calibration, radiometric calibration, atmospheric correction, and mosaicking of this data, the UAV hyperspectral reflectance data of the study area was obtained, as shown in Table 1:
[0039] Table 1 Original band information
[0040]
[0041]
[0042] Step 2: Calculation of measured equivalent water thickness. The field observation data mainly take the unobstructed trees that can be directly photographed above the canopy as the sampling criterion. A total of 151 sample trees are selected, and 440 fresh leaves are collected. During the field sampling process, the fresh weight of the leaves is weighed on-site, and the leaf area is obtained by scanning with a Wanshen plant image analyzer. The obtained fresh leaves are dried in an indoor oven until they reach a constant weight and then weighed again to obtain the dry weight of the leaves. The equivalent water thickness index is calculated from the three parameters of the obtained fresh weight, dry weight, and leaf area. The formula for calculating the equivalent water thickness is:
[0043]
[0044] In the formula, EWT is the equivalent water thickness of the leaf, with the unit of g cm -2 ; FW is the fresh weight of the leaf, with the unit of g; DW is the dry weight of the leaf, with the unit of g; A is the leaf area of the leaf, with the unit of cm -2 ;
[0045] Step 3: Obtaining sensitive bands. The equivalent water thickness in the field measurement data is divided into four range intervals: 0.001 - 0.01, 0.01 - 0.015, 0.015 - 0.02, and 0.02 g cm -2 above. Then, various methods are used to analyze the spectral characteristics of the spectral curve, highlighting the spectral characteristics between different levels of equivalent water thickness and the spectral curve of hyperspectral data, and obtaining the bands sensitive to the change of vegetation equivalent water thickness. This part of the operation is implemented in ENVI5.3;
[0046] Step 4: Calculation of characteristic indices. The vegetation characteristic indices are calculated using existing index equations. At the same time, according to the empirical formula, the corresponding vegetation indices are calculated. In this study, eight texture features including mean, var, cor, hom, dis, con, ent, and sec (see Table 3) are also selected as the input variable parameters of the model to avoid the possible saturation phenomenon of vegetation indices. The gray level setting for texture feature selection is 32, the window size is 5×5, and the direction is (0, 1). The indices used in this paper mainly include: empirical indices, texture indices, and sensitive band indices, totaling 45 characteristic indices. See Tables 2 - 3:
[0047] Table 2 Vegetation Indices
[0048]
[0049]
[0050]
[0051]
[0052] Note: In the table, R i represents the reflectance at i nm; the letter + number refers to the vegetation index calculated directly using the existing formula in the original text, that is, the empirical index; the letter + _ + number refers to the index calculated from the sensitive bands, that is, the sensitive band index; the calculation formula of the vegetation index used mainly comes from the research of C.T. de Almeida et al.
[0053] Table 3 Texture Index
[0054]
[0055]
[0056] Step 5: Construction of the machine learning model. First, in this study, the feature indices calculated above and the original hyperspectral bands were resampled to 0.5 m, 1 m, and 2 m respectively, and the coordinate information of the measured data was used to extract the above data to establish a training dataset. Second, according to the variable importance ranking results of the XGBoost model, the input variables of the model were selected, and at the same time, two machine learning algorithms, RF and XGBoost, were combined to construct 6 machine learning models for estimating EWT. Finally, the ten-fold cross-validation method was used to verify the accuracy of all models.
[0057] The verification method is to divide the data into 10 parts, each part having almost the same number of samples, leaving one part as the test data, and taking the other 9 parts in turn as the training set. A total of ten models are established, and the average prediction accuracy is taken as the final accuracy. The coefficient of determination R 2 (Equation 2) and the root mean square error RMSE (Equation 3) are used as the model accuracy evaluation criteria;
[0058]
[0059]
[0060] In the formula, is the predicted value; is the mean value; y i is the measured value; n is the number of samples.
[0061] In the present invention, in this study, the key vegetation parameter (EWT) of the sample points was calculated based on the field survey data. At the same time, the spectral information of the hyperspectral data was extracted according to the geographical coordinates of the measured sample plots. Furthermore, methods such as first derivative, second derivative, and envelope removal were used to perform spectral feature analysis on the spectral curve, identify the sensitive bands, calculate the feature indices, obtain environmental variable data with different resolutions, and realize the construction of a machine learning model for data variables with different spatial resolutions.
[0062] In the present invention, the spectral features of hyperspectral remote sensing data with different spatial resolutions are fully exploited, and a machine learning estimation model with multiple spatial resolutions suitable for the study area is constructed, which has important reference significance for future similar research work.
[0063] Example 2:
[0064] Results and analysis mainly include:
[0065] 1. Measured equivalent water thickness: By statistically analyzing the observation data, the average value of EWT in the study area is 0.014 g / cm -2 ; the maximum value of EWT is 0.027 g / cm -2 ; the minimum value of EWT is 0.002 g / cm -2 . The results of the EWT density curve show that the measured data presents a nearly normal distribution characteristic (see Figure 2 ), and the distribution interval is relatively large, meeting the modeling requirements.
[0066] 2. Sensitive bands: Multiple characteristic indices such as the vegetation enhanced index, normalized vegetation index, and vegetation water index are calculated. After dividing the measured data into four categories, the original bands of the data can better distinguish the characteristics of different ranges of EWT values at 770 nm and 904 nm (see Figure 3 ). After removing the envelope line, obvious differences are shown at the peak at 533 nm, and the valleys at 504 nm and 680 nm in the spectrum. Through spectral analysis of the original spectrum and the envelope line removal, finally, 504 nm, 533 nm, 680 nm, 770 nm, and 904 nm, namely B48, B61, B127, B167, and B227, are selected as the sensitive bands that can highlight the spectral characteristics of different EWT vegetation.
[0067] 3. Characteristic variables: Based on the XGBoost algorithm, the input variables (original bands, vegetation indices, texture characteristic indices) of data with different spatial resolutions are sorted by importance to determine the optimal variables participating in the modeling (see Table 4). To highlight the differences between different models, the two models with the same resolution input the same variables. Among them, when the characteristic indices with a spatial resolution of 0.5 m are used as input variables, the various indices constructed by the sensitive bands do not contribute much to the model; the 1 m spatial resolution model shows that the VOG_1 index constructed by the sensitive bands (770 nm, 904 nm) participates in the modeling; among the input variables with a 2 m spatial resolution, the REP_1 and Cor(B167) indices constructed by the sensitive bands contribute significantly to the model. It can be seen that as the spatial resolution decreases, the contribution of the indices constructed by the near-infrared 770 nm and 904 nm bands to the model training increases.
[0068] 4. Model accuracy: According to the training results of the RF and XGBoost models (see Table 4), for RF (R 2 = 0.49, RMSE = 0.0032 g cm -2 ), and for XGBoost (R 2 = 0.41, RMSE = 0.0034 g cm -2 ), both have the highest accuracy at a spatial resolution of 2 m. It can be seen that the model accuracy increases as the spatial resolution decreases, and when the spatial resolution reaches 2 m, the simulation accuracy reaches the highest. By comparing the accuracies of the two models at the same resolution, it is found that RF has a higher accuracy than XGBoost. In addition, at spatial resolutions of 1 m and 2 m, the REP_1, Cor(B167), and VOG_1 indices constructed from sensitive bands significantly improve the model accuracy. As the resolution of the input data for the machine learning model decreases, the number of feature indices involved in modeling increases.
[0069] Table 4 Model accuracies with input variables of different spatial resolutions
[0070]
[0071]
[0072] In the present invention, the high-spatial-resolution data obtained by the unmanned aerial vehicle has a "salt-and-pepper effect", making the canopy structures of different tree species complex and the spatial heterogeneity large. Therefore, selecting appropriate data sources is beneficial for estimating the key parameters of vegetation. By comparing and analyzing the model efficiencies of the RF and XGBoost models with input variables of different spatial resolutions in the estimation of vegetation EWT in the Dujiangyan area. The research results show that the spatial resolution of the input variables affects the model accuracy, and the data with a spatial resolution of 2 m obtains the highest model accuracy. In addition, the measured EWT data range is 0.009 - 0.027 g cm -2 , and the EWT estimated by the RF model with a spatial resolution of 2 m is distributed between 0.006 - 0.023 g cm -2 . The two can effectively reflect the changes in the moisture of vegetation leaves. It can be seen that the optimal spatial resolution of the model input variables is 2 meters.
[0073] In the present invention, it is also found that during the estimation process of vegetation EWT, when the spatial resolution of the input variables is reduced from 0.5 m to 2 m, the accuracy of the RF model is improved from 0.21 to 0.49. This indicates that selecting an appropriate spatial resolution plays an important role in estimating EWT. The vegetation canopy in high-spatial-resolution data consists of multiple pixels, increasing the spatial heterogeneity of canopy pixels. In this case, the spectral and spatial texture features of the canopy of the same type of vegetation cannot be characterized by a single pixel. Because hyperspectral remote sensing image data has characteristics such as high dimensionality, high spatial resolution, and information redundancy, the spectral and spatial correlations between a single pixel and its surrounding pixels are relatively large. When the sample size is fixed, the increase in the number of data bands will cause the accuracy of the model to exhibit the Hughes phenomenon of "increasing first and then decreasing".
[0074] In the present invention, the existence of this phenomenon reduces the model efficiency at a spatial resolution of 0.5 m. Analyzing the reasons for its formation, it is mainly because the input variables of this model are the original bands of hyperspectral data, and the bands are relatively close to each other, unable to accurately respond to the surface vegetation moisture. The feature index method based on sensitive bands can effectively achieve dimensionality reduction of hyperspectral data. This process mainly combines the spectral characteristics of ground objects, uses band calculations to highlight the sensitive features related to vegetation moisture, and promotes the improvement of model accuracy. For example, the vegetation feature indices with spatial resolutions of 1 m and 2 m constructed based on the feature index method contribute to the increase in model accuracy.
[0075] In the present invention, it is further found that changing the spatial resolution of the input data can more accurately characterize the vegetation canopy information and the constructed feature index is more effective. The number of pixels within the same vegetation canopy increases with the increase in spatial resolution. Due to the existence of spatial correlations between pixels, there should be a high degree of similarity among the pixels of the vegetation canopy. However, affected by terrain, slope aspect, observation time, and angle, there are obvious spatial heterogeneities among the pixels of the vegetation canopy with different spatial resolutions, and the phenomena of "same object with different spectra" and "same spectra with different objects" are common. The increase in spatial resolution causes an increase in the image noise of the data, making the image spectrum "distorted" and reducing the accuracy of the estimation model. This "distortion" of the data can be suppressed by reducing the resolution. In the observation area, the land use types are mainly forest land, orchards, tea gardens, etc. Although the sizes of the vegetation canopies vary, the vegetation crown widths are similar to the optimal spatial scale of 2 m in this study. Therefore, the spatial resolution of different study areas should be determined according to the crown widths of the local vegetation canopies.
[0076] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. A method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning, characterized in that, It includes the following steps: Step 1: Acquisition of field observation data. First, in the study area, 151 trees that are not blocked above the canopy and can be directly photographed are randomly selected for sampling. Three fresh leaves are collected from each tree. Second, an unmanned aerial vehicle (UAV) is used to conduct an aerial survey of the study area to obtain image data containing multiple spectral channels, and the data is preprocessed to obtain the UAV hyperspectral reflectance data of the study area; Step 2: Calculation of measured equivalent water thickness. First, when collecting fresh leaves, the surface area and fresh weight data of the leaves are statistically analyzed. Second, the obtained fresh leaves are dried in an indoor oven until they reach a constant weight and then weighed again to obtain the dry weight of the leaves. Finally, the measured equivalent water thickness is calculated from the three parameters of the obtained fresh weight, dry weight, and leaf area. The formula for calculating the equivalent water thickness is: Equation 1 In the formula, is the equivalent water thickness of the leaf, with the unit of ; is the fresh weight of the leaf, with the unit of ; is the dry weight of the leaf, with the unit of ; is the leaf area of the leaf, with the unit of ; Step 3: Selection of sensitive bands. The measured equivalent water thickness is divided into multiple range intervals, and the spectral characteristics between different grades of equivalent water thickness and the spectral curves of hyperspectral data are analyzed to obtain the bands that are sensitive to changes in the equivalent water thickness of vegetation; Step 4: Calculation of characteristic indices. According to existing empirical formulas, empirical indices are calculated, and at the same time, combined with the selected sensitive bands, vegetation characteristic indices based on sensitive bands are calculated; Step 5: Machine learning model. First, after resampling the calculated characteristic indices and the original hyperspectral bands to 0.5 m, 1 m, and 2 m respectively, the above data is extracted using the coordinate information of the measured data to establish a training dataset. Second, according to the variable importance ranking results of the XGBoost model, the input variables of the model are selected, and at the same time, combining two machine learning algorithms, RF and XGBoost, six machine learning models for estimating EWT are constructed. Finally, the ten-fold cross-validation method is used to verify the accuracy of all models.
2. The method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning according to claim 1, wherein In Step 1, the UAV is the DJI Matrice 600 multi-rotor UAV, and a Nano-Hyperspec micro-hyperspectral imager is carried on the UAV.
3. The method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning according to claim 1, wherein In Step 1, the data preprocessing includes reflectance calibration, radiometric calibration, atmospheric correction, and mosaicking.
4. The method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning according to claim 1, wherein In Step 2, the sampling criterion is the trees that are not blocked above the canopy and can be directly photographed.
5. The method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning according to claim 1, wherein In Step 2, during the field sampling process, the fresh weight of the leaves is weighed on-site, and the leaf area is obtained by scanning using a Wanshen plant image analyzer; after drying to a constant weight indoors, the dry weight is obtained, and the measured equivalent water thickness with high precision is calculated.
6. The method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning according to claim 1, characterized in that, In step 3, the equivalent water thickness is divided into four ranges: 0.001 - 0.01, 0.01 - 0.015, 0.015 - 0.02, and 0.02 The above four range intervals are used, and in ENVI 5.3, the first derivative, second derivative, and envelope line removal methods are used to further highlight the spectral characteristics between the measured equivalent water thickness and the hyperspectral data, and eliminate the influence of the similar spectral curve situation.
7. The method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning according to claim 1, characterized in that In Step 4, 37 vegetation indices are selected, and 8 texture features, namely mean, variance, correlation, co-occurrence, dissimilarity, contrast, information entropy, and second-order moment, are also selected as model input variable parameters to avoid possible saturation phenomena of vegetation indices; The gray level setting of the selected texture features is 32, the window size is 5×5, and the direction is (0, 1). The empirical indices used include: vegetation indices, texture indices, and sensitive band indices, a total of 45 characteristic indices.
8. The method for estimating the equivalent water thickness of vegetation based on hyperspectral and machine learning according to claim 1, characterized in that In the verification method of step 5, the data is divided into 10 parts, each part having almost the same number of samples. One part is left as the test data, and the other 9 parts are taken in turn as the training set. A total of ten models are established, and the average prediction accuracy is taken as the final accuracy. The coefficient of determination R2 (Equation 2) and the root mean square error RMSE (Equation 3) are used as the model accuracy evaluation criteria; Equation 2 Equation 3 In the formula, is the predicted value; is the mean value; is the measured value; n is the number of samples.