Method and system for estimating above-ground biomass of vegetation

By using drones equipped with hyperspectral and LiDAR sensors combined with a random forest estimation model, the problem of accuracy in estimating aboveground biomass of vegetation in complex environments has been solved, achieving efficient and low-cost vegetation biomass estimation, which is suitable for refined monitoring of complex ecosystems.

CN121789047APending Publication Date: 2026-04-03NORTH CHINA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately estimate aboveground biomass in complex environments, especially in restoration ecosystems with poor soil, uneven nutrient distribution, and strong spatial heterogeneity in vegetation configuration. Traditional methods are time-consuming, costly, and difficult to obtain spatially continuous biomass distribution data.

Method used

By using drones equipped with hyperspectral and LiDAR sensors and combining them with a random forest estimation model, the model is constructed to accurately estimate the aboveground biomass of vegetation by acquiring information such as vegetation location, diameter at breast height (DBH), vegetation height, spectral index, texture features, and canopy structure features.

Benefits of technology

It enables accurate estimation of vegetation aboveground biomass in complex environments, improves the model's robustness and estimation ability in complex environments, simplifies the operation process, reduces costs, and improves efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789047A_ABST
    Figure CN121789047A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vegetation intelligent monitoring, and discloses a vegetation above-ground biomass estimation method and system, and the method comprises the following steps: obtaining the above-ground biomass of a sampled vegetation; obtaining point cloud data of the to-be-detected area and a hyperspectral image with the spatial resolution greater than a set threshold value; a preprocessed hyperspectral image is obtained; calculating to obtain a canopy height model, and extracting vegetation canopy structure features; obtaining spectral indexes and texture features of the sampled vegetation; the features with the importance scores larger than a set threshold value are screened out to serve as model feature variables; constructing a random forest estimation model based on the model characteristic variables and the above-ground biomass of the sampled vegetation; and estimating the above-ground biomass of the vegetation in the to-be-measured area by using the random forest estimation model and the model characteristic variables. The problem that in the prior art, accurate estimation of the biomass on the vegetation ground is difficult to achieve is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent vegetation monitoring technology, specifically a method and system for estimating aboveground biomass of vegetation. Background Technology

[0002] Vegetation has become the primary form of ecological restoration in mining areas, serving as a crucial indicator of ecosystem health and a significant source of carbon sequestration. Developing rapid methods for estimating aboveground biomass is essential for quantifying the carbon sequestration capacity of restored ecosystems and evaluating restoration effectiveness. Traditional methods for estimating aboveground biomass rely on ground sampling surveys, which are time-consuming, costly, and struggle to obtain spatially continuous biomass distribution data. Compared to ground surveys, the rapid development of remote sensing technology has made it possible to obtain spatially continuous aboveground biomass data. However, satellite remote sensing suffers from spatial resolution limitations, making it difficult to achieve refined monitoring and assessment of vegetation growth status in complex environments such as poor soil with uneven nutrient distribution and strong spatial heterogeneity in vegetation configuration.

[0003] In recent years, near-ground remote sensing technology using unmanned aerial vehicles (UAVs) has become a key technology for bridging the scale gap between satellite remote sensing and ground-based surveys due to its high spatial resolution, flexible operation, and relatively low cost. UAV platforms can flexibly carry various sensors according to monitoring needs, providing new solutions for intelligent vegetation monitoring in complex tailings pond ecosystems. UAV hyperspectral imaging technology can capture fine spectral features of vegetation across hundreds of consecutive bands, providing rich information for identifying the physiological and ecological status of vegetation. However, due to spectral saturation, estimation accuracy decreases in areas with high vegetation cover; while hyperspectral imaging provides rich spectral information, its ability to characterize the vertical structure of vegetation is limited.

[0004] UAVs equipped with LiDAR, through their ability to penetrate vegetation canopies, can compensate for the lack of information on spatial vertical structure in hyperspectral data. The synergy between hyperspectral and LiDAR point cloud multi-source remote sensing data has already been applied in biomass estimation (e.g., Chinese patent CN118570677A, "A Method for Estimating Individual Tree Biomass Using Hyperspectral and Air-to-Ground Synergistic LiDAR," for individual tree biomass estimation; Chinese patent CN116773464A, "A Method for Monitoring Larch Caterpillar Pests Based on UAV Hyperspectral and LiDAR," for larch pest and disease monitoring; and Chinese patent CN113030903A, "A Method for Inverting the Nutrient Level of Sedge Based on UAV Hyperspectral and LiDAR," for sedge nutrient level monitoring). However, existing technologies still face challenges in accurately estimating aboveground biomass, particularly in complex environments such as infertile soils with uneven nutrient distribution and strong spatial heterogeneity in vegetation configuration (e.g., tailings dam ecological restoration systems). Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention provides a method and system for estimating aboveground biomass of vegetation, solving the problems of difficulty in accurately estimating aboveground biomass of vegetation in existing technologies.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: A method for estimating aboveground biomass of vegetation includes the following steps: Collect location information, diameter at breast height (DBH), vegetation height, and vegetation type of the sampled vegetation, and obtain the aboveground biomass of the sampled vegetation. Acquire point cloud data and hyperspectral images of the area to be tested with a spatial resolution greater than a set threshold. The hyperspectral images are sequentially subjected to radiometric calibration, atmospheric correction, geometric correction, image stitching, and noise reduction filtering to obtain the preprocessed hyperspectral images. Using the preprocessed hyperspectral image as a reference, the acquired point cloud data is sequentially registered, resampled, denoised, and classified to generate a digital surface model and a digital elevation model. Then, the point cloud data is normalized using the digital elevation model to obtain normalized point cloud data. The difference between the digital surface model and the digital elevation model is calculated to obtain the canopy height model. Based on the normalized point cloud data, vegetation canopy structure features are extracted. Based on the location information of the sampled vegetation and the preprocessed hyperspectral image, the spectral index and texture features of the sampled vegetation are obtained. The importance scores of the aboveground biomass, canopy height model, canopy structure features, spectral index, and texture features of the sampled vegetation were ranked, and features with importance scores greater than a set threshold were selected as model feature variables. A random forest estimation model is constructed based on model feature variables and aboveground biomass of sampled vegetation. The aboveground biomass of vegetation in the test area was estimated using a random forest estimation model and model characteristic variables.

[0007] The beneficial effects of this invention are: This invention fully leverages vegetation spectral indices, texture features, and LiDAR characteristics to overcome the limitations of traditional single-data sources in estimating aboveground biomass in complex environments such as tailings dam ecological restoration systems, where soil nutrient deficiencies are uneven and vegetation configurations are highly spatially heterogeneous. It also addresses the shortcomings of single-source LiDAR data (low spectral resolution) and single-source hyperspectral data (spectral saturation and inability to detect vertical structures). By combining hyperspectral and LiDAR technologies, it achieves accurate estimation of aboveground biomass. Furthermore, by acquiring parameters such as sampled vegetation height and diameter at breast height (DBH), and combining multi-source heterogeneous hyperspectral and LiDAR data, the process is simple, cost-effective, and efficient, effectively overcoming the limitations of traditional ground sampling surveys, which suffer from spatial discontinuity, low efficiency, and high cost. Finally, by combining a random forest estimation model with hyperspectral and LiDAR data, this invention effectively uncovers the complex nonlinear relationship between multi-source heterogeneous information and biomass, improving the model's robustness and estimation capabilities in complex environments.

[0008] Based on the above technical solution, the present invention can be further improved as follows.

[0009] As a preferred technical solution, a diameter at breast height (DBH) measuring rod is used to measure the diameter at breast height (DBH), and a laser rangefinder is used to measure the vegetation height.

[0010] The beneficial effects of adopting the above-mentioned preferred technical solution are: It provides a convenient and accurate way to obtain diameter at breast height (DBH) and vegetation height.

[0011] As a preferred technical solution, the aboveground biomass of the sampled vegetation is obtained based on the allometric growth equation applicable to each vegetation type in the area to be tested.

[0012] The beneficial effects of adopting the above-mentioned preferred technical solution are: The allometric growth equation simplifies the complex relationship between an organism's morphological characteristics, physiological functions, and body size into a quantifiable mathematical expression, making it easier to calculate the aboveground biomass of sampled vegetation more efficiently.

[0013] As a preferred technical solution, a drone equipped with a lidar system is used to acquire point cloud data of the area to be measured, and a drone equipped with a hyperspectral imaging system is used to acquire hyperspectral images with a spatial resolution greater than a set threshold.

[0014] The beneficial effects of adopting the above-mentioned preferred technical solution are: It facilitates the efficient and flexible acquisition of point cloud data and hyperspectral images using drones.

[0015] As a preferred technical solution, the density of point cloud data is greater than 200 points / m².

[0016] The beneficial effects of adopting the above-mentioned preferred technical solution are: It facilitates the acquisition of high-precision point cloud data.

[0017] As a preferred technical solution, hyperspectral images have a spatial resolution greater than 5 cm.

[0018] The beneficial effects of adopting the above-mentioned preferred technical solution are: It facilitates the acquisition of high spatial resolution hyperspectral images.

[0019] As a preferred technical solution, the method of ranking the importance scores of aboveground biomass, canopy height model, canopy structure characteristics, spectral index, and texture characteristics of the sampled vegetation, and selecting features with importance scores greater than a set threshold as model feature variables, includes the following steps: The feature optimization method was used to rank the importance scores of the aboveground biomass, canopy height model, canopy structure features, spectral index, and texture features of the sampled vegetation. The features ranked from largest to smallest importance score were then classified as determined features, undetermined features, and rejected features. After removing rejected features, the determined features are used as model feature variables for the random forest estimation model, and the discriminant function is used to evaluate whether the undetermined features should be used as model feature variables.

[0020] The beneficial effects of adopting the above-mentioned preferred technical solution are: Identifying the model feature variables of the random forest estimation model through feature optimization methods facilitates the construction of a more accurate random forest estimation model.

[0021] As a preferred technical solution, after constructing the random forest estimation model, the coefficient of determination, root mean square error, relative root mean square error, mean absolute error, and residual estimation bias are calculated.

[0022] The beneficial effects of adopting the above-mentioned preferred technical solution are: This facilitates effective evaluation of the stability of random forest estimation models.

[0023] As a preferred technical solution, the steps include: The aboveground biomass of vegetation in the test area was estimated using a random forest estimation model and model characteristic variables, and a distribution map of aboveground biomass of vegetation in the test area was drawn.

[0024] The beneficial effects of adopting the above-mentioned preferred technical solution are: It facilitates the acquisition of spatially continuous vegetation aboveground biomass distribution maps, avoiding the shortcomings of not being able to fully reflect the distribution of vegetation aboveground biomass in the area under test, and making it easier to observe vegetation aboveground biomass more intuitively.

[0025] Based on the above technical solutions, the present invention also provides a vegetation aboveground biomass estimation system.

[0026] A vegetation aboveground biomass estimation system, used to implement the aforementioned vegetation aboveground biomass estimation method, includes the following modules connected in sequence: The aboveground biomass acquisition module of the sampled vegetation is used to: collect the location information, diameter at breast height, vegetation height, and vegetation type of the sampled vegetation, and obtain the aboveground biomass of the sampled vegetation; The point cloud data and hyperspectral image acquisition module is used to: acquire point cloud data and hyperspectral images with a spatial resolution greater than a set threshold for the area to be measured; The preprocessing module is used to perform radiometric calibration, atmospheric correction, geometric correction, image stitching, and noise reduction filtering on the hyperspectral image in sequence to obtain the preprocessed hyperspectral image. The canopy height model and vegetation canopy structure feature acquisition module is used to: use preprocessed hyperspectral images as a reference, perform registration, resampling, denoising, and point cloud classification on the acquired point cloud data in sequence to generate digital surface model and digital elevation model, then use the digital elevation model to normalize the point cloud data to obtain normalized point cloud data, calculate the difference between the digital surface model and the digital elevation model to obtain the canopy height model, and extract vegetation canopy structure features based on the normalized point cloud data; The spectral index and texture feature acquisition module is used to obtain the spectral index and texture features of the sampled vegetation based on the location information of the sampled vegetation and the preprocessed hyperspectral image. The feature selection module is used to: rank the importance scores of the aboveground biomass, canopy height model, canopy structure features, spectral index, and texture features of the sampled vegetation, and select features with importance scores greater than a set threshold as model feature variables; The random forest estimation model building module is used to: construct a random forest estimation model based on model feature variables and aboveground biomass of sampled vegetation; The aboveground biomass estimation module is used to estimate the aboveground biomass of vegetation in the area to be tested using a random forest estimation model and model feature variables.

[0027] Compared with the prior art, the present invention has the following advantages: This invention fully leverages vegetation spectral indices, texture features, and LiDAR characteristics to overcome the limitations of traditional single-data sources in estimating and restoring vegetation biomass in complex environments such as tailings dam ecological restoration systems, where soil nutrient deficiencies are uneven and vegetation configurations are highly spatially heterogeneous. It also addresses the shortcomings of single-source LiDAR data (low spectral resolution) and single-source hyperspectral data (spectral saturation and inability to detect vertical structures). By combining hyperspectral and LiDAR technologies, it achieves accurate estimation of vegetation biomass. Furthermore, by acquiring parameters such as sampled vegetation height and diameter at breast height (DBH), and combining multi-source heterogeneous hyperspectral and LiDAR data, the process is simple, cost-effective, and efficient, effectively overcoming the limitations of traditional ground sampling surveys, which suffer from spatial discontinuity, low efficiency, and high cost. Finally, by combining a random forest estimation model with hyperspectral and LiDAR data, this invention effectively uncovers the complex nonlinear relationship between multi-source heterogeneous information and biomass, improving the model's robustness and estimation capabilities in complex environments. Conveniently and accurately obtain diameter at breast height (DBH) and vegetation height; The allometric growth equation simplifies the complex relationship between the morphological characteristics, physiological functions and body size of organisms into a quantifiable mathematical expression, which facilitates more efficient calculation of aboveground biomass of sampled vegetation. It facilitates the efficient and flexible acquisition of point cloud data and hyperspectral images using drones; It facilitates the acquisition of high-precision point cloud data; It facilitates the acquisition of high spatial resolution hyperspectral images; Identifying the model feature variables of the random forest estimation model through feature optimization methods facilitates the construction of a more accurate random forest estimation model. It facilitates effective evaluation of the stability of random forest estimation models; It facilitates the acquisition of spatially continuous vegetation aboveground biomass distribution maps, avoiding the shortcomings of not being able to fully reflect the distribution of vegetation aboveground biomass in the area under test, and making it easier to observe vegetation aboveground biomass more intuitively. Attached Figure Description

[0028] Figure 1 This is a flowchart of one embodiment of the present invention; Figure 2 This is a distribution map of sampling points for vegetation restoration in a tailings dam according to an embodiment of the present invention; Figure 3 This is a map showing the distribution of aboveground biomass in the restored vegetation of a tailings dam according to an embodiment of the present invention. Detailed Implementation

[0029] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0030] The principles and features of the present invention are described below. The embodiments given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0031] In the following examples, the vegetation is arbor vegetation.

[0032] Example 1 like Figures 1 to 3 As shown, the technical problem to be solved by the present invention is to provide a method for estimating the aboveground biomass of tailings pond vegetation using UAV hyperspectral combined with LiDAR data, addressing the deficiency of existing technologies in being unable to obtain precise distribution of aboveground biomass of restored vegetation in tailings ponds.

[0033] Based on the above-mentioned problems, this invention uses UAV remote sensing technology equipped with a hyperspectral synergistic LiDAR sensor and employs a random forest estimation model to estimate the aboveground biomass of vegetation restored in tailings ponds. This enables refined monitoring of vegetation restoration in tailings pond ecosystems and provides technical support for the evaluation of tailings pond ecological restoration effects and carbon sequestration management.

[0034] Based on the advantages of low-altitude UAV remote sensing, such as low cost and easy operation of non-destructive sampling, this invention makes full use of multi-source heterogeneous data from hyperspectral imagery and lidar point clouds. It combines random forest estimation models to mine spectral, textural, and canopy structure features that are sensitive to aboveground biomass in tailings pond restoration, captures the complex nonlinear relationship between multi-source heterogeneous information and biomass, and establishes a random forest estimation model that can efficiently process high-dimensional, heterogeneous remote sensing data and accurately invert aboveground biomass in tailings pond restoration.

[0035] like Figure 1 As shown, the technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows: Step 1: Investigate the vegetation distribution in the test area, set up quadrats and plan the UAV flight path based on the vegetation distribution. Before aerial photography, ground control points (GCPs) are established using a GNSS receiver. Within each quadrat of the test area, the location information of the sampled vegetation is acquired using a high-precision GNSS receiver, the tree species are recorded, and the diameter at breast height (DBH) is measured using a diameter-at-breast height (DBH) measuring rod and the tree height is measured using a laser rangefinder. Based on the above field sampling parameters, the aboveground biomass of different types of restored vegetation is calculated according to the allometric growth equation applicable to each tree species type in the test area, and the aboveground biomass value of the sampled vegetation is obtained. Step 2: The drone is equipped with a LiDAR system. Operational parameters such as flight altitude, flight speed, and flight path overlap are set to obtain high-precision LiDAR point cloud data with a point cloud density exceeding 200 points / m². On the same day, the drone is equipped with a hyperspectral imaging system with a sampling interval and spectral resolution at the nanometer level. Flight parameters such as flight altitude, forward overlap, and lateral overlap are set to obtain hyperspectral images with a spatial resolution greater than 5cm.

[0036] Step 3: Perform radiometric calibration, atmospheric correction, geometric correction, image stitching, and noise reduction filtering on the hyperspectral images obtained in Step 2 in sequence to obtain preprocessed hyperspectral images.

[0037] Step 4: Select corresponding ground features as ground control points on the LiDAR point cloud data acquired in Step 2 and the preprocessed hyperspectral image in Step 3. Using the preprocessed hyperspectral image as a reference, register, align, and resample the LiDAR point cloud data to maintain the same coordinate system and spatial resolution as the preprocessed hyperspectral image.

[0038] Step 5: Denoise the registered, aligned, and resampled LiDAR point cloud data obtained in Step 4 and then classify the point cloud; generate a Digital Surface Model (DSM) and a Digital Elevation Model (DEM); use the DEM to normalize the LiDAR point cloud data to obtain denoised and normalized LiDAR point cloud data.

[0039] Step 6: Calculate the difference between the DSM and DEM obtained in Step 5 to obtain the canopy height model CHM. Extract vegetation canopy structure features based on the denoised and normalized LiDAR point cloud data obtained in Step 5, thus obtaining LiDAR features including CHM and vegetation canopy structure features.

[0040] Step 7: Calculate the spectral index based on the hyperspectral image preprocessed in Step 3; perform principal component transformation on the hyperspectral image preprocessed in Step 3 to extract the first principal component component, and then extract multi-scale texture features based on the first principal component component combined with the gray-level co-occurrence matrix (using different moving window sizes); Step 8: Overlay the location information of the sampled vegetation from Step 1, the LiDAR features obtained in Step 6, and the spectral index and texture features obtained in Step 7, and extract the LiDAR features, spectral index and texture feature values ​​of the corresponding locations to obtain the LiDAR features, spectral index and texture features of the sampled vegetation.

[0041] Step 9: Based on the aboveground biomass of the sampled vegetation from Step 8 and the corresponding LiDAR features, spectral indices, and texture feature parameters, Boruta features are screened. Features are ranked according to their importance and categorized into defined features, undefined features, and rejected features. Undefined features are further evaluated using a discriminant function. Feature variables with high importance to the measured aboveground biomass of the restored vegetation in the tailings dam are selected as model feature variables.

[0042] Step 10: Based on the model feature variables obtained in Step 9 and the aboveground biomass data of the sampled vegetation, construct a random forest estimation model for estimating the aboveground biomass of vegetation for tailings dam restoration, and use leave-one-out cross-validation to evaluate the stability and accuracy of the random forest estimation model.

[0043] Step 11: Using the random forest estimation model established in Step 10 and the model feature variables selected in Step 9, the aboveground biomass of the restored vegetation in the test area is inverted, and the accuracy of the model is verified by independently obtained aboveground biomass samples of the sampled vegetation.

[0044] This invention fully leverages vegetation spectral indices, texture features, and LiDAR characteristics to overcome the limitations of traditional single-data sources in estimating the accuracy of aboveground biomass restoration in complex environments such as poor soil with uneven nutrient distribution and strong spatial heterogeneity in vegetation configuration (e.g., tailings dam ecological restoration systems). It also addresses the shortcomings of single-source LiDAR data (low spectral resolution) and single-source hyperspectral data (spectral saturation and inability to detect vertical structures). By combining hyperspectral and LiDAR technologies, it achieves accurate estimation of aboveground biomass.

[0045] This invention obtains parameters such as the height and diameter at breast height of the sampled vegetation in a limited manner, and combines hyperspectral and LiDAR multi-source heterogeneous data. The processing flow is simple, easy to operate, low-cost and highly efficient, effectively overcoming the limitations of traditional ground sampling surveys, such as spatial discontinuity, low efficiency and high cost.

[0046] This invention combines a random forest estimation model to process hyperspectral and LiDAR data, effectively mining the complex nonlinear relationship between multi-source heterogeneous information and biomass, and improving the model's robustness and estimation ability in complex environments.

[0047] Example 2 like Figures 1 to 3 As shown, based on Example 1, this example provides a more detailed implementation method.

[0048] Step 1: Investigate the vegetation distribution in the area to be tested, set up quadrats and plan UAV flight routes based on the vegetation distribution. Before aerial photography, ground control points (GCPs) are set up using a GNSS receiver. Within each 5m×5m quadrat in the iron tailings dam area to be tested, the location information of the sampled vegetation is obtained using a high-precision GNSS receiver. For vegetation with a diameter at breast height (DBH) of 10cm or more, the DBH of standing trees at 1.37m above the ground is measured using a DBH measuring tape, and the tree height is measured using a laser rangefinder. The tree species type is recorded. Based on the field sampling parameters, the aboveground biomass of different types of restored vegetation is calculated according to the corresponding allometric growth equation in "Carbon Storage-Biomass Equation of Chinese Forest Ecosystems" (as shown in Table 1, where D is DBH and H is tree height). The aboveground biomass value of the sampled vegetation is obtained, and the geographical location of each tree is recorded using a GNSS receiver (e.g., ...). Figure 2 (Solid circle mark).

[0049] Table 1. Estimation of allometric growth equations for different tree species within the tailings dam area. Step Two: Data was acquired on the same day using both a lidar system and a hyperspectral imaging system. An AlphaAir450 lidar system was used on a drone, with flight parameters set at a 500m flight altitude, 10m / s flight speed, and 60% flight path overlap, to obtain high-precision data with a point cloud density exceeding 200 points / m². Simultaneously, a GaiaSky-mini3-VN / NIR+POS hyperspectral imaging system (400-1000nm spectral range, 224 bands, 1.35nm sampling interval, 5nm@32μm spectral resolution) was used, with flight parameters set at a 500m flight altitude, 75% forward overlap, and 65% lateral overlap, to obtain hyperspectral images with a spatial resolution of 5cm, covering an area of ​​approximately 2km² in a single flight.

[0050] Step 3: Perform data preprocessing on the hyperspectral images obtained in Step 2 in sequence: perform radiometric calibration based on 50% gray cloth reference data; perform atmospheric correction using an atmospheric correction model; complete geometric correction based on the deployed ground control points; finally, use the Seamless Mosaic algorithm to stitch the images together, and then perform noise reduction filtering to obtain the preprocessed hyperspectral images.

[0051] Step 4: Select corresponding ground features as ground control points on the LiDAR point cloud data acquired in Step 2 and the preprocessed hyperspectral image in Step 3. Using the preprocessed hyperspectral image as a reference, register, align, and resample the LiDAR point cloud data to a resolution of 5cm to ensure that it maintains the same coordinate system and spatial resolution as the preprocessed hyperspectral image.

[0052] Step 5: The registered, aligned, and resampled LiDAR point cloud data obtained in Step 4 are processed as follows: flight strip stitching and denoising filtering based on the K-nearest neighbor distance statistical method are performed; the improved progressive densification triangulation algorithm is used to classify the point cloud, generating a digital surface model (DSM) and a digital elevation model (DEM) with a spatial resolution of 5cm; the DEM is normalized to obtain denoised and normalized point cloud data, which is the final DSM, DEM, and denoised and normalized point cloud data.

[0053] Step Six: Calculate the difference between the DSM and DEM obtained in Step Five to obtain the canopy height model CHM. Based on the denoised and normalized LiDAR point cloud data obtained in Step Four, extract vegetation canopy structure features, i.e., obtain LiDAR features including CHM and vegetation canopy structure features, specifically including three types of features: structural features (CHM and leaf area index); height features (H1 / H5 / H10 / H20 / H25 / H30 / H40 / H50 / H60 / H70 / H75 / H80 / H90 / H95 / H99 percentiles, cumulative distribution values ​​of AIH1 / AIH5 / AIH10 / AIH20 / AIH25 / AIH30 / AIH40 / AIH50 / AIH60 / AIH70 / AIH75 / AIH80 / AIH90 / AIH95 / AIH99, etc.); density features (10 density distribution indicators from D0 to D9).

[0054] Step 7: Based on the hyperspectral image preprocessed in Step 3, calculate 46 spectral indices reflecting different vegetation physiological and ecological characteristics (as shown in Table 2); perform principal component transformation on the hyperspectral image preprocessed in Step 3 to extract the first principal component, and then extract 24 texture features based on the first principal component combined with the gray-level co-occurrence matrix. In this embodiment, three moving window sizes of 3×3, 5×5, and 7×7 pixels are used to calculate eight gray-level co-occurrence matrix texture features, including mean, variance, homogeneity, heterogeneity, contrast, information entropy, second moment of angle, and correlation, ultimately obtaining 24 texture feature variables.

[0055] Table 2 Contents of the Spectral Index Dataset Table 2. Contents of the Spectral Index Dataset (continued) In Table 2, , , These represent the band numbers, Indicates band reflectivity, Indicates band reflectivity, Indicates band The reflectance was analyzed by performing mathematical operations on any two or three bands of the preprocessed hyperspectral image, and the correlation between the combined spectral index and the aboveground biomass of vegetation was analyzed. By comparing the correlation coefficients, the most significant characteristic band combinations were selected.

[0056] Step 8: Overlay the location information of the sampled vegetation from Step 1, the 58 LiDAR features obtained in Step 6 including CHM and vegetation canopy structure features, and the 46 spectral indices and 24 texture features obtained in Step 7, and extract the LiDAR features, spectral indices and texture feature values ​​at the corresponding locations to obtain the canopy structure features, spectral indices and texture features of the sampled vegetation.

[0057] Step Nine: As shown in Tables 3 and 4, based on the aboveground biomass of the sampled vegetation from Step Eight and the corresponding 58 LiDAR features, 46 spectral indices, and 24 texture features (a total of 128 feature variables), Boruta feature screening was performed. The iteration count was set to 99. Features were sorted according to their importance scores, and Z-scores were obtained through standardization. The Z-score with the highest shadow feature was used as the importance score threshold. Features with Z-scores significantly greater than the importance score threshold were accepted as confirmed features; those significantly less than the threshold were marked as rejected features. Features with Z-scores not significantly different from the threshold or with large fluctuations that could not be clearly judged were temporarily designated as pending features. A discriminant function was used to perform a secondary evaluation on these pending features to determine whether they were confirmed or rejected. Thirty-eight confirmed features (33 spectral indices, 4 LiDAR features, and 1 texture feature) with high importance to the measured aboveground biomass of the restored tree vegetation in the tailings dam were selected as model feature variables for the random forest estimation model.

[0058] Table 3. Results of Boruta Feature Screening and Feature Determination Table 4. Results of Secondary Screening of Undetermined Features in Boruta Feature Filtering Step 10: Based on the model feature variables of the 38 random forest estimation models obtained in Step 9 and the sampled aboveground biomass data of vegetation, construct random forest estimation models to estimate the aboveground biomass of vegetation in the tailings dam restoration project. Calculate the coefficient of determination (R²) using leave-one-out cross-validation. 2 The root mean square error (RMSE), relative root mean square error (rRMSE), mean absolute error (MAE), and residual estimation bias (RPD) (Table 5) were used to evaluate the stability and accuracy of the random forest estimation model.

[0059] Table 5 Accuracy of Random Forest Estimation Model Step 11: Using the random forest estimation model established in Step 10 and the model feature variables of the 38 random forest estimation models selected in Step 9, the aboveground biomass of the restored vegetation in the tailings dam is inverted. The 38 model feature variables are input into the random forest estimation model for prediction. The spatial distribution data of vegetation is used for masking, and finally, a distribution map of aboveground biomass of the restored arbor vegetation in the tailings dam is generated. Figure 3 ).

[0060] More detailed explanations are as follows: I. Regarding preprocessing: 1. Radiometric calibration: Also known as radiometric standardization, it is the process of converting the raw digital quantization values ​​(DN values) recorded by a drone's camera sensor into physically meaningful absolute radiance (or apparent reflectance). Simply put, it is converting the grayscale values ​​"seen" by the camera into the actual physical energy values ​​"emitted" by the object itself.

[0061] 2. Atmospheric Correction: Atmospheric correction is the process of eliminating or reducing the influence of atmospheric molecules, aerosols, water vapor, etc., on the scattering and absorption of electromagnetic waves. Its goal is to restore the "apparent reflectivity" or "radiance" received by the sensor to the true "surface reflectivity" of the ground object.

[0062] 3. Geometric Correction: Geometric correction is the process of correcting geometric distortions in an image caused by factors such as sensor orientation, terrain undulation, and Earth curvature, so that the pixel positions in the image correspond one-to-one with their actual geographic coordinates on the Earth's surface. It mainly includes two levels: coarse geometric correction and orthorectification.

[0063] 4. Image stitching: Image stitching is the process of seamlessly merging multiple single UAV images with overlapping areas, after geometric correction, to generate a complete, large-format orthophoto map covering the entire survey area.

[0064] 5. Noise Reduction Filtering: Noise reduction filtering is a process that uses digital image processing techniques to suppress or eliminate noise in an image while preserving as much of the original image details and information as possible. Noise is unwanted random signals introduced during imaging, transmission, or recording.

[0065] II. Boruta Feature Filtering Algorithm Principle and Process: 1. Core Algorithm Principles: Boruta is a wrapper-style fully correlated feature selection algorithm based on random forests. Its core idea is to construct "shadow features" to simulate the impact of random noise on the model, thereby establishing a dynamic feature importance benchmark. The algorithm uses statistical tests and discriminant functions to select features with importance scores significantly higher than this random benchmark, ensuring that the selected features have a true explanatory power for the dependent variable.

[0066] 2. Algorithm Implementation Steps: Step 1: Constructing the Shadow Feature Space. First, the original feature matrix is ​​copied. Then, each column of the copied feature data is independently and randomly rearranged. This process eliminates the original correlation between the features and the dependent variable, generating "shadow features" that retain only the data distribution characteristics but have no actual predictive power.

[0067] Step 2: Model Training and Importance Measurement. Concatenate the original features and shadow features to form an augmented matrix, and input this matrix into the random forest model for iterative training. After training, calculate the importance scores for all features.

[0068] Step 3: Determine the random baseline threshold. In each iteration, extract the maximum value among all shadow feature importance scores. This value represents the highest level of importance that random noise can achieve under the current data distribution, and serves as the random baseline threshold for determining whether a feature is effective.

[0069] Step 4: Significance Determination and Iteration. Using the binomial distribution test, the maximum value between the importance score of each original feature and the importance score of the shadow feature is statistically inferred: if the number of hits of a feature is significantly higher than expected, it is marked as "determined feature"; if the number of hits of a feature is significantly lower than expected, it is marked as "rejected feature"; if a significance determination cannot be made at the current confidence level, it is temporarily marked as "undetermined feature" and the next round of iteration continues.

[0070] Step 5: Secondary discrimination of undetermined features. After reaching the maximum number of iterations, if there are still features in the "undetermined" state, the algorithm will introduce a discriminant function to perform secondary screening. By comparing the statistical difference between the historical importance distribution of undetermined features and the shadow feature distribution, the remaining undetermined features are finally classified as "determined features" or "rejected features".

[0071] 3. Feature classification results: After multiple rounds of iteration and final selection by the discriminant function, all original features will be clearly divided into two categories: ① Recognized features: those confirmed by statistical tests or the discriminant function to contain significantly more information than random noise, and are retained. ② Rejected features: irrelevant features that are determined to have no significant contribution to the model's prediction are removed.

[0072] III. Random Forest Estimation Model: 1. Core Algorithm Principle: Random Forest is an ensemble learning method based on decision trees. For the biomass estimation problem of this invention, the model constructs multiple regression decision trees and uses the arithmetic mean of the predictions from all decision trees as the final output to effectively reduce model variance and control overfitting. Its core mechanism includes: Bootstrap resampling: Random sampling with replacement is performed on the training set, and each decision tree is trained based on a different subset of samples, which enhances the model's robustness to sample perturbations.

[0073] Feature random subspace: During the node splitting process of each tree, only a portion of features are randomly selected from the entire feature set as candidate splitting terms, which further reduces the correlation between trees.

[0074] 2. Model training and validation process: Step 1: Data Preprocessing and Standardization. Key features selected using the Boruta algorithm (including spectral indices, texture features, and LiDAR features) are used as independent variables X, and aboveground biomass obtained in the field is used as the dependent variable Y. X and Y are standardized to a standard normal distribution with a mean of 0 and a variance of 1 to eliminate dimensional differences.

[0075] Step 2: Noise Enhancement. To improve the model's generalization ability and anti-interference capability, a small amount of Gaussian noise is injected into the independent variable X.

[0076] Step 3: Hyperparameter Grid Optimization. Define the hyperparameter search space and employ a grid search strategy to minimize prediction error, selecting the optimal hyperparameter combination.

[0077] Step 4: Leave-one-out cross-validation and accuracy evaluation. Using the selected optimal hyperparameter combination, leave-one-out cross-validation is performed for out-of-sample prediction. After restoring the predicted values ​​through inverse standardization, the coefficient of determination is calculated. Root mean square error Mean absolute error Relative root mean square error and prediction deviation ratio To evaluate the model's accuracy, the calculation formula is as follows: (1) (2) (3) (4) (5) in, Represents the total number of samples. Indicates the sample number. This represents the estimated biomass value. This represents the actual value of biomass. This represents the average actual biomass.

[0078] Step 5: Final Model Construction. Based on the optimal hyperparameter combination, the final model is trained and packaged using all the original data.

[0079] 3. Model parameter settings: This invention constructs a search space that includes parameters such as the number of decision trees and the maximum depth of the trees. The optimal parameter combination determined by grid search is shown in Table 6.

[0080] Table 6 Optimal parameter combinations determined by grid search As described above, the present invention can be implemented well.

[0081] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0082] In the description of this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0083] In the description of this invention, although embodiments of the invention have been shown and described herein, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this invention.

[0084] In the description of this invention, all features disclosed in all embodiments of this specification, or steps in all methods or processes implied in the disclosure, may be combined and / or extended or replaced in any way, except for mutually exclusive features and / or steps.

[0085] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Based on the technical essence of the present invention, any simple modifications, equivalent substitutions, and improvements made to the above embodiments within the spirit and principles of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for estimating aboveground biomass of vegetation, characterized in that, Includes the following steps: Collect location information, diameter at breast height (DBH), vegetation height, and vegetation type of the sampled vegetation, and obtain the aboveground biomass of the sampled vegetation. Acquire point cloud data and hyperspectral images of the area to be tested with a spatial resolution greater than a set threshold. The hyperspectral images are sequentially subjected to radiometric calibration, atmospheric correction, geometric correction, image stitching, and noise reduction filtering to obtain the preprocessed hyperspectral images. Using the preprocessed hyperspectral image as a reference, the acquired point cloud data is sequentially registered, resampled, denoised, and classified to generate a digital surface model and a digital elevation model. Then, the point cloud data is normalized using the digital elevation model to obtain normalized point cloud data. The difference between the digital surface model and the digital elevation model is calculated to obtain the canopy height model. Based on the normalized point cloud data, vegetation canopy structure features are extracted. Based on the location information of the sampled vegetation and the preprocessed hyperspectral image, the spectral index and texture features of the sampled vegetation are obtained. The importance scores of the aboveground biomass, canopy height model, canopy structure features, spectral index, and texture features of the sampled vegetation were ranked, and features with importance scores greater than a set threshold were selected as model feature variables. A random forest estimation model is constructed based on model feature variables and aboveground biomass of sampled vegetation. The aboveground biomass of vegetation in the test area was estimated using a random forest estimation model and model characteristic variables.

2. The method for estimating aboveground biomass of vegetation according to claim 1, characterized in that, The diameter at breast height (DBH) was measured using a DBH measuring tape, and the vegetation height was measured using a laser rangefinder.

3. The method for estimating aboveground biomass of vegetation according to claim 2, characterized in that, The aboveground biomass of the sampled vegetation was obtained based on the allometric growth equation applicable to each vegetation type within the test area.

4. The method for estimating aboveground biomass of vegetation according to claim 1, characterized in that, Use a drone equipped with a lidar system to acquire point cloud data of the area to be measured, and use a drone equipped with a hyperspectral imaging system to acquire hyperspectral images with a spatial resolution greater than a set threshold.

5. The method for estimating aboveground biomass of vegetation according to claim 1, characterized in that, The density of point cloud data is greater than 200 points / m².

6. The method for estimating aboveground biomass of vegetation according to claim 1, characterized in that, The spatial resolution of hyperspectral images is greater than 5 cm.

7. The method for estimating aboveground biomass of vegetation according to claim 1, characterized in that, The process of ranking the importance scores of aboveground biomass, canopy height model, canopy structure features, spectral index, and texture features of the sampled vegetation, and selecting features with importance scores greater than a set threshold as model feature variables, includes the following steps: The feature optimization method was used to rank the importance scores of the aboveground biomass, canopy height model, canopy structure features, spectral index, and texture features of the sampled vegetation. The features ranked from largest to smallest importance score were then classified as determined features, undetermined features, and rejected features. After removing rejected features, the determined features are used as model feature variables for the random forest estimation model, and the discriminant function is used to evaluate whether the undetermined features should be used as model feature variables.

8. The method for estimating aboveground biomass of vegetation according to claim 1, characterized in that, After constructing the random forest estimation model, the coefficient of determination, root mean square error, relative root mean square error, mean absolute error, and residual estimation bias are calculated.

9. A method for estimating aboveground biomass of vegetation according to any one of claims 1 to 8, characterized in that, Includes the following steps: The aboveground biomass of vegetation in the test area was estimated using a random forest estimation model and model characteristic variables, and a distribution map of aboveground biomass of vegetation in the test area was drawn.

10. A vegetation aboveground biomass estimation system, characterized in that, A method for estimating aboveground biomass of vegetation as described in any one of claims 1 to 9, comprising the following modules connected in sequence: The aboveground biomass acquisition module of the sampled vegetation is used to: collect the location information, diameter at breast height, vegetation height, and vegetation type of the sampled vegetation, and obtain the aboveground biomass of the sampled vegetation; The point cloud data and hyperspectral image acquisition module is used to: acquire point cloud data and hyperspectral images with a spatial resolution greater than a set threshold for the area to be measured; The preprocessing module is used to perform radiometric calibration, atmospheric correction, geometric correction, image stitching, and noise reduction filtering on the hyperspectral image in sequence to obtain the preprocessed hyperspectral image. The canopy height model and vegetation canopy structure feature acquisition module is used to: use preprocessed hyperspectral images as a reference, perform registration, resampling, denoising, and point cloud classification on the acquired point cloud data in sequence to generate digital surface model and digital elevation model, then use the digital elevation model to normalize the point cloud data to obtain normalized point cloud data, calculate the difference between the digital surface model and the digital elevation model to obtain the canopy height model, and extract vegetation canopy structure features based on the normalized point cloud data; The spectral index and texture feature acquisition module is used to obtain the spectral index and texture features of the sampled vegetation based on the location information of the sampled vegetation and the preprocessed hyperspectral image. The feature selection module is used to: rank the importance scores of the aboveground biomass, canopy height model, canopy structure features, spectral index, and texture features of the sampled vegetation, and select features with importance scores greater than a set threshold as model feature variables; The random forest estimation model building module is used to: construct a random forest estimation model based on model feature variables and aboveground biomass of sampled vegetation; The aboveground biomass estimation module is used to estimate the aboveground biomass of vegetation in the area to be tested using a random forest estimation model and model feature variables.

Citation Information

Patent Citations

  • Crex nutrition level inversion method based on unmanned aerial vehicle hyperspectrum and laser radar

    CN113030903A

  • Fallen leaf pine moth pest monitoring method based on unmanned aerial vehicle hyperspectrum and laser radar

    CN116773464A

  • Individual tree biomass estimation method based on hyperspectrum and air-ground collaborative LiDAR

    CN118570677A

  • Method for joint inversion of forest aboveground biomass by integrating three data sources

    CN108921885A

  • Tree species classification method based on an unmanned aerial vehicle hyperspectral image and LiDAR point cloud

    CN109492563A