A forest canopy height remote sensing estimation method and computer readable medium
By employing a lightweight gradient boosting machine learning method with spatial filtering, and combining multi-source remote sensing data and topographic data, a forest canopy height prediction model is constructed. This addresses the issue of existing technologies not considering spatial effects and achieves higher accuracy in forest canopy height estimation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN UNIV
- Filing Date
- 2023-03-17
- Publication Date
- 2026-05-01
AI Technical Summary
Existing remote sensing methods for estimating forest canopy height fail to effectively consider the spatial effects of vegetation canopy height and environmental factors, resulting in poor model transferability and insufficient accuracy.
A lightweight gradient boosting machine learning method with spatial filtering is adopted. By constructing the feature vector of the spatial weight matrix and combining multi-source remote sensing data and topographic data, a forest canopy height prediction model is built, taking into account the complex nonlinear relationship and spatial effects between vegetation canopy height and environmental factors.
It improves the accuracy and robustness of forest canopy height prediction, enabling more accurate estimation of forest canopy height.
Smart Images

Figure CN116381700B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of remote sensing technology, and in particular relates to a remote sensing estimation method for forest canopy height and a computer-readable medium. Background Technology
[0002] Global climate change has become a focus of international concern, and forest carbon cycling is a crucial aspect of global climate change research. Approximately 77% of vegetation carbon in terrestrial ecosystems is stored in forests. To mitigate the impacts of global warming and protect forest ecosystems, my country has launched a national carbon emissions trading system. Accurately estimating the spatial distribution pattern and dynamic changes of forest carbon storage is fundamental to terrestrial ecosystem carbon budget accounting. Canopy height is a vital factor and fundamental data for estimating forest carbon storage, and a key link in terrestrial ecosystem carbon budget accounting. Traditional methods for surveying forest canopy height are time-consuming and labor-intensive, and can only obtain information over small areas. LiDAR (Light Detection and Ranging) can penetrate the canopy to accurately obtain information on the vertical structure of the forest canopy, measuring the height of the ground surface and canopy, as well as average canopy area and tree density. It demonstrates significant advantages in estimating forest canopy height, especially spaceborne LiDAR data, which has unique advantages in large-area forest canopy height inversion and mapping. Remote sensing technology can objectively and rapidly acquire forest parameters at various scales, thus possessing significant advantages in forest canopy height mapping. Currently, optical remote sensing data is the main data source for estimating forest canopy height. It can accurately and quickly estimate the height of forest canopy over a large area. It makes full use of reflectivity, spatial characteristics, texture, vegetation index information and spectral information of different bands to reflect more detailed forest remote sensing information. At the same time, the impact of cloud cover on image quality can be solved by using annual composite indicators and multi-year image completion methods.
[0003] Common modeling methods for estimating forest canopy height include traditional statistical regression and machine learning methods. Statistical regression models are simple and computationally efficient, but they can only describe the linear relationship between forest canopy height and multi-source remote sensing indicators; the most representative model is the multiple linear regression model. Machine learning methods can better reflect the complex nonlinear relationship between remote sensing indicators and vegetation canopy height. A representative algorithm is the random forest algorithm, which constructs a series of decision trees by randomly sampling samples and derives predictions through voting. This algorithm can also evaluate the importance of variables and is often used for feature selection. Support vector regression models maintain good performance for high-dimensional features and can explain nonlinear regression problems by setting different kernel functions, exhibiting strong generalization ability; however, its computational cost is also very high when the sample size is very large. Lightweight Gradient Boosting Machine (LightGBM) learning machines have advantages such as robustness against interference and overfitting, fast computation speed, and strong stability. Compared with deep learning algorithms, they have lower computational cost, require fewer parameters and less data input, and can achieve excellent results. It also demonstrates strong model transferability, and can adjust the weights of different models based on the input data during training, making it applicable in fields such as ecology and environment.
[0004] According to the first law of geography, both vegetation canopy height and the spatial distribution of environmental factors exhibit spatial autocorrelation. However, remote sensing inversion of vegetation canopy height has not considered the influence of spatial effects on the model. Griffith's eigenvector spatial filtering method, through eigenvalue decomposition of the spatial weight matrix constructed from geographic units, maps spatial effects into eigenvectors. By filtering out significant eigenvector sets, it identifies the spatial effects influencing the distribution of geographic variables, representing the spatial distribution patterns of geographic variables and the spatial influence of geographic units. These are then incorporated as independent variables into the model. This method considers the variance inflation effect and regression coefficient offset effect caused by spatial autocorrelation in statistical modeling, thereby reducing the impact of spatial effects on the model and improving model accuracy. The advantage of this method lies in using eigenvectors of the spatial weight matrix to express spatial influence, exhibiting strong scalability and direct applicability to linear and generalized linear regression. It has been applied in areas such as air pollution, vegetation cover, and landslide disasters, with results showing that the eigenvector spatial filtering method significantly improves model accuracy.
[0005] In summary, machine learning-based remote sensing estimation and prediction of vegetation canopy height does not consider the spatial effects of environmental factors on vegetation canopy height, and the model's transferability is also poor. Therefore, there is an urgent need for a lightweight gradient boosting machine learning method based on spatial effects to predict vegetation canopy height, providing important support for forest carbon storage and global climate change analysis. Summary of the Invention
[0006] The purpose of this invention is to provide a method and computer-readable medium for estimating forest canopy height, enabling the efficient and accurate establishment of a forest canopy height estimation model based on machine learning and remote sensing data, and completing canopy height mapping.
[0007] The technical solution of this invention is a method for estimating forest canopy height using spatial filtering, and the specific steps are as follows:
[0008] Step 1: Obtain the coordinates and relative height parameters of the footprint points of the spaceborne lidar data in the forest area at multiple historical moments, and obtain the preprocessed coordinates and relative height parameters of the footprint points of the spaceborne lidar data at multiple historical moments through preprocessing.
[0009] Step 2: Acquire near-ground vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data. Through projection transformation and matching, obtain vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data that are the same as the coordinates of the footprint points of the preprocessed spaceborne lidar data from multiple historical moments.
[0010] Step 3: Construct a candidate variable set by combining the relative height parameter data of the spaceborne lidar data from multiple historical moments, land use classification data, vegetation cover data, and topographic data that are the same as the coordinates of the footprint points of the preprocessed spaceborne lidar data from multiple historical moments. Using the variable screening method, filter the variables that are significantly related to the vegetation canopy height data collected near the ground to obtain a set of feature variables that are significantly related to the vegetation canopy height data collected near the ground.
[0011] Step 4: Construct a random forest learning model with optimal model parameters. Calculate the relative height parameters of the preprocessed spaceborne lidar data from multiple historical moments using the random forest learning model with optimal model parameters to obtain the calibrated vegetation canopy height of each footprint point in the spaceborne lidar data at each historical moment.
[0012] Step 5: Acquire passive optical remote sensing data, synthetic aperture radar data, interferometric radar data, and terrain data from multiple historical moments. Using radar remote sensing data preprocessing methods, obtain preprocessed passive optical remote sensing data, synthetic aperture radar data, interferometric radar data, and terrain data from multiple historical moments.
[0013] Step 6: Calculate the set of remote sensing vegetation indices and the set of annual phenological indices using preprocessed passive optical remote sensing data from multiple historical moments. Combine the preprocessed passive optical remote sensing data, synthetic aperture radar data, and interferometric radar data from multiple historical moments to calculate the set of annual image statistics. Construct a set of terrain feature indices using terrain data from multiple historical moments, and further construct a set of non-spatial feature vectors.
[0014] Step 7: Construct a spatial weight matrix by using the spatial distance relationship between the coordinates of the footprint points of the preprocessed spaceborne lidar data from multiple historical moments. Then, center the spatial weight matrix and calculate its eigenvalues and eigenvectors. Arrange the eigenvectors of the spatial weight matrix in descending order of their corresponding eigenvalues to obtain the sorted eigenvectors of the spatial weight matrix.
[0015] Step 8: Filter the eigenvectors of the sorted spatial weight matrices whose eigenvalues are greater than a threshold to construct an initial set of spatial eigenvectors; combine the non-spatial eigenvector set with the initial set of spatial eigenvectors to construct an environmental variable set.
[0016] Step 9: Construct an optimized lightweight gradient boosting machine learning model. The optimized lightweight gradient boosting machine learning model is used to calculate the predicted value of forest canopy height by combining preprocessed passive optical remote sensing data, synthetic aperture radar data, and interferometric radar data from multiple historical moments.
[0017] Preferably, the preprocessing procedure described in step 1 is as follows:
[0018] The coordinates and relative height parameters of the footprint points of the spaceborne lidar data at multiple historical moments are sequentially subjected to projection transformation, outlier processing, and noise point removal.
[0019] As a preferred embodiment, the projection transformation matching in step 2 is specifically as follows:
[0020] The vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data are projected and transformed to the coordinate system corresponding to the coordinate positions of the footprint points of the spaceborne lidar data. The coordinate positions of the preprocessed footprint points of the spaceborne lidar data at multiple historical moments described in step 1 are then matched to obtain vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data that are identical to the coordinate positions of the preprocessed footprint points of the spaceborne lidar data at multiple historical moments.
[0021] As a preferred embodiment, step 4, which involves constructing a random forest learning model with optimal model parameters, is as follows:
[0022] A random forest learning model is constructed by sequentially inputting each element from a set of feature variables significantly correlated with near-ground collected vegetation canopy height data as a sample to predict the vegetation canopy height. A loss function model is then constructed using the near-ground collected vegetation canopy height data as the true label. Root mean square error and coefficient of determination are used as evaluation metrics. The optimal random forest learning model with the best parameters is obtained through optimized training.
[0023] Preferably, the radar remote sensing data preprocessing method in step 5 is as follows:
[0024] The following steps are performed sequentially: radiometric correction, missing value processing, resampling, stitching and segmentation, and projection conversion.
[0025] Preferably, the set of remote sensing vegetation indices in step 6 consists of multiple types of remote sensing vegetation index indicators.
[0026] The annual phenological index set described in step 6 consists of various types of annual phenological indicators;
[0027] The annual image statistics set mentioned in step 6 consists of various types of annual image statistics indicators;
[0028] The set of terrain feature indicators described in step 6 consists of various types of terrain feature indicators;
[0029] Step 6, which involves constructing a non-spatial feature vector set, is detailed below:
[0030] The remote sensing vegetation index set, annual phenological index set, annual image statistical value set, and topographic feature index set are sequentially filtered using variable screening methods. Variables that are significantly correlated with the calibrated vegetation canopy height of each footprint point in the spaceborne lidar data at each historical moment described in step 5 are then filtered to obtain the filtered remote sensing vegetation index set, the filtered annual phenological index set, the filtered annual image statistical value set, and the filtered topographic feature index set.
[0031] As a preferred embodiment, the optimized lightweight gradient boosting machine learning model described in step 9 is as follows:
[0032] A lightweight gradient boosting machine learning model is constructed. The set of environmental variables is used as samples to predict the height of the vegetation canopy. The height of the vegetation canopy after calibration at each footprint point of the satellite lidar data at each historical moment is used as the true label to construct a loss function model. The root mean square error and the coefficient of determination are used as evaluation indicators. The optimized lightweight gradient boosting machine learning model is obtained through optimization training.
[0033] The present invention also provides a computer-readable medium storing a computer program executed by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of the forest canopy height estimation method of the spatial filter value.
[0034] The beneficial effects of this invention are as follows: This invention provides a lightweight gradient boosting machine for remote sensing prediction of forest canopy height based on spatial filtering. In the estimation of vegetation canopy height, it considers the complex nonlinear relationship between vegetation canopy height and remote sensing multi-band and environmental factors. At the same time, it considers the influence of spatial effects and adds them to the lightweight gradient boosting machine model in the form of spatial feature vectors. This can more accurately construct the vegetation canopy height model and improve the robustness of the model, thereby improving the accuracy of vegetation canopy height prediction. Attached Figure Description
[0035] Figure 1 : Flowchart of the method according to an embodiment of the present invention.
[0036] Figure 2 : A schematic diagram of the height calibration model according to an embodiment of the present invention.
[0037] Figure 3 : A schematic diagram of a lightweight gradient boosting machine model according to an embodiment of the present invention. Detailed Implementation
[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0039] In specific implementation, the method proposed in the technical solution of this invention can be automatically executed by those skilled in the art using computer software technology. System devices for implementing the method, such as computer-readable storage media storing the corresponding computer program of the technical solution of this invention and computer equipment including the computer program running the corresponding computer program, should also be within the protection scope of this invention.
[0040] The following is combined with Figure 1-3 The technical solution of this invention is a method for estimating forest canopy height using spatial filtering, as detailed below:
[0041] The flowchart of the method of the present invention is as follows: Figure 1 As shown.
[0042] Step 1: Obtain the coordinates and relative height parameters of footprint points from the spaceborne lidar data in the forest area at multiple historical time points. Preprocess the data to obtain the preprocessed coordinates and relative height parameters of the footprint points at multiple historical time points. The acquired spaceborne lidar metadata includes, but is not limited to, footprint point coordinates, percentage-related height parameters, data acquisition time, laser model, acquisition attitude, footprint point terrain, and other parameters affecting the data acquisition quality of the spaceborne lidar.
[0043] The preprocessing process described in step 1 is as follows:
[0044] The coordinates and relative height parameters of the footprint points of the spaceborne lidar data at multiple historical moments are sequentially subjected to projection transformation, outlier processing, and noise point removal.
[0045] Step 2: Acquire near-ground vegetation canopy height data, land use classification data, vegetation cover data, and topographic data. Through projection transformation and matching, obtain vegetation canopy height data, land use classification data, vegetation cover data, and topographic data that match the coordinates of preprocessed spaceborne lidar footprints from multiple historical time points. The ground-collected vegetation canopy height data is divided into four categories: vegetation canopy height models collected by lidar on aircraft, vegetation canopy height models collected by lidar on drones, vegetation canopy height models collected by ground-based lidar, and manually measured individual tree heights.
[0046] Step 2, the projection transformation matching, is as follows:
[0047] The vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data are projected and transformed to the coordinate system corresponding to the coordinate positions of the footprint points of the spaceborne lidar data. The coordinate positions of the preprocessed footprint points of the spaceborne lidar data at multiple historical moments described in step 1 are then matched to obtain vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data that are identical to the coordinate positions of the preprocessed footprint points of the spaceborne lidar data at multiple historical moments.
[0048] Step 3: Construct a candidate variable set by combining the relative height parameters of the spaceborne lidar data from multiple historical moments, land use classification data, vegetation cover data, and topographic data that have the same coordinate positions as the footprint points of the preprocessed spaceborne lidar data from multiple historical moments. Then, use variable selection methods to filter variables that are significantly correlated with the vegetation canopy height data collected near the ground. Commonly used variable selection methods include random forest, minimum shrinkage and selection operators, recursive feature elimination, LightGBM-based selection model, stepwise regression model, principal component analysis, subset selection, etc., to obtain a set of feature variables that are significantly correlated with the vegetation canopy height data collected near the ground.
[0049] Step 4: Construct a random forest learning model with optimal model parameters. Calculate the relative height parameters of the preprocessed spaceborne lidar data from multiple historical moments using the random forest learning model with optimal model parameters to obtain the calibrated vegetation canopy height of each footprint point in the spaceborne lidar data at each historical moment.
[0050] Step 4, which involves constructing a random forest learning model with optimal parameters, is detailed below:
[0051] A random forest learning model is constructed by sequentially inputting each element from a set of feature variables significantly correlated with near-ground collected vegetation canopy height data as a sample to predict the vegetation canopy height. A loss function model is then constructed using near-ground collected vegetation canopy height data as the true label. Root mean square error and coefficient of determination are used as evaluation metrics. The optimal random forest learning model with the best parameters is obtained through optimized training. Figure 2 As shown.
[0052] Step 5: Acquire passive optical remote sensing data, synthetic aperture radar data, interferometric radar data, and terrain data from multiple historical moments. Passive optical remote sensing data includes, but is not limited to, Landsat series satellite data, Sentinel series satellite data, PALSAR series satellite data, and other continuous active and passive remote sensing data. Through radar remote sensing data preprocessing methods, obtain preprocessed passive optical remote sensing data, synthetic aperture radar data, interferometric radar data, and terrain data from multiple historical moments.
[0053] The radar remote sensing data preprocessing method described in step 5 is as follows:
[0054] The following steps are performed sequentially: radiometric correction, missing value processing, resampling, stitching and segmentation, and projection conversion.
[0055] Step 6: Calculate the set of remote sensing vegetation indices and the set of annual phenological indices using preprocessed passive optical remote sensing data from multiple historical moments. Combine the preprocessed passive optical remote sensing data, synthetic aperture radar data, and interferometric radar data from multiple historical moments to calculate the set of annual image statistics. Construct a set of terrain feature indices using terrain data from multiple historical moments, and further construct a set of non-spatial feature vectors.
[0056] The remote sensing vegetation index set described in step 6 consists of various types of remote sensing vegetation index indicators, including commonly used remote sensing indices (EVI, EVI2, RVI), tassel cap transformation related indicators (brightness, humidity, greenness, humidity-greenness difference), etc.
[0057] The annual phenological index set mentioned in step 6 consists of various types of annual phenological indexes, including basic phenological parameters such as the NDVI value at the beginning of summer and the NDVI value during the leaf-fall season.
[0058] The annual image statistical value set mentioned in step 6 consists of various types of annual image statistical value indicators, including basic statistical indicators such as maximum value, minimum value, average value, standard deviation, and median value.
[0059] The set of terrain feature indicators described in step 6 consists of various types of terrain feature indicators, including slope, aspect, and elevation.
[0060] Step 6, which involves constructing a non-spatial feature vector set, is detailed below:
[0061] The remote sensing vegetation index set, annual phenological index set, annual image statistical value set, and terrain feature index set are sequentially filtered using variable screening methods. Variables that are significantly correlated with the calibrated vegetation canopy height of each footprint point in the spaceborne lidar data at each historical moment, as described in step 5, are screened. Commonly used variable screening methods include random forest, minimum shrinkage and selection operator, recursive feature elimination, LightGBM-based selection model, stepwise regression model, principal component analysis, subset selection, etc., to obtain the filtered remote sensing vegetation index set, the filtered annual phenological index set, the filtered annual image statistical value set, and the filtered terrain feature index set.
[0062] Step 7: Construct a spatial weight matrix by analyzing the spatial distance relationships between the coordinates of footprint points in the preprocessed spaceborne lidar data from multiple historical time points. Center the spatial weight matrix and calculate its eigenvalues and eigenvectors. Arrange the eigenvectors of the spatial weight matrix in descending order of their corresponding eigenvalues to obtain the sorted eigenvectors of the spatial weight matrix. Spatial weight matrices are divided into two categories: distance-based weight matrices and topology-based weight matrices. Distance-based weight matrices primarily target the locations of spaceborne lidar footprint points and can use Gaussian, exponential, double-square, or triple-cubic weight generation functions. Topology-based weight matrices primarily target raster data of relevant ground information acquired by remote sensing sensors and can use adjacency methods such as Rook and Queen to construct the weight matrix.
[0063] Step 8: Filter the eigenvectors of the sorted spatial weight matrices whose eigenvalues are greater than a threshold to construct an initial set of spatial eigenvectors; combine the non-spatial eigenvector set with the initial set of spatial eigenvectors to construct an environmental variable set.
[0064] Step 9: Construct an optimized lightweight gradient boosting machine learning model. The optimized lightweight gradient boosting machine learning model is used to calculate the predicted value of forest canopy height by combining preprocessed passive optical remote sensing data, synthetic aperture radar data, and interferometric radar data from multiple historical moments.
[0065] The optimized lightweight gradient boosting machine learning model described in step 9 is as follows:
[0066] A lightweight gradient boosting machine learning model is constructed. The set of environmental variables is used as samples to predict the height of the vegetation canopy. The height of the vegetation canopy after calibration at each footprint point of the satellite lidar data at each historical moment is used as the true label to construct a loss function model. The root mean square error and the coefficient of determination are used as evaluation indicators. The optimized lightweight gradient boosting machine learning model is obtained through optimization training.
[0067] A specific embodiment of the present invention also provides a computer-readable medium.
[0068] The computer-readable medium is a server workstation;
[0069] The server workstation stores a computer program executed by the electronic device. When the computer program runs on the electronic device, it causes the electronic device to perform the steps of the forest canopy height estimation method of the inter-filter value according to the present invention.
[0070] It should be understood that any parts not described in detail in this specification belong to the prior art.
[0071] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.
Claims
1. A method for estimating forest canopy height using spatial filtering, characterized in that: A candidate variable set was constructed, and variable screening methods were used to screen variables that were significantly correlated with the vegetation canopy height data collected near the ground, so as to obtain a set of feature variables that were significantly correlated with the vegetation canopy height data collected near the ground. A set of terrain feature indicators is constructed using terrain data from multiple historical moments, and a set of non-spatial feature vectors is further constructed. Construct a spatial weight matrix, center the spatial weight matrix and calculate the eigenvalues and eigenvectors of the spatial weight matrix, and arrange the eigenvectors of the spatial weight matrix in descending order of their corresponding eigenvalues to obtain the sorted eigenvectors of the spatial weight matrix. The eigenvectors of the sorted spatial weight matrices whose eigenvalues are greater than a threshold are filtered to construct an initial set of spatial eigenvectors; the non-spatial eigenvector set and the initial set of spatial eigenvectors are combined to form an environmental variable set. An optimized lightweight gradient boosting machine learning model was constructed to calculate the predicted forest canopy height using preprocessed passive optical remote sensing data, synthetic aperture radar data, and interferometric radar data from multiple historical moments.
2. The method for estimating forest canopy height using spatial filtering according to claim 1, characterized in that, Includes the following steps: Step 1: Obtain the coordinates and relative height parameters of the footprint points of the spaceborne lidar data in the forest area at multiple historical moments. Then, through projection transformation, outlier processing, and noise point removal, obtain the preprocessed coordinates and relative height parameters of the footprint points of the spaceborne lidar data at multiple historical moments. Step 2: Acquire near-ground vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data. Through projection transformation and matching, obtain vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data that are the same as the coordinates of the footprint points of the preprocessed spaceborne lidar data from multiple historical moments. Step 3: Construct a candidate variable set by combining the relative height parameter data of the spaceborne lidar data from multiple historical moments, land use classification data, vegetation cover data, and topographic data that are the same as the coordinates of the footprint points of the preprocessed spaceborne lidar data from multiple historical moments. Using the variable screening method, filter the variables that are significantly related to the vegetation canopy height data collected near the ground to obtain a set of feature variables that are significantly related to the vegetation canopy height data collected near the ground. Step 4: Construct a random forest learning model with optimal model parameters. Calculate the relative height parameters of the preprocessed spaceborne lidar data from multiple historical moments using the random forest learning model with optimal model parameters to obtain the calibrated vegetation canopy height of each footprint point in the spaceborne lidar data at each historical moment. Step 5: Acquire passive optical remote sensing data, synthetic aperture radar data, interferometric radar data, and terrain data from multiple historical moments. Using radar remote sensing data preprocessing methods, obtain preprocessed passive optical remote sensing data, synthetic aperture radar data, interferometric radar data, and terrain data from multiple historical moments. Step 6: Calculate the set of remote sensing vegetation indices and the set of annual phenological indices using preprocessed passive optical remote sensing data from multiple historical moments. Combine the preprocessed passive optical remote sensing data, synthetic aperture radar data, and interferometric radar data from multiple historical moments to calculate the set of annual image statistics. Construct a set of terrain feature indices using terrain data from multiple historical moments, and further construct a set of non-spatial feature vectors. Step 7: Construct a spatial weight matrix by using the spatial distance relationship between the coordinates of the footprint points of the preprocessed spaceborne lidar data from multiple historical moments. Then, center the spatial weight matrix and calculate its eigenvalues and eigenvectors. Arrange the eigenvectors of the spatial weight matrix in descending order of their corresponding eigenvalues to obtain the sorted eigenvectors of the spatial weight matrix. Step 8: Filter the eigenvectors of the sorted spatial weight matrices whose eigenvalues are greater than a threshold to construct an initial set of spatial eigenvectors; combine the non-spatial eigenvector set with the initial set of spatial eigenvectors to construct an environmental variable set. Step 9: Construct an optimized lightweight gradient boosting machine learning model. The optimized lightweight gradient boosting machine learning model is used to calculate the predicted value of forest canopy height by combining preprocessed passive optical remote sensing data, synthetic aperture radar data, and interferometric radar data from multiple historical moments.
3. The forest canopy height estimation method based on spatial filter values according to claim 2, characterized in that: Step 2, the projection transformation matching, is as follows: The vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data are projected and transformed to the coordinate system corresponding to the coordinate positions of the spaceborne lidar data footprint points. These coordinates are then matched with the coordinate positions of the preprocessed spaceborne lidar data footprint points from multiple historical moments as described in step 1 to obtain vegetation canopy height data, land use classification data, vegetation coverage data, and topographic data that are identical to the coordinate positions of the preprocessed spaceborne lidar data footprint points from multiple historical moments.
4. The forest canopy height estimation method based on spatial filter values according to claim 3, characterized in that: Step 4, which involves constructing a random forest learning model with optimal parameters, is detailed below: A random forest learning model is constructed by taking each element of the feature variables that are significantly related to the vegetation canopy height data collected near the ground as the sample input in turn, predicting the vegetation canopy height result, and constructing a loss function model by combining the vegetation canopy height data collected near the ground as the true label. The root mean square error and the coefficient of determination are used as evaluation indicators, and the random forest learning model with the best model parameters is obtained through optimization training.
5. The forest canopy height estimation method based on spatial filter values according to claim 4, characterized in that: The radar remote sensing data preprocessing method described in step 5 is as follows: The following steps are performed sequentially: radiometric correction, missing value processing, resampling, stitching and segmentation, and projection conversion.
6. The forest canopy height estimation method based on spatial filter value according to claim 5, characterized in that: The remote sensing vegetation index set described in step 6 consists of various types of remote sensing vegetation index indicators; The annual phenological index set described in step 6 consists of various types of annual phenological indicators; The annual image statistics set mentioned in step 6 consists of various types of annual image statistics indicators; The set of terrain feature indicators described in step 6 consists of various types of terrain feature indicators; Step 6, which involves constructing a non-spatial feature vector set, is detailed below: The remote sensing vegetation index set, annual phenological index set, annual image statistical value set, and terrain feature index set are sequentially filtered using variable screening methods. Variables that are significantly correlated with the calibrated vegetation canopy height of each footprint point in the spaceborne lidar data at each historical moment, as described in step 5, are then filtered to obtain the filtered remote sensing vegetation index set, the filtered annual phenological index set, the filtered annual image statistical value set, and the filtered terrain feature index set.
7. The forest canopy height estimation method based on spatial filter values according to claim 6, characterized in that: The optimized lightweight gradient boosting machine learning model described in step 9 is as follows: A lightweight gradient boosting machine learning model is constructed. The set of environmental variables is used as samples to predict the height of the vegetation canopy. The height of the vegetation canopy after calibration at each footprint point of the satellite lidar data at each historical moment is used as the true label to construct a loss function model. The root mean square error and the coefficient of determination are used as evaluation indicators. The optimized lightweight gradient boosting machine learning model is obtained through optimization training.
8. A computer-readable medium, characterized in that, It stores a computer program executed by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Model for depicting forest photosynthetically active radiation distribution by utilizing three-dimensional point cloud data
CN112415537A
Snow space-time analysis and prediction method based on random forest
CN114972984A