A forest canopy height estimation method, system, terminal and storage medium

By combining LSTM regression models with GEDI and Sentinel-1SAR data, the problem of insufficient accuracy and efficiency in forest canopy height estimation in existing technologies has been solved, achieving more efficient and accurate forest structure monitoring.

CN119478656BActive Publication Date: 2025-12-26GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411358140.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-12-26
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

Existing technologies have failed to fully utilize the potential of deep learning models in InSAR multi-temporal coherence data analysis, resulting in insufficient accuracy and efficiency in forest canopy height estimation. Furthermore, relying on ground-based measured data is time-consuming and labor-intensive, and cannot comprehensively cover and reflect the dynamic changes in forest structure in a timely manner.

Method used

By combining LSTM regression model with GEDI and Sentinel-1SAR data, outliers and non-forest areas are removed through preprocessing to generate coherent multi-temporal sequences. A two-layer LSTM regression model is then used for training, reducing reliance on ground-based measured data and improving estimation accuracy and efficiency.

Benefits of technology

It improves the accuracy and efficiency of forest canopy height estimation, reduces reliance on ground-based measurement data, lowers data collection costs and difficulties, and better captures dynamic changes in forest structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478656B_ABST
    Figure CN119478656B_ABST
Patent Text Reader

Abstract

The application discloses a forest canopy height estimation method, system, terminal and storage medium, and the method comprises the steps of: obtaining GEDI data of a forest for pretreatment, removing outliers and eliminating non-forest areas, and preparing a data set according to non-forest samples and forest samples; obtaining SAR image data to form an interferogram, generating a coherence image, generating a coherence upper triangular matrix of multiple time phases according to the coherence image, and converting the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-time phase sequence; inputting the processed coherence multi-time phase sequence as a characteristic variable into an LSTM regression model, training the LSTM regression model using the data set; inputting image data of a forest to be identified into the trained LSTM regression model, and outputting the forest canopy height of the forest to be identified. The application improves the accuracy, reliability and efficiency of forest height estimation, and makes forest monitoring more efficient and extensive.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer application, and particularly relates to a forest canopy height estimation method and system, a terminal and a computer readable storage medium. BACKGROUND

[0002] Forests, as a vital terrestrial ecosystem on Earth, play a decisive role in global carbon cycling. Forest canopy height, as one of the forest structure parameters, is an important indicator for studying aboveground biomass and a key indicator for studying forest productivity and biodiversity. With the development of remote sensing technology, especially the application of synthetic aperture radar (SAR) technology, forest structure can be more accurately monitored and evaluated. SAR technology has been widely used in forest structure research due to its high resolution, ability to work day and night, and insensitivity to atmospheric conditions. In particular, interferometric synthetic aperture radar (InSAR) and polarimetric InSAR (PolInSAR) are recognized as effective means for measuring forest canopy height due to their deep penetration in long-wave and cross-polarization modes. In addition, InSAR coherence shows great potential in land cover classification and vegetation classification, which can significantly improve the classification performance. The change of InSAR coherence is not only related to different land cover types, but also closely related to changes in the land surface, such as natural processes of vegetation growth.

[0003] In the field of forest canopy height measurement, existing technologies mainly focus on using SAR data, especially data obtained from satellite systems such as Sentinel-1 and TanDEM-X. These technologies use time series analysis of SAR images to analyze forest structure changes, with less application of InSAR coherence. In existing research, SAR data, especially C-band SAR data, combined with ground reference data, have been used for forest height prediction. Existing technologies have made some progress in using InSAR data and GEDI (Global Ecosystem Dynamics Investigation) measurement data to study forest canopy height, such as analyzing InSAR data through random forest regression models and deep learning networks, and combining Sentinel-1 and Sentinel-2 data to improve the accuracy of land cover classification and forest height estimation.

[0004] However, the existing technology still has some limitations. First, traditional InSAR-based methods usually rely on ground truth data for calibration and validation, which is not only time-consuming and labor-intensive, but also may not be able to obtain comprehensive coverage due to the limitations of ground conditions. Second, existing technology often ignores the dynamic changes of InSAR coherence when processing SAR data, which limits its application in capturing seasonal changes in forest structure. In addition, existing methods may have problems in data fusion and consistency when combining multi-source remote sensing data for forest structure inversion, affecting the final estimation accuracy. These limitations indicate that the existing technology may not fully utilize the potential of deep learning models in analyzing InSAR multi-temporal coherence data, especially in continuous monitoring and prediction of forest canopy height, and more advanced methods are needed to improve the accuracy and efficiency of forest height estimation.

[0005] Therefore, the existing technology still needs to be improved and developed. SUMMARY

[0006] The main purpose of the present application is to provide a forest canopy height estimation method, system, terminal and computer readable storage medium, which aims to solve the problem that the existing technology fails to fully utilize the potential of deep learning models in analyzing InSAR multi-temporal coherence data, and cannot improve the accuracy and efficiency of forest height estimation.

[0007] To achieve the above-mentioned purpose, the present application provides a forest canopy height estimation method, which comprises the following steps:

[0008] Obtain GEDI data of the forest, preprocess the GEDI data, remove outliers and exclude non-forest areas, and prepare a data set according to non-forest samples and forest samples;

[0009] Obtain SAR image data, form an interferogram according to the SAR image data, generate a coherence image according to the interferogram, generate a multi-temporal coherence upper triangular matrix according to the coherence image, and convert the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-temporal sequence;

[0010] Input the processed coherence multi-temporal sequence as a feature variable into an LSTM regression model, train the LSTM regression model using the data set, and obtain a trained LSTM regression model;

[0011] Obtain image data of the forest to be identified, input the image data of the forest to be identified into the trained LSTM regression model, and output the forest canopy height of the forest to be identified.

[0012] Optionally, the forest canopy height estimation method, wherein the GEDI data of the forest is obtained, the GEDI data is preprocessed, outliers are removed and non-forest areas are removed, a data set is made according to non-forest samples and forest samples, and specifically comprises:

[0013] GEDIL2A height data in the same time range as the coherence data is obtained through the NASA Global Ecosystem Dynamics Investigation project;

[0014] Land cover types are distinguished through a global land cover map, outliers in the GEDIL2A height data are removed, and intermediate GEDIL2A height data is obtained;

[0015] Non-forest samples and forest samples in the intermediate GEDIL2A height data are selected and used as a data set of a classification model.

[0016] Optionally, the forest canopy height estimation method, wherein the GEDI data of the forest is obtained, the GEDI data is preprocessed, outliers are removed and non-forest areas are removed, a data set is made according to non-forest samples and forest samples, and then further comprises:

[0017] Multispectral images of a Sentinel-2 satellite are obtained, and a vegetation index is calculated according to the multispectral images;

[0018] The vegetation index is used as a feature vector, and a classification model is constructed based on a support vector machine algorithm, the classification model is used to filter out non-forest areas and only keep GEDI data points within the detection range that are all forest;

[0019] A relative height index is extracted based on GEDI echoes, and the relative height index is used as the real forest canopy height.

[0020] Optionally, the forest canopy height estimation method, wherein the vegetation index comprises an enhanced vegetation index, a normalized vegetation index, a red edge vegetation index, and a difference vegetation index.

[0021] Optionally, the forest canopy height estimation method, wherein the SAR image data is obtained, an interferogram is formed according to the SAR image data, a coherence image is generated according to the interferogram, a coherence upper triangular matrix of multiple time phases is generated according to the coherence image, and each sampling point corresponding to the coherence upper triangular matrix is converted into a coherence multi-time phase sequence, and specifically comprises:

[0022] Sentinel-1A SAR image data with a revisit time of a preset time is obtained, the Sentinel-1A SAR image data is paired for interference two by two to form an interferogram, and a corresponding coherence image is generated using a multi-window method;

[0023] wherein the calculation of the interferometric coherence is:

[0024]

[0025] wherein γ represents the interferometric coherence, with a value ranging from 0 to 1, S1 and S2 represent two registered co-polarized SAR complex images, E{·} represents the mathematical expectation, * represents the complex conjugate operator, |S1| represents the amplitude of S1, and |S2| represents the amplitude of S2;

[0026] From the coherent image point set, select the coherence points in the resolution range where the GEDI footprint exists;

[0027] After screening, arrange each sampling point in order of interferometric interval time to generate a coherence upper triangular matrix of multiple time phases;

[0028] If there are multiple GEDI footprints, the tree height values of the multiple GEDI footprints are averaged according to the distance to determine the GEDI tree height value of the coherence point where the GEDI footprint exists;

[0029] wherein each sampling point has a coherence upper triangular matrix of multiple time phases and a corresponding GEDI tree height value;

[0030] In the coherence upper triangular matrix, the values with the same interferometric time interval are distributed in the diagonal line, and for the vegetation area, the relatively high coherence is maintained near the main diagonal line;

[0031] Extract the upper triangular part of the coherence upper triangular matrix along the diagonal line, convert the upper triangular part into a coherence sequence with different time spans of unequal lengths, and convert the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-time phase sequence.

[0032] Optionally, the forest canopy height estimation method, wherein the processed coherence multi-time phase sequence is input as a feature variable into the LSTM regression model, the LSTM regression model is trained using the data set, and a trained LSTM regression model is obtained, specifically comprising:

[0033] The processed coherence multi-time phase sequence is input as a feature variable into the LSTM regression model, and the prediction result of the LSTM regression model is verified using the preprocessed GEDI sampling data;

[0034] The dataset is proportionally divided into a training set, a validation set and a test set, the training set is used to train the LSTM regression model, so that the LSTM regression model learns the mapping relationship from the coherence multi-temporal sequence to the canopy height, the validation set is used to adjust the parameters and hyperparameters of the LSTM regression model in the training process, to avoid overfitting, and the test set is independent of the training process, and is used for final evaluation of the generalization ability of the LSTM regression model.

[0035] When the training of the LSTM regression model is completed using the training set and the validation set, the test set is used to test the LSTM regression model, and the trained LSTM regression model is obtained after meeting the test requirements.

[0036] Optionally, the forest canopy height estimation method, wherein the LSTM regression model is a double-layer LSTM regression model, each layer contains 256 units, and uses ReLU as an activation function, and a Dropout layer is added.

[0037] The loss function of the LSTM regression model training is mean square error, calculated as:

[0038]

[0039] Wherein, MSE represents the mean square error, y i represents the actual canopy height of the i-th sample, represents the canopy height of the i-th sample predicted by the model, and N represents the number of samples.

[0040] The evaluation indexes R square and root mean square error RMSE are used as the judgment standard for the accuracy evaluation of the final model prediction canopy height, calculated as:

[0041]

[0042] Wherein, R 2 represents the evaluation index R square, and RMSE represents the root mean square error, represents the average value of the canopy height.

[0043] In addition, in order to achieve the above purpose, the present application also provides a forest canopy height estimation system, wherein the forest canopy height estimation system comprises:

[0044] The data processing module is used for acquiring GEDI data of the forest, pre-processing the GEDI data, removing outliers and excluding non-forest areas, and preparing a dataset according to non-forest samples and forest samples.

[0045] The sequence generation module is configured to obtain SAR image data, form an interferogram according to the SAR image data, generate a coherence image according to the interferogram, generate a coherence upper triangular matrix of multiple time phases according to the coherence image, and convert the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-time phase sequence.

[0046] The model training module is configured to input the processed coherence multi-time phase sequence into an LSTM regression model as a feature variable, train the LSTM regression model using the data set, and obtain a trained LSTM regression model.

[0047] The height estimation module is configured to obtain image data of a forest to be identified, input the image data of the forest to be identified into the trained LSTM regression model, and output a forest canopy height of the forest to be identified.

[0048] In addition, to achieve the above object, the present application also provides a terminal, wherein the terminal comprises a memory, a processor, and a forest canopy height estimation program stored in the memory and executable on the processor, and the forest canopy height estimation program implements the steps of the forest canopy height estimation method when executed by the processor.

[0049] In addition, to achieve the above object, the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a forest canopy height estimation program, and the forest canopy height estimation program implements the steps of the forest canopy height estimation method when executed by a processor.

[0050] In the present application, GEDI data of a forest is obtained, the GEDI data is preprocessed, outliers are removed and non-forest areas are excluded, a data set is prepared according to non-forest samples and forest samples, SAR image data is obtained, an interferogram is formed according to the SAR image data, a coherence image is generated according to the interferogram, a coherence upper triangular matrix of multiple time phases is generated according to the coherence image, and the coherence upper triangular matrix corresponding to each sampling point is converted into a coherence multi-time phase sequence. The processed coherence multi-time phase sequence is input into an LSTM regression model as a feature variable, the LSTM regression model is trained using the data set, a trained LSTM regression model is obtained, image data of a forest to be identified is obtained, the image data of the forest to be identified is input into the trained LSTM regression model, and a forest canopy height of the forest to be identified is output. The present application can better capture the dynamic changes of forest structure by combining the changes of InSAR coherence and the accurate measurement of GEDI, improve the accuracy, reliability and efficiency of forest height estimation, reduce the dependence on ground measurement data, reduce the cost and difficulty of data collection, and make forest monitoring more efficient and extensive. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 is a flowchart of a preferred embodiment of the forest canopy height estimation method of the present application;

[0052] Figure 2 is a schematic diagram of the process of forest canopy height estimation based on InSAR time series coherence and LSTM in the preferred embodiment of the forest canopy height estimation method of the present application;

[0053] Figure 3 is a structural diagram of a preferred embodiment of the forest canopy height estimation system of the present application;

[0054] Figure 4 is a structural diagram of a preferred embodiment of the terminal of the present application. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical scheme and advantages of the present application clearer and more explicit, the present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0056] The present application proposes a new method to estimate forest canopy height by using InSAR coherence or its change and training a deep learning model, Long Short-Term Memory (LSTM), with the forest canopy height values measured by NASA Global Ecosystem Dynamics Investigation (GEDI), in view of the limitations of the prior art in forest canopy height estimation. The LiDAR instrument carried by GEDI has provided accurate forest canopy height data since 2018, which can be used as an alternative source of ground truth data. The prior art usually relies on ground measurement data for calibration, which has the problems of low efficiency and high cost, and may not be able to reflect the dynamic changes of forest structure in a timely manner. In addition, the prior art also faces challenges in processing SAR data with irregular time intervals. By combining the change of InSAR coherence and the accurate measurement of GEDI, the dynamic changes of forest structure can be better captured, and the accuracy and reliability of the estimation can be improved. Secondly, this method can reduce the dependence on ground measurement data, reduce the cost and difficulty of data collection, and make forest monitoring more efficient and extensive.

[0057] The forest canopy height estimation method according to the preferred embodiment of the present application, as shown in Figure 1 and Figure 2 , comprises the following steps:

[0058] Step S10, obtain GEDI data of the forest, preprocess the GEDI data, remove outliers and exclude non-forest areas, and make a data set according to non-forest samples and forest samples.

[0059] Specifically, first is the GEDI data acquisition and screening strategy, the GEDI L2A elevation data (i.e. spaceborne laser radar height measurement data GEDI L2A level data) in the same time range as the coherence data is obtained through the NASA global ecosystem dynamics survey project; GEDI data is divided into different levels, from the most original data to the finished product data after more processing and analysis. Among them, the L1A level data contains the original GEDI laser radar waveform data, which has not been subjected to any geographic positioning or correction. The L1B level data is the waveform data containing geographic positioning, which is the result of preliminary processing of L1A data, providing geographic coordinate information of the waveform. The L2A level data required in the method of the present application is generated by analyzing the peak value and distribution of L1B waveform data, containing elevation and height indicators derived from geographically positioned waveforms. These data provide ground elevation and relative height indicators related to tree canopy structure. These data provide more detailed information on surface and vegetation structure, suitable for vegetation height, biomass estimation and other applications. GEDI contains three lasers and a full-wave recording lidar instrument working in the near-infrared region. A single laser is launched along the orbit at an interval of 60 meters and has a circular detection area of about 25 meters. Therefore, each data point represents a GEDI waveform return indicator for every 25-meter diameter. Since GEDI data provides more accurate vegetation information, it is used as a data source to verify the prediction accuracy. For each data point, the relative height indicator RH100 based on the L2A processing algorithm is extracted, i.e. the 100th percentile of the energy return height relative to the ground, representing the height difference between the highest point of reflection (usually the top of the tree canopy) and the lowest mode detected reflection (usually the ground). This value can be considered as an approximation of the canopy height. However, GEDI measurement data of cities, bare land and water areas are unreliable, and non-forest areas usually exhibit high coherence, resulting in prediction errors. Therefore, it is necessary to preprocess the GEDI data to remove these outliers.

[0060] The land cover types are distinguished by a global land cover map (for example, a global land cover (GLC) map provided by the European Space Agency), and the outliers in the GEDIL2A elevation data are removed, that is, most of the outliers are removed, to obtain intermediate GEDIL2A elevation data; the non-forest samples and forest samples in the intermediate GEDIL2A elevation data are selected and used as a data set of a classification model. That is, since the GLC image cannot completely exclude all non-forest areas in the detection area of the GEDI sampling points, a large number of non-forest surfaces remain in the screened area, which directly affects the purity of the data and the accuracy of subsequent analysis. Therefore, it is necessary to further remove the non-forest areas from the remaining GEDI point set. In order to achieve this, non-forest and forest samples can be manually selected as a data set of a classification model.

[0061] The Sentinel-2 satellite has 13 spectral bands, of which four bands (B2, B3, B4 and B8) with a spatial resolution of 10 meters correspond to different vegetation reflection characteristics. Using the reflection values of these bands, some commonly used vegetation indices can be calculated, such as the enhanced vegetation index (EVI), the normalized difference vegetation index (NDVI), the red edge vegetation index (RVI) and the difference vegetation index (DVI). The present application uses the B2, B3, B4 and B8 bands of Sentinel-2 and the NDVI, EVI, RVI and DVI calculated according to the four bands as feature variables, and constructs a classification model based on the support vector machine (SVM) algorithm. The classification model is used to filter out non-forest areas and only keep GEDI data points that are all forest within the detection range (the GLC map can only do a rough screening based on the latitude and longitude of the data points, and cannot remove the non-forest areas within the detection range of the data points. Therefore, after screening the GLC map, an SVM is used here for a fine screening, and the combination of the two ensures that the data points are all in forest areas within the detection range). In addition, considering the possibility of abnormal forest canopy height in the pre-processed data set, the GEDI data is further screened. For example, GEDI cannot accurately predict the height of sparse vegetation below 3 meters. Therefore, the forest canopy height of the GEDI sample with RH100<3 meters is assigned a value of 0.

[0062] The relative height indicator is extracted based on GEDI echoes (GEDI echoes refer to signals received by the GEDI lidar sensor, after the laser pulse of GEDI is launched from the International Space Station and hits the Earth's surface, part of the light is reflected back. These reflected signals carry information about the characteristics of the Earth's surface, including the vertical structure of vegetation and the terrain of the ground) and taken as the true forest canopy height. The relative height indicator is extracted based on the GEDI echoes, mainly by analyzing the reflection pattern in the lidar waveform, determining the highest and lowest reflection points in the waveform, and then calculating the height difference between the vegetation canopy and the ground.

[0063] In step S20, SAR image data is acquired, an interferometric image is formed according to the SAR image data, a coherence image is generated according to the interferometric image, a coherence upper triangular matrix of multiple time phases is generated according to the coherence image, and a coherence upper triangular matrix corresponding to each sampling point is converted into a coherence multi-time phase sequence.

[0064] Specifically, the multi-view coherence image generation and processing take the Sentinel-1A data obtained from the European Space Agency's Copernicus Global Earth Observation Project as one of the data sources, and the revisit period is 12 days. The Sentinel-1A single look complex (SLC) image adopts a cross-polarization (Vertical polarization Horizontal polarization, VV) and co-polarization (Vertical polarization, VV) polarization mode, and the ground resolution is 5 meters x 20 meters in the interferometric wide swath (IW) mode. The co-polarization VV data of Sentinel-1A is selected for calculation. In order to obtain coherence data, Sentinel-1A SAR image data with a revisit time of 12 days is obtained, the Sentinel-1A SAR image data is paired for interference two by two to form an interferometric image, and a corresponding coherence image is generated using a multi-view window method of 20x5 pixels.

[0065] wherein the interferometric coherence is a normalized cross-correlation coefficient between two registered SAR images, and the value range is 0 to 1, and the calculation formula is as follows:

[0066]

[0067] wherein γ represents the interferometric coherence, the value range is 0 to 1, S1 and S2 represent two registered co-polarization SAR complex images, E{·} represents mathematical expectation, * represents complex conjugate operator, |S1| represents the amplitude of S1, and |S2| represents the amplitude of S2.

[0068] From the coherence image point set, select the coherence points with GEDI footprints in the resolution range; after screening, arrange each sampling point in order according to the interference interval time to generate a multi-temporal coherence upper triangular matrix; if there are multiple GEDI footprints, the tree height values of multiple GEDI footprints are averaged according to the distance weighting to determine the GEDI tree height value of the coherence point with the GEDI footprint; therefore, after the above screening strategy, each sampling point has a multi-temporal coherence upper triangular matrix and a corresponding GEDI tree height value.

[0069] Considering that in the coherence upper triangular matrix, the values with the same interference time interval are distributed in the diagonal line, and for the vegetation area, the relatively high coherence remains near the main diagonal line; therefore, the upper triangular part of the coherence upper triangular matrix is extracted along the diagonal line, and the upper triangular part is converted into a coherence sequence with different time spans of unequal lengths, so as to convert the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-temporal sequence.

[0070] In order to make the sequence length consistent, equal-length pre-padding processing is performed on these features, and time coding is ingeniously added to fully utilize these information in analysis. At the same time, the masking layer is used to mark the padding value as a mask, so as to ensure that the model ignores these padding values when calculating. In this way, each sequence not only maintains its original time dynamic characteristics, but also can be uniformly input into the model for processing.

[0071] Step S30, input the processed coherence multi-temporal sequence as a feature variable into the LSTM regression model, train the LSTM regression model using the data set, and obtain the trained LSTM regression model.

[0072] Specifically, the forest canopy height estimation model (i.e. Figure 2The LSTM regression model in the application, LSTM, Long Short-Term Memory, first, the processed coherence multi-temporal sequence is input into the LSTM regression model as a characteristic variable, and the preprocessed GEDI sampling data is used to verify the prediction result of the LSTM regression model. Second, the data set is divided into a training set (70%), a verification set (15%) and a test set (15%) in proportion, the training set is used to train the LSTM regression model, so that the LSTM regression model learns the mapping relationship from the coherence multi-temporal sequence to the canopy height, the verification set is used to adjust the parameters and hyperparameters of the LSTM regression model in the training process to avoid overfitting, and the test set is independent of the training process and is used to finally evaluate the generalization ability of the LSTM regression model. When the training of the LSTM regression model is completed using the training set and the verification set, the test set is used to test the LSTM regression model, and after meeting the test requirements, the trained LSTM regression model is obtained.

[0073] Among them, the LSTM regression model is a double-layer LSTM regression model, each layer contains 256 units, and uses ReLU as the activation function, while adding a Dropout layer to prevent overfitting, randomly discarding the output of a part of neurons, and enhancing the generalization ability of the model. Finally, model training and testing are performed, the optimizer in the training process is selected as Adam, which automatically adjusts the learning rate and helps the model converge faster. And use the early stopping strategy to monitor the loss change of the verification set, stop training in advance when the loss on the verification set no longer improves, and prevent overfitting.

[0074] The loss function of the LSTM regression model training is the mean squared error (MSE), which is calculated as:

[0075]

[0076] Among them, MSE represents the mean squared error, y i represents the actual canopy height of the i-th sample, represents the canopy height predicted by the model for the i-th sample, and N represents the number of samples.

[0077] The input shape of the model is determined by the time step and the number of features of the training data. In addition, the evaluation indicators R square and root mean square error RMSE((Root Mean Square Error, RMSE)) are used as the judgment standard for the accuracy evaluation of the final model prediction canopy height, which is calculated as:

[0078]

[0079] Among them, R 2R-square and RMSE represents the root mean square error, The average value of the canopy height is represented.

[0080] Step S40, obtain the image data of the forest to be identified, input the image data of the forest to be identified into the trained LSTM regression model, and output the forest canopy height of the forest to be identified.

[0081] Finally, when the trained LSTM regression model is obtained after the model training is completed, the trained LSTM regression model can be directly used to directly estimate the forest canopy height of the forest to be identified, and the accuracy and reliability of the forest canopy height estimation are improved.

[0082] The beneficial effects of the present application are:

[0083] (1) A detailed GEDI data acquisition and preprocessing strategy is proposed, GEDI L2A level data of NASA is used, relative height index RH100 is extracted to obtain accurate forest canopy height. At the same time, GLC map and SVM algorithm are introduced to combine Sentinel-2 extracted data to assist in screening forest area, and non-forest surface interference data is filtered out to ensure the purity of GEDI data.

[0084] (2) By analyzing the multi-temporal coherence of Sentinel-1A SAR data, including extracting the upper triangular part of the coherence matrix and adding time coding, and performing equal-length pre-padding processing on the unequal-length sequence, these feature engineering steps enhance the model's ability to capture time dynamics, which helps to capture the dynamic changes of forest structure.

[0085] (3) A double-layer LSTM regression model is used, the processed coherence multi-temporal sequence is used as an input feature, and GEDI sampling data is used for prediction by training the forest canopy height estimation model. And through the early stopping strategy and Dropout layer, overfitting is prevented, and the generalization ability of the model is improved.

[0086] Further, the Sentinel-1 InSAR data used in the present application can consider using other types of InSAR data such as TanDEM-X, ALOS-2 PALSAR-2 or future SAR satellite systems, which may have different wave bands, resolutions or revisit periods, which may affect the performance of the model; the LSTM model can be replaced by other types of deep learning models such as Gated Recurrent Unit (GRU), Transformer or Convolutional LSTM, which may have different advantages in processing time series data; after the model prediction, spatial smoothing or filtering techniques can be used to reduce prediction noise, or fused with other remote sensing data (such as optical images) to further improve the estimation accuracy of forest canopy height; at the same time, this method of the present application can also be applied to various scenarios such as ground biomass estimation, land cover type change, etc.; in addition, transfer learning can also be tried, that is, first pre-train the model on a large-scale data set, and then fine-tune it on the data of a specific region to improve the adaptability of the model in a new region.

[0087] Further, as shown in Figure 3 based on the above forest canopy height estimation method, the present application also correspondingly provides a forest canopy height estimation system, wherein the forest canopy height estimation system comprises:

[0088] The data processing module 51 is configured to obtain GEDI data of a forest, pre-process the GEDI data, remove outliers and exclude non-forest areas, and prepare a data set according to non-forest samples and forest samples.

[0089] The sequence generation module 52 is configured to obtain SAR image data, form an interferogram according to the SAR image data, generate a coherence image according to the interferogram, generate a coherence upper triangular matrix of multiple time phases according to the coherence image, and convert the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-time phase sequence.

[0090] The model training module 53 is configured to input the processed coherence multi-time phase sequence as a feature variable into an LSTM regression model, train the LSTM regression model using the data set, and obtain a trained LSTM regression model.

[0091] The height estimation module 54 is configured to obtain image data of a forest to be identified, input the image data of the forest to be identified into the trained LSTM regression model, and output the forest canopy height of the forest to be identified.

[0092] Further, as shown in Figure 4As shown, based on the above forest canopy height estimation method and system, the application also correspondingly provides a terminal, which comprises a processor 10, a memory 20 and a display 30. Figure 4 Only some components of the terminal are shown, but it should be understood that all the shown components are not required, and more or less components can be alternatively implemented.

[0093] The memory 20 can be an internal storage unit of the terminal in some embodiments, such as a hard disk or a memory of the terminal. The memory 20 can also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Further, the memory 20 can include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software and various data installed on the terminal, such as program codes of the terminal, etc. The memory 20 can also be used to temporarily store data that has been output or will be output. In an embodiment, the memory 20 stores a forest canopy height estimation program 40, which can be executed by the processor 10 to implement the forest canopy height estimation method in the application.

[0094] The processor 10 can be a central processing unit (CPU), a microprocessor or other data processing chip in some embodiments, which is used to run program codes or process data stored in the memory 20, such as to execute the forest canopy height estimation method, etc.

[0095] The display 30 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, etc. in some embodiments. The display 30 is used to display information of the terminal and to display visualized user interfaces. The components 10-30 of the terminal communicate with each other through a system bus.

[0096] In an embodiment, the processor 10 implements the steps of the forest canopy height estimation method when executing the forest canopy height estimation program 40 in the memory 20.

[0097] The application also provides a computer readable storage medium, wherein the computer readable storage medium stores a forest canopy height estimation program, which implements the steps of the forest canopy height estimation method as described above when executed by a processor.

[0098] In summary, the present application provides a forest canopy height estimation method, system, terminal and storage medium, the method comprising: obtaining GEDI data of a forest, preprocessing the GEDI data, removing outliers and excluding non-forest areas, and preparing a dataset according to non-forest samples and forest samples; obtaining SAR image data, forming an interferogram according to the SAR image data, generating a coherence image according to the interferogram, generating a coherence upper triangular matrix of multiple time phases according to the coherence image, and converting the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-time phase sequence; inputting the processed coherence multi-time phase sequence as a characteristic variable into an LSTM regression model, training the LSTM regression model using the dataset, and obtaining a trained LSTM regression model; obtaining image data of a forest to be identified, inputting the image data of the forest to be identified into the trained LSTM regression model, and outputting the forest canopy height of the forest to be identified. The present application can better capture the dynamic changes of forest structure by combining the changes of InSAR coherence and the accurate measurement of GEDI, improve the accuracy, reliability and efficiency of forest height estimation, reduce the dependence on ground measurement data, reduce the cost and difficulty of data collection, and make forest monitoring more efficient and extensive.

[0099] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or terminals including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles or terminals. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or terminal including the element.

[0100] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable computer-readable storage medium, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.

[0101] It should be understood that the application of the present application is not limited to the above examples, and those skilled in the art can improve or modify the above description, all of which should be within the scope of the appended claims of the present application.

Claims

1. A method for estimating forest canopy height, characterized in that, The forest canopy height estimation method comprises: Obtaining GEDI data of a forest, preprocessing the GEDI data, removing outliers and excluding non-forest areas, and preparing a data set according to non-forest samples and forest samples; Obtaining SAR image data, forming an interferogram according to the SAR image data, generating a coherence image according to the interferogram, generating a coherence upper triangular matrix of multiple time phases according to the coherence image, and converting the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-time phase sequence; The processed coherence multi-time phase sequence is input into an LSTM regression model as a characteristic variable, the LSTM regression model is trained using the data set, and a trained LSTM regression model is obtained; Obtaining image data of a forest to be identified, inputting the image data of the forest to be identified into the trained LSTM regression model, and outputting the forest canopy height of the forest to be identified; Obtaining Sentinel-1A SAR image data with a preset revisit time, pairing the Sentinel-1A SAR image data two by two for interference to form an interferogram, and generating a corresponding coherence image using a multi-window method; The calculation of the interference coherence is: ; wherein, denotes the interferometric coherence, taking values in the range 0 to 1, S 1 and S 2 denote two registered co-polarized SAR complex images, E {·} denotes the mathematical expectation, * denotes the complex conjugate operator, S 1 | denotes S 1 the amplitude of, S 2 | denotes S 2 the amplitude of. From the coherence image point set, select the coherence points with GEDI footprints within the resolution range; After screening, arrange each sampling point in order according to the interference interval time to generate a coherence upper triangular matrix of multiple time phases; If there are multiple GEDI footprints, the tree height values of the multiple GEDI footprints are averaged according to the distance to determine the GEDI tree height value of the coherence point with the GEDI footprint; Each sampling point has a coherence upper triangular matrix of multiple time phases and a corresponding GEDI tree height value; In the coherence upper triangular matrix, the values with the same interference time interval are diagonally distributed, and for vegetation areas, the relatively high coherence remains near the main diagonal line; Extract the upper triangular part of the coherence upper triangular matrix along the diagonal line, convert the upper triangular part into a coherence sequence with different time spans of different lengths, and convert the coherence upper triangular matrix corresponding to each sampling point into a coherence multi-time phase sequence; The LSTM regression model is a double-layer LSTM regression model, each layer contains 256 units, and uses ReLU as the activation function, while adding a Dropout layer; The loss function of the LSTM regression model training is the mean square error, which is calculated as: ; where MSE represents mean square error, represents the actual canopy height of the i th sample, represents the canopy height predicted by the model for the i th sample, N represents the number of samples; The evaluation indicators R square and root mean square error RMSE are used as the judgment standard for the accuracy evaluation of the final model prediction canopy height, which is calculated as: ; ; wherein, R2denotes the evaluation indicator R square sum, RMSE denotes the root mean square error, R2denotes the average value of the canopy height.

2. The forest canopy height estimation method of claim 1, wherein, The GEDI data of the forest is obtained, the GEDI data is preprocessed, outliers are removed and non-forest areas are excluded, and a data set is prepared according to non-forest samples and forest samples, which specifically comprises: Obtaining GEDI L2A elevation data of the same time range as the coherence data through the NASA Global Ecosystem Dynamics Investigation project; Distinguish land cover types through global land cover maps, remove outliers in the GEDI L2A elevation data, and obtain intermediate GEDI L2A elevation data; Select non-forest samples and forest samples in the intermediate GEDI L2A elevation data as a data set of a classification model.

3. The forest canopy height estimation method of claim 2, wherein, The GEDI data of the forest is acquired, the GEDI data is preprocessed, outliers are removed and non-forest areas are removed, a data set is made according to the non-forest samples and the forest samples, and then the method further comprises: A multispectral image of a Sentinel-2 satellite is acquired, and a vegetation index is calculated according to the multispectral image; The vegetation index is taken as a feature vector, and a classification model is constructed based on a support vector machine algorithm, the classification model being used to filter out non-forest areas and only keep GEDI data points in the detection range that are all forest; A relative height index is extracted based on GEDI echoes, and the relative height index is taken as a true forest canopy height.

4. The forest canopy height estimation method of claim 3, wherein, The vegetation index comprises an enhanced vegetation index, a normalized vegetation index, a red edge vegetation index and a difference vegetation index.

5. The forest canopy height estimation method of claim 1, wherein, The processed coherent multi-temporal sequence is input as a feature variable into an LSTM regression model, and the data set is used to train the LSTM regression model to obtain a trained LSTM regression model, and the method specifically comprises: The processed coherent multi-temporal sequence is input as a feature variable into an LSTM regression model, and the prediction result of the LSTM regression model is verified using the preprocessed GEDI sampling data; The data set is divided into a training set, a validation set and a test set in proportion, the training set is used to train the LSTM regression model, so that the LSTM regression model learns a mapping relationship from the coherent multi-temporal sequence to the canopy height, the validation set is used to adjust parameters and hyperparameters of the LSTM regression model in the training process to avoid overfitting, and the test set is independent of the training process and is used to finally evaluate the generalization ability of the LSTM regression model; After the training set and the validation set are used to complete the training of the LSTM regression model, the test set is used to test the LSTM regression model, and a trained LSTM regression model is obtained after the test requirement is met.

6. A forest canopy height estimation system characterized by, The forest canopy height estimation system is used to implement the forest canopy height estimation method of any one of claims 1-5, and the forest canopy height estimation system comprises: A data processing module is configured to acquire GEDI data of a forest, pre-process the GEDI data, remove outliers and non-forest areas, and make a data set according to non-forest samples and forest samples; A sequence generation module is configured to acquire SAR image data, form an interferogram according to the SAR image data, generate a coherence image according to the interferogram, generate a coherent upper triangular matrix of multiple time phases according to the coherence image, and convert the coherent upper triangular matrix corresponding to each sampling point into a coherent multi-temporal sequence; A model training module is configured to input the processed coherent multi-temporal sequence as a feature variable into an LSTM regression model, train the LSTM regression model using the data set, and obtain a trained LSTM regression model. The height estimation module is configured to acquire image data of a forest to be identified, input the image data of the forest to be identified into the trained LSTM regression model, and output a forest canopy height of the forest to be identified.

7. A terminal, characterized by comprising: The terminal comprises a memory, a processor, and a forest canopy height estimation program stored on the memory and executable on the processor, and the forest canopy height estimation program, when executed by the processor, implements the steps of the forest canopy height estimation method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a forest canopy height estimation program, and the forest canopy height estimation program, when executed by a processor, implements the steps of the forest canopy height estimation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Remote sensing data feature conversion method and system based on image processing, and intelligent terminal

    CN117876802A