Large-scale forest height remote sensing inversion method and system based on dynamic time sequence feature optimization

By constructing a remote sensing inversion method for forest height optimized by dynamic temporal features, the problems of spectral saturation and lack of historical information in tall forest stands have been solved, and high-precision estimation of forest height at large scale has been achieved, especially significantly improving the estimation accuracy in areas of tall forest stands and forest stands in the recovery period.

CN121811269APending Publication Date: 2026-04-07HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies cannot solve the problems of spectral saturation in tall forest stands and accurate estimation of historical forest stands with complex interference through a dynamic, adaptive, and physically meaningful feature system. In particular, they suffer from insufficient accuracy and lack of historical information in large-scale forest height monitoring.

Method used

A remote sensing inversion method for forest height based on dynamic temporal feature optimization is constructed. By acquiring spaceborne lidar and multi-temporal optical remote sensing image data, a composite temporal feature set is constructed, including temporal spectral stability and perturbation recovery trajectory index. The optimal time length and feature set are adaptively selected, a machine learning model is trained, and a spatially continuous forest height distribution map is generated.

Benefits of technology

It significantly improved the accuracy of forest height estimation, especially in tall forest stands and low/restored forest stands, with accuracy improvements of over 35% and 45% respectively, enabling large-scale, high-precision forest height monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811269A_ABST
    Figure CN121811269A_ABST
Patent Text Reader

Abstract

The invention provides a large-scale forest height remote sensing inversion method and system based on dynamic time sequence feature optimization. The method comprises the following steps: acquiring discrete point-like forest canopy height truth value data measured by a satellite-borne laser radar in a research area; acquiring a multi-temporal optical remote sensing image set spanning at least a preset time length in the research area, and performing preprocessing and annual median synthesis to generate a plurality of annual time sequence data sets with different time lengths; for each annual time sequence data set, constructing a composite time sequence feature set pixel by pixel, and training a preset machine learning regression model by taking the corresponding composite time sequence feature set as input and taking discrete point forest canopy height true value data as output to obtain a corresponding forest height prediction model; and evaluating and comparing the precision of the model under all time durations, and based on a precision optimal principle, determining an optimal time duration and a composite time sequence feature set so as to generate a forest height distribution diagram with continuous space in the research area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of remote sensing information processing and forestry monitoring technology, and in particular to a method and system for large-scale forest height remote sensing inversion based on dynamic temporal feature optimization. Background Technology

[0002] Forest canopy height is a key parameter for characterizing the three-dimensional structure of forests, assessing forest biomass, estimating ecosystem carbon sinks, and studying biodiversity. Traditional ground survey methods, such as using altimeters, total stations, or laser rangefinders to measure individual trees, while highly accurate, are time-consuming, labor-intensive, costly, and difficult to implement on large regional or even global scales.

[0003] The emergence of satellite remote sensing technology has made it possible to achieve large-scale, periodic forest height monitoring. Currently, the mainstream forest height remote sensing acquisition schemes are mainly divided into two categories: active lidar measurement and passive optical remote sensing inversion. Spaceborne lidar systems, such as ICESat-2, directly measure the distance from the ground surface to the canopy using laser pulses, enabling the acquisition of high-precision forest height information. However, due to limitations in satellite orbit design, its data is discrete and sparsely distributed, making it impossible to directly generate spatially continuous, high-resolution forest height products. Optical remote sensing imagery, such as Landsat and Sentinel-2, has significant advantages in terms of spatial continuity and long temporal series. Therefore, the current mainstream research direction is to establish statistical models between optical remote sensing characteristics and lidar altimetry values ​​to achieve spatially continuous estimation of forest height.

[0004] These methods have roughly gone through the following stages: Parametric models: Early studies typically assumed a linear relationship between forest height and spectral indices, texture features, etc., and used parametric methods such as multiple regression to build models. These models are simple in form, but they are difficult to characterize the complex nonlinear relationship between forest structure and remote sensing signals, and they have poor universality and are prone to overfitting.

[0005] Non-parametric machine learning models: With the development of machine learning, non-parametric methods such as random forests and neural networks have been widely used. They can effectively handle high-dimensional features and nonlinear relationships, significantly improving estimation accuracy. However, most of these methods rely on static spectral, textural, topographic, and climatic features averaged over a single or multiple time period. In dense, tall forests, these optical features are prone to "spectral saturation," meaning that signals such as vegetation indices tend to saturate as canopy height increases, leading to a systematic underestimation of high-value tree heights. More importantly, static features cannot capture historical information about forest growth, succession, and natural disturbances (such as fires, windfalls, and logging). For stands that have been disturbed and are in different recovery stages, estimating based solely on a "snapshot" of the current state will inevitably result in significant bias.

[0006] Models that take spatial heterogeneity into account: To improve the regional adaptability of models, some studies have introduced spatial units such as ecological zones and forest types, attempting to reduce the impact of environmental heterogeneity through zonal modeling. While this approach has made progress in handling spatial differentiation, the features it uses are essentially static and fail to solve the fundamental problems of spectral saturation and lack of historical information.

[0007] Furthermore, while a few studies have attempted to utilize time-series data, these have primarily focused on forest cover change detection or phenological analysis, typically employing fixed time window lengths when used for height estimation. Forest growth cycles and disturbance recovery histories vary considerably, and the optimal response length for fast-growing plantations and virgin old-growth forests to time information clearly differs. Fixed time windows cannot adapt to the dynamic characteristics of different forest ecosystems, making it difficult to optimally extract historical information most relevant to the current canopy height, thus limiting further improvements in model accuracy.

[0008] In summary, existing technologies share a common core deficiency: they cannot simultaneously solve the problem of spectral saturation in tall forest stands and the problem of accurately estimating complex and disturbed historical forest stands through a dynamic, adaptive, and physically meaningful feature system. Summary of the Invention

[0009] To address the dual limitations of existing forest height estimation methods—namely, that static features cannot reflect the dynamic history of forests and that fixed time windows cannot adapt to different ecosystems—this invention provides a large-scale forest height remote sensing inversion method and system based on dynamic temporal feature optimization. This invention aims to fundamentally solve the spectral saturation problem of tall forest stands by constructing a novel "feature-time window adaptive optimization framework," and accurately capture the recovery state of disturbed forest stands, thereby achieving high-precision, adaptive estimation of large-scale forest height with spatial continuity.

[0010] In a first aspect, the present invention provides a large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization, comprising: Step 1: Obtain true-value data of discrete point forest canopy heights obtained from spaceborne lidar measurements within the study area; Step 2: Acquire a set of multi-temporal optical remote sensing images spanning at least a preset duration within the study area; Step 3: Preprocess and synthesize the annual median of the multi-temporal optical remote sensing image set to generate multiple annual time series datasets with different time lengths; Step 4: For each annual time series dataset of the specified time length, construct a composite time series feature set pixel by pixel. The composite time series feature set includes a time series spectral stability index and a time series perturbation recovery trajectory index. Step 5: For each annual time series dataset of the specified time length, take the corresponding composite time series feature set as input and the discrete point forest canopy height ground value data as output to train the preset machine learning regression model to obtain the corresponding forest height prediction model. Step 6: Evaluate and compare the accuracy of the forest height prediction models for all time lengths. Based on the principle of optimal accuracy, adaptively determine the best time length for inversion and its corresponding composite time series feature set. Step 7: Using the composite time-series feature set under the optimal time length and the corresponding forest height prediction model, generate a spatially continuous forest height distribution map of the study area.

[0011] Furthermore, in step 1, the spaceborne lidar is an ICESat-2 / ATL08 product; Correspondingly, in step 5, before training, the discrete point-like forest canopy height ground truth data is subjected to quality control; the quality control process includes: strong and weak beam filtering, day and night photon filtering, photon number threshold filtering, ICESat-2 / ATL08 terrain height data filtering based on SRTM data, and non-forest area masking based on land cover products.

[0012] Furthermore, in step 2, the multi-temporal optical remote sensing image set is the Landsat series image; Correspondingly, in step 3, the preprocessing includes: performing radiometric normalization on the surface reflectance data from different sensors to eliminate spectral inconsistencies caused by sensor differences.

[0013] Furthermore, in step 3, the Medoid multidimensional median synthesis method is used to synthesize the annual median, and the forest growing season of each year is used as the image time window during synthesis.

[0014] Furthermore, in step 4, when constructing the time-series spectral stability index, the spectral data used includes at least three of the following: visible light band, near-infrared band, short-wave infrared band, normalized difference vegetation index (NDVI), enhanced vegetation index (EVI), and cap change brightness index (TCB).

[0015] Furthermore, the time-series spectral stability index includes skewness, which reflects the asymmetry of the spectral distribution, and interquartile range, which resists the influence of outliers.

[0016] Furthermore, in step 4, when constructing the time-series disturbance recovery trajectory index, the LandTrendr time series segmentation algorithm is used to extract the index.

[0017] Furthermore, the temporal disturbance recovery trajectory indicators include: the maximum recovery magnitude and the disturbance recovery signal-to-noise ratio.

[0018] Furthermore, in step 6, cross-validation is used to evaluate the accuracy of the forest height prediction model. The accuracy evaluation metrics used include the coefficient of determination R² and the root mean square error RMSE. Correspondingly, the principle of optimal accuracy specifically includes: prioritizing the time length corresponding to the forest height prediction model with the smallest RMSE as the optimal time length; when the RMSE difference between forest height prediction models under multiple time lengths is less than a preset threshold, the shortest time length is selected as the optimal time length.

[0019] Secondly, the present invention provides a large-scale forest height remote sensing inversion system based on dynamic temporal feature optimization, comprising: The forest height acquisition module is used to acquire true data of discrete point forest canopy height obtained from spaceborne lidar measurements within the study area; The remote sensing image acquisition module is used to acquire a set of multi-temporal optical remote sensing images of the study area spanning at least a preset duration. The data preprocessing and synthesis module is used to preprocess and synthesize the multi-temporal optical remote sensing image set to generate multiple annual time series datasets of different time lengths. The composite time series feature construction module is used to construct a composite time series feature set pixel by pixel for each annual time series dataset of the time length. The composite time series feature set includes a time series spectral stability index and a time series perturbation recovery trajectory index. The model training and adaptive optimization module is used to train a preset machine learning regression model for each annual time series dataset of the specified time length, taking the corresponding composite time series feature set as input and the discrete point forest canopy height ground truth data as output, to obtain the corresponding forest height prediction model; and to evaluate and compare the accuracy of the corresponding forest height prediction model under all time lengths, and adaptively determine the best time length for inversion and its corresponding composite time series feature set based on the principle of optimal accuracy. The forest height mapping module is used to generate a spatially continuous forest height distribution map of the study area using the composite temporal feature set under the optimal time length and the corresponding forest height prediction model.

[0020] The beneficial effects of this invention are as follows: Overcoming the problem of spectral saturation: By introducing time-series spectral stability indices, the "instantaneous snapshot" of a single time phase is expanded into a "dynamic portrait" reflecting long-term spectral behavior. Indices such as skewness and interquartile range can effectively capture subtle spectral changes in tall and dense forest stands over long-term scales, breaking the saturation relationship between static characteristics and tree height.

[0021] Precise quantification of historical impact: The introduction of the temporal disturbance recovery trajectory index transforms the key ecological process of forest disturbance and recovery history into quantifiable model input features. This enables the model to understand the degree of disturbance and the duration of recovery that led to the current tree height, thereby achieving unprecedentedly accurate estimation of disturbed forest stands.

[0022] Achieving Adaptive and Precise Optimization: By comparing model accuracy under variable time lengths, this invention can find the optimal observation history length driven by data from different regions and forest types. This avoids the subjectivity and limitations of manually setting fixed windows, ensuring that the model can access the most relevant historical information in any scenario, maximizing both global and local accuracy. Experiments have demonstrated that the method of this invention significantly improves overall accuracy compared to models based on static features (specifically, RMSE is reduced by more than 35%). In particular, in traditionally challenging areas such as tall stands and low / recovery stands, the estimation accuracy improvement can reach over 46% and 45%, respectively.

[0023] Advantages of automation and engineering: The entire process of this invention can be highly automated and deployed on cloud computing platforms such as Google Earth Engine, enabling efficient and rapid production of large-scale, high-precision, and spatially continuous forest height products, which has significant application value. Attached Figure Description

[0024] Figure 1 This is one of the flowcharts illustrating a large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization provided in an embodiment of the present invention; Figure 2 The second flowchart illustrates a large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization provided in this embodiment of the invention. Figure 3 Description of a single forest pixel Landsat time series trajectory provided in this embodiment of the invention: (a) pixel values ​​extracted from the time series image; (b) an example of time series TCB spectral value fitting and segmentation based on the Landtrendr algorithm, where A, B, C, and D are segmentation breakpoints; a and d represent the spectral variation range of perturbation and recovery (labeled MAG), respectively, and b and c represent the duration of perturbation and recovery states (labeled DUR), respectively. Figure 4 This is a graph showing the variation of model accuracy under different time window lengths provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of the structure of a large-scale forest height remote sensing inversion system based on dynamic temporal feature optimization, provided in an embodiment of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0026] This invention provides a large-scale forest height remote sensing inversion method and system based on dynamic temporal feature optimization. The core of this method lies in addressing the spectral saturation and historical information loss problems caused by the use of static features and fixed time windows in existing methods. The method includes: acquiring ground truth tree height values ​​from spaceborne lidar and long-term multi-temporal optical remote sensing images; generating multiple annual sequence datasets of different time lengths; innovatively constructing a composite temporal feature set for each time length, which couples a "temporal spectral stability index" measuring long-term spectral stability from a statistical dynamics perspective with a "temporal disturbance recovery trajectory index" quantifying the maximum disturbance recovery event from an ecological process perspective; using the composite feature set as input and the ground truth tree height as output, training multiple machine learning models in parallel; adaptively determining the optimal time series length by comparing the model accuracy across all time lengths; and finally, using the features and models at the optimal length to achieve spatially continuous estimation of forest height. This invention effectively overcomes the spectral saturation problem of tall forest stands by using the composite temporal feature system and adaptive optimization mechanism, and accurately characterizes the recovery state of disturbed forest stands. Compared with traditional static feature methods, it significantly improves the estimation accuracy of forest height, especially high and low extreme values.

[0027] like Figure 1 As shown, this embodiment of the invention provides a large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization, including the following steps: S101: Acquire true-value data of discrete point forest canopy heights measured by spaceborne lidar (such as ICESat-2 / ATL08 products) within the study area; Specifically, subsequent processes require the use of measured discrete point-like ground-value data of forest canopy height to train machine learning models. However, the quality of directly obtained ground-value height data varies, which can affect model accuracy. Therefore, to ensure the accuracy of the trained model, quality control of the directly obtained ground-value height data should be implemented in practical applications. In this embodiment, a multi-level serial quality control process is adopted, including: strong and weak beam filtering, day and night photon filtering, photon number threshold filtering, ICESat-2 / ATL08 terrain height data filtering based on SRTM data, and non-forest area masking based on land cover products, thereby ensuring the purity and reliability of training samples.

[0028] S102: Acquire a multi-temporal optical remote sensing image set spanning at least a preset time period (e.g., 30 years) within the study area; Specifically, in this embodiment, Landsat series images within the study area are collected to construct a multi-temporal optical remote sensing image set. The Landsat series images include TM, ETM+, and OLI sensor data.

[0029] S103: Preprocess and synthesize the annual median of the multi-temporal optical remote sensing image set to generate multiple annual time series datasets with different time lengths; Specifically, firstly, since data from different sensors differ, the preprocessing includes radiometric normalization to eliminate spectral inconsistencies caused by sensor differences; in addition, it includes applying cloud and cloud shadow detection algorithms (such as CFMask) to remove invalid pixels.

[0030] Next, the Medoid multidimensional median composite method was used to synthesize images of the forest growing season each year (e.g., May-October) to best represent the interannual spectral conditions.

[0031] Finally, by truncating the long-term time series data, a set of annual time series datasets with different time lengths are generated. These include various scales such as 5 years, 10 years, 15 years, 20 years, 25 years, and 30 years, to cover various forest life processes from short-term disturbance recovery to long-term ecological succession.

[0032] S104: For each annual time series dataset of the specified time length, construct a composite time series feature set pixel by pixel. The composite time series feature set includes a time series spectral stability index and a time series perturbation recovery trajectory index. Specifically, time-series spectral stability indices are used to quantify the overall spectral statistical characteristics and stability of pixels over time. These indices can include skewness and interquartile range calculated from vegetation indices (preferably NDVI, EVI, and TCB) or spectral band values ​​(such as visible, near-infrared, and short-wave infrared bands) in the time series; they can also include the mean and standard deviation calculated from vegetation indices or spectral band values ​​in the time series. Skewness is used to capture potential trends in forest growth or decline; interquartile range is used to resist interference from outliers (such as residual clouds and aerosols); and the mean and standard deviation characterize central tendency and dispersion. Together, these indices reflect the long-term health status of the forest and the stability of biomass accumulation.

[0033] The temporal perturbation recovery trajectory index is used to quantify the physical quantities and rhythm of the maximum perturbation event and subsequent recovery process experienced by a pixel. In this embodiment, the LandTrendr time series segmentation algorithm is used to extract the index, decomposing the spectral trajectory into a series of continuous straight line segments, thereby accurately extracting the physical parameters of the perturbation and recovery events. The key parameters of the LandTrendr algorithm are set as follows: the maximum number of segments (maxSegments) is 4 to 8, the peak threshold (spikeThreshold) is 0.7 to 0.95, and the p-value threshold (pvalThreshold) is 0.01 to 0.1. This index can be divided into three categories: perturbation features, recovery features, and process signal-to-noise ratio (SNR). Among them, perturbation features include the maximum perturbation magnitude and the maximum perturbation duration; recovery features include the maximum recovery magnitude and the maximum recovery duration; and process SNR includes the perturbation-recovery SNR, which is used to measure the significance of the event in the entire time trajectory. These indices directly record the "life events" of the forest and are closely related to the formation mechanism of the current canopy height.

[0034] S105: For each annual time series dataset of the specified time length, take the corresponding composite time series feature set as input, take the discrete point forest canopy height ground value data as output, train the preset machine learning regression model (preferably random forest, with the number of decision trees set to not less than 2000) to obtain the corresponding forest height prediction model. S106: Evaluate and compare the accuracy of the forest height prediction models for all time lengths, and adaptively determine the optimal time length for inversion and its corresponding composite time series feature set based on the principle of optimal accuracy. Specifically, k-fold cross-validation can be used to evaluate each model, and the accuracy evaluation metrics include both the coefficient of determination R² and the root mean square error RMSE.

[0035] The principle of optimal accuracy specifically includes: prioritizing the time length corresponding to the forest height prediction model with the smallest RMSE as the optimal time length; when the RMSE difference between forest height prediction models under multiple time lengths is less than a preset threshold (such as less than 5%), then the shortest time length is selected as the optimal time length.

[0036] S107: Using the composite temporal feature set under the optimal time length and the corresponding forest height prediction model, generate a spatially continuous forest height distribution map of the study area.

[0037] Specifically, forest height is estimated for each pixel in the entire study area to generate a final, spatially continuous forest height distribution map.

[0038] The large-scale forest height remote sensing inversion method provided in this invention is a multi-level, closed-loop feedback technical solution. The core of this solution is to construct a dynamic feature system that integrates long-term spectral behavior and physical disturbance history, and to introduce a data-driven adaptive optimizer to find the optimal feature set and observation duration combination for a specific region.

[0039] To verify the effectiveness of the present invention, the entire forest area of ​​China was used as the research area. The method of the present invention was implemented on the Google Earth Engine (GEE) cloud computing platform. The overall process is as follows: Figure 2 As shown.

[0040] S201: Data Preparation and Preprocessing Forest height ground truth data: Download the 2019-2020 ICESat-2 ATL08 product (version 5) data and perform the following cascaded filtering process to obtain high-precision tree height sample points (h). canopy ): beam type Filter out data other than strong beams to ensure signal quality.

[0041] night flag Prioritize nighttime data (value 1) to reduce solar background noise.

[0042] Filter weak signal spots with fewer than 10 photons.

[0043] The terrain slope at the location of each photon is calculated based on the SRTM DEM, and potential error points with a slope greater than 10° are eliminated.

[0044] Using the global 30-meter land cover product (FROM-GLC), sample points were excluded from non-forest categories (such as farmland, water bodies, and cities).

[0045] In the end, approximately 1.65 million high-quality forest height sample points and their latitude and longitude coordinates were obtained.

[0046] Optical remote sensing time series data: Collect all Landsat TM / ETM+ / OLI surface reflectance data from 1986 to 2020.

[0047] Preprocessing: The coefficients proposed by Roy et al. were used to normalize the TM / ETM+ data to the OLI sensor reference, eliminating system spectral differences. A data quality layer (QA Band) generated using the CFMask algorithm was used to mask cloud, cloud shadow, and cirrus cloud pixels to complete cloud removal.

[0048] Annual composite: All available imagery from May 1st to October 31st each year (the main growing season in the Northern Hemisphere) is selected. The Medoid composite method is employed. This method calculates the median vector of all pixels within the growing season in a multi-dimensional space composed of multiple bands, and selects the observation with the closest Euclidean distance to this median vector as the representative pixel for that year. This method is more resistant to outliers than the ordinary median method and preserves the physical relationships between bands.

[0049] Construct a variable-length time series dataset: The time series is truncated backwards with 2020 as the endpoint. Six annual synthetic image datasets with lengths of 5 years (2016-2020), 10 years (2011-2020), 15 years (2006-2020), 20 years (2001-2020), 25 years (1996-2020), and 30 years (1991-2020) are constructed respectively.

[0050] S202: Calculation of Composite Time Series Features For the six time-series datasets mentioned above, two types of features are calculated pixel by pixel.

[0051] (1) Calculation of temporal spectral stability index: Three spectral quantities, Green band, NDVI and TCB (cap change brightness), were selected for index calculation.

[0052] NDVI sequence within an n-year time window For example, calculate the following indicators: average value: Standard deviation: Skewness: Quartiles (25th and 75th percentiles): Q1, Q3 For each spectral quantity, the above four indicators are calculated, resulting in a total of 3 (spectral quantities) × 4 (indicators) = 12 time-series spectral stability indicators.

[0053] (2) Extraction of time-series disturbance recovery trajectory indicators: The LandTrendr algorithm was called in GEE. Key parameters were set as follows: maximum number of segments maxSegments=6, peak threshold spikeThreshold=0.9, p-value threshold pvalThreshold=0.05, recovery value threshold recoveryThreshold=0.25. These parameters were optimized through pre-experimentation and can effectively capture the main disturbance recovery events without overfitting the noise. Computational basis and indicator extraction: The algorithm fits the time series of the TCB index, and the fitting curve is shown in Figure 1. Figure 3As shown. The most significant disturbance-recovery event cycle is extracted from the fitting results, and the following metrics are calculated: Maximum interference level (Mag) D ): Mag D = Value_End Vertex - Value_Start Vertex Maximum disturbance duration (Dur) D ): Dur D = Year_End Vertex - Year_Start Vertex Maximum recovery level (Mag) R ): This represents the difference between the spectral value at the end of the recovery process and the spectral value at the beginning of the recovery process; i.e., Mag R = Value_End Vertex - Value_Start Vertex Perturbation recovery signal-to-noise ratio (SNR): , where RMSE is the root mean square error between the LandTrendr fitted curve and the original TCB sequence.

[0054] A total of four temporal perturbation recovery trajectory indicators were extracted. Therefore, for each pixel at each time length, the total dimension of the generated composite temporal feature set is 12 + 4 = 16 features.

[0055] S203: Model Training and Adaptive Optimization Sample matching: Using the latitude and longitude of ICESat-2 sample points obtained in S201, feature vectors for corresponding locations are extracted from feature sets at six different time lengths, and compared with the tree height ground truth h. canopy Pairing them up creates six training datasets.

[0056] Model training: For each time period of the dataset, a random forest regression model (decision tree = 5000) is used, and the dataset is divided into training and test sets at a ratio of 70% and 30% for hold-out validation. The out-of-bag error is used to monitor the model performance during the training process.

[0057] Accuracy assessment and optimal window selection: Calculate the coefficient of determination (R²) for each model on the test set, as shown in the figure. Figure 4 As shown.

[0058] Adaptive Decision Making: Based on the above results, increasing the time series length can improve estimation accuracy. The optimal time series length does not vary across different ecoregions. In some ecoregions (e.g., i24), the improvement is small, with R² increasing from 0.33 to 0.35, a mere 5.71% improvement. However, in other ecoregions (e.g., i1 to i4), the improvement is significant, with R² values ​​increasing from 0.34 to 0.73, a 53.42% improvement. Therefore, the optimal time series length varies depending on the relationship between the FCH and the category of the modeled feature and the characteristics of the ecoregion. Specifically, applications requiring information on how the FCH estimation accuracy changes over time may benefit from choosing the shortest possible time series length that meets the accuracy requirements. If the research goal is to obtain the most accurate estimate under the current model conditions, then the time series length with the highest accuracy can be selected.

[0059] S204: Forest Height Mapping and Effect Verification Global mapping: The 16-dimensional composite time-series features calculated for each 30-meter pixel across China within a 20-year time window are input into a pre-trained optimal random forest model to estimate forest height values ​​and generate a spatially continuous map of China's forest height distribution in 2020.

[0060] Performance Validation: The present invention (Method C) is compared with Method A (comparative example: using ecological zoning and static spectral, topographic, and climatic features) and Method B (comparative example: using a fixed 20-year window of time-series spectral stability indicators, but without perturbation recovery trajectory indicators). On an independent validation set (approximately 500,000 ICESat-2 sample points not used in training), the accuracy comparison of the three methods is as follows: As can be seen from the table above, compared with existing methods, the method of the present invention has the highest accuracy in the whole region, tall forest stands, and low / recovery forest stands.

[0061] Based on the same inventive concept, such as Figure 5 As shown, this embodiment of the invention provides a large-scale forest height remote sensing inversion system based on dynamic temporal feature optimization, including: a forest height acquisition module, a remote sensing image acquisition module, a data preprocessing and synthesis module, a composite temporal feature construction module, a model training and adaptive optimization module, and a forest height mapping module.

[0062] Specifically, the forest height acquisition module is used to acquire discrete point-like true-value data of forest canopy height obtained from spaceborne lidar measurements within the study area; the remote sensing image acquisition module is used to acquire a multi-temporal optical remote sensing image set spanning at least a preset duration within the study area; the data preprocessing and synthesis module is used to preprocess and synthesize the multi-temporal optical remote sensing image set to generate multiple annual time series datasets of different time lengths; the composite time series feature construction module is used to construct a composite time series feature set pixel by pixel for each annual time series dataset of the specified time length, the composite time series feature set including a time series spectral stability index and a time series perturbation recovery trajectory index; model training and self-training... The adaptive optimization module is used to train a preset machine learning regression model for each annual time series dataset of the specified time length, taking the corresponding composite time series feature set as input and the discrete point-like forest canopy height ground truth data as output, to obtain the corresponding forest height prediction model; and to evaluate and compare the accuracy of the corresponding forest height prediction models under all time lengths, and adaptively determine the optimal time length for inversion and its corresponding composite time series feature set based on the principle of optimal accuracy; the forest height mapping module is used to generate a spatially continuous forest height distribution map of the study area using the composite time series feature set under the optimal time length and the corresponding forest height prediction model.

[0063] It should be noted that the large-scale forest height remote sensing inversion system provided in this embodiment of the invention is for implementing the above method. Its specific functions can be referred to in the above method embodiments, and will not be repeated here.

[0064] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization, characterized in that, include: Step 1: Obtain true-value data of discrete point forest canopy heights obtained from spaceborne lidar measurements within the study area; Step 2: Acquire a set of multi-temporal optical remote sensing images spanning at least a preset duration within the study area; Step 3: Preprocess and synthesize the annual median of the multi-temporal optical remote sensing image set to generate multiple annual time series datasets with different time lengths; Step 4: For each annual time series dataset of the specified time length, construct a composite time series feature set pixel by pixel. The composite time series feature set includes a time series spectral stability index and a time series perturbation recovery trajectory index. Step 5: For each annual time series dataset of the specified time length, take the corresponding composite time series feature set as input and the discrete point forest canopy height ground value data as output to train the preset machine learning regression model to obtain the corresponding forest height prediction model. Step 6: Evaluate and compare the accuracy of the forest height prediction models for all time lengths. Based on the principle of optimal accuracy, adaptively determine the best time length for inversion and its corresponding composite time series feature set. Step 7: Using the composite time-series feature set under the optimal time length and the corresponding forest height prediction model, generate a spatially continuous forest height distribution map of the study area.

2. The large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization according to claim 1, characterized in that, In step 1, the spaceborne lidar is an ICESat-2 / ATL08 product; Correspondingly, in step 5, before training, the discrete point-like forest canopy height ground truth data is subjected to quality control; the quality control process includes: strong and weak beam filtering, day and night photon filtering, photon number threshold filtering, ICESat-2 / ATL08 terrain height data filtering based on SRTM data, and non-forest area masking based on land cover products.

3. The large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization according to claim 1, characterized in that, In step 2, the multi-temporal optical remote sensing image set is the Landsat series image; Correspondingly, in step 3, the preprocessing includes: performing radiometric normalization on the surface reflectance data from different sensors to eliminate spectral inconsistencies caused by sensor differences.

4. The large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization according to claim 1, characterized in that, In step 3, the Medoid multidimensional median synthesis method is used to synthesize the annual median, and the forest growing season of each year is used as the image time window during the synthesis.

5. The large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization according to claim 1, characterized in that, In step 4, when constructing the time-series spectral stability index, the spectral data used includes at least three of the following: visible light band, near-infrared band, short-wave infrared band, normalized difference vegetation index (NDVI), enhanced vegetation index (EVI), and cap change brightness index (TCB).

6. The large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization according to claim 5, characterized in that, The time-series spectral stability indices include skewness, which reflects the asymmetry of the spectral distribution, and interquartile range, which resists the influence of outliers.

7. The large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization according to claim 1, characterized in that, In step 4, when constructing the time-series disturbance recovery trajectory index, the LandTrendr time series segmentation algorithm is used to extract the index.

8. The large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization according to claim 7, characterized in that, The time-series disturbance recovery trajectory indicators include: maximum recovery magnitude and disturbance recovery signal-to-noise ratio.

9. The large-scale forest height remote sensing inversion method based on dynamic temporal feature optimization according to claim 1, characterized in that, In step 6, cross-validation is used to evaluate the accuracy of the forest height prediction model. The accuracy evaluation metrics used include the coefficient of determination R² and the root mean square error RMSE. Correspondingly, the principle of optimal accuracy specifically includes: prioritizing the time length corresponding to the forest height prediction model with the smallest RMSE as the optimal time length; when the RMSE difference between forest height prediction models under multiple time lengths is less than a preset threshold, the shortest time length is selected as the optimal time length.

10. A large-scale forest height remote sensing inversion system based on dynamic temporal feature optimization, characterized in that, include: The forest height acquisition module is used to acquire true data of discrete point forest canopy height obtained from spaceborne lidar measurements within the study area; The remote sensing image acquisition module is used to acquire a set of multi-temporal optical remote sensing images of the study area spanning at least a preset duration. The data preprocessing and synthesis module is used to preprocess and synthesize the multi-temporal optical remote sensing image set to generate multiple annual time series datasets of different time lengths. The composite time series feature construction module is used to construct a composite time series feature set pixel by pixel for each annual time series dataset of the time length. The composite time series feature set includes a time series spectral stability index and a time series perturbation recovery trajectory index. The model training and adaptive optimization module is used to train a preset machine learning regression model for each annual time series dataset of the specified time length, taking the corresponding composite time series feature set as input and the discrete point forest canopy height ground truth data as output, to obtain the corresponding forest height prediction model; and to evaluate and compare the accuracy of the corresponding forest height prediction model under all time lengths, and adaptively determine the best time length for inversion and its corresponding composite time series feature set based on the principle of optimal accuracy. The forest height mapping module is used to generate a spatially continuous forest height distribution map of the study area using the composite temporal feature set under the optimal time length and the corresponding forest height prediction model.