A forest ecosystem carbon storage accounting method and system
By combining spaceborne lidar data processing with optical remote sensing, the problems of coverage and cost of airborne lidar in large-scale carbon storage monitoring have been solved, enabling efficient and continuous carbon storage accounting, which is suitable for forest carbon storage surveys at the national scale.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 湖南省第二测绘院
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-24
AI Technical Summary
Traditional airborne lidar is limited by its small coverage area, high cost, and poor environmental adaptability when monitoring carbon storage in forest ecosystems. It is difficult to achieve large-scale, efficient, and continuous carbon storage monitoring, and the data standardization is low, which cannot meet the needs of macro-level carbon sink management.
Preprocessing of spaceborne lidar data was performed to extract waveform length and height features. Combined with optical remote sensing image features and a random forest regression model, continuous feature parameters of the study area were generated. Texture factors were extracted using first-order and second-order texture filtering tools. A random forest machine learning model was constructed to predict canopy height and calculate carbon storage.
It enables continuous observation at the regional level, with a single-track coverage width of up to hundreds of kilometers, reducing data acquisition costs, minimizing the impact of terrain and weather, and providing long-term continuous time-series data, which facilitates the analysis of interannual variations in forest carbon pools.
Smart Images

Figure CN121479740B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of carbon monitoring technology, and in particular to a method and system for calculating carbon storage in forest ecosystems. Background Technology
[0002] The current method for calculating forest ecosystem carbon storage is based on UAV laser point cloud data, field sampling data, digital elevation models, sentinel remote sensing data, and land cover data of the monitoring area. Combined with software such as ArcGIS and Origin, the process begins by collecting field plot data using UAV lidar and interpreting it to obtain the true value of the canopy height. Simultaneously, different types of remote sensing data products are generated based on multi-source remote sensing data, and meteorological, temperature, and related geographic information data are collected and used as independent variables. Then, a random forest (RF) model and a stepwise regression model are used to find the independent variable with the highest correlation to the true value of canopy height, forming a multiple linear regression equation. The equation is used to estimate the height of vegetation across the entire area. Then, the allometric growth equation is used to calculate the aboveground biomass (AGB) of the entire vegetation area, and the root-to-stem ratio is used to calculate the belowground biomass (BGB) from the aboveground biomass. Finally, the aboveground and belowground biomass are summed to obtain the total vegetation biomass, which is then multiplied by the carbon content coefficient to obtain the vegetation carbon storage of the forest ecosystem in the monitoring area.
[0003] However, traditional airborne lidar is limited by its coverage area (a single flight can only monitor a local area), high cost (reliant on aircraft leasing, flight permits, etc.), and poor environmental adaptability (highly affected by terrain and weather). This makes it difficult to efficiently conduct large-scale forest parameter monitoring, and its data continuity and standardization are poor, failing to meet the long-term dynamic data requirements of macro-level carbon sink management. Furthermore, airborne lidar data is significantly affected by flight plans, weather (heavy rain, dense fog), and terrain (mountainous areas, canyons), resulting in poor data continuity in long-term dynamic monitoring. The low standardization of data from different batches due to differences in flight parameters makes it difficult to support carbon sink trend analysis. In addition, airborne lidar data has diverse formats and limited scale, requiring complex coordinate transformations and scale matching for fusion with optical remote sensing data. It is also difficult to efficiently integrate with regional auxiliary data (DEM, slope), leading to limited feature dimensions and incomplete forest structure characterization. Summary of the Invention
[0004] To address the above problems, this invention proposes a method for calculating carbon storage in forest ecosystems, the method comprising:
[0005] The acquired spaceborne lidar data is preprocessed, and then spaceborne feature parameters are extracted from the preprocessed data. The spaceborne feature parameters include waveform length features and waveform height features.
[0006] Features extracted from optical remote sensing images that are spatially matched with the location of the spaceborne lidar are used as input variables. Combined with a random forest regression model, the spaceborne feature parameters are extended to the entire study area to generate continuous spaceborne feature parameters for the entire study area.
[0007] Optical remote sensing data of the entire study area is acquired, and the optical remote sensing data is processed to obtain red, green, blue and near-infrared bands. The red, green, blue and near-infrared bands are converted into surface reflectance products, and then multiple vegetation indices are calculated based on the surface reflectance products of each band.
[0008] Based on first-order and second-order probabilistic statistical filtering tools, texture filtering is performed on the red, green, blue and near-infrared bands respectively to extract first-order texture factors and second-order texture factors. The first-order texture factors include data range, mean, variance, information entropy and skewness. The second-order texture factors include mean, variance, coherence, contrast, dissimilarity, information entropy, second moment and correlation.
[0009] The importance of continuous satellite-borne feature parameters, various vegetation indices, first-order texture factors, and second-order texture factors in the entire study area is ranked using the random forest algorithm. The feature with the highest importance ranked by the random forest algorithm is fixed as the first variable. Then, other features are added one by one in order of decreasing importance, and the second and third features are fixed in turn until all features are involved in the calculation. After each new feature is added, the root mean square error of the model is recorded. The combination of variables corresponding to the minimum root mean square error of the model is selected as the optimal feature variable set.
[0010] By combining the optimal set of feature variables with elevation model data and its derived slope data, a random forest machine learning model is constructed and trained to predict and generate continuous distribution results of canopy height throughout the study area.
[0011] The continuously distributed canopy height results are input into the allometric growth equation to estimate the forest biomass of the entire study area. Then, based on the forest system land type and tree species, the forest type and carbon content coefficient are determined. The forest biomass is multiplied by the tree species carbon content coefficient to calculate the carbon storage of different forest types. Finally, the vegetation carbon storage of the entire forest system in the study area is obtained.
[0012] Furthermore, the acquired spaceborne lidar data undergoes preprocessing, specifically including:
[0013] First, the background noise threshold is set by the noise average value and noise standard deviation to denoise the spaceborne lidar data; then, Gaussian filtering is used to smooth the denoised waveform data to remove glitches, and Gaussian decomposition algorithm is used to filter out non-forest information.
[0014] Furthermore, the waveform length features include the waveform start position, waveform end position, waveform total height, first peak position, last peak position, waveform leading edge width, waveform trailing edge width, and peak length; the waveform height features include relative height 25%, relative height 50%, relative height 75%, relative height 95%, relative height 97%, and the difference between relative height 75% and relative height 25%.
[0015] Furthermore, the various vegetation indices include: visible light resistance to atmospheric conditions index, red-green vegetation index, ratio vegetation index, atmospheric resistance vegetation index, difference vegetation index, normalized vegetation index, enhanced vegetation index, and soil-regulating vegetation index.
[0016] Furthermore, the window size parameter of the first-order and second-order probabilistic statistical filtering tools is set to 3.
[0017] Furthermore, after performing first-order probabilistic statistical filtering, 20 texture factors are obtained; after performing second-order probabilistic statistical filtering, 32 texture factors are obtained.
[0018] Furthermore, the random forest machine learning model is trained using cross-validation, with 90% of the data used for training and 10% used for validation.
[0019] This application also provides a forest ecosystem carbon storage accounting system for implementing the accounting method described above, the system comprising:
[0020] The satellite-borne feature parameter acquisition module is used to preprocess the acquired satellite-borne lidar data and then extract satellite-borne feature parameters from the preprocessed data. The satellite-borne feature parameters include waveform length features and waveform height features. The features extracted from optical remote sensing images that are spatially matched with the satellite-borne lidar points are used as input variables. Combined with a random forest regression model, the satellite-borne feature parameters are extended to the entire study area to generate continuous satellite-borne feature parameters for the entire study area.
[0021] The vegetation index acquisition module is used to acquire optical remote sensing data of the entire study area, process the optical remote sensing data to obtain red, green, blue and near-infrared bands, and convert the red, green, blue and near-infrared bands into surface reflectance products respectively. Then, based on the surface reflectance products of each band, multiple vegetation indices are calculated.
[0022] The texture factor acquisition module, based on first-order and second-order probabilistic statistical filtering tools, performs texture filtering on the red, green, blue, and near-infrared bands respectively to extract first-order and second-order texture factors. The first-order texture factors include data range, mean, variance, information entropy, and skewness, while the second-order texture factors include mean, variance, coherence, contrast, dissimilarity, information entropy, second moment, and correlation.
[0023] The feature variable selection module uses a random forest algorithm to rank the importance of continuous satellite-borne feature parameters, various vegetation indices, first-order texture factors, and second-order texture factors across the entire study area. The feature with the highest importance, as ranked by the random forest algorithm, is fixed as the first variable. Then, other features are added one by one in descending order of importance, with the second and third features fixed sequentially, until all features are involved in the calculation. After each new feature is added, the root mean square error of the model is recorded, and the combination of variables corresponding to the minimum root mean square error of the model is selected as the optimal feature variable set.
[0024] The canopy height acquisition module is used to combine the optimal set of feature variables with the elevation model data and the slope data derived therefrom, to construct and train a random forest machine learning model, and to predict and generate continuous distribution results of canopy height throughout the study area.
[0025] The forest system vegetation carbon storage calculation module is used to input the continuously distributed canopy height results into the allometric growth equation to estimate the forest biomass of the entire study area. Then, based on the forest system land type and tree species, the forest type and carbon content coefficient are determined. The forest biomass is multiplied by the tree species carbon content coefficient to calculate the carbon storage of different forest types. Finally, the carbon storage of the forest system vegetation of the entire study area is obtained by summarizing the results.
[0026] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the accounting method described above.
[0027] The forest ecosystem carbon storage accounting method provided by this invention enables continuous regional observation, with a single-track coverage width of hundreds of kilometers, allowing for rapid acquisition of large-scale forest data. This invention extends satellite-borne features to the entire study area through a full-training method, perfectly adapting to the wide-area coverage capabilities of carbon satellites and overcoming the limitations of airborne lidar's "small-scale, fragmented" monitoring. It is particularly suitable for national-scale forest carbon storage surveys. Using the method of this invention, no complex ground preparation is required, the cost of a single data acquisition is far lower than that of airborne methods, and it is unaffected by short-term terrain and weather conditions, enabling stable acquisition of large-scale data. Relying on carbon satellite data, this invention significantly lowers the threshold for large-scale forest parameter estimation, making carbon monitoring possible in underdeveloped or ecologically fragile areas. The method of this invention provides long-term, continuous time-series data with a high degree of data standardization, facilitating the analysis of interannual variations in forest carbon pools. Through waveform feature extraction and model optimization, this invention can fully utilize carbon satellite time-series data to achieve dynamic monitoring of carbon storage and biomass, providing support for carbon sink trend prediction—something that airborne lidar cannot efficiently achieve. Attached Figure Description
[0028] Figure 1 This is a flowchart of a method for calculating carbon storage in a forest ecosystem, as described in the embodiment.
[0029] Figure 2 The image shows the canopy height fitting diagram of the random forest machine learning model in the example scheme.
[0030] Figure 3 Canopy height residuals of the random forest machine learning model in the example implementation scheme.
[0031] Figure 4 This is a block diagram of a forest ecosystem carbon storage accounting system involved in the implementation scheme. Detailed Implementation
[0032] To better understand the above technical solutions, exemplary embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. While exemplary embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of this disclosure to those skilled in the art.
[0033] As one implementation scheme, Figure 1 This is a flowchart of the accounting method involved in the embodiments of this application.
[0034] like Figure 1 As shown, the accounting method includes:
[0035] The first step is to preprocess the acquired spaceborne lidar data, and then extract spaceborne feature parameters from the preprocessed data. The spaceborne feature parameters include waveform length features and waveform height features.
[0036] This step involves denoising the large amount of background noise contained in the spaceborne lidar data, setting invalid signals to 0, and thus retaining valid signals. The noise threshold is calculated using the formula: threshold = mean ± n * sd_corrected, where threshold is the set background noise threshold, mean is the average noise value, sd_corrected is the noise standard deviation, and n is given as a predetermined multiplier. As an empirical constant, the value of n is usually determined based on the actual waveform. If the value of n is set too large, valid signals within the waveform will be discarded during processing, thus affecting the accuracy of subsequent waveform decomposition. Conversely, if the value of n is set too small, it will be affected by noise from the ground shrub layer, leading to misjudgment as valid signals.
[0037] The denoised waveform data is smoothed using Gaussian filtering to remove "glitch" signals. Simultaneously, to filter out non-forest information and highlight the effective signal, a Gaussian algorithm is used for waveform decomposition. A Gaussian component is determined based on two inflection points and one maximum point. These Gaussian components are then merged according to certain principles, resulting in the initial parameters of the obtained Gaussian components.
[0038] Waveform length and height features were extracted using Python. Waveform length features include the waveform start position, waveform end position, total waveform height, first peak position, last peak position, leading edge width, trailing edge width, and peak length. Waveform height features include relative height percentages of 25%, 50%, 75%, 95%, 97%, and the difference between 75% and 25% of the relative height. These features reflect the proportion of vegetation cover above the forest floor and the three-dimensional structure of the forest.
[0039] To estimate canopy height, several satellite-borne features from lidar data are used in conjunction with significant correlations with canopy height. For example, the last peak position refers to the location of the lidar signal's ground echo, reflecting forest structure information. The 97% waveform relative height refers to the position where the lidar signal's return waveform reaches 97% of its height. This feature is also closely related to canopy height, as it reflects the upper part of the canopy and provides an estimate that can distinguish low vegetation or other obstructions. The first peak position is the location of the first significant peak in the lidar signal's return after encountering the canopy; this position is usually associated with the top of the canopy and is an important indicator of canopy height. The half-wave height refers to the height at which the energy in the lidar signal's return waveform reaches half of its peak value; this feature reflects the structural characteristics of the canopy.
[0040] The second step involves using the features extracted from the optical remote sensing images that are spatially matched with the location of the spaceborne lidar as input variables. Combined with a random forest regression model, the spaceborne feature parameters are extended to the entire study area to generate continuous spaceborne feature parameters for the entire study area.
[0041] Because lidar data points are spatially discontinuous, obtaining continuous forest structure information over large areas remains challenging. Optical imagery, however, offers wide coverage and high resolution. By combining optical imagery with lidar data, the spatial limitations of lidar data can be compensated for by utilizing the features of optical imagery. Using features obtained from optical imagery as input variables, and incorporating a random forest algorithm, the numerical values of the spaceborne features are extended to the entire study area, generating a continuous surface across the entire region. This ensures comprehensive spaceborne features throughout the region, facilitating the acquisition of comprehensive forest structure parameters.
[0042] The third step is to acquire optical remote sensing data of the entire study area, process the optical remote sensing data to obtain red, green, blue and near-infrared bands, and convert the red, green, blue and near-infrared bands into surface reflectance products respectively. Then, based on the surface reflectance products of each band, multiple vegetation indices are calculated.
[0043] For continuous optical remote sensing features, this step uses ArcGIS and ENVI to preprocess Landsat 9 L2SP data. Based on the conversion coefficients provided in the Landsat 9 product manual, each band is converted into a surface reflectance product. Then, the surface reflectance product and the Band Math tool are used to calculate eight commonly used vegetation indices, including the Visible Atmosphere Resistance Index (VARI), Red-Green Vegetation Index (RGVI), Ratio Vegetation Index (RVI), Atmospheric Resistance Vegetation Index (ARVI), Difference Vegetation Index (DVI), Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Soil-Regulated Vegetation Index (SAVI). These vegetation indices can reflect the vegetation and surface characteristics of the entire study area.
[0044] The fourth step involves performing texture filtering on the red, green, blue, and near-infrared bands using first-order and second-order probabilistic statistical filtering tools, respectively, to extract first-order and second-order texture factors. The first-order texture factors include data range, mean, variance, information entropy, and skewness, while the second-order texture factors include mean, variance, coherence, contrast, dissimilarity, information entropy, second moment, and correlation.
[0045] In this step, the Occurrence Measures tool in ENVI is used based on the first-order probabilistic statistical filtering tool. The window size parameter is set to 3 to obtain the texture features of the optical remote sensing image. Each band has 5 texture factors, namely data range, mean, variance, entropy, and skewness. The probabilistic statistical filtering uses the number of occurrences of each gray level in the processing window for texture calculation. Then, the red, green, blue, and near-infrared bands of the Landsat image are used to obtain a multi-band texture factor image through the above steps. The multi-band image is then split into single-band images, resulting in a total of 20 texture factor images for subsequent data processing.
[0046] In addition, to obtain the characteristic variables of forest stock volume, biomass, and carbon storage, the Co-occurrence Measures tool based on second-order probability statistics in ENVI was used. The window size parameter was set to 3 to obtain the texture features of the optical remote sensing image. For each band of the optical image, there are 8 texture factors, namely mean, variance, homogeneity, contrast, dissimilarity, entropy, second moment, and correlation. Then, the red, green, blue, and near-infrared bands of Landsat image were used to obtain a multi-band texture factor image through the above steps. Then, the multi-band image was split into single-band images, resulting in a total of 32 texture factor images.
[0047] The fifth step involves using a random forest algorithm to rank the importance of continuous satellite-borne feature parameters, various vegetation indices, first-order texture factors, and second-order texture factors across the entire study area, thereby obtaining the optimal set of feature variables.
[0048] The most important feature, sorted by the random forest algorithm, is fixed as the first variable. Then, other features are added one by one in order of importance from high to low. The second and third features are fixed in turn until all features are involved in the calculation. After each new feature is added, the root mean square error of the model is recorded. The combination of variables corresponding to the minimum root mean square error of the model is selected as the optimal set of feature variables.
[0049] The sixth step involves combining the optimal set of feature variables with the elevation model data and the slope data derived from it to construct and train a random forest machine learning model, predicting and generating continuous distribution results of canopy height for the entire study area.
[0050] In this embodiment, 90% of the data in the Random Forest (RF) tree is selected for training the model, and the remaining 10% is used for model validation. The selected features are combined with ground survey plots and airborne data in a certain area to construct a Random Forest machine learning model. Its canopy height fitting map and canopy height residual map are shown below. Figure 2 , 3 As shown, where Figure 2 The figure shows the fitting effect of using the random forest machine learning model for canopy height prediction. The horizontal axis represents the observed value, i.e. the measured canopy height, and the vertical axis represents the predicted value, i.e. the predicted canopy height. The scatter points in the figure represent sample points, and the black diagonal line is the 1:1 reference line. Figure 3 This is used to evaluate the distribution and trend of model prediction errors. The horizontal axis represents the observed values, i.e., the measured canopy height, and the vertical axis represents the residuals. The scatter plots represent the residuals for each sample point. Figure 2and Figure 3 As can be seen, RMSE = 2.19m, EA = 86.34%, R 2 =0.85, the residual distribution of canopy height conforms to a random distribution, indicating that the random forest machine learning model can accurately capture and interpret the changing trend of the data, and the prediction results have high reliability and accuracy, and can capture the true distribution of the data well.
[0051] The seventh step involves inputting the continuously distributed canopy height results into the allometric growth equation to estimate the forest biomass of the entire study area. Then, based on the forest system land types and tree species, the forest type and carbon content coefficient are determined. The forest biomass is multiplied by the tree species carbon content coefficient to calculate the carbon storage of different forest types. Finally, the total vegetation carbon storage of the forest system in the entire study area is obtained.
[0052] The allometric growth equation is a mathematical model describing the growth of organisms, showing the relationship between their growth rate and biomass indicators such as volume, weight, and length. In this step, we first calculate the aboveground biomass (AGB) of the entire vegetation area, and then use the root-to-rhizome ratio to calculate the belowground biomass (BGB) from the aboveground biomass. The specific formula is as follows:
[0053] AGB stem =a*H b
[0054] AGB above =m*AGBstem n
[0055] Where: AGB stem Stem biomass (Mg·hm) -2 ); AGB above Aboveground biomass (Mg·hm) -2 H represents the canopy height (m); a, b, m, and n are the power function parameters of the allometric growth model, and the parameter values need to be differentiated according to different ecological regions and vegetation types in China.
[0056] Forest carbon storage equals forest biomass multiplied by the carbon content coefficient of tree species. The internationally used carbon content coefficient is 0.45~0.5. Based on the forest system land type and tree species, the forest type is determined, and the carbon storage of different forest types is calculated. The total carbon storage of forest system vegetation in the monitoring area is obtained by summing them up.
[0057]
[0058] In the formula: C is the total carbon storage of forest vegetation (Mg); B i Biomass (Mg / hm) of the i-th forest land type 2); S i Let be the area of the i-th forest type (hm). 2 ); α i denoted as the carbon content coefficient of the i-th forest land type, obtained from literature and sample plot survey data; i represents the number of forest land types.
[0059] In the calculation method provided in this embodiment, obtaining accurate characteristic variables is a crucial step before estimating carbon storage and canopy height. These characteristic variables include optical remote sensing data, radar data, vegetation indices, texture factors, and spaceborne characteristic parameters. These variables reflect forest structure and ecological conditions. Optical remote sensing data provides land cover and vegetation information, spaceborne lidar data can penetrate the tree canopy to obtain topographic and canopy height information, vegetation indices (such as NDVI) reflect vegetation health, texture factors describe surface roughness and structural features, and spaceborne characteristic parameters provide precise height information. Obtaining these characteristic variables provides a rich information foundation for subsequent modeling and analysis.
[0060] After obtaining a large number of feature variables, they need to be screened to determine the features that have the greatest influence on the target variable. The purpose of feature screening is to remove redundant and irrelevant variables and retain the variables that contribute most to the model's predictive performance, thereby improving the model's accuracy and stability. Through feature screening, the complexity of the model can be reduced, overfitting can be avoided, the model's generalization ability can be improved, and the relationship between feature variables and the target variable can be better understood.
[0061] Building a reserve model based on the selected feature variables is a crucial step. Random Forest (RF) can be used to construct the model. Through model building, accurate estimation of canopy height can be achieved.
[0062] Then, the aboveground biomass (AGB) of the entire vegetation area is calculated using the allometric growth equation, and the belowground biomass (BGB) is calculated from the aboveground biomass using the root-to-stem ratio. Finally, the aboveground and belowground biomass are summed to obtain the total vegetation biomass, which is then multiplied by the carbon content coefficient to obtain the vegetation carbon storage of the forest ecosystem.
[0063] In another embodiment, a forest ecosystem carbon storage accounting system is provided to implement the accounting method described above, such as... Figure 4 As shown, the system includes:
[0064] The satellite-borne feature parameter acquisition module is used to preprocess the acquired satellite-borne lidar data and then extract satellite-borne feature parameters from the preprocessed data. The satellite-borne feature parameters include waveform length features and waveform height features. The features extracted from the optical remote sensing images that are spatially matched with the satellite-borne lidar points are used as input variables. Combined with the random forest regression model, the satellite-borne feature parameters are extended to the entire study area to generate continuous satellite-borne feature parameters for the entire study area.
[0065] The vegetation index acquisition module is used to acquire optical remote sensing data of the entire study area, process the optical remote sensing data to obtain red, green, blue and near-infrared bands, convert the red, green, blue and near-infrared bands into surface reflectance products respectively, and then calculate multiple vegetation indices based on the surface reflectance products of each band.
[0066] The texture factor acquisition module, based on first-order and second-order probabilistic statistical filtering tools, performs texture filtering on the red, green, blue and near-infrared bands respectively to extract first-order texture factors and second-order texture factors. The first-order texture factors include data range, mean, variance, information entropy and skewness. The second-order texture factors include mean, variance, coherence, contrast, dissimilarity, information entropy, second moment and correlation.
[0067] The feature variable selection module uses a random forest algorithm to rank the importance of continuous satellite-borne feature parameters, multiple vegetation indices, first-order texture factors, and second-order texture factors in the entire study area. The feature with the highest importance ranked by the random forest algorithm is fixed as the first variable. Then, other features are added one by one in order of importance from high to low. The second and third features are fixed in turn until all features are involved in the calculation. After each new feature is added, the root mean square error of the model is recorded. The combination of variables corresponding to the minimum root mean square error of the model is selected as the optimal feature variable set.
[0068] The canopy height acquisition module is used to combine the optimal set of feature variables with the elevation model data and the slope data derived therefrom, to construct and train a random forest machine learning model, and to predict and generate a continuously distributed canopy height result for the entire study area.
[0069] The forest system vegetation carbon storage calculation module is used to input the continuously distributed canopy height results into the allometric growth equation to estimate the forest biomass of the entire study area. Then, based on the forest system land type and tree species, the forest type and carbon content coefficient are determined. The forest biomass is multiplied by the tree species carbon content coefficient to calculate the carbon storage of different forest types. Finally, the carbon storage of the forest system vegetation of the entire study area is obtained by summarizing the results.
[0070] This invention provides a method and system for calculating carbon storage in forest ecosystems. Combining the random forest algorithm, it extends discrete spaceborne features to a continuous regional surface, perfectly adapting to the wide-area coverage advantage of spaceborne lidar. This overcomes the limitations of airborne lidar's "small-scale, fragmented" monitoring, making it particularly suitable for national-scale forest carbon storage surveys. The "hierarchical estimation model of forest carbon pool parameters" (canopy height → biomass → carbon storage) is based on spaceborne wide-area data and can directly output regional-scale spatial distribution results of carbon storage. Spaceborne lidar operates periodically on a fixed orbit, providing long-term continuous time-series data with a high degree of standardization, facilitating the analysis of interannual variations in forest carbon pools.
Claims
1. A method for calculating carbon storage in forest ecosystems, characterized in that, The method includes: The acquired spaceborne lidar data is preprocessed, and then spaceborne feature parameters are extracted from the preprocessed data. The spaceborne feature parameters include waveform length features and waveform height features. Features extracted from optical remote sensing images that are spatially matched with the location of the spaceborne lidar are used as input variables. Combined with a random forest regression model, the spaceborne feature parameters are extended to the entire study area to generate continuous spaceborne feature parameters for the entire study area. Optical remote sensing data of the entire study area is acquired, and the optical remote sensing data is processed to obtain red, green, blue and near-infrared bands. The red, green, blue and near-infrared bands are converted into surface reflectance products respectively. Then, based on the surface reflectance products of each band, multiple vegetation indices are calculated. Based on first-order and second-order probabilistic statistical filtering tools, texture filtering is performed on the red, green, blue and near-infrared bands respectively to extract first-order texture factors and second-order texture factors. The first-order texture factors include data range, mean, variance, information entropy and skewness. The second-order texture factors include mean, variance, coherence, contrast, dissimilarity, information entropy, second moment and correlation. The importance of continuous satellite-borne feature parameters, multiple vegetation indices, first-order texture factors, and second-order texture factors in the entire study area is ranked using the random forest algorithm. The feature with the highest importance ranked by the random forest algorithm is fixed as the first variable. Then, other features are added one by one in order of importance from high to low. The second and third features are fixed in turn until all features are involved in the calculation. After each new feature is added, the root mean square error of the model is recorded. The combination of variables corresponding to the minimum root mean square error of the model is selected as the optimal set of feature variables. The optimal set of feature variables is combined with elevation model data and its derived slope data to construct and train a random forest machine learning model, predicting and generating continuous distribution results of canopy height throughout the study area. The continuously distributed canopy height results are input into the allometric growth equation to estimate the forest biomass of the entire study area. Then, based on the forest system land type and tree species, the forest type and carbon content coefficient are determined. The forest biomass is multiplied by the tree species carbon content coefficient to calculate the carbon storage of different forest types. Finally, the vegetation carbon storage of the entire forest system in the study area is obtained.
2. The accounting method according to claim 1, characterized in that, The acquired spaceborne lidar data undergoes preprocessing, specifically including: First, the background noise threshold is set by the noise average value and noise standard deviation to denoise the spaceborne lidar data; then, Gaussian filtering is used to smooth the denoised waveform data to remove glitches, and Gaussian decomposition algorithm is used to filter out non-forest information.
3. The accounting method according to claim 1, characterized in that, The waveform length features include the waveform start position, waveform end position, waveform total height, first peak position, last peak position, waveform leading edge width, waveform trailing edge width, and peak length. The waveform height features include relative height 25%, relative height 50%, relative height 75%, relative height 95%, relative height 97%, and the difference between relative height 75% and relative height 25%.
4. The accounting method according to claim 1, characterized in that, The various vegetation indices include: visible light resistance to atmospheric conditions index, red-green vegetation index, ratio vegetation index, atmospheric resistance vegetation index, difference vegetation index, normalized vegetation index, enhanced vegetation index, and soil-regulating vegetation index.
5. The accounting method according to claim 1, characterized in that, The window size parameter for the first-order and second-order probabilistic statistical filtering tools is set to 3.
6. The accounting method according to claim 1, characterized in that, After performing first-order probabilistic statistical filtering, 20 texture factors are obtained; after performing second-order probabilistic statistical filtering, 32 texture factors are obtained.
7. The accounting method according to claim 1, characterized in that, The random forest machine learning model is trained using cross-validation, with 90% of the data used for training and 10% used for validation.
8. A carbon storage accounting system for forest ecosystems, characterized in that, The system for implementing the accounting method according to any one of claims 1-7 comprises: The satellite-borne feature parameter acquisition module is used to preprocess the acquired satellite-borne lidar data and then extract satellite-borne feature parameters from the preprocessed data. The satellite-borne feature parameters include waveform length features and waveform height features. The features extracted from the optical remote sensing images that are spatially matched with the satellite-borne lidar points are used as input variables. Combined with the random forest regression model, the satellite-borne feature parameters are extended to the entire study area to generate continuous satellite-borne feature parameters for the entire study area. The vegetation index acquisition module is used to acquire optical remote sensing data of the entire study area, process the optical remote sensing data to obtain red, green, blue and near-infrared bands, convert the red, green, blue and near-infrared bands into surface reflectance products respectively, and then calculate multiple vegetation indices based on the surface reflectance products of each band. The texture factor acquisition module, based on first-order and second-order probabilistic statistical filtering tools, performs texture filtering on the red, green, blue and near-infrared bands respectively to extract first-order texture factors and second-order texture factors. The first-order texture factors include data range, mean, variance, information entropy and skewness. The second-order texture factors include mean, variance, coherence, contrast, dissimilarity, information entropy, second moment and correlation. The feature variable selection module uses a random forest algorithm to rank the importance of continuous satellite-borne feature parameters, multiple vegetation indices, first-order texture factors, and second-order texture factors in the entire study area. The feature with the highest importance ranked by the random forest algorithm is fixed as the first variable. Then, other features are added one by one in order of importance from high to low. The second and third features are fixed in turn until all features are involved in the calculation. After each new feature is added, the root mean square error of the model is recorded. The combination of variables corresponding to the minimum root mean square error of the model is selected as the optimal feature variable set. The canopy height acquisition module is used to combine the optimal set of feature variables with the elevation model data and the slope data derived therefrom, to construct and train a random forest machine learning model, and to predict and generate a continuously distributed canopy height result for the entire study area. The forest system vegetation carbon storage calculation module is used to input the continuously distributed canopy height results into the allometric growth equation to estimate the forest biomass of the entire study area. Then, based on the forest system land type and tree species, the forest type and carbon content coefficient are determined. The forest biomass is multiplied by the carbon content coefficient to calculate the carbon storage of different forest types. Finally, the total forest system vegetation carbon storage of the entire study area is obtained.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the accounting method as described in any one of claims 1 to 7.