Subtropical forest aboveground biomass estimation method and related device
By comprehensively utilizing multi-source remote sensing data and field survey data, a multivariate regression model based on XGBoost was constructed, which solved the problem of accuracy in estimating aboveground biomass in subtropical forests, achieved higher-precision biomass estimation, and supported research on ecosystem carbon cycle and forest management.
Patent Information
- Application Number
- CN202510981707.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies are insufficient to accurately estimate subtropical forest aboveground biomass across multiple dimensions, impacting global carbon cycle research and sustainable forest development.
By combining multi-source remote sensing data and field survey data, we constructed a multivariate regression model based on XGBoost by extracting forest attributes, vegetation index features, texture features and topographic factors, and used machine learning methods to estimate biomass.
It improves the accuracy and precision of aboveground forest biomass estimation, supporting research on ecosystem carbon cycles and forest management.
Smart Images

Figure CN120877103A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of quantitative remote sensing technology, and in particular to a method and related apparatus for estimating aboveground biomass in subtropical forests. Background Technology
[0002] Forest biomass, as a crucial indicator of forest carbon storage, is a vital measure of forest carbon sequestration capacity and a foundation for evaluating regional forest carbon balance. Forest biomass consists of aboveground biomass (AGB) and belowground biomass (BGB), but in most cases, it is referred to as AGB, which can be measured with a certain degree of accuracy within a certain range. Accurate estimation of forest biomass and its changes can enhance our understanding of the global carbon cycle and reduce the uncertainty in carbon emission estimates caused by human activities or natural disturbances. In today's rapidly changing climate, conducting regional multi-scale forest biomass estimation research can not only provide a theoretical basis for ecosystem carbon cycle research but also promote sustainable forest development. However, how to more accurately estimate forest biomass based on multiple dimensions remains a significant challenge. Summary of the Invention
[0003] The purpose of this application is to provide a method and related apparatus for estimating aboveground biomass in subtropical forests, which can more accurately estimate aboveground forest biomass.
[0004] To achieve the above objectives, this application provides the following solution:
[0005] Firstly, this application provides a method for estimating aboveground biomass in subtropical forests, including:
[0006] Acquire multi-source remote sensing and field survey data and historical ground feature spectral information for the target subtropical forest area; the multi-source remote sensing and field survey data includes lidar data, optical image data, ground survey data and topographic data.
[0007] Based on the photons in the lidar data, the attributes of the target subtropical forest region are extracted; the attributes include the longitude, latitude, cloud aerosols, canopy height, canopy height uncertainty, absolute canopy height, best fit of segmented terrain height, and best surface elevation of the target subtropical forest region.
[0008] Based on the spectral information in the optical image data, and using the historical spectral information of the target subtropical forest area, the vegetation index characteristics of the target subtropical forest area are calculated.
[0009] Based on the band information in the optical image data, the surface reflectance of the target subtropical forest area is determined.
[0010] Based on the surface reflectance of each single band, the gray-level co-occurrence matrix obtained by second-order statistical filtering is used to extract the texture features of the target subtropical forest region; the texture features include mean, variance, contrast, dissimilarity, information entropy, cooperativity, second moment and correlation.
[0011] Based on the topographic data, 3D analysis tools are used to extract the topographic factors of the target subtropical forest area; the topographic factors include slope factor, aspect factor and elevation factor.
[0012] Based on the attributes, vegetation index characteristics, texture characteristics, and topographic factors of the target subtropical forest region, a multivariate regression model is constructed; the multivariate regression model is a machine learning model based on XGBoost; the machine learning model based on XGBoost is constructed by adding a biomass conservation equation constraint term to the XGBoost loss function to construct a physics-guided hybrid machine learning model.
[0013] The aboveground biomass of the target subtropical forest area is estimated based on a multivariate regression model, and the estimation results are output.
[0014] Secondly, this application provides a device for estimating aboveground biomass in subtropical forests, comprising:
[0015] The data acquisition module is used to acquire multi-source remote sensing and field survey data and historical ground feature spectral information of the target subtropical forest area; the multi-source remote sensing and field survey data includes lidar data, optical image data, ground survey data and topographic data.
[0016] The attribute extraction module is used to extract the attributes of the target subtropical forest region based on the photons in the lidar data; the attributes include the longitude, latitude, cloud aerosols, canopy height, canopy height uncertainty, absolute canopy height, best fit of segmented terrain height, and best surface elevation of the target subtropical forest region.
[0017] The vegetation index feature calculation module is used to calculate the vegetation index features of the target subtropical forest area based on the spectral information in the optical image data and the historical ground cover spectral information of the target subtropical forest area.
[0018] The surface reflectance determination module is used to determine the surface reflectance of the target subtropical forest area based on the band information in the optical image data.
[0019] The texture feature extraction module is used to extract the texture features of the target subtropical forest area based on the gray-level co-occurrence matrix obtained by second-order statistical filtering according to the surface reflectance of each single band; the texture features include mean, variance, contrast, dissimilarity, information entropy, cooperativity, second moment and correlation.
[0020] The terrain factor extraction module is used to extract terrain factors of the target subtropical forest area based on terrain data using 3D analysis tools; the terrain factors include slope factor, aspect factor and elevation factor.
[0021] The model building module is used to construct a multivariate regression model based on the attributes, vegetation index characteristics, texture characteristics, and topographic factors of the target subtropical forest region. The multivariate regression model is a machine learning model based on XGBoost. The machine learning model based on XGBoost is constructed by adding a biomass conservation equation constraint term to the XGBoost loss function to build a physics-guided hybrid machine learning model.
[0022] The estimation module is used to estimate the aboveground biomass of a target subtropical forest area based on a multivariate regression model and output the estimation results.
[0023] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a method for estimating aboveground biomass in a subtropical forest as described in any of the above-mentioned methods.
[0024] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for estimating aboveground biomass in a subtropical forest as described above.
[0025] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements a method for estimating aboveground biomass in subtropical forests as described above.
[0026] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0027] This application provides a method and related apparatus for estimating aboveground biomass in subtropical forests. First, multi-source remote sensing and field survey data, as well as historical spectral information of ground features, are acquired, providing a rich data foundation for estimation. Lidar data, optical imagery data, ground survey data, and topographic data each contain different forest characteristic information; their combined use can comprehensively reflect the complex structure of the forest. Next, forest attributes, including longitude, latitude, cloud aerosols, and canopy height, are extracted based on lidar data. These attributes help understand the spatial distribution and vertical structure of the forest, which is crucial for biomass estimation. Then, vegetation index characteristics are calculated using optical imagery data, reflecting the forest's vegetation cover and health status, serving as an important reference for biomass estimation. Simultaneously, determining surface reflectance and extracting texture features can further reveal the forest's fine structure and composition, improving the accuracy of the estimation. Furthermore, topographic factors, such as slope, aspect, and altitude, are extracted from topographic data. These factors have a significant impact on forest growth and distribution and are therefore indispensable factors in biomass estimation. Finally, a multivariate regression model based on XGBoost was constructed. This model comprehensively considers all the above factors and automatically finds the optimal estimation parameters through machine learning, thereby achieving accurate estimation of aboveground biomass. In summary, this application improves the accuracy of aboveground forest biomass estimation through comprehensive analysis and modeling. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is an application environment diagram of a method for estimating aboveground biomass in subtropical forests according to an embodiment of this application.
[0030] Figure 2 This is a flowchart illustrating a method for estimating aboveground biomass in subtropical forests, provided as an embodiment of this application.
[0031] Figure 3 This is a single-band surface reflectance map of GF-2 imagery provided in an embodiment of this application.
[0032] Figure 4 This is a schematic diagram of vegetation index provided in one embodiment of this application.
[0033] Figure 5 This is a texture feature factor result diagram provided in an embodiment of this application.
[0034] Figure 6Correlation diagram among various factors of Masson pine provided in an embodiment of this application.
[0035] Figure 7 Correlation diagram among various factors of slash pine provided in an embodiment of this application.
[0036] Figure 8 Correlation diagram among various factors of Quercus acutissima provided in an embodiment of this application.
[0037] Figure 9 Correlation diagram among various factors in a mixed forest provided in an embodiment of this application.
[0038] Figure 10 A model fitting effect diagram provided in an embodiment of this application.
[0039] Figure 11 Distribution of biomass estimation results for different models provided in one embodiment of this application.
[0040] Figure 12 This is a schematic diagram of the functional modules of a subtropical forest aboveground biomass estimation device provided in an embodiment of this application.
[0041] Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0043] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] The method for estimating aboveground biomass in subtropical forests provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on other servers. Terminal 102 can send the multi-source remote sensing and field survey data and historical ground object spectral information to be processed to server 104. After receiving the multi-source remote sensing and field survey data and historical ground object spectral information, server 104 extracts the attributes of the target subtropical forest area based on photons in the lidar data; calculates the vegetation index characteristics of the target subtropical forest area based on the spectral information in the optical image data and the historical ground object spectral information of the target subtropical forest area; determines the surface reflectance of the target subtropical forest area based on the band information in the optical image data; extracts the texture features of the target subtropical forest area based on the gray-level co-occurrence matrix obtained by second-order statistical filtering based on the surface reflectance of each single band; extracts the topographic factors of the target subtropical forest area based on the topographic data using 3D analysis tools; and estimates the aboveground biomass of the target subtropical forest area based on a multivariate regression model and outputs the estimation results; the multivariate regression model is a machine learning model based on XGBoost. Server 104 can feed back the obtained estimation results to terminal 102. Furthermore, in some embodiments, the method for estimating subtropical forest aboveground biomass can also be implemented independently by server 104 or terminal 102. For example, terminal 102 can directly process the multi-source remote sensing and field survey data and historical ground cover spectral information to be processed, or server 104 can obtain the multi-source remote sensing and field survey data and historical ground cover spectral information to be processed from the data storage system and process the data accordingly.
[0045] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.
[0046] In one exemplary embodiment, such as Figure 2As shown, a method for estimating aboveground biomass in subtropical forests is provided. This method is executed by computer equipment, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 208. Wherein:
[0047] Step 201: Obtain multi-source remote sensing and field survey data and historical ground feature spectral information of the target subtropical forest area; the multi-source remote sensing and field survey data includes lidar data, optical image data, ground survey data and topographic data.
[0048] Step 202: Based on the photons in the lidar data, extract the attributes of the target subtropical forest region; the attributes include the longitude, latitude, cloud aerosols, canopy height, canopy height uncertainty, absolute canopy height, best fit of segmented terrain height, and best surface elevation of the target subtropical forest region.
[0049] Step 203: Based on the spectral information in the optical image data and the historical ground cover spectral information of the target subtropical forest area, calculate the vegetation index characteristics of the target subtropical forest area.
[0050] Step 204: Determine the surface reflectance of the target subtropical forest area based on the band information in the optical image data.
[0051] Step 205: Based on the surface reflectance of each single band, the gray-level co-occurrence matrix obtained by second-order statistical filtering is used to extract the texture features of the target subtropical forest region; the texture features include mean, variance, contrast, dissimilarity, information entropy, cooperativity, second moment and correlation.
[0052] Step 206: Based on the terrain data, use 3D analysis tools to extract the terrain factors of the target subtropical forest area; the terrain factors include slope factor, aspect factor and elevation factor.
[0053] Step 207: Based on the attributes, vegetation index characteristics, texture characteristics, and topographic factors of the target subtropical forest area, construct a multivariate regression model; the multivariate regression model is a machine learning model based on XGBoost; the machine learning model based on XGBoost is constructed by adding a biomass conservation equation constraint term to the XGBoost loss function to construct a physics-guided hybrid machine learning model.
[0054] Step 208: Estimate the aboveground biomass of the target subtropical forest area based on a multivariate regression model and output the estimation results.
[0055] In this embodiment, the National Forest Park of City A is taken as the study area, with geographical coordinates between 117°58'-118°03'E and 32°17'-32°25'N. It is located in the northern subtropical humid monsoon climate zone, in the northeastern part of Anhui Province, with a total area of 3551.5 km2. According to the 2017 forest sub-compartment survey data, the forest types include coniferous forest, deciduous broad-leaved forest, broad-leaved mixed forest, and coniferous-broad-leaved mixed forest. Among them, pure forest sub-compartments represented by Masson pine, slash pine, and Quercus acutissima account for more than 78.2% of the entire forest area.
[0056] In one exemplary embodiment, step 201 can be performed as follows:
[0057] When acquiring multi-source remote sensing and field survey data for the target subtropical forest area, the process is specifically divided into four aspects: acquiring lidar data, optical image data, ground survey data, and topographic data.
[0058] For optical imagery data acquisition, ICESAT-2 / ATLAS data was downloaded from the Snow and Ice Data Center. This center provides 21 standard ATLAS data products across 4 levels. This embodiment uses ATL08 data, a land vegetation elevation product derived from the photon-counting laser altimeter ATLAS on the ICESAT-2 satellite. The ATL08 product provides a relative height measurement in 100-meter (along-orbit) × 12-meter (across-orbit) segments to ensure a sufficient number of photons are available for canopy height estimation. For canopy height, the h_Canopy (Rh98) metric is used, which provides the relative canopy height at 98% of the energy return height relative to the ground data. Due to uncertainties in the signal-to-noise ratio at the canopy top, Rh98 is used instead of the maximum height to represent the canopy top height.
[0059] Because seasonal changes in vegetation and ground snow and ice can affect the number of canopy photons and the vertical structure information of the canopy in ATLAS data, this study selected ATL08 data products from June to October 2023 within the National Forest Park area of City A, focusing primarily on summer ATLAS photon point cloud data to reduce the impact of seasonal changes on photon data. The official documentation provides relevant modules in Python for extracting light spot information from ATL08; this embodiment uses a Python script to process the downloaded data.
[0060] For acquiring lidar data, this embodiment selects four GF-2PMS1 images from October 31, 2023, and two GF-6WFV images from April 30, 2023, which can uniformly cover the area of this embodiment. The Gaofen-2 (GF-2) image has a spatial resolution of 0.8m and possesses four bands: red, green, blue, and near-infrared. The GF-6WFV image has a spatial resolution of 16m and, in addition to the traditional multispectral bands (B1-B4: 0.45-0.89), adds Red Edge I (B5: 0.69-0.73), Red Edge II (0.73-0.77), Purple Edge (B7: 0.40-0.45), and Yellow Edge (B8: 0.59-0.62) bands. The images selected in this embodiment can completely cover the study area.
[0061] The GF-2 and GF-6 imagery were processed using ENVI 5.6 software. Radiometric calibration and atmospheric correction were performed based on the absolute radiometric calibration coefficients and spectral response functions of the GF-2PMS1 and GF-6 sensors, as officially released by the China Satellite Resources and Application Center. Orthorectification used 30m DEM data as terrain correction data, selecting correction points based on Google's offset-free imagery to correct geometric distortions caused by the system and reduce topographic-induced geometric distortions. A rational polynomial geometric fine correction model was used during the correction process; this is a general imaging sensor model for high-resolution satellite imagery, achieving a corrected GF-6 spatial resolution of 16m. The NNDiffuse Pan Sharpening fusion method was used to fuse the corrected GF-2 4m multispectral and 1m panchromatic data.
[0062] For the acquisition of ground survey data, this embodiment uses the 2017 Class II forest resource survey vector data of City A National Forest Park as the basis, and refers to the results of the Ninth National Forest Resources Inventory (2014) to obtain the "Statistical Table of Average Annual Growth Consumption of Arbor Forests by Different Tree Species". Based on the average annual growth of different tree species in each age group, the standing volume per hectare in the 2017 Class II survey data of City A National Forest Park was converted to the standing volume per hectare in 2023, and then corrected using samples from the field survey.
[0063] A two-stage sampling method was used for plot selection and layout. First, the survey was based on sub-compartments of arbor species in the forest farm, with all forested land as the total sampling base. Stratified sampling was then conducted for sub-compartments with different sites, slope aspects, and forest types to ensure a sufficient number of samples for each tree species. The plots were distributed as evenly as possible in terms of slope and aspect. In this embodiment, *Pinus massoniana*, *Pinus slashii*, and *Quercus acutissima* were selected as dominant tree species. The remaining tree species in the forest area were calculated as mixed coniferous and broad-leaved forests. A total of 452 sub-compartments were selected for biomass estimation.
[0064] Based on the second-class forest resource survey data, a single-dimensional volume table was used to invert the volume of individual trees, yielding the volume of each individual tree. The volume of all individual trees within the sample plot was then summed to obtain the volume of the entire plot. In this embodiment, the Fang et al. volume-biomass conversion equation and parameters were used to calculate the biomass of tree species (biomass in this study refers only to tree biomass, excluding shrub and herbaceous biomass). The regression equation is:
[0065] B = aV + b.
[0066] In the formula, B is the biomass per unit area (t / hm2); V is the volume per unit area (m3 / hm2); a and b are parameters, and the estimation model is selected according to different forest stand types. The corresponding values of parameters a and b are shown in Table 1.
[0067] Table 1 Biomass Conversion Coefficient
[0068]
[0069] For topographic data acquisition, 1:30m DEM data was used to extract elevation, slope, and aspect information of the study area, which were then used as topographic factors in the construction of the feature set. In this embodiment, the ground reference system used for spatial data is World Geodetic System-1984, and the projection method is Universal Transverse Mercator.
[0070] Furthermore, in this embodiment, three-dimensional canopy structure parameters (such as vertical stratified biomass distribution and leaf area density profiles) can be established by introducing ICESat-2 satellite data and UAV LiDAR point cloud data. Specifically, firstly, ICESat-2 altimeter waveform data and related geographic information, as well as UAV LiDAR data covering the same or adjacent areas, are collected. Then, both types of data are preprocessed, including coordinate transformation, data format unification, and noise reduction, to ensure data consistency and quality. Next, spatial registration is performed using Geographic Information System (GIS) tools, and spatial resolution is matched using interpolation algorithms to achieve data fusion, laying the foundation for subsequent analysis.
[0071] Building upon this foundation, large-scale vegetation height was estimated using ICESat-2 data, and detailed vegetation structure information was obtained by combining high-precision point cloud data from UAV LiDAR. Biomass distribution at different height levels was estimated through statistical analysis and modeling methods. Simultaneously, leaf area index (LAI) was extracted from LiDAR data and corrected and expanded using ICESat-2 data to establish a leaf area density profile model. Subsequently, ground-based measured data were used to verify the model's accuracy, and parameters were adjusted to optimize the estimation results. Finally, the spatial distribution characteristics of three-dimensional canopy structure parameters were analyzed, exploring their relationship with environmental factors (such as topography and climate), and the research findings were applied to fields such as ecological monitoring, forest management, and carbon cycle research.
[0072] In one exemplary embodiment, step 202 may be performed as follows:
[0073] In this embodiment, for the acquired ATL08 data, the following attributes were extracted for each photon: longitude, latitude, cloud aerosol, canopy height, canopy height uncertainty, absolute canopy height, best fit of segmented terrain height, and best surface elevation. The extracted attributes and their specific descriptions are listed in Table 2.
[0074] Table 2. Features and Detailed Description of ATL08
[0075]
[0076] In one exemplary embodiment, when performing steps 203-204, the specific steps may be as follows:
[0077] Spectral information was extracted from both GF-2 and GF-6 imagery. In addition to the original four bands from GF-2, eight bands from GF-6 were added, for a total of twelve bands. The surface reflectance of the four multispectral bands extracted from the GF-2 imagery is shown below. Figure 3 As shown.
[0078] Specifically, based on the existing ground cover spectral information, in order to better estimate forest biomass, seven common vegetation indices were added, including Normalized Difference Vegetation Index (NDVI), Ratio Vegetation Index (RVI), Difference Vegetation Index (DVI), Enhanced Vegetation Index (EVI), Soil-Adjusted Vegetation Index (SAVI), Greenness Normalized Difference Vegetation Index (GNDVI), and Infrared Percentage Vegetation Index (IPVI). The newly added red-edge bands of the domestic GF-6 satellite are similar to the 5th and 6th bands of Sentinel-2, so two vegetation indices that are highly sensitive to changes in biomass and used in Sentinel-2 data, MTCI and NDRE1, were introduced. MTCI is particularly sensitive to chlorophyll in plant leaves, and its ratio is directly proportional to the chlorophyll content in plant leaves. NDREI is often used for vegetation classification and leaf area index calculation by normalizing the peaks and troughs of the red edge. Nine vegetation indices were obtained by using the BandMath tool in ENVI to calculate the bands of the preprocessed Gaofen-6 image.
[0079] The formulas for calculating each vegetation index are as follows:
[0080]
[0081] NDVI = ρ NIR -ρ RED / ρ NIR +ρ RED .
[0082]
[0083] DVI = ρ NIR -ρ RED .
[0084]
[0085] IPVI = ρ NIR / ρ NIR +ρ RED .
[0086]
[0087] Among them, RVI is the ratio vegetation index, NDVI is the normalized difference vegetation index, EVI is the enhanced vegetation index, DVI is the difference vegetation index, SAVI is the soil-modified vegetation index, GNDVI is the green normalized difference vegetation index, IPVI is the infrared percentage vegetation index, MTCI is the MERIS terrestrial chlorophyll index, NDRE1 is the normalized difference red edge 1 index, and ρ NIR For the near-infrared band of GF-2 imagery, ρ REDFor the red edge band of the GF-2 image, L is set to 0.5, ρ BLUE For the blue-edge band of GF-2 imagery, ρ GREEN For the green edge band of GF-2 imagery, ρ 0.75 For GF-6 imagery in the red-edge II band, ρ 0.71 For the red-edge I band of GF-6 imagery, ρ in the MTCI equation RED1 This is the red-edge band of the GF-6 image.
[0088] Specifically, the extracted vegetation indices include, for example: Figure 4 As shown.
[0089] In one exemplary embodiment, step 205 may be performed as follows:
[0090] Studies have shown that incorporating spectral texture features can effectively improve the estimation accuracy of the model. This embodiment applies the Gray-Level Co-occurrence Matrix (GLCM) obtained based on second-order statistical filtering to extract common texture features from the four single bands of GF-2 and the newly added bands (B5, B6) of GF-6. Eight commonly used texture features are extracted, including: Mean, Variance, Contrast, Dissimilarity, Entropy, Homogeneity, Second Moment, and Correlation. The calculation formulas for each texture feature are as follows:
[0091] Mean = ∑ i ∑ j p(i,j)*i.
[0092] VAR = ∑ i ∑ j (iu) 2 p(i,j).
[0093]
[0094] CON = ∑ i ∑ j (ij) 2 p(i,j).
[0095] DIS = ∑ i ∑ j p(i,j)|ij|.
[0096] ENT=-∑ i ∑ j p(i,j)logp(i,j).
[0097] ASM=∑i ∑ j p(i,j) 2 .
[0098]
[0099] Where i and j represent the number of rows and columns of the matrix, P(i,j) represents the probability that the two gray values corresponding to the i-th row and j-th column appear simultaneously, and u i u j This represents the mean of the rows and columns. This represents the variance of rows and columns.
[0100] The extraction windows were 3x3, 5x5, 7x7, 9x9, 11x11, and 13x13, with a movement step of 1 pixel. To avoid window edge effects and maintain the stability and clarity of image texture features, 64 gray levels were selected. ENVI was used to extract eight texture feature factors from the 0° direction 3x3 window of the newly added red edge 1 (B5) in GF-6, as follows: Figure 5 As shown.
[0101] Furthermore, in some embodiments, texture feature extraction can employ anti-interference enhancement algorithms, namely, constructing a Cloud Mask Detection Network (CMDN) and using a Generative Adversarial Network (GAN) to repair texture features obscured by clouds. The repair process is mainly implemented through a GAN, and its core process includes cloud mask localization and GAN repair. First, the Cloud Mask Detection Network (CMDN) accurately identifies areas in the image obscured by clouds (generating a "mask"). Subsequently, the two core components of the GAN work together: the generator takes the cloud mask and the surrounding unobscured image texture as input, learns the statistical regularity and spatial structure of the texture within the region, and predicts and fills in reasonable texture features for the cloud-covered area; the discriminator judges the difference between the area repaired by the generator and the real clean image, and through repeated games, forces the generator to output a repair result consistent with the statistical distribution of real texture features.
[0102] The generator's restoration results must meet the dual constraints of local consistency and global rationality, ensuring a natural transition of texture at the edges of the restored area with the surrounding clear areas, while the overall texture features after restoration conform to the typical patterns of the land cover type. Through iterative optimization, the network ultimately generates restoration results that preserve the spatial details of the original image and conform to the statistical regularities of spectral texture, effectively eliminating cloud interference. This technique utilizes the adversarial learning mechanism of GANs, relying on prior knowledge of clear areas to reconstruct textures obscured by clouds, ensuring that the restored features are consistent with the real surface landscape in both local structure and global statistical properties, thereby improving the reliability of subsequent analysis.
[0103] In one exemplary embodiment, step 206 can be performed as follows:
[0104] This embodiment extracts terrain factors based on 30m resolution DEM data and uses 3D analysis tools in ArcMap to extract slope, aspect and elevation factors in the study area.
[0105] In one exemplary embodiment, step 207 can be performed as follows:
[0106] Extreme Gradient Boosting (XGBoost) is one of the most widely used and effective machine learning models. It's a scalable and highly accurate implementation of gradient boosting, pushing the computational limits of boosting tree algorithms. It's primarily used to improve the performance and speed of machine learning models. In regression fitting, a regression tree is first constructed, and then the boosting algorithm is used to continuously refine the fitting results. Furthermore, by using simple score discrimination for pruning and directly adding a regularization term to the penalty function, XGBoost effectively controls overfitting while maintaining accuracy, thus gaining increasing attention in quantitative remote sensing. During construction, XGBoost generates a new tree through feature segmentation to accommodate the residuals of the previous prediction in each iteration. The final prediction result is achieved by adding the output values of all samples predicted by each tree. An XGBoost model with K trees trained on dataset D can be represented as:
[0107]
[0108] in, It is the prediction result for the i-th sample; F is the set space of the base learners, f k Let q(x) represent the functional relationship between the structure q(x) and weight w of the k-th tree. The loss function can be expressed as:
[0109]
[0110] Where n represents the number of biomass samples, l represents the loss function representing the error between the predicted and actual biomass values, and Ω represents the complexity of the tree model with a sample size of n. The formula shows that the objective function consists of the sum of two functions, l(·) and Ω(·). Specifically, the first part is the sum of the errors between the predicted and actual values.
[0111] In addition, in this embodiment, a multivariate regression model was constructed based on two other algorithms, specifically a machine learning model based on random forest and a machine learning model based on gradient boosting decision tree.
[0112] Specifically, the Random Forest (RF) model, proposed by Breiman et al., combines the Bagging ensemble learning algorithm with a random subspace approach. This model is suitable for classification and regression analysis and can be used to predict categorical or continuous dependent variables. The Random Forest model is a model that incorporates decision trees trained using multiple bagging ensemble learning. It first uses bootstrapping to randomly and repeatedly select k samples from the original training set, where k represents the same size as the original training set. For each of these k samples, a decision tree model is constructed, resulting in k classification results. The final classification result is determined by aggregating the votes from each record based on these k classification results. The Random Forest model is simple to implement, exhibits strong generalization ability, and significantly outperforms adaptive boosting algorithms in terms of speed.
[0113] Gradient Boosting Decision Tree (GBDT) is a supervised learning algorithm in machine learning. It is an iterative decision tree algorithm, a high-performance nonlinear regression prediction algorithm, and can also be used for classification tasks. Its main idea is to gradually reduce the residuals through a continuous iterative process, constructing a multivariate regression decision tree through gradient optimization. Finally, the conclusions of all regression trees are summarized to form the final model, aiming to simultaneously reduce the model's variance and bias. In addition to strong interpretability and the ability to effectively handle mixed models, tree models also exhibit strong predictive performance and good stability.
[0114] In one exemplary embodiment, after performing step 208, the following may also be included:
[0115] The accuracy of the estimation results was evaluated using the coefficient of determination R. 2 The root mean square error (RMSE) is used as an evaluation metric, and the calculation formula is shown below:
[0116]
[0117] Where: y i Modeling values for biomass in small classes. This is the predicted value of the biomass of the sample plot. is the average biomass of the sample plots, and N is the number of sample plots.
[0118] This application also provides an application scenario that utilizes the aforementioned method for estimating aboveground biomass in subtropical forests. Specifically, in this embodiment, 319 data points were selected, including spaceborne lidar feature parameters, optical remote sensing single-band data (GF-2: B1-B4; GF-6: B1-B8), and single-band texture features, vegetation index features, and topographic features at different windows. Based on forest age classification, outliers were removed using a three-standard-deviation principle, ultimately selecting 311 data points for modeling: 142 from Masson pine, 87 from slash pine, and 82 from Quercus acutissima. 70% of the sample data was randomly selected as the training set for model construction, and the remaining 30% was used as the training set for model accuracy verification. Since the number of modeling factors exceeds the number of modeling samples, using all feature factors in modeling increases the risk of overfitting; therefore, feature factor selection is usually necessary before modeling. This embodiment utilizes a random forest feature algorithm built with Python for feature optimization, performing correlation analysis between feature factors and biomass for different tree species, and selecting factors significantly correlated with biomass.
[0119] like Figure 6 As shown, the biomass of *Pinus massoniana* showed better correlation with the h_canopy and h_canopy_abs features of ICESAT-2 / ATLAS images, and the correlations were all positive, with the h_canopy feature showing a correlation of 0.65. The biomass of *Pinus massoniana* showed the best correlation with the MTCI features of GF-2 & GF-6 images, with the vegetation index features of GF-6 images showing a better correlation than those of GF-2 images. The biomass of *Pinus massoniana* also showed high correlation with NDRE1, GF-2-B2, and h_canopy features of ICESAT-2 / ATLAS and GF-2 & GF-6 data, with absolute values of correlation coefficients ranging from 0.67 to 0.71.
[0120] like Figure 7 As shown, the biomass of slash pine had the strongest correlation with h_canopy of ICESAT-2 / ATLAS images, while the correlation with other features was not significant. The correlation with GF-2-B4, MTCI, and GF-6-B2 of GF-2 & GF-6 images was more significant. The feature correlation between slash pine biomass and ICESAT-2 / ATLAS and GF-2 & GF-6 data was significantly higher than that of single data sources, with the absolute value of the correlation coefficient of MTCI and GF-6-B5 reaching 0.7.
[0121] like Figure 8As shown, the biomass of *Quercus acutissima* showed no significant correlation with features extracted from ICESAT-2 / ATLAS imagery; however, it showed a more significant correlation with GF-6-B5, GF-6-B2, GF-2-B4, b4_mean_5, and MTCI from GF-2 & GF-6 imagery, with correlation coefficients ranging from 0.66 to 0.72; the correlation with combined ICESAT-2 / ATLAS and GF-2 & GF-6 data was even better, with the highest correlation of 0.7 with GF-6-B5 and the second highest correlation of 0.69 with MTCI. Topographic feature factors showed no significant correlation.
[0122] like Figure 9 As shown, the biomass of mixed coniferous and broad-leaved forests showed a high correlation with the h_canopy and h_canopy_abs features of ICESAT-2 / ATLAS images; the correlation with MTCI, GF-2-B4, and b4_mean_3 of GF-2 & GF-6 images was better, above 0.7, while the correlation of other features was around 0.6; the correlation between each feature and biomass was significantly improved when ICESAT-2 / ATLAS and GF-2 & GF-6 data were combined, with the correlation of MTCI, h_canopy, and b4_mean_5 above 0.70, among which MTCI showed a significant correlation with a correlation coefficient of 0.75; the correlation between biomass and features varied among different tree species under different data source combinations.
[0123] In summary, the feature selection based on ICESAT-2 / ATLAS and GF-2 & GF-6 data showed the best correlation, with GF-2-B2, MTCI, GF-6-B5, and h_canopy features showing significant correlation.
[0124] Given that different tree species exhibit varying adaptability to models under different data source combinations, this embodiment takes Masson pine, slash pine, Quercus acutissima, and mixed forests as examples. Using ICESAT-2 / ATLAS and GF-2&GF-6 as data sources, and combining measured data from forest resource surveys, RF, GBDT, and XGB models for different tree species are established to estimate forest biomass and compare model accuracy.
[0125] For Masson pine, in a single ICESAT-2 / ATLAS data source, the optimal model XGB has an R² of 0.82 and an RMSE of 10.184 t / ha; in a single GF-2 & GF-6 data source, it has better adaptability to RF, with an R² of 0.79 and an RMSE of 6.378 t / ha; under the combined ICESAT-2 / ATLAS and GF-2 & GF-6 data, the optimal model is RF with an accuracy of 0.88 and an RMSE of 4.685 t / ha.
[0126] For slash pine, the optimal model XGB has an R² of 0.81 and an RMSE of 7.886 t / ha in a single ICESAT-2 / ATLAS data source; in a single GF-2 & GF-6 data source, the modeling accuracy is best in GBDT with an R² of 0.86 and an RMSE of 8.912 t / ha; and the optimal model combining ICESAT-2 / ATLAS and GF-2 & GF-6 data is XGB with an accuracy of 0.89 and an RMSE of 9.587 t / ha.
[0127] For *Quercus acutissima*, the optimal model XGB has an R² of 0.85 and an RMSE of 19.447 t / ha in a single ICESAT-2 / ATLAS data source; the GBDT model also shows good adaptability in single GF-2 & GF-6 data with an R² of 0.85 and an RMSE of 18.235 t / ha; the optimal model for combining ICESAT-2 / ATLAS and GF-2 & GF-6 data is also XGB with an accuracy of 0.90 and an RMSE of 19.126 t / ha.
[0128] For mixed forests, the optimal model GBDT with a single ICESAT-2 / ATLAS data source has an R² of 0.79 and an RMSE of 24.724 t / ha; the model with the single GF-2 & GF-6 data source has the strongest RF suitability, with an accuracy of 0.82 and an RMSE of 18.440 t / ha; and the optimal model with the combined ICESAT-2 / ATLAS and GF-2 & GF-6 data source has an accuracy of 0.88 and an RMSE of 19.926 t / ha.
[0129] Specifically, the optimal model R for each tree species is based on different data sources. 2 The RMSE is shown in Table 3:
[0130] Table 3 shows the optimal model R for each tree species based on different data sources. 2 and RMSE
[0131]
[0132]
[0133] Specifically, a comparative analysis of the model suitability for different tree species based on the same data source was conducted. In the ICESAT-2 / ATLAS data, the XGB model showed better suitability for different tree species. In the GF-2 & GF-6 data, the GBDT model consistently showed the highest accuracy across different tree species, indicating that this model better adapts to the relationship between optical image features and AGB, achieving a better fit. Combining ICESAT-2 / ATLAS and GF-2 & GF-6 data across different tree species, the XGB model generally demonstrated stronger suitability, and R...2 Compared to a single data source, the RMSE varies due to different models overfitting or underfitting the data. Specifically, the optimal model RSE for different tree species based on the same data varies. 2 The RMSE is shown in Table 4:
[0134] Table 4 shows the optimal model R for different tree species based on the same data. 2 and RMSE
[0135]
[0136] Among them, comparative analyses were conducted on the same tree species based on different data sources and different models, and on different tree species and different models under the same data source. Finally, a model was constructed and the AGB of the study area was predicted based on the combined ICESAT-2 / ATLAS and GF-2 & GF-6 data, and forest biomass mapping and analysis were carried out. Figure 10 To assess the fitting performance of different models for the whole forest based on three tree species and mixed forests by combining ICESAT-2 / ATLAS and GF-2 & GF-6 data, the XGB test set R... 2 The XGB model yielded the best RMSE results, followed by GBDT and RF. In this embodiment, the XGB model will be selected to predict forest biomass in the study area and to generate maps and perform analysis.
[0137] In summary, the total forest biomass of National Forest Park in City A in 2023 was approximately 838,000 tons, of which Masson pine biomass was approximately 179,800 tons, slash pine biomass was approximately 113,300 tons, Quercus acutissima biomass was approximately 230,900 tons, and mixed coniferous and broad-leaved forest biomass was approximately 313,900 tons. Masson pine biomass accounted for 21.5% of the total forest biomass, slash pine biomass accounted for 13.5%, Quercus acutissima biomass accounted for 28.5%, and mixed coniferous and broad-leaved forest biomass accounted for 37.5%.
[0138] Based on the estimation of the three main tree species in the forest area—Massonia massoniana, slash pine, and Quercus acutissima—the remaining few tree species were treated as mixed coniferous and broad-leaved forests. For these four forest stand types, biomass distribution per unit area was mapped using the optimal model XGB. Figure 11The results of the three models for estimating biomass per unit area show differences. The RF and XGB models estimate values that are relatively close, ranging from 8 t / ha to 272 t / ha. The GBDT model estimates higher values, ranging from 26 t / ha to 413 t / ha. The RF model estimates lower values for the low biomass values, while the GBDT model estimates higher values for the high biomass values. In comparison, the XGB model's biomass estimation is more reasonable and effective. In this example, the DEM values in the area gradually decrease from the center outwards. The biomass values predicted by the three models consistently show high values in the central forest area and low values in the north and south, which corresponds to the DEM values. The terrain in the north and south is more fragmented, resulting in relatively lower biomass; the terrain in the central area is more complete, resulting in relatively higher biomass.
[0139] Based on the same inventive concept, this application also provides a data processing apparatus for implementing the above-described method for estimating aboveground biomass in subtropical forests. The solution provided by this apparatus is similar to the solution described in the above-described method; therefore, the specific limitations in one or more data processing apparatus embodiments provided below can be found in the limitations of the above-described method for estimating aboveground biomass in subtropical forests, and will not be repeated here.
[0140] In one exemplary embodiment, such as Figure 12 As shown, a data processing apparatus is provided, comprising:
[0141] The data acquisition module 1201 is used to acquire multi-source remote sensing and field survey data and historical ground feature spectral information of the target subtropical forest area; the multi-source remote sensing and field survey data includes lidar data, optical image data, ground survey data and topographic data.
[0142] The attribute extraction module 1202 is used to extract the attributes of the target subtropical forest region based on the photons in the lidar data; the attributes include the longitude, latitude, cloud aerosols, canopy height, canopy height uncertainty, absolute canopy height, best fit of segmented terrain height, and best surface elevation of the target subtropical forest region.
[0143] The vegetation index feature calculation module 1203 is used to calculate the vegetation index features of the target subtropical forest area based on the spectral information in the optical image data and the historical ground cover spectral information of the target subtropical forest area.
[0144] The surface reflectance determination module 1204 is used to determine the surface reflectance of the target subtropical forest area based on the band information in the optical image data.
[0145] The texture feature extraction module 1205 is used to extract the texture features of the target subtropical forest area based on the gray-level co-occurrence matrix obtained by second-order statistical filtering according to the surface reflectance of each single band; the texture features include mean, variance, contrast, dissimilarity, information entropy, cooperativity, second moment and correlation.
[0146] The terrain factor extraction module 1206 is used to extract terrain factors of the target subtropical forest area based on terrain data using 3D analysis tools; the terrain factors include slope factor, aspect factor and elevation factor.
[0147] The model building module 1207 is used to construct a multivariate regression model based on the attributes, vegetation index characteristics, texture characteristics and topographic factors of the target subtropical forest area; the multivariate regression model is a machine learning model based on XGBoost; the machine learning model based on XGBoost is constructed by adding a biomass conservation equation constraint term to the XGBoost loss function to build a physics-guided hybrid machine learning model.
[0148] The estimation module 1208 is used to estimate the aboveground biomass of the target subtropical forest area based on a multivariate regression model and output the estimation results.
[0149] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 13 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data for data processing. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When executed by the processor, the computer program implements a method for estimating aboveground biomass in subtropical forests.
[0150] Those skilled in the art will understand that Figure 13The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0151] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0152] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0153] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0155] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0156] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0157] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0158] This embodiment uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application; at the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. In summary, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for estimating aboveground biomass in subtropical forests, characterized in that, The estimation method includes: Acquire multi-source remote sensing and field survey data and historical ground feature spectral information for the target subtropical forest area; the multi-source remote sensing and field survey data includes lidar data, optical image data, ground survey data, and topographic data; Based on the photons in the lidar data, the attributes of the target subtropical forest region are extracted; the attributes include the longitude, latitude, cloud aerosols, canopy height, canopy height uncertainty, absolute canopy height, best fit of segmented terrain height, and best surface elevation of the target subtropical forest region. Based on the spectral information in the optical image data, and using the historical ground cover spectral information of the target subtropical forest area, the vegetation index characteristics of the target subtropical forest area are calculated. Based on the band information in the optical image data, the surface reflectance of the target subtropical forest area is determined; Based on the surface reflectance of each single band, the gray-level co-occurrence matrix obtained by second-order statistical filtering is used to extract the texture features of the target subtropical forest region; the texture features include mean, variance, contrast, dissimilarity, information entropy, cooperativity, second moment and correlation. Based on the topographic data, 3D analysis tools were used to extract the topographic factors of the target subtropical forest area; the topographic factors include slope factor, aspect factor and elevation factor. Based on the attributes, vegetation index characteristics, texture characteristics, and topographic factors of the target subtropical forest region, a multivariate regression model is constructed; the multivariate regression model is a machine learning model based on XGBoost; the machine learning model based on XGBoost is constructed by adding a biomass conservation equation constraint term to the XGBoost loss function to build a physics-guided hybrid machine learning model. The aboveground biomass of the target subtropical forest area is estimated based on a multivariate regression model, and the estimation results are output.
2. The method for estimating aboveground biomass in subtropical forests according to claim 1, characterized in that, The lidar data is ATL08 data.
3. The method for estimating aboveground biomass in subtropical forests according to claim 1, characterized in that, The optical image data includes GF-2PMS1 image data and GF-6WFV image data.
4. The method for estimating aboveground biomass in subtropical forests according to claim 1, characterized in that, The vegetation index features include normalized vegetation index, ratio vegetation index, difference vegetation index, enhanced vegetation index, soil-regulated vegetation index, greenness normalized vegetation index, infrared percentage vegetation index, MERIS terrestrial chlorophyll index, and normalized difference red edge 1 index.
5. The method for estimating aboveground biomass in subtropical forests according to claim 1, characterized in that, The specific calculation formulas for each index in the vegetation index features are as follows: NDVI=p NIR -r RED / r NIR +r RED ; DVI=ρ NIR -r RED ; IPVI=ρ NIR / r NIR +r RED ; Among them, RVI is the ratio vegetation index, NDVI is the normalized difference vegetation index, EVI is the enhanced vegetation index, DVI is the difference vegetation index, SAVI is the soil-modified vegetation index, GNDVI is the green normalized difference vegetation index, IPVI is the infrared percentage vegetation index, MTCI is the MERIS terrestrial chlorophyll index, NDRE1 is the normalized difference red edge 1 index, and ρ NIR For the near-infrared band of GF-2 imagery, ρ RED For the red edge band of the GF-2 image, L is set to 0.5, ρ BLUE For the blue-edge band of GF-2 imagery, ρ GREEN For the green edge band of GF-2 imagery, ρ 0.75 For GF-6 imagery in the red-edge II band, ρ 0.71 For the GF-6 image, red-edge band I, ρ RED1 This is the red-edge band of the GF-6 image.
6. The method for estimating aboveground biomass in subtropical forests according to claim 1, characterized in that, The XGBoost-based machine learning model in the multivariate regression model is specifically as follows: in, It is the prediction result of the i-th sample; F is the set space of the basic learners, f k The function represents the relationship between the structure q(x) and weight w of the k-th tree.
7. A device for estimating aboveground biomass in subtropical forests, characterized in that, The estimation device includes: The data acquisition module is used to acquire multi-source remote sensing and field survey data and historical ground feature spectral information of the target subtropical forest area; the multi-source remote sensing and field survey data includes lidar data, optical image data, ground survey data and topographic data; The attribute extraction module is used to extract the attributes of the target subtropical forest region based on the photons in the lidar data; the attributes include the longitude, latitude, cloud aerosols, canopy height, canopy height uncertainty, absolute canopy height, best fit of segmented terrain height, and best surface elevation of the target subtropical forest region. The vegetation index feature calculation module is used to calculate the vegetation index features of the target subtropical forest area based on the spectral information in the optical image data and the historical ground cover spectral information of the target subtropical forest area. The surface reflectance determination module is used to determine the surface reflectance of the target subtropical forest area based on the band information in the optical image data. The texture feature extraction module is used to extract the texture features of the target subtropical forest area based on the gray-level co-occurrence matrix obtained by second-order statistical filtering according to the surface reflectance of each single band; the texture features include mean, variance, contrast, dissimilarity, information entropy, cooperativity, second moment and correlation. The terrain factor extraction module is used to extract terrain factors of the target subtropical forest area based on terrain data using 3D analysis tools; the terrain factors include slope factor, aspect factor and elevation factor. The model building module is used to construct a multivariate regression model based on the attributes, vegetation index characteristics, texture characteristics and topographic factors of the target subtropical forest area; the multivariate regression model is a machine learning model based on XGBoost; the machine learning model based on XGBoost is constructed by adding a biomass conservation equation constraint term to the XGBoost loss function to build a physics-guided hybrid machine learning model. The estimation module is used to estimate the aboveground biomass of a target subtropical forest area based on a multivariate regression model and output the estimation results.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement a method for estimating aboveground biomass in a subtropical forest according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements a method for estimating aboveground biomass in a subtropical forest as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements a method for estimating aboveground biomass in a subtropical forest as described in any one of claims 1-6.