Methods and systems for estimating forest aboveground biomass density driven by multiple ecological neighborhoods
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]针对现有森林地上生物量密度估测方法主要依赖样地与遥感变量之间的全局统计关系、对样本局部特征相似关系和地理空间邻近关系利用不足、难以适应不同森林类型和生态背景下遥感变量响应差异、在空间独立测试和跨区域外推条件下容易出现精度下降及高值低估等问题,本发明提供一种多源生态邻域驱动的森林地上生物量密度估测方法及系统
(1) 本发明中,提供一种多源生态邻域驱动的森林地上生物量密度估测模型,通过表格特征编码模块、邻域记忆编码模块、环境路由编码模块、环境专家修正模块、门控融合模块和残差修正模块,实现多源遥感变量到森林 AGBD 的智能化非线性建模,提升了森林地上生物量密度估测的精度和稳定性。
Smart Images

Figure CN122571328A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of forest parameter remote sensing inversion and geospatial intelligent modeling technology, and in particular to a method and system for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods. Background Technology
[0002] Above-Ground Biomass Density (AGBD) is a crucial parameter characterizing the carbon storage, productivity, and ecological functions of forest ecosystems. It is also a fundamental indicator for forest carbon sequestration accounting, ecological protection and restoration, refined forest resource management, and natural capital accounting. Accurately obtaining forest AGBD at regional scales and even large-scale continuous spatial distributions is essential for supporting forest management decisions. Traditional forest above-ground biomass estimation primarily relies on plot surveys. Individual tree biomass is calculated using parameters such as diameter at breast height (DBH), tree height, tree species, and timber density, combined with allometric growth equations. This biomass is then aggregated by plot area to obtain plot-scale AGBD. While this method offers high accuracy in point measurements, plot surveys are costly, time-consuming, and have limited spatial coverage, making it difficult to meet the application requirements of large-scale, high-timeliness, and continuous spatial mapping.
[0003] With the development of remote sensing technology, multi-source remote sensing data has become an important data foundation for forest AGBD estimation. Optical remote sensing images can reflect vegetation chlorophyll, water content, canopy cover, and phenological status. Multispectral data such as Landsat and Sentinel-2, along with their constructed vegetation indices such as NDVI, EVI, and NDMI, and tassel-cap transformation characteristics, have been widely used to characterize forest structure and growth status. Synthetic Aperture Radar (SAR) has a certain canopy penetration capability and can reflect forest structure, water content, and scattering characteristics. Sentinel-1 C-band backscattering coefficient, ALOS / PALSARL band HH / HV polarization characteristics, and their ratios, differences, and texture features can provide important supplements to optical data. LiDAR can directly characterize the vertical structure and canopy height of forests. Airborne LiDAR, ICESat / GLAS, ICESat-2, and Global Ecosystem Dynamics Investigation (GEDI) data provide more direct three-dimensional structural information for forest AGBD estimation. In addition, auxiliary environmental variables such as digital elevation models, slope, land cover, canopy cover, vegetation height, leaf area index, and photosynthetically active radiation absorption ratio are often used in conjunction with optical, radar, and LiDAR data to enhance the model's ability to express differences in forest types, topographic backgrounds, and canopy structure. Saatchi et al. integrated ground forest plots, ICESat / GLAS LiDAR, MODIS, QuikSCAT, and SRTM data for spatial extrapolation to generate a benchmark map of forest carbon storage covering three major tropical regions. The NASA GEDI mission used the full-waveform space lidar on the International Space Station to acquire information on forest vertical structure, and generated footprint-level and 1km grid-level AGBD products based on waveform relative height, plant functional types, and regional stratification models, providing a new observational benchmark for global forest biomass estimation, model training, and product validation.
[0004] Based on the aforementioned multi-source remote sensing data, forest AGBD estimation methods have evolved from traditional empirical statistical models to machine learning, deep learning, and multi-source fusion modeling. Early studies mostly employed empirical models such as linear regression, exponential regression, and power function regression. These methods are simple in structure and highly interpretable, but are significantly affected by forest type, canopy saturation, topographic conditions, and sensor noise, resulting in limited accuracy in high-biomass forests and cross-regional applications. Subsequently, machine learning methods such as Random Forest (RF), Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), LightGBM, and CatBoost have been widely applied to the fusion modeling of multi-source remote sensing variables and sample plot observation data, capable of characterizing the complex nonlinear relationship between forest structure parameters and AGBD. Shu et al. fused Landsat 8 OLI, Sentinel-1A, and continuous forest resource inventory sample plot data, using linear regression, random forest, and XGBoost models to estimate subtropical forest AGB, validating the estimation advantages of machine learning models combined with optical-radar fusion data. With the development of geospatial artificial intelligence (GeoAI) and deep learning technologies, convolutional neural networks, attention mechanisms, Transformers, tabular deep learning networks, and multi-task learning models have been gradually introduced into the joint estimation of forest AGBD, forest height, and canopy cover. Deep learning methods can automatically learn implicit representations from high-dimensional, multi-source, and nonlinear features. However, under conditions of limited sample plots, uneven spatial distribution of samples, and significant differences in ecological backgrounds among different sites, problems such as insufficient spatial generalization ability, unstable regional extrapolation, underestimation of high-value biomass, and insufficient model interpretability still easily arise. Summary of the Invention
[0005] To address the shortcomings of existing forest aboveground biomass density estimation methods, which primarily rely on global statistical relationships between sample plots and remote sensing variables, fail to adequately utilize local feature similarities and geospatial proximity, struggle to adapt to differences in remote sensing variable responses across different forest types and ecological backgrounds, and are prone to accuracy degradation and underestimation of high values under conditions of independent spatial testing and cross-regional extrapolation, this invention provides a multi-source ecological neighborhood-driven method and system for estimating forest aboveground biomass density. This method constructs multi-source remote sensing table features, neighborhood memory features, and environmental routing features, and introduces multi-branch coding, environmental expert correction, gated fusion, and residual correction mechanisms into the multi-source ecological neighborhood network model MSEN-Net. This enables the forest aboveground biomass density estimation process to simultaneously utilize the nonlinear expressive power of multi-source remote sensing variables, constraints of locally similar samples, constraints of spatially adjacent samples, and differentiated mapping relationships under different environmental backgrounds. This achieves high-precision estimation of forest aboveground biomass density at the sample plot scale and generation of regional-scale forest aboveground biomass density raster products.
[0006] According to one aspect of the present invention, a method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods is provided, comprising: Obtain forest plot survey data and construct true values of forest aboveground biomass density at the plot scale based on individual tree measurement information and plot area; Acquire multi-source remote sensing and auxiliary environmental data corresponding to the spatial location and observation time of the sample plot, extract predictive variables related to optical, radar, topography, land cover and vegetation structure, and construct multi-source remote sensing tabular features; The true values of forest aboveground biomass density at the plot scale are correlated with the features of the multi-source remote sensing table according to the plot identification, observation time and spatial location to obtain the modeling sample table, and feature processing is performed on continuous predictive variables and categorical predictive variables. Based on the nearest neighbor relationships of training samples in feature space and geographic space, a neighborhood memory feature is constructed. Environmental prototype clustering is performed based on the multi-source remote sensing table features of the training samples, and environmental routing features are constructed based on the distance between the samples and the environmental prototypes and the soft attribution probability. The multi-source remote sensing table features, neighborhood memory features, and environmental routing features are input into the MSEN-Net model to output the predicted value of forest aboveground biomass density at the plot scale. The feature columns, standardized parameters, and category coding rules saved during the training phase are applied to the multi-source remote sensing raster data of the target area to perform pixel-level inference and generate regional-scale forest aboveground biomass density raster products.
[0007] Furthermore, the construction of the true values of aboveground biomass density in the forest at the plot scale includes: Obtain forest plot survey data including plot identifier, plot spatial location, observation time, individual tree identifier, tree species, diameter at breast height (DBH), basal diameter, tree height, measured height, and survival status; Based on the individual tree measurement fields, diameter range, and applicable conditions of the equation, select the allometric growth equation and calculate the aboveground biomass of the individual tree; Effective aboveground biomass per tree was aggregated according to plot size and observation time, and then standardized by plot area to obtain the true value of forest aboveground biomass density at the plot size.
[0008] Furthermore, the multi-source remote sensing and auxiliary environmental data includes at least one of optical remote sensing data, synthetic aperture radar data, digital elevation model, land cover data, canopy cover data, vegetation height data, leaf area index data, and photosynthetically active radiation absorption ratio data; the multi-source remote sensing table features include at least one of optical bands, vegetation index, moisture index, tassel cap transformation features, radar backscattering and polarization derivative features, topographic variables, land cover variables, and vegetation structure variables.
[0009] Furthermore, the feature processing includes: using the true value of forest aboveground biomass density at the sample plot scale as the target variable, using the latitude and longitude coordinates of the sample plot for spatial nearest neighbor retrieval and spatial division, standardizing the continuous predictive variables, and performing category coding or category embedding on the land cover category variables.
[0010] Furthermore, the construction of the neighborhood memory features includes: Calculate the feature space distance between samples based on the multi-source remote sensing table features of the samples, and retrieve the nearest neighbor samples in the feature space; Calculate the geospatial distance between samples based on the latitude and longitude coordinates of the sample plots, and retrieve geospatial nearest neighbor samples; Based on the feature space nearest neighbor samples and the geographic space nearest neighbor samples, the nearest neighbor prediction value, the nearest neighbor distance statistics, and the difference between the two are calculated to form a neighborhood memory feature vector. The nearest neighbor prediction value includes the feature space nearest neighbor prediction value and the geographic space nearest neighbor prediction value, and the nearest neighbor distance statistics include the feature space nearest neighbor distance statistics and the geographic space nearest neighbor distance statistics.
[0011] Furthermore, the construction of the environmental routing features includes: Clustering of multi-source remote sensing table features of training samples yields multiple environmental prototypes; Calculate the distance between the sample and each environmental prototype, and calculate the soft attribution probability of the sample to each environmental prototype based on the distance; The soft-attribution probability and environmental distance statistics are combined to form an environmental routing feature vector.
[0012] Furthermore, the MSEN-Net model includes a table feature encoding module, a neighborhood memory encoding module, an environment routing encoding module, an environment expert correction module, a gating module, and a residual correction module; The table feature encoding module is used to perform nonlinear encoding on multi-source remote sensing table features; the neighborhood memory encoding module is used to encode the feature space nearest neighbor prediction value, the geospatial nearest neighbor prediction value and their distance statistics; the environmental routing encoding module is used to encode the environmental prototype soft attribution probability and environmental distance statistics; the three types of features obtained from the three encoding modules are fused to obtain a fused representation; The environmental expert correction module includes a basic table prediction branch and multiple environmental expert correction branches to generate environmental perception prediction values. The basic table prediction branch takes the table-encoded feature vector as input and outputs a basic prediction value. The environmental expert correction branches include... Each expert sub-branch corresponds to an environmental prototype or a type of ecological background. The expert sub-branch takes the table-encoded feature vector as input and outputs the prediction correction amount corresponding to the environmental prototype. The basic prediction value is added to the environmental correction amount to obtain the environmental perception prediction value. The gating module is used to adaptively coordinate three types of candidate prediction results: environmental perception prediction, feature space nearest neighbor prediction, and geospatial nearest neighbor prediction, to obtain a gated fusion prediction value. The residual correction module corrects the gated fusion predictions to obtain the final predictions of the MSEN-Net model.
[0013] Furthermore, the gating module generates gating weights based on the fusion representation and gating statistical vectors, and uses these gating weights to adaptively weight and fuse the environmental perception prediction value, the feature space nearest neighbor prediction value, and the geospatial nearest neighbor prediction value. The gating statistical vector includes a neighborhood memory statistic and an environmental distance statistic. The neighborhood memory statistic is composed of the feature space nearest neighbor prediction value, the geospatial nearest neighbor prediction value, the feature space nearest neighbor distance statistic, the geospatial nearest neighbor distance statistic, and the difference term between the two types of nearest neighbor prediction values and the distance statistic. The environmental distance statistic is calculated from the minimum distance from the sample to each environmental prototype.
[0014] Furthermore, the generation of the regional-scale forest aboveground biomass density raster product includes: Obtain multi-source remote sensing prediction variable raster of the target area, and construct pixel-level table features according to the feature columns, standardized parameters and category coding rules saved during the training phase; Based on the feature spatial nearest neighbor relationship, geospatial nearest neighbor relationship, and environmental prototype soft attribution relationship between training samples and target pixels, pixel-level neighborhood memory features and environmental routing features are constructed. The pixel-level table features, pixel-level neighborhood memory features, and pixel-level environmental routing features are input into the trained MSEN-Net model to obtain pixel-level forest aboveground biomass density prediction values, which are then written into the target area raster to generate a regional-scale forest aboveground biomass density raster product.
[0015] The present invention also provides a forest aboveground biomass density estimation system driven by multiple ecological neighborhoods, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the program instructions in the memory to execute the forest aboveground biomass density estimation method driven by multiple ecological neighborhoods described in the above technical solution.
[0016] The application of the technical solution of the present invention has the following beneficial effects: (1) In this invention, a multi-source ecological neighborhood-driven forest aboveground biomass density estimation model is provided. Through a table feature encoding module, a neighborhood memory encoding module, an environmental routing encoding module, an environmental expert correction module, a gated fusion module and a residual correction module, intelligent nonlinear modeling of multi-source remote sensing variables into forest AGBD is realized, thereby improving the accuracy and stability of forest aboveground biomass density estimation.
[0017] (2) In this invention, by constructing neighborhood memory features, the feature spatial similarity relationship and geographic spatial proximity relationship between samples are utilized to introduce feature nearest neighbor prediction, spatial nearest neighbor prediction and distance statistics information, so that the model can make full use of local similar sample constraints and spatial dependence information, and improve the problem that traditional models mainly rely on global statistical relationships and are insufficient in expressing local forest stand differences.
[0018] (3) In this invention, by constructing environmental routing features and environmental expert correction modules, the environmental prototype soft attribution probability is used to characterize the environmental differences under different forest types, terrain backgrounds, canopy structures and remote sensing response states, and adaptive weighted correction is performed on the environmental expert branch, which improves the model's adaptability to the differences in remote sensing variable responses under different ecological backgrounds and enhances the cross-regional extrapolation and spatial generalization capabilities.
[0019] (4) In this invention, by organically combining multi-source remote sensing feature construction, MSEN-Net model training, and regional-scale raster inference, a regional-scale forest aboveground biomass density raster product can be generated. This method can integrate multi-source information such as optical, radar, topographic, land cover, and vegetation structure, providing reliable data support for forest carbon storage assessment, forest resource monitoring, carbon sink accounting, and ecological environment management.
[0020] The implementation method of this invention is simple and easy to implement, highly practical, and has a clear overall process structure. It can realize automated processing from sample plot true value processing, multi-source remote sensing variable construction, model training to regional scale AGBD mapping, effectively improving the accuracy, generalization ability and product reliability of forest aboveground biomass density estimation in complex forest scenarios. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the overall process of a multi-source ecological neighborhood-driven method and system for estimating forest aboveground biomass density in an embodiment of the present invention.
[0022] Figure 2 This is a schematic diagram of the MSEN-Net multi-source ecological neighborhood network model in an embodiment of the present invention.
[0023] Figure 3 This is a flowchart of steps 8 and 9 of the embodiments of the present invention for the raster prediction, product output and accuracy verification of forest aboveground biomass density at the regional scale. Detailed Implementation
[0024] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0025] This invention addresses the problems of traditional forest aboveground biomass density estimation methods, which rely heavily on the global statistical relationship between sample plots and remote sensing variables, fail to adequately utilize the local feature similarity and geospatial proximity relationships between sample plots, struggle to adapt to differences in remote sensing variable responses under different forest types and ecological backgrounds, and are prone to accuracy degradation and underestimation of high biomass under conditions of spatially independent testing and cross-regional extrapolation. To resolve these issues, this invention establishes the MSEN-Net hybrid deep learning model based on multi-source remote sensing table features, neighborhood memory features, and environmental routing features. Through table feature encoding modules, neighborhood memory encoding modules, environmental routing encoding modules, environmental expert modules, gating fusion modules, and residual correction modules, it jointly characterizes the nonlinear relationship between multi-source remote sensing variables and forest AGBD, local similarity sample constraints, spatial proximity sample constraints, and differentiated response relationships under different ecological backgrounds. This improves the AGBD estimation accuracy, spatial generalization ability, and regional raster product generation stability in complex forest scenarios.
[0026] like Figure 1 As shown, the process of an embodiment of the present invention specifically includes the following steps: Step 1: Constructing ground truth data for sample plots.
[0027] Step 1.1: Acquisition of Plot Survey Data. Acquire vegetation structure survey data, including fields such as plot number, plot location, observation time, individual tree number, tree species, diameter at breast height (DBH), tree height, survival status, timber density, or species information that can be used to match timber density. Organize the forest plot survey data into an individual tree record table and a plot record table. The individual tree record table is used to calculate the aboveground biomass of individual trees, and the plot record table provides plot location, observation year, plot area, and plot-scale true value correlation information.
[0028] Step 1.2: Selection of Allometric Growth Equation. For each individual tree record, a measurement field integrity check is performed. Based on the individual tree measurement fields, diameter at breast height (DBH) range, and equation applicability conditions, a corresponding allometric growth equation is selected for each valid individual tree. When DBH measurements exist and meet the equation applicability conditions, allometric growth equations such as Jenkins or Chojnacky can be used; when DBH is within a smaller diameter class, calculations can be performed using a combination of taper function and correction coefficients; when DBH is unavailable but basal diameter and tree height are available, allometric growth equations such as Annighöfer can be used; when key fields are missing or cannot be matched with an applicable equation, a corresponding quality marker is set, and the tree is not included in the effective individual tree aboveground biomass summary. The first of the sample plots The aboveground biomass of a single tree can be summarized as follows: (1) In the formula, Indicates the first The first of the sample plots Aboveground biomass of a single tree; Indicates the first The first of the sample plots Diameter at breast height (DBH) of a single tree; Indicates the first The first of the sample plots The base diameter of a single tree; Indicates the first The first of the sample plots The height of a single tree; Indicates the first The first of the sample plots A single tree species or functional type of plant; This represents the allometric growth equation determined based on individual tree measurement fields, tree species, plant functional type, or regional experience.
[0029] In practical implementation, the allometric growth equations include, but are not limited to, the Jenkins, Chojnacky, and Annighöfer equations. As an example, when breast diameter measurements are available and the applicable conditions of the corresponding equations are met, the Jenkins equation can be expressed as:
[0030] The Chojnacky equation can be expressed as:
[0031] When measurements of basal diameter and tree height are available, the Annighöfer equation can be expressed as:
[0032] In the formula, , and The aboveground biomass of a single tree is represented by the Jenkins equation, the Chojnacky equation, and the Annighöfer equation, respectively. Indicates the diameter at breast height (DBH) of a single tree; Indicates the diameter of a single wood stalk; Indicates the height of a single tree; , , , , and Indicates empirical parameters determined by tree species, plant functional type, regional experience, or corresponding allometric growth equation parameter tables; Represents an exponential function; This represents the natural logarithm function.
[0033] Step 1.3: Aboveground biomass aggregation at the plot scale. Grouping and summing the individual tree records marked "valid" within the same plot and the same observation year yields the total aboveground biomass of the plot. The total aboveground biomass of the i-th plot can be expressed as:
[0034] In the formula, Indicates the first Total aboveground biomass of each sample plot; Indicates the first The first of the sample plots Aboveground biomass of a single tree; Indicates the first The number of valid individual trees included in the statistics within each sample plot; This represents the summation of all valid individual trees within the sample plot.
[0035] Step 1.4: Plot Area Matching and Area Standardization. The total aboveground biomass of each plot is matched with its corresponding plot area. The total aboveground biomass is then standardized based on the effective plot area to obtain the true value of forest aboveground biomass density at the plot scale. The true value of aboveground biomass density for each sample plot can be expressed as:
[0036] In the formula, Indicates the first True values of forest aboveground biomass density in each sample plot; Indicates the first The total cumulative aboveground biomass of all effective individual trees within each sample plot; Indicates the first The effective area of each sample plot; This represents the area unit conversion factor, used to convert the total aboveground biomass of a sample plot to aboveground biomass density per unit area.
[0037] Step 1.5: Quality Check and Table Compilation of Plot-Scale True Values. A quality check is performed on the AGBD true values at the plot scale, eliminating plot records with missing plot areas or insufficient effective single tree counts. The AGBD true values at the plot scale that pass the quality check are correlated with the plot identifier, observation year, and latitude and longitude coordinates to form a plot true value table. The plot true value table includes information such as plot identifier, observation time, spatial location, plot area, and true values of forest aboveground biomass density at the plot scale.
[0038] Step 2: Construction of multi-source remote sensing predictive variables.
[0039] Multi-source remote sensing and auxiliary environmental data corresponding to the spatial location and observation time of the sample plots were acquired, and multi-source remote sensing predictive variables required for forest AGBD estimation were constructed. The multi-source remote sensing and auxiliary environmental data include, but are not limited to, multispectral remote sensing data such as Landsat and Sentinel-2, synthetic aperture radar data such as Sentinel-1 and ALOS / PALSAR, and auxiliary environmental data such as digital elevation models, land cover, canopy cover, vegetation height, leaf area index, and photosynthetically active radiation absorption ratio. Cloud masking, radiometric calibration, topographic correction, spatial resampling, coordinate system 1, and variable screening were performed on the above data. Using the spatial location and observation year as indexes, the multi-source remote sensing predictive variables were extracted to the sample plot scale. For missing remote sensing predictive variables, valid observations from adjacent years at the same spatial location with a time interval not exceeding 2 years were preferentially used to replace them, resulting in sample plot-scale predictive variable records corresponding to the spatial location and observation time of each sample plot.
[0040] Step 3: Modeling sample table generation and feature preprocessing.
[0041] Step 3.1: Variable Preprocessing and Multi-Source Remote Sensing Feature Construction. The ground truth values of the sample plot-scale AGBD obtained in Step 1 are correlated with the multi-source remote sensing predictive variables obtained in Step 2 to obtain a modeling sample table. The sample plot-scale AGBD ground truth values are used as the target variable; the latitude and longitude coordinates of the sample plots are used to support spatial nearest neighbor retrieval and spatial partitioning; discrete fields such as land cover are treated as categorical predictive variables; and the remaining optical, radar, topographic, and ecological environment variables are treated as continuous predictive variables, forming a multi-source remote sensing table feature for model input. The continuous predictive variables are standardized, and the calculation method can be expressed as follows:
[0042] In the formula, Indicates the first The first sample The standardized values of a continuous variable; This represents the variable value before standardization; Indicates the first training set The mean of a continuous variable; Indicates the first training set The standard deviation of a continuous variable.
[0043] For discrete variables such as land cover, class embedding is used to convert them into continuous vectors:
[0044] In the formula, Indicates the first The values of the categorical variable for each sample; This represents a category embedding mapping function; This represents the embedding vector corresponding to the categorical variable.
[0045] After the standardization of continuous predictor variables and the embedding of categorical predictor variables described above, the standardized continuous features and categorical embedding vectors are combined according to a preset feature order to form a multi-source remote sensing table feature vector for model input:
[0046] In the formula, Indicates the first Multi-source remote sensing table feature vectors of each sample; Indicates the first Standardized values of each continuous predictor variable in each sample; Indicates the first The class embedding vector of each sample; This indicates the number of continuous predictor variables.
[0047] Step 3.2: Target Variable Transformation. To reduce the scale difference of the target values and enhance the model's learning ability on high biomass samples, the target values are logarithmically transformed and standardized:
[0048] In the formula, Indicates the first The original AGBD ground truth values for each sample; This represents the transformed training objective; Indicates training set The mean; Indicates training set Standard deviation; This represents the natural logarithm function.
[0049] Step 4: Constructing neighborhood memory features.
[0050] Step 4.1: Nearest Neighbor Search in Feature Space. Based on the standardized continuous and categorical variable encoding results obtained in Step 3, calculate the similarity of samples in the feature space. The sample and the first The feature space distance between samples can be expressed as:
[0051] In the formula, Indicates the first The sample and the first The distance between each sample in the feature space; and They represent the first The first sample and the first The feature vector of each sample after standardization and encoding; Represents Euclidean distance. Searching for the [number]th [item] in ascending order of feature space distance. one sample A set of nearest neighbor samples is formed in the feature space. .
[0052] Step 4.2: Geospatial Nearest Neighbor Search. Calculate the geospatial distance between samples based on the latitude and longitude coordinates of the sample plots. The sample and the first The geographic spatial distance between samples can be expressed as:
[0053] In the formula, Indicates the first The sample and the first Distance in the geographic space of each sample; and They represent the first The first sample and the first The spatial location vector of this sample; Represents Euclidean distance. Searches for the [number]th [item] in ascending order of geographic spatial distance. one sample A set of nearest neighbor samples is formed in the feature space. .
[0054] Step 4.3: Nearest Neighbor Prediction Calculation. Based on the feature space nearest neighbor set and the geospatial nearest neighbor set, calculate the predicted value of the first nearest neighbor. Feature spatial nearest neighbor prediction and geospatial nearest neighbor prediction for each sample:
[0055] In the formula, Indicates the first The feature space nearest neighbor prediction value of each sample; Indicates the first Geographically predicted nearest neighbor values for each sample; Indicates the first The set of nearest neighbor samples of a sample in the feature space; Indicates the first The set of nearest neighbor samples in geographic space for each sample; Represents the first in the feature space The nearest neighbor sample pair of the nth The weights of each sample; Represents the first in geographic space The nearest neighbor sample pair of the nth The weights of each sample; Indicates the first The true value or the nearest neighbor baseline prediction of the nearest neighbor sample.
[0056] Step 4.4: Neighborhood Memory Feature Construction. The predicted neighbor values, distance statistics, and differences between the two are combined to form the next... Neighborhood memory feature vectors of each sample:
[0057] In the formula, Indicates the first The neighborhood memory feature vector of each sample; Indicates the first Distance statistics between a sample and its nearest neighbors in the feature space; Indicates the first Distance statistics between a sample and its geographic nearest neighbor samples; This indicates the difference between two types of nearest neighbor predictions; This indicates the difference between spatial distance statistics and feature distance statistics.
[0058] Step 5: Constructing environmental routing features.
[0059] Step 5.1: Environmental Prototype Clustering. Based on the multi-source remote sensing table features of the training samples, perform K-means clustering or other clustering processes to obtain... An environmental prototype. The environmental prototype is used to represent the typical sample distribution under different forest types, topographic backgrounds, canopy structures, and remote sensing response states.
[0060] Step 5.2: Environmental Distance Calculation. Calculate the distance between the sample to be processed and the centers of each environmental prototype to characterize the proximity of the sample to different environmental prototypes. The distances from the sample to each environmental prototype center and the set of distances to all environmental prototype centers can be represented as follows:
[0061] In the formula, Indicates the first Multi-source remote sensing table feature vectors of each sample; Indicates the first The center vector of an environmental prototype; Indicates the first The sample to the first The distance to the environmental prototype; Indicates the first The set of distances from each sample to the entire environmental prototype; This represents the Euclidean distance.
[0062] Step 5.3: Calculation of environmental soft attribution probability. The nth sample pair The soft attribution probability of an environmental prototype can be expressed as:
[0063] In the formula, Indicates the first The nth sample pair The soft attribution probability of an environmental prototype; Indicates the first The sample to the first The distance to the environmental prototype; Indicates the first The sample to the first The distance to the environmental prototype; Temperature parameter indicating the sharpness of the soft attribution probability distribution; Indicates the number of environment prototypes; This indicates summation over all environment prototypes; Indicates the summation index.
[0064] Step 5.4: Constructing Environmental Routing Features. The environmental routing feature vector of a sample can be represented as:
[0065] In the formula, Indicates the first Environmental routing feature vectors for each sample; They represent the first The sample pairs from the 1st to the 1st The soft attribution probability of an environmental prototype; Indicates the first The minimum distance from each sample to all environmental prototypes.
[0066] Step 6: MSEN-Net model construction and training, where the network structure of the MSEN-Net model is as follows: Figure 2 As shown.
[0067] Step 6.1: Multi-branch coding and fusion representation construction. The multi-source remote sensing tabular features obtained in Step 3 are used... The neighborhood memory features obtained in step 4 and the environmental routing features obtained in step 5 Input a multi-source ecological neighborhood network model. The model includes a table feature encoding module, a neighborhood memory encoding module, an environmental routing encoding module, an environmental expert correction module, a gating module, and a residual correction module. Specifically, the table feature encoding module performs nonlinear encoding on multi-source remote sensing table features such as optical, radar, topographic, land cover, and vegetation structure features; the neighborhood memory encoding module encodes the predicted nearest neighbors in the feature space, the predicted nearest neighbors in the geospatial space, and their distance statistics; and the environmental routing encoding module encodes the soft-attribution probability of environmental prototypes and environmental distance statistics. The three-class encoded feature vectors of a sample can be represented as:
[0068] Three types of representations are obtained through the corresponding encoding module, and then the vectors are concatenated. The fusion representation of the samples can be expressed as:
[0069] In the formula, This represents the table feature encoding module. This indicates the neighborhood memory encoding module. This indicates the environment routing coding module; Indicates the first Fusion representation of individual samples; This represents the multi-source remote sensing table feature representation output by the table feature encoding module; This represents the neighborhood memory representation output by the neighborhood memory encoding module; This represents the environmental routing representation output by the environmental routing coding module; This indicates a vector concatenation operation.
[0070] Step 6.2: Construction of the Environmental Expert Correction Module. This module includes a basic table prediction branch and multiple environmental expert correction branches. The basic table prediction branch encodes feature vectors using tables. As input, it is used to learn the basic mapping relationship between multi-source remote sensing table features and forest AGBD, and its output is the basic prediction value:
[0071] In the formula, Indicates the predicted branch of the base table; Indicates the first The base predicted value for each sample.
[0072] Environmental experts' revision branches include There are several expert sub-branches, each corresponding to an environmental prototype or a type of ecological background. Each environmental expert branch uses tabular encoding of feature vectors. As input, the output is the prediction correction amount corresponding to this environment prototype:
[0073] In the formula, Indicates the first An environmental expert revised the branch; Indicates the first The environmental expert branch on the first The correction amount for each sample output.
[0074] Based on the environmental soft affiliation probability obtained in step 5 ,right The outputs of each environmental expert branch are weighted and aggregated to obtain the first... Environmental correction amount per sample:
[0075] Adding the base prediction value to the environmental correction value yields the environmental perception prediction value:
[0076] In the formula, Indicates the first Environmental correction amount for each sample; Indicates the first The nth sample pair The soft attribution probability of an environmental prototype; Indicates the environmental correction factor; This represents the predicted value based on environmental perception.
[0077] Step 6.3: Gating Module Construction. Based on the feature space nearest neighbor results and geospatial nearest neighbor results obtained in Step 4, and the environmental prototype distance obtained in Step 5, a gating statistical vector is constructed. This is used as a low-dimensional statistic directly input into the gating module. This includes neighborhood memory statistics and environmental distance statistics. The neighborhood memory statistics are derived from the feature space nearest neighbor prediction value, the geographic space nearest neighbor prediction value, and the feature space nearest neighbor distance statistics obtained in step 4. Geographical proximity distance statistics , as well as the difference between the two types of nearest neighbor predictions and distance statistics. and Preferably, an inverse distance-weighted statistic of the sample distance within the corresponding nearest neighbor set is used, and logarithmic transformation and standardization are applied to reduce the impact of different distance units and extreme distance values on the gating network. The environmental distance statistic is derived from step 5... Distance from each sample to each environmental prototype The calculation yields the optimal choice based on the minimum environmental distance. After logarithmic transformation and standardization, the environment term is used in the gating statistical vector. The gating statistical vector can be expressed as:
[0078] In the formula, Indicates the first The gating statistical vector of each sample; and These represent the feature space nearest neighbor prediction value and the geographic space nearest neighbor prediction value, respectively. ) and These represent the feature space nearest neighbor distance statistics and the geographic space nearest neighbor distance statistics, respectively. Indicates the first The minimum distance from each sample to all environmental prototypes.
[0079] To adaptively coordinate the three types of candidate prediction results—environmental perception prediction, feature space nearest neighbor prediction, and geospatial nearest neighbor prediction—the fused representation, neighborhood memory vector, and environmental routing vector are input into the gating module to generate the first... The gating weights corresponding to each sample. These gating weights can be expressed as:
[0080] In the formula, Indicates the first The gating weight vector of each sample; Indicates fusion representation; Represents the gating statistics vector; Indicates the gate control module; This represents the normalized exponential function.
[0081] Based on the aforementioned gating weights, the environmental perception prediction value, the feature space nearest neighbor prediction value, and the geospatial nearest neighbor prediction value are weighted and fused to obtain the gating fused prediction value:
[0082] In the formula, This represents the predicted value after gating and fusion; , and Let represent the gating weights corresponding to environmental perception prediction, feature space nearest neighbor prediction, and geospatial nearest neighbor prediction, respectively, and satisfy . ; This represents the environmental perception prediction value obtained by the environmental expert correction module; Represents the nearest neighbor prediction value in the feature space; This represents the predicted value of the geographic nearest neighbor.
[0083] Step 6.4: Residual Correction Module Construction. To further compensate for any potential systematic biases after gated fusion prediction, a residual correction module is constructed to perform secondary correction on the gated fusion prediction value. The residual correction module takes the concatenated result of the fusion representation vector, gated statistical vector, and gated weight vector as input. It outputs a sample-level residual correction amount, which is then weighted according to the residual correction coefficient and superimposed onto the gated fusion prediction value to obtain the final prediction value.
[0084] In the formula, This represents the final predicted value output by the MSEN-Net model; This represents the gated fusion prediction value; Indicates the residual correction strength coefficient; This indicates the residual correction module; This indicates that the residual correction input is formed by concatenating the fusion representation vector, the gated statistical vector, and the gated weight vector.
[0085] Step 6.5: Loss Function Construction. Forest aboveground biomass density samples exhibit a clear long-tail distribution characteristic, with relatively few samples of medium to high values, but these samples have a significant impact on the identification of high biomass areas, extreme value estimation, and upper limits of regional mapping. This invention constructs a weighted robust loss function that balances the focus on high-value samples with the suppression of outlier residuals, based on the... The target value of each sample is used to construct sample weights, so that samples with larger target values contribute more to the training. The training weights of each sample can be expressed as:
[0086] In the formula, Indicates the first Training weights for each sample; This represents the weight adjustment coefficient for high-value samples; This represents the weight truncation threshold, which limits the upper bound of the weights of extremely high-value samples, thus preventing a small number of extreme samples from exerting an excessive pull on the overall training process. This represents the target value used for weight construction, which is the target value of forest aboveground biomass density after logarithmic transformation; This indicates a function that takes the smaller value.
[0087] A weighted robust regression loss is constructed to reduce the adverse effects of outliers, noisy observations, or locally mismatched residuals on parameter updates, thereby improving the model's extreme value estimation ability and regional mapping robustness. The weighted robust regression loss function can be expressed as:
[0088] In the formula, Represents the total loss function; Indicates the number of training samples; This represents the robust regression loss function; This represents the model's predicted value in the transformed target space; Indicates the first The true value of each sample in the transformed target space.
[0089] Step 7: Model learning, training, and accuracy verification.
[0090] The modeling sample set obtained in step 3 is divided into training, validation, and test sets according to a preset random partitioning rule. The MSEN-Net model constructed in step 6 is then used to initialize parameters and iteratively train using the training set to obtain a converged forest AGBD estimation model. The validation set is used for model tuning and monitoring of the training process, and the test set is used to evaluate the model's estimation ability under the condition of identically distributed samples. Furthermore, samples from different spatial regions are divided into training, validation, and test sets based on the latitude and longitude coordinates of the sample plots for spatially independent validation to evaluate the model's spatial generalization ability under cross-regional extrapolation conditions. Model accuracy can be evaluated using metrics such as coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), and bias. It can also be compared with random forest (RF), extreme gradient boosting (XGBoost), support vector regression (SVR), and tabular deep neural network models, including but not limited to TabNet and FT-Transformer.
[0091] Step 8: Regional-scale forest AGBD raster prediction and mapping.
[0092] Obtain multi-source remote sensing prediction variable raster data for the target area, ensuring consistency with the feature column order, missing value handling rules, standardization parameters, and category coding rules used during model training. For effective forest pixels within the target area, extract multi-source remote sensing variables, spatial location, and environmental routing information to construct pixel-level table features, neighborhood memory features, and environmental routing features. Input these features into the MSEN-Net model trained in step 7 to achieve pixel-by-pixel inference prediction of forest AGBD in the target area. The AGBD prediction value for each effective pixel can be expressed as:
[0093] In the formula, Indicates the first Forest AGBD prediction for each effective cell; This represents the completed training of the MSEN-Net model; Indicates model parameters; Indicates the first Multi-source remote sensing table feature vectors of 1 effective pixel; Indicates the first Neighborhood memory vectors of effective pixels; Indicates the first An environment routing vector for each valid cell.
[0094] The prediction results are written back to the target area raster, and invalid values are assigned to pixels outside the non-forest area and the prediction area to generate a regional-scale forest aboveground biomass density raster product. The product is output as a separate GeoTIFF file, and the coordinate system, spatial resolution, invalid values, units, and effective forest mask information are recorded simultaneously.
[0095] Step 9: Verify the accuracy of AGBD products at the regional scale.
[0096] like Figure 3 As shown, the regional-scale forest AGBD raster product generated in step 8 is compared with the true values from the sample plot survey and external forest biomass reference products. Error indices and spatial consistency indices of the regional-scale product are calculated to evaluate the reliability of the AGBD raster product under different forest types, topographic conditions, and biomass levels. The coefficient of determination is calculated based on the true values from the sample plot survey. Accuracy indicators include root mean square error (RMSE), mean absolute error (MAE), and mean deviation (Bias).
[0097] In one embodiment, the external forest biomass reference product is the FIA BIGMAP aboveground biomass product. First, the FIA BIGMAP is processed by unit conversion, spatial registration, resampling, cell alignment, and effective forest masking to ensure consistency with the AGBD raster product generated in step 8 in terms of units, coordinate system, spatial resolution, and evaluation range. Based on the consistent effective evaluation pixels, the correlation coefficient, average difference, relative deviation, and spatial distribution consistency between the AGBD product generated by this invention and the FIA BIGMAP product are calculated.
[0098] In summary, this embodiment provides a multi-source ecological neighborhood-driven method for estimating forest aboveground biomass density. Starting from sample plot ground truth construction and multi-source remote sensing predictive variable extraction, it organically combines tabular feature encoding, neighborhood memory modeling, environmental routing modeling, environmental expert correction, gating fusion, and residual correction. It fully utilizes the nonlinear expressive power of multi-source remote sensing variables, the local feature similarity between samples, geospatial proximity, and response differences under different ecological and environmental backgrounds to achieve intelligent estimation and raster mapping of forest aboveground biomass density at the regional scale. In the test set of this embodiment, the MSEN-Net model achieved [results] under randomized test conditions. The prediction accuracy was 0.7503 with an RMSE of 86.93 Mg / ha; compared to the two best-performing comparative models besides our method, Tree Ensemble Ridge and TabPFN, These represent improvements of approximately 28.1% and 4.2% respectively, while RMSE decreased by approximately 22.4% and 5.5% respectively. Under spatially independent partitioning test conditions, the MSEN-Net model achieved... The concentration was 0.3902, and the RMSE was 109.58 Mg / ha; compared to TreeEnsemble Ridge and TabPFN, The results showed improvements of approximately 45.2% and 27.3% respectively, and reductions of approximately 8.7% and 6.2% in RMSE respectively, indicating that the method has good estimation performance in both routine tests and spatial extrapolation tests.
[0099] In practice, the method proposed in this invention can be automated by those skilled in the art using computer software technology, including processes such as sample plot ground truth processing, multi-source remote sensing feature construction, model training, model verification, regional grid inference, and product output.
[0100] This invention also provides a multi-source ecological neighborhood-driven forest aboveground biomass density estimation system, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the program instructions in the memory to execute the multi-source ecological neighborhood-driven forest aboveground biomass density estimation method described above.
[0101] Finally, it should be noted that the specific embodiments described herein are merely illustrative of the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can modify, supplement, or substitute the described specific embodiments without departing from the spirit and substance of the present invention, but such modifications, supplements, or substitutions should all fall within the scope of protection defined by the appended claims.
Claims
1. A method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods, characterized in that, include: Obtain forest plot survey data and construct true values of forest aboveground biomass density at the plot scale based on individual tree measurement information and plot area; Acquire multi-source remote sensing and auxiliary environmental data corresponding to the spatial location and observation time of the sample plot, extract predictive variables related to optical, radar, topography, land cover and vegetation structure, and construct multi-source remote sensing tabular features; The true values of forest aboveground biomass density at the plot scale are correlated with the features of the multi-source remote sensing table according to plot identification, observation time and spatial location to obtain a modeling sample table, and feature processing is performed on continuous predictive variables and categorical predictive variables. Based on the nearest neighbor relationships of training samples in feature space and geographic space, a neighborhood memory feature is constructed. Environmental prototype clustering is performed based on the multi-source remote sensing table features of the training samples, and environmental routing features are constructed based on the distance between the samples and the environmental prototypes and the soft attribution probability. The multi-source remote sensing table features, neighborhood memory features, and environmental routing features are input into the MSEN-Net model to output the predicted value of forest aboveground biomass density at the plot scale. The feature columns, standardized parameters, and category coding rules saved during the training phase are applied to the multi-source remote sensing raster data of the target area to perform pixel-level inference and generate regional-scale forest aboveground biomass density raster products.
2. The method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods according to claim 1, characterized in that: The construction of the true values of forest aboveground biomass density at the plot scale includes: Obtain forest plot survey data including plot identifier, plot spatial location, observation time, individual tree identifier, tree species, diameter at breast height (DBH), basal diameter, tree height, measured height, and survival status; Based on the individual tree measurement fields, diameter range, and applicable conditions of the equation, select the allometric growth equation and calculate the aboveground biomass of the individual tree; Effective aboveground biomass per tree was aggregated according to plot size and observation time, and then standardized by plot area to obtain the true value of forest aboveground biomass density at the plot size.
3. The method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods according to claim 1, characterized in that: The multi-source remote sensing and auxiliary environmental data include at least one of optical remote sensing data, synthetic aperture radar data, digital elevation model, land cover data, canopy cover data, vegetation height data, leaf area index data, and photosynthetically active radiation absorption ratio data; the multi-source remote sensing table features include at least one of optical bands, vegetation index, moisture index, tassel cap transformation features, radar backscattering and polarization derivative features, topographic variables, land cover variables, and vegetation structure variables.
4. The method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods according to claim 1, characterized in that: The feature processing includes: using the true value of forest aboveground biomass density at the sample plot scale as the target variable, using the latitude and longitude coordinates of the sample plot for spatial nearest neighbor retrieval and spatial division, standardizing continuous predictive variables, and performing category coding or category embedding on land cover category variables.
5. The method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods according to claim 1, characterized in that: The construction of the neighborhood memory features includes: Calculate the feature space distance between samples based on the multi-source remote sensing table features of the samples, and retrieve the nearest neighbor samples in the feature space; Calculate the geospatial distance between samples based on the latitude and longitude coordinates of the sample plots, and retrieve geospatial nearest neighbor samples; Based on the feature space nearest neighbor samples and the geographic space nearest neighbor samples, the nearest neighbor prediction value, the nearest neighbor distance statistics, and the difference between the two are calculated to form a neighborhood memory feature vector. The nearest neighbor prediction value includes the feature space nearest neighbor prediction value and the geographic space nearest neighbor prediction value, and the nearest neighbor distance statistics include the feature space nearest neighbor distance statistics and the geographic space nearest neighbor distance statistics.
6. The method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods according to claim 1, characterized in that: The construction of the environmental routing features includes: Clustering of multi-source remote sensing table features of training samples yields multiple environmental prototypes; Calculate the distance between the sample and each environmental prototype, and calculate the soft attribution probability of the sample to each environmental prototype based on the distance; The soft-attribution probability and environmental distance statistics are combined to form an environmental routing feature vector.
7. The method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods according to claim 1, characterized in that: The MSEN-Net model includes a table feature encoding module, a neighborhood memory encoding module, an environmental routing encoding module, an environmental expert correction module, a gating module, and a residual correction module. The table feature encoding module is used to perform nonlinear encoding on multi-source remote sensing table features; the neighborhood memory encoding module is used to encode the feature space nearest neighbor prediction value, the geospatial nearest neighbor prediction value and their distance statistics; the environmental routing encoding module is used to encode the environmental prototype soft attribution probability and environmental distance statistics; the three types of features obtained from the three encoding modules are fused to obtain a fused representation; The environmental expert correction module includes a basic table prediction branch and multiple environmental expert correction branches to generate environmental perception prediction values. The basic table prediction branch takes the table-encoded feature vector as input and outputs a basic prediction value. The environmental expert correction branches include... Each expert sub-branch corresponds to an environmental prototype or a type of ecological background. The expert sub-branch takes the table-encoded feature vector as input and outputs the prediction correction amount corresponding to the environmental prototype. The basic prediction value is added to the environmental correction amount to obtain the environmental perception prediction value. The gating module is used to adaptively coordinate three types of candidate prediction results: environmental perception prediction, feature space nearest neighbor prediction, and geospatial nearest neighbor prediction, to obtain a gated fusion prediction value. The residual correction module corrects the gated fusion predictions to obtain the final predictions of the MSEN-Net model.
8. The method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods according to claim 7, characterized in that: The gating module generates gating weights based on the fusion representation and gating statistical vectors, and uses these gating weights to adaptively weight and fuse the environmental perception prediction value, the feature space nearest neighbor prediction value, and the geospatial nearest neighbor prediction value. The gating statistical vector includes neighborhood memory statistics and environmental distance statistics. The neighborhood memory statistics consist of the feature space nearest neighbor prediction value, the geospatial nearest neighbor prediction value, the feature space nearest neighbor distance statistics, the geospatial nearest neighbor distance statistics, and the difference term between the two types of nearest neighbor prediction values and the distance statistics. The environmental distance statistics are calculated from the minimum distance from the sample to each environmental prototype.
9. The method for estimating forest aboveground biomass density driven by multi-source ecological neighborhoods according to claim 1, characterized in that: The generation of the regional-scale forest aboveground biomass density raster product includes: Obtain multi-source remote sensing prediction variable raster of the target area, and construct pixel-level table features according to the feature columns, standardized parameters and category coding rules saved during the training phase; Based on the feature spatial nearest neighbor relationship, geospatial nearest neighbor relationship, and environmental prototype soft attribution relationship between training samples and target pixels, pixel-level neighborhood memory features and environmental routing features are constructed. The pixel-level table features, pixel-level neighborhood memory features, and pixel-level environmental routing features are input into the trained MSEN-Net model to obtain pixel-level forest aboveground biomass density prediction values, which are then written into the target area raster to generate a regional-scale forest aboveground biomass density raster product.
10. A forest aboveground biomass density estimation system driven by multiple ecological neighborhoods, characterized in that: It includes a processor and a memory, the memory being used to store program instructions, and the processor being used to call the program instructions in the memory to execute the forest aboveground biomass density estimation method driven by the multi-source ecological neighborhood as described in any one of claims 1-9.