Forest aboveground biomass downscaling method based on multi-scale geographically weighted regression

By employing multi-scale geographically weighted regression and Kriging interpolation methods, the problem of unconsidered spatial differentiation effects in the AGB prediction model was addressed, resulting in high-precision, high-resolution AGB distribution maps that support the sustainable development and protection of forest resources.

CN116091939BActive Publication Date: 2025-12-12NANJING FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310002639.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-03
Publication Date
2025-12-12
Estimated Expiration
2043-01-03

AI Technical Summary

Technical Problem

Existing AGB prediction models fail to fully consider the differences in the spatial differentiation effects of different predictor variables, resulting in a lack of authenticity in the results. Traditional field surveys are time-consuming and laborious, and it is difficult to provide spatiotemporally clear AGB distribution information in a wide range of areas.

Method used

We employ a multi-scale geographic weighted regression method, combined with multispectral image fusion and Kriging interpolation, to capture the spatial heterogeneity of predictor variables through the MGWR model, construct a high-resolution AGB distribution map, and use multivariate stepwise regression, random forest importance ranking, and Pearson correlation coefficient to screen variables, while Kriging interpolation is combined to improve prediction accuracy.

Benefits of technology

It enables reliable downscaling from low-resolution remote sensing data to high-resolution AGB distribution maps, improves prediction accuracy, provides more targeted forest management solutions, and supports sustainable development and the protection of rare and endangered flora and fauna.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091939B_ABST
    Figure CN116091939B_ABST
Patent Text Reader

Abstract

The application discloses a forest above-ground biomass downscaling method based on multi-scale geographic weighted regression, comprising collecting research area data, extracting prediction variables; fusing a full-color band with a corresponding multi-spectral band to generate a fused multi-spectral image; determining characteristic variables having a closer relationship with AGB values, and meanwhile eliminating redundant variables from a modeling process; capturing differences in spatial heterogeneity levels of various prediction variables through MGWR; directly applying an MGWR statistical regression model constructed by using a coarse resolution data set to a prediction variable set to complete a downscaling task; and separating a structural component of AGB residual obtained by MGWRD through Kriging interpolation, and superimposing the separated component on a corresponding MGWRD prediction AGB value on space to form a final distribution mode of AGB. A low spatial resolution remote sensing image obtained at a lower cost is subjected to a statistical regression downscaling method to obtain a higher resolution AGB distribution map, and Kriging is used to improve prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of forest above-ground biomass estimation, and particularly to a forest above-ground biomass downscaling method based on multi-scale geographic weighted regression. BACKGROUND

[0002] Forest above-ground biomass (AGB) is a basic indicator for assessing forest ecosystem productivity, measuring carbon storage and carbon sink potential, and estimating carbon emissions caused by land use and climate change. Understanding the spatial distribution of AGB helps to understand the impact of global climate change on forests and provides basic data support for developing sustainable forest management strategies. Although traditional field surveys of forest AGB are accurate, this method of collecting AGB information is time-consuming, labor-intensive, costly, and destructive, and it is difficult to obtain in remote and steep slopes, so it is difficult to provide spatial and temporal explicit AGB distribution information in a wide range. Remote sensing images contain rich spectral and texture information, have appropriate temporal and spatial resolution, wide coverage, and good timeliness, and can overcome the limitations of traditional field surveys to a certain extent when predicting AGB. However, due to data limitations, cost limitations, time and spatial scale, the sources of remote sensing data used for AGB quantitative estimation are different. Medium and high resolution remote sensing images are often used for local and regional AGB estimation, and for national or even global AGB estimation, low spatial resolution remote sensing data is often used, but the estimation accuracy is relatively low. Due to the obvious spatial heterogeneity of AGB and the differences in remote sensing data types and estimation methods, different applications may have different random / systematic errors in the AGB estimation process. Among them, the spatial difference (or scale mismatch) between the size of the field sampling plot and the size of the remote sensing image pixel may lead to AGB estimation error. Therefore, how to derive a more reliable and higher spatial resolution AGB distribution map from easily available low-cost, low-resolution remote sensing images is a promising method worth considering.

[0003] Downscaling converts low-resolution information into high-resolution information, which is currently widely used in the study of climate change, precipitation and other aspects. However, there are few studies on downscaling low-resolution AGB using high-resolution multispectral data. Existing studies mainly give the representative biomass value of each vegetation type based on the forest vegetation distribution map, and downscale it to the average biomass of each vegetation type to obtain high-resolution biomass data. However, this downscaling method is not a statistical spatial downscaling method. In the field of geography and spatial analysis, the instability of the relationship between spatially distributed variables is called non-stationarity. Machine learning algorithms such as artificial neural networks (ANN), support vector machines (SVM), and random forests (RF) can simulate the non-linear relationship between AGB and related prediction variables, thereby achieving reliable downscaling from low resolution to high resolution. However, these methods assume that the prediction variables have the same response to AGB at different spatial locations. Due to the autocorrelation and non-stationarity of forest ecosystems, the prediction variables will change with the spatial location. Most existing AGB prediction models do not fully consider the differences in the spatial differentiation effects of different prediction variables on AGB, resulting in a lack of authenticity of the results. SUMMARY

[0004] The present application provides a forest aboveground biomass downscaling method based on multiscale geographically weighted regression, which solves the problem that most existing AGB prediction models do not fully consider the differences in the spatial differentiation effects of different prediction variables on AGB, resulting in a lack of authenticity of the results.

[0005] To achieve the above purpose, the present application provides the following technical scheme: a forest aboveground biomass downscaling method based on multiscale geographically weighted regression, comprising the following steps:

[0006] S1, collecting research area data and extracting prediction variables;

[0007] S2, fusing the panchromatic band with the corresponding multispectral band to generate a fused multispectral image;

[0008] S3, determining the characteristic variables that have a closer relationship with AGB values, and simultaneously eliminating the redundant variables from the modeling process;

[0009] S4, capturing the differences in the spatial heterogeneity levels of various prediction variables through the MGWR model;

[0010] S5, directly applying the MGWR statistical regression model constructed using the coarse resolution dataset to the prediction variable set to complete the downscaling task;

[0011] S6, and the structural component of the AGB residual derived by the MGWRD is separated by Kriging interpolation, and the separated component is superimposed on the corresponding MGWRD predicted AGB value in space to form the final distribution pattern of AGB.

[0012] Preferably, in S2, the image fusion method includes Brovey transformation, GS transformation, NNDiffuse method and PCA fusion, and the best fusion method is determined by calculating the standard deviation, correlation coefficient, average gradient and information entropy and other indicators to quantitatively evaluate the fusion effect.

[0013] Preferably, in S3, the method for determining the variable set includes multivariate stepwise regression screening method, random forest importance sorting algorithm, and method combining Pearson correlation coefficient and variance inflation factor.

[0014] Preferably,

[0015] 1) The multivariate stepwise regression screening method, a one-dimensional regression model is established by using the dependent variable by multivariate stepwise screening method, and the F value corresponding to each variable is calculated;

[0016] 2) The random forest importance sorting algorithm, the feature importance is compared by the random forest importance sorting algorithm to select the prediction factor with high importance;

[0017] 3) The method combining Pearson correlation coefficient and variance inflation factor, the feature variables highly correlated with AGB are screened out by using Pearson correlation coefficient, and then whether the screened feature variables violate multicollinearity is tested, and the variance inflation factor is used to judge:

[0018]

[0019] Where, R i is the composite correlation coefficient of the ith variable to the remaining k-1 prediction variables, if VIF is between 0 and 10, there is no multicollinearity; if VIF≥10, there is high multicollinearity between variables, and part of the variables should be removed from the model.

[0020] Preferably, the random forest algorithm includes a first index and a second index for measuring the importance of the variable, the first index is the increase of the predicted mean square error percentage of each tree, and the second index is the total reduction of the Inc Node Purity of the average variable split on all trees.

[0021] Preferably, in S4, the MGWR model is represented as follows:

[0022]

[0023] where y i is the i-th observation of the response variable, x ij is the j-th explanatory variable at position i, β bwj represents the regression coefficients of different variables j at different bandwidths, (u i , v i ) represents the spatial geographic coordinates of the sample points, k represents the number of predictors, and ε is the model regression residual.

[0024] Preferably, the AGB downscaling process is:

[0025] AGB high (OV) = MGWRD(OV)

[0026] where AGB high represents the fine-scale AGB prediction value; OV represents the optimal predictor set determined by MGWR analysis, MGWRD represents the AGB statistical regression model created by the optimal predictor set, and the resolution is coarser.

[0027] Preferably, the Kriging interpolation method includes ordinary Kriging and co-Kriging with aspect as a covariate.

[0028] Preferably, in S6, the final distribution pattern of AGB is:

[0029] R(x i ) = AGB(x i ) - AGB MGWRD (x i )

[0030] AGB MGWRD-OK / MGWRD-CK (x i ) = AGB MGWRD (x i ) + R OK / CK (x i )

[0031] where R(x i ) is the AGB residual value at sample position i, AGB(x i ) is the AGB observation value at sample position i, AGB MGWRD (x i ) is the MGWR prediction value at sample position i, AGB MGWRD-OK / MGWRD-CK (x i ) is the AGB prediction value obtained by MGWRD-OK or MGWRD-CK model, and R OK / CK (x i ) is the residual structure component interpolated by OK or CK at sample position i.

[0032] Compared with the prior art, the present application has the beneficial effects of:

[0033] 1、In the present application, the method used utilizes an integrated model of MGWR and Kriging interpolation, by considering the spatial non-stationarity and spatial autocorrelation existing in the complex structure of forest ecosystems, a low spatial resolution remote sensing image obtained at a lower cost is used to obtain a higher resolution AGB distribution map through a statistical regression downscaling method, compared with many existing machine learning models such as support vector machine, random forest, BP neural network, etc., when predicting AGB, they often do not consider the influence of spatial position on AGB distribution, and a large error is generated in the prediction process, the method used in the present application can improve this defect and improve the prediction accuracy.

[0034] 2、The present application proposes a statistical downscaling model framework considering spatial non-stationarity to solve the problem of limited basic data and the difficulty of effectively estimating by traditional methods, and by applying the relationship between 30m AGB and the prediction variables to the 15m prediction variable set, a good accuracy is obtained. Compared with existing AGB downscaling research, the present application can obtain a fine AGB distribution map per pixel, and accurately drawing the AGB distribution map of the region will help the forestry and environmental protection departments to develop more targeted forest management plans, better protect the rare and endangered animals and plants protected by the state in their territory, realize the sustainable development of forest resources, and actively respond to global climate change.

[0035] 3、In the present application, a two-stage variable selection strategy is carried out, so that the obtained AGB distribution pattern is more reliable and realistic, by introducing the concept of action scale on the basis of three variable screening methods of multivariate stepwise regression, random forest importance sorting and Pearson correlation coefficient, the spatial non-stationarity is screened for characteristic variables, and finally the variable set with the highest spatial stationarity is determined as the input prediction variable set for subsequent downscaling operation.

[0036] 4、By considering the spatial autocorrelation of AGB, the different models of the ordinary Kriging method and the collaborative Kriging method are compared, including the exponential model, the spherical model and the Gaussian model, the fitting effect of the AGB residual error is compared, by calculating the spatial stationarity of each prediction variable, the slope direction is selected as the covariate, and is superimposed on the previous downscaling result, further improving the prediction accuracy of the present application, which provides a reference for predicting the AGB distribution of complex mountainous areas in the future. BRIEF DESCRIPTION OF DRAWINGS

[0037] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, and are used to explain the present application together with embodiments of the present application, and do not constitute a limitation on the present application.

[0038] In the drawings:

[0039] Figure 1 is a schematic diagram of the change of downscaled pixel under the influence of the bandwidth of the prediction variable of the present application;

[0040] Figure 2 is a significance ranking diagram of the first 15 variables of the present application;

[0041] Figure 3 is a correlation heat map between the selected feature variables and AGB of the present application;

[0042] Figure 4 is a diagram of the present application by different variable screening methods: (a) random forest importance ranking; (b) multivariate stepwise regression; (c) different prediction variables determined by Pears-Vif.GWR and MGWR;

[0043] Figure 5 is a diagram of the training and validation performance of the MGWRD model of the present application: (a) training sample (n = 118) and (b) validation sample (n = 50);

[0044] Figure 6 is a diagram of the AGB downscaled result based on the MGWR model of the present application;

[0045] Figure 7 is a histogram of the distribution of the AGB residual predicted by the MGWRD of the present application;

[0046] Figure 8 is a diagram of the improved AGB prediction result based on the MGWRD-CK model of the present application;

[0047] Figure 9 is a flowchart of the method of the present application. DETAILED DESCRIPTION

[0048] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0049] Embodiment: As shown in the above, the forest aboveground biomass downscaled method based on multi-scale geographic weighted regression comprises the following steps: Figure 9

[0050] S1, collect the data of the study area, and extract the prediction variables;

[0051] S2, fuse the panchromatic band with the corresponding multispectral band to generate a fused multispectral image; the image fusion method includes Brovey transformation, GS transformation, NNDiffuse method and PCA fusion, and the fusion effect is quantitatively evaluated by calculating the standard deviation, correlation coefficient, average gradient and information entropy, etc. to determine the best fusion method; ​

[0052] S3, determine the characteristic variable with a closer relationship with the AGB value, and eliminate the redundant variables from the modeling process; the method for determining the variable set includes a multivariate stepwise regression screening method, a random forest importance sorting algorithm, and a method combining a Pearson correlation coefficient and a variance inflation factor;

[0053] 1) The multivariate stepwise regression screening method, a one-dimensional regression model is established using the dependent variable by the multivariate stepwise screening method, and the F value corresponding to each variable is calculated;

[0054] 2) The random forest importance sorting algorithm, the feature importance is compared by the random forest importance sorting algorithm to select the prediction factor with high importance; the random forest algorithm includes a first index and a second index for measuring the importance of the variable, the first index is the increase of the mean square error percentage of prediction of each tree, and the second index is the total reduction of the Inc Node Purity of the average variable split on all trees;

[0055] 3) The method combining the Pearson correlation coefficient and the variance inflation factor, the Pearson correlation coefficient is used to screen the characteristic variables highly correlated with the AGB, then whether the screened characteristic variables violate the multicollinearity is tested, and the variance inflation factor is used to determine:

[0056]

[0057] wherein, R i is the composite correlation coefficient of the i-th variable to the remaining k-1 prediction variables, if the VIF is between 0 and 10, there is no multicollinearity; if the VIF is greater than or equal to 10, there is a high multicollinearity between the variables, and part of the variables should be eliminated from the model;

[0058] S4, capture the difference in the spatial heterogeneity level of various prediction variables by the MGWR model; the MGWR model is represented as follows:

[0059]

[0060] wherein y i is the i-th observation value of the response variable, x ij is the observation value of the j-th explanatory variable at position i, β bwj represents the regression coefficient of different variables j under different bandwidths, (u i , v i ) represents the spatial geographic coordinates of the sample point, k represents the number of prediction variables, and ε is the model regression residual;

[0061] S5, the MGWR statistical regression model constructed by using the coarse resolution dataset is directly applied to the prediction variable set to complete the task of downscaling; the AGB downscaling process is:

[0062] AGB high (OV) = MGWRD(OV)

[0063] AGB high , wherein AGB i represents the fine-scale AGB prediction value; OV represents the optimal prediction variable determined by the MGWR analysis, MGWRD represents the AGB statistical regression model created by the optimal prediction variable set, and the resolution is coarse;

[0064] S6, and the structure component of the AGB residual obtained by the MGWRD is separated out by the Kriging interpolation method including the ordinary Kriging method and the collaborative Kriging method with aspect as the covariant, and the separated component is superimposed on the corresponding MGWRD predicted AGB value in space to form the final distribution pattern of AGB, and the formula of the final distribution pattern of AGB is:

[0065] R(x i ) = AGB(x i )- AGB MGWRD (x i )

[0066] AGB MGWRD-OK / MGWRD-CK (x i ) = AGB MGWRD (x i ) + R OK / CK (x i )

[0067] AGB i (x i ) is the AGB observation value of sample position i, AGB MGWRD (x i ) is the MGWR prediction value of sample position i, AGB MGWRD-OK / MGWRD-CK (x i ) is the AGB prediction value obtained by the MGWRD-OK or MGWRD-CK model, and R OK / CK (x i ) is the residual structure component at sample position i interpolated by OK or CK.

[0068] 1. The specific method steps for realizing functions or results:

[0069] 1.1 Data source and pretreatment

[0070] Two scenes of Landsat 8 OLI images on July 22, 2014 were used in this example, with band numbers 119 / 040 and 119 / 041, and cloud cover of 6.89% and 4.02%, respectively. Landsat 8 OLI surface reflectance data was directly generated by the U.S. Geological Survey EROS Data Center from the Landsat Surface Reflectance Code (LaSRC), which uses MODIS ancillary climate data and a unique radiative transfer model, with aerosol retrieval testing along the coastal aerosol belt. In addition, terrain correction was performed on the Landsat surface reflectance image by using the C correction model to minimize the impact of terrain shading on surface reflectance. The present invention also used a free digital elevation model (DEM) with a resolution of 30 meters from (http: / / earthexplorer.usgs.gov), and slope and aspect information was derived in the ArcGIS 10.3 software package. Based on the DEM data, orthorectification was performed on the Landsat 8 OLI image by using the RPC module in the ENVI 5.3 software package to further refine the positioning accuracy of the pixels and minimize the impact of terrain undulations within the study area.

[0071] The ground forest inventory data used in this example was compiled from the results of the 2014 National Forest Inventory (NFI) in Zhejiang Province. The NFI in Zhejiang Province began in 1979, and a series of square sample plots with an area of 0.08 hectares (28.28 m x 28.28 m) were established using systematic sampling methods. The AGB of each sample plot was obtained by summing the AGB of each measured tree within the sample plot, and the AGB of each measured tree was calculated using the allometric equation corresponding to the tree species. For some tree species that did not have an allometric equation for biomass, an equation for the same genus was used as a substitute for approximate calculation. In order to obtain more accurate inversion results, sample plots covered by cloud pixels in the Landsat 8 OLI image were excluded from the analysis. In addition, although the pixel size of the OLI image (30 m x 30 m) is very close to the size of each sample plot (28.28 m x 28.28 m), there is still a small difference in spatial scale, and it is necessary to convert the AGB per unit area to the pixel-level AGB with units of t / 900m2for ease of modeling and mapping processes. After completion, the statistical characteristics of the AGB values of the filtered samples were analyzed using SPSS software (version 21.0), and those AGB values greater than or less than the average value of the sample plot AGB plus / minus 3 times the standard deviation were considered to be outliers and were further excluded from the analysis. Finally, a total of 168 sample plots were retained, of which 70% were randomly determined for model training and the remaining 30% were used for model validation. The following table summarizes the statistical characteristics of the AGB values of the final sample plots after screening: 2 ​

[0072]

[0073] 1.2 Independent variable extraction

[0074] Four categories of modeling variables were extracted in this study, including original bands and their transformations, vegetation indices, texture features, and terrain features. Spectral feature transformations included the results of principal component analysis (PCA), tasseled cap transformation (TC), and minimum noise fraction (MNF), which were implemented to remove the redundancy among original bands and to generate synthetic features that might be highly correlated with AGB. The details of the techniques used to generate features are shown in the following table of AGB modeling variables, which included a total of 132 feature variables in the modeling analysis.

[0075]

[0076]

[0077]

[0078] 1.3 Image fusion

[0079] To obtain a higher resolution set of predictor variables to support subsequent downscaling modeling, four image fusion methods including Brovey transform, GS transform, NNDiffuse method, and PCA fusion were applied to fuse the panchromatic band with the corresponding multispectral bands to generate a fused multispectral image with a spatial resolution of 15 meters. The quantitative evaluation of fusion results was conducted by calculating the standard deviation, correlation coefficient, mean gradient, and information entropy to determine the best fusion method. On this basis, those features with a spatial resolution of 15 meters that might be included in the downscaling model were accordingly generated from the fused image and the resampled DEM.

[0080] 1.4 Determination of variable set

[0081] Variable selection refers to the process of determining those feature variables that have a closer relationship with AGB values through certain criteria, while removing redundant variables from the modeling process. To compare the overall impact of different features or variable combinations on AGB, three variable selection methods were implemented and compared, including multiple stepwise regression, random forest importance ranking, and the combination of Pearson correlation coefficient and variance inflation factor.

[0082] 1.4.1 Multiple Stepwise Regression (MSR)

[0083] Multiple stepwise selection method combines the advantages of variable forward selection and reverse selection method, respectively, to establish a one-dimensional regression model with dependent variables, and to calculate the F value corresponding to each variable. In the establishment of stepwise regression model, the largest F value of the selected predictor variable is selected from the predictor variable not yet included in the model, and then the t test is performed on each selected predictor variable. If a selected predictor variable is no longer significant, it will be removed. Repeat the above steps until there is no significant explanatory variable selected into the regression equation, and there is no significant explanatory variable removed from the regression equation. The best subset retained is used as a characteristic variable to establish a model. In this analysis, stepwise regression was performed in SPSS software (Version 25, Armonk, NY, USA), and the significance level of F test was set to 0.05 and 0.10 for determining variable entry and deletion.

[0084] 1.4.2 Random Forest Importance Ranking (RFR)

[0085] Random Forest Importance Ranking algorithm was used to compare feature importance to select predictors with high importance. Random Forest algorithm has two indicators to measure the importance of variables. The first indicator (calculated from the ranking of out-of-bag data) is the increase in the mean squared error percentage (% IncMSE) of each tree prediction, and the second indicator is the total reduction of Inc Node Purity of average variable split on all trees. Higher % IncMSE and Inc Node Purity values indicate more important predictor variables.

[0086] 1.4.3 Pearson-Vif

[0087] Pearson correlation coefficient was used to screen the feature variables highly correlated with AGB, and then whether the screened feature variables violated the multicollinearity was tested, and the variance inflation factor (VIF) was used to determine.

[0088]

[0089] where R i is the composite correlation coefficient of the ith variable with the remaining k-1 predictor variables. If VIF is between 0 and 10, there is no multicollinearity. If VIF ≥ 10, there is a high multicollinearity between variables, and some variables should be removed from the model.

[0090] 1.5 Multiscale Geographic Weighted Regression (MGWR)

[0091] The advantage of MGWR over classic GWR is that it removes the assumption of a single bandwidth for all variables in GWR, resulting in a spatial process that is closer to the actual state. Considering spatial heterogeneity and scale of action, MGWR can be applied to study various influencing factors and their corresponding processes. In theory, MGWR allows different spatial scales for different predictors, indicating that each predictor has a different spatial smoothing range. In the MGWR model, this spatial smoothing range is reflected by the bandwidth of each predictor. A small bandwidth means that the relationship changes within a relatively local range, while a large bandwidth indicates that the relationship is stable over a larger range. The closer the bandwidth is to the global, the higher the spatial smoothing degree, indicating that the predictor has less variation in the results produced in the study area. In this work, MGWR is used to capture the differences in the level of spatial heterogeneity of various predictors. The MGWR model is represented as follows:

[0092]

[0093] Here y i is the i-th observation of the response variable, x ij is the j-th explanatory variable at location i, β bwj represents the regression coefficient of different variables j at different bandwidths, (u i , v i ) represents the spatial geographical coordinates of the sample points, k represents the number of predictors, and ε is the model regression residual.

[0094] 1.6 AGB downscaling

[0095] 1.6.1 Multiscale geographically weighted regression downscaling model (MGWRD model)

[0096] The MGWR model can determine the optimal combination of predictors by considering the spatial stationarity of the predictors, and analyze the spatial scale of action (bandwidth) of each predictor. A smaller change in the spatial position of a predictor, or a larger scale of action (bandwidth), indicates that the influence is similar in space, with less heterogeneity, which will result in a more smooth downscaling output, as shown in Figure 1 (b). Conversely, when the scale of action (bandwidth) of a predictor is small, more heterogeneous downscaling results will be produced, as shown in Figure 1 (a). Once the optimal combination of predictors is determined, the MGWR statistical regression model constructed using the coarse resolution dataset (30-meter resolution) will be directly applied to the 15-meter resolution predictor set to complete the downscaling task. The AGB downscaling process is as follows:

[0097] AGB high (OV) = MGWRD(OV)

[0098] Here AGB high denotes the fine-scale AGB prediction; OV denotes the optimal predictor variables determined by MGWR analysis, and MGWRD denotes the AGB statistical regression model created by the optimal predictor variable set with a coarser resolution;

[0099] 1.6.2 Multiscale geographically weighted regression downscaling-co-kriging model (MGWRD-CK model)

[0100] Since the MGWRD model does not take into account the spatial autocorrelation of AGB samples, we used a combination of MGWRD and kriging interpolation to more accurately determine the spatial distribution of AGB. The kriging interpolation method used in this paper includes ordinary kriging (OK) and co-kriging with covariates (CK). The specific steps are as follows: (1) separate the structural component of the 15-meter resolution AGB residual obtained by MGWRD through ordinary kriging and co-kriging interpolation; (2) superimpose the separated component on the corresponding MGWRD predicted AGB value in space to form the final distribution pattern of 15-meter AGB, and the specific principles of the two steps are explained as follows:

[0101] R(x i ) = AGB(x i ) - AGB MGWRD (x i )

[0102] AGB MGWRD-OK / MGWRD-CK (x i ) = AGB MGWRD (x i ) + R OK / CK (x i )

[0103] where R(x i ) is the AGB residual value of sample location i, AGB(x i ) is the observed value of AGB at sample location i, and AGB MGWRD (x i ) is the MGWR prediction value at sample location i. AGB MGWRD-OK / MGWRD-CK (x i ) is the AGB prediction value obtained by the MGWRD-OK or MGWRD-CK model, and R OK / CK (x i) is the residual structural component interpolated by OK or CK at sample location i. Here, we focused on comparing the residual prediction accuracy of ordinary kriging (OK) interpolation and co-kriging (CK) interpolation with aspect as the covariate, and we selected the exponential, spherical and Gaussian models of semivariogram to fit the residuals. To test the performance of the MGWR algorithm (Equation 2), we compared its results with those of the random forest model (RF), geographically weighted regression (GWR) and traditional ordinary least squares regression (OLS) methods.

[0104] 1.7 Cross-validation

[0105] Here, the ten-fold cross-validation method was adopted to evaluate the average prediction performance of the four models based on different predictor variables. On this basis, the statistical data including mean absolute error (MAE), root mean square error (RMSE), average value of determination coefficient (R2) and bias were derived from ten-fold cross-validation to show the performance of the models. Their formulas are as follows:

[0106]

[0107]

[0108]

[0109]

[0110] where n is the total number of validation observations, is the AGB predicted by the model, y i is the AGB observed on the ground, is the arithmetic mean of all observed AGB values.

[0111] 2. Experimental data for obtaining the performance of the product or result:

[0112] 2.1 Determination of the set of predictor variables

[0113] 2.1.1 Multiple stepwise regression (MSR)

[0114] The following table of multiple stepwise regression results summarizes the predictor variables selected by the MSR method and the corresponding fitted regression equation. MSR finally retained 12 predictor variables, with the number of textures accounting for the largest proportion, and some spectral features, topographic factors (including slope and aspect) were also selected;

[0115]

[0116] Note: Y represents AGB, X1, X2, X3, …, X 11slope, cor_B1_3, B15, EVI, sec_B4_3, aspect, var_B4_3, dis_B4_3, homo_B6_5, SAVI, B4, and mean_B5_3

[0117] 2.1.2 Random Forest Importance Ranking (RFR)

[0118] Figure 2 The importance scores of the top 15 variables after running the RF model for 100 times are shown. By comparing the IncMSE and IncNodePurity metrics, con_B4_3, B2, and con_B2_5 are removed from the top 15, and the final input variables are selected as elevation, aspect, sec_B3_5, B357, B14, B7, var_B5_3, mean_B7_5, brightness, B16, B5, and cor_B5_5 for subsequent analysis.

[0119] 2.1.3 Pearson-Vif

[0120] Figure 3 The correlation heat map between the top 12 predictor variables and AGB derived by the Pearson-Vif method is shown, including two topographic factors, five texture features, and five spectral features. Among them, the two topographic factors, i.e., aspect and elevation, have the strongest correlation with AGB, while the remaining predictor variables have no significant difference in correlation with AGB.

[0121] 2.2 Scale of Action

[0122] Before building the MGWR model, the spatial autocorrelation test was conducted to determine whether there was a significant relationship between the samples. As a result, the Moran's I index was calculated to be 0.23, the Z-score was 3.6, and the P-value was 0.0003. The results of the Z-score and the P-value together indicate that the AGB in the study area presents a statistically significant clustering pattern rather than a random distribution pattern. Table 3 shows the prediction performance indicators of OLS, RF, GWR, and MGWR models under different variable sets determined by the three variable screening methods, which shows the different prediction performances of the four AGB modeling algorithms under the three groups of prediction variables. When implementing OLS, GWR, MGWR, and RF regression analysis, the RFR-selected variable set obtained the lowest RSS, AIC, AICc, and CV values, and the fitting degree (R square value) of the RFR-selected variable set was also the highest accordingly. Therefore, we believe that the RFR-selected variable set is the best variable combination, which has the highest explanatory power for the AGB change. On this basis, by comparing the regression results of OLS, GWR, MGWR, and RF, MGWR is superior to OLS, GWR, and RF in terms of RSS, AIC, AICc, CV, and R square statistics. Therefore, it is further concluded that the MGWR regression is the best regression in the current analysis;

[0123]

[0124] Note: RSS is the residual sum of squares; AIC is the Akaike information criterion; AICc is the corrected Akaike information criterion; CV is the coefficient of variation; R square is the determination coefficient ("*" indicates statistical significance at the 5% level).

[0125] As Figure 4 The figure shows the action scale or bandwidth determined by GWR and MGWR for different prediction variables. In different variable groups, the identified bandwidth of the classic GWR is between 121 and 130 kilometers. This single bandwidth assumes that all variables affect AGB within the same regional range, which has great limitations. In contrast, MGWR can capture the differences in the spatial heterogeneity level of the prediction variables. For example, in Figure 5 In (a), the bandwidths of elevation and aspect are small, indicating that these two variables affect AGB on a relatively local scale, respectively. The relationship between the six variables of sec_B3_5, NDVI, var_B5_3, B7, mean_B4_5, B16, and AGB shows spatial smoothness, but the process is different within a wide regional range. Other variables affect AGB globally because their optimal bandwidth is close to the maximum possible number of neighbors 168 (total sample size). The RFR-selected prediction variable set shows a larger action scale, with an average of 130 kilometers and a median of 139 kilometers, which is closer to the global range or total sample size (168 kilometers) Figure 5(a)). These results further confirm that the set of predictors selected by RFR is optimal and produces downscaling results with less spatial heterogeneity.

[0126] 2.3 Multiscale Geographically Weighted Regression Downscaling Model (MGWRD Model)

[0127] The current method consistently employs a 10x cross-validation strategy. Figure 5 (a) and (b) show the fitting accuracy and validation accuracy of the MGWRD model built based on variables selected using RFR. Figure 5 As shown, the fitting accuracy and validation accuracy (R²) of the MGWRD model are... 2 The values ​​were 0.57 and 0.58, respectively; the deviation values ​​were -2.78t / ha and -2.74t / ha, respectively. This indicates that the MGWRD model constructed in this invention did not exhibit overfitting and achieved high AGB estimation accuracy.

[0128] like Figure 6 The spatial distribution pattern of AGB predicted by MGWRD is shown. The predicted AGB values ​​range from 12.53 to 306.81 tons / hectare. The predicted low AGB values ​​are mainly distributed in the northern, southwestern, southeastern, and central parts of the study area, while the predicted high values ​​are mainly concentrated in the southern, central-eastern, and central-western parts of the study area.

[0129] 2.3 Multiscale Geographically Weighted Regression Downscaling-Co-Kriging Model (MGWRD-CK Model)

[0130] like Figure 7 A bar chart showing the AGB residuals predicted by MGWRD is displayed. The standard deviation of the residuals is 63.36 tons / hectare, the absolute kurtosis is -0.36, and the absolute skewness is close to 1, indicating that the residuals are approximately normally distributed. Therefore, the preconditions for kriging interpolation are met, and subsequent kriging interpolation analysis can be performed on the AGB prediction residuals. Figure 4 As can be seen from the scale of action of the three sets of characteristic variables, the slope aspect has strong spatial heterogeneity. Therefore, it is chosen as a covariate for kriging interpolation.

[0131] The table below shows the semivariance function statistics for different models under the OK and CK methods, displaying the residual fitting results and related statistics for the semivariance function models based on the exponential, spherical, and Gaussian models. According to the GS+ modeling results, overall, the exponential model outperforms the other two models under both the OK and CK methods. For the CK method, the exponential model achieves higher accuracy, with an R-squared of 0.81 in the table, higher than 0.78 (OK). Therefore, the CK exponential function model was chosen to further improve the accuracy of downscaling AGB.

[0132]

[0133] As Figure 8 The improved AGB prediction results based on the MGWRD-CK model are shown. Obviously, there is a clear spatial similarity between the prediction of MGWRD-CK ( Figure 8 ) and the prediction of MGWRD ( Figure 6 ), and in the prediction of MGWRD-CK, the minimum AGB drops to 5.82 tons / ha, and the maximum AGB increases to 328.18 tons / ha. Compared with the prediction of the MGWR model, the range of AGB prediction of the MGWRD-CK model is expanded, indicating that the model is more adaptable to AGB prediction. The model is verified by using 30% of the independent sample, and the relevant statistical data are shown in the following table. The verification accuracy statistics of the MGWRD-OK and MGWRD-CK models based on the verification data show that the accuracy of the MGWRD-OK and MGWRD-CK models in MAE, RMSE, R square and Bias and other specific statistical data is improved compared with the MGWRD model;

[0134]

[0135] The data provided by the common sensor currently often has the contradiction of high spatial resolution and low temporal resolution, or high temporal resolution and low spatial resolution, and the trade-off between the space-time resolution limits the acquisition of remote sensing data with high frequency and high spatial resolution. For example, when estimating the aboveground biomass of forest in a small area scale, scholars mostly use high-resolution satellite data, which often has the characteristics of high cost and large workload. For the estimation of biomass in a national scale or even a global scale, high temporal resolution and low spatial resolution remote sensing data are mostly used, and the accuracy of the estimation result is relatively low, which cannot present the spatial distribution pattern of the aboveground biomass of forest in detail and cannot meet the needs of forest investigation and dynamic change monitoring. Based on this, the statistical regression downscaling idea is applied to the forest aboveground biomass data to generate a high spatial resolution AGB spatial pattern map at a lower cost. In recent years, the research results of downscaling methods in the fields of land surface temperature and precipitation are rich, but the research on the downscaling of forest aboveground biomass is rare, and the existing research is mostly based on the forest vegetation distribution map, and the representative value of the biomass of each vegetation type is given, which is taken as the average biomass of each vegetation type at a large scale, and is down-scaled to the grid, so that high-resolution biomass data is obtained, which is not a traditional statistical downscaling, but a downscaling based on the assignment strategy, and cannot effectively express the local detail characteristics of AGB. The statistical downscaling model framework considering spatial non-stationarity is proposed to solve the problem of limited basic data and difficulty in effective estimation by traditional methods. Specifically, the relationship between 30m AGB and the prediction variable is applied to the 15m prediction variable set, and good accuracy is obtained. Compared with the existing AGB downscaling research, the fine AGB distribution map of each pixel can be obtained, and accurate drawing of the AGB distribution map of the region will help the forestry and environmental protection departments to formulate more targeted forest management schemes, better protect the rare and endangered animals and plants protected by the state in their territory, realize the sustainable development of forest resources, and actively respond to global climate change.

[0136] Due to the complex structure of forest ecosystem and the difference of remote sensing data types and estimation methods, different applications may have different random / systematic errors in the AGB estimation process. Among them, the spatial difference (or scale mismatch) between the sample size of the original sample and the pixel size of the remote sensing image may cause AGB estimation error. For example, when using standard size to estimate the AGB of a large area, local spatial variation will cause large sampling error. In the present application, the resolution of remote sensing image data is 30m x 30m, and the ground investigation data of 28.28m x 28.28m is converted into 30m x 30m biomass measured data per unit area. Through the conversion, the error caused by the scale mismatch is reduced as much as possible.

[0137] The two-stage variable selection strategy is carried out in the application, so that the obtained AGB distribution pattern is more reliable and realistic. The spatial non-stationarity generated by the complex structure of the forest ecosystem cannot be ignored. In order to obtain smoother downscaling results, on the basis of three variable screening methods of multivariate stepwise regression, random forest importance ranking and Pearson correlation coefficient, the concept of action scale is further introduced, the spatial non-stationarity is screened from the perspective of the characteristic variable, and finally the variable set with the highest spatial stationarity is determined as the input prediction variable set for subsequent downscaling operation. The existing AGB distribution prediction model mostly selects one variable screening method, which may have the problem of collinearity and redundancy between different prediction variables. At the same time, the spatial properties of the characteristic variables are not considered in the existing AGB prediction model, while the method compares three characteristic variable screening methods, and considers the spatial position property, which can effectively improve the accuracy of AGB prediction. In addition, by considering the action scale of the characteristic variable, the method can analyze which is the local variable and which is the global variable by statistically analyzing the response degree of each characteristic variable to AGB in space, which can provide a reference for the selection of AGB prediction variables in the future.

[0138] The application compares the fitting effect of different models (exponential model, spherical model and Gaussian model) of ordinary kriging method and collaborative kriging method on AGB residual error by considering the spatial autocorrelation of AGB, selects the slope direction as the covariate by calculating the spatial stationarity of each prediction variable, and superimposes it on the previous downscaling result, which further improves the prediction accuracy of the application. This provides a reference for predicting the AGB distribution of complex mountainous areas in the future.

[0139] Finally, it should be noted that: the above only for the preferred examples of the application, and not for limiting the application, although the application has been described in detail with reference to the foregoing examples, for those skilled in the art, it still can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the application shall be included in the protection scope of the application.

Claims

1. A method for downscaling forest aboveground biomass based on multi-scale geographically weighted regression, comprising the following steps: S1, collecting data of the study area, and extracting the prediction variables; S2, fusing the panchromatic band with the corresponding multispectral bands to generate a fused multispectral image; S3, determining the characteristic variables that have a closer relationship with the AGB value, and removing the redundant variables from the modeling process; S4, capturing the differences in the spatial heterogeneity level of various prediction variables through the MGWR model; S5, directly applying the MGWR statistical regression model constructed using the coarse resolution dataset to the prediction variable set to complete the downscaling task; S6, separating the structural components of the AGB residual derived from the MGWRD by Kriging interpolation, and superimposing the separated components on the corresponding MGWRD predicted AGB value to form the final distribution pattern of AGB; In S4, the MGWR model is represented as follows: ; wherein is the ith observation of the response variable, is the observation of the jth explanatory variable at position i, represents the regression coefficients of different variables j at different bandwidths, represents the spatial geographic coordinates of the sample points is, k represents the number of prediction variables, is the model regression residual; In S5, the AGB downscaling process is as follows: ; wherein, represents a fine-scale AGB predictor; represents the best predictor variables determined by the MGWR analysis, and MGWRD represents the AGB statistical regression model created from the best predictor variable set, at a coarser resolution.

2. The multi-scale geographically weighted regression-based forest aboveground biomass downscaling method according to claim 1, characterized in that: In S2, the image fusion method includes Brovey transformation, GS transformation, NNDiffuse method and PCA fusion, and the fusion effect is quantitatively evaluated by calculating the standard deviation, correlation coefficient, average gradient and information entropy index to determine the best fusion method.

3. The multi-scale geographically weighted regression-based forest aboveground biomass downscaling method according to claim 1, characterized in that: In S3, the method for determining the variable set includes multivariate stepwise regression screening method, random forest importance sorting algorithm and method combining Pearson correlation coefficient and variance inflation factor. 4.The method for downscaling forest aboveground biomass based on multi-scale geographically weighted regression according to claim 3, characterized in that: 1) the multivariate stepwise regression screening method, a one-dimensional regression model is established using the dependent variable through multivariate stepwise screening method, and the F value corresponding to each variable is calculated; 2) the random forest importance sorting algorithm, the feature importance is compared through the random forest importance sorting algorithm to select the prediction factor with high importance; 3) the method combining Pearson correlation coefficient and variance inflation factor, the characteristic variables highly correlated with AGB are screened out by using Pearson correlation coefficient, then whether the screened characteristic variables violate the multicollinearity is tested, and the variance inflation factor is used to judge: , =1, 2, …, k. wherein, is the composite correlation coefficient of the i-th variable with the remaining k-1 predicted variables. If the VIF is between 0 and 10, there is no multicollinearity; if VIF > 10, there is a high multicollinearity between variables, and some variables should be removed from the model.

5. The multi-scale geographically weighted regression-based forest aboveground biomass downscaling method according to claim 4, characterized in that: The random forest importance sorting algorithm includes a first index and a second index for measuring the importance of the variable, the first index is the increase of the prediction mean square error percentage of each tree, and the second index is the total reduction of the average variable split Inc Node Purity on all trees.

6. The multi-scale geographically weighted regression-based forest aboveground biomass downscaling method according to claim 1, wherein: The Kriging method includes ordinary Kriging method and collaborative Kriging method with aspect as a covariate.

7. The multi-scale geographically weighted regression-based forest aboveground biomass downscaling method according to claim 1, wherein: In S6, the formula of the final distribution pattern of AGB is as follows: ; ; wherein AGBresiduals for sample position i, is the AGB observation for sample position i, is the MGWR prediction for sample position i, is the AGB prediction obtained by the MGWRD-OK or MGWRD-CK model, is the residual structural component interpolated at sample position i by OK or CK.