Forest aboveground biomass remote sensing measurement method based on improved boruta algorithm

By improving the Boruta algorithm and combining it with SVR, the optimal feature variables were selected, and a remote sensing estimation model for forest ground biomass was constructed. This solved the problem of the neglect of the importance of feature variables in existing technologies and achieved high-precision remote sensing estimation of forest ground biomass.

CN115376011BActive Publication Date: 2025-12-09CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210594831.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-12-09
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

Existing technologies for estimating forest aboveground biomass overemphasize feature variables in terms of model fitting accuracy, neglecting their importance and leading to inaccurate estimation results.

Method used

A remote sensing estimation model for forest ground biomass was constructed by combining the improved Boruta algorithm with support vector regression (SVR) through feature variable selection. Vegetation index, texture features, and topographic features were extracted using Sentinel2 image data and SRTM-DEM data. The optimal feature variables were selected by combining the Boruta algorithm, and the SVR model was established for estimation.

Benefits of technology

The accuracy of remote sensing estimation of forest ground biomass was improved, and high-precision measurement of forest ground biomass was achieved. The correlation coefficient reached 0.79, the coefficient of determination R2 was 0.63, the RMSE was 28.28 t/ha, and the MAE was 23.04 t/ha.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115376011B_ABST
    Figure CN115376011B_ABST
Patent Text Reader

Abstract

The present application relates to forest aboveground biomass measurement, specifically relates to a kind of forest aboveground biomass remote sensing measurement method based on improved Boruta algorithm, improved Boruta algorithm's forest aboveground biomass remote sensing measurement method, it is characterized in that, the method includes the following steps: (1) selecting the remote sensing image data of the region to be measured and data pretreatment;(2) design improved Boruta algorithm;(3) SVR model is established using improved Boruta algorithm;(4) the forest aboveground biomass of the region to be measured is measured using model, the precision of the present application is higher, the AGB distribution condition measured and the actual surveying research ground topography and forest vegetation distribution condition basically coincide, provide a new method for forest measurement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the measurement of forest aboveground biomass, and particularly to a forest aboveground biomass remote sensing measurement method improved by Boruta algorithm. BACKGROUND

[0002] Forest fire must have three main conditions: forest fuel, fire weather, and fire source. Forest fuel is the material basis for forest fire. As the main component of forest fire fuel, the estimation of forest aboveground biomass lays a solid theoretical and practical foundation for forest fire prediction. Zhang et al. used Landsat5 TM image as research data, and used multiple linear regression and BP neural network two models to estimate the biomass of Tahe and Amur River area in the north slope of Yilu Lake in Northeast Heilongjiang Province. Colin et al. used random forest, linear mixed effect regression, stereoscopic and support vector regression four models to estimate the biomass. Zhang et al. integrated LiDAR data and Landsat8 image through a deep learning-based workflow, and developed a new method for collaborative estimation of biomass. Pablito et al. used Landsat8 OLI as data source, and used RF model and support vector model to estimate the forest biomass observed in Madrid, Spain, Mexico. The results show that the SVR model is the best for estimating forest biomass. Eelis et al. used Gaussian process (GPR) and support vector regression (SVR) regression to estimate forest biomass, and the results show that the performance of GPR is slightly worse than SVR. Comprehensive analysis of these methods, they mainly have their own shortcomings in modeling, overemphasizing the fitting accuracy of the model, and ignoring the number and importance of the modeling feature variables. SUMMARY

[0003] Based on this, the present application uses forest planning and design survey data as basic data, extracts a large number of variables using geographic remote sensing information technology, takes forest aboveground biomass as the estimation target, combines the improved Boruta algorithm feature variable screening with SVR, and completes the construction of a high-precision remote sensing estimation model of forest aboveground biomass, thereby providing a reference basis for the combination of machine learning and forest aboveground biomass variable selection.

[0004] The present application provides a forest aboveground biomass remote sensing measurement method improved by Boruta algorithm, which comprises the following steps:

[0005] (1) selecting remote sensing image data of the area to be measured and pre-processing the data;

[0006] (2) designing an improved Boruta algorithm;

[0007] (3) establishing an SVR model by using the improved Boruta algorithm;

[0008] (4) using the model to measure the forest aboveground biomass of the region to be measured.

[0009] Another preferred scheme of the present application is that step (1) is to select Sentinel2 image data. The Sentinel2 image data contains 13 bands of multispectral data, wherein the spatial resolution of Band2, Band3, Band4 and Band8 is 10m, the spatial resolution of Band5, Band6, Band7, Band8a, Band11 and Band12 is 20m, and the spatial resolution of Band1, Band9 and Band10 is 60m.

[0010] Another preferred scheme of the present application is that the radiation calibration and atmospheric correction of the LIC level standard product data of Sentinel2 in step (1) is processed by using the Sencor plug-in in the SNAP software provided by ESA. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a DEM map of the Lutou Forest Farm;

[0012] Figure 2 is a Boruta algorithm flowchart;

[0013] Figure 3 is a Boruta-SVR model accuracy verification result based on the improved Boruta algorithm;

[0014] Figure 4 is an improved Boruta algorithm flowchart;

[0015] Figure 5 is a Boruta-SVR model accuracy verification result based on the improved Boruta algorithm. DETAILED DESCRIPTION

[0016] The present application will be further described below through specific embodiments to help understand the technical solutions of the present application, but the protection scope of the present application is not limited to these embodiments.

[0017] I. Overview of the study area

[0018] The Lutou Forest Farm is located in Jiayi Town, Pingjiang County, Yueyang City, Hunan Province, located in the northern section of the Loxia Mountains, adjacent to the Lianyun Mountain National Forest Park, and overall as the west slope of the Lianyun Mountains, separated from the Mufu Mountains in Pingjiang by the Miluo River, and overlooking the back of the Dawei Mountain in Liuyang. The forest farm is located between 113°51′52″E-113°58′24″E and 28°31′27″N-28°38′00″N, with a total area of 4762hm 2 .

[0019] II. Data source and pre-processing

[0020] (1) Sentinel2 data

[0021] Sentinel 2 is a high-resolution multi-spectral imaging satellite, which carries the Multi-Spectral Instrument (MSI) device, mainly used for land monitoring, which is a way to obtain vegetation growth, soil coverage, coastal environment and inland waterways, etc. It not only has important and special significance for improving agriculture and forest planting, predicting food production and even ensuring food security, but also can monitor volcanic eruptions, mudslides, floods or other major natural disasters.

[0022] The spatial resolution between visible light to near-infrared to short-wave infrared is different, there are 60m, 10m, 20m, a total of three according to the processing level of Sentinel2 remote sensing image, the data product is divided into L0, L1A, L1B and L1C, a total of four processing levels, of which the original data is L0; The data product containing metadata after geometric rough correction is L1A; L1B data product does not have proper geometric correction, but embeds the geometric model after GCP optimization; L1C belongs to the apparent reflectance product, and has undergone sub-pixel level geometric correction and orthographic correction, and is the only product that can be downloaded in China at present. The European Space Agency (ESA) only publishes L1C level product data of Sentinel2, which has not yet realized radiation calibration and atmospheric correction processing. The Sentinel2 band information is as follows Table 1:

[0023] Table 1 Sentinel2 band information

[0024]

[0025]

[0026] (2) SRTM-DEM data

[0027] SRTM-DEM data comes from the joint measurement of the United States Federal Aviation Administration (NASA) and the Military Department National Mapping Bureau (NIMA). SRTM is the abbreviation of Shuttle Radar Topography Mission, which means Space Shuttle Radar Topography Mission. This space mapping has the characteristics of wide coverage, large amount of data information and high precision, which is unprecedented in the history of mapping. It takes about two years to process all the raw data collected within 10 days. After data processing, the global DEM data can improve the accuracy by about 30 times. At present, the measurement data has covered the whole territory of China. In this study, SRTM-DEM dataFigure 1 is DEM map of Lutou Forest Farm, and its spatial resolution is 30 m, which is used as the data of terrain feature variables.

[0028] (3) Plot data and processing

[0029] The test area is Lutou Forest Farm. The data comes from the forest resource planning and design survey data in 2020, with a total of 1236 subplots. In this experiment, 1 plot of 20 m x 20 m size is extracted from each sub-plot, i.e. a total of 1236 fixed plot data is collected. Plot data includes forest type, soil type, soil thickness, elevation, slope, aspect, canopy density, etc. The height and diameter of trees with a diameter of more than 5 cm are obtained by measuring each tree. According to the allometric growth model and regression model, the AGB of different tree species is calculated according to Table 2, and the AGB of different tree species is added up as the forest AGB of the plot, which is divided by the plot area, and then the unit area forest AGB is calculated.

[0030] Table 2 AGB calculation formula of different tree species

[0031]

[0032] Note: W T , W S , W P , W B , W L , D, H respectively represent aboveground biomass, stem biomass, bark biomass, branch biomass, leaf biomass, diameter at breast height, tree height

[0033] Data preprocessing

[0034] (1) Radiometric calibration

[0035] Remote sensing image data usually records the quantitative value (DN value) of the sensor to reflect the spectral characteristics of the ground object. The DN value increases with the increase of the radiation capacity. Due to the restriction of factors such as atmosphere and solar elevation angle, the quantitative value of the sensor cannot truly reflect the surface reflectance. Therefore, remote sensing image data needs to go through the process of radiometric calibration. The formula of radiometric calibration is as follows:

[0036] L λ = Gain x DN + Bias (1)

[0037] In the formula, L λ , Gain, DN, Bias respectively represent the radiance value, gain value, calibrated image value, and bias.

[0038] (2) Atmospheric correction

[0039] In the process of remote sensing data processing, atmospheric correction has become an important link. Because the atmosphere exists between the satellite sensor and the ground, the radiation energy received by the satellite sensor includes not only the ground radiation information, but also the atmospheric radiation information. Atmospheric molecules, aerosol scattering and absorption of ground radiation information reduce the quality of remote sensing images and affect quantitative research. After atmospheric correction, the errors caused by the atmosphere can be effectively reduced or eliminated, thereby greatly improving the authenticity of ground information expression.

[0040] Due to the format and other problems of image data (Sentinel2), other image processing software does not support the processing of Sentinel2 image data, and only SNAP software provided by ESA can be used to process the LIC level standard product data of Sentinel2 using sencor plug-in for radiation calibration and atmospheric correction. After atmospheric correction, Band2, Band3, Band4 and Band8 not only need to retain the original 10m resolution for calculating texture features, but also need to be resampled to 20m resolution, so as to keep consistent with other six 20m resolution bands in the next step of band combination and vegetation index calculation. With the support of Layer Stacking function (ENVI software), each band is combined into two groups of images, 10m spatial resolution bands Band2, Band3, Band4 and Band8; 20m spatial resolution bands Band2, Band3, Band4, Band8, Band5, Band6, Band7, Band8a, Band11 and Band12.

[0041] (3) Geometric correction

[0042] Due to the influence of speed, attitude, atmosphere, shape of the earth and ground rotation, satellite images will be distorted to a certain extent, resulting in some geometric distortion of satellite images. Therefore, before image analysis and processing, the image needs to be geometrically corrected to reduce the influence of distortion. Geometric correction is the process of eliminating or correcting geometric errors in remote sensing images. Geometric correction is divided into two categories: geometric rough correction and geometric fine correction. Geometric fine correction is mainly realized through the Image to image tool in ENVI. First, select the image data that has completed geometric correction as the reference image, then select the obvious and subtle intersection on the artificial surface as the control point, and register another raster file, so that the same target object appears in the same area of the corrected image, realizing the geometric fine correction of the image of Luxi Forest Farm

[56] .

[0043] (4) Image mosaic

[0044] Image mosaic, also known as image stitching, is a process of mosaicking adjacent multiple remote sensing images into a seamless image of large size on the basis of certain mathematics. The image stitching function of ENVI can combine several images together to generate a composite image, providing an interactive method, regardless of whether there are geographic coordinates or not. In general, after geometric correction, each image is planned in a completely unified coordinate system, and after appropriate cutting and removal of overlapping parts, the cut images are combined into a large format image.

[0045] (5) Image cropping

[0046] In daily remote sensing applications, many researchers are only interested in information within a certain range, so it is necessary to crop the remote sensing image into a certain size of the study area. This study takes the Lutou Forest Farm as the study area,

[0047] The ENVI5.3 region of interest cropping tool is used to crop Sentinel2 and DEM images, and finally the remote sensing images and DEM map of the study area are obtained.

[0048] III. Methods and principles

[0049] 3.1 Forest AGB remote sensing characteristic variable extraction

[0050] Based on Sentinel2 data and SRTM-DEM data, this study uses ENVI software to extract vegetation index, texture features, single-band information, and topographic features, etc. variables, making full use of the spectral, spatial, and topographic information resources provided by Sentinel2 remote sensing images, and extracting remote sensing characteristic variables as basic variables for experiments.

[0051] 3.1.1 Vegetation index extraction

[0052] Satellite remote sensing images contain rich vegetation information, which can be expressed and presented through the spectral characteristics, differences, and change rules of green plant leaves and vegetation canopy. Different types of vegetation exhibit different spectral information at different spectral bands. In different wavelength ranges in remote sensing satellites, vegetation exhibits different states. In the red light band 0.6-0.7um of visible light, vegetation has a strong absorption band, and in the near-infrared band 0.7-1.10um, vegetation index has many reflection peaks, so the data of these two bands are usually combined and operated in a linear or nonlinear manner, such as addition, subtraction, multiplication, and division to produce some values about vegetation growth rate and biomass. These values are called vegetation index. Vegetation index can distinguish plants from soil, water, and other backgrounds, and is an important indicator reflecting the existence, quantity, quality, growth status, and spatial distribution characteristics of vegetation, which can better and more effectively describe the green vegetation information on the ground.

[0053] The choice of vegetation index is important in the process of estimating forest AGB. Different vegetation indices will lead to different results. Through research, the vegetation index suitable for this study is as follows:

[0054] (1) Ratio Vegetation Index (RVI)

[0055] Ratio Vegetation Index (Ratio Vegetation Index, RVI) refers to the ratio between the near-infrared band and the red band reflectance. RVI is the earliest developed vegetation index, which is used to evaluate and monitor the degree of vegetation coverage. Taking 50% of vegetation coverage as the standard for measuring vegetation sensitivity, the coverage range of green and healthy vegetation is greater than 50% RVI, the vegetation index sensitivity is higher, and it is proportional to the vegetation coverage. If the vegetation coverage is less than 50%, its sensitivity will be significantly reduced. RVI will be affected by the surrounding atmospheric environmental conditions, such as atmospheric effects, which can greatly reduce the sensitivity of vegetation. Therefore, the atmospheric correction step is essential in the process of calculating RVI.

[0056] Because Sentinel2 data has a red edge band, the red edge band is considered to replace the red band to establish a new vegetation index RVI re5 Many scientific studies have shown that among the three bands of 705, 740, and 783 nm, the reflectance at 705 nm has a great correlation with chlorophyll content. Therefore, Band5 is mainly used as the red edge band in this study to calculate the vegetation index.

[0057]

[0058]

[0059] (2) Normalized Difference Vegetation Index (NDVI)

[0060] Normalized Difference Vegetation Index (Normalized Difference Vegetation Index, NDVI) is a widely used vegetation index, which is defined as the difference between the reflectance of the infrared band and the red band divided by the sum of the two. The numerical range is -1~1, negative value indicates that the surface cover of the study area is cloud, water, snow, etc.; the NDVI of green vegetation plants is 0.2~0.8; when NDVI=0, it indicates that the area has bare soil and rock.

[0061] NDVI is sensitive to green vegetation, and its sensitivity increases with the increase of vegetation coverage. However, when the vegetation coverage reaches a certain level, the growth rate will slow down. In this study, NDVI was used as a reference to evaluate and simulate remote sensing images, ground measurements, or new vegetation indices.

[0062] ① Normalized Difference Vegetation Index (NDVI)

[0063]

[0064] ② Normalized Difference Red Edge Vegetation Index (NDVIre)

[0065]

[0066]

[0067] The normalized red edge vegetation index has certain advantages compared to other vegetation indices, especially in terms of sensitivity to vegetation changes, mainly in the fine changes of leaf canopy such as forest gaps and aging. Many scientific research results show that the normalized red edge vegetation index has a close relationship with forest AGB.

[0068] ③ Modified Normalized Difference Vegetation Index (mNDVI)

[0069] mNDVI = (B8 - B4) / (B8 + B4 - 2B2) (7)

[0070] The modified normalized vegetation index adds the blue band in the calculation, which weakens the influence of factors such as water vapor, thereby improving the accuracy of the index.

[0071] ④ Modified Normalized Difference Red Edge Vegetation Index (mNDVI)

[0072] Similarly, the unique vegetation red edge band of Sentinel2 is used instead of the red band

[0073] mNDVI re5 = (B8 - B5) / (B8 + B5 - 2B2) (8)

[0074] ⑤ Normalized Difference Infrared Index (NDII)

[0075] NDII = (B8 - B11) / (B8 + B11) (9)

[0076] NDII is very sensitive to changes in crop canopy water content, and the index increases with the increase of water content, which can be used to monitor forest canopy.

[0077] (3) Green Chlorophyll Vegetation Index (Cl green )

[0078] Chlorophyll is the most important vegetation color in the process of photosynthesis, and it is also an important feature of the ability of photosynthesis. It has been used as the most important physiological parameter to monitor the growth and health of green vegetation.

[0079] ① Greenness Chlorophyll Index

[0080] Cl green = (B8 / B3)-1 (10)

[0081] ② Red Edge Chlorophyll Vegetation Index

[0082] Similarly, by replacing the green band with Band5 as the red edge band, the red edge chlorophyll vegetation index is obtained

[0083] Cl re5 = (B8 / B5)-1 (11)

[0084] (4) Enhanced Vegetation Index (EVI)

[0085] Enhanced Vegetation Index (EVI) is an "optimized" index, which aims to eliminate the interference of plant canopy and reduce the interference of atmosphere on plants, to enhance the vegetation signal in high biomass areas, so as to improve the vegetation sensitivity, improve the vegetation monitoring, and more accurately reflect the growth, development and change process of vegetation. The increase of blue band can enhance the vegetation signal, and thus greatly reduce the influence of soil and aerosol, and is often used in vegetation dense areas.

[0086] EVI = 2.5 (B8-B4) / B8 + 6B4-7.5B2 + 1 (12)

[0087] (5) Modified Simple Ratio (MSR)

[0088] The main purpose of Modified Simple Ratio (MSR) is to improve the problem of light saturation caused by changes in plant biochemical parameters. Its main feature is to fully consider the ratio of near-infrared and red band reflectance. Therefore, when detecting vegetation biochemical components, the error influence of atmosphere, soil and background can be eliminated, so as to achieve a good linear relationship.

[0089]

[0090] (6) Difference Vegetation Index (DVI)

[0091] Difference Vegetation Index (DVI) is the difference between the values of near-infrared band and visible band, DVI is more sensitive to the change of soil background than RMI, and its sensitivity to vegetation decreases when the vegetation coverage is high, so it can better identify the surrounding vegetation and water body, and reflect the change trend of vegetation coverage. In previous scientific research, DVI is mainly applied to monitor and evaluate the regional ecological conditions. On the contrary, RMI is more sensitive to high vegetation coverage and is suitable for monitoring natural forests.

[0092] DVI = B8 - B4 (14)

[0093] (7) Nonlinear Index (NLI)

[0094] NLI = ((B8) 2 -B4) / ((B8) 2 +B4) (15)

[0095] Similarly, Band5 is used as the red edge band to participate in the calculation of the nonlinear index, and the calculation formula is as follows:

[0096] NLI re5 =((B8) 2 -B5) / ((B8) 2 +B5) (16)

[0097] (8) New Inverted Red Edge Chlorophyll Index (IRECI)

[0098] New Inverted Red Edge Chlorophyll Index (IRECI) can quantitatively reflect the chlorophyll content of plants, and has good correlation with the chlorophyll content and leaf area index of vegetation layer.

[0099] IRECI = (B7 - B4) / (B5 - B6) (17)

[0100] Texture feature extraction

[0101] Texture is a visual feature that can represent the homogeneity in the image and is not affected by color and brightness, and it realizes the gray space distribution within the image pixel field. The surface of an object has its own inherent characteristics, and different objects form different textures, such as trees, textiles, and stone materials, etc. In addition, texture features can also reflect important information such as the surface structure, arrangement, and organization of objects. The human visual system mainly produces different perceptions of the environment by analyzing the texture features of matter. Therefore, some characteristics and changes of objects can be reflected through surface texture features. Therefore, texture feature information can be used as an important parameter and basis for distinguishing the attributes of ground objects.

[0102] In this study, the gray level co-occurrence matrix method was used to extract the texture feature information of Sentinel2 images to identify the ground object properties in the Lutou Forest Farm and further accurately distinguish the forest AGB.

[0103] (1) Gray level co-occurrence matrix theory

[0104] The gray level co-occurrence matrix is actually defined as the probability of another pixel with a gray level of j, starting from a pixel with a gray level of i, and a distance of (d x , d y ). The mathematical expression is P(i, j | d, θ) = #{(x, y) | f(x, y) = i, f(x + dx, y + dy) = j; x, y = 0, 1, 2, …, N-1}. In the formula, d is the relative distance expressed in pixel number; θ generally considers four directions, namely 0°, 45°, 90°, and 135°; # represents the set; i, j = 0, 1, 2, …, L-1; (x, y) is the pixel coordinate in the image, and L is the number of image gray levels. In short: in a gray image, the number of pixel pairs of a certain shape appearing in the entire image. The gray level co-occurrence matrix provides information about the direction, interval, change amplitude, and speed of the image gray level, but it cannot directly provide texture features, so it is necessary to extract statistical properties based on the gray level co-occurrence matrix to quantitatively describe the texture features.

[0105] (2) Factors affecting the gray level co-occurrence matrix

[0106] The extraction of texture features is not only affected by the method of extraction, but also by factors such as the gray level of the image, the window size and moving step, and the spectral band. According to the basic spatial correlation theory, the spatial correlation intensity is related to the distance, and the closer the distance, the stronger the spatial correlation. Therefore, we can choose Band2, Band3, Band4, and Band8 with a spatial resolution of 10m in Sentinel2 data for texture feature extraction, and set the window size to 3x3, 5x5, 7x7, 9x9, and 11x11, respectively, use the mean value of 45°, 90°, 135°, and 180°, and use a moving step of 1 to calculate the texture features.

[0107] (3) Calculation formula of texture feature extraction

[0108] Haralikc et al. proposed 14 kinds of texture features. In remote sensing image texture information, the commonly used feature statistics mainly include correlation measure, mean measure, contrast measure, synergy measure, dissimilarity measure, entropy measure, second moment measure, and variance measure. The calculation formula is shown in Table 3:

[0109] Table 3 Calculation table of gray level co-occurrence matrix texture features

[0110]

[0111]

[0112] Note: M, N size of gray level co-occurrence matrix, in general case M = N, and M, N equal to the set of gray level

[0113] (4) Principal component analysis

[0114] Principal component analysis is a dimension reduction method. Through principal component analysis, it attempts to construct some comprehensive variables with as few variables as possible, independent of each other, containing as much overall information as possible contained in the original variables, using a few new indicators to replace the original indicators, thereby reducing the dimension of data and the amount of calculation. Principal component analysis is based on variance to extract the most valuable information. The amount of data information changes with the size of variance. The greater the variance, the greater the amount of information contained in the data. The smaller the variance, the less the amount of information contained in the data. According to the amount of information contained, the first principal component and the second principal component are selected as the main variables. Each comprehensive variable is a linear combination of the original variables, and they are not related to each other.

[0115] The correlation between the texture feature variables under different windows and the forest AGB is compared, and the 7x7 window with a higher degree of correlation of most variables is selected. According to the experimental process, it is found that the texture information of Band2, Band3, Band4 and Band8 has strong correlation, so principal component operation is performed on Band2, Band3, Band4 and Band8 of Sentinel2 multispectral image in remote sensing image processing software ENVI5.3, and finally 11 new texture features are obtained: 1 mean principal component (P 1MEAN ), 2 variance principal components (P 1VARI , P 2VARI ), 1 contrast principal component (P 1CONT ), 2 correlation principal components (P 1CORR , P 2CORR ), 1 homogeneity principal component (P 1HOMO ), 2 dissimilarity principal components (P 1DISS , P 2DISS ), 1 entropy principal component (P 1ENTR ), 1 second moment principal component (P 1SECD ).

[0116] Single band extraction

[0117] According to the spectral characteristics of plants, the vegetation has different absorption and reflectivity to different wave bands, and the pixels reflected to the image are different. Therefore, the amount of information contained in each single wave band pixel is different. Therefore, as much as possible, the signal reflecting the vegetation information from each wave band needs to be extracted. Vegetation clarity is related to the spatial resolution of remote sensing images, and vegetation clarity increases with the increase of spatial resolution. Therefore, 10 wave bands with 20-meter and 10-meter spatial resolution in Sentinel-2 data are selected as the estimation of forest above-ground biomass remote sensing variables.

[0118] Terrain variable extraction

[0119] Among the numerous remote sensing variables, terrain variables are the most frequently used variables. Everything is affected by terrain features, which mainly include slope, elevation, aspect, and slope position. Through reading a large number of articles by scholars, it is known that forest AGB is mainly affected by terrain variables such as elevation, slope, and aspect. DEM data with a spatial resolution of 30 meters is selected, and the 3DAnalyst tool under ArcToolbox in Arcgis software is used to extract the slope, aspect, and elevation of the study area.

[0120] Correlation analysis of forest AGB remote sensing variables

[0121] Correlation analysis principle

[0122] Forest AGB estimation is related to many factors, and the variables related to different regions are different. However, the results of selecting independent variables using different remote sensing data sources are also different. Therefore, determining the best variable is an important step to improve the estimation accuracy of forest AGB. In this study, SPSS 25.0 software was used to analyze the correlation between the 40 selected characteristic variables and forest AGB. Correlation analysis refers to the analysis of two or more variables with correlation, so as to measure the degree of correlation between two variable factors. Correlation coefficient is a representative index that measures the degree and direction of linear correlation.

[0123] Pearson correlation coefficient (Pearson correlation coefficient) is used to reflect the degree of linear correlation between two random variables, and reflects the linear correlation between two variables. The value of r ranges from -1 to 1. When the value is 1, it indicates that there is an absolute positive correlation between the two random variables; when the value is -1, it indicates that there is an absolute negative correlation between the two random variables; and when the value is 0, it indicates that there is no linear relationship between the two random variables. The calculation formula of Pearson correlation coefficient is as follows:

[0124]

[0125] (wherein: n, X i , Y i , respectively represent the sample number, the i-point observation value corresponding to the variable X, Y, the X sample average, and the Y sample average)

[0126] Correlation analysis results

[0127] Generally speaking, each type of forest AGB has different forest structure and characteristics. Therefore, 16 vegetation indexes, 11 texture features, 10 single-band information, and 3 terrain variables, a total of 40 characteristic variables, were selected as remote sensing preparation variables for estimating forest AGB. Based on these data, the correlation between the preparation variables and forest aboveground biomass was determined.

[0128] The correlation degree between the selected characteristic variable group and forest AGB was analyzed using SPSS software. The results showed that the terrain features, vegetation index, and forest AGB had high correlation, while the texture features and single-band had lower correlation with forest AGB. In terms of specific situations, terrain features DEM (correlation coefficient absolute value greater than 0.133), single-band Band2, Band3, Band4, Band5, Band6, Band11, Band12 (correlation coefficient absolute value greater than 0.102), texture features P 1ENTR , P 1SECD , P 1VARI , P 2VARI , P 1CONT , P 1CORR , P 2CORR , P 1DISS (correlation coefficient absolute value greater than 0.079), and vegetation index NDVI, NDVI re5 , NDVI re6 , RVI, RVI re5 , Cl green , EVI, mNDVI, MSR, mNDVI re5 , NLI re5 , IRECI (correlation coefficient absolute value greater than 0.079) had high correlation with forest AGB and the relationship was significant. Therefore, 12 vegetation indexes, 8 texture features, 7 single-band variables, and 1 terrain variable with strong correlation with forest AGB were selected, totaling 28, to prepare for further screening of forest AGB variables.

[0129] Through various methods, the vegetation index, texture features, single band, terrain variables and other characteristic variables in remote sensing data were extracted, and their correlation analysis was carried out. According to the Pearson correlation results, there is a good correlation between forest AGB and vegetation index. According to the various vegetation indexes calculated by different bands, the vegetation index obtained by the red edge band has a strong correlation with forest AGB. Compared with the red edge band, the correlation coefficient of the vegetation index obtained by the red band is low. Among the many red edge vegetation indexes related to forest AGB, RVI re5 and NDVI re5 have the highest correlation. Single band information is used to reflect vegetation information, and Band2, Band3 and Band4 are the most obvious, which may be because the study area is a forest farm, the trees are tall and the leaves are full of chlorophyll, and Band2, Band3 and Band4 are extremely sensitive to leaf green. Texture feature variables have been used to identify ground object attributes. According to the correlation analysis results, the correlation between the texture feature variables obtained by principal component analysis and forest AGB is low. Because the elevation of the study area is very different, the elevation in the terrain variables has a strong correlation with forest aboveground biomass.

[0130] Boruta algorithm + SVR

[0131] Boruta algorithm principle

[0132] Feature selection can capture the importance of all features related to the target variable in the data, and can better select important features to analyze the complex relationship between forest AGB and remote sensing variables. Boruta algorithm is a feature selection algorithm based on random forest, which compares the importance of initial features and shadow features to achieve the purpose of obtaining the best feature combination. The basic idea of the algorithm is to copy the initial feature dataset, randomly shuffle the row order of each feature variable to form a shadow feature, and then mix the initial feature and the shadow feature to form a new feature dataset. Based on the multiple iterations of the random forest model, the importance score of each feature is obtained by using the random forest importance score calculation formula. The maximum value of the importance of the shadow feature is defined as Z score , and the importance of the initial feature is higher than Z score , which is marked as the most important, otherwise, it is marked as unimportant and the initial feature is deleted, and finally the shadow feature is deleted to obtain the best feature combination for modeling. According to the out-of-bag error of the RF model, the importance score calculation formula of the Boruta algorithm is determined, and the formula is:

[0133]

[0134] MSE OOB , y i , represent the out-of-bag error of random forest, sample value, predicted value of out-of-bag sample of sample y i , respectively.

[0135]

[0136] Z score , SDMSE OOB represent the Z-score, mean of out-of-bag error, standard deviation of out-of-bag error, respectively.

[0137] Boruta algorithm flow

[0138] The Boruta feature selection algorithm flow is shown in Figure 2 and described as follows:

[0139] For the initial feature data set A (m rows n columns, m groups of samples, n initial features, m>1, n>1), clone the initial feature data set A and perform random row transformation independently for each column as a shadow feature B (m rows n columns), merge the initial feature data set A and the shadow feature data set B to obtain a new data set C (m rows 2n columns). Train the random forest model based on the data set C to obtain the importance score of each initial feature and shadow feature, and the maximum value of the importance score of the shadow feature is defined as Z score The initial feature variables with an importance degree significantly greater than Z score are marked as important. The initial feature variables with an importance degree close to Z score are marked as tentative. The initial feature variables with an importance degree significantly less than Z score are marked as unimportant and permanently deleted. The initial feature variables marked as tentative are retained, and then all shadow features are deleted. Until all initial features are assigned importance, the best optimal features are obtained.

[0140] Boruta feature selection

[0141] Based on the 28 variables screened out by Pearson correlation analysis, the Boruta feature selection algorithm was used to select the best feature variable combination (about forest AGB). The Boruta() function in R language was used to extract the features. In the process of Boruta algorithm, set.seed() was used to set a random number seed. A specific seed could generate a specific pseudo-random sequence. The main purpose of this function was to make the simulation reproducible. doTrace: It refers to the level of detail. Setting the doTrace parameter to 1 or 2 reports the progress of the process. 0 means no tracking, 1 means reporting decisions as soon as attributes are cleaned, and 2 means all 1 plus reporting each iteration. The default is generally 0. maxRuns: The maximum number of random forest runs. If temporary attributes are retained, you can consider increasing this parameter, which defaults to 100. holdHistory: Stores the full history of importance runs when it is set to TRUE by default. Based on the final results of Boruta feature selection, 13 features were retained, and the 13 retained features were: NDVI re5 , RVI re5 , mNDVI, P 1CONT , EVI, Cl green , P 1DISS , RVI, NDVI, MSR, DEM, NDVI re6 , Band11, a total of 13 features.

[0142] Accuracy verification and result analysis

[0143] Using the forest AGB calculated from the 861 sample plot survey data as the dependent variable, the Boruta algorithm screened NDVI re5 , RVI re5 , mNDVI, P 1CONT , EVI, Cl green , P 1DISS , RVI, NDVI, MSR, DEM, NDVI re6 , Band11, a total of 13 features as independent variables, RBF, grid search method, and ten-fold cross-validation were selected as the kernel function, parameter optimization method (C=2, g=0.01) of the Boruta-SVR forest AGB estimation model. The reserved 371 sample plot data were used to verify the accuracy of the Boruta-SVR model. From the accuracy verification results of the Boruta-SVR model in Figure 5-4, there was a certain degree of correlation between the estimated value and the sample plot AGB measured value, and the pearson correlation coefficient was 0.79. The fitting determination coefficient R 2The RMSE is 0.63, the RMSE is 28.28 t / ha, and the MAE is 23.04 t / ha. The estimated sample AGB result is 74.56 t / ha, which is close to the measured sample aboveground biomass of 74.12 t / ha, with a difference of 0.44 t / ha. The verification results are shown in Table 4 and Figure 3 :

[0144] Table 4 Forest aboveground biomass Boruta-SVR model precision verification

[0145]

[0146] From the above table and Figure 3 , although the Boruta-SVR model improves the model precision in estimating forest AGB, making the model have better estimation performance, in the above experimental application, the Boruta algorithm has problems such as high complexity of shadow feature samples and slow iteration calculation time. In order to solve the above problems, a parameter is added to control the proportion of shadow features in the Boruta algorithm in the following research process, and the model trained by the data set is changed to XGBoost model, which can reduce the sample complexity, reduce the iteration time of the algorithm, and improve the estimation performance of the model.

[0147] Improved Boruta algorithm

[0148] The improved Boruta algorithm in this study is a feature selection algorithm based on XGBoost model, which compares the importance of initial features and shadow features to achieve the purpose of obtaining the optimal feature combination. The basic idea of this algorithm is to create a shadow feature for the initial feature (m rows n columns, m groups of samples, n initial features, m>1, n>1) first. The shadow feature is obtained by extracting [m*p]*n groups of samples from the initial sample in proportion P (0<=p<1), and then putting back after random row transformation. Then combine the initial feature and the shadow feature into a new feature combination. Based on the multiple iterations of the XGBoost model, the importance score calculation formula of the XGBoost model is used to obtain the importance score of each feature. The maximum value of the importance of the shadow feature is defined as Z score , and the importance of the initial feature is higher than Z score , which is marked as the most important, otherwise, it is marked as unimportant and the initial feature is deleted, and finally the shadow feature is deleted to obtain the best feature combination for modeling.

[0149] Feature importance is based on XGBoost model, which constructs K classification trees. In each tree, the importance of a feature is obtained by using the number of times the split node is used, and the F Score value is sorted. The F scoreThe maximum value of Z score .

[0150] Improved Boruta algorithm process

[0151] For the initial feature data set A (m rows n columns, m groups of samples, n initial features, m>1, n>1), clone the initial feature data set A to obtain the shadow feature data set B. The shadow feature data set B extracts [m*p]*n groups of samples according to the proportion P (0<=p<1), mixes and shuffles the [m*p]*n groups of data to obtain the shadow feature data set C, randomly shuffles the row sequence to obtain the shadow feature data set D, and merges the initial feature data set A and the shadow feature data set D to obtain a new set E. Train the XGBoost model based on the data set E to obtain the importance score of each initial feature and shadow feature. The maximum value of the importance score of the shadow feature is defined as Z score The importance of each variable of the initial feature is obviously greater than Z score , which is marked as important. The importance of each variable of the initial feature is close to Z score , which is marked as tentative. The importance of each variable of the initial feature is obviously less than Z score , which is marked as unimportant and permanently deleted. Keep the initial feature variables marked as tentative, and then delete all shadow features. Until all initial features are assigned importance, the best optimal features are obtained. The improved Boruta algorithm process diagram is as Figure 4 .

[0152] Improved Boruta algorithm feature selection

[0153] According to the 28 variables screened out by pearson correlation analysis, the improved Boruta feature selection algorithm was used to select the best feature variable combination (about forest AGB). The Borutapy () function of Python language was used to extract the features. In the process of improving the Boruta algorithm, the shadow features ([m*p]*n) were created first, and then the XGBoost was defined as a random forest classifier. The improved Boruta feature selection parameters were set as follows: estimators: set the evaluator. Since XGBoost supports decision trees, it is set to Forest. n-estimators: set the number of estimators. Generally, n_estimators is too small, which is easy to underfit, and n_estimators is too large, which is easy to overfit. A moderate value is generally set, and the value of auto is set in this study. Verbose: control output. 0 means output the training process, 1 means output occasionally, and >1 means output for each. The value of 2 was set in this study, that is, each variable was output. Random_state: the seed used by the random number generator, which is set to the default value of 1. max_iter: the maximum number of iterations. In order to compare the performance of the improved Boruta algorithm, the number of iterations is set to 99, which is the same as the Boruta algorithm. After many experiments, it was found that there were 10 rejected features, NDVI re6 , P 2CORR , IRECI, P 1SECD , Band6, Band4, Band3, P 2VARI ,, Band2, NLI re5 There are 8 pending features, Band11, P 1VARI , mNDVI re5 , NDVI re5 , RVI re5 , Band5, Band12, P 1CORR There are a total of 8, and the reserved features are mNDVI, P 1CONT , EVI, Cl green , P 1DISs , RVI, NDVI, MSR, DEM, P 1ENTR There are a total of 10 features.

[0154] By adjusting the n-estimators parameter, the 8 pending features, Band11, P 1VARI , mNDVI re5 , NDVI re5 , RVI re5 , Band5, Band12, P 1CORR were further tested to see if they needed to be retained. As a result, NDVI was retained.re5 , RVI re5 , Band11, mNDVI re5 These four features, reject P 1VARI , Band5, P 1CORR , Band12 four features. Improved Boruta feature selection results for 14 features: mNDVI, P 1CONT , EVI, Cl green , P 1DISS , RVI, NDVI, MSR, DEM, P 1ENTR , NDVI re5 , RVI re5 , Bandll, mNDVI re5 .

[0155] Improved Boruta variable selection results can be seen that 1) the terrain features have the greatest impact on the study of AGB estimates, which is most significant in the elevation variable. 2) RVI, NDVI, RVI re5 , NDVI re5 , Cl green Used to monitor the distinction between the study of land cover and vegetation growth and health, EVI, MSR, mNDVI, mNDVI re5 To reduce the impact of atmospheric, soil, background and other factors, can more accurately distinguish biomass, therefore, RVI, NDVI, MSR and other variables with forest AGB estimation correlation. 3) P 1DISS , P 1CONT , P 1ENTR Mainly used to identify important parameters and basis of the property of the object, easy to distinguish forest AGB, therefore, P 1DISS , P 1CONT , P 1VARI , P 1ENTR Also have a certain correlation with forest AGB. 4) Band11 on forest AGB part sensitive, used to distinguish snow, cloud and ice index, also with forest AGB estimation. Improved Boruta algorithm selected variables and forest AGB most relevant and comprehensive, in line with the actual situation.

[0156] In summary, after 99 iterations, the improved Boruta algorithm selected 14 features, rejected 14 features, Boruta algorithm selected 13 features, rejected 15 features, about the two kinds of Boruta algorithm before and after the improvement will be marked as important features of the set of features reserved. Through the number of features, the length of time when the iteration algorithm, R 2 Three indicators as the judgment of the improved Boruta algorithm before and after the two kinds of algorithm feature set. The evaluation results are shown in Table 5:

[0157] Table 5 Comparison of results before and after improvement of Boruta algorithm

[0158]

[0159] Note: The average time of 10 experimental data of the iteration calculation before and after the improvement of Boruta in the table

[0160] According to Table 5, the initial number of features is 28, the number of features selected by the Boruta algorithm is 13, and the number of features selected by the improved Boruta is 14, which is one more than the number of features selected by the Boruta algorithm. At the same time, the iteration calculation time and R 2 of the improved Boruta algorithm are better than those of the Boruta algorithm to some extent. It is shown that under the overall framework of the Boruta algorithm, introducing a regular parameter to control the proportion of shadow features of the Boruta algorithm and changing the model trained by the data set to the XGBoost model indeed reduces the iteration time of the algorithm, reduces the complexity of the shadow features, and improves the prediction performance of the model.

[0161] Accuracy verification and result analysis

[0162] The forest AGB calculated by the data of 861 sample plots is the dependent variable, and the NDVI, NDVI re5 , RVI, RVI re5 , Cl green , EVI, mNDVI, MSR, mNDVI re5 , DEM, Band11, P 1ENTR , P 1CONT , P 1DISS selected by the improved Boruta algorithm are the independent variables, which are the kernel function of SVM. The grid search method and ten-fold cross-validation are used as the parameter optimization method (c=2, g=0.01), and the RF-SVR model is established for forest AGB estimation. The reserved 371 sample plot data are used to verify the accuracy of the improved Boruta-SVR model. From the accuracy verification results of the improved Boruta-SVR model, there is a certain degree of correlation between the estimated value and the measured value of the sample plot AGB. The pearson correlation coefficient is 0.82, the fitting determination coefficient R 2 between the estimated value and the measured value is 0.68, the RMSE is 24.31 t / ha, and the MAE is 17.97 t / ha. The average of the estimated sample plot AGB is 74.13 t / ha, which is only 0.01 t / ha different from the measured average of the aboveground forest biomass, which is 74.12 t / ha. The accuracy verification results are shown in Table 6 and Figure 5 :

[0163] Table 6 Improved Boruta-SVR precision verification of forest aboveground biomass

[0164]

[0165] The improved Boruta-SVR model has certain advantages in estimating forest AGB. There are many machine learning feature selection algorithms. In order to reflect the advantages of the improved Boruta algorithm in feature selection, the commonly used Lasso and RF algorithms are introduced to optimize the SVR model to estimate forest aboveground biomass AGB.

[0166] Through the variable selection of the improved Boruta algorithm, first, the shadow variables are created for the initial variables, and then a regular parameter is introduced to control the proportion of shadow variables. The feature importance is measured by the number of feature splits F score in the XGBoost model, and by comparing the importance of the initial variables and the shadow variables, the optimal combination of variable selection algorithm is obtained.

[0167] The improved Boruta-SVR model is used to retrieve the AGB of Lutou Forest Farm in May 2020, and the AGB distribution of the study area is obtained. The total AGB of Lutou Forest Farm is 4.53*10 5 t. The spatial distribution of forest AGB in Lutou Forest Farm shows a north-south trend, gradually increasing and becoming dense from south to north, mainly distributed in the northwest and central areas of the forest farm with gentle terrain, and less distributed in the southwest and southeast of the forest farm. Due to the high altitude of the surrounding area and the sparse vegetation, the AGB is less. The distribution of AGB in the study area is roughly consistent with the actual geological and geomorphic conditions and the distribution of forest vegetation.

[0168] In this study, the variables most closely related to forest AGB are selected to optimize the model and improve the estimation accuracy of forest AGB. To solve this problem, the machine learning Boruta algorithm is used for variable selection, and the difficulties encountered in the running process are improved to obtain the improved Boruta algorithm. At the same time, the model fitting degree, the number of feature variables, and the length of iterative calculation before and after the improvement of the machine learning Boruta algorithm are compared and analyzed. In order to reflect the advantages of the improved Boruta algorithm in variable selection algorithm, the SVR estimation model of forest aboveground biomass constructed by the improved Boruta algorithm has high precision and the smallest root mean square error, which shows that the improved Boruta algorithm plays a role in estimating the AGB of the study area.

[0169] The results show that the feature variables selected by the improved Boruta algorithm are more comprehensive, and the estimation accuracy of the optimized SVR model is improved compared with the SVR model optimized by the Boruta algorithm. The improved Boruta algorithm reduces the complexity of shadow features in the feature selection process and reduces the iteration time. The NDVI, NDVI re5 , RVI, RVI re5 , Cl green , EVI, mNDVI, MSR, DEM, Band11, P 1ENTR , mNDVI re5 , P 1CONT , P 1DISS 14 features are selected as input variables of the SVR model, and the improved Boruta-SVR forest AGB estimation model is constructed. The reserved validation data is used to test the accuracy of the improved Boruta-SVR model, and the results show that the R 2 of the improved Boruta-SVR model is 0.68, the RMSE is 24.31 t / ha, and the MAE is 17.97 t / ha.

[0170] By comparing the Boruta-SVR forest AGB estimation model, the results show that the SVR model optimized by the improved Boruta algorithm has high precision and superior performance in variable selection.

[0171] The improved Boruta-SVR optimal model is selected to invert the forest AGB of Lutou Forest Farm in May 2020, and the AGB distribution map of the study area is obtained. The results show that the spatial distribution of AGB in the study area presents a north-south trend, gradually increases and is dense from south to north, and mainly distributes in the northwest, northeast and central areas of the forest farm with gentle terrain. The southwest and southeast of the forest farm have less aboveground biomass. The surrounding area of the forest farm has high elevation and sparse vegetation, and the AGB is less. The AGB distribution of the study area is roughly consistent with the actual investigation of the geological and geomorphic conditions and the distribution of forest vegetation in the study area.

[0172] The above-described embodiments only express several embodiments of the present application, and the description is more specific and detailed, but it cannot be understood as limiting the scope of the present patent. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present patent should be subject to the appended claims.

Claims

1. A remote sensing method for forest aboveground biomass measurement using an improved Boruta algorithm, characterized in that, The method includes the following steps: (1) Select remote sensing image data of the area to be measured and preprocess the data; (2) Design an improved Boruta algorithm; (3) Establish an SVR model using the improved Boruta algorithm; (4) Use the model to measure the aboveground biomass of the forest in the area to be tested; Specifically, step (2) of improving the Boruta algorithm is as follows: For the initial feature dataset A, the initial dataset has m rows and n columns, with m groups of samples and n initial features, m>1, n>1, and then the initial feature dataset A is cloned to obtain the shadow feature dataset B; The shadow feature dataset B extracts [m*p]*n groups of samples according to the proportion P, where 0<=p<1. The [m*p]*n groups of data are mixed and shuffled to obtain the shadow feature dataset C. The row order is randomly shuffled to obtain the shadow feature dataset D. The initial feature dataset A and the shadow feature dataset D are merged to obtain a new set E. The XGBoost-based model was trained using dataset E to obtain importance scores for each initial feature and shadow feature. The maximum value of the importance score of the shadow feature was defined as Z. score The importance of each variable in the initial feature is significantly greater than that of Z. score Marked as important; for the initial features, the degree of importance of each variable is close to Z. score Marked as provisional; the importance of each variable in the initial feature is significantly less than that of Z. score Mark the initial features as unimportant and delete them permanently; retain the initial feature variables marked as provisional and then delete all shadow features; continue until all initial features are assigned importance to obtain the best and optimal features.

2. The forest aboveground biomass remote sensing measurement method based on the improved Boruta algorithm according to claim 1, characterized in that, Step (1) is to select Sentinel2 image data.

3. The forest aboveground biomass remote sensing measurement method based on the improved Boruta algorithm according to claim 1, characterized in that, Step (1) The Sentinel2 image data contains multispectral data of 13 bands, of which Band2, Band3, Band4 and Band8 have a spatial resolution of 10m, Band5, Band6, Band7, Band8a, Band11 and Band12 have a spatial resolution of 20m, and Band1, Band9 and Band10 have a spatial resolution of 60m.

4. The forest aboveground biomass remote sensing measurement method based on the improved Boruta algorithm according to claim 1, characterized in that, Step (1) involves data preprocessing using the Sencor plugin in the SNAP software provided by ESA to perform radiometric calibration and atmospheric correction on the LIC-level standard product data of Sentinel 2.