Forest tree height inversion method and system based on multi-source remote sensing data and Unet network

Through the improved Unet deep learning model, the problem of forest tree height measurement is solved, high-precision and large-scale forest tree height calculation is achieved, and the scientificity and accuracy of forest resource management is improved.

CN120124012APending Publication Date: 2025-06-10NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202510585406.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has difficulties in measuring high-resolution forest trees, traditional methods are time-consuming and labor-intensive and data acquisition is difficult. Although remote sensing technology has the advantages of large-area observation, it has problems such as poor penetration ability and high noise.

Method used

The forest tree height inversion method based on multi-source remote sensing data and improved Unet deep learning neural network model is adopted, and the satellite-based lidar data and optical remote sensing image data are integrated, and the model is optimized through the two-dimensional dense connection module and attention mechanism module to achieve highly accurate forest prediction in large-scale continuous areas.

Benefits of technology

High-precision and large-scale forest tree height calculations are realized, the effect of forest tree height inversion model is improved, sparse data can be better processed, and the scientificity and accuracy of forest resource management and protection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124012A_ABST
    Figure CN120124012A_ABST
Patent Text Reader

Abstract

The invention discloses a forest height inversion method and system based on multi-source remote sensing data and a Unet network, and particularly relates to the technical field of remote sensing data and deep learning, and the method comprises the steps: obtaining multi-source remote sensing data, including satellite-borne laser radar data and optical remote sensing image data; preprocessing the acquired multi-source remote sensing data, and constructing a network training data set; a Unet deep learning network model is improved, model parameters are optimized through training data, and a forest tree height inversion model is generated; based on a Unet basic deep learning network model, replacing a conventional convolution module with a two-dimensional dense connection module 2D DenseBlock, and adding an attention mechanism module CBAM; and utilizing the generated forest tree height inversion model to invert the forest tree height, and evaluating the inversion precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing data and deep learning, and particularly to a method and system for forest tree height inversion based on multi-source remote sensing data and Unet network. Background Art

[0002] Forests play a crucial role in the global carbon cycle. As one of the core parameters of the forest vertical structure, forest tree height is of great significance in estimating above-ground forest biomass and wood volume, monitoring the impact of forest degradation, evaluating the effectiveness of forest restoration, simulating other key ecosystem variables, and helping to understand the forest carbon cycle process.

[0003] In the past, forest inventory was the main way to estimate forest biomass and height. It could provide reliable statistical information for large areas of forests, but it could not generate high-resolution maps. Traditional forest height surveys used sampling surveys and manual measurements, which were not only time-consuming and laborious, but also extremely difficult to obtain data.

[0004] In recent decades, remote sensing technology has been widely used in the field of forest parameter acquisition due to its characteristics of large-area synchronous observation, long-term continuous observation, and rich information. The research on estimating forest tree height based on different remote sensing sources mainly includes the following three categories: Based on optical remote sensing data: Optical remote sensing images can only obtain spectral information in the horizontal direction, have poor penetration ability, have great limitations in obtaining forest vertical structure parameters, and are significantly affected by weather; Based on synthetic aperture radar data: Although synthetic aperture radar can achieve continuous observation of the observed object all day and all weather, it is greatly affected by terrain, has a lot of noise, high interpretation difficulty, complex formula derivation, and is difficult to play an effective role in practical applications; Based on lidar data: Lidar has unique penetration ability, can well obtain the vertical structure information of trees, is less affected by weather and has high precision, and has become the mainstream means of obtaining tree height at present. Lidar mainly includes three types: ground-based, airborne, and spaceborne. The first two can only obtain tree height data in a small area, and the operation cost is relatively high, which is not suitable for obtaining large-area tree height data. In contrast, spaceborne lidar is more suitable for large-area tree height mapping, but its data is a series of discontinuous light spots, and it is impossible to obtain regional and densely covered tree height data, and deep learning needs to be used to convert the discrete light spot data into regional surface data.

[0005] In recent years, deep learning has developed rapidly, and it has become a reality to autonomously learn high-level features of images. As a well-known model in deep learning, the convolutional neural network performs excellently in the field of image processing. Remote sensing images have the characteristics of rich information, high resolution, continuous imaging, etc., and can reflect ground object information to a certain extent. Therefore, by using the ability of the convolutional neural network to autonomously learn high-level features of images, features reflecting tree height information can be extracted from remote sensing images. At present, there are few studies on using convolutional neural networks combined with remote sensing images for forest height prediction, and most of them still remain in the machine learning stage. However, machine learning cannot effectively utilize spatial information, resulting in generally low prediction accuracy. Summary of the Invention

[0006] In view of this, the present invention improves the Unet deep learning neural network model, integrates lidar tree height data and optical remote sensing images, and designs a forest tree height inversion method and system based on multi-source remote sensing data and the Unet network to achieve accurate prediction of forest heights in large-scale and continuous areas, so as to solve the problems raised in the background technology.

[0007] The specific technical solution is as follows: The forest tree height inversion method based on multi-source remote sensing data and the Unet network includes the following steps: (1) Obtain multi-source remote sensing data, including spaceborne lidar data and optical remote sensing image data; (2) Preprocess the obtained multi-source remote sensing data to construct a network training data set; (3) Improve the Unet deep learning network model, and optimize the model parameters through training data to generate a forest tree height inversion model; Based on the Unet basic deep learning network model, use the two-dimensional dense connection module 2D DenseBlock to replace the conventional convolutional module, and add the attention mechanism module CBAM; (4) Use the generated forest tree height inversion model to invert the forest tree height and evaluate the inversion accuracy.

[0008] Preferably, step (2) includes the following steps: (2.1) Spaceborne lidar data preprocessing: To ensure the quality of spaceborne lidar data, the screening rules are as follows: 1) Exclude low-sensitivity footprints; 2) Exclude footprints with low elevation accuracy; 3) Exclude any data with low data quality; 4) Exclude footprint points that may have various errors.

[0009] (2.2) Optical remote sensing image data preprocessing: Mask the pixels of the remote sensing image according to the quality assessment band QA_PIXEL to complete the image cloud removal; All images are superimposed, and their median values ​​are taken as the final pixel values, and they are synthesized into one image, namely, the annual image median synthesis.

[0010] (2.3) Calculation of vegetation index: The calculated vegetation index includes: Normalized Difference Vegetation Index NDVI, Normalized Difference Wetness Index NDMI, Normalized Difference Water Index NDWI, Atmospheric Impedance Vegetation Index ARVI, Enhanced Vegetation Index EVI and Ratio Vegetation Index RVI. The specific calculation formula is as follows: ; ; ; ; ; ; Among them, BLUE, GREEN, RED, NIR, and SWIR represent the reflectance values ​​of the blue band, green band, red band, near-infrared band, and short-wave infrared band, respectively.

[0011] (2.4) Matching of space-borne lidar data and optical remote sensing image data: The satellite-borne lidar data is rasterized using ArcGIS 10.8 software. At this time, the grids with light spots are valid grids, and their values ​​are valid values. The grids without light spots are invalid grids, and their values ​​are invalid values ​​NAN. The grid resolution was set to 30 meters, and bilinear interpolation technology was used to resample the sample data that could not be directly mapped to a specific grid cell; The optical remote sensing image data is loaded and aligned with the rasterized and resampled satellite-borne lidar data to obtain a 13-band data set, of which the first 12 bands are 30-meter-resolution optical remote sensing image data and vegetation index data, and the 13th band is the satellite-borne lidar data resampled to 30 meters.

[0012] (2.5) Network training data set preparation: The above complete 13 band data are cut into multiple data groups according to the size of 256×256 pixels, and the data groups without valid ground truth values ​​are eliminated to ensure the quality of the data set for model training; Through random sampling, the screened data is divided into a training set, a test set, and a validation set according to the ratio of 40%:30%:30%. A total of 3,774 data groups are retained, namely 1,510 training sets, 1,132 test sets, and 1,132 validation sets. The first 12 bands are the input data of the model in the following text, and the 13th band is the label of the model.

[0013] Preferably, the expression of the two-dimensional dense connection module 2D DenseBlock is: ; where refers to the concatenation of the feature maps generated from the 0th layer to the layer; H l represents the l- 1st layer of feature maps to the l layer of feature maps, including convolution, batch normalization, and non-linear activation operations; The output of the 2D DenseBlock is composed of the feature maps of all internal layers. This integrated feature representation provides richer information for forest tree height inversion.

[0014] Preferably, the attention mechanism module CBAM includes a channel attention mechanism CAM and a spatial attention mechanism SAM: The channel attention weight and the spatial attention weight The calculation formulas are respectively: ; ; where, represents the input feature map, represents the sigmoid function, MLP represents a two-layer fully connected neural network, represents the global average pooling operation, represents the global maximum pooling operation, represents a convolution operation with a 7*7 convolution kernel. The channel attention mechanism is performed in the spatial dimension, and the spatial attention mechanism is performed in the channel dimension; The feature map output after passing through CBAM enables the Unet deep learning network model to focus on the features most relevant to the forest tree height inversion task, so as to improve the learning ability of the forest tree height inversion model for sparse data.

[0015] Preferably, the loss function of the training data in step (3) is as follows: For the prediction of forest tree height, the mean absolute error MAE is selected as the loss function, and the calculation formula is: ; Wherein, n is the number of true value points, is the model predicted value, is the true value, and the range of the mean absolute error is , which is equal to 0 when the predicted value and the true value are exactly the same, that is, perfect prediction; the worse the prediction effect, the greater the error value between the predicted value and the true value.

[0016] Preferably, step (4) includes the following steps: (4.1) Forest tree height inversion: Integrate the optical remote sensing data and vegetation index data in the study area, input the optical remote sensing data and vegetation index data into the forest tree height inversion model, and perform forest height inversion in this area; (4.2) Inversion accuracy evaluation: Compare the model predicted value with the true value to evaluate the inversion accuracy; use the root mean square error RMSE and the determination coefficient R 2 as the accuracy evaluation index, and the formula is as follows: Root mean square error RMSE: ; Determination coefficient R 2 : ; Wherein, n is the total number of samples, is the model predicted value, is the true value, is the average value of the true value; the range of the root mean square error is , which is equal to 0 when the predicted value and the true value are exactly the same, that is, perfect prediction; the worse the prediction effect, the greater the root mean square error value; the range of the determination coefficient is , which is equal to 1 when the predicted value and the true value are exactly the same, that is, perfect prediction; the worse the prediction effect, the smaller the determination coefficient.

[0017] The present invention has the following advantages: 1. The present invention integrates spaceborne lidar data and optical remote sensing image data, combines the accuracy advantage of spaceborne lidar data at the point scale with the spatio-temporal continuity advantage of optical remote sensing image data at the surface scale, and realizes high-precision and large-scale forest tree height calculation; 2. The present invention improves the Unet deep learning neural network model, combines a two-dimensional dense connection module and an attention mechanism module, and optimizes the calculation method of the loss function for sparse true values, so that the constructed forest tree height inversion model has a better effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic flow chart of the present invention; Figure 2Schematic diagram of the improved Unet network in the present invention; Figure 3 Structural diagram of 2D DenseBlock in the present invention; Figure 4 Structural diagram of CBAM in the present invention. Specific implementation manners

[0019] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0020] As Figure 1 shown, this embodiment discloses a forest tree height inversion method based on multi-source remote sensing data and the Unet network, including the following steps: (1) Obtain multi-source remote sensing data, including spaceborne lidar data GEDI and optical remote sensing image data Landsat; Download the spaceborne lidar GEDI forest tree height product and the optical remote sensing image Landsat surface reflectance product in the embodiment area, and keep the time ranges of the two consistent; This embodiment downloads the GEDI L2A product and the Landsat8 L2 product between January 2020 and December 2020.

[0021] For the GEDI L2A product, select Rh98 as the height inversion index, which represents the 98% relative height in the energy return process and is usually considered the true value of the tree height. In addition, to ensure the accuracy of the data, a series of indicators are extracted as shown in Table 1, including longitude, latitude, and a series of quality flags.

[0022] Table 1 Extracted parameters of spaceborne lidar GEDI

[0023] For the Landsat8 L2 product, select the six spectral band data of SR_B2~SR_B7 as the main input indicators. At the same time, since the L2 product has completed atmospheric correction, there is no need to further process the image. In addition to the six band data, to ensure the accuracy of the data, the QA_PIXEL band is also extracted, which indicates which pixels may be affected by factors such as instruments or clouds.

[0024] Table 2 Extracted parameters of optical remote sensing images

[0025] (2) Preprocess the acquired multi-source remote sensing data to construct a network training dataset; (2.1) Preprocessing of spaceborne lidar data To ensure the quality of GEDI spaceborne lidar data, the acquired GEDI data is screened using the criteria in Table 3, and the rules are as follows: 1) Exclude low-sensitivity footprints (sensitivity < 0.95); 2) Exclude low-elevation-accuracy footprints (the elevation difference between the GEDI-derived elevation and the SRTM elevation is greater than 30 m, approximately the 75% confidence interval of the elevation difference in this embodiment; 3) Exclude any data with low data quality (degrade > 0, stale_return_flag = 1, degrade_flag = 1, quality_flag = 0); 4) Exclude footprint points that may have various errors (rx_assess_flag = 1) Table 3 Spaceborne lidar GEDI screening conditions

[0026] After the above screening, the spaceborne lidar data for the whole year of 2020 is retained.

[0027] (2.2) Preprocessing of optical remote sensing image data: To ensure the quality of Landsat optical remote sensing data, the acquired Landsat images are screened using the criteria in Table 4, and the rules are as follows: 1) Exclude pixel points where the 3rd bit of QA_PIXEL is equal to 1, which represents possible interference from clouds; 2) Exclude pixel points where the 4th bit of QA_PIXEL is equal to 1, which represents possible interference from cloud shadows; Table 4 Optical remote sensing image screening conditions

[0028] Extract all Landsat-8 images during 2020, mask the pixel points where the 3rd bit of QA_PIXEL is equal to 1 and the 4th bit of QA_PIXEL is equal to 1 to complete the cloud removal of the images; then stack all the images and take the median value as the final pixel value, and synthesize all the images into one image, that is, median value synthesis of annual images.

[0029] (2.3) Matching of spaceborne lidar data and optical remote sensing image data: Considering the limitations of Landsat satellite imagery in band coverage, which only includes three visible light bands, one near-infrared band, and two shortwave infrared bands, and being restricted by the relatively low spatial resolution, such a band range and resolution level may pose significant challenges in direct feature extraction. To overcome these limitations, vegetation indices are introduced as auxiliary information to enrich the content of feature extraction and enhance the robustness of the model. Vegetation indices can comprehensively reflect information such as vegetation cover status, growth conditions, and biomass, providing more environmental background information and physical characteristic indicators for forest tree height inversion.

[0030] The vegetation indices adopted include the Normalized Difference Vegetation Index (NDVI), the Normalized Difference Moisture Index (NDMI), the Normalized Difference Water Index (NDWI), the Atmospherically Resistant Vegetation Index (ARVI), the Enhanced Vegetation Index (EVI), and the Ratio Vegetation Index (RVI). These indices comprehensively utilize different band information in Landsat imagery to reflect vegetation and other surface cover characteristics from multiple perspectives. The specific calculation formulas for the vegetation formulas are as follows: The Normalized Difference Vegetation Index (NDVI) is one of the most commonly used indices to reflect vegetation vitality and biomass, and it evaluates the health status of vegetation through the spectral reflectance difference between the infrared band and the near-infrared band.

[0031] ; The Normalized Difference Moisture Index (NDMI) utilizes the information in the near-infrared band and the shortwave infrared band to monitor the humidity situation in the area.

[0032] ; The Normalized Difference Water Index (NDWI) uses the green light and the near-infrared band to detect water information.

[0033] ; The Atmospherically Resistant Vegetation Index (ARVI) is an index calculated to reflect vegetation thickness and density. For high-density vegetation such as forests, ARVI can better reflect its characteristics.

[0034] ; The Enhanced Vegetation Index (EVI) is an improvement over NDVI, reducing the influence of the atmosphere, background, and soil, and is more suitable for the analysis of high-density vegetation areas.

[0035] ; The Ratio Vegetation Index (RVI) reflects the biomass and growth status of vegetation.

[0036] ; Among them, BLUE, GREEN, RED, NIR, and SWIR represent the reflectance values of the blue band, green band, red band, near-infrared band, and short-wave infrared band respectively, specifically referring to the SR_B2, SR_B3, SR_B4, SR_B5, and SR_B6 values in the aforementioned Landsat8 products.

[0037] By introducing these vegetation indices as auxiliary information for feature extraction, it helps to make up for the deficiencies when directly extracting features from Landsat images. By combining spectral features and vegetation indices, a more abundant and effective feature set is provided for complex land cover types, thereby improving the classification performance and robustness of the entire model. The application of this method not only lies in improving the accuracy of forest tree height inversion, but more importantly, it can provide more scientific and accurate data support for the management and protection of forest resources, laying a solid foundation for achieving sustainable utilization and ecological protection goals.

[0038] (2.4) Matching of spaceborne lidar data and optical remote sensing image data: The spaceborne lidar data GEDI is sparse spot data, with a spatial resolution of 25 meters for the spots and discontinuous spatial parts; the optical remote sensing image Landsat is raster data with a standard raster resolution of 30 meters, and the two cannot be directly aligned.

[0039] First, rasterize the GEDI spaceborne lidar data through the Arcgis10.8 software. At this time, the raster with spots is a valid raster, and its value is the valid value, while the raster without spots is an invalid raster, and its value is the invalid value (NAN).

[0040] Secondly, set the raster resolution to 30 meters, which has a small difference from its own resolution. For the sample point data that cannot be directly mapped to specific raster cells, bilinear interpolation technology is used for resampling.

[0041] Finally, load the Landsat optical remote sensing image data and align it with the rasterized and resampled GEDI spaceborne lidar data. At this time, the completed dataset has 13 bands. The first 12 bands are the Landsat optical remote sensing image data and vegetation index data with a resolution of 30 meters, which are the input data of the model; the 13th band is the GEDI spaceborne lidar data resampled to 30 meters, which is the label of the model.

[0042] (2.5)Production of network training dataset: Cut the above complete 13-band data into multiple data groups according to the size of 256×256 pixels. Eliminate the data groups without valid ground truth to ensure the quality of the dataset for model training.

[0043] Through random sampling, the filtered data is divided into a training set, a test set, and a validation set according to the ratio of 40%:30%:30%. Through the above steps, a total of 3774 data groups are retained, namely 1510 training sets, 1132 test sets, and 1132 validation sets. The first 12 bands are the input data of the model in the following text, and the 13th band is the label of the model.

[0044] (3)Improve the Unet deep learning network model and optimize the model parameters through training data to generate a forest tree height inversion model; (3.1)Based on the Unet basic deep learning network model, use a two-dimensional dense connection module (2D DenseBlock) to replace the conventional convolution module and add an attention mechanism module (CBAM); Unet is a popular convolutional neural network architecture. It is characterized by a symmetric "U" shape structure, including an encoder (downsampling path) and a decoder (upsampling path). The encoder gradually reduces the spatial dimension of the data, while the decoder gradually restores the spatial dimension and details of the data, and at the same time maintains the context information of the original input data through skip connections. This structure makes Unet very suitable for handling pixel-level image prediction tasks.

[0045] In this embodiment, both the encoder and decoder of the Unet structure have 5 layers, as Figure 2 shown. In the encoder (downsampling path), each layer first passes through a two-dimensional dense connection module, and then through an attention mechanism module to double the number of feature channels. Then continue to perform downsampling through a max pooling operation with a stride of 2 to reduce the size of the image and pass it to the next layer to repeat the above operation. The number of feature channels from the first layer to the fifth layer is 12, 32, 64, 128, 256 in sequence, and the image sizes are 256, 128, 64, 32, 16 in sequence.

[0046] In the decoder (upsampling path), the image is first upsampled by bilinear interpolation to expand the image size. Then it is skip-connected with the feature maps of the corresponding layers in the encoder (downsampling path) to achieve feature channel combination. Subsequently, it passes through a 2D dense connection module, and then through an attention mechanism module to reduce the number of feature channels. Then it continues to be upsampled by bilinear interpolation, repeating the above operations. The number of feature channels from the fifth layer to the first layer is 256, 128, 64, 32, 1 in sequence, and the image sizes are 16, 32, 64, 128, 256 in sequence.

[0047] In the original Unet architecture, the encoder and decoder are mainly composed of convolutional layers and pooling layers (encoder part) and upsampling layers (decoder part). To further improve the performance of the model, the standard convolutional layers in Unet are replaced by 2D DenseBlocks, as Figure 3 shown. The dense connection pattern of the 2D DenseBlock promotes the reuse of features, enhances the information flow, helps to solve the vanishing gradient problem, and improves the parameter efficiency of the model at the same time.

[0048] The process of replacing the convolutional layers in Unet with 2D dense connection modules involves introducing 2D DenseBlocks at each level of the encoder and decoder. This replacement enables the network to retain more feature information when performing downsampling and upsampling operations, because the layers within each 2D DenseBlock can directly access the feature maps of the previous layers. In addition, since the output of the 2D DenseBlock is composed of the feature maps of all internal layers, this integrated feature representation provides richer information for forest tree height inversion.

[0049] Each layer in the 2D DenseBlock is connected to all the previous layers. In the 2D DenseBlock, the feature maps from the previous layer are concatenated, allowing for rich information flow.

[0050] ; where refers to the concatenation of the feature maps generated from layer 0 to layer ; H l represents the feature maps from layer l- 1 to layer lThe variation function of the layer feature map includes convolution, batch normalization, and non-linear activation operations. By introducing the 2D DenseBlock, efficient information flow within the network, deep reuse of features, and significant reduction of model parameters are effectively achieved. The design of the 2D DenseBlock allows each layer to directly access the feature maps of all previous layers, which not only promotes unobstructed information flow and maximizes the utilization of features, but also reduces the learning burden of duplicate features, thereby enhancing the training efficiency and generalization ability of the model while reducing the risk of overfitting. In addition, the 2D DenseBlock also achieves a relatively low number of parameters while increasing the network depth, further improving the performance and computational efficiency of the model. These characteristics enable the 2D DenseBlock to exhibit significant advantages in processing complex data and improving classification performance, providing an effective solution framework for optimizing deep learning models.

[0051] The attention mechanism can enable the convolutional neural network to pay attention to important regions or features. CBAM combines channel attention and spatial attention. By sequentially focusing on the channel and spatial dimensions, CBAM effectively enhances the feature representation ability. As Figure 4 shown in (a) of

[0052] 1) Channel attention mechanism The goal of the channel attention mechanism is to identify which channels are more important, that is, which feature channels are more critical for the current task. It aggregates the global spatial information of the input feature map, learns the importance weights of different channels, and uses these weights to enhance the response of important feature channels while suppressing unimportant channels. Specifically, the data processing flow of the channel attention module is as Figure 4 shown in (b) of

[0053] Channel attention weight The calculation formula is: ; where represents the input feature map, represents the sigmoid function, MLP represents a two-layer fully connected neural network, represents the global average pooling operation, represents the global max pooling operation; The specific processing flow is: Assume that the size of the input feature map F is . First, perform the global average pooling and global max pooling operations on the input feature layer respectively to obtain two feature maps.

[0054] Then, these aggregated features are further processed by a shared two-layer fully connected neural network (MLP). The first layer of this network contains C / r neurons (where r is a reduction factor used to reduce the number of parameters and computational complexity), and ReLU is used as the activation function. The number of neurons in the second layer is restored to C, which is the same as the number of original feature channels.

[0055] Finally, the outputs of the two-layer MLP are merged by element-wise addition and activated by a sigmoid function to generate the final channel attention weights. 。

[0056] By performing element-wise multiplication of with the original input feature map F, the adjusted input features are provided for the spatial attention module, enhancing the model's attention to important channels.

[0057] 2) Spatial attention mechanism The goal of the spatial attention mechanism is to identify which spaces are more important, that is, which spaces are more critical for the current task. It aggregates the global feature channel information of the input feature map, learns the importance weights of different spaces, and uses these weights to enhance the response of important spaces while suppressing unimportant spaces. Specifically, the data processing flow of the spatial attention module is as shown in (c) of Figure 4 。

[0058] Spatial attention weights The calculation formula is: ; where F represents the input feature map, represents the sigmoid function, represents the global average pooling operation, represents the global max pooling operation, represents the convolution operation with a 7*7 convolution kernel; The specific processing flow is as follows: First, by performing channel-based global max pooling and global average pooling on F, two feature maps with dimensions of are generated respectively. These two feature maps capture the extreme values and average information of different responses in space, helping the model to focus on different spatial regions. Subsequently, through the concatenation operation in the channel dimension, these two feature maps are merged into a feature map, thus integrating the information of global max pooling and global average pooling.

[0059] Subsequently, the merged feature map passes through a convolutional layer, which not only performs feature fusion but also compresses the dimension, making the output feature map become again. Compared with the conventional convolutional kernel, the convolutional kernel can cover a wider spatial range, enabling the model to better understand the spatial context information, thus enhancing the effect of the spatial attention mechanism.

[0060] By processing the output of the above convolutional layer through the sigmoid activation function, the final spatial attention feature map is obtained. This feature map reflects the attention degree of the model to different spatial positions. Through the element-wise multiplication operation with the input feature map F of the spatial attention module, the final feature output is obtained. This process effectively guides the model to concentrate resources on processing those spatial regions that are more critical to the classification task, thereby improving the discriminative ability and overall performance of the model.

[0061] 3) The role of the attention mechanism in the model In this embodiment, the convolutional block attention module (CBAM) is introduced into the Unet architecture. CBAM works through two main components: the channel attention mechanism and the spatial attention mechanism, which respectively focus on which channels of the input feature map are important and which regions in the feature map are significant. The introduction of this dual attention mechanism enables the network to adaptively focus on the features most relevant to the current task, thereby improving the model's learning ability for complex and sparse data.

[0062] In the sparse ground truth data training stage, CBAM guides the Unet model to automatically discover and preferentially learn the features that are more critical to this task, which is crucial for extracting effective information from limited labeled data. This not only improves the learning efficiency of the model but also enhances the ability to capture subtle ground object features. Compared with the traditional Unet model, the Unet with CBAM can establish a regression formula more accurately and significantly improve the accuracy when dealing with the tree height inversion task with similar spectral features. When the trained model performs per-pixel prediction of a large area, through the role of CBAM, the model can still effectively invert the forest tree height in the face of unknown or complex terrain conditions.

[0063] In the Unet model, the final result is output using a convolutional layer. First, the convolutional operation can provide accurate regression values for each pixel while maintaining the spatial structure information of the image, which is crucial for pixel-level tasks. It allows the model to accurately map the boundaries and features of ground objects, such as the height changes of the terrain, ensuring spatial consistency and the accuracy of the results. Second, the weight sharing mechanism of the convolutional layer significantly reduces the number of model parameters, which not only reduces the risk of overfitting but also improves the computational efficiency and generalization ability of the model. Through the encoder-decoder structure, the Unet model shows great flexibility in processing information at different scales. The use of convolutional layers further promotes the fusion of features at different levels, enabling the model to integrate global and local information and perform more accurate regression for complex ground object types. In addition, directly using the results output by the convolutional layer eliminates the complex steps of mapping high-dimensional features back to the original resolution, simplifies the model structure, and improves the processing speed. Combining the above advantages, the application of convolutional layers in Unet greatly improves the performance of the model in remote sensing image processing, especially in pixel-level regression tasks, meeting the high standards of remote sensing data analysis.

[0064] (3.2) Modify the training function of the deep learning network model so that the model is optimized for valid values.

[0065] For the prediction of forest tree height, the mean absolute error (MAE) is selected as the loss function, and the calculation formula is: ; In the formula, n is the number of true value points, is the model prediction value, is the true value, and the range of the mean absolute error is all , which is equal to 0 when the prediction value exactly matches the true value, that is, perfect prediction; the worse the prediction effect, the larger the above two error values.

[0066] (4) Use the generated forest tree height inversion model to invert the forest tree height and evaluate the inversion accuracy; specifically as follows: (4.1) Forest tree height inversion: Integrate the optical remote sensing data and vegetation index data within the implementation example area, and input the above data into the forest tree height model to perform the forest height inversion in this area.

[0067] (4.2) Inversion accuracy assessment: Compare the model prediction value with the true value to evaluate the inversion accuracy. The root mean square error (RMSE) and the coefficient of determination (R 2 ) are used as the accuracy assessment indicators, and the formulas are as follows: Root mean square error (RMSE): ; Coefficient of determination (R 2): ; where n is the total number of samples, is the model prediction value, is the true value, is the average value of the true values. The range of the root mean square error is , equal to 0 when the predicted value exactly matches the true value, i.e., perfect prediction; the worse the prediction effect, the larger the root mean square error value. The range of the coefficient of determination is , equal to 1 when the predicted value exactly matches the true value, i.e., perfect prediction; the worse the prediction effect, the smaller the coefficient of determination.

[0068] The coefficient of determination is an index to measure the goodness of fit of the model to the data, reflecting to what extent the independent variable can explain the change of the dependent variable. Its value is usually between 0 and 1, and the value closer to 1 indicates stronger explanatory power of the model. The root mean square error is a commonly used index to measure the deviation between the model predicted value and the actual observed value, especially suitable for regression problems. It gives higher punishment to larger errors. The mean absolute error is another index to measure the difference between the predicted value and the actual observed value, and it gives the same weight to all errors.

[0069] In this embodiment, comparative experiments and ablation experiments are set up to prove the effectiveness of the proposed deep learning algorithm. The experimental results are shown in Tables 5 and 6. The improved U-Net proposed in this embodiment has an accuracy improvement of more than 25% compared with the traditional machine learning algorithm, and the proposed two-dimensional dense connection module and attention mechanism module enhance the effectiveness of the algorithm.

[0070] Table 5 Comparative Experiments

[0071] Table 6 Ablation Experiments Provided in this Embodiment

[0072] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A forest tree height inversion method based on multi-source remote sensing data and Unet network, characterized by: The following steps are involved: (1) Acquire multi-source remote sensing data, including space-borne lidar data and optical remote sensing image data; (2) Preprocess the acquired multi-source remote sensing data and construct a network training data set; (3) Improve the Unet deep learning network model and optimize the model parameters through training data to generate a forest tree height inversion model; Based on the Unet basic deep learning network model, the conventional convolution module is replaced by the 2D dense connection module 2D DenseBlock, and the attention mechanism module CBAM is added; (4) Use the generated forest tree height inversion model to invert forest tree height and evaluate the inversion accuracy.

2. The method for inverting forest tree height based on multi-source remote sensing data and Unet network according to claim 1, characterized in that: The expression of the two-dimensional dense connection module 2D DenseBlock is: ; in Refers to the layers from 0 to Concatenation of feature maps generated in layers; H l Indicates l- 1st layer feature map to the l The function of changing the layer feature map, including convolution, batch normalization, and nonlinear activation operations; The output of 2D DenseBlock is composed of the feature maps of all internal layers, and this integrated feature representation provides richer information for forest tree height inversion.

3. The method for inverting forest tree height based on multi-source remote sensing data and Unet network according to claim 1, characterized in that: The attention mechanism module CBAM includes the channel attention mechanism CAM and the spatial attention mechanism SAM: Channel attention weight and spatial attention weights The calculation formulas are: ; ; in, represents the input feature map, represents the sigmoid function, MLP represents a two-layer fully connected neural network, represents the global average pooling operation, represents the global maximum pooling operation, The convolution kernel is 7*7, the channel attention mechanism is performed on the spatial dimension, and the spatial attention mechanism is performed on the channel dimension; Feature map output after CBAM , so that the Unet deep learning network model focuses on the features most relevant to the forest tree height inversion task, so as to improve the learning ability of the forest tree height inversion model for sparse data.

4. The method for inverting forest tree height based on multi-source remote sensing data and Unet network according to claim 1, characterized in that: The loss function of the training data in step (3) is as follows: For the prediction of forest tree height, the mean absolute error MAE is selected as the loss function, and the calculation formula is: ; In the formula, n is the number of true value points, is the model prediction value, is the true value, and the range of mean absolute error is , when the predicted value is completely consistent with the true value, it is equal to 0, that is, perfect prediction; the worse the prediction effect, the greater the error between the predicted value and the true value.

5. The method for inverting forest tree height based on multi-source remote sensing data and Unet network according to claim 1, characterized in that: Step (4) includes the following steps: (4.1) Forest tree height inversion: Integrate the optical remote sensing data and vegetation index data in the study area, input the optical remote sensing data and vegetation index data into the forest tree height inversion model, and perform forest height inversion in the area; (4.2) Inversion accuracy assessment: The model prediction value is compared with the true value to assess the inversion accuracy; the root mean square error (RMSE) and the coefficient of determination (R) are used to 2 As an accuracy evaluation indicator, the formula is as follows: Root mean square error RMSE: ; Coefficient of determination R 2 : ; In the formula, n is the total number of samples, is the model prediction value, is the true value, is the average value of the true value; the range of the root mean square error is , when the predicted value is completely consistent with the true value, it is equal to 0, that is, perfect prediction; the worse the prediction effect, the larger the root mean square error value; the range of the determination coefficient is , when the predicted value is completely consistent with the true value, it is equal to 1, that is, a perfect prediction; the worse the prediction effect, the smaller the determination coefficient.

6. A forest tree height inversion system based on multi-source remote sensing data and Unet network, characterized by: Used to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Power transmission line channel tree height inversion method based on laser radar and optical remote sensing

    CN111414891A

  • Large-scale forest canopy height inversion method based on ICESat-2 and multi-source remote sensing data

    CN119251697A

  • Large-scale forest height remote sensing retrieval method considering ecological zoning

    US20230213337A1

Cited By

  • Tree height and biomass collaborative inversion method, system and device based on codec double focusing and medium

    CN120522692A

  • Forest stand boundary extraction method and system

    CN120931948A

  • A method and system for extracting forest stand boundaries

    CN120931948B

  • Scanning microwave radiometer precipitation inversion method based on weight attention mechanism

    CN121305393A

  • Multi-scale-based space-time double-flow network and tree height parameter prediction method

    CN121706062A