Tree crown segmentation method

By combining the hybrid encoder structure of RGB images and CHM images, the problem of difficult identification of crown boundaries in complex forest environments is solved, high-precision crown segmentation is achieved, and the robustness and segmentation accuracy of the crown segmentation model are improved.

CN120599249APending Publication Date: 2025-09-05NANJING FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510635515.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional crown segmentation methods have difficulty accurately identifying crown boundaries in complex forest environments, and the segmentation accuracy is low. Especially in the Northeast cold-temperate coniferous forest and temperate mixed forest areas, where crown overlap is severe, canopy morphology is diverse, and understory vegetation is complex, existing technologies find it difficult to effectively integrate multi-source remote sensing data to improve segmentation accuracy.

Method used

A tree crown segmentation model with a hybrid encoder structure is proposed, which combines RGB images and CHM images, performs feature extraction and fusion through CNN modules and Transformer modules, and uses a two-branch convolutional neural network to extract texture and height features respectively. The shallow feature fusion module and the improved ViT-Adapter framework are used to enhance feature interaction, and the lightweight convolutional attention module is used to improve segmentation accuracy.

Benefits of technology

It effectively improves the robustness of crown segmentation under complex backgrounds, reduces the mis-segmentation rate, enhances the edge and morphological representation of the crown, improves the crown segmentation model's ability to perceive spatial information at different scales, and achieves high-precision crown segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599249A_ABST
    Figure CN120599249A_ABST
Patent Text Reader

Abstract

The invention provides a crown segmentation method, and belongs to the technical field of remote sensing image tree parameter information extraction. Comprising the steps that RGB images and CHM images of a tree area are collected, tree crown labeling is carried out, a tree crown segmentation data set is formed, division is carried out according to the proportion, and a training set, a verification set and a test set are generated; a crown segmentation model is constructed, training is carried out, and model parameters are optimized; and obtaining an RGB image and a CHM image of the target tree area, and performing crown segmentation on the target tree area by using the trained crown segmentation model. According to the method, the RGB image and the CHM image of the tree are processed through the tree crown segmentation model, multi-modal data fusion is achieved, the RGB image has rich texture and color features, the CHM image supplements the defect of the RGB image in the aspect of vertical structure information, the robustness of tree crown segmentation under the complex background is effectively improved through fusion of the RGB image and the CHM image, and the occurrence rate of wrong segmentation is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a tree crown segmentation method and belongs to the technical field of remote sensing image tree parameter information extraction. Background Art

[0002] Traditional methods for tree crown segmentation often rely on a single remote sensing data source and often struggle to cope with complex forest environments. While optical images offer good resolution and visual intuitiveness, they are easily affected by lighting variations, shadow interference, and background elements such as snow, making it difficult to distinguish between tree crowns and ground cover. While LiDAR point clouds can provide accurate three-dimensional structural information, they are limited by factors such as high acquisition costs and complex processing procedures, making them difficult to apply widely across large forest areas. Particularly in the cold-temperate coniferous forests and temperate mixed forests of Northeast my country, the severe overlap of tree crowns, diverse canopy morphology, and complex understory vegetation further complicate the identification of tree crown boundaries. Therefore, how to fuse multi-source remote sensing data and deeply explore the feature correlations between them has become a key issue in improving the accuracy of tree crown segmentation. Summary of the Invention

[0003] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a crown segmentation method that solves the problems of difficult crown boundary identification and low crown segmentation accuracy in complex forest environments.

[0004] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0005] The present invention provides a crown segmentation method, comprising: collecting RGB images and CHM images of a tree area, annotating the crowns, and forming a crown segmentation data set; dividing the crown segmentation data set proportionally to generate a training set, a validation set, and a test set; constructing a crown segmentation model, and training the crown segmentation model using the training set and the validation set, optimizing model parameters, and testing the crown segmentation model using the test set; obtaining RGB images and CHM images of a target tree area, and performing crown segmentation on the target tree area using the trained crown segmentation model.

[0006] Furthermore, the crown segmentation model includes: a hybrid encoder and decoder composed of a CNN module and a Transformer module; the CNN module adopts a two-branch convolutional neural network; its first branch is used to extract features from RGB images to obtain a texture feature map, and the second branch is used to extract features from CHM images to obtain a height feature map.

[0007] Furthermore, the first branch and the second branch both include: convolutional layers, ; The first branch The convolutional layer outputs Level RGB shallow features; the first The convolutional layer outputs Level CHM shallow features, among which Level RGB shallow features and level Level CHM shallow features are shallow features of the same scale; ; The hybrid encoder further includes: a shallow feature fusion module; the shallow feature fusion module is used to Level RGB shallow features and level The shallow features of the first-level CHM are fused to obtain the Level shallow fusion features.

[0008] Furthermore, the process of feature fusion performed by the shallow feature fusion module includes: Level RGB shallow features and level The shallow features of the first-level CHM are subjected to global average pooling, residual block processing, convolution operation, ReLU activation function and Sigmoid activation function respectively to obtain the first Level RGB intermediate features and level Level CHM intermediate features; the first Level RGB intermediate features and level The CHM intermediate features are weighted and element-wise added to obtain the Level shallow fusion features.

[0009] Furthermore, the process of obtaining the height feature map includes: the CHM image is sequentially passed through the second branch The convolutional layers are downsampled to obtain the first feature; the first feature and the The shallow fusion features are weighted fused to obtain a high-level feature map.

[0010] Further, The shallow fusion features of the first layer The shallow fusion features of each level are combined with the feature maps of the same resolution in the decoder upsampling path through skip connections.

[0011] Furthermore, the Transformer module adopts the improved ViT-Adapter framework, which has the following improvements over the ViT-Adapter framework: tokenizing the texture feature map and the height feature map and adding position embedding to obtain the patch sequence and After that, the multi-head attention mechanism is introduced to and Perform update calculations to obtain the updated patch sequence and , the update calculation process includes: ; ; ; ; ; Where, Respectively The query weight matrix, key weight matrix and value weight matrix of each head; Respectively The query matrix, key matrix and value matrix of each head; is the dimension of the key vector, is the self-attention function; is the transpose, is the activation function; is the multi-head self-attention function; For the Output of the head; is the number of heads; For splicing operation; is the output weight matrix; They are and Updated sequence representation.

[0012] Furthermore, the improvements of the improved ViT-Adapter framework over ViT-Adapter include:

[0013] The updated patch sequence and Splice them into input features; use the RGB shallow features output by any convolutional layer in the first branch as spatial prior features, input them into the spatial feature injection module of the improved ViT-Adapter framework, and inject the spatial prior features into the input features through the cross-attention mechanism.

[0014] Furthermore, a lightweight convolutional attention module is added to the decoder to enhance the performance of convolutional neural networks in feature learning by introducing channel attention and spatial attention.

[0015] Compared with the prior art, the present invention has the following beneficial effects:

[0016] (1) The crown segmentation method provided by the present invention uses a crown segmentation model to process the RGB image and CHM image of the tree to achieve multimodal data fusion. The RGB image has rich texture and color features, and the CHM image supplements the deficiency of the RGB image in vertical structure information. The feature fusion of the two effectively improves the robustness of crown segmentation under complex backgrounds and reduces the incidence of missegmentation.

[0017] (2) The crown segmentation model provided by the present invention adopts a hybrid encoder composed of a CNN module and a Transformer module. The CNN module uses a two-branch convolutional neural network to extract features from RGB images and CHM images respectively to obtain texture feature maps and height feature maps. The shallow feature fusion module also performs weighted fusion of the RGB shallow features and CHM shallow features of the same scale output by the corresponding convolution layers in the first branch and the second branch to enhance the crown edge and morphological representation, so as to solve the problem of difficult recognition of crown boundaries in complex environments.

[0018] (3) The crown segmentation model provided by the present invention adopts the improved ViT-Adapter framework that introduces the multi-head attention mechanism as the Transformer module of the hybrid encoder; through multi-head attention, the deep feature interaction between texture feature map patches and height feature map patches is captured, breaking through the barriers of modal heterogeneity and enhancing the crown segmentation model's ability to perceive spatial information at different scales. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the structure of the tree crown segmentation model in Example 1 of the present invention;

[0020] Figure 2 Schematic diagram of the structure of the shallow feature fusion module in Example 1 of the present invention;

[0021] Figure 3 This is a heat map of the tree crown segmentation accuracy of different models in Example 3 of the present invention;

[0022] Figure 4 This is a comparison chart of tree crown segmentation results of different models in a sparse forest in Example 3 of the present invention;

[0023] Figure 5 This is a comparison chart of the tree crown segmentation results of different models under dense forest in Example 3 of the present invention. DETAILED DESCRIPTION

[0024] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices. The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described are only some embodiments of the present application, not all embodiments.

[0025] Example 1

[0026] This embodiment provides a tree crown segmentation method. By acquiring RGB images and CHM images of the target tree area and performing multimodal data fusion, the method deeply explores the deep connection between the texture and color features extracted from the RGB images and the vertical height information extracted from the CHM images to improve the accuracy of tree crown segmentation in complex forest environments. The method includes:

[0027] Step 1: Collect RGB images and CHM images of the tree area, perform tree crown annotation, and form a tree crown segmentation dataset.

[0028] Specifically, the crown segmentation dataset includes: RGB images and CHM images of the tree area, as well as the crown segmentation results of the area obtained by crown annotation. The RGB image is obtained by extracting the RGB value of the image, and the CHM image is obtained by the canopy height model, which is obtained by subtracting the digital elevation model from the digital surface model.

[0029] In some specific embodiments, an edge-energy function automated dataset annotation algorithm is used for tree crown annotation, which not only retains the pixel information of the original image, but also retains the boundary information of most tree crowns compared to direct cropping, making the morphology of most tree crowns complete.

[0030] Step 2: Divide the crown segmentation dataset proportionally to generate training set, validation set, and test set.

[0031] In some specific embodiments, the ratio of the training set, the validation set, and the test set is 6:2:2.

[0032] Step 3: Construct a crown segmentation model, train the crown segmentation model using a training set and a validation set, optimize model parameters, and test the crown segmentation model using a test set;

[0033] refer to Figure 1The tree crown segmentation model provided in this embodiment is based on the TransUNet structure, including: a hybrid encoder and decoder composed of a CNN module and a Transformer module.

[0034] As is known, TransUNet is a Transformer-based segmentation network derived from the U-Net model. It introduces a hybrid encoder with a hybrid CNN-Transformer architecture based on the U-Net model. By combining CNN and Transformer, TransUNet addresses the limitations of traditional convolutional neural networks in modeling long-range dependencies and processing large images. The decoder of TransUNet uses a cascaded upsampling strategy to upsample the encoded features and combine them with high-resolution CNN feature maps to achieve precise positioning. Therefore, TransUNet not only preserves detailed features but also leverages global contextual information to improve segmentation accuracy.

[0035] In particular, in this embodiment, in order to use CNN to simultaneously extract features of RGB images and CHM images and generate corresponding texture feature maps and height feature maps, the CNN module adopts a two-branch convolutional neural network; its first branch is used to extract features of RGB images to obtain texture feature maps, and its second branch is used to extract features of CHM images to obtain height feature maps.

[0036] In some embodiments, the hybrid encoder further includes: a shallow feature fusion module, which is used to fuse shallow features of the same scale output by each convolutional layer in the first branch and the second branch to obtain shallow fusion features of the corresponding scale.

[0037] Specifically, the first branch and the second branch of the dual-branch convolutional neural network both include: convolutional layers, ; The first branch The convolutional layer outputs Level RGB shallow features; the first The convolutional layer outputs Level CHM shallow features, .

[0038] It should be noted that the Level RGB shallow features and level The shallow features of the CHM level are shallow features of the same scale, that is, the resolution of the two is the same.

[0039] No. Level RGB shallow features and level The shallow features of the first-level CHM are fused through the shallow feature fusion module to obtain the first Level shallow fusion features.

[0040] In some specific embodiments, each branch of the hybrid encoder consists of four convolutional layers for extracting multi-scale features, and the convolution kernel size of each layer is 3×3 and the stride is 2.

[0041] refer to Figure 2 In some more specific embodiments, the process of feature fusion by the shallow feature fusion module includes: Level RGB shallow features and level The shallow features of the first-level CHM are globally averaged pooled to aggregate global information; then, they are processed using the Res-block residual block, and then two convolution operations with a kernel size of 1×1 are used for feature compression and activation processing, followed by applying the ReLU activation function and the Sigmoid activation function respectively to obtain the first Level RGB intermediate features and level Finally, the first Level RGB intermediate features and level The CHM intermediate features are weighted and element-wise added to obtain the Level shallow fusion features.

[0042] In some embodiments, The shallow fusion features of the first layer The shallow fusion features of the first layer are combined with the feature maps of the same resolution in the upsampling path in the decoder through skip connections. The shallow fusion features of the first level are used to fuse with the features obtained by downsampling in the second branch to obtain a height feature map.

[0043] That is to say, the process of obtaining the height feature map includes: the CHM image passes through the second branch in turn The convolutional layers are downsampled to obtain the first feature; the first feature and the The shallow fusion features are weighted fused to obtain a high-level feature map.

[0044] The texture feature map and height feature map obtained by the dual-branch convolutional neural network are input into the Transformer module for encoding to obtain the encoded features.

[0045] Next, the decoder upsamples the encoded features to the original image size and generates pixel-level segmentation results. In the upsampling path, the encoded features are combined with high-resolution shallow fusion features to enrich semantic information and achieve more accurate positioning.

[0046] In some embodiments, the crown segmentation model is trained using an AdamW optimizer; the loss function is a weighted combination of cross entropy and Dice loss.

[0047] Step 4: Obtain the RGB image and CHM image of the target tree area, and use the trained crown segmentation model to perform crown segmentation on the target tree area.

[0048] In the crown segmentation method provided in this embodiment, a crown segmentation model is constructed by improving the TransUNet structure. On the one hand, a two-branch convolutional neural network is used to extract the feature information of RGB images and CHM images respectively. On the other hand, a shallow feature fusion module is used to fuse the multi-scale shallow features of the RGB images and CHM images, and the obtained shallow fusion features are jump-connected to the upsampling path of the decoder for feature combination to maintain the local details and contextual information of the image.

[0049] Example 2

[0050] Based on Example 1, this embodiment provides a more comprehensive hybrid encoder structure.

[0051] As is well known, the Vision Transformer (ViT) introduced the Transformer architecture from natural language processing to computer vision for image processing. ViT extracts and classifies image features by segmenting the image into a series of small blocks and treating them as serialized tokens that are input into the Transformer encoder. However, ViT's performance is not ideal when handling dense prediction tasks such as object detection, instance segmentation, and semantic segmentation. Therefore, ViT-Adapter was proposed. The core idea of ​​ViT-Adapter is to compensate for ViT's shortcomings by introducing inductive bias.

[0052] Inductive bias is a kind of prior knowledge introduced during the training process that helps the model better adapt to specific tasks.

[0053] In ViT-Adapter, this inductive bias is achieved through specialized adapters, which are additional architectures that can introduce prior information about the data and task into the model.

[0054] In this embodiment, the Transformer module of the hybrid encoder adopts the improved ViT-Adapter. The improvements of ViT-Adapter include: tokenizing the texture feature map and the height feature map and adding position embedding to obtain the patch sequence. and After that, the multi-head attention mechanism is introduced to and Perform update calculations.

[0055] In the multi-head self-attention mechanism, by computing the input patch sequence and The attention relationship between them is used to capture the dependency between them. First, the weight matrix transformed and Defined as query Q, key K and value V; Next, the scaled dot product attention is calculated, and multiple attention heads are used to capture feature interactions at different levels. The output of each attention head is concatenated and linearly transformed using an output weight matrix to obtain the updated patch sequence and In this process, the multi-head self-attention layer updates the input patch sequence, preserving the spatial relationship and cross-modal interaction information between the two patch sequences.

[0056] Specifically, the update calculation process includes: ; ; ; ; ; Where, Respectively The query weight matrix, key weight matrix and value weight matrix of each head; Respectively The query matrix, key matrix and value matrix of each head; is the dimension of the key vector, is the self-attention function; is the transpose, is the activation function; is the multi-head self-attention function; For the Output of the head; is the number of heads; For splicing operation; is the output weight matrix, which is used to project the spliced ​​multi-head output into the final feature space; They are and Updated sequence representation.

[0057] In some embodiments, the improved ViT-Adapter also includes the following improvements relative to the ViT-Adapter: and Splice them into input features; use the shallow features of the RGB image extracted by any convolutional layer in the first branch as spatial prior features, input them into the spatial feature injection module of the improved ViT-Adapter, inject the spatial prior features into the input features through the cross-attention mechanism, and then input them into the ViT model.

[0058] In some specific embodiments, reference Figure 1 , the shallow features output by the third convolution layer in the first branch are used as spatial prior features.

[0059] In some embodiments, a lightweight Convolutional Block Attention Module (CBAM) is incorporated into the decoder to enhance the performance of convolutional neural networks in feature learning by introducing channel attention and spatial attention. This helps the network automatically focus on more important feature regions in the image, thereby improving the model's performance in tasks, particularly object segmentation.

[0060] The CBAM module consists of two parts: channel attention and spatial attention. The role of the channel attention mechanism is to enhance the learning of useful features by focusing on the importance of each channel. Specifically, it generates two descriptors through global average pooling and global maximum pooling, and fuses them through the fully connected layer. Finally, it outputs the weight of each channel through the Sigmoid activation function, which enables the model to automatically determine which channels are more important for the current task; the spatial attention mechanism helps the network identify the most important areas in the image by focusing on the features of specific spatial positions. It performs a pooling operation on each channel to generate a two-dimensional feature map, and then extracts spatial features through convolution operations to enhance attention to the target area.

[0061] In tree crown segmentation tasks, especially when using remote sensing data or drone imagery, tree crowns are often complex and difficult to distinguish from the background. The channel and spatial attention mechanisms provided by CBAM significantly improve the decoder's ability to recover details and focus on target areas. This enables the model to more accurately segment tree crown areas when processing images with complex backgrounds, irregular tree crown shapes, or dense tree crowns, thereby improving the accuracy of segmentation results.

[0062] Example 3

[0063] This embodiment provides a comparative experiment to verify the applicability and segmentation accuracy of the method of the present invention for tree crown segmentation.

[0064] In this experiment, 50 2560×2560 automatically annotated data sets were selected and cropped into 2375 512×512 data sets with complete boundary information. The experimental set and the test set were allocated in an 8:2 ratio.

[0065] The configuration of the experimental equipment used in this experiment is shown in Table 1, and the parameter settings are shown in Table 2.

[0066] surface Experimental equipment configuration table

[0067] Configuration Name Configuring specific values CPU Intel i5-13600KF Graphics card NVIDIA GeForce RTX 3090 Video Memory 24GB Experimental Platform Ubuntu 20.04 Deep Learning Framework Pytorch 2.0.1 Python version 3.9 CUDA version 11.3

[0068] surface Experimental parameter table

[0069] Parameter name Configuring specific values Optimizer AdamW Initial learning rate 0.0002 Attenuation coefficient 0.05 Batch size 8 momentum 0.9 Training rounds 120

[0070] All experiments were implemented based on the PyTorch framework on an NVIDIA GeForce RTX 3090. The models were trained using the AdamW optimizer with an initial learning rate of 0.0002, a decay coefficient of 0.05, a batch size of 8, and a momentum of 0.9. After collecting samples in the sliding window, random rotation and flipping were applied.

[0071] In the constructed crown segmentation model, the first and second branches of the dual-branch convolutional neural network are composed of two ResNet50 models, each of which consists of 4 stages.

[0072] The ViT-Adapter module contains 12 Transformer layers in total. The channel size is set to =512.

[0073] In the comparative experiments, all models maintained the same number of iterations, and no pre-trained models were introduced. The remaining hyperparameters were set according to the default configuration of each model.

[0074] In order to verify the performance of the crown segmentation model, this experiment uses two indicators: the average Accuracy of all classes (mAcc) and the average IoU of all classes (mIoU) to evaluate the results.

[0075] mAcc is the most basic metric for measuring model prediction accuracy. It calculates the proportion of pixels correctly predicted by the model to all predicted pixels. In the tree crown segmentation task, accuracy reflects the overall degree to which the model can correctly distinguish between the tree crown and the background.

[0076] mIoU is a more rigorous indicator for evaluating model accuracy in semantic segmentation tasks. It evaluates the segmentation performance of the model by calculating the overlap between the model predicted area and the true area. In crown segmentation, mIoU measures the ratio of the intersection to the union between the model predicted crown area and the true crown area. The value of mIoU ranges from 0 to 1. The closer the value is to 1, the better the model performs in crown segmentation.

[0077] This experiment was conducted in two stand structures, sparse forest and dense forest, to verify the crown segmentation effects of different comparison models. The selected comparison models include: the U-Net model, Mask-RCNN model and DeepLab model using the traditional CNN network, and the HRNet model, TransUNet model and Mask2Former model using the combination of CNN and Transformer. The results of the comparison experiment are shown in Table 3, among which Ours is the model provided by the present invention.

[0078]

[0079] In the table, the best metrics are in bold. As can be seen, among CNN-based models, U-Net, a classic segmentation model, performs favorably on both sparse and dense forests. In particular, in sparse forests, it achieves a mAcc of 70.53% and a mIoU of 68.18%. However, compared with more advanced models in recent years, U-Net performs slightly worse on dense forest segmentation tasks.

[0080] In comparison, models like Mask-RCNN and DeepLab demonstrate superior segmentation and recognition of complex objects, particularly in scenes with highly overlapping trees and complex structures. Mask-RCNN achieves mAcc and mIoU of 66.03% and 63.31% in dense forests, respectively, while DeepLab outperforms with mAcc of 67.18% and mIoU of 64.74%.

[0081] This comparison shows that when faced with scenes with complex stand structures and blurred canopy boundaries, models that combine stronger feature expression and context perception capabilities have obvious advantages.

[0082] The introduction of Transformer enhances the model's ability to capture global context and long-range dependencies, thereby effectively improving the accuracy of image segmentation.

[0083] In the tree crown segmentation task of sparse forests and dense forests, both HRNet and TransUNet showed obvious performance advantages, especially TransUNet, which achieved 84.26% mAcc and 81.45% mIoU in sparse forest environments, and 75.93% mAcc and 72.69% mIoU in dense forests. This shows that after integrating the Transformer and U-Net architectures, not only the segmentation effect is improved, but also the ability to cope with complex forest stand structures is possessed.

[0084] HRNet also demonstrated good adaptability, particularly in the dense forest segmentation task, achieving mAcc and mIoU of 74.75% and 71.25%, respectively. Mask2Former achieved the best overall performance among all compared models, particularly adept at handling dense forest areas, achieving 79.57% mAcc and 76.51% mIoU in this scenario. This model's focus on instance segmentation and detail refinement enables more accurate segmentation results in complex environments, with particularly strong performance in dense areas, significantly outperforming other compared models.

[0085] The model proposed in this paper outperforms Mask2Former, particularly in dense forest segmentation tasks, achieving mAcc and mIoU of 80.63% and 77.12%, respectively. Although slightly inferior to Mask2Former in sparse forest tasks, the model generally demonstrates strong generalization and practical value across different forest stand types, providing an effective solution for high-precision tree crown segmentation in complex forest environments.

[0086] Experimental results show that with the continuous optimization of the network structure, especially after the introduction of the Transformer module, the segmentation accuracy of the model has been significantly improved. Mask2Former and TransUNet performed well in complex scenes, especially in the segmentation of dense forest areas, demonstrating strong feature modeling and regional recognition capabilities. The model provided by the present invention achieved high segmentation accuracy under different forest stand structures, not only possessing good adaptability but also demonstrating excellent overall performance, further verifying its potential and reliability in practical applications.

[0087] refer to Figure 3 The color change in the figure can intuitively show the quality of the segmentation effect: the redder the color, the better the segmentation performance, and the bluer the color, the weaker the model performance.

[0088] In the sparse forest scenario, the heatmaps for U-Net and Mask-RCNN are generally lighter, indicating relatively low mAcc and mIoU. Starting with DeepLab, however, the colors shift significantly toward red, indicating improved model performance. Mask2Former and the model proposed by this invention perform particularly well in the sparse forest scenario, with the deep red areas clearly highlighting their excellent segmentation accuracy.

[0089] In dense forest environments, the overall heat map color is blue, and the performance of traditional CNN models is particularly biased towards dark blue, reflecting their limitations in processing complex canopy structures. However, as the models gradually evolved from HRNet and TransUNet to Mask2Former and the model provided by the present invention, the color in the figure gradually transitioned from dark blue to light red, clearly demonstrating the steady improvement in segmentation performance. In particular, the light-colored area of ​​the model provided by the present invention has significantly expanded, further highlighting its strong adaptability and stability in complex forest stand structures.

[0090] The color changes in this heatmap not only vividly illustrate the evolution of model performance, but also directly reflect the trend from traditional CNN to a fusion of CNN and Transformer architectures. This fusion strategy not only improves segmentation accuracy but also enhances the model's expressiveness and robustness in complex scenarios such as dense forests.

[0091] The U-Net, Mask-RCNN, and DeepLab used in this experiment only rely on single drone image data for tree crown segmentation. Although they are effective in some scenarios, their segmentation accuracy is significantly limited in complex environments, especially in sparse forest areas with overlapping snow cover.

[0092] refer to Figure 4-5 Against a snowy background, the texture and color differences between trees and the ground in the image are weakened, making it easy for the model to misidentify trees as ground, significantly reducing the accuracy and robustness of segmentation. In contrast, HRNet, TransUNet, Mask2Former, and the model proposed in this paper integrate more diverse data sources, making comprehensive use of drone imagery and CHM generated from point clouds.

[0093] This demonstrates that the multivariate data fusion strategy not only effectively addresses the difficulty in distinguishing between snow and trees, but also leverages height information to precisely locate the tree crown region, significantly improving segmentation. Among these improved models, the one provided by this invention excels in processing crown boundary details, enabling more precise depiction of crown contours and providing high-quality data support for subsequent crown feature extraction and biomass inversion.

[0094] Example 4

[0095] In multimodal data processing and feature fusion, the selection and combination of modules have a crucial impact on model performance. In order to analyze the contribution of each module to model performance, this embodiment provides an ablation experiment to explore the individual role of each module in the crown segmentation model provided by the present invention and its combined effect.

[0096] Table 4-4 shows the results of the ablation experiment, where SFF is the shallow feature fusion module. Indicates that the corresponding module has been added.

[0097] surface Ablation experiment results

[0098] Experimental results show that the introduction of the shallow feature fusion module (SFF) significantly improves model performance. When added alone, mAcc increases from 62.75% to 70.53%, and mIoU increases from 60.02% to 68.74%, demonstrating the positive role of the SFF module in enhancing shallow feature fusion.

[0099] ViT-Adapter is a multimodal deep feature fusion module that further strengthens the deep information interaction between different modal data. Experiments show that after introducing this module, mAcc increased to 75.95% and mIoU increased by 3.87%, verifying its effectiveness in improving the model's cross-modal understanding capabilities.

[0100] In addition, after integrating the lightweight attention mechanism CBAM into the encoder, the model performance is further improved, with mAcc and mIoU reaching 80.63% and 77.12% respectively, indicating that CBAM guides the model to focus more on key areas by adaptively assigning importance to different feature maps, thereby improving the overall performance.

[0101] Overall, the synergy of SFF, ViT-Adapter, and CBAM provides the model with efficient and comprehensive feature fusion capabilities, significantly improving its accuracy and segmentation performance in multimodal remote sensing data processing.

[0102] Ablation experiments show that the three modules, SFF, ViT-Adapter and CBAM, each play an important role in feature fusion at different levels, and their combination can effectively improve the performance of the Vit-AUNet model, thereby achieving the best segmentation effect.

[0103] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A tree crown segmentation method, characterized in that: include: Collect RGB images and CHM images of tree areas, perform tree crown annotation, and form a tree crown segmentation dataset; Divide the crown segmentation dataset proportionally to generate training set, validation set and test set; Constructing a crown segmentation model, training the crown segmentation model using a training set and a validation set, optimizing model parameters, and testing the crown segmentation model using a test set; Obtain the RGB image and CHM image of the target tree area, and use the trained crown segmentation model to perform crown segmentation on the target tree area.

2. The tree crown segmentation method according to claim 1, characterized in that: The tree crown segmentation model includes: a hybrid encoder and a decoder composed of a CNN module and a Transformer module; The CNN module adopts a two-branch convolutional neural network; the first branch is used to extract features from RGB images to obtain a texture feature map, and the second branch is used to extract features from CHM images to obtain a height feature map.

3. The tree crown segmentation method according to claim 2, characterized in that: The first branch and the second branch both include: convolutional layers, where ; The first branch The convolutional layer outputs Level RGB shallow features; the first The convolutional layer outputs Level CHM shallow features, among which Level RGB shallow features and level Level CHM shallow features are shallow features of the same scale; ; The hybrid encoder further includes: a shallow feature fusion module; The shallow feature fusion module is used to Level RGB shallow features and level The shallow features of the first-level CHM are fused to obtain the Level shallow fusion features.

4. The tree crown segmentation method according to claim 3, characterized in that: The process of feature fusion performed by the shallow feature fusion module includes: The first Level RGB shallow features and level The shallow features of the first-level CHM are subjected to global average pooling, residual block processing, convolution operation, ReLU activation function and Sigmoid activation function respectively to obtain the first Level RGB intermediate features and level Level CHM intermediate features; The said Level RGB intermediate features and level The CHM intermediate features are weighted and element-wise added to obtain the Level shallow fusion features.

5. The tree crown segmentation method according to claim 3, characterized in that: The process of obtaining the height feature map includes: the CHM image is sequentially passed through the second branch The convolutional layers are downsampled to obtain the first feature; the first feature and the The shallow fusion features are weighted fused to obtain a high-level feature map.

6. The tree crown segmentation method according to claim 5, characterized in that: No. The shallow fusion features of the first layer The shallow fusion features of the first and second levels are combined with the feature maps of the same resolution in the decoder upsampling path through skip connections.

7. The tree crown segmentation method according to claim 2, characterized in that: The Transformer module adopts the improved ViT-Adapter framework. Its improvements over the ViT-Adapter framework include: tokenizing the texture feature map and the height feature map and adding position embedding to obtain the patch sequence and After that, the multi-head attention mechanism is introduced to and Perform update calculations to obtain the updated patch sequence and , the update calculation process includes: ; ; ; ; ; Where, Respectively The query weight matrix, key weight matrix and value weight matrix of each head; Respectively The query matrix, key matrix and value matrix of each head; is the dimension of the key vector, is the self-attention function; is the transpose, is the activation function; is the multi-head self-attention function; For the Output of the head; is the number of heads; For splicing operation; is the output weight matrix; They are and Updated sequence representation.

8. The tree crown segmentation method according to claim 7, characterized in that: The improvements of the improved ViT-Adapter framework over ViT-Adapter also include: The updated patch sequence and Splice into input features; The RGB shallow features output by any convolutional layer in the first branch are used as spatial prior features and input into the spatial feature injection module of the improved ViT-Adapter framework, and the spatial prior features are injected into the input features through the cross-attention mechanism.

9. The tree crown segmentation method according to claim 2, characterized in that: A lightweight convolutional attention module is added to the decoder to enhance the performance of convolutional neural networks in feature learning by introducing channel attention and spatial attention.

10. The tree crown segmentation method according to claim 1, characterized in that: The crown segmentation model was trained using the AdamW optimizer; the loss function was a weighted combination of cross entropy and Dice loss.