Single tree segmentation and tree species identification method based on multi-source multi-scale mask network

By fusing features from hyperspectral images and LiDAR data using a multi-source, multi-scale mask network, the time-consuming and inconsistent problems of single-tree segmentation and tree species identification in existing technologies are solved, achieving higher accuracy in single-tree segmentation and tree species identification.

CN121170565APending Publication Date: 2025-12-19SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510276976.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing technologies for single-tree segmentation and tree species identification are characterized by being time-consuming, labor-intensive, highly inconsistent, sensitive to noise, and difficult to fully integrate multi-source remote sensing data, resulting in poor segmentation performance and insufficient generalization ability.

Method used

A multi-source, multi-scale mask network is employed, which integrates features from hyperspectral images, high-resolution RGB images, and LiDAR data through a deformable fusion attention and query-constrained decoder. This enables multi-scale feature learning and interaction, generating accurate single-tree masks and tree species labels.

Benefits of technology

It improves the accuracy of single tree segmentation and tree species identification, reduces the computational cost and complexity of the attention mechanism, and enhances the robustness and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170565A_ABST
    Figure CN121170565A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a single tree segmentation and tree species identification method based on a multi-source multi-scale mask network, and the method comprises the steps: carrying out the effective fusion of a hyperspectral image, a high-resolution RGB image and a LiDAR data feature through a designed multi-source multi-scale fusion encoder; due to the designed deformable fusion attention, the complexity and the calculation cost of an attention mechanism are greatly reduced, and multi-source and multi-scale feature learning of each feature point is realized; according to the query constraint decoder, spatial features of a tree data set are explicitly embedded in the initialization process of object query through a query constraint module, the learning space of object query is constrained, and the convergence process of query is refined, so that the individual tree segmentation precision is improved. Therefore, the single tree segmentation and tree species identification method provided by the invention has good use value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, specifically relating to a method for single-tree segmentation and tree species identification based on multi-source, multi-scale masking networks. Background Technology

[0002] The management and protection of forest ecosystems are crucial for maintaining biodiversity, mitigating climate change, and supporting a variety of ecosystem services. In forest management and protection, accurate tree segmentation allows researchers to obtain detailed information on tree species, size, distribution, and health status, facilitating a more comprehensive assessment of forest ecosystems. Furthermore, individual tree segmentation and species identification enable precise estimation of forest biomass and carbon storage, contributing to an understanding of forests' role in carbon sequestration and climate change mitigation. Unlike traditional segmentation tasks that focus on depicting homogeneous areas, individual tree segmentation aims to divide the vegetation canopy into distinct canopies and identify their species types, with each canopy representing a single tree. This task presents several unique challenges, including variations in tree species, canopy shape, and shading from surrounding objects, making it a challenging yet vital problem in forestry and ecology.

[0003] Traditional methods of forest inventory and tree segmentation are often time-consuming, labor-intensive, and have limited spatial coverage. Furthermore, they can be subjective and inconsistent in data collection and analysis. Remote sensing technology offers a promising solution to these challenges by providing high-resolution and multimodal data that can be used to describe the characteristics of forest ecosystems in detail. In recent years, remote sensing data has been widely used in tree segmentation and species identification as a means of accurately quantifying forest attributes at the individual tree level. In particular, hyperspectral images (HSI), high-resolution RGB images, and LiDAR data have become valuable tools for forest monitoring and management. Hyperspectral data provides detailed spectral information over a wide wavelength range, allowing for precise differentiation of different tree species based on their unique spectral characteristics. By analyzing the spectral reflectance patterns captured by hyperspectral sensors, tree species can be identified, and this species information can be used to improve the accuracy of tree segmentation. High-resolution RGB images provide detailed spatial information about the canopy and its surrounding environment, allowing for the identification of fine-scale canopy features such as branch structure and leaf texture. These images are well-suited for accurately depicting the canopy. LiDAR data can capture detailed three-dimensional information about forest canopy and topography, providing additional insights into tree morphology and canopy structure. By analyzing LiDAR-derived metrics such as canopy height, canopy density, and vertical vegetation distribution, researchers can improve the accuracy of individual tree segmentation. A watershed algorithm was used to segment the canopy height model generated from LiDAR data, followed by normalized segmentation to achieve individual tree segmentation.

[0004] Numerous methods for individual tree segmentation and species identification have been proposed, which can be categorized based on the data source used: hyperspectral image-based segmentation, high-resolution RGB image-based segmentation, LiDAR data-based segmentation, and multi-source data fusion-based segmentation. These methods employ traditional or deep learning approaches to propose various segmentation models for individual tree segmentation and species identification, such as region growing based on high-resolution images, watershed segmentation algorithms based on the Canopy Height Model (CHM), and deep learning algorithms based on point cloud and spectral fusion. Among these methods, multi-source data fusion-based methods demonstrate better segmentation performance than those based on a single data source. Extensive research has shown that utilizing multi-source data enhances the model's ability to identify features. These complementary data sources provide detailed information about canopy features, spatial distribution, and structural attributes, improving the accuracy and robustness of tree segmentation algorithms. Therefore, effectively integrating hyperspectral data, high-resolution RGB data, and LiDAR data for multi-source feature extraction is crucial for achieving high-precision individual tree segmentation and species identification.

[0005] However, most existing multi-source data fusion methods require manual feature design. Considering the large number of features from multi-source remote sensing data, manually selecting effective features and designing extraction methods requires forestry expertise and extensive experimentation. This process is not only time-consuming but may also fail to fully capture the complex high-dimensional patterns present in remote sensing data, lacking generalization and robustness. Researchers have proposed using deep learning methods to mitigate this problem. While deep learning is fully capable of handling tasks from a single data source, there is no relevant experience in fusing hyperspectral images, high-resolution RGB images, and LiDAR data. These three types of data can provide feature information about trees from different perspectives, but existing research only performs feature fusion by extracting and then combining features. This fails to fully utilize the characteristics of the three types of data to optimize the network, limiting the effectiveness of single-tree segmentation and tree species identification. Therefore, it is necessary to analyze the characteristics of the three types of data and design the network accordingly. Furthermore, the three types of data exhibit scale inconsistencies. Simply fusing data at a single scale makes it difficult to simultaneously consider both global and local feature information. At a large scale, the focus is on the spatial distribution information of different tree species globally, while at a small scale, more attention is paid to information such as tree geometric texture. Therefore, cross-scale feature extraction is needed to achieve full fusion of information at multiple scales.

[0006] In existing technologies, such as in Reference 1 (Dalponte M, Bruzzone L, Gianelle D. Tree species classification in the Southern Alps based on the fusion of very high geometrical resolution multispectral / hyperspectral images and LiDAR data[J]. Remote sensing of environment, 2012, 123: 258-270.), Dalponte combined hyperspectral images and LiDAR data to classify trees in the Southern Alps, achieving the highest accuracy; in Reference 2 (Alonzo M, Bookhagen B, Roberts DA. Urban tree species mapping using hyperspectral and lidar data fusion[J]. Remote sensing of environment, 2014, 148: 70-83.), Alonzo et al. combined hyperspectral images generated from lidar data with CHM images to classify 29 common tree species in the United States, improving the accuracy by 4.2 percentage points compared to using spectral data alone; in Reference 3 (Qin H, Zhou W, Yao In the paper "Individual tree segmentation and tree species classification in subtropical broadleaf forests using UAV-based LiDAR, hyperspectral, and ultrahigh-resolution RGB data[J]. Remote Sensing of Environment, 2022, 280: 113-143.", Qin et al. proposed a watershed-spectral-texture controlled normalized segmentation algorithm for single-tree segmentation, combining UAV LiDAR, hyperspectral, and high-resolution RGB data. However, the methods provided in the above literature require manual design of canopy features, which may not be able to fully capture the complex high-dimensional patterns in remote sensing data, and may not be able to fully mine the feature information of tree species in hyperspectral images, high-resolution RGB images, and LiDAR data, thus hindering the effective interaction and fusion of various feature information. In addition, designing effective tree features also requires expertise in forestry and a large number of experiments, making the process very time-consuming and prone to suboptimal choices.Furthermore, handcrafted features are often tailored to specific tree seed sets, thus limiting the generalization ability of traditional algorithms. Additionally, these algorithms may be sensitive to noise, artifacts, or variations in the tree data, leading to unstable or inaccurate segmentation results.

[0007] In addition, the existing processing methods also have the following problems: 1) Additional post-processing steps are required, such as non-maximum suppression (NMS), which requires parameter design and may introduce additional errors; 2) They can perform well for single-tree level recognition tasks, but are difficult to achieve good results for finer-grained single-tree segmentation tasks.

[0008] In view of this, the present invention is proposed. Summary of the Invention

[0009] To address the shortcomings of existing technologies, the present invention aims to provide a single-tree segmentation and tree species identification method based on a multi-source, multi-scale mask network. This method fully integrates features from different types of data and achieves multi-source, multi-scale feature learning for each feature point, thereby greatly reducing the attention mechanism and computational cost while improving segmentation accuracy.

[0010] To achieve the above objectives, the present invention adopts the following technical solution:

[0011] This invention provides a method for single-tree segmentation and tree species identification based on a multi-source, multi-scale mask network, comprising the following steps:

[0012] S1. Acquire hyperspectral images, high-resolution RGB images, and CHM images generated from lidar data;

[0013] S2. The hyperspectral image, high-resolution RGB image, and CHM image obtained in step S1 are processed by different projection modules to obtain hyperspectral features, high-resolution RGB image features, and CHM image features. Then, a fusion encoder based on deformable fusion attention is used to fuse the above hyperspectral features, high-resolution RGB image features, and CHM image features to obtain multi-source multi-scale fusion features.

[0014] S3. Process the hyperspectral image to obtain the normalized vegetation index, and integrate it with CHM image features to generate a spatial mask, thereby initializing the object query;

[0015] S4. The decoder obtains multi-source multi-scale fusion features, performs cross-attention interaction between the query object and the multi-source multi-scale fusion features, adjusts the region of interest of the query layer by layer, and feeds back the prediction mask of the current layer to the attention calculation of the next layer.

[0016] S5. The instance segmentation module generates masks and classifies the optimized object queries, and outputs pixel-level segmentation results with tree species labels by combining the pixel-level features of the encoder.

[0017] As a preferred technical solution, in step S2, the fusion encoding of the fusion encoder is represented as follows:

[0018]

[0019] in, and These represent the outputs of the l-th encoder and the (l-1)-th encoder, respectively.

[0020] DFA stands for Deformable Fusion Attention;

[0021] LN represents the layer normalization module;

[0022] FFN stands for Feedforward Neural Network;

[0023] x HSI Indicates hyperspectral characteristics, x RGB Represents the features of a high-resolution RGB image, x CHM This represents the features of a CHM image.

[0024] As a preferred technical solution, x HSI =ProH(I HSI );x RGB =ProR(I RGB );x CHM =ProR(I CHM );

[0025] Among them, I HSI For hyperspectral images, I RGB For high-resolution RGB images, I CHM This is a CHM image.

[0026] As a preferred technical solution,

[0027]

[0028] Where K is the number of sampling points, L is the number of feature map layers, E is the number of data types, and D is the number of attention heads (lowercase letters).

[0029] k, l, e, and d represent their respective indices;

[0030] H d It is used to distinguish the weights of different attention heads;

[0031] S de It is the weight of the e-th data feature in the d-th attention head;

[0032] A dlk This represents the attention weight at the k-th sampling point on the l-th feature map in the d-th attention head;

[0033] f el From x through a linear layer RGB or x CHM The obtained feature map of layer l;

[0034] p refers to all pixels on the feature map;

[0035] Off dlk This represents the offset of the k-th sampling point on the l-th feature map in the d-th attention head;

[0036] p+Off dlk Generate the two-dimensional coordinates of all sampling points on the l-th layer feature map in the d-th attention head.

[0037] As a preferred technical solution,

[0038]

[0039] in, This represents the k-th feature sampling point on the l-th layer hyperspectral feature map in the d-th attention head.

[0040] As a preferred technical solution,

[0041]

[0042] in, This represents the CHM feature map of the l-th layer. This represents the RGB feature map of the l-th layer.

[0043] As a preferred technical solution,

[0044]

[0045] in, This represents the k-th feature sampling point on the l-th layer hyperspectral feature map in the d-th attention head;

[0046]

[0047] As a preferred technical solution, in step S3, the spatial mask is represented as follows:

[0048]

[0049] Where x and y represent the horizontal and vertical coordinates of the pixels on the mask, respectively.

[0050] As a preferred technical solution, in step S4, the prediction mask generated by the object query after cross-attention interaction is represented as follows:

[0051]

[0052] Where x and y represent the horizontal and vertical coordinates of the pixels on the mask, respectively;

[0053] X l This represents the feature map of the l-th layer.

[0054] As a preferred technical solution,

[0055]

[0056] in, This represents the prediction mask for the (l-1)th layer;

[0057] Q l V l , These represent the three input features of the attention mechanism, respectively, represented by X. l-1 Each is obtained by passing through a linear layer;

[0058] This represents the l-th layer of input features that covers the CHM mask;

[0059] X l-1 Represented as the feature map of layer l-1;

[0060] ⊙ represents element-wise multiplication.

[0061] Compared with the prior art, the present invention has the following beneficial effects:

[0062] (1) In the segmentation and tree species identification method provided by the present invention, a multi-source multi-scale fusion encoder is designed to effectively fuse the features of hyperspectral images, high-resolution RGB images and LiDAR data. Projection modules of the three types of data are constructed respectively, and data features are fully fused on multiple data sources and multiple scales, effectively realizing data interaction in different feature spaces and interaction of global and local feature information.

[0063] (2) In the segmentation and tree species identification method provided by this invention, a deformable fusion attention mechanism is designed. By selecting a small number of sparse feature points on different data and scales to calculate attention, the complexity and computational cost of the attention mechanism are greatly reduced. It also considers the feature representation of hyperspectral data, high-resolution RGB data and CHM images at different scales, and realizes multi-source, multi-scale feature learning for each feature point.

[0064] (3) In the segmentation and tree species identification method provided by the present invention, a query constraint decoder is proposed and the traditional Transformer design is improved. Through the query constraint module, the spatial features of the tree dataset are explicitly embedded in the object query initialization process to constrain the learning space of the object query and refine the convergence process of the query, thereby improving the segmentation accuracy of a single tree. Attached Figure Description

[0065] Figure 1 For M 3 Former model framework diagram;

[0066] Figure 2 This is a deformable fusion attention structure diagram;

[0067] Figure 3 This is a diagram of the query constraint decoder structure. Detailed Implementation

[0068] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0069] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0070] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0071] Reference Figures 1-3 This invention provides a method for single-tree segmentation and tree species identification based on a multi-source, multi-scale mask network, comprising the following steps:

[0072] S1. Acquire hyperspectral images, high-resolution RGB images, and CHM images generated from lidar data;

[0073] S2. The hyperspectral image, high-resolution RGB image, and CHM image obtained in step S1 are processed by different projection modules to obtain hyperspectral features, high-resolution RGB image features, and CHM image features. Then, a fusion encoder based on deformable fusion attention is used to fuse the above hyperspectral features, high-resolution RGB image features, and CHM image features to obtain multi-source multi-scale fusion features.

[0074] S3. Process the hyperspectral image to obtain the normalized vegetation index, and integrate it with CHM image features to generate a spatial mask, thereby initializing the object query;

[0075] S4. The decoder obtains multi-source multi-scale fusion features, performs cross-attention interaction between the query object and the multi-source multi-scale fusion features, adjusts the region of interest of the query layer by layer, and feeds back the prediction mask of the current layer to the attention calculation of the next layer.

[0076] S5. The instance segmentation module generates masks and classifies the optimized object queries, and outputs pixel-level segmentation results with tree species labels by combining the pixel-level features of the encoder.

[0077] The above technical solution implements a multi-source and multi-scale mask Transformer network that uses hyperspectral images, high-resolution RGB images, and LiDAR data. 3 Former) provides a better technical solution for single-tree segmentation and tree species identification tasks. First, it generates multi-scale feature maps for three types of data through three projection modules. Then, M 3 Former uses Deformable Fusion Attention (DFA) to encode these feature maps at different scales, obtaining multi-source fused features with spectral, textural, and spatial characteristics. The decoder acquires these multi-source and multi-scale fused features and interacts with the target query, adjusting the region of interest for each target query to generate a prediction mask. It's important to note that before the object query interacts with the fused features, it needs to be initialized using, but not limited to, a Query Constraint Decoder (QCD). The QCD extracts Normalized Difference Vegetation Index (NDVI) features from the hyperspectral image and integrates them with CHM image features to generate spatial masks, using these spatial masks to constrain the computational range of the attention. Furthermore, the prediction mask generated from the object query is added to the attention of the next layer of the decoder; this process establishes an initial feature learning space for the object query. After decoder iteration, the object query is obtained by, but not limited to, an instance segmentation module. The instance segmentation module allows the object query to obtain a prediction mask and predicted class through two feedforward networks. The prediction mask requires combining the pixel-level features provided by the encoder to generate a pixel-level mask for the image, and then combining it with the prediction classification to obtain the final segmentation result. Among them, the instance segmentation module is a core component in computer vision that combines object detection and semantic segmentation. It aims to distinguish different object instances in the same category and generate a pixel-level mask for each instance.

[0078] In some implementations, in step S1, the hyperspectral image, high-resolution RGB image, and lidar data can be obtained through general means. The obtained lidar data can be processed into a CHM image through steps such as data acquisition, data preprocessing (e.g., denoising), point cloud classification, and model generation. The calculation model for canopy height involved in the classification and model generation steps is general knowledge that is known to those skilled in the art, and it is not discussed in detail in this invention. It is based on the fact that those skilled in the art can implement it.

[0079] Of course, in step S1, to better process the image, the corresponding image can also be cropped to obtain several blocks. Let the hyperspectral image be... High-resolution RGB images are CHM image is Where H and W represent the image height and width, respectively, B represents the spectral dimension of the hyperspectral image, and C and D represent the channel dimensions of the high-resolution RGB image and CHM image, respectively.

[0080] In some implementations, see Figure 1 In step S2, the image data obtained in step S1 is first processed (e.g., by different projection modules) and mapped to the same feature space to obtain the hyperspectral feature representation as x. HSI =ProH(I HSI High-resolution RGB image features are represented as x RGB =ProR(I RGB CHM image features are represented as x CHM =ProR(I CHM The three features each have four scales, corresponding to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image size. The three projection modules consist of convolutional layers, pooling layers, and residual connections. ProH includes a 7×7 convolutional layer, a 3×3 pooling layer, and four 3×3 convolutional layers. The final four convolutional layers output four hyperspectral feature maps at different scales for subsequent multi-scale fusion. ProR, after generating multi-scale features, adds a 5×5 pooling layer and a 3×3 convolutional layer for dimensionality reduction. This adjustment better matches and fuses high-resolution and low-resolution features. ProC is similar to ProH but reduces the size parameters of the intermediate layers. It is important to emphasize that the projection modules mentioned above are a superior processing method; however, 3D-ResNet, ViT (VisionTransformer), and other processing methods can also be used to process the image. The fusion encoding of the above fusion encoder is represented as follows:

[0081]

[0082] in, and These represent the outputs of the l-th encoder and the (l-1)-th encoder, respectively.

[0083] DFA stands for Deformable Fusion Attention;

[0084] LN represents the layer normalization module;

[0085] FFN stands for Feedforward Neural Network;

[0086] x HSI Indicates hyperspectral characteristics, x RGB Represents the features of a high-resolution RGB image, x CHM This represents the features of a CHM image.

[0087] Deformable Fusion Attention (DFA) is a hybrid attention model that combines deformable convolution with multi-head attention mechanisms. It aims to enhance the model's ability to fuse features from multiple sources and at multiple scales by dynamically adjusting sampling positions and weight allocation. Its core idea is to introduce learnable offsets so that the attention mechanism can adaptively focus on key regions, making it particularly suitable for multimodal data processing in complex scenarios (such as hyperspectral, LiDAR, and RGB fusion).

[0088] LN is a feature layer normalization technique that solves the problem of internal covariate shift in deep neural networks by normalizing all features of a single sample.

[0089] Feedforward Neural Network (FFN) is a fundamental and widely used neural network model. Its core feature is unidirectional information flow (from the input layer to the output layer), with no feedback or recurrent connections. As a foundational architecture for deep learning, FFN plays a crucial role in fields such as image recognition and natural language processing.

[0090] As can be seen, in multi-source, multi-scale fusion encoders, when connecting feature maps of different scales and inputting them into the attention mechanism, deformable fusion attention is used to reduce computational cost at high resolutions. In some implementations,

[0091] Where K is the number of sampling points, L is the number of layers in the feature map, E is the number of data types, D is the number of attention heads, and lowercase letters k, l, e, and d represent their respective indices;

[0092] H d It is used to distinguish the weights of different attention heads;

[0093] S deIt is the weight of the e-th data feature in the d-th attention head;

[0094] A dlk This represents the attention weight at the k-th sampling point on the l-th feature map in the d-th attention head;

[0095] f el From x through a linear layer RGB or x CHM The obtained feature map of layer l;

[0096] p refers to all pixels on the feature map;

[0097] Off dlk This represents the offset of the k-th sampling point on the l-th feature map in the d-th attention head;

[0098] p+Off dlk Generate the two-dimensional coordinates of all sampling points on the l-th layer feature map in the d-th attention head.

[0099] In some implementations...

[0100] in, This represents the k-th feature sampling point on the l-th layer hyperspectral feature map in the d-th attention head.

[0101] In some implementations...

[0102] in, This represents the CHM feature map of the l-th layer. This represents the RGB feature map of the l-th layer.

[0103]

[0104] in, This represents the k-th feature sampling point on the l-th layer hyperspectral feature map in the d-th attention head;

[0105]

[0106] In deformable fusion attention, make x HSI Each pixel on the surface only focuses on the x-axis. RGB and x CHM Instead of focusing on all pixels, it focuses on a sparse subset of the pixels, which reduces the computational burden of the attention mechanism. HSIA linear layer generates k offsets, each including horizontal and vertical offsets. These K offsets are added to the 2D coordinates of each feature point, resulting in K sampling points. Sampling is performed from feature map f based on the coordinates of these sampling points, sampling K points on each feature layer of f. Since each f has L feature layers, each point in p receives 2LK sampling points. In other words, each point in p only needs to fuse the features of these 2LK sampling points, and because K is much smaller than the total number of pixels in the feature map, this method achieves the goals of reducing attention computation and accelerating network convergence.

[0107] Furthermore, in the aforementioned Query Constraint Decoder (QCD), see [link to QCD documentation]. Figure 3 It consists of a filter and a decoder with masking-constrained attention. The NDVI index is extracted from the hyperspectral image. Based on its value in the CHM image, pixels with an NDVI index greater than v (v = 0.2) are filtered out, and the remaining pixels are set to zero, occluding interfering elements such as land, buildings, and water bodies. This process initially constructs a spatial mask for trees. The masked image is input into a convolutional network to extract features, and then passes through two linear layers to obtain the key and value of the attention mechanism. In addition, spatial relationships between feature points are enhanced through two-dimensional spatial location embedding. After multiplying the object query with the key to obtain attention weights, a mask is applied to these attention weights.

[0108] The mask consists of two parts: one part is a CHM image mask filtered by the NDVI index, which serves as the initial feature space state of interest, represented as:

[0109]

[0110] Where x and y represent the horizontal and vertical coordinates of the pixels on the mask, respectively.

[0111] The other part is the prediction mask generated by the object query after the interaction, represented as:

[0112]

[0113] Where x and y represent the horizontal and vertical coordinates of the pixels on the mask, respectively;

[0114] X l The feature map of the l-th layer can be further represented as:

[0115]

[0116] in, This represents the prediction mask for the (l-1)th layer;

[0117] Q l V l , These represent the three input features of the attention mechanism, respectively, represented by X. l-1 Each is obtained by passing through a linear layer;

[0118] This represents the l-th layer of input features that covers the CHM mask;

[0119] X l-1 Represented as the feature map of layer l-1;

[0120] ⊙ represents element-wise multiplication.

[0121] In the above technical solutions, the Sigmoid function is a non-linear activation function widely used in machine learning and deep learning, often used for binary classification probability output and non-linear feature transformation in neural networks; the Softmax function is the core activation function used for multi-classification tasks in machine learning and deep learning, and its core function is to convert any real number vector into a probability distribution, with the sum of the output values ​​being 1.

[0122] Clearly, based on the processing in steps S4 and S5, the object query, after passing through each layer of the decoder with limited attention, enters the segmentation inference module to generate the prediction mask. pre The inference module consists of two parts: predicting a single tree mask and the tree category, respectively. This predicted mask is then fed into the next layer of the decoder's constrained attention mechanism, where it interacts with the mask. CHM The intersection process adds a learnable iterative constraint mask to the attention weights. Under the influence of the mask, the range of features learned in the query is limited, enabling it to more accurately identify the features that need attention, thus improving the accuracy of single-tree segmentation and tree species identification in multi-source, multi-scale mask networks.

[0123] Based on the method provided by this invention, the following has been achieved:

[0124] (1) Using a multi-source, multi-scale segmentation network based on mask Transformer is beneficial for the interaction of spatial and local information of different data and fully integrates the features of various types of data.

[0125] (2) It uses deformable fusion attention based on sparse feature points to calculate attention by selecting a small number of sparse feature points on different data and scales, which greatly reduces the complexity and computational cost of the attention mechanism. It also considers the feature representation of hyperspectral data, high-resolution RGB data and CHM images at different scales, and realizes multi-source, multi-scale feature learning for each feature point.

[0126] (3) Using a query constraint decoder based on CHM image spatial mask, the spatial features of the tree dataset are explicitly embedded during the initialization process of object query, constraining the learning space of object query, refining the convergence process of query, thereby improving the segmentation accuracy of single tree.

[0127] Therefore, the single-tree segmentation and tree species identification method provided by this invention has good practical value.

[0128] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for single-tree segmentation and tree species identification based on multi-source, multi-scale masking networks, characterized in that, Includes the following steps: S1. Acquire hyperspectral images, high-resolution RGB images, and CHM images generated from lidar data; S2. The hyperspectral image, high-resolution RGB image, and CHM image obtained in step S1 are processed by different projection modules to obtain hyperspectral features, high-resolution RGB image features, and CHM image features. Then, a fusion encoder based on deformable fusion attention is used to fuse the above hyperspectral features, high-resolution RGB image features, and CHM image features to obtain multi-source multi-scale fusion features. S3. Process the hyperspectral image to obtain the normalized vegetation index, and integrate it with CHM image features to generate a spatial mask, thereby initializing the object query; S4. The decoder obtains multi-source multi-scale fusion features, performs cross-attention interaction between the query object and the multi-source multi-scale fusion features, adjusts the region of interest of the query layer by layer, and feeds back the prediction mask of the current layer to the attention calculation of the next layer. S5. The instance segmentation module generates masks and classifies the optimized object queries, and outputs pixel-level segmentation results with tree species labels by combining the pixel-level features of the encoder.

2. The method for single-tree segmentation and tree species identification based on multi-source multi-scale masking networks according to claim 1, characterized in that, In step S2, the fusion encoding of the fusion encoder is represented as follows: in, and These represent the outputs of the l-th encoder and the (l-1)-th encoder, respectively. DFA stands for Deformable Fusion Attention; LN represents the layer normalization module; FFN stands for Feedforward Neural Network; x HSI Indicates hyperspectral characteristics, x RGB Represents the features of a high-resolution RGB image, x CHM This represents the features of a CHM image.

3. The method for single-tree segmentation and tree species identification based on multi-source multi-scale mask networks according to claim 2, characterized in that, x HSI =ProH(I HSI );x RGB =ProR(I RGB );x CHM =ProR(I CHM ); Among them, I HSI For hyperspectral images, I RGB For high-resolution RGB images, I CHM This is a CHM image.

4. The method for single-tree segmentation and tree species identification based on multi-source multi-scale mask networks according to claim 3, characterized in that, Where K is the number of sampling points, L is the number of layers in the feature map, E is the number of data types, D is the number of attention heads, and lowercase letters k, l, e, and d represent their respective indices; H d It is used to distinguish the weights of different attention heads; S de It is the weight of the e-th data feature in the d-th attention head; A dlk This represents the attention weight at the k-th sampling point on the l-th feature map in the d-th attention head; f el From x through a linear layer RGB or x CHM The obtained feature map of layer l; p refers to all pixels on the feature map; Off dlk This represents the offset of the k-th sampling point on the l-th feature map in the d-th attention head; p+Off dlk Generate the two-dimensional coordinates of all sampling points on the l-th layer feature map in the d-th attention head.

5. The method for single-tree segmentation and tree species identification based on multi-source multi-scale masking networks according to claim 4, characterized in that, in, This represents the k-th feature sampling point on the l-th layer hyperspectral feature map in the d-th attention head.

6. The method for single-tree segmentation and tree species identification based on multi-source multi-scale mask networks according to claim 4, characterized in that, in, This represents the CHM feature map of the l-th layer. This represents the RGB feature map of the l-th layer.

7. The method for single-tree segmentation and tree species identification based on multi-source multi-scale mask networks according to claim 4, characterized in that, in, This represents the k-th feature sampling point on the l-th layer hyperspectral feature map in the d-th attention head; 8. The method for single-tree segmentation and tree species identification based on multi-source multi-scale mask networks according to claim 1, characterized in that, In step S3, the spatial mask is represented as: Where x and y represent the horizontal and vertical coordinates of the pixels on the mask, respectively.

9. The method for single-tree segmentation and tree species identification based on multi-source multi-scale masking networks according to claim 1, characterized in that, In step S4, the prediction mask generated after the cross-attention interaction for object query is represented as follows: Where x and y represent the horizontal and vertical coordinates of the pixels on the mask, respectively; X l This represents the feature map of the l-th layer.

10. The method for single-tree segmentation and tree species identification based on multi-source multi-scale mask networks according to claim 9, characterized in that, in, This represents the prediction mask for the (l-1)th layer; Q l V l , These represent the three input features of the attention mechanism, respectively, represented by X. l-1 Each is obtained by passing through a linear layer; This represents the l-th layer of input features that covers the CHM mask; X l-1 Represented as the feature map of layer l-1; ⊙ represents element-wise multiplication.