A three-dimensional point cloud tree segmentation method, system, terminal and storage medium based on skip connection gate and query attention

CN122574004BActive Publication Date: 2026-09-29GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611041068.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29
Estimated Expiration
2046-07-14

AI Technical Summary

Technical Problem

[0004]本发明的主要目的在于提供一种基于跳连门控与查询注意力的三维点云单木分割方法、系统、终端及计算机可读存储介质,旨在解决现有的单木分割方法由于在跳跃连接处直接拼接特征、缺乏动态注意力机制、损失函数对边界点的错误不敏感导致单木分割结果精度较低的问题

Benefits of technology

[0015]本发明中,构建初始单木分割模型,并采用多任务协同损失函数对所述初始单木分割模型进行训练,得到单木分割模型;获取初始森林三维点云,并对所述初始森林三维点云进行预处理,得到森林三维点云;通过所述单木分割模型的改进稀疏3D U-Net,对所述森林三维点云进行特征提取、基于跳连双门控融合模块的特征融合和特征解码映射,得到点级特征;通过所述单木分割模型的查询注意力模块,对所述点级特征进行特征增强,得到增强点级特征;根据所述森林三维点云中每个点的三维空间坐标和所述增强点级特征,通过所述单木分割模型的多任务分割头进行语义分类和偏移回归,得到每个所述点的语义类别预测结果和指向实例中心的三维偏移向量;根据每个点的语义类别预测结果和指向实例中心的三维偏移向量,对每个所述点进行坐标偏移修正和基于HDBSCAN密度聚类算法的聚类,得到单木实例分割结果。本发明通过引入Query-based注意力机制(基于查询的注意力机制)突破局部感受野限制,实现对树木全局结构的感知;通过在稀疏3D U-Net网络跳跃连接处引入面向单木分割的跳连双门控融合模块,对编码器特征进行自适应筛选,有效抑制了浅层噪声和无关细节,同时显著增强了网络对分叉点、树冠交叠区和实例边界等高价值区域的特征表达;通过设计融合Focal Loss(焦点损失)、Smooth L1(平滑L1损失)与实例一致性约束的多任务协同损失函数,强化边界难分样本学习并提升实例内部预测的连贯性与回归鲁棒性。与现有技术相比,本发明有效克服了复杂林分中树冠严重交叠带来的“欠分割”与“过分割”难题,通过自适应门控筛选与全局注意力显著提升了特征的有效性与模型在关键区域的判别力,最终生成的分割结果具有更准确的几何形态和更清晰的轮廓,即最终生成的分割结果精度更高,为后续的林业参数(如胸径、冠幅、生物量)估算提供了更高质量的数据基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574004B_ABST
    Figure CN122574004B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data processing, and discloses a three-dimensional point cloud tree segmentation method, system, terminal and storage medium based on skip connection gating and query attention, the method comprising: training an initial tree segmentation model by using a multi-task cooperative loss function to obtain a tree segmentation model; performing feature fusion on forest three-dimensional point clouds based on a skip connection double-gating fusion module by improving a sparse 3D U-Net to obtain point-level features; performing feature enhancement on the point-level features by a query attention module to obtain enhanced point-level features; performing semantic classification and offset regression on the enhanced point-level features by a multi-task segmentation head to obtain semantic category prediction results and three-dimensional offset vectors; and finally, obtaining tree instance segmentation results through coordinate offset correction and clustering. The application improves the effectiveness of features and the discriminability of a model in a key area through adaptive gating filtering and global attention, so that the final generated tree segmentation results are more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, system, terminal, and computer-readable storage medium for three-dimensional point cloud single-tree segmentation based on jump gating and query attention. Background Technology

[0002] In recent years, forest 3D point cloud data acquisition technology based on LiDAR (Light Detection and Ranging) has been widely used. LiDAR can quickly acquire large-scale, high-density, and high-precision forest spatial structure information, providing an important data foundation for forest structure parameter extraction, biomass estimation, carbon storage assessment, and forest management decisions. Accurately segmenting individual tree instances from forest 3D point clouds is a crucial preliminary step for extracting parameters such as tree height, crown width, diameter at breast height (DBH), and number of trees. However, existing deep learning-based single-tree segmentation methods have significant drawbacks. First, existing sparse 3D U-Nets directly splice features at skip connections, easily introducing shallow noise and weakening the representation of high-value regions. Second, existing models lack a dynamic attention mechanism to guide the network to focus limited computational resources on these structurally complex "high-information" regions, and cannot adaptively enhance feature channels that are more discriminative in distinguishing between wood and leaves. Finally, mainstream segmentation loss functions tend to optimize the large, dominant region while being insensitive to errors at the few boundary points, resulting in segmented boundaries that are often blurry, smooth, or jagged. These drawbacks lead to low accuracy in individual tree segmentation, severely impacting the accuracy of subsequent estimations of key forestry parameters such as diameter at breast height (DBH) and crown width.

[0003] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0004] The main objective of this invention is to provide a method, system, terminal, and computer-readable storage medium for 3D point cloud single-tree segmentation based on jump connection gating and query attention. This aims to solve the problems of low accuracy in existing single-tree segmentation methods, which suffer from direct feature splicing at jump connections, lack of dynamic attention mechanism, and insensitivity of loss function to errors at boundary points.

[0005] To achieve the above-mentioned objectives, this invention provides a method for segmenting 3D point clouds into individual trees based on jump-connection gating and query attention. The method includes: An initial single-tree segmentation model is constructed, and the initial single-tree segmentation model is trained using a multi-task collaborative loss function to obtain a single-tree segmentation model; An initial 3D point cloud of the forest is obtained, and the initial 3D point cloud of the forest is preprocessed to obtain a 3D point cloud of the forest. The improved sparse 3D U-Net of the single tree segmentation model is used to extract features from the forest 3D point cloud, perform feature fusion based on the jump-connected dual-gated fusion module, and perform feature decoding mapping to obtain point-level features. The point-level features are enhanced by the query attention module of the single-tree segmentation model to obtain enhanced point-level features. Based on the three-dimensional spatial coordinates of each point in the forest three-dimensional point cloud and the enhanced point-level features, semantic classification and offset regression are performed through the multi-task segmentation head of the single tree segmentation model to obtain the semantic category prediction result and the three-dimensional offset vector pointing to the instance center for each point. Based on the semantic category prediction results of each point and the three-dimensional offset vector pointing to the instance center, coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm are performed on each point to obtain the single-tree instance segmentation results.

[0006] Optionally, the step of constructing an initial single-tree segmentation model and training the initial single-tree segmentation model using a multi-task collaborative loss function to obtain a single-tree segmentation model specifically includes: Construct an initial single-tree segmentation model, wherein the initial single-tree segmentation model includes an initial improved sparse 3D U-Net, an initial query attention module, and an initial multi-task segmentation head, and the initial improved sparse 3D U-Net includes an initial jump connection dual-gated fusion module set at the jump connection; We construct a semantic classification loss based on Focal Loss, a bias loss based on Smooth L1 loss, and an instance consistency loss based on instance consistency constraints. Combining these three losses, we obtain a multi-task collaborative loss function: ; in, For multi-task collaborative loss function, For semantic classification loss, For offset loss, For instance consistency loss, For semantic classification loss weights, As the weight for offset loss, Weight the instance consistency loss; The error between the output of the initial single-tree segmentation model and the true label is calculated using a multi-task collaborative loss function. The parameters of the initial single-tree segmentation model are then iteratively optimized using a backpropagation algorithm until the initial single-tree segmentation model converges, resulting in the single-tree segmentation model: The single-tree segmentation model includes an improved sparse 3D U-Net, a query attention module, and a multi-task segmentation head. The improved sparse 3D U-Net includes a jump-connection dual-gated fusion module set at the jump connection.

[0007] Optionally, the improved sparse 3D U-Net also includes an encoder and a decoder; The improved sparse 3D U-Net of the single-tree segmentation model is used to extract features from the forest 3D point cloud, perform feature fusion based on the jump-connected dual-gated fusion module, and perform feature decoding mapping to obtain point-level features, specifically including: Multi-scale feature extraction is performed on the forest 3D point cloud using the sparse convolutional layer and residual block of the encoder to obtain encoder features; The encoder features are progressively upsampled through the sparse deconvolution layer of the decoder to obtain the decoder features; The jump-connect dual-gated fusion module performs feature fusion on the encoder features and the decoder features based on encoder feature adaptive filtering and channel response adaptive enhancement to obtain gated fusion features. The gated fusion features are upsampled and mapped using the remaining decoding layers and output mapping layers in the decoder to obtain point-level features corresponding to points in the forest 3D point cloud: ; in, Indicates gating fusion characteristics, This indicates upsampling and mapping operations. This represents point-level features.

[0008] Optionally, the jump-connect dual-gated fusion module includes a spatial gated branch and a channel gated branch; The method of using the jump-connect dual-gated fusion module to perform feature fusion on the encoder features and decoder features based on encoder feature adaptive filtering and channel response adaptive enhancement to obtain gated fusion features specifically includes: The encoder features and the decoder features are concatenated for the first time to obtain the first intermediate fused features: ; in, Indicates the first The first intermediate fusion feature of the layer, Indicates the first Encoder features of the layer Indicates the first Decoder features of the layer; The first intermediate fusion feature is input into the spatial gating branch, and the spatial gating branch generates jump selection weights based on the first intermediate fusion feature through a spatial perception multilayer perceptron: ; in, This indicates the weight for selecting the jump connection. This represents the Sigmoid activation function. This represents a multilayer perceptron for spatial perception. The spatial gating branch uses the jump connection selection weights to adaptively weight the encoder features point-by-point and channel-by-channel, resulting in enhanced encoder features: ; in, For the first Enhanced encoder features after spatial gating branch filtering; The enhanced encoder features and the decoder features are concatenated a second time to obtain the second intermediate fused features: ; in, Indicates the first The second intermediate fusion feature of the layer; The second intermediate fusion feature is input into the channel gating branch, and the channel gating branch performs global pooling on the second intermediate fusion feature to obtain the channel description vector; The channel gating branch generates channel recalibration weights based on the channel description vector using a channel-aware multilayer perceptron. ; in, This indicates that the channel recalibrates its weights. and Let represent the first and second learnable weight matrices in the channel-aware multilayer perceptron, respectively. Represents a non-linear activation function. Represents the channel description vector; The channel-gated branch uses the channel recalibration weights to dynamically recalibrate the channel dimension of the second intermediate fused feature, thereby obtaining the gated fused feature: ; in, Indicates the first Layer-gated fusion features.

[0009] Optionally, the query attention module includes multiple attention layers, wherein the attention layer includes cross attention units, self attention units, and feedforward networks; The step of enhancing the point-level features through the query attention module of the single-tree segmentation model to obtain enhanced point-level features specifically includes: The three-dimensional spatial coordinates of points in the forest three-dimensional point cloud are processed by position encoding to generate the position codes of the points: ; in, Represents the three-dimensional spatial coordinates of a point. This represents the position encoding function. Indicates positional encoding; A learnable query state is introduced. Based on the query state, the location encoding of the point, and the point-level features, attention is calculated through the cross-attention unit to obtain the query state after aggregating global information. ; ; in, Indicates by the first Layer query status The query vector obtained from the mapping This represents the key vector obtained from the point-level feature mapping after fusion of positional encoding. This represents the value vector obtained from point-level feature mapping. , and These represent the learnable linear projection matrices corresponding to the query vector, key vector, and value vector, respectively. Indicates the first The query status after the layer aggregates global information. This means converting similarity into attention weights. This represents the dimension of the key vector. Indicates matrix transpose; Based on the query status after aggregating global information, the self-attention unit interactively deduplicates multiple query statuses after aggregating global information to obtain the deduplicated query status. Based on the deduplicated query state, the feedforward network is used to perform a nonlinear transformation on the deduplicated query state to obtain an enhanced query state. After iterative refinement through multiple attention layers, the final attention layer outputs the final enhanced query status. Based on the association between points in the forest 3D point cloud and the query state, the final enhanced query state is fed back and fused into the corresponding point's point-level features to obtain the enhanced point-level features of the point: ; in, This indicates enhanced point-level features. Indicates The final enhanced query state after iterative refinement by the attention layer. This indicates a feature fusion operation.

[0010] Optionally, the multi-task segmentation head includes a semantic classification branch and an offset regression branch; The step involves performing semantic classification and offset regression using the multi-task segmentation head of the single-tree segmentation model based on the 3D spatial coordinates of each point in the forest 3D point cloud and the enhanced point-level features, to obtain the semantic category prediction result and the 3D offset vector pointing to the instance center for each point. Specifically, this includes: Based on the three-dimensional spatial coordinates corresponding to each point in the forest three-dimensional point cloud and the enhanced point-level features, semantic classification is performed through the semantic classification branch to obtain the probability that each point belongs to a tree or the background, and based on the probability, the semantic category prediction result of each point in the forest three-dimensional point cloud is obtained, wherein the semantic category prediction result includes trees and background; Based on the three-dimensional spatial coordinates corresponding to the points in the forest three-dimensional point cloud whose semantic category prediction result is trees and the enhanced point-level features, offset regression is performed through the offset regression branch to obtain the three-dimensional offset vector of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees pointing to the center of the instance to which the point belongs.

[0011] Optionally, the step of performing coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm on each point according to the semantic category prediction result of each point and the three-dimensional offset vector pointing to the instance center to obtain the single-tree instance segmentation result specifically includes: Points in the forest 3D point cloud whose semantic category prediction results are in the background are filtered out; The three-dimensional spatial coordinates of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees are added to the three-dimensional offset vector to obtain the corrected coordinates of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees. Using the corrected coordinates as the clustering basis, the HDBSCAN density clustering algorithm is used to aggregate points belonging to the same tree into the same instance cluster, and an instance number is assigned to each point to obtain the single tree instance segmentation result.

[0012] To achieve the above-mentioned objectives, the present invention also provides a 3D point cloud single-tree segmentation system based on jump-connection gating and query attention, wherein the 3D point cloud single-tree segmentation system based on jump-connection gating and query attention includes: The training module is used to construct an initial single-tree segmentation model and train the initial single-tree segmentation model using a multi-task collaborative loss function to obtain a single-tree segmentation model. The data acquisition module is used to acquire the initial forest 3D point cloud and preprocess the initial forest 3D point cloud to obtain the forest 3D point cloud. The point-level feature acquisition module is used to extract features from the forest 3D point cloud through the improved sparse 3D U-Net of the single tree segmentation model, perform feature fusion based on the jump-connected dual-gated fusion module, and perform feature decoding mapping to obtain point-level features; The feature enhancement module is used to enhance the point-level features through the query attention module of the single-tree segmentation model to obtain enhanced point-level features; The prediction module is used to perform semantic classification and offset regression based on the three-dimensional spatial coordinates of each point in the forest three-dimensional point cloud and the enhanced point-level features, through the multi-task segmentation head of the single tree segmentation model, to obtain the semantic category prediction result of each point and the three-dimensional offset vector pointing to the instance center. The single-tree segmentation module is used to perform coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm for each point according to the semantic category prediction result of each point and the three-dimensional offset vector pointing to the instance center, so as to obtain the single-tree instance segmentation result.

[0013] To achieve the above-mentioned objectives, the present invention also provides a terminal, the terminal comprising: a memory, a processor, and a 3D point cloud single-tree segmentation program based on jump-connection gating and query attention stored in the memory and executable on the processor, wherein when the 3D point cloud single-tree segmentation program based on jump-connection gating and query attention is executed by the processor, the steps of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described above are implemented.

[0014] To achieve the above-mentioned objectives, the present invention also provides a computer-readable storage medium storing a 3D point cloud single-tree segmentation program based on jump-connection gating and query attention. When the 3D point cloud single-tree segmentation program based on jump-connection gating and query attention is executed by a processor, it implements the steps of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described above.

[0015] In this invention, an initial single-tree segmentation model is constructed and trained using a multi-task collaborative loss function to obtain a single-tree segmentation model; an initial forest 3D point cloud is acquired and preprocessed to obtain a forest 3D point cloud; through an improved sparse 3D U-Net of the single-tree segmentation model, feature extraction, feature fusion based on a jump-connected dual-gated fusion module, and feature decoding mapping are performed on the forest 3D point cloud to obtain point-level features; through the query attention module of the single-tree segmentation model, feature enhancement is performed on the point-level features to obtain enhanced point-level features; based on the 3D spatial coordinates of each point in the forest 3D point cloud and the enhanced point-level features, semantic classification and offset regression are performed through the multi-task segmentation head of the single-tree segmentation model to obtain the semantic category prediction result and the 3D offset vector pointing to the instance center for each point; based on the semantic category prediction result and the 3D offset vector pointing to the instance center for each point, coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm are performed on each point to obtain the single-tree instance segmentation result. This invention overcomes the limitations of local receptive fields by introducing a query-based attention mechanism, enabling perception of the global structure of trees. By introducing a jump-connection dual-gated fusion module for single-tree segmentation at the jump connections of the sparse 3D U-Net network, it adaptively filters encoder features, effectively suppressing shallow noise and irrelevant details, while significantly enhancing the network's feature representation of high-value regions such as bifurcation points, canopy overlap areas, and instance boundaries. Furthermore, by designing a multi-task collaborative loss function that integrates Focal Loss, Smooth L1, and instance consistency constraints, it strengthens the learning of boundary-difficult samples and improves the coherence and robustness of predictions within instances. Compared with existing technologies, this invention effectively overcomes the problems of "undersegmentation" and "oversegmentation" caused by severe canopy overlap in complex forest stands. Through adaptive gating and global attention, it significantly improves the effectiveness of features and the discriminative power of the model in key areas. The final segmentation results have more accurate geometric shapes and clearer contours, that is, the final segmentation results have higher accuracy, providing a higher quality data foundation for subsequent forestry parameter (such as diameter at breast height, crown width, and biomass) estimation. Attached Figure Description

[0016] Figure 1 This is a flowchart of a preferred embodiment of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention of the present invention; Figure 2 This is another flowchart of a preferred embodiment of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention of the present invention; Figure 3This is a structural diagram of a preferred embodiment of the 3D point cloud single-tree segmentation system based on jump-connection gating and query attention of the present invention; Figure 4 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0018] In recent years, with the increasing demand for smart forestry, ecological monitoring, and refined forest resource management, LiDAR-based 3D forest point cloud data acquisition technology has been widely applied. LiDAR can quickly acquire large-scale, high-density, and high-precision forest spatial structure information, providing an important data foundation for forest structure parameter extraction, biomass estimation, carbon storage assessment, and forest management decisions. Accurately segmenting individual tree instances from the 3D forest point cloud is a crucial preliminary step for extracting parameters such as tree height, crown width, diameter at breast height (DBH), and number of trees. Existing single-tree segmentation methods mainly include traditional methods based on manual rules and data-driven methods based on deep learning. Traditional methods typically rely on the geometric features, spatial neighborhood relationships, or clustering rules of the point cloud for segmentation, such as segmentation methods based on height models, Euclidean clustering, region growing, or graph structures. These methods can achieve certain results in scenarios with simple forest stand structures and clear canopy separation. However, in complex forest environments with high canopy closure, severe canopy overlap, and significant understory vegetation disturbance, they are prone to problems such as incorrect merging of adjacent trees and fragmentation of individual trees, exhibiting drawbacks such as parameter sensitivity, weak generalization ability, and insufficient robustness. Deep learning methods have been gradually introduced into 3D point cloud single-tree segmentation tasks. Existing methods typically use point-based networks to directly process unstructured point clouds, or use voxel-based networks to map point clouds to 3D voxel meshes for feature extraction. Compared to traditional rule-based methods, deep learning methods can automatically learn the spatial features of point clouds, improving the automation level and adaptability to complex scenarios of single-tree segmentation to some extent. However, although these data-driven methods significantly improve the automation level and accuracy of segmentation, their underlying design paradigms still have inherent limitations due to mutual constraints in the field of single-tree segmentation.

[0019] First, existing sparse 3D U-Nets directly stitch features together at skip connections, easily introducing shallow noise and weakening the representation of high-value regions. Second, current mainstream methods generally use fixed-size voxels to discretize point clouds, a strategy fundamentally contradicting the inherent uneven density of tree point clouds (i.e., dense trunks and sparse canopies). Larger voxel sizes blur the fine geometry of the canopy tips, leading to the loss of crucial information; while smaller voxel sizes generate massive amounts of redundant data in dense areas like the trunk, resulting in huge computational and storage overhead. Standard 3D convolutional networks homogenize all spatial locations and feature channels when processing feature maps, failing to distinguish their contribution to the segmentation task. For objects with extremely complex morphology like trees, their bifurcation points, branch connections, and canopy boundaries intersecting with other trees clearly contain more crucial information for distinguishing different instances than the smooth trunk surface. Existing models lack a dynamic attention mechanism to guide the network to focus limited computational resources on these structurally complex, "high-information" regions, and cannot adaptively enhance feature channels that are more discriminative in distinguishing wood from leaves. This "feature blindness" makes it difficult for the model to learn sufficiently robust feature representations, especially prone to confusion and misclassification in complex scenes. Finally, mainstream segmentation loss functions, such as cross-entropy or Dessian loss, are essentially based on the measurement of region overlap. These loss functions tend to optimize the large proportion of the main body region, while being insensitive to errors at the few boundary points. In forest scenes, the number of boundary points defining tree outlines is far fewer than the number of interior points, resulting in a lack of strong, explicit supervision signals for boundary accuracy during model training. Therefore, although the segmented main body region may be macroscopically correct, its boundaries often appear blurry, smooth, or jagged, which seriously affects the accuracy of subsequent estimations of key forestry parameters such as diameter at breast height (DBH) and crown width.

[0020] To address the aforementioned technical problems, this invention provides a 3D point cloud tree segmentation method based on jump-connection gating and query attention. The method involves constructing an initial tree segmentation model and training it using a multi-task collaborative loss function. An initial forest 3D point cloud is then acquired and preprocessed to obtain a forest 3D point cloud. Using an improved sparse 3D U-Net based on the tree segmentation model, feature extraction, feature fusion based on a jump-connection dual-gating fusion module, and feature decoding mapping are performed on the forest 3D point cloud to obtain point-level features. The query attention module of the tree segmentation model is used to enhance these point-level features, resulting in enhanced point-level features. Based on the 3D spatial coordinates of each point in the forest 3D point cloud and the enhanced point-level features, semantic classification and offset regression are performed using the multi-task segmentation head of the tree segmentation model to obtain the semantic category prediction result and a 3D offset vector pointing to the instance center for each point. Finally, based on the semantic category prediction result and the 3D offset vector pointing to the instance center for each point, coordinate offset correction and HDBSCAN (Hierarchical Density-Based Spatial) algorithm are applied to each point. Clustering of Applications with Noise (density-based hierarchical spatial clustering algorithm) is used to obtain single-tree instance segmentation results through density clustering. This invention overcomes the limitations of local receptive fields by introducing a query-based attention mechanism to achieve perception of the global structure of trees; by introducing a jump-connection dual-gated fusion module for single-tree segmentation at the jump connections of the sparse 3D U-Net network, the encoder features are adaptively filtered, effectively suppressing shallow noise and irrelevant details, while significantly enhancing the network's feature representation of high-value regions such as bifurcation points, canopy overlap areas, and instance boundaries; by designing a multi-task collaborative loss function that integrates Focal Loss, Smooth L1, and instance consistency constraints, the learning of boundary-difficult samples is strengthened, and the coherence and regression robustness of intra-instance predictions are improved. Compared with existing technologies, this invention effectively overcomes the problems of "undersegmentation" and "oversegmentation" caused by severe canopy overlap in complex forest stands. Through adaptive gating and global attention, it significantly improves the effectiveness of features and the discriminative power of the model in key areas. The final segmentation results have more accurate geometric shapes and clearer contours, that is, the final segmentation results have higher accuracy, providing a higher quality data foundation for subsequent forestry parameter (such as diameter at breast height, crown width, and biomass) estimation.

[0021] The application content will be further explained below with reference to the accompanying drawings and the description of the embodiments.

[0022] A preferred embodiment of the present invention is a 3D point cloud single-tree segmentation method based on jump-gating and query attention, such as... Figure 1 As shown, it specifically includes: S1. Construct an initial single-tree segmentation model and train the initial single-tree segmentation model using a multi-task collaborative loss function to obtain a single-tree segmentation model.

[0023] In one implementation of this embodiment, the step of constructing an initial single-tree segmentation model and training the initial single-tree segmentation model using a multi-task collaborative loss function to obtain a single-tree segmentation model specifically includes: Construct an initial single-tree segmentation model, wherein the initial single-tree segmentation model includes an initial improved sparse 3D U-Net, an initial query attention module, and an initial multi-task segmentation head, and the initial improved sparse 3D U-Net includes an initial jump connection dual-gated fusion module set at the jump connection; We construct a semantic classification loss based on Focal Loss, a bias loss based on Smooth L1 loss, and an instance consistency loss based on instance consistency constraints. Combining these three losses, we obtain a multi-task collaborative loss function: ; in, For multi-task collaborative loss function, For semantic classification loss, For offset loss, For instance consistency loss, For semantic classification loss weights, As the weight for offset loss, Weight the instance consistency loss; The error between the output of the initial single-tree segmentation model and the true label is calculated using a multi-task collaborative loss function. The parameters of the initial single-tree segmentation model are then iteratively optimized using a backpropagation algorithm until the initial single-tree segmentation model converges, resulting in the single-tree segmentation model: The single-tree segmentation model includes an improved sparse 3D U-Net, a query attention module, and a multi-task segmentation head. The improved sparse 3D U-Net includes a jump-connection dual-gated fusion module set at the jump connection.

[0024] Specifically, training data is acquired, including 3D point clouds of the forest and their corresponding ground truth labels / classes. Based on this training data, the initial tree segmentation model is trained using a multi-task collaborative loss function to obtain the tree segmentation model. The tree segmentation model includes an improved sparse 3D U-Net, a query attention module, and a multi-task segmentation head. The improved sparse 3D U-Net, based on the existing sparse 3D U-Net, introduces a skip-connection dual-gated fusion module at skip connections to adaptively filter encoder features, effectively suppressing shallow noise and irrelevant details, while significantly enhancing the network's feature representation of high-value regions such as bifurcation points, canopy overlap areas, and instance boundaries. The query attention module, based on a query-based attention mechanism, helps overcome the limitations of local receptive fields and achieve perception of the global tree structure. The multi-task segmentation head is used for semantic classification prediction and offset regression prediction.

[0025] To effectively handle difficult-to-separate samples in forest point cloud segmentation, improve the robustness of offset prediction, and establish instance-level global constraints during training, this invention designs a collaborative loss function. Traditional cross-entropy loss treats all samples equally, leading to the model overemphasizing easily separable samples (such as the central region of the tree trunk) while neglecting difficult-to-predict points such as the tree crown boundary. Therefore, this invention uses Focal Loss instead of standard cross-entropy to construct a semantic classification loss function. : ; in, Let γ be the model's predicted probability of the true class, and γ be the focusing parameter. This is the class balance factor. When γ > 0, easily separable samples ( Loss weights that approach 1 Samples approaching 0 are significantly compressed; conversely, difficult-to-distinguish samples ( The loss weights (which are relatively small) are kept close to 1. This allows the training process to automatically focus on regions with low prediction confidence, such as canopy boundaries and canopy overlaps, enabling end-to-end hard sample mining without manual annotation.

[0026] In the task of predicting the offset of tree instance centers, the distance from the crown edge points to the center is relatively large, and the annotations are ambiguous. Traditional L2 loss excessively penalizes large errors and is easily dominated by outliers and annotation noise. This invention uses SmoothL1 loss to improve robustness and constructs an offset loss function. : ; in, β is the smoothing threshold used to calculate the difference between the predicted offset and the true offset. express L2 norm. Small error region ( L2 loss is used to ensure accurate convergence; large error region ( The loss degenerates into L1 loss, and the linearly increasing gradient suppresses the impact of large error points. This piecewise design allows the model to accurately optimize the offset prediction of normal points while tolerating the ambiguity of canopy edge annotations and outliers. The offset loss is calculated only for samples whose semantic label is tree.

[0027] While point-query association matrices can incorporate global instance information into predictions, they lack explicit instance-level consistency constraints. Therefore, this invention uses instance consistency loss. Make full use of query state in query-based attention mechanisms: ; in, For the first A set of points within a real tree instance The total number of real tree instances. For the first Point-query association distribution of all points within a real tree instance, For the first The average association distribution of real tree instances for Divergence.

[0028] Therefore, the overall loss function (i.e., the multi-task collaborative loss function) ): .

[0029] S2. Obtain the initial forest 3D point cloud and preprocess the initial forest 3D point cloud to obtain the forest 3D point cloud.

[0030] Specifically, the original forest scene point cloud is acquired, and the original data is normalized, segmented (sliding window slicing), and data augmentation is performed to convert the disordered large-scale point cloud into a standardized input format suitable for deep neural network processing, laying the data foundation for subsequent feature extraction.

[0031] S3. Using the improved sparse 3D U-Net of the single-tree segmentation model, feature extraction, feature fusion based on the jump-connected dual-gated fusion module, and feature decoding mapping are performed on the forest 3D point cloud to obtain point-level features.

[0032] In one implementation of this embodiment, the improved sparse 3D U-Net further includes an encoder and a decoder; The improved sparse 3D U-Net of the single-tree segmentation model is used to extract features from the forest 3D point cloud, perform feature fusion based on the jump-connected dual-gated fusion module, and perform feature decoding mapping to obtain point-level features, specifically including: Multi-scale feature extraction is performed on the forest 3D point cloud using the sparse convolutional layer and residual block of the encoder to obtain encoder features; The encoder features are progressively upsampled through the sparse deconvolution layer of the decoder to obtain the decoder features; The jump-connect dual-gated fusion module performs feature fusion on the encoder features and the decoder features based on encoder feature adaptive filtering and channel response adaptive enhancement to obtain gated fusion features. The gated fusion features are upsampled and mapped using the remaining decoding layers and output mapping layers in the decoder to obtain point-level features corresponding to points in the forest 3D point cloud: ; in, Indicates gating fusion characteristics, This indicates upsampling and mapping operations. This represents point-level features.

[0033] In one implementation of this embodiment, the jump-connect dual-gated fusion module includes a spatial gated branch and a channel gated branch; The method of using the jump-connect dual-gated fusion module to perform feature fusion on the encoder features and decoder features based on encoder feature adaptive filtering and channel response adaptive enhancement to obtain gated fusion features specifically includes: The encoder features and the decoder features are concatenated for the first time to obtain the first intermediate fused features: ; in, Indicates the first The first intermediate fusion feature of the layer, Indicates the first Encoder features of the layer Indicates the first Decoder features of the layer; The first intermediate fusion feature is input into the spatial gating branch, and the spatial gating branch generates jump selection weights based on the first intermediate fusion feature through a spatial perception multilayer perceptron: ; in, This indicates the weight for selecting the jump connection. This represents the Sigmoid activation function, used to map weights to the (0, 1) interval; This represents a spatial perception multilayer perceptron, consisting of linear layers, batch normalization layers, and activation functions. The spatial gating branch uses the jump connection selection weights to adaptively weight the encoder features point-by-point and channel-by-channel, resulting in enhanced encoder features: ; in, For the first Enhanced encoder features after spatial gating branch filtering; The enhanced encoder features and the decoder features are concatenated a second time to obtain the second intermediate fused features: ; in, Indicates the first The second intermediate fusion feature of the layer; The second intermediate fusion feature is input into the channel gating branch, and the channel gating branch performs global pooling on the second intermediate fusion feature to obtain the channel description vector; The channel gating branch generates channel recalibration weights based on the channel description vector using a channel-aware multilayer perceptron. ; in, This indicates that the channel recalibrates its weights. and Let represent the first and second learnable weight matrices in the channel-aware multilayer perceptron, respectively. Represents a non-linear activation function. Represents the channel description vector; The channel-gated branch uses the channel recalibration weights to dynamically recalibrate the channel dimension of the second intermediate fused feature, thereby obtaining the gated fused feature: ; in, Indicates the first Layer-gated fusion features.

[0034] Specifically, the Sparse 3D U-Net, used as the backbone network of this invention, adopts a classic encoder-decoder architecture. The encoder consists of a series of sparse convolutional layers and residual blocks, progressively reducing the spatial resolution of the point cloud and expanding the receptive field through stride convolution operations; the decoder progressively restores the spatial resolution through sparse inverse convolution. At various scales of the network, the encoder is primarily responsible for extracting low-level geometric features containing rich details, while the decoder outputs high-level semantic context features containing global information. The two are fused across multiple scales via skip connections. Existing skip connections typically fuse encoder and decoder features through direct concatenation, which easily introduces shallow noise.

[0035] To this end, this invention innovatively introduces a jump-connection dual-gated fusion module at the jump connection point. For the first... Encoder features of the layer and decoder features This module first concatenates the two to obtain the fused input. Subsequently, Input spatial gating branch, calculate the jump connection selection weights Next, the calculated weights are used. Features of the original encoder Adaptive weighting is performed point-by-point and channel-by-channel to filter spatial details and suppress noise, resulting in enhanced encoder features after spatial gating. Then, the filtered encoder features With decoder features The data is then concatenated again and input into the channel gating branch. In the channel gating branch, the concatenated non-empty voxel feature matrix (i.e., the concatenated enhanced encoder features) is first processed. With decoder features Perform global pooling (such as global average pooling or global max pooling) on ​​the non-empty voxel dimension to aggregate global context information and obtain the channel description vector. : ; ; in, This indicates a global pooling operation. Indicates the number of non-empty voxels. Represents a non-empty voxel characteristic matrix. Indicates the number of feature channels. Represents a matrix.

[0036] Then, channel recalibration weights are generated using a channel-sensing multilayer perceptron. The final gating fusion feature By analyzing splicing features This is obtained through dynamic recalibration of the channel dimension.

[0037] The skip-connect dual-gated fusion module obtains gated fusion features at the decoder skip connection. These features continue to pass through subsequent decoding layers and output mapping layers, transforming them into final point-level features corresponding to points or non-empty voxels. ,thus, It includes both local geometric details at the encoding end and high-level semantic information at the decoding end, as well as the key region responses filtered by the skip-connected dual-gated fusion module. It should be noted that in this embodiment, the method of the present invention can be directly applied to the forest 3D point cloud, where the processing object is the points in the forest 3D point cloud. Alternatively, the forest 3D point cloud can be mapped to a 3D voxel mesh first, and then the method of the present invention can be applied, where the processing object is a non-empty voxel. Whether it is a point or a non-empty voxel, the method flow of the present invention is consistent; therefore, points and non-empty voxels can be understood as equivalent.

[0038] Through the aforementioned dual-gating mechanism, not only is dynamic spatial filtering of shallow encoder features achieved, effectively suppressing the transmission of point cloud noise and irrelevant background details, but channel recalibration also enhances the more discriminative channel responses in multi-scale fusion features for segmentation tasks. This mechanism enables the network to adaptively enhance feature representation of key regions such as tree bifurcation points, canopy overlap areas, and instance boundaries, thereby significantly improving the segmentation accuracy of individual trees in complex forest stand scenarios.

[0039] S4. The point-level features are enhanced by the query attention module of the single-tree segmentation model to obtain enhanced point-level features.

[0040] In one implementation of this embodiment, the query attention module includes multiple attention layers, and the attention layers include cross attention units, self attention units, and feedforward networks; The step of enhancing the point-level features through the query attention module of the single-tree segmentation model to obtain enhanced point-level features specifically includes: The three-dimensional spatial coordinates of points in the forest three-dimensional point cloud are processed by position encoding to generate the position codes of the points: ; in, Represents the three-dimensional spatial coordinates of a point. The position coding function can be represented by sinusoidal position coding or multilayer perceptron mapping. This represents the location code generated from the coordinates, used to incorporate spatial location information into attention calculations; A learnable query state is introduced. Based on the query state, the location encoding of the point, and the point-level features, attention is calculated through the cross-attention unit to obtain the query state after aggregating global information. ; ; in, Indicates by the first Layer query status The query vector obtained from the mapping This represents the key vector obtained from the point-level feature mapping after fusion of positional encoding. This represents the value vector obtained from point-level feature mapping. , and These represent the learnable linear projection matrices corresponding to the query vector, key vector, and value vector, respectively. This means combining point-level features with spatial location encoding, so that attention computation considers both feature similarity and spatial location information simultaneously. Indicates the first The query state after the aggregation of global information in the layer (i.e., the query state after cross-attention aggregation). This represents the similarity matrix between query vectors and key vectors, reflecting the strength of the association between each query and each point. This means converting similarity into attention weights. This represents the dimension of the key vector in a single attention head. Used to scale the dot product result to prevent excessively large values ​​from causing... Gradient instability Represents matrix transpose; value vector This represents the point-level feature information that has been weighted and aggregated. Points with higher weights can be regarded as potential tree instance regions of interest to the corresponding query. Based on the query status after aggregating global information, the self-attention unit interactively deduplicates multiple query statuses after aggregating global information to obtain the deduplicated query status. Based on the deduplicated query state, the feedforward network is used to perform a nonlinear transformation on the deduplicated query state to obtain an enhanced query state. After iterative refinement through multiple attention layers, the final attention layer outputs the final enhanced query status. Based on the association between points in the forest 3D point cloud and the query state, the final enhanced query state is fed back and fused into the corresponding point's point-level features to obtain the enhanced point-level features of the point: ; in, This represents the enhanced point-level features after fusing global instance semantic information. Indicates The final enhanced query state after iterative refinement by the attention layer. This indicates a feature fusion operation based on point-query relationships.

[0041] Specifically, to overcome the limitations of local receptive fields in traditional point cloud segmentation methods and enable the model to accurately distinguish point clouds belonging to different trees in the overlapping canopy region within a global view, this invention employs an attention mechanism based on learnable query states. (Introduction) Learnable query status The slots represent potential tree instances. This represents the initial learnable query state matrix. This represents the number of query states, which can also be understood as the number of potential tree instances representing slots. This represents the feature dimension of each query state. Unlike traditional methods that rely solely on the features of the point cloud itself for local modeling, this method learns query states as network parameters to participate in end-to-end training, enabling the automatic learning of abstract representations of tree instances. Initially, these query states are randomly initialized using a Gaussian distribution; as training progresses, they gradually learn global semantic patterns from different tree instances or instance regions.

[0042] The query attention module designed in this invention adopts an iterative refinement strategy and stacks... The attention module (i.e., attention layer). Let the [number]th [unit] be an attention module (i.e. The query status of the layer is ,in Indicates the number of stacked attention modules. This indicates the current attention layer index. Each attention module mainly consists of three parts: cross-attention (i.e., cross-attention unit), self-attention (i.e., self-attention unit), and a feedforward network. Cross-attention is used to make the query state transition from point-level features. The system aggregates global point cloud information. After aggregating the global point cloud information, the self-attention unit further establishes interactions between query states, enabling different instance queries to compete and cooperate with each other, reducing the occurrence of multiple queries focusing on the same tree instance. Subsequently, the feedforward network performs a nonlinear transformation on the query states to enhance their expressive power. This iterative process can be summarized as follows: ; ; ; in, This indicates the query status after cross-attention and residual normalization. This represents the query state after self-attention and residual normalization. Indicates the first The query status output by the layer; Presentation layer normalization operation, This represents the cross-attention operation, used to aggregate global contextual information from point-level features; This represents a multi-head self-attention operation, used to establish interaction relationships between different query states. This represents a feedforward network used to enhance the non-linear expressive power of the query state. The addition in the above formula represents residual connections, which helps maintain stable feature propagation and alleviate the degradation problem during training of deep networks. The feedforward network can be further represented as: ; in, This indicates the query status of the input feedforward network. and Let these represent the first and second learnable weight matrices in the feedforward network. This represents a non-linear activation function. The feedforward network is used to perform a non-linear transformation on the query state after attention updates, further enhancing the expressive power of the query state.

[0043] For large-scale point clouds in forest scenes, if direct calculation is performed... The attention matrix may result in high memory overhead. Indicates the number of query states. Indicates the number of points or non-empty voxels. This means that the association weight between each query and each point needs to be calculated. Therefore, when there are many points, a block processing strategy can be adopted to divide the point cloud into several batches, calculate the attention relationship between each batch and the query state, and then accumulate or merge the results to control the peak memory usage.

[0044] go through After layer-by-layer iterative refinement, the enhanced query state is finally obtained. This contains global semantic information about potential tree instances in the scene. Subsequently, the query state will be ultimately enhanced. The point-query relationships are fed back and fused into the point-level features to obtain globally enhanced point-level features. The final result This will be used as input to the subsequent multi-task segmentation head to assist in semantic classification and point-to-instance center offset prediction, thereby improving the segmentation effect of single trees in the overlapping areas of tree canopies and at instance boundaries.

[0045] S5. Based on the three-dimensional spatial coordinates of each point in the forest three-dimensional point cloud and the enhanced point-level features, semantic classification and offset regression are performed through the multi-task segmentation head of the single-tree segmentation model to obtain the semantic category prediction result and the three-dimensional offset vector pointing to the instance center for each point.

[0046] In one implementation of this embodiment, the multi-task segmentation head includes a semantic classification branch and an offset regression branch; The step involves performing semantic classification and offset regression using the multi-task segmentation head of the single-tree segmentation model based on the 3D spatial coordinates of each point in the forest 3D point cloud and the enhanced point-level features, to obtain the semantic category prediction result and the 3D offset vector pointing to the instance center for each point. Specifically, this includes: Based on the three-dimensional spatial coordinates corresponding to each point in the forest three-dimensional point cloud and the enhanced point-level features, semantic classification is performed through the semantic classification branch to obtain the probability that each point belongs to a tree or the background, and based on the probability, the semantic category prediction result of each point in the forest three-dimensional point cloud is obtained, wherein the semantic category prediction result includes trees and background; Based on the three-dimensional spatial coordinates corresponding to the points in the forest three-dimensional point cloud whose semantic category prediction result is trees and the enhanced point-level features, offset regression is performed through the offset regression branch to obtain the three-dimensional offset vector of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees pointing to the center of the instance to which the point belongs.

[0047] Specifically, the input for this stage is the point-level features enhanced by the global attention module, and the corresponding 3D spatial coordinates of the points. The 3D spatial coordinates and enhanced point-level features are fed into the multi-task segmentation head for semantic classification and offset regression, respectively: the semantic classification branch predicts the probability that each point belongs to a tree or the background, and the offset regression branch predicts the 3D offset vector of each tree point pointing to the center of its instance.

[0048] S6. Based on the semantic category prediction results of each point and the three-dimensional offset vector pointing to the instance center, perform coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm on each point to obtain the single-tree instance segmentation results.

[0049] In one implementation of this embodiment, the step of performing coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm on each point according to the semantic category prediction result of each point and the three-dimensional offset vector pointing to the instance center to obtain the single-tree instance segmentation result specifically includes: Points in the forest 3D point cloud whose semantic category prediction results are in the background are filtered out; The three-dimensional spatial coordinates of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees are added to the three-dimensional offset vector to obtain the corrected coordinates of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees. Using the corrected coordinates as the clustering basis, the HDBSCAN density clustering algorithm is used to aggregate points belonging to the same tree into the same instance cluster, and an instance number is assigned to each point to obtain the single tree instance segmentation result.

[0050] Specifically, background points are filtered out based on the semantic prediction results, and the original coordinates of the tree points are added to the predicted offset vector to obtain the offset-corrected point coordinates. Using the corrected coordinates as the clustering basis, the HDBSCAN density clustering algorithm is used to aggregate points belonging to the same tree into the same instance cluster, and an instance number is assigned to each point, outputting the final single-tree instance segmentation result.

[0051] In summary, to address the problems of existing technologies in processing 3D point clouds in complex forest scenes, such as instance adhesion in overlapping canopy areas due to limited local receptive fields, the introduction of shallow noise and weakening of high-value region representation by directly stitching features at skip connections in existing sparse 3D U-Nets, and insufficient constraints on canopy boundaries and instance-level global consistency, this invention proposes a 3D point cloud single-tree segmentation method based on skip connection gating and query attention. Figure 2 As shown, its main processing flow includes the following five core steps: 1. Data Preprocessing: As the system's entry point, this module receives point clouds of the original forest scene as input. It normalizes, segments (sliding window slicing), and performs data augmentation operations on the raw data, transforming the disordered large-scale point cloud into a standardized input format suitable for deep neural network processing, laying the data foundation for subsequent feature extraction.

[0052] 2. Backbone Network: The preprocessed data is fed into a backbone network based on Sparse 3D U-Net to extract multi-scale deep features. This network includes an encoder and a decoder, and introduces a skip-connection gated fusion module (i.e., a skip-connection dual-gated fusion module) at the corresponding skip connections to adaptively filter the features at the encoding end, thereby effectively suppressing shallow noise and enhancing the feature representation of key regions such as overlapping canopies and bifurcation points.

[0053] 3. Global Feature Enhancement: The multi-scale features output by the backbone network are then enhanced by the query-based attention module. This module uses learnable query states to aggregate point cloud features from the entire scene, breaking the limitations of the local receptive field of traditional convolution and giving the model a global perspective, thereby effectively distinguishing adjacent tree individuals with similar local features in complex forest stands.

[0054] 4. Multi-task Segmentation Head and Post-processing Generation: The features refined by the global attention module are fed into the network's multi-task segmentation head. The multi-task segmentation head outputs a semantic category prediction (distinguishing between trees and background) and a 3D offset vector pointing to the instance center for each point in parallel. These outputs are then fed into the post-processing and instance generation module. After coordinate offset correction, points belonging to the same tree are grouped together using the HDBSCAN density clustering algorithm, forming the final high-precision single-tree instance segmentation result.

[0055] 5. Composite Loss Function Training: During the model training phase, the network's predicted output is compared with the manually labeled ground truth. This invention employs a multi-task collaborative loss function to calculate the error and optimizes the network parameters through backpropagation. This loss function enhances the learning of difficult-to-distinguish samples such as tree canopy boundaries through Focal Loss, improves the robustness of offset predictions using Smooth L1 loss, and introduces instance consistency constraints to ensure global consistency of predictions within the same tree.

[0056] In this invention, the jump-gated fusion mechanism provides cleaner multi-scale features, the query-based attention mechanism extracts instance representations with a more global perspective from these features, and the multi-task collaborative loss function provides better supervision signals to guide the learning of these representations.

[0057] Furthermore, based on the aforementioned method for segmenting individual trees in 3D point clouds using jump-connection gating and query attention, this invention also provides a system for segmenting individual trees in 3D point clouds using jump-connection gating and query attention. A preferred embodiment of the system for segmenting individual trees in 3D point clouds using jump-connection gating and query attention is as follows: Figure 3 As shown, it specifically includes: Training module 01 is used to construct an initial single-tree segmentation model and train the initial single-tree segmentation model using a multi-task collaborative loss function to obtain a single-tree segmentation model. Data acquisition module 02 is used to acquire an initial forest 3D point cloud and preprocess the initial forest 3D point cloud to obtain a forest 3D point cloud. The point-level feature acquisition module 03 is used to extract features from the forest three-dimensional point cloud through the improved sparse 3D U-Net of the single tree segmentation model, perform feature fusion based on the jump-connected dual-gated fusion module, and perform feature decoding mapping to obtain point-level features; Feature enhancement module 04 is used to enhance the point-level features through the query attention module of the single-tree segmentation model to obtain enhanced point-level features; Prediction module 05 is used to perform semantic classification and offset regression based on the three-dimensional spatial coordinates of each point in the forest three-dimensional point cloud and the enhanced point-level features, through the multi-task segmentation head of the single tree segmentation model, to obtain the semantic category prediction result of each point and the three-dimensional offset vector pointing to the instance center. The single-tree segmentation module 06 is used to perform coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm on each point according to the semantic category prediction result of each point and the three-dimensional offset vector pointing to the instance center, so as to obtain the single-tree instance segmentation result.

[0058] Furthermore, based on the aforementioned method and system for segmenting 3D point clouds into individual trees based on jump-gating and query attention, this invention also provides a terminal, wherein a preferred embodiment of the terminal is as follows: Figure 4 As shown, it specifically includes a processor 10, a memory 20, and a display 30. Figure 4 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0059] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, or flash card. Furthermore, the memory 20 may include both internal and external storage units. The memory 20 is used to store application software and various types of data installed on the terminal, such as the terminal's program code. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a 3D point cloud single-tree segmentation program 40 based on jump-connection gating and query attention. This 3D point cloud single-tree segmentation program 40 can be executed by the processor 10 to implement the steps of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention in this application.

[0060] In some embodiments, processor 10 may be a central processing unit (CPU), microprocessor or other data processing chip, used to run program code stored in memory 20 or process data, such as executing a 3D point cloud single-tree segmentation program 40 based on jump gating and query attention.

[0061] In some embodiments, display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. Display 30 is used to display information on the terminal and to display a visual user interface.

[0062] In one embodiment, when the processor 10 executes the 3D point cloud single-tree segmentation program 40 based on jump-connection gating and query attention stored in the memory 20, it implements the steps of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described above.

[0063] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a 3D point cloud single-tree segmentation program based on jump-connection gating and query attention, and when the 3D point cloud single-tree segmentation program based on jump-connection gating and query attention is executed by a processor, it implements the steps of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described above.

[0064] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0065] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0066] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for segmenting a single tree in a 3D point cloud based on jump-connection gating and query attention, characterized in that, The 3D point cloud single-tree segmentation method based on jump-connection gating and query attention includes: An initial single-tree segmentation model is constructed, and the initial single-tree segmentation model is trained using a multi-task collaborative loss function to obtain a single-tree segmentation model; An initial 3D point cloud of the forest is obtained, and the initial 3D point cloud of the forest is preprocessed to obtain a 3D point cloud of the forest. The improved sparse 3D U-Net of the single tree segmentation model is used to extract features from the forest 3D point cloud, perform feature fusion based on the jump-connected dual-gated fusion module, and perform feature decoding mapping to obtain point-level features. The point-level features are enhanced by the query attention module of the single-tree segmentation model to obtain enhanced point-level features. Based on the three-dimensional spatial coordinates of each point in the forest three-dimensional point cloud and the enhanced point-level features, semantic classification and offset regression are performed through the multi-task segmentation head of the single tree segmentation model to obtain the semantic category prediction result and the three-dimensional offset vector pointing to the instance center for each point. Based on the semantic category prediction results of each point and the three-dimensional offset vector pointing to the instance center, coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm are performed on each point to obtain the single-tree instance segmentation results. The improved sparse 3D U-Net includes an encoder and a decoder, wherein encoder features are obtained through the encoder and decoder features are obtained through the decoder. The jump-connect dual-gated fusion module includes a spatial gated branch and a channel gated branch; The jump-connect dual-gated fusion module performs feature fusion on the encoder features and the decoder features based on encoder feature adaptive filtering and channel response adaptive enhancement to obtain gated fusion features, specifically including: The encoder features and the decoder features are concatenated for the first time to obtain the first intermediate fused features: ; in, Indicates the first The first intermediate fusion feature of the layer, Indicates the first Encoder features of the layer Indicates the first Decoder features of the layer; The first intermediate fusion feature is input into the spatial gating branch, and the spatial gating branch generates jump selection weights based on the first intermediate fusion feature through a spatial perception multilayer perceptron: ; in, This indicates the weight for selecting the jump connection. This represents the Sigmoid activation function. This represents a multilayer perceptron for spatial perception. The spatial gating branch uses the jump connection selection weights to adaptively weight the encoder features point-by-point and channel-by-channel, resulting in enhanced encoder features: ; in, For the first Enhanced encoder features after spatial gating branch filtering; The enhanced encoder features and the decoder features are concatenated a second time to obtain the second intermediate fused features: ; in, Indicates the first The second intermediate fusion feature of the layer; The second intermediate fusion feature is input into the channel gating branch, and the channel gating branch performs global pooling on the second intermediate fusion feature to obtain the channel description vector; The channel gating branch generates channel recalibration weights based on the channel description vector using a channel-aware multilayer perceptron. ; in, This indicates that the channel recalibrates its weights. and Let represent the first and second learnable weight matrices in the channel-aware multilayer perceptron, respectively. Represents a non-linear activation function. Represents the channel description vector; The channel-gated branch uses the channel recalibration weights to dynamically recalibrate the channel dimension of the second intermediate fused feature, thereby obtaining the gated fused feature: ; in, Indicates the first Layer-gated fusion features.

2. The 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described in claim 1, characterized in that, The process of constructing an initial single-tree segmentation model and training it using a multi-task collaborative loss function to obtain a single-tree segmentation model specifically includes: Construct an initial single-tree segmentation model, wherein the initial single-tree segmentation model includes an initial improved sparse 3D U-Net, an initial query attention module, and an initial multi-task segmentation head, and the initial improved sparse 3D U-Net includes an initial jump connection dual-gated fusion module set at the jump connection; We construct a semantic classification loss based on Focal Loss, a bias loss based on Smooth L1 loss, and an instance consistency loss based on instance consistency constraints. Combining these three losses, we obtain a multi-task collaborative loss function: ; in, For multi-task collaborative loss function, For semantic classification loss, For offset loss, For instance consistency loss, For semantic classification loss weights, As the weight for offset loss, Weight the instance consistency loss; The error between the output of the initial single-tree segmentation model and the true label is calculated using a multi-task collaborative loss function. The parameters of the initial single-tree segmentation model are then iteratively optimized using a backpropagation algorithm until the initial single-tree segmentation model converges, resulting in the single-tree segmentation model: The single-tree segmentation model includes an improved sparse 3D U-Net, a query attention module, and a multi-task segmentation head. The improved sparse 3D U-Net includes a jump-connection dual-gated fusion module set at the jump connection.

3. The 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described in claim 2, characterized in that, The improved sparse 3D U-Net of the single-tree segmentation model is used to extract features from the forest 3D point cloud, perform feature fusion based on the jump-connected dual-gated fusion module, and perform feature decoding mapping to obtain point-level features, specifically including: Multi-scale feature extraction is performed on the forest 3D point cloud using the sparse convolutional layer and residual block of the encoder to obtain encoder features; The encoder features are progressively upsampled through the sparse deconvolution layer of the decoder to obtain the decoder features; The jump-connect dual-gated fusion module performs feature fusion on the encoder features and the decoder features based on encoder feature adaptive filtering and channel response adaptive enhancement to obtain gated fusion features. The gated fusion features are upsampled and mapped using the remaining decoding layers and output mapping layers in the decoder to obtain point-level features corresponding to points in the forest 3D point cloud: ; in, Indicates gating fusion characteristics, This indicates upsampling and mapping operations. This represents point-level features.

4. The 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described in claim 3, characterized in that, The query attention module includes multiple attention layers, and the attention layer includes cross attention units, self attention units, and feedforward networks. The step of enhancing the point-level features through the query attention module of the single-tree segmentation model to obtain enhanced point-level features specifically includes: The three-dimensional spatial coordinates of points in the forest three-dimensional point cloud are processed by position encoding to generate the position codes of the points: ; in, Represents the three-dimensional spatial coordinates of a point. This represents the position encoding function. Indicates positional encoding; A learnable query state is introduced. Based on the query state, the location encoding of the point, and the point-level features, attention is calculated through the cross-attention unit to obtain the query state after aggregating global information. ; ; in, Indicates by the first Layer query status The query vector obtained from the mapping This represents the key vector obtained from the point-level feature mapping after fusion of positional encoding. This represents the value vector obtained from point-level feature mapping. , and These represent the learnable linear projection matrices corresponding to the query vector, key vector, and value vector, respectively. Indicates the first The query status after the layer aggregates global information. This means converting similarity into attention weights. This represents the dimension of the key vector. Indicates matrix transpose; Based on the query status after aggregating global information, the self-attention unit interactively deduplicates multiple query statuses after aggregating global information to obtain the deduplicated query status. Based on the deduplicated query state, the feedforward network is used to perform a nonlinear transformation on the deduplicated query state to obtain an enhanced query state. After iterative refinement through multiple attention layers, the final attention layer outputs the final enhanced query status. Based on the association between points in the forest 3D point cloud and the query state, the final enhanced query state is fed back and fused into the corresponding point's point-level features to obtain the enhanced point-level features of the point: ; in, This indicates enhanced point-level features. Indicates The final enhanced query state after iterative refinement by the attention layer. This indicates a feature fusion operation.

5. The 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described in claim 1, characterized in that, The multi-task segmentation head includes a semantic classification branch and an offset regression branch; The step involves performing semantic classification and offset regression using the multi-task segmentation head of the single-tree segmentation model based on the 3D spatial coordinates of each point in the forest 3D point cloud and the enhanced point-level features, to obtain the semantic category prediction result and the 3D offset vector pointing to the instance center for each point. Specifically, this includes: Based on the three-dimensional spatial coordinates corresponding to each point in the forest three-dimensional point cloud and the enhanced point-level features, semantic classification is performed through the semantic classification branch to obtain the probability that each point belongs to a tree or the background, and based on the probability, the semantic category prediction result of each point in the forest three-dimensional point cloud is obtained, wherein the semantic category prediction result includes trees and background; Based on the three-dimensional spatial coordinates corresponding to the points in the forest three-dimensional point cloud whose semantic category prediction result is trees and the enhanced point-level features, offset regression is performed through the offset regression branch to obtain the three-dimensional offset vector of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees pointing to the center of the instance to which the point belongs.

6. The 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described in claim 5, characterized in that, The step of performing coordinate offset correction and HDBSCAN density clustering on each point based on the semantic category prediction result and the 3D offset vector pointing to the instance center to obtain the single-tree instance segmentation result specifically includes: Points in the forest 3D point cloud whose semantic category prediction results are in the background are filtered out; The three-dimensional spatial coordinates of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees are added to the three-dimensional offset vector to obtain the corrected coordinates of the points in the forest three-dimensional point cloud whose semantic category prediction result is trees. Using the corrected coordinates as the clustering basis, the HDBSCAN density clustering algorithm is used to aggregate points belonging to the same tree into the same instance cluster, and an instance number is assigned to each point to obtain the single tree instance segmentation result.

7. A 3D point cloud single-tree segmentation system based on jump-connection gating and query attention, wherein the 3D point cloud single-tree segmentation system based on jump-connection gating and query attention is applied to the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described in any one of claims 1-6, characterized in that, The 3D point cloud single-tree segmentation system based on jump-gating and query attention includes: The training module is used to construct an initial single-tree segmentation model and train the initial single-tree segmentation model using a multi-task collaborative loss function to obtain a single-tree segmentation model. The data acquisition module is used to acquire the initial forest 3D point cloud and preprocess the initial forest 3D point cloud to obtain the forest 3D point cloud. The point-level feature acquisition module is used to extract features from the forest 3D point cloud through the improved sparse 3D U-Net of the single tree segmentation model, perform feature fusion based on the jump-connected dual-gated fusion module, and perform feature decoding mapping to obtain point-level features; The feature enhancement module is used to enhance the point-level features through the query attention module of the single-tree segmentation model to obtain enhanced point-level features; The prediction module is used to perform semantic classification and offset regression based on the three-dimensional spatial coordinates of each point in the forest three-dimensional point cloud and the enhanced point-level features, through the multi-task segmentation head of the single tree segmentation model, to obtain the semantic category prediction result of each point and the three-dimensional offset vector pointing to the instance center. The single-tree segmentation module is used to perform coordinate offset correction and clustering based on the HDBSCAN density clustering algorithm for each point according to the semantic category prediction result of each point and the three-dimensional offset vector pointing to the instance center, so as to obtain the single-tree instance segmentation result.

8. A terminal, characterized in that, The terminal includes: a memory, a processor, and a 3D point cloud single-tree segmentation program based on jump-connection gating and query attention, stored in the memory and executable on the processor. When the 3D point cloud single-tree segmentation program based on jump-connection gating and query attention is executed by the processor, it implements the steps of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a 3D point cloud single-tree segmentation program based on jump-connection gating and query attention. When the 3D point cloud single-tree segmentation program based on jump-connection gating and query attention is executed by a processor, it implements the steps of the 3D point cloud single-tree segmentation method based on jump-connection gating and query attention as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Point cloud individual tree segmentation method and system based on collaborative attention and density adaptive voxelization, terminal and storage medium

    CN121095576A

  • Three-dimensional point cloud instance segmentation method for extracting single tree from forest scene

    CN121280718A