3D scene segmentation method and device based on layered multi-label converter

Through the hierarchical multi-marking converter method, the multi-scale features of 3D scenes are extracted using local and global attention modules, which solves the problem of high global attention calculation cost, and achieves efficient 3D scene segmentation performance improvement and memory optimization.

CN120339608APending Publication Date: 2025-07-18CHINA ACADEMY OF ELECTRONICS AND INFORMATION TECHNOLOGY OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510382522.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing 3D scene segmentation method based on transformers is difficult to effectively capture long-distance dependencies due to the high global attention calculation cost, and the resource consumption is large.

Method used

The hierarchical multi-label transformer method is adopted to aggregate local geometric and context information through the point embedding layer, multi-scale feature extraction is performed using the LFGS-Attention module, and features are propagated through the interpolation layer, combining local feature attention and global semantic attention modules to capture long-distance context dependence.

Benefits of technology

While reducing resource consumption, it improves the performance of 3D scene segmentation, achieves superior segmentation effect, and significantly reduces the memory consumption of global information capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339608A_ABST
    Figure CN120339608A_ABST
Patent Text Reader

Abstract

The invention provides a 3D scene segmentation method and device based on a layered multi-mark converter, and the method comprises the steps: employing a point embedding layer to aggregate the local geometric and context information of each point for an input point cloud; initial features corresponding to the initialization marks are transmitted and updated to obtain features of the point cloud and the multi-class marks; sampling the original point cloud for multiple times, and extracting corresponding multi-scale features; the multi-scale features extracted from the sampling points are propagated back to the original points by using an interpolation layer; in each interpolation layer, the interpolation features and the point features of the corresponding coding stage are connected through jump connection, and the output of the last interpolation layer is sent to a classifier layer to obtain a prediction label of each point; and performing 3D scene segmentation based on the prediction label. According to the method, superior segmentation performance is achieved on an existing data set, and compared with standard global attention of the same number of input points, memory consumption of multiple marks for capturing global information is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D vision technology, and in particular to a 3D scene segmentation method and device based on a hierarchical multi-label transformer. Background Art

[0002] Networks based on the transformer architecture have recently become the focus of research in the field of 3D vision. However, directly applying the transformer to 3D tasks, especially 3D scene segmentation, results in a huge computational cost due to the quadratic computational complexity of the input point number. Therefore, most existing transformer-based networks focus on aggregating local features but fail to capture long-range dependencies due to the high computational cost of global attention. Summary of the Invention

[0003] The technical problem to be solved by the present invention is how to improve the segmentation performance and reduce resource consumption. In view of this, the present invention provides a 3D scene segmentation method and device based on a hierarchical multi-label transformer.

[0004] The technical solution adopted by the present invention is a 3D scene segmentation method based on a hierarchical multi-label transformer, including: For the input point cloud, use a point embedding layer to aggregate the local geometry and context information of each point; Initialize multiple learnable labels as multi-class labels, and transfer the corresponding initial features to LFGS-Attention to obtain the updated features of the point cloud and multi-class labels; Sample the original point cloud multiple times, and use the LFGS-Attention module at each sampling layer to extract corresponding multi-scale features; Use an interpolation layer to propagate the multi-scale features extracted from the sampled points of the sampling layer back to the original points to obtain interpolation features; At each interpolation layer, connect the interpolation features with the point features in the corresponding encoding stage through skip connections, where the output of the last interpolation layer is fed into a classifier layer to obtain the predicted label of each point; Perform 3D scene segmentation based on the predicted labels.

[0005] In one embodiment, the multi-class labels aggregate point features from the sampled points of the sampling layer at each stage, thereby capturing long-range context dependencies.

[0006] In one embodiment, the LFGS-Attention module is configured to: normalize the features of the point cloud through layer normalization for a given input point cloud and its features; Input the obtained features into a local feature attention module and a global semantic attention module respectively; Add the outputs of the local feature attention module and the global semantic attention module, and obtain the fused features through a linear layer; After layer normalization and MLP, the output is added with a residual connection to obtain the final features.

[0007] In one embodiment, the local feature attention module is configured to learn local features through divided cube windows, as shown in the following formula; Where, MSA represents multi-head self-attention, and LN represents layer normalization.

[0008] In one embodiment, the global semantic attention module is configured as: Among multiple class tokens where C is the number of classes and D is the dimension of the feature map, apply two linear layers to the point cloud features to generate queries and values , apply another two linear layers to the multi-class token features to generate keys and values ; Calculate the global attention and through the matrix multiplication between , as shown in the following formula: Based on the attention , calculate the global class-specific features respectively through the matrix multiplication with , and calculate the assignment of the whole points through Gumbel-Softmax ; After assigning each point to the corresponding class token, merge the features of all points belonging to the same class to generate new features of multiple class tokens , as shown in the following formula: Where, represents soft assignment, represents hard assignment, is an independent and identically distributed random sample drawn from the Gumbel(0,1) distribution, is the temperature parameter, is the stop gradient operator, represents the number of points included in the multi-class token; Update the features of the point cloud by adding the global class-specific features to the residual connection, as shown in the following formula: .

[0009] On the other hand, the present invention also provides a 3D scene segmentation device based on a hierarchical multi-token transformer, including: An aggregation unit configured to, for the input point cloud, use a point embedding layer to aggregate the local geometry and context information of each point; An initialization unit configured to initialize a plurality of learnable tokens as multi-class tokens and pass the corresponding initial features to LFGS-Attention to obtain the updated features of the point cloud and the multi-class tokens; A multi-scale feature unit configured to sample the original point cloud multiple times and use the LFGS-Attention module to extract corresponding multi-scale features at each sampling layer; An interpolation feature unit configured to use an interpolation layer to propagate the multi-scale features extracted from the sampled points of the sampling layer back to the original points to obtain interpolation features; A prediction unit configured to, at each of the interpolation layers, connect the interpolation features with the point features of the corresponding encoding stage through skip connections, wherein the output of the last interpolation layer is fed into a classifier layer to obtain the prediction labels of each point; A scene segmentation unit configured to perform 3D scene segmentation based on the prediction labels.

[0010] On the other hand, the present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and the processor is configured to execute the computer program to implement the 3D scene segmentation method based on a hierarchical multi-token transformer as described in any one of the above.

[0011] On the other hand, the present invention also provides a computer storage medium storing a computer program, and the computer program is executed to implement the 3D scene segmentation method based on a hierarchical multi-token transformer as described in any one of the above.

[0012] Compared with the prior art, the present invention has at least the following advantages: The method provided by the present invention captures long-range dependencies at low cost by introducing a plurality of learnable tokens at different levels. It achieves superior segmentation performance on existing data sets, and compared with the standard global attention with the same number of input points, the memory consumption of the multi-tokens for capturing global information is significantly reduced. Description of the Drawings

[0013] Figure 1 It is a schematic flowchart of the 3D scene segmentation method based on a hierarchical multi-token transformer according to an embodiment of the present invention; Figure 2 It is a schematic diagram of the overall architecture according to an embodiment of the present invention; Figure 3Schematic diagram of the LFGS-Attention module according to an embodiment of the present invention; Figure 4 Schematic diagram of the sampling and interpolation layer according to an embodiment of the present invention; Figure 5 Schematic diagram of the visualization display of the semantic segmentation result on the S3DIS dataset according to an embodiment of the present invention; Figure 6 Schematic diagram of the real segmentation result, the segmentation result of the model of the present invention without the GSA module, and a segmentation result of the model of the present invention including the GSA module in a randomly selected scene; Figure 7 Schematic diagram of the real segmentation result, the segmentation result of the model of the present invention without the GSA module, and another segmentation result of the model of the present invention including the GSA module in a randomly selected scene; Figure 8 Schematic diagram of the composition of the 3D scene segmentation device based on the hierarchical multi-label transformer according to an embodiment of the present invention; Figure 9 Schematic diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners

[0014] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the present invention will be described in detail as follows in combination with the accompanying drawings and preferred embodiments.

[0015] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the ordinary understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms (such as those defined in a common dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the related art and will not be interpreted in an idealized or overly formal sense unless clearly defined herein.

[0016] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0017] An embodiment of the present invention, a 3D scene segmentation method based on a hierarchical multi-label transformer, as Figure 1 shown, includes: For the input point cloud, use a point embedding layer to aggregate the local geometry and context information of each point; Initialize multiple learnable tokens as multi-class tokens, and transfer the corresponding initial ones to LFGS-Attention to obtain the updated features of the point cloud and the multi-class tokens; Sample the original point cloud multiple times, and use the LFGS-Attention module at each sampling layer to extract corresponding multi-scale features; Use the interpolation layer to propagate the multi-scale features extracted from the sampled points of the sampling layer back to the original points to obtain interpolation features; At each interpolation layer, connect the interpolation features with the point features in the corresponding encoding stage through skip connections, where the output of the last interpolation layer is fed into the classifier layer to obtain the prediction label for each point; Perform 3D scene segmentation based on the prediction labels.

[0018] In one embodiment, multi-class labels aggregate point features from the sampled points of the sampling layer at each stage, thereby capturing long-range context dependencies.

[0019] In one embodiment, the LFGS-Attention module is configured to: for a given input point cloud and its features, normalize the features of the point cloud through layer normalization; Input the obtained features into the local feature attention module and the global semantic attention module respectively; Add the outputs of the local feature attention module and the global semantic attention module, and obtain the fused features through a linear layer; After layer normalization and MLP, the output is added with a residual connection to obtain the final features.

[0020] In one embodiment, the local feature attention module is configured to learn local features through divided cube windows, as shown in the following formula; where MSA represents multi-head self-attention and LN represents layer normalization.

[0021] In one embodiment, the global semantic attention module is configured to: Among multiple class labels where C is the number of classes and D is the dimension of the feature map, apply two linear layers to the point cloud features to generate queries and values , apply another two linear layers to the multi-class label features to generate keys and values ; Calculate the global attention through the matrix multiplication between and , as shown in the following formula: Based on the attention , respectively through Calculate the global class-specific features by matrix multiplication , and calculate the assignment of the entire point through Gumbel-Softmax ; After assigning each point to the corresponding class label, merge the features of all points belonging to the same class to generate new features for multiple class labels , specifically as follows: where, represents soft assignment, represents hard assignment, is an independent and identically distributed random sample drawn from the Gumbel(0,1) distribution, is the temperature parameter, is the stop gradient operator, represents the number of points included in the multi-class label; Update the features of the point cloud by adding the global class-specific features to the residual connection, as shown in the following formula: .

[0022] Next, the method provided in this embodiment will be described in detail in conjunction with the attached Figures 2 to 7 drawings.

[0023] A. Overall architecture The present invention proposes a transformer-based network for point cloud scene segmentation, named HMTT. Figure 1 Shows the overall architecture of the proposed HMTT.

[0024] For the input point cloud, the present invention first uses a point embedding layer to aggregate the local geometry and context information of each point.

[0025] Then, the present invention initializes multiple learnable labels as multi-class labels. Next, by passing their initial features to the LFGS-Attention module, the updated features of the point cloud and multi-class labels can be obtained.

[0026] In order to learn the multi-scale features of the point cloud, the present invention samples the original point cloud three times and uses the LFGS-Attention module to extract the corresponding features after each sampling layer.

[0027] In the decoding stage, the present invention propagates the features learned from the sampled points back to the original points according to the interpolation layer.

[0028] After each interpolation layer, the interpolated features are connected to the point features in the corresponding encoding stage through a skip connection.

[0029] Finally, the output of the last interpolation layer is fed into the classifier layer to obtain the predicted label for each point.

[0030] The final prediction is supervised by the cross-entropy loss function of the point ground truth label.

[0031] Meanwhile, multiple class tokens also update their features by aggregating corresponding features from the sampled point clouds at each stage, further enabling the network to capture long-range context dependencies.

[0032] To ensure that each token effectively learns the representation of the corresponding class, the present invention applies an average pooling layer to the final features of the tokens to obtain class scores, and then uses a multi-label soft margin loss function to supervise them with the scene-level ground truth labels.

[0033] B. LFGS-Attention Module Standard self-attention calculation needs to compute the global attention between one point and all other points, resulting in a quadratic complexity with respect to the number of points, which makes it difficult to directly apply to large-scale point cloud tasks such as 3D scene segmentation. To solve this problem, the proposed LFGS-Attention module adopts a Local Feature Attention (LFA) module and a Global Semantic Attention (GSA) module to learn local geometric features and global semantic features. On the one hand, the LFA module restricts self-attention within a locally partitioned window for efficient computation. On the other hand, the GSA module computes the global attention between the entire input points and a few class learnable tokens, which enables these class tokens to capture global class-specific information and long-range dependencies at a moderate computational cost.

[0034] Figure 2 FIG. shows the schematic diagram of the proposed overall HMTT architecture. The input point cloud first aggregates local geometric information through a point embedding layer and then is input into the proposed LFGS-Attention module for feature extraction. Subsequently, the learned features pass through three sampling layers, each followed by an LFGS-Attention module, to obtain multi-scale features. Finally, the encoded features are interpolated and fed into the classifier for prediction. Meanwhile, multi-class tokens aggregate point features from the sampled points at each stage, thereby capturing long-range context dependencies. The cross-entropy loss and the multi-label soft margin loss are jointly adopted to optimize the network through the supervision of the point ground truth labels and the scene-level ground truth labels.

[0035] The proposed LFGS-Attention module first normalizes the input features through layer normalization and then feeds the normalized features into its two core modules: the Local Feature Attention (LFA) module and the Global Semantic Attention (GSA) module. Next, the local and global features extracted by these two modules are added together and passed through a linear layer to obtain the fused features. Finally, the fused features pass through layer normalization and an MLP (Multi-Layer Perceptron), and a residual connection is added to the output to obtain the final features.

[0036] Please refer again to Figure 2 , the input point cloud first aggregates local geometric information through a point embedding layer and then is fed into the proposed LFGS-Attention module for feature extraction. Next, the learned features pass through three sampling layers, with an LFGS-Attention module following each sampling layer, to obtain multi-scale features. Finally, the encoded features are interpolated and fed into a classifier for prediction. Meanwhile, multi-class labels aggregate point features from the sampled points at each stage, thus capturing long-range context dependencies. The cross-entropy loss and the multi-label soft margin loss are adopted together to optimize the network through the supervision of point ground truth labels and scene-level ground truth labels.

[0037] Please refer to Figure 3 , the LFGS-Attention module proposed in this embodiment first normalizes the input features through layer normalization and then feeds the normalized features into the LFA module and the GSA module respectively. Next, the outputs of these two modules are added together and passed through a linear layer to obtain the fused features. Finally, the fused features pass through layer normalization and an MLP, and the residual connection is added to the output to obtain the final features.

[0038] Given an input point cloud and its features , where N represents the number of points and D represents the dimension of the feature map. The present invention first normalizes the feature F of the point cloud through layer normalization

[37] and then feeds the obtained features into the LFA module and the GSA module respectively. Next, the present invention adds the outputs of these two modules together and obtains the fused features through a linear layer . Finally, after layer normalization and an MLP, the residual connection is added to the output to obtain the final features . The present invention formulates all steps as: where LN represents the layer normalization operator.

[0039] Local Feature Attention (LFA) module: The Local Feature Attention (LFA) module aims to learn local features through divided cube windows. Specifically, the input cloud First, it is divided into non - overlapping cubic windows, and then multi - head self - attention is performed on each cubic window. In this way, the local features of neighbors within each point window are aggregated. In addition, inspired by Swin Transformer

[19] , the present invention further establishes cross - window connections based on a shifted window scheme. Different from moving windows on 2D regular images, which requires complex operations such as cyclic shift and inverse cyclic shift, moving windows on 3D irregular point clouds can be easily achieved by moving the point cloud. Therefore, the present invention moves the point cloud by half of the window size and re - divides the moved point cloud into new cubic windows. Next, the present invention performs self - attention again on each new cubic window. The entire calculation process is as shown in Figure 2 the upper - right corner and is formulated as: where MSA represents multi - head self - attention and LN represents layer normalization.

[0040] 2) Global Semantic Attention (GSA) module: The present invention proposes a Global Semantic Attention (GSA) module to capture global semantic features and long - range context dependencies. Figure 2 The lower - right corner of shows the process of the GSA module. First, the present invention introduces multiple class tokens and values , applies two linear layers to the point cloud features to generate queries and keys for the multi - class token features, and applies another two linear layers to generate keys and values . Next, the present invention calculates the global attention through the matrix multiplication between and . The last two steps can be formulated as: Based on the attention , the present invention calculates the global class - specific features through the matrix multiplication with respectively, and calculates the assignment of the whole point through Gumbel - Softmax. Here, the assignment is used to replace the one - hot assignment operation through argmax, which enables the GSA module to be differentiable and trainable in an end - to - end manner. After assigning each point to the corresponding class token, the present invention combines the features of all points belonging to the same class to generate new features of multiple class tokens. The present invention formulates the above process as: where Indicates soft assignment, Indicates hard assignment, is an independent and identically distributed random sample drawn from the Gumbel(0,1) distribution, is the temperature parameter, is the stop-gradient operator, represents the number of points included in the multi-class label. Finally, by adding the global class-specific feature to the residual connection to update the features of the point cloud, the detailed process can be expressed as: As can be seen from the present invention, the GSA module can enable each point to aggregate the global class features from multiple class labels. At the same time, the features of multiple class labels are updated by aggregating the local features from the corresponding points.

[0041] 3) Complexity analysis: In this section, the present invention compares the memory complexity of the proposed LFGS-Attention module with that of the standard global self-attention. For the features of the input point cloud , the complexity of the standard global self-attention is: For the LFA module, assuming that the input point cloud is divided into V windows, and each window contains an average of k points on average, the complexity of the LFA module is: The complexity of the GSA module is: To further reduce the computational cost, the present invention makes the weights of the two linear layers between the LFA module and the GSA module shared, and both of these two linear layers are applied to F to generate queries and values. Therefore, the final complexity of the proposed LFGS-Attention module is: It should be noted that k is the average number of points within the window , C is the number of classes in the scene . Therefore, the proposed module, while extracting effective local and global features, reduces the memory complexity from to .

[0042] Sampling and interpolation layers, for details, please refer to Figure 4 .

[0043] 1) Sampling layer: Given the input point cloud and its features , the present invention first uses the farthest point sampling (FPS) algorithm to obtain from a subset of points . Based on the indices in , the present invention can also obtain the corresponding features of . Then the present invention uses the K-nearest neighbor (kNN) algorithm to search for neighbors for each point in and groups the features of these neighbors. Next, the present invention performs layer normalization on the grouped features and inputs the normalized features into a linear layer. Finally, the present invention performs max pooling to aggregate the features of neighboring points to obtain .

[0044] 2) Interpolation layer: Given a point cloud , a subset and its features , the objective of the present invention is to obtain and the features of < ). To this end, the present invention first uses the K-nearest neighbor (kNN) algorithm to select neighbor points for each point in for and calculates the distance between each point and its neighbors. At the same time, the present invention inputs the features into a layer normalization layer and a linear layer. Then the present invention obtains the interpolation features according to the inverse distance weighting method. Next, the present invention passes the corresponding features in the encoding stage to a layer normalization layer and a linear layer. Finally, the present invention adds the interpolation features and the transformed features of to obtain the output features.

[0045] C. Loss function As introduced before in the present invention, two loss functions are adopted herein, namely the cross-entropy loss function of the point true label and the multi-label soft margin loss function of the scene-level true label, which can be expressed as: where represents the cross-entropy loss, represents the multi-label soft margin loss, and λ represents the balancing weight. The cross-entropy loss function is usually used to measure the classification and segmentation performance of the model, and the formula is: where is the predicted probability of class i, is the true label, and C is the number of categories. The multi-label soft margin loss function aims to measure the performance of the model in predicting multiple labels. In this paper, the present invention introduces multiple category tokens to hierarchically aggregate the global semantic features of the entire point cloud, which is applicable to this loss function. By supervising each token with the corresponding category label, the present invention can ensure the effective learning of category representations. The formula for the multi-label soft margin loss function is: where is the predicted probability of category i, is the true label, and C is the number of categories.

[0046] To further illustrate the performance improvement brought by this embodiment, the present invention conducts point cloud segmentation experiments on two popular large-scale scene datasets, namely S3DIS and ScanNetV2. The S3DIS dataset collects the point clouds of 271 rooms in 6 regions from 3 buildings. Each point is labeled as one of 13 category labels, including doors, tables, bookshelves, floors, etc. According to the common setting, the rooms in Area 5 are used for testing, and the other rooms are used for training. The ScanNetV2 dataset contains 1,513 scans from 706 different scenes, including 2.5 million RGB-D frames. Each scan is divided into 20 categories and labeled. The present invention uses 1,201 scans for training, 312 scans for validation, and 100 scans for testing according to the official division. The evaluation metrics use the mean intersection over union (mIoU) of each category, the mean accuracy (mAcc) of each category, and the overall point accuracy (OA).

[0047] The network architecture is as Figure 2 shown. The proposed HMTT receives the point cloud with xyz coordinates and rgb colors as input, and uses KPConv as the point embedding layer to aggregate the local features of each point. The sampling layer is implemented by the farthest point sampling algorithm, and the interpolation layer is performed based on inverse distance weighting. Considering the number of input points, the present invention adds an additional sampling layer between the point embedding layer and the LFGS module for the ScanNetV2 dataset. The downsampling ratio of S3DIS is set to 8, and that of ScanNetV2 is 4. The feature dimensions of the LFGS module for both datasets are set to [48, 96, 192, 384] at each stage. In addition, the number of multi-category tokens is set to 13 for S3DIS and 20 for ScanNetV2, corresponding to the number of categories.

[0048] The present invention implements HMTT on 4 Tesla V100 GPUs based on the Pytorch framework. The present invention uses the AdamW optimizer with a base learning rate of 0.006 and a weight decay of 0.05 for network training. For S3DIS, the input points of the room are first sampled with a grid size of 0.04 m, and the maximum number of points is set to 80,000. The window sizes of the LFA module at different sampling stages are set to [0.16 m, 0.32 m, 0.64 m, 1.28 m] respectively. For ScanNetV2, the input points of the room are first sampled with a grid size of 0.02 m, and the maximum number of points is set to 120,000. The window sizes of the LFA module at different sampling stages are set to [0.1 m, 0.2 m, 0.4 m, 0.8 m] respectively. The networks of both datasets are trained for 100 epochs with a batch size of 8.

[0049] The present invention first evaluates the scene segmentation performance of the proposed HMTT on the S3DIS dataset. Table I and Figure 5 show the segmentation results quantitatively and qualitatively respectively. As can be seen from Figure 5 it, the segmentation results of the model of the present invention are roughly similar to the ground truth segmentation labels. In addition, to demonstrate the superiority of the model of the present invention, some state-of-the-art segmentation methods are selected as baseline methods, which can be divided into projection-based, voxel-based, point-wise MLP-based, convolution-based, graph-based and transformer-based methods. Table II shows the scene segmentation results of the proposed HMTT and these baseline methods. As can be seen from Table I, the model of the present invention improves by more than 4% in terms of mAcc and mIoU metrics compared with other non-transformer-based methods. Compared with the recent transformer-based methods, the model of the present invention still improves by 0.8% in terms of OA and mIoU respectively. In addition, it can be found that the model of the present invention achieves the state-of-the-art performance on 6 out of 13 categories of the S3DIS dataset. Although the present invention does not achieve the best accuracy on other categories, it can be seen from Figure 5 that satisfactory segmentation results are obtained for these categories. In particular, the segmentation of doors and tables is very similar to the ground truth segmentation labels.

[0050] E. Results on the ScanNetV2 dataset.

[0051] Meanwhile, the present invention further tested the model of the present invention for scene segmentation on the ScanNetV2 dataset. Table 1 lists the segmentation results of the HMTT model of the present invention and various baseline models, including point-wise MLP-based, convolution-based, graph-based, and transformer-based models. It can be found that the model of the present invention still achieves better segmentation performance than the existing state-of-the-art methods. Specifically, the model of the present invention reaches 75.5% and 73.9% mIoU on the validation set and the test set respectively, which represents a significant improvement over the performance of previous models. It is worth noting that compared with the S3DIS data scanned by RGB-D cameras, the ScanNet data is collected by handheld devices, and there are more clutters and missing object parts in the dataset. Therefore, segmentation on ScanNet data is more challenging.

[0052] Figure 6 and Figure 7 shows the segmentation results of the model of the present invention and the ground truth segmentation. It can be found that although there are some segmentation errors in the model of the present invention, in most cases, the segmentation results are consistent with the ground truth segmentation.

[0053] Table 1 Semantic Segmentation Results for Dataset Validation and Testing In this section, the present invention conducted a series of ablation studies on the S3DIS dataset to verify the effectiveness of each module of the proposed HMTT model. Specifically, the present invention first discussed the effectiveness of the two sub-modules in the LFGS-Attention module, the local feature attention (LFA) module and the global semantic attention (GSA) module. Then the present invention discussed the segmentation performance of models using different loss functions. Finally, the present invention compared the memory consumption of the model using standard global attention to calculate global attention on the same GPU with the model of the present invention.

[0054] Effects of different modules.

[0055] As introduced earlier in the present invention, the present invention proposes an LFGS-Attention module that includes a Local Feature Attention (LFA) module and a Global Semantic Attention (GSA) module to learn local geometric features and global semantic features. To verify the effectiveness of the proposed LFA module and GSA module, the present invention designed an ablation experiment, comparing four models: Model A only uses the LFA module without using a moving window operation to construct cross-window connections; Model B uses the LFA module but does not include the GSA module; Model C adds the GSA module compared with Model A, and Model D adds the GSA module compared with Model B. Table III lists the segmentation mIoU of these models on the S3DIS dataset. Based on the data in the table, the present invention can conduct the following analysis. First, compared with Model A, the mIoU of Model B increased by 0.3. Similarly, compared with Model C, the mIoU of Model D also increased by 0.3. This indicates that introducing a moving window operation in the LFA module can enhance the segmentation performance of the model. As described earlier, this improvement can be attributed to the fact that the moving window operation allows the point cloud to exchange information across the divided cube windows. By applying the moving window operation, the model can effectively capture long-range dependencies and context information, enabling it to better understand the relationships between different parts of the point cloud. Therefore, the model can improve its segmentation performance by leveraging the enhanced context information obtained through the interaction between the point clouds within the moving window. In contrast, models without the moving window operation can only capture the information within a single local window, limiting their ability to learn context information and features. In addition, by comparing the segmentation results of these models, the present invention can further prove the effectiveness of the GSA (Global Semantic Aggregation) module. Specifically, compared with Model A without the GSA module, the mIoU of Model C with the GSA module increased by 0.8. Similarly, compared with Model B without the GSA module, the mIoU of Model D with the GSA module also increased by 0.8. This clearly shows that including the GSA module significantly enhances the segmentation performance of the model. To visually demonstrate the effectiveness of the GSA module, the present invention provides an intuitive comparison of the segmentation results of the model with and without the GSA module in Figure 6 . It can be observed from Figure 6 that without the GSA module, although some bookshelf points are correctly segmented, many bookshelf points are misclassified as clutter points. This may be attributed to the distant spatial relationships between these bookshelves, which cannot be captured by the model without the GSA module from the perspective of context information. On the other hand, the model with the GSA module effectively aggregates the information of each category at the corresponding class label level. Therefore, it achieves superior segmentation of most bookshelf points. These observations support the conclusion that the GSA module significantly improves the segmentation performance of the model. By hierarchically aggregating category-specific information, the GSA module enables the model to capture key context information, thereby obtaining more accurate and consistent segmentation results.

[0056] Effect of different loss functions.

[0057] As mentioned above, the network of the present invention is trained using cross-entropy loss and multi-label soft margin loss, considering both point ground truth labels and scene-level ground truth labels. While cross-entropy loss is necessary for supervised segmentation tasks, the present invention will focus on the effectiveness of multi-label soft margin loss in this section.

[0058] Complexity analysis.

[0059] To further demonstrate the efficiency of the model of the present invention in integrating multi-tokens for global attention calculation, the present invention compares it with a traditional standard global attention method commonly used in previous works, where each point performs an attention operation with all other points. To ensure a fair comparison, the present invention keeps the LFA module in the LFGS-Attention module unchanged and only focuses on the segmentation performance of the model of the present invention under different global attention calculation methods. The present invention tests the segmentation results of these models on a Tesla V100 GPU equipped with 16GB of memory. Specifically, the present invention compares the segmentation results of two models under a batch input of 10,000 points. The multi-token model of the present invention achieves a segmentation accuracy of 65.8% while consuming only 1.9GB of GPU memory. On the other hand, the model using standard global attention occupies 10.3GB of GPU memory but has an accuracy of only 60.8%. This comparison clearly shows that the introduced multi-tokens not only achieve higher segmentation accuracy but also consume less GPU memory. In addition, the present invention presents the best experimental results of the model of the present invention, aiming to maximize the number of input points in a batch within the available GPU memory. These results improve the performance achieved by training with a 10,000-point batch by an additional 5.4%. This improvement can be attributed to including more points, which contains more comprehensive scene information. Therefore, the model is more capable of learning potential semantic features, thus improving its segmentation performance.

[0060] The second embodiment of the present invention, a 3D scene segmentation device based on a hierarchical multi-token transformer, as Figure 8 shown, can be understood as an entity device for implementing the method provided in the first embodiment. The device includes: An aggregation unit configured to, for the input point cloud, use a point embedding layer to aggregate the local geometry and context information of each point; An initialization unit configured to initialize a plurality of learnable tokens as multi-class tokens and transfer the corresponding initial ones to LFGS-Attention to obtain the updated features of the point cloud and the multi-class tokens; A multi-scale feature unit, configured to sample the original point cloud multiple times and extract corresponding multi-scale features using the LFGS-Attention module at each sampling layer; An interpolation feature unit, configured to use an interpolation layer to propagate the multi-scale features extracted from the sampled points of the sampling layer back to the original points to obtain interpolation features; A prediction unit, configured to connect the interpolation features with the point features of the corresponding encoding stage through skip connections at each interpolation layer, wherein the output of the last interpolation layer is fed into a classifier layer to obtain a prediction label for each point; A scene segmentation unit, configured to perform 3D scene segmentation based on the prediction labels.

[0061] In the third embodiment of the present invention, an electronic device, as Figure 9 shown, includes a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to implement the 3D scene segmentation method based on the hierarchical multi-label transformer as described in the first embodiment.

[0062] In the fourth embodiment of the present invention, a computer storage medium stores a computer program, and the computer program is executed to implement the 3D scene segmentation method based on the hierarchical multi-label transformer as described in the first embodiment.

[0063] Through the description of the specific implementation manners, it should be possible to understand more deeply and specifically the technical means and effects adopted by the present invention to achieve the predetermined purpose. However, the accompanying drawings are only for reference and illustration, and are not used to limit the present invention.

Claims

1. A 3D scene segmentation method based on a hierarchical multi-label transformer, characterized in that Including: For the input point cloud, use a point embedding layer to aggregate the local geometry and context information of each point; Initialize multiple learnable tokens as multi-class tokens, and pass the corresponding initial features to LFGS-Attention to obtain the updated features of the point cloud and multi-class tokens; Sample the original point cloud multiple times, and use the LFGS-Attention module at each sampling layer to extract corresponding multi-scale features; Use an interpolation layer to propagate the multi-scale features extracted from the sampled points of the sampling layer back to the original points to obtain interpolation features; At each of the interpolation layers, connect the interpolation features with the point features of the corresponding encoding stage through skip connections, where the output of the last interpolation layer is fed into a classifier layer to obtain the predicted label for each point; Perform 3D scene segmentation based on the predicted labels.

2. The 3D scene segmentation method based on a hierarchical multi-label transformer according to claim 1, wherein The multi-class tokens aggregate point features from the sampled points of the sampling layer at each stage, thereby capturing long-range context dependencies.

3. The 3D scene segmentation method based on a hierarchical multi-label transformer according to claim 2, wherein The LFGS-Attention module is configured to: for a given input point cloud and its features, normalize the features of the point cloud through layer normalization; Respectively input the obtained features into a local feature attention module and a global semantic attention module; Add the outputs of the local feature attention module and the global semantic attention module, and obtain fused features through a linear layer; After layer normalization and MLP, add a residual connection to the output to obtain the final features.

4. The 3D scene segmentation method based on a hierarchical multi-label transformer according to claim 3, characterized in that, The local feature attention module is configured to learn local features through divided cube windows, as shown in the following formula; where MSA represents multi-head self-attention and LN represents layer normalization.

5. The 3D scene segmentation method based on a hierarchical multi-label transformer according to claim 3, wherein The global semantic attention module is configured to: Among multiple class tokens where C is the number of classes and D is the dimension of the feature map, two linear layers are applied to the point cloud features to generate queries and values ; another two linear layers are applied to the multi-class token features to generate keys and values ; Calculate the global attention through and matrix multiplication as follows: , as shown in the following formula: Attention-based , respectively calculate the global class-specific features through matrix multiplication with , and calculate the assignment of the entire point through Gumbel-Softmax ; after assigning each point to the corresponding class label, merge the features of all points belonging to the same class to generate new features for multiple class labels , as follows: Specifically, the following formula: Among them, represents soft assignment, represents hard assignment, is an independent and identically distributed random sample drawn from the Gumbel(0, 1) distribution, is the temperature parameter, is the stop gradient operator, represents the number of points included in the multi-class label; Updating the features of the point cloud by adding global class-specific features as shown in the following equation: 。 6. A 3D scene segmentation device based on a hierarchical multi-label transformer, characterized in that, Including: An aggregation unit, configured to use a point embedding layer to aggregate the local geometry and context information of each point for the input point cloud; An initialization unit, configured to initialize multiple learnable tokens as multi-class tokens, and pass the corresponding initial features to LFGS-Attention to obtain the updated features of the point cloud and multi-class tokens; A multi-scale feature unit, configured to sample the original point cloud multiple times, and use the LFGS-Attention module at each sampling layer to extract corresponding multi-scale features; An interpolation feature unit, configured to use an interpolation layer to propagate the multi-scale features extracted from the sampled points of the sampling layer back to the original points to obtain interpolation features; A prediction unit, configured to at each of the interpolation layers, connect the interpolation features with the point features of the corresponding encoding stage through skip connections, where the output of the last interpolation layer is fed into a classifier layer to obtain the predicted label for each point; A scene segmentation unit, configured to perform 3D scene segmentation based on the predicted labels.

7. An electronic device, characterized in that, Including a memory and a processor, the memory stores a computer program, and the processor is used to execute the computer program to implement the 3D scene segmentation method based on a hierarchical multi-token transformer according to any one of claims 1 to 5.

8. A computer storage medium, characterized in that, The medium stores a computer program, and the computer program is executed to implement the 3D scene segmentation method based on the hierarchical multi-token transformer according to any one of claims 1 to 5.