Point cloud semantic segmentation method based on stage information fusion transformer and grouping normalization

By introducing stage information fusion and grouping normalization methods in point cloud semantic segmentation, the problems of high computing costs and insufficient segmentation accuracy in the prior art are solved, and higher segmentation accuracy and robustness are achieved, suitable for augmented reality and autonomous driving.

CN120259652APending Publication Date: 2025-07-04HEBEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510315566.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing point cloud semantic segmentation method based on Transformer has shortcomings in terms of calculation cost and segmentation accuracy, especially ignoring the fusion problem of information before and after encoding and decoding, resulting in high computational cost and limited segmentation accuracy.

Method used

Using the Transformer and grouping normalization method based on stage information fusion, multi-level information fusion is carried out in the encoding and decoding stages, and combining maximum pooling and learnable weight pooling, feature loss is reduced and segmentation accuracy is improved.

Benefits of technology

It significantly improves the accuracy of point cloud semantic segmentation and the robustness of the model, reduces the computational cost, and is suitable for augmented reality and autonomous driving tasks with high accuracy requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259652A_ABST
    Figure CN120259652A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud semantic segmentation method based on stage information fusion transformer and grouping normalization, and aims to improve the accuracy of point cloud semantic segmentation and reduce the operation cost. A traditional U-net network is adopted to divide the process into a coding stage and a decoding stage for five times, coding firstly adopts core point convolution to extract initial features of point clouds, then downsampling is carried out, features before and after sampling and grouping are utilized to carry out mutual normalization, then the features are linearly combined, and finally, the initial features of the point clouds are extracted; and finally, carrying out adaptive combination of maximum pooling and learnable weight pooling. And then transform attention operation is carried out, information of different coding stages is fused into an attention result, operation opposite to coding is carried out during decoding, attention calculation is carried out after up-sampling is carried out, information fusion is added, and finally segmentation is carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and artificial intelligence, and mainly aims at a large-scale indoor point cloud semantic segmentation method, especially normalizing grouped data before and after sampling, and fusing information in different encoding and decoding stages for segmentation. Background Art

[0002] Point cloud semantic segmentation is an important task in the fields of computer vision and robotics. The goal is to accurately understand and analyze the semantic information of three-dimensional point cloud data. Traditional point cloud processing methods mainly rely on manual feature extraction and local geometric information. However, these methods usually have difficulty effectively capturing the global context relationships in point clouds. For example, deep learning methods such as PointNet achieve good segmentation performance through global feature aggregation, but still have limitations when dealing with complex scenes. In addition, methods such as graph convolutional networks (GCN) perform well in dealing with the local structure of point clouds, but also face the problem of insufficient global feature modeling.

[0003] To solve this problem, research on point cloud semantic segmentation based on Transformer has been proposed, aiming to address the deficiencies of traditional methods in dealing with global dependencies and geometric detail preservation of point clouds. Point Transformer proposes an adaptive attention mechanism for point clouds, which can dynamically capture local and global relationships between points; Point Cloud Attention Network (PCAN) uses a self-attention mechanism to select key features of point clouds, improving the performance of segmentation and classification; Stratiffed Transformer for 3D Point Cloud Segmentation uses a hierarchical transformer and maintains an encoding table for semantic segmentation, achieving good results.

[0004] These Transformer-based methods have achieved good results in terms of segmentation accuracy, but calculating the attention for Q (Query), K (Key), and V (Value) of each point and performing multiple repeated calculations result in a large computational cost. Moreover, only feature fusion is performed at the current stage, ignoring the problem of information reduction before and after encoding and decoding, which limits the segmentation accuracy. The present invention aims to solve the limitations of the existing technologies mentioned above, introduce cross-stage transformers for information fusion, and perform sufficient normalization to improve the accuracy of point cloud semantic segmentation and reduce the computational cost. Summary of the Invention

[0005] The object of the present invention is to overcome the accuracy and computational cost problems in existing transformer-based point cloud semantic segmentation methods, and to propose a point cloud semantic segmentation method based on stage information fusion of transformers and group normalization. Based on the ordinary point transformer, this method improves the accuracy and robustness of semantic segmentation by introducing information from different encoding and decoding layers to perform multi-level fusion on the results of each attention operation.

[0006] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present invention provides a point cloud semantic segmentation method based on stage information fusion of transformers and group normalization, and the method includes the following steps:

[0008] Use kernel point convolution to extract the preliminary features of the input point cloud;

[0009] Use farthest point sampling and K-nearest neighbor algorithm to downsample and group the point cloud;

[0010] Perform mutual normalization and feature fusion on the points before and after sampling;

[0011] Use max pooling and pooling with learnable weights to pool the points after sampling;

[0012] Use stage information fusion to fuse the information of point features at different encoding / decoding stages;

[0013] Use mapping and interpolation methods to upsample the encoded points;

[0014] The process of the stage information fusion is as follows: the feature of the current layer is H n , and the feature H n+1 of the previous layer is the feature of the previous encoding / decoding stage. First, obtain Q, K (Key), and V of the transformer through 3 linear layers. After multiplying Q and K, add the feature P obtained by dimensionality raising of the position information through MLP to perform an attention operation to obtain the attention score Attention1. Then multiply Attention1 and V to obtain the attention H n ' of the current layer; then use the feature H n+1 of the previous layer to perform another set of linear layer processing to obtain Q1, K1, and V1. Multiply Q and K1 and add the feature P, that is, use the Q of the current layer to query the K1 of the previous layer to obtain the attention score Attention2 of the two-stage relationship. Multiply Attention2 and the V1 of the previous layer to obtain the attention H n+1 '; finally, sum the two attentions to obtain the information feature H' fused with the previous stage of this stage;

[0015] In the decoding stage, five decoded features X1 to X5 in different dimensions are obtained. The five decoded features in different dimensions are mapped to the categories to be segmented through different MLPs, so as to be compared with the correct labels for the final result output and loss function optimization;

[0016] The output feature X decoded in each stage i After passing through the MLP, the result is compared with the true result to obtain the loss function of each stage. Summing all the loss functions gives the finally optimized loss function; the final classification result and classification accuracy obtained by performing MLP mapping on the decoded feature X5 output in the last stage of the decoder.

[0017] Furthermore, the process of pooling the sampled points using max pooling and learnable-weight pooling is as follows: Learnable-weight pooling first initializes a learnable weight with a shape of W ∈ R 1×16 using Kaming uniform initialization:

[0018]

[0019] Multiply the feature F1 element-wise with the W matrix to obtain Y:

[0020] Y = F1 ⊙ W

[0021] Then, sum each sample on the intermediate dimension to obtain the learnable pooling feature F3. The normalized feature F1 simultaneously passes through a max pooling to obtain the pooled feature F2. Add F2 and F3 and then pass through a multi-layer perceptron to obtain a scaling factor Z. Sum the result of multiplying F2 by Z and the result of multiplying F3 by 1 - Z to obtain the adaptive pooling feature F4.

[0022] Furthermore, the method is used for semantic segmentation of the large indoor dataset S3DIS, and the mean intersection over union, total accuracy, and mean accuracy are 71.9, 91.4%, and 78.0% respectively.

[0023] Furthermore, the method is used for semantic segmentation of 6-fold cross validation, and the mean intersection over union, total accuracy, and mean accuracy are 76.3, 93.4%, and 84.8% respectively.

[0024] In a second aspect, the present invention provides a computer program product, which includes a storage medium and computer-executable instructions stored on the storage medium. The instructions are configured to implement the steps of the method when executed.

[0025] In a third aspect, the present invention provides a point cloud semantic segmentation system based on stage information fusion of transformers and group normalization, which includes a preliminary feature extraction module, a downsampling fusion pooling and result normalization module, a stage information fusion module, and a loss function optimization module.

[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0027] The present invention uses kernel point convolution for preliminary feature extraction, and adopts the method of adaptive fusion of max pooling and learnable weight pooling as well as sampling normalization, which can greatly reduce the problem of feature loss in the pooling process. The subsequent stage information fusion transformers in the encoding and decoding process can effectively aggregate the upper layer information and the current layer information, reduce the calculation of attention, improve the segmentation accuracy, and also improve the robustness of the model. This method has important value in tasks that require high-precision segmentation such as augmented reality and autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a schematic flowchart of the method of the present invention;

[0029] Figure 2 is a diagram of the kernel point convolution process of the present invention;

[0030] Figure 3 is a schematic diagram of the sampling group normalization of the present invention;

[0031] Figure 4 is a schematic diagram of the fusion of max pooling and learnable weight pooling of the present invention;

[0032] Figure 5 is a diagram of the stage information fusion of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The following further elaborates on the content of the present invention in conjunction with the accompanying drawings.

[0034] Embodiment 1

[0035] As Figure 1 shown, the point cloud semantic segmentation method based on stage information fusion transformers and group normalization of the present invention includes the following steps:

[0036] S1: Preliminary feature extraction;

[0037] Using kernel point convolution, as Figure 2 shown, for the input point cloud P = {p1, p2,..., p N}, where N is the number of points in the input point cloud. First, the original point cloud P is subjected to keypoint selection to select the keypoints P'. For each keypoint P', a ball query is performed to find the local neighborhood that needs to perform keypoint convolution, and then the distance from each neighborhood point P j to the keypoint P' is calculated to obtain a weight function:

[0038] W = φ(∥p j - p'∥)

[0039] where φ(·) is a weight function, and the Gaussian function is used to weight the distance, and p j and p' are respectively -

[0040] The feature f of each neighborhood point is multiplied by the corresponding weight W, and then the weighted features of all neighborhood points are summed to obtain the new feature F ∈ R N×c . The preliminary features extracted by keypoint convolution contain rich local information, and the local neighborhood features of each point are fused for subsequent attention operations of the transformer.

[0041] S2: Downsampling;

[0042] For the feature F obtained by keypoint convolution, downsampling calculates the number of points to be sampled according to different strides, and then performs farthest point sampling to obtain the required point coordinates n_p, and the point feature n_x ∈ R N×c . Then, according to the calibrated number of neighbor points, the neighbor point features X ∈ R N×16×c of each point are obtained by grouping through the K-nearest neighbor algorithm, where c is the number of channels, further increasing the local features of each point, and being able to sample a small number of points that are closest to the features of the original point cloud, reducing the problem of sampling information loss.

[0043] S3: Sampling normalization;

[0044] As Figure 3 shown, the mean and standard deviation μ 1、 μ2, S1, and S2 are respectively obtained for the point feature n_x before downsampling grouping and the point feature X after grouping. Then, a normalization operation is performed, and the grouped feature X minus the mean μ1 before grouping is divided by the sum of the standard deviation S1 before grouping and a bias to obtain the new normalized feature X ∈ R N×16×c :

[0045]

[0046] Similarly, the new normalized feature n_x ∈ R N×c before grouping is obtained:

[0047]

[0048] The normalized features are all subjected to a linear transformation by multiplying by an α and adding a bias β. In this way, mutual normalization is performed on the features before grouping and the features after grouping. After concatenating the last two features and passing them through a linear layer, the features F1 ∈ R of the dimension required for subsequent operations can be obtained. N×16×c Performing such mutual normalization can not only reduce the problem of feature loss before and after, but also ensure that the model converges quickly, improve its robustness, and enable the model to converge rapidly.

[0049] S4: Max pooling and learnable pooling;

[0050] As Figure 4 shown, on the one hand, the normalized feature F1 passes through a max pooling to obtain the pooled feature F2 ∈ R N ×c ; on the other hand, it passes through a learnable weight pooling to obtain the feature F3 ∈ R N×c . For the learnable weight pooling, a learnable weight with a shape of W ∈ R 1×16 is first initialized, and Kaming uniform initialization is used:

[0051]

[0052] The feature F1 is multiplied element-wise with the W matrix to obtain Y:

[0053] Y = F1 ⊙ W

[0054] Then, for each sample, the sum is taken over the intermediate dimension to obtain the feature F3 of the learnable pooling. Subsequently, after adding F2 and F3 and passing them through a multi-layer perceptron to obtain a scaling factor Z, the sum of the result of multiplying F2 by Z and the result of multiplying F3 by 1 - Z is calculated to obtain the adaptive pooling feature F4 ∈ R N×c The pooling feature obtained in this way not only contains the most significant features of the points but also some detailed features, providing sufficient significant information for the subsequent stage information fusion and attention mechanism. The max pooling and learnable pooling are adaptively fused, enabling the grouped points in the downsampling to extract the most significant features and reducing the loss of local features, thereby increasing the accuracy of the model.

[0055] S5: Stage information fusion;

[0056] The entire network is divided into an encoding stage and a decoding stage, which are repeated five times respectively. The current layer feature is H n , and the previous layer feature H n+1 is the feature of the previous encoding (decoding) stage. As Figure 5As shown, the Q, K (Key), and V of the Transformer are obtained through 3 linear layers. After multiplying Q and K, the feature P obtained by dimensionality elevation of the position information through the MLP is added to perform the attention operation to obtain the attention score Attention1. Then, Attention1 is multiplied by V to obtain the attention H of the current layer. n '; Then, the feature H of the previous layer is used n+1 to perform another set of linear layer processing to obtain Q1, K1, and V1. Multiply Q and K1 and add the feature P, that is, use the Q of the current layer to query the K1 of the previous layer to obtain the attention score Attention2 of the two-stage relationship. Attention2 is then multiplied by the V1 of the previous layer to obtain the attention H'. n+1 '; Finally, the two attentions are summed to obtain the information feature H' that fuses the current stage and the previous stage. In this way, different-stage information is added under the condition of sufficient global features, improving the accuracy of the final segmentation network.

[0057] S6: Upsampling;

[0058] It mainly works through mapping and interpolation. When only one point cloud is input (i.e., the first layer of decoding): By processing, aggregating, and transforming the point features within the group, and then performing a linear transformation on the transformed features to achieve feature conversion; When two point clouds are input (i.e., the second layer to the fifth layer): By performing a linear transformation on the features of each point cloud and using the interpolation operation to map the features of point cloud 2 to point cloud 1, finally obtaining the fused features. This process enables the Transformer to perform effective upsampling on irregular point cloud data while maintaining the global and local structural information between points, thereby improving the performance in point cloud processing tasks.

[0059] S7: MLP (Multi-Layer Perceptron);

[0060] During the decoding stage, 5 decoded features X1 to X5 with different dimensions are obtained. Then, the features with five different dimensions are mapped to the categories to be segmented through different MLPs, so as to compare with the correct labels for the final result output and loss function optimization.

[0061] S8: Output result and optimization;

[0062] The result obtained by mapping the feature X5 output in the last stage of the decoder is compared with the label of the true result, and the final accuracy Miou (Mean Intersection over Union) of the semantic segmentation of the present invention is obtained through calculation. Then, the output features X of each stage of decoding i After going through the steps of S7, the results are compared with the true results to obtain the loss function of each stage. Summing all the loss functions gives the finally optimized loss function.

[0063] After all processes are completed, performance evaluation is carried out, mainly using the metrics Miou (mean intersection over union) and OA (overall accuracy) on the large indoor dataset S3DIS for evaluation.

[0064] S3DIS (Stanford 3D Indoor Spaces Dataset) is a large-scale point cloud dataset for indoor scenes, which is widely used in the research of semantic segmentation and instance segmentation of point clouds. This dataset contains indoor scenes from 6 different regions, covering 271 rooms, including various scene types such as offices, meeting rooms, corridors, etc. The point clouds in the S3DIS dataset are collected by a Matterport camera, which combines 3 structured light sensors with different spacings and can capture RGB and depth images by rotating 360° at each scanning position. Each point in it is assigned a semantic label, such as chair, table, floor, wall, etc.

[0065] In the experiment, the average intersection over union, overall accuracy, and average accuracy of the present invention and other networks are compared on the S3DIS test set and 6-fold cross validation.

[0066] Table 1 Comparison of experimental results between the present invention and other networks on the S3DIS test set

[0067]

[0068] Table 2 Comparison of experimental results between the present invention and other networks on 6-fold cross validation of S3DIS

[0069] network OA (Overall Accuracy) mAcc (mean accuracy) Miou (mean intersection over union) PointNet 78.5 66.2 47.6 RSNet - 66.5 56.5 SPGraph 85.5 73.0 62.1 PAT - 76.5 64.3 PointCNN 88.1 75.6 65.4 PointWeb 87.3 76.2 66.7 ShellNet 87.1 - 66.8 RandLA-Ne 88.0 82.0 70.0 KPConv - 79.1 70.6 PointTransformer 90.2 81.9 73.5 the method of the present invention 93.4 84.8 76.3

[0070] As can be seen from the above, whether it is the accuracy or the average intersection over union, the results of the model of the present application on the S3DIS dataset are higher than those of other methods. Moreover, compared with the current popular point transformer model, the time cost is only half of it, and the results are more accurate than it. This shows that the accuracy and robustness of the method of the present invention are quite good, indicating that this stage fusion and normalization method has a good effect in semantic segmentation.

[0071] Through these evaluations, the method of the present invention can clearly identify the advantages of the new method and potential improvement directions, providing guidance for future research. At the same time, these results also provide valuable references for researchers in related fields, promoting the further development of point cloud semantic segmentation technology.

[0072] Example 2

[0073] The point cloud semantic segmentation system of the present invention based on stage information fusion transformer and group normalization includes a preliminary feature extraction module, a downsampling fusion pooling and result normalization module, a stage information fusion module, and a loss function optimization module.

[0074] The preliminary feature extraction module uses lightweight core point convolution to obtain rich local features of each point, and performs dimensionality increase to facilitate subsequent feature extraction. The features extracted for each point contain semantic information of the point and multiple points in its neighborhood, which is convenient for subsequent feature extraction.

[0075] The downsampling fusion pooling and result normalization module uses different strides for farthest point sampling and grouping, and after grouping, calculates the center and variance of the points before and after grouping, performs mutual normalization on the two groups of points, then performs large pooling and a learnable weight pooling, and adaptively sums the two results. The global features obtained by max pooling may lose some information. Adaptive summing it with the pooling of learnable weights can make the obtained features more comprehensive.

[0076] The stage information fusion module performs subtraction operation on the query of the current encoding (decoding) stage and the key of the previous stage, and then performs multiplication operation with the value of the previous stage to obtain a semantic information that fuses the relationship between the current stage and the previous stage, and fuses it with the features of the current layer to improve the final segmentation accuracy. The stage information fusion module is added based on pointtransformer, which adds information of different stages when the global features are sufficient, thereby improving the accuracy of the final segmentation network.

[0077] The loss function optimization module returns the features of a point and the index of the point after each downsampling. In the decoding stage, the points obtained in each decoding process are segmented and compared with the labels corresponding to the returned indices to obtain the loss function of each process, and the loss functions of each process are summed to obtain a new loss function.

[0078] The overall process is divided into an encoding process and a decoding process. Each process contains 5 stages. The encoding stage includes downsampling and attention operations, and the decoding stage includes upsampling and attention operations. Each stage performs five operations with different strides, and finally segments the result obtained by decoding, and finally calculates the accuracy of the segmentation.

[0079] Using the traditional U-net network, the process is divided into five encoding and decoding stages respectively. In the encoding stage, first, core point convolution is used to extract the initial features of the point cloud, then downsampling is performed. The features before and after sampling grouping are mutually normalized, and then the features are linearly combined. Finally, max pooling and pooling with learnable weights are adaptively combined. Then, the transformer attention operation is performed while fusing the information of different encoding stages into the attention result. During decoding, the opposite operations of encoding are taken, first upsampling is performed, then attention calculation is carried out and information fusion is added, and finally segmentation is performed. This method provides an effective tool for point cloud semantic segmentation and also provides valuable reference for relevant researchers in this field, promoting the further development of point cloud semantic segmentation technology

[0080] Matters not described in the present invention are applicable to the prior art

Claims

1. A point cloud semantic segmentation method based on the fusion of stage information, transformer, and group normalization, characterized in that The method includes the following steps: Using keypoint convolution to extract the preliminary features of the input point cloud; Using farthest point sampling and K-nearest neighbor algorithm to downsample and group the point cloud; Performing mutual normalization and feature fusion on the points before and after sampling; Pooling the points after sampling using max pooling and pooling with learnable weights; Using stage information fusion to fuse the information of point features at different encoding / decoding stages; Upsampling the encoded points using mapping and interpolation methods; The process of fusing the stage information is as follows: the feature of the current layer is H n , and the feature H n+1 of the previous layer is the feature of the previous encoding / decoding stage. First, Q, K (Key), and V of the transformer are obtained through 3 linear layers. After multiplying Q and K, the feature P obtained by dimensionality elevation of the position information through MLP is added to perform attention operation to obtain the attention score Attention1. Then, Attention1 is multiplied by V to obtain the attention H n ' of the current layer; then, another set of linear layer processing is performed using the feature H n+1 of the previous layer to obtain Q1, K1, and V1. Multiply Q and K1 and then add the feature P, that is, use Q of the current layer to query K1 of the previous layer to obtain the attention score Attention2 of the two-stage relationship. Attention2 is then multiplied by V1 of the previous layer to obtain the attention H n+1 '; finally, the two attentions are summed to obtain the information feature H' fused with the previous stage; In the decoding stage, five decoded features X1 to X5 with different dimensions are obtained. The five decoded features with different dimensions are mapped to the categories to be segmented through different MLPs, so as to compare with the correct labels for the final result output and loss function optimization; The output feature X decoded in each stage i After passing through the MLP, the result is compared with the true result to obtain the loss function for each stage, and the sum of all loss functions is obtained to get the final optimized loss function; the final classification result and classification accuracy obtained by performing MLP mapping on the decoded feature X5 output by the last stage of the decoder.

2. The method according to claim 1, characterized in that, The process of pooling the sampled points using max pooling and learnable-weight pooling is as follows: Learnable-weight pooling first initializes a learnable weight with a shape of W ∈ R 1×16 using Kaming uniform initialization: Multiplying the feature F1 element-wise with the W matrix to obtain Y: Y = F1 ⊙ W Then, the sum of each sample in the intermediate dimension is calculated to obtain the learnable pooling feature F3. The normalized feature F1 also passes through a max pooling to obtain the pooled feature F2. The sum of F2 and F3 is passed through a multi-layer perceptron to obtain a scaling factor Z. The result of multiplying F2 by Z is summed with the result of F3 multiplied by 1-Z to obtain the adaptive pooling feature F4.

3. The method according to claim 1, wherein The method is used for semantic segmentation of the large indoor dataset S3DIS, and the average intersection over union, overall accuracy, and mean accuracy are 71.9, 91.4%, and 78.0% respectively.

4. The method according to claim 1, wherein The method is used for semantic segmentation of 6-fold cross validation, and the average intersection over union, overall accuracy, and mean accuracy are 76.3, 93.4%, and 84.8% respectively.

5. A computer program product for performing the method according to any one of claims 1 to 4, characterized in that, The product includes a storage medium and computer-executable instructions stored on the storage medium, and the instructions are configured to implement the steps of the method when executed.

6. A point cloud semantic segmentation system based on the fusion of stage information, transformer, and group normalization, characterized in that, The system includes a preliminary feature extraction module, a downsampling fusion pooling and result normalization module, a stage information fusion module, and a loss function optimization module.