Unsupervised three-dimensional point cloud semantic segmentation method based on progressive expansion super voxel

By adopting the progressively expanded super voxel and clustering method in unsupervised three-dimensional point cloud semantic segmentation, providing pseudo-labels as supervision signals, solving the problem of weakening feature representations and fuzzy segmentation results in the existing methods, and achieving efficient and accurate semantic segmentation.

CN120125825APending Publication Date: 2025-06-10NORTHEASTERN UNIV AT QINHUANGDAO
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510543454.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing unsupervised three-dimensional point cloud semantic segmentation method is difficult to learn feature representations with strong discrimination, resulting in low segmentation accuracy and fuzzy correspondence between the segmentation result and the specific semantic category.

Method used

Using a method based on progressive expansion super voxel, by constructing a progressive expansion super voxel, the network can gradually learn more global semantic information from local features, and combines the clustering method to provide pseudo-labels for the model to guide the learning as a supervised signal.

Benefits of technology

It effectively solves the problem of lack of semantic information in small local point sets, improves the accuracy and efficiency of semantic segmentation, and can learn meaningful semantic categories without supervision without manual annotation or pre-training models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125825A_ABST
    Figure CN120125825A_ABST
Patent Text Reader

Abstract

The invention provides an unsupervised three-dimensional point cloud semantic segmentation method based on progressive expansion super voxels, and relates to the technical field of three-dimensional point cloud processing. A progressive expansion super voxel is constructed, and a pseudo label is provided for the model in combination with a clustering technology, and is used as a supervision signal to guide the model to learn. In the training process, the gradually expanded super voxels are adopted to guide the training of the semantic segmentation model, so that the semantic segmentation model can gradually learn more global semantic information from local features, and the problems of fuzzy boundaries and insufficient understanding of complex scenes caused by lack of semantic information in a small local point set can be effectively solved. According to the mode, the network can learn meaningful semantic categories under the unsupervised condition without any manual annotation or pre-training model, time can be greatly saved, the adaptability to a new scene can be improved, semantic information can be accurately and efficiently extracted from the three-dimensional point cloud, and the processing method is clear and has high operability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional point cloud processing, and in particular to an unsupervised three-dimensional point cloud semantic segmentation method based on progressively expanded supervoxels. Background Art

[0002] As the core means of obtaining three-dimensional semantic information of scenes, 3D point cloud semantic segmentation technology has been widely used in fields such as autonomous driving, robot navigation, and 3D building modeling. According to different learning methods, existing methods are mainly divided into supervised and unsupervised categories. Supervised methods rely on large-scale labeled data for training and can usually achieve higher segmentation accuracy. However, in actual industrial applications, the acquisition of 3D point cloud data often faces challenges such as scene diversity and data scarcity, and the cost of refined semantic annotation is high, which seriously restricts the application scope and generalization ability of supervised methods. In contrast, unsupervised methods do not need to rely on labeled data and have certain advantages in data-constrained scenarios. However, due to the lack of clear semantic category guidance, existing unsupervised methods often find it difficult to learn feature representations with strong discriminative power, resulting in low segmentation accuracy and fuzzy correspondence between segmentation results and specific semantic categories. Usually, complex post-processing steps are required to meet actual application needs, and timeliness is low. Summary of the invention

[0003] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned prior art and to provide an unsupervised three-dimensional point cloud semantic segmentation method based on progressively expanding supervoxels. By constructing progressively expanding supervoxels, the network can gradually learn more global semantic information from local features, effectively solving the problem of lack of semantic information in small local point sets, thereby realizing model training and optimization. At the same time, the clustering method is combined to provide pseudo labels for the model, so that the network can learn meaningful semantic categories in an unsupervised manner, thereby accurately and efficiently extracting semantic information from the three-dimensional point cloud.

[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is: An unsupervised 3D point cloud semantic segmentation method based on progressively expanding supervoxels includes a feature extraction module that extracts feature information of each point from point cloud data through a neural network, a supervoxel construction module that gradually generates larger supervoxels during training to guide semantic learning, and a semantic primitive clustering module that uses a clustering algorithm to group and classify basic elements of semantic categories. The specific steps are as follows: Step 1: Set the number of supervoxel layers, the number of supervoxels at different levels, the number of training rounds, the voxel grid size, the voxel cloud connection segmentation algorithm parameters, the region growing algorithm parameters, and the K-means clustering algorithm parameters; Step 2: Initialize the supervoxel by using the voxel cloud connection segmentation algorithm and the region growing algorithm; perform preliminary segmentation of the point cloud by using the voxel cloud connection segmentation algorithm to generate multiple small connected regions; apply the region growing algorithm to expand these small regions to form the initial supervoxel; Step 3: The feature extraction module processes the x, y, z three-dimensional coordinates of each point in the input point cloud data, extracts features and generates a high-dimensional feature representation; Step 4: The semantic primitive clustering module clusters all supervoxels in the dataset and assigns a semantic label to each 3D point and supervoxel; Step 5: Using the cross entropy loss function, the pseudo labels generated by the semantic primitive clustering module are used as supervisory signals to train and optimize the neural network of the feature extraction module; Step 6: After a certain number of trainings, the level of the supervoxels is increased, the number of supervoxels is reduced, and K-means clustering is reapplied to the point cloud in combination with the features of the current network output to further divide the initial superpoints into a small number of new superpoints; Step 7: Repeat the above steps 5 and 6 until a predetermined supervoxel level is reached; Step 8: Complete model training and generate semantic segmentation results.

[0005] Furthermore, in step 2, the voxel cloud connection segmentation algorithm first voxelizes the input point cloud into a voxel grid of 5 cm×5 cm×5 cm; then, a set of seed points are evenly distributed in the voxelized point cloud, with an interval of 50 cm between the seed points, and these seed points serve as the initial centers of the supervoxels; for each seed point, P adjacent points are searched within a sphere with a radius of 50 cm, and the distance between each adjacent point and the current seed point is calculated using the following formula: ; Among them, D n , D c , D s Respectively represent the spatial Euclidean distance, color Euclidean distance and normal Euclidean distance between adjacent points, ω n ,ω c ,ω s is the corresponding weight coefficient, is the interval between seed points.

[0006] Furthermore, in step 2, the region growing algorithm first defines the radius as The neighborhood points of are used to evaluate the similarity between adjacent points; the similarity of adjacent points is expressed by the similarity measure of the normal line, and the formula is: ; Among them, n i and nj They are point p i and p j The normal of the cluster; the smoothness threshold, curvature threshold and residual threshold are set to 3, 1, and 1 respectively to determine whether to include adjacent points in the current cluster. Specifically, if If it is greater than the smoothness threshold, then p i and p j These two points are similar; The size of each cluster is limited by the minimum and maximum number of points, and the number of neighboring points is set to , used to control the range of similarity judgment; the clustering condition is expressed as: and ; Among them, D(p i , p j ) is point p i and p j The Euclidean distance between Curvature(p j ) is the curvature of point pj; residual_threshold is the residual threshold, and curvature_threshold is the curvature threshold.

[0007] Finally, a set of clustering results is output, representing the segmentation of similar areas in the point cloud.

[0008] Furthermore, in step 3, the feature extraction module gradually extracts local and global features of the point cloud through feature transformation, multi-layer perceptron with shared weights, and global pooling operations, including the following steps: Step 3.1: The input point cloud data X is transformed into features; A 3×3 fully connected network is used to generate the transformation matrix T, which maps the input points to the new coordinate space. The transformed point cloud is represented as , and the dimension of the data remains unchanged; Step 3.2: Input the transformed point cloud into a multi-layer perceptron with shared weights to extract local features point by point, thereby generating a local feature F with a dimension of n×64 1 , where n represents the number of points in the input point cloud; Step 3.3: Generate a new feature transformation matrix T' through a 64×64 fully connected network, transform the local features output by the multilayer perceptron, and generate a new feature representation , the dimension of the data remains unchanged at n×64; Step 3.4: Use the multi-layer perceptron to extract high-dimensional features from local features to obtain high-dimensional features, and obtain the global feature F with a dimension of 1×1024 through the global maximum pooling layer. global ; Step 3.4: The local feature F with dimension n×64 obtained above 1 is concatenated with n global features F with dimension 1×1024 global to generate a point feature F with dimension n×1088 concact =Concat(F 1 ’, F global_broadcast ), where F global_broadcast is the feature obtained by broadcasting the global feature F global to each point, with dimension n×1024; through the concatenation operation, the complete feature representation F of each point is obtained concact ; Step 3.5: Compress the concatenated features through a multi-layer perceptron to extract high-dimensional features, and finally generate the high-dimensional feature F of each point point , which is a point feature with dimension n×128 that integrates global features.

[0009] The beneficial effects of adopting the above technical solution are as follows: The unsupervised 3D point cloud semantic segmentation method based on progressive expansion of supervoxels provided by the present invention constructs progressive expansion of supervoxels and combines clustering technology to provide pseudo-labels for the model as supervision signals to guide the model to learn. During the training process, progressive expansion of supervoxels is used to guide the training of the semantic segmentation model, so that it can gradually learn more global semantic information from local features, and can effectively solve the problems of blurred boundaries caused by the lack of semantic information in small local point sets and insufficient understanding of complex scenes. This method enables the network to learn meaningful semantic categories without any manual annotation or pre-trained model, can greatly save time and improve the adaptability to new scenes, and thus accurately and efficiently extract semantic information from 3D point clouds. This processing method is clear and has strong operability. Description of the Drawings

[0010] Figure 1 is a flowchart of the unsupervised 3D point cloud semantic segmentation method based on progressive expansion of supervoxels provided by an embodiment of the present invention; Figure 2 is a neural network architecture diagram of the feature extraction module provided by an embodiment of the present invention. Detailed Embodiment

[0011] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0012] This embodiment proposes an unsupervised 3D point cloud semantic segmentation method based on progressively expanding supervoxels. The overall architecture includes the following three parts: The feature extraction module extracts the feature information of each point from the point cloud data through a neural network; the supervoxel construction module gradually generates larger supervoxels during the training process to effectively guide the progress of semantic learning; the semantic primitive clustering module uses a clustering algorithm to group and classify the basic elements of semantic categories. As Figure 1 shown, the method of this embodiment is described as follows.

[0013] Step 1: Set the number of layers of supervoxels, the number of supervoxels at different levels, the number of training epochs, the voxel grid size, the parameters of the voxel cloud connectivity segmentation algorithm, the parameters of the region growing algorithm, and the parameters of the K-means clustering algorithm.

[0014] Step 2: Initialize the supervoxels through the voxel cloud connectivity segmentation algorithm and the region growing algorithm; perform a preliminary segmentation of the point cloud through the voxel cloud connectivity segmentation algorithm to generate multiple small connected regions; apply the region growing algorithm to expand these small regions to form initial supervoxels.

[0015] Step 3: The feature extraction module processes the x, y, and z three-dimensional coordinates of each point in the input point cloud data, extracts features, and generates a high-dimensional feature representation.

[0016] Step 4: The semantic primitive clustering module clusters all the supervoxels in the dataset and assigns a semantic label to each 3D point and supervoxel.

[0017] Step 5: Using the cross-entropy loss function, take the pseudo-labels generated by the semantic primitive clustering module as the supervision signal to train and optimize the neural network of the feature extraction module.

[0018] Step 6: Every time 7 rounds of training are completed, increase the level of the supervoxels, reduce the number of supervoxels, combine the features output by the current network, reapply K-means clustering to the point cloud, and further divide the initial superpoints into a small number of new superpoints.

[0019] Step 7: Repeat the above Steps 5 and 6 until the predetermined level of the supervoxels is reached.

[0020] Step 8: Complete the training of the model and generate the semantic segmentation result.

[0021] As Figure 2 shown, the specific implementation steps of the neural network of the feature extraction module are as follows: Step 3.1: The input of the feature extraction module is the point cloud data represented as , where n is the number of points, and 3 represents the three-dimensional coordinates (x, y, z) of each point. Through the input transformation matrix , align the point cloud data to a new coordinate space, and the transformed point cloud representation , this process keeps the data dimension unchanged while enhancing the rotational invariance of the point cloud features.

[0022] Step 3.2: Input the transformed point cloud X’ into a multi-layer perceptron with shared weights to extract local features point by point. This process uses a two-layer multi-layer perceptron: the output feature dimension of the first layer is n×64, and the dimension of the second layer is still n×64 to obtain local features .

[0023] Step 3.3: Through the feature transformation matrix , align the local features to a new coordinate space, and the transformed features are represented as . This process further enhances the rotational invariance of the features.

[0024] Step 3.4: Use a three-layer multi-layer perceptron with shared weights to extract high-dimensional features from the local features, and the output dimensions are n×64, n×128, and n×1024 in sequence, and finally obtain high-dimensional features . Through the global pooling operation, extract the global features , thereby capturing the global geometric information of the entire point cloud.

[0025] Step 3.5: Broadcast the global feature F global to each point and concatenate it with the local features to generate the feature representation of the point: ; where, is the feature obtained by broadcasting the global feature to each point, is the local feature. Through the concatenation operation, obtain the complete feature representation F of each point concact .

[0026] Step 3.6: Use a three-layer multi-layer perceptron with shared weights to extract high-dimensional features from the concatenated features, and the output feature dimensions are n×512, n×256, and n×128 in sequence, and finally generate the high-dimensional features of each point .

[0027] The implementation steps of the voxel cloud connectivity segmentation algorithm are as follows: Step 2.1.1: Voxelize the input point cloud data. Specifically, the point cloud data is converted into a voxel grid, and the size of the voxel is set to 5cm×5cm×5cm.

[0028] Step 2.1.2: After voxelization, each voxel in the point cloud will be assigned to a group. The distance between voxels within each group is set to 50cm. The center points of these groups are defined as supervoxels and used as the starting centers for subsequent calculations.

[0029] Step 2.1.3: For each supervoxel, search for adjacent points within a sphere with a radius of 50 cm. Specifically, calculate the spatial distance, color distance, and normal distance between supervoxels through weighted combination: ; where D n , D c , D s represent the spatial Euclidean distance, color Euclidean distance, and normal Euclidean distance between adjacent points respectively, and ω n , ω c , ω s are the corresponding weight coefficients, set to 0.7, 0.2, and 0.1 respectively, is the interval between seed points, set to 50 cm here.

[0030] Step 2.1.4: According to the calculated distances, the algorithm connects similar superpoints together to form a complete segmentation region. By continuously iterating this process, the algorithm can effectively segment the point cloud data into multiple regions. Finally, the algorithm outputs the segmented point cloud data, and each region corresponds to a segmentation label, which can be used for subsequent point cloud analysis and processing.

[0031] The implementation steps of the region growing algorithm are as follows: Step 2.2.1: Define the parameters of the region growing algorithm, including the neighborhood radius and similarity metric. The neighborhood radius is set to 3, and the similarity metric is calculated through the following formula: ; where n i and n j are the normals of points p i and p j respectively.

[0032] Step 2.2.2: Set the clustering conditions: The maximum number of points in each cluster is limited to 10000, and the upper limit of the number of clusters is 15. The residual threshold and curvature threshold are both set to 1.

[0033] Step 2.2.3: During the clustering process, check the points p i in the neighborhood of each point p j , and calculate the Euclidean distance D(p i , p j ) and curvature Curvature(p j ) between them. Only when both conditions that the distance is less than the residual threshold and the curvature is less than the curvature threshold are satisfied, the point p j will be included in the current cluster, that is: and ; where D(p i , p j ) is the Euclidean distance between points p i and p j , and Curvature(p j ) is the curvature of point pj; residual_threshold is the residual threshold, and curvature_threshold is the curvature threshold.

[0034] Step 2.2.4: Output a set of clustering results, with each cluster corresponding to a region.

[0035] The implementation steps of the supervoxel construction module and the semantic primitive clustering module are as follows: Step 4.1: Perform a preliminary segmentation of the point cloud through the voxel cloud connectivity segmentation algorithm to generate multiple small connected regions.

[0036] Step 4.2: Apply the region growing algorithm to expand these small regions to form initial supervoxels.

[0037] Step 4.3: During the training process, after a certain number of rounds, extract the features of all points within each supervoxel and calculate their average value, which is used as the feature of the supervoxel.

[0038] Step 4.4: Use the K-means clustering algorithm in the semantic primitive clustering module to cluster the features of these supervoxels into fewer categories, with each category representing a new superpoint. This process gradually expands the scale of the supervoxels.

[0039] Step 4.5: In subsequent training rounds, the semantic primitive clustering module assigns pseudo-labels to the supervoxels. The neural network in the feature extraction module calculates the cross-entropy loss based on the pseudo-labels and uses gradient descent to update and optimize the network.

[0040] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. An unsupervised 3D point cloud semantic segmentation method based on progressively dilated supervoxels, characterized by: It includes a feature extraction module that extracts feature information of each point from point cloud data through a neural network, a supervoxel construction module that gradually generates larger supervoxels during training to guide semantic learning, and a semantic primitive clustering module that uses a clustering algorithm to group and classify the basic elements of semantic categories. The specific steps are as follows: Step 1: Set the number of supervoxel layers, the number of supervoxels at different levels, the number of training rounds, the voxel grid size, the voxel cloud connection segmentation algorithm parameters, the region growing algorithm parameters, and the K-means clustering algorithm parameters; Step 2: Initialize the supervoxel by using the voxel cloud connection segmentation algorithm and the region growing algorithm; perform preliminary segmentation of the point cloud by using the voxel cloud connection segmentation algorithm to generate multiple small connected regions; apply the region growing algorithm to expand these small regions to form the initial supervoxel; Step 3: The feature extraction module processes the x, y, z three-dimensional coordinates of each point in the input point cloud data, extracts features and generates a high-dimensional feature representation; Step 4: The semantic primitive clustering module clusters all supervoxels in the dataset and assigns a semantic label to each 3D point and supervoxel; Step 5: Using the cross entropy loss function, the pseudo labels generated by the semantic primitive clustering module are used as supervisory signals to train and optimize the neural network of the feature extraction module; Step 6: After a certain number of trainings, the level of the supervoxels is increased, the number of supervoxels is reduced, and K-means clustering is reapplied to the point cloud in combination with the features of the current network output to further divide the initial superpoints into a small number of new superpoints; Step 7: Repeat the above steps 5 and 6 until a predetermined supervoxel level is reached; Step 8: Complete model training and generate semantic segmentation results.

2. The unsupervised 3D point cloud semantic segmentation method based on progressively dilated supervoxels according to claim 1, characterized in that: In step 2, the voxel cloud connection segmentation algorithm first voxelizes the input point cloud into a voxel grid of 5 cm×5 cm×5 cm; then, a set of seed points are evenly distributed in the voxelized point cloud, with an interval of 50 cm between the seed points, and these seed points are used as the initial centers of the supervoxels; for each seed point, P adjacent points are searched within a sphere with a radius of 50 cm, and the distance between each adjacent point and the current seed point is calculated using the following formula: ; Among them, D n , D c , D s Respectively represent the spatial Euclidean distance, color Euclidean distance and normal Euclidean distance between adjacent points, ω n ,ω c ,ω s is the corresponding weight coefficient, is the interval between seed points.

3. The unsupervised 3D point cloud semantic segmentation method based on progressively dilated supervoxels according to claim 1, characterized in that: In step 2, the region growing algorithm first defines the radius as Neighborhood points of to evaluate the similarity between adjacent points; The similarity of adjacent points is expressed by the similarity measure of normals, and the formula is: ; Among them, n i and n j They are point p i and p j The normal of the cluster; the smoothness threshold, curvature threshold and residual threshold are set to 3, 1, and 1 respectively to determine whether to include adjacent points in the current cluster. Specifically, if If it is greater than the smoothness threshold, then p i and p j These two points are similar; The size of each cluster is limited by the minimum and maximum number of points, and the number of neighboring points is set to , used to control the range of similarity judgment; the clustering condition is expressed as: and ; Among them, D(p i , p j ) is point p i and p j The Euclidean distance between Curvature(p j ) is the curvature of point pj; residual_threshold is the residual threshold, curvature_threshold is the curvature threshold; Finally, a set of clustering results is output, representing the segmentation of similar areas in the point cloud.

4. The unsupervised 3D point cloud semantic segmentation method based on progressively dilated supervoxels according to claim 1, characterized in that: In step 3, the feature extraction module gradually extracts local and global features of the point cloud through feature transformation, multi-layer perceptron with shared weights, and global pooling operations, including the following steps: Step 3.1: The input point cloud data X is transformed into features; A 3×3 fully connected network is used to generate the transformation matrix T, which maps the input points to the new coordinate space. The transformed point cloud is represented as , and the dimension of the data remains unchanged; Step 3.2: Input the transformed point cloud into a multi-layer perceptron with shared weights to extract local features point by point, thereby generating a local feature F1 with a dimension of n×64, where n represents the number of points in the input point cloud; Step 3.3: Generate a new feature transformation matrix T' through a 64×64 fully connected network, transform the local features output by the multilayer perceptron, and generate a new feature representation , the dimension of the data remains unchanged at n×64; Step 3.4: Use the multi-layer perceptron to extract high-dimensional features from local features to obtain high-dimensional features, and obtain the global feature F with a dimension of 1×1024 through the global maximum pooling layer. global ; Step 3.4: Combine the local feature F1 with dimension n×64 obtained above with the global feature F with dimension n×1024 global Spliced ​​together to generate a point feature F with a dimension of n×1088 concact =Concat(F1', F global_broadcast ), F global_broadcast The global feature F global The features obtained by broadcasting to each point have a dimension of n×1024; through the splicing operation, the complete feature representation F of each point is obtained concact ; Step 3.5: Compress the concatenated features through a multi-layer perceptron to extract high-dimensional features, and finally generate high-dimensional features F for each point point , point features with dimension n×128 that are integrated with global features.

Citation Information

Cited By

  • Multispectral point cloud classification method based on space-spectrum self-supervision pre-training

    CN121190883A

  • A Multispectral Point Cloud Classification Method Based on Spatial-Spectral Self-Supervised Pre-training

    CN121190883B