A Point Cloud Semantic Segmentation Method Based on Cross-Point Cloud Context Information

By introducing a memory mechanism that cross-point cloud context information in point cloud semantic segmentation, the problem that sparse points in the prior art is difficult to obtain rich context information, which significantly improves the accuracy of point cloud semantic segmentation.

CN115984567BActive Publication Date: 2025-06-03TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310148017.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2025-06-03
Estimated Expiration
2043-02-22

AI Technical Summary

Technical Problem

The existing point cloud semantic segmentation algorithm only uses the context information within the neighborhood of the point, making it difficult for points in sparse parts to obtain rich context information, which in turn affects the accuracy of the segmentation results.

Method used

The point cloud semantic segmentation method based on cross-point cloud context information is adopted to enhance the characteristics of points in the point cloud through memory storing and utilizing cross-point cloud context information, and provide richer context information.

Benefits of technology

It realizes that sparse partial points obtain richer context information, thereby improving the accuracy and effect of point cloud semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984567B_ABST
    Figure CN115984567B_ABST
Patent Text Reader

Abstract

The present invention relates to a point cloud semantic segmentation method based on cross-point cloud context information. The point cloud semantic segmentation network adopted is divided into an encoder and a decoder, and includes the following steps: performing data preprocessing and data augmentation on the data set; passing the preprocessed and data-augmented point cloud through the encoder. Each stage of the encoder will output the point cloud features of that stage. Take the neighborhood of each point in each stage, calculate the local geometric features of each point according to the neighborhood, and update the memory of each stage; in the third step, pass the point cloud features output by the encoder through the decoder. Each stage of the decoder will output the point cloud features of that stage; input the enhanced point cloud features of the last stage into a multi-layer perceptron to obtain an output; calculate the cross-entropy loss between the output result and the semantic label of the point cloud, and use the backpropagation algorithm and the gradient descent algorithm to optimize the parameters in the encoder and decoder; complete the training of the entire neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and computer vision, and relates to point cloud semantic segmentation technology. Specifically, it is a point cloud semantic segmentation method based on cross-point cloud context information. Background Art

[0002] In recent years, with the maturity of three-dimensional acquisition devices such as lidar, the acquisition of point clouds has become increasingly easy. Since point clouds have extensive applications in fields such as autonomous driving and virtual reality, the research on related algorithms for point clouds has also become more and more. Among them, point cloud semantic segmentation is an important research topic.

[0003] Point cloud semantic segmentation aims to infer the category of each point and is one of the important tasks in the field of computer vision. In recent point cloud semantic segmentation algorithms, PointNet++ [1] proposed an encoder-decoder architecture. In the encoder stage, the context information of the point cloud is obtained through multiple downsampling and neighborhood aggregation operations. In the decoder stage, the downsampled point cloud is gradually restored through multiple interpolation and skip connections, and finally the point cloud semantic segmentation result is obtained; PointNeXt [2] designed and stacked a large number of inverse residual multi-layer perceptron modules to obtain richer context information. Although these algorithms have achieved good results, these algorithms only utilize the context information within the neighborhood range of the point. Due to the sparse nature of the point cloud data itself, the number of points around some points is very small, resulting in the inability of the point to extract rich context information to complete the category inference. While in other point clouds, there may be many points around points of the same category, and rich context information can be extracted. By utilizing the cross-point cloud context information, the sparse points can obtain rich context information, thus achieving an accurate point cloud semantic segmentation result.

[0004] References:

[0005] [1] Qi C R, Yi L, Su H, et al. Pointnet++: Deep hierarchical feature learning on point sets in a metric space[J]. Advances in neural information processing systems, 2017, 30.

[0006] [2]Qian G,Li Y,Peng H,et al.PointNeXt:Revisiting PointNet++withImproved Training and Scaling Strategies[J].arXiv preprint arXiv:2206.04670,2022. Summary of the Invention

[0007] The present invention provides a point cloud semantic segmentation method based on cross-point cloud context information, which uses the cross-point cloud context information stored in the memory to enhance the point features in the point cloud and provide richer context information. This method can be used in most current point cloud semantic segmentation neural networks, can be plug-and-play, and achieve more accurate point cloud semantic segmentation results. The present invention is realized through the following technical solutions:

[0008] A point cloud semantic segmentation method based on cross-point cloud context information, the point cloud semantic segmentation network used is divided into two parts: an encoder and a decoder. The encoder consists of multiple stages, and each stage has a module for downsampling the point cloud, namely the SA module. The decoder also consists of multiple stages, and the number of stages is the same as that of the encoder. Each stage has a module for upsampling the point cloud, namely the FP module; its characteristics include the following steps:

[0009] First step, perform data preprocessing and data augmentation on the data set;

[0010] Second step, pass the preprocessed and data-augmented point cloud through the encoder. Each stage of the encoder will output the point cloud features of that stage. Denote the point cloud features output by the i-th stage of the encoder as For each point in , take the neighborhood, and calculate the local geometric features of each point according to the neighborhood. Denote the local geometric features of all points as For each semantic category, sample in and , and update the memory of this stage; the method is:

[0011] (1) Pass the preprocessed and data-augmented point cloud through the encoder. Denote the point cloud features output by the i-th stage of the encoder as where N represents the number of points and D represents the number of channels of the features;

[0012] (2) For each point in , take the k nearest points around it as the neighborhood of this point. Denote the k nearest points around it as {x 1 ,x 2 ,…,x k}, the normal vector estimation algorithm is used to calculate the normal vector of each point according to the neighborhood of each point, and the normal vector of each point is used as the local geometric feature of the point. Denote the local geometric features of all points as

[0013] (3) For each category, random sampling is performed in and Specifically, for category c, randomly sample the features of L points in to obtain In sample the local geometric features of the corresponding points to obtain

[0014] (4) Set the memory in the i-th stage, where L represents the capacity of the memory, C represents the number of categories, and D′ represents the number of channels of the features, D′ = D + 3. Split the context information of each category in the memory into two parts. For category c, denote the two parts as the point cloud feature and the local geometric feature Use the following method to update and :

[0015]

[0016]

[0017] Here, μ takes 0.9; the above update process is iterated once for each category. After the update, the cross-point cloud context information of this stage is stored in the memory;

[0018] Thirdly, pass the point cloud feature output by the encoder through the decoder. Each stage of the decoder will output the point cloud feature of this stage. Denote the point cloud feature output by the i-th stage of the decoder as After that, take the neighborhood of each point in and calculate the local geometric feature of each point according to this neighborhood. Denote the local geometric features of all points as After that, split the context information of each category in the memory of this stage into two parts: the point cloud feature and the local geometric feature, and then splice the corresponding parts of all categories to obtain and After that, use and as the position encoding, and use the attention mechanism between and to enhance to obtain the enhanced point cloud feature Input the enhanced point cloud feature of the last stage into the multi-layer perceptron to obtain the output

[0019] Step 4: Output the results and semantic labels of point clouds The cross entropy loss is calculated between the point cloud features output at each stage of the encoder The sampling anchor point in the corresponding stage of memory Positive and negative samples are sampled, and then the InfoNCE loss is calculated between the anchor points, positive samples and negative samples; finally, the cross entropy loss and the InfoNCE loss are weighted summed, and the back propagation algorithm and the gradient descent algorithm are used to optimize the parameters in the encoder and decoder;

[0020] Step 5: Complete the training of the entire neural network model;

[0021] In the sixth step, the point cloud to be tested is input into the trained neural network model to predict the probability distribution of each point, and the maximum probability value in the probability distribution of each point is taken as the predicted category of the point to obtain the final point cloud semantic segmentation result.

[0022] The beneficial effects of the technical solution provided by the present invention are:

[0023] Existing point cloud semantic segmentation methods only use the context information within the neighborhood of the point, which makes it difficult to obtain rich context information for the sparse points in the point cloud. The present invention adopts a point cloud semantic segmentation method based on cross-point cloud context information, and uses the cross-point cloud context information stored in the memory to enhance the features of the points in the point cloud, so that the sparse parts of the point cloud can obtain richer context information, thereby achieving more accurate point cloud semantic segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flow chart of a point cloud semantic segmentation method based on cross-point cloud context information;

[0025] Figure 2 and Figure 3 The experimental results of the method of the present invention are compared with the existing best method; DETAILED DESCRIPTION

[0026] The technical solution of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Based on the technical solution of the present invention, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present invention.

[0027] The point cloud semantic segmentation basic network adopted by the present invention is the point cloud semantic segmentation network involved in the document [1]. The point cloud semantic segmentation network is improved and the memory is increased.

[0028] The point cloud semantic segmentation network in Document 1 is divided into an encoder and a decoder. The encoder consists of multiple stages, with a SA module in each stage. The decoder also consists of multiple stages, and the number of stages is the same as that of the encoder, with an FP module in each stage. The SA module is used for point cloud downsampling. Given the input point cloud features, it first samples a batch of points using the Farthest Point Sampling algorithm, and this batch of points is called the center points. Then, it finds the k nearest points around each center point, inputs the features of these k points into a multi-layer perceptron for mapping, and then uses max pooling for aggregation to output the downsampled point cloud features. The FP module is used for point cloud upsampling. Given the input point cloud features and the downsampled point cloud features, for each point in the point cloud features, it finds the 3 nearest points in the downsampled point cloud features, aggregates their features, concatenates the aggregated features with the point features in the point cloud, and obtains the upsampled point cloud features through a multi-layer perceptron.

[0029] The datasets used for experimental feasibility verification in this invention are the S3DIS (Stanford Large-Scale 3D Indoor Spaces) dataset and the ShapeNetPart dataset respectively. The Stanford Large-Scale 3D Indoor Spaces dataset contains the point clouds of 271 rooms from 6 regions, with a total of 13 semantic categories. Among them, the point clouds of 203 rooms are used for training, and the point clouds of 68 rooms are used for testing. The ShapeNetPart dataset contains 16 types of objects, with a total of 16,880. Each type of object is divided into 2 to 6 parts, and all objects contain 50 parts in total. Among them, 14,006 objects are used for training, and 2,874 objects are used for testing.

[0030] First step, preprocess and augment the data of the S3DIS or ShapeNetPart dataset. For the training set of the S3DIS dataset, use the grid downsampling algorithm to first downsample the point cloud, and then randomly select 24,000 points from the downsampled point cloud as the result of preprocessing. After the preprocessing is completed, perform data augmentation. Geometrically, perform online data augmentation on the point cloud by means of point cloud scaling, point cloud rotation, and point cloud jittering; color-wise, perform online data augmentation on the point cloud by means of contrast enhancement and color dropout.

[0031] For the training set of the ShapeNetPart dataset, randomly select 2,048 points from the point cloud as the result of preprocessing. After the preprocessing is completed, perform data augmentation. Geometrically, perform online data augmentation on the point cloud by means of point cloud scaling and point cloud rotation. Data augmentation in terms of color is not used here.

[0032] In the second step, the preprocessed and data-augmented point cloud is passed through the encoder. Each stage of the encoder outputs the point cloud features of that stage. Denote the point cloud features output by the $i$-th stage of the encoder as After that, for each point in, take its neighborhood, and calculate the local geometric features of each point according to this neighborhood. Denote the local geometric features of all points as Finally, for each semantic category, sample in and and update the memory of this stage using momentum update. The specific method is as follows:

[0033] (1) Pass the preprocessed and data-augmented point cloud through the encoder. Denote the point cloud features output by the $i$-th stage of the encoder as where $N$ represents the number of points and $D$ represents the number of channels of the features.

[0034] (2) For each point in , take the $k$ nearest points around it as the neighborhood of this point. Denote the $k$ nearest points as $\{x 1 , x 2 , \cdots, x k \}$. Then use the normal vector estimation algorithm to calculate the normal vector of this point according to the neighborhood of each point:

[0035]

[0036]

[0037] Solve using the normal vector estimation algorithm. Obtain the normal vectors of all points, and take the normal vector of each point as the local geometric feature of this point. Denote the local geometric features of all points as Here $N$ represents the number of points.

[0038] (3) For each category, perform random sampling in and Specifically, for category $c$, randomly sample the features of $L$ points in to obtain Sample the local geometric features of the corresponding points in to obtain

[0039] (4) Set the memory at the $i$-th stage as where $L$ represents the capacity of the memory, $C$ represents the number of categories, and $D'$ represents the number of channels of the features. Split the context information of each category in the memory into two parts. For category $c$, denote the two parts as and ​Update and using momentum update:

[0040]

[0041]

[0042] Here, μ is taken as 0.9. The above update process is iterated once for each category. After the update, the cross-point cloud context information of this stage is stored in the memory.

[0043] In the third step, pass the point cloud features output by the encoder through the decoder. Each stage of the decoder will output the point cloud features of that stage. Denote the point cloud features output by the i-th stage of the decoder as After that, take the neighborhood of each point in , and calculate the local geometric features of each point according to this neighborhood. Denote the local geometric features of all points as After that, split the context information of each category in the memory of this stage into two parts, and then splice the corresponding parts of all categories to obtain and After that, use and as the position encoding, and use the attention mechanism between and to enhance to obtain the enhanced point cloud features Input the enhanced point cloud features of the last stage into the multi-layer perceptron to obtain the output The specific method is as follows:

[0044] (1) Pass the point cloud features output by the encoder through the decoder. Denote the point cloud features output by the i-th stage of the decoder as where N represents the number of points and D represents the number of channels of the features. For each point in , take the k nearest points around it as the neighborhood of this point. After that, use the normal vector estimation algorithm to calculate the normal vector of this point according to the neighborhood of each point. Take the normal vector of each point as the local geometric feature of this point, and denote the local geometric features of all points as Here, N represents the number of points.

[0045] (2) Split the cross-point cloud context information of each category in the corresponding memory of this stage into two parts. For category c, denote the two parts as and Then splice the corresponding parts of all categories to obtain and

[0046]

[0047]

[0048] (3) Use and as positional encodings, and use the attention mechanism between and to enhance to obtain enhanced point cloud features

[0049]

[0050]

[0051]

[0052]

[0053]

[0054] where linear() represents the linear mapping function, softmax() represents the normalized exponential function, Norm() represents the normalization function, and MLP() represents the multi-layer perceptron.

[0055] (4) Input the enhanced point cloud features of the last stage into the multi-layer perceptron to obtain the output

[0056] Fourthly, calculate the cross-entropy loss between the output result and the semantic label of the point cloud . Sample anchor points in the point cloud features output at each stage of the encoder, sample positive and negative samples in the memory corresponding to the corresponding stage, and then calculate the InfoNCE loss between the anchor points, positive samples and negative samples. Finally, perform a weighted sum of the cross-entropy loss and the InfoNCE loss, and use the backpropagation algorithm and the gradient descent algorithm to optimize the parameters in the encoder and decoder. The specific method is:

[0057] (1) Denote o j,c as the prediction score of the c-th category of the j-th point, and denote Y j,c as the true score of the c-th category of the j-th point, and then calculate the cross-entropy loss:

[0058]

[0059]

[0060] (2) The point cloud features output at each stage of the encoder Sample anchor points. For class c, the anchor points are denoted as where V represents the number of anchor points for class c, and D represents the number of channels; sample positive and negative samples in the memory at the corresponding stage. For class c, the positive samples are denoted as and the negative samples are denoted as where W represents the number of positive samples for class c, and W(C - 1) represents the number of negative samples for class c. Then calculate the InfoNCE loss:

[0061]

[0062]

[0063] (3) Perform weighted summation of the cross-entropy loss and the InfoNCE loss:

[0064]

[0065] Then use the backpropagation algorithm and the gradient descent algorithm to optimize the parameters in the encoder and decoder.

[0066] Step 5: Repeat steps 1 to 4 for 100 rounds to complete the training of the entire neural network model.

[0067] Step 6: Input the point cloud to be tested into the trained neural network model, predict the probability distribution of each point, and take the maximum probability value in the probability distribution of each point as the predicted class of the point to obtain the final point cloud semantic segmentation result.

[0068] The following verifies the feasibility of the method of the present invention in combination with specific examples:

[0069] Comparative experiments were carried out on two public datasets, namely the S3DIS (Stanford Large-Scale 3D Indoor Spaces) dataset and the ShapeNetPart dataset. The Stanford Large-Scale 3D Indoor Spaces dataset contains the point clouds of 271 rooms from 6 regions, with a total of 13 semantic classes. Among them, the point clouds of 203 rooms are used for training, and the point clouds of 68 rooms are used for testing; the ShapeNetPart dataset contains 16 types of objects, with a total of 16,880. Each object is divided into 2 to 6 parts, and all objects contain 50 parts in total. Among them, 14,006 objects are used for training, and 2,874 objects are used for testing.

[0070] ​On the S3DIS dataset, the mean intersection over union (mIoU), mean class accuracy (mAcc), and overall accuracy (OA) are used to quantitatively evaluate the results of point cloud semantic segmentation. On the ShapeNetPart dataset, the instance mean intersection over union (Ins.mIoU) and the category mean intersection over union (Cat.mIoU) are used to quantitatively evaluate the results of point cloud semantic segmentation.

[0071] According to Figure 2 and Figure 3 The experimental results of the proposed method and the existing state-of-the-art point cloud semantic segmentation methods on different datasets shown in

Claims

1. A point cloud semantic segmentation method based on cross-point cloud context information. The point cloud semantic segmentation network adopted is divided into an encoder and a decoder. The encoder consists of multiple stages, and each stage has a module for downsampling the point cloud, namely the SA module. The decoder also consists of multiple stages, and the number of stages is the same as that of the encoder. Each stage has a module for upsampling the point cloud, namely the FP module; Characterized in that, It includes the following steps: The first step is to perform data preprocessing and data augmentation on the data set; In the second step, the preprocessed and data-augmented point cloud is passed through the encoder. Each stage of the encoder outputs the point cloud features of that stage. Denote the point cloud features output by the $i$-th stage of the encoder as Take the neighborhood of each point in , and calculate the local geometric features of each point according to the neighborhood. Denote the local geometric features of all points as Sample each semantic category in and , and update the memory of this stage; the method is as follows: (1) Pass the preprocessed and data-augmented point cloud through the encoder. Denote the point cloud features output by the $i$-th stage of the encoder as where $N$ represents the number of points and $D$ represents the number of channels of the features; (2) For each point in , take the k nearest points in the vicinity as the neighborhood of this point. The k nearest points in the vicinity are denoted as {x 1 , x 2 , …, x k}}. Use the normal vector estimation algorithm to calculate the normal vector of this point according to the neighborhood of each point. Take the normal vector of each point as the local geometric feature of this point. Denote the local geometric features of all points as (3) For each category, random sampling is performed within and Specifically, for category c, L points' features are randomly sampled within to obtain The local geometric features of the corresponding points are sampled within to obtain (4) Set the memory in the i-th stage where L represents the capacity of the memory, C represents the number of categories, and D ′ represents the number of channels of the features, and D ′ = D + 3. Split the context information of each category in the memory into two parts. For category c, denote the two parts as the point cloud features and the local geometric features Update and in the following way: Here, μ is the weight coefficient; the above update process is iterated for each category once. After the update, the cross-point cloud context information of this stage is stored in the memory; Step 3: Pass the point cloud features output by the encoder through the decoder. Each stage of the decoder outputs the point cloud features of that stage. Denote the point cloud features output by the i-th stage of the decoder as After that, take the neighborhood of each point in and calculate the local geometric features of each point according to this neighborhood. Denote the local geometric features of all points as After that, split the context information of each category in the memory of this stage into two parts: point cloud features and local geometric features, and then splice the corresponding parts of all categories to obtain and After that, use and as the position encoding, and use the attention mechanism between and to enhance to obtain the enhanced point cloud features Input the enhanced point cloud features of the last stage into a multi-layer perceptron to obtain the output The fourth step is to determine the loss function and use the backpropagation algorithm and the gradient descent algorithm to optimize the parameters in the encoder and decoder; The fifth step is to complete the training of the entire neural network model; The sixth step is to input the point cloud to be tested into the trained neural network model, predict the probability distribution of each point, and take the maximum probability value in the probability distribution of each point as the predicted category of this point to obtain the final point cloud semantic segmentation result.

2. The point cloud semantic segmentation method according to claim 1, Characterized in that, μ is taken as 0.

9.

3. The point cloud semantic segmentation method according to claim 1, Characterized in that, The method for determining the loss function is as follows: Calculate the cross-entropy loss between the output result and the semantic labels of the point cloud ; Sample anchor points from the point cloud features output at each stage of the encoder , sample positive and negative samples from the memory at the corresponding stage , and then calculate the InfoNCE loss between the anchor points, positive samples, and negative samples; Perform a weighted sum of the cross-entropy loss and the InfoNCE loss.

Citation Information

Patent Citations

  • Semantic segmentation method for point cloud in outdoor large scene

    CN112560865A

  • End-to-end three-dimensional point cloud registration method

    CN114332176A