A context- and attention-based 3D point cloud semantic segmentation method
By introducing context and attention mechanisms into three-dimensional point cloud semantic segmentation, combining relational shape networks and multi-layer perceptron classifiers, the challenge of semantic segmentation of three-dimensional point cloud models is solved, achieving a more efficient and robust segmentation effect.
Patent Information
- Application Number
- CN202210221944.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-09
AI Technical Summary
The existing technology is difficult to effectively solve the semantic segmentation problem of three-dimensional point cloud models, especially when dealing with problems such as disorder, sparseness, downsampling, noise, etc. of point cloud data, and lacks robust and efficient segmentation methods.
A three-dimensional point cloud semantic segmentation method based on context and attention is adopted to extract point cloud features through a relational shape network, and the features are constrained and strengthened by context and attention modules. Finally, a multi-layer perceptron classifier is used for semantic segmentation.
The semantic segmentation accuracy of the three-dimensional point cloud model is improved, especially the segmentation effect in edge areas, enhances the robustness and adaptability of the model, and can process complex point cloud data more effectively.
Smart Images

Figure CN114693923B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer image processing and computer graphics, and in particular relates to a context- and attention-based three-dimensional point cloud semantic segmentation method. Background Art
[0002] In recent years, with the continuous development and popularization of 3D data acquisition equipment, 3D model data has exploded, and it has also attracted researchers' interest in understanding and processing 3D model data. The main forms of 3D models are point clouds, voxels, patches, etc. Among them, due to the many advantages of point cloud data, such as being easily acquired through simple equipment and being insensitive to external factors such as lighting, the analysis of 3D point cloud models has become a hot research field. However, point cloud data also has some characteristics: irregular, disordered, and relatively sparse. These characteristics make it very difficult to process and understand point cloud data. At present, deep learning technology has achieved excellent results in the field of 2D images. However, unlike 2D images that naturally have positional structures, the disorder of 3D point cloud models makes it impossible to directly apply convolution operations on 2D images to 3D point cloud models, which makes it difficult to apply deep learning methods to the analysis of 3D models.
[0003] Although the semantic segmentation problem of 3D point cloud models is very basic, it is very challenging for the following reasons:
[0004] 1. Point clouds belonging to the same part must be correctly labeled with the same semantic label;
[0005] 2. Global and local features must be effectively aggregated and analyzed to achieve better segmentation results;
[0006] 3. The analysis method must be robust to downsampling, noise, and diversity of similar models.
[0007] In recent years, many methods have emerged in the field of 3D point cloud semantic segmentation, which can be roughly divided into the following four categories: multi-layer perceptron-based methods, point cloud convolution-based methods, recurrent neural network-based methods, graph-based methods, etc.
[0008] Methods based on multi-layer perceptrons use shared multi-layer networks to share parameters. For example, references 1C.R.Qi, H.Su, K.Mo, and LJGuibas.PointNet: Deep Learning on Point Sets for 3DClassification and Segmentation.2017., reference 2C.R.Qi, L.Yi, H.Su, and L.J.Guibas.Pointnet++: Deep hierarchical feature learning on point sets in ametric space.Advances in neural information processing systems, 2017, 30., etc. use shared multi-layer perceptrons to extract features from each point cloud information by fusing multi-scale information, but it is difficult for shared multi-layer perceptrons to focus on the local geometric connections of point clouds.
[0009] The point cloud convolution-based method extracts point cloud features by directly performing convolution operations on the input point cloud data. For example, reference 3 S.B.Hua, KMTran, and KSYeung.Pointwise convolutional neural networks.Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2018:984-993., reference 4 Y.Li, R.Bu, M.Sun, W.Wu, X.Di. and B.Chen, Pointcnn:Convolution on x-transformed points.Advances in neural information processing systems.2018;31. proposes a method of using point-by-point convolution for point clouds, by performing sliding convolution calculations on the entire point cloud area and giving each point cloud within the convolution kernel the same weight. Reference 5 H. Thomas, C. R. Q., J. E. Deschaud, B. Marcotegui, and Goulette. Kpconv: Flexible and deformable convolution for point clouds. Proceedings of the IEEE / CVF international conference on computer vision. 2019: 6411-6420. It is proposed to obtain the value of the kernel transformation matrix by establishing a distribution instead of calculating the similarity, thereby realizing the dot product. Reference 6 Y. Liu, B. Fan, S. Xiang, C. Pan. Relation-shape convolutional neural network for point cloud analysis. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 8895-8904. The feature representation capability is enhanced by constructing a local neighborhood shape convolution. Compared with the traditional method of constructing a kernel transformation matrix, this method better adapts to the complex geometric shape changes of point clouds.
[0010] The recurrent neural network-based method can improve the segmentation accuracy by recursively transferring and utilizing the contextual features implicitly present in the point cloud and using these features to enhance the feature representation capability of the point cloud. Document 7Z. Zhao, M. Liu, K. Ramani. DAR-Net: Dynamic aggregation network for semantic scenesegmentation. arXiv preprint ar learningframework for semantic parsing of large-scale 3D point clouds.Proceedings of the IEEE international conference on computer vision.2017:5678-5687., Document 9 vision(ECCV).2018:403-417. etc. designed a dynamic feature aggregation method to fuse local and global features.
[0011] The graph-based method first determines the adjacency relationship of all points in the point cloud model according to the position of the point cloud, and constructs the point cloud data into a graph-structured data. As a more natural data structure, the graph is very suitable for processing irregular data such as point clouds. Reference 10Y.Shen, C.Feng, Y.Yang, and D.Tian.Mining point cloud localstructures by kernel correlation and graph pooling.Proceedings of the IEEE conference on computer vision and pattern recognition.2018:4548-4557. The adjacency relationship of the point cloud set is determined by the geometric similarity of the kernel correlation measure, and convolution is implemented on each node and its neighboring nodes. Document 11D.Boscaini, J.Masci, S.Melzi, MMBronstein, U.Castellani, andP.Vandergheynst.Learning class-specific descriptors for deformable shapes using localized spectral convolutional networks.Computer Graphics Forum.2015,34(5):13-23., Document 12L.Yi,H.Su,X.Guo,and JLGuibas.Syncspeccnn:Synchronizedspectral cnn for 3d shape segmentation.Proceedings of the IEEE Conference onComputer Vision and Pattern Recognition.2017:2282-2290., Document 13D.K.Hammond,P.Vandergheynst,R.Gribonval.Wavelets on graphs via spectral graphtheory.Applied and Computational Harmonic Analysis, 2011, 30(2): 129-150. et al. define convolution on graphs in the spectral domain. However, these methods usually require the calculation of a large number of parameters.
[0012] Recently, attention mechanisms have been widely used in various fields such as machine translation, object detection, and semantic segmentation. In the field of 3D model segmentation, graph convolutional neural networks first introduced the attention mechanism. References 14 L. Wang, Y. Huang, Y. Hou, S. Zhang, and J. Shan. Graph attention convolution for point clouds semantic segmentation. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 10296-10305., Reference 15 J. Yang, Q. Zhang, B. Ni, L. Li, J. Liu, M. Zhou, and Q. Tian. Modeling point clouds with self-attention and gumbel subset sampling. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 3323-3332. etc. understand point clouds by constructing a point cloud self-attention transformation network. In addition, contextual information has also become the focus of research related to 3D point clouds. Reference 16M. Defferrard, X. Bresson, P. Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. Advances in neural information processing systems, 2016, 29., Reference 17G. Yu, K. Liu, Y. Zhang, C. Zhu, and K. Xu. Partnet: Arecursive part decomposition network for fine-grained and hierarchical shapes segmentation. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 9491-9500. etc. respectively use graph convolution and recurrent neural networks to utilize contextual information to enhance feature representation capabilities.However, these methods embed context or attention into deep networks, thus limiting the universality of these modules. Summary of the invention
[0013] Purpose of the invention: The technical problem to be solved by the present invention is to provide a context- and attention-based 3D point cloud semantic segmentation method in view of the deficiencies of the prior art, comprising the following steps:
[0014] Step 1, collecting data from the input 3D point cloud model data set;
[0015] Step 2, extracting point-by-point features of the point cloud data through a relational shape network to obtain point cloud features containing shape information;
[0016] Step 3: Use the context and attention modules to constrain and strengthen the extracted point cloud features within and between classes, and obtain point cloud features with contextual priors and global semantic associations;
[0017] Step 4: Use a multi-layer perceptron classifier to classify the point cloud features and obtain the final predicted label for each point in the point cloud data.
[0018] Step 1 includes the following steps:
[0019] Step 1-1: Input 3D point cloud model dataset S = {S Train , S Test} is divided into training set S Train ={s1, s2, ...s i , ..., s n} and test set S Test ={s n+1 ,s n+2 , ..., s n+ j,...,s n+m}, where s i represents the i-th model in the training set, s n+j represents the jth model in the test set;
[0020] Step 1-2, set the input single 3D point cloud model s i (records the coordinates of all points of the 3D model, which is taken from the ShapeNet standard 3D point cloud model semantic segmentation dataset containing 16 types of 3D models) and the label set l for the parts to which all points belong i (The label of the component type to which each point of the model belongs is recorded. There are 50 components in this data set). N points are randomly sampled from all point cloud data as the network input point set P. i ={p1, p2, ... p i , ..., p N}, from the label set li Take out the point P with the i-th point i The corresponding labels form a new label set g i , i takes values of 1 to N; the data set in step 1-1 is sampled to obtain a new data set P = {P Train , P Test}, so that the characteristic shapes of different point cloud models can be kept consistent during the network segmentation process. The experiment found that sampling N points can effectively take into account the performance of the hardware GPU and; P Train represents the sampled point cloud training set, P Test Represents the sampled point cloud test set;
[0021] Step 1-3, for the training set P obtained in step 1-2 Train Random scaling and translation are performed, where the scaling factor u is sampled from a uniform distribution U(0.8, 1.25) and the translation amount is sampled from a uniform distribution U(-0.1, 0.1).
[0022] Wherein, step 1-2 comprises the following steps:
[0023] Step 1-2-1, for a single 3D point cloud model i , whose point cloud set is s i ={s i1 ,s i2 ,...s ij ,..,s in}, where s ij Represents point cloud model s i The j-th point data, j is 1 to n; sampling with replacement is performed from the index set Q = {1, 2, ..., n}, and repeated N times to obtain the sampled index set Q1 = {q1, q2, ...q k , ..., q N},i k ∈I, where q k represents the index of the kth sampling from the set Q;
[0024] Step 1-2-2, collect the point cloud in step 1-2-1 i The point cloud subscript in Q1 and the point cloud corresponding to the element in Q1 are added to the sampling point set P to obtain the new point cloud model data P i ={p1, p2, ... p k , ..., p N}, where p k For step 1-2-1 ij j take q k ,Right now
[0025] Step 1-2-3, repeat steps 1-2-1 and 1-2-2 until all 3D point cloud models in the training set have been sampled.
[0026] In steps 1-3, the coordinates of each point cloud data, that is, the first three dimensions of the point cloud data, are randomly scaled and translated, which can improve the model training effect and robustness.
[0027] Step 2 includes the following steps:
[0028] Step 2-1, for the sampled point cloud training set P Train = {P1, P2, ... P i , ..., P n}, collect the true labels G of each point Train = {G1, G2, ... G i , ..., G n} and point cloud data are input into the relational shape network for training, and the encoder extracts high-dimensional point cloud features, where P i Refers to the data of the i-th point cloud model, G i Refers to the true label set of each point in the i-th point cloud model;
[0029] Step 2-2, upsample and decode the point cloud features extracted in step 2-1 to obtain point cloud features that conform to the input shape and contain relationship information. Use bilinear interpolation to gradually increase the point cloud cardinality until the input shape N is reached, and finally obtain an N×512-dimensional feature matrix.
[0030] Wherein, step 2-1 comprises the following steps:
[0031] Step 2-1-1, for a single point cloud model data P i , group the point cloud data according to the farthest point sampling strategy, and iteratively select the point with the largest Euclidean distance to all point cloud data as the sphere center to obtain the point cloud grouping PG i ={pg1, pg2, ..., pg i , ..., pg m}, where pg i ={p i1 ,..,p ik ,..p in}, pg i represents the i-th point cloud group, p ik Indicates pg i For the kth point in , the farthest point sampling can maximize the coverage of the sampling point cloud on the original point cloud data;
[0032] Step 2-1-2, PG in step 2-1-1 iAfter the forward propagation convolution operation, it is extracted into an m×512-dimensional feature matrix f i ;
[0033] Step 2-1-3, repeat steps 2-1-1 and 2-1-2 for a total of 3 times, in each repetition, m is 512, 128, 1, n is 32, 32, 128, respectively, to obtain the first stage of point cloud grouping PG i -1. Point cloud grouping PG in the second stage i -2. The third stage of point cloud grouping PG i -3 and the point cloud feature matrix f of the first stage i -1. The point cloud feature matrix f of the second stage i -2. The point cloud feature matrix f of the third stage i -3.
[0034] Step 3 includes the following steps:
[0035] Step 3-1: For a single point cloud model data P i And its corresponding true label G, the feature matrix obtained in step 2 is passed through the context module to obtain the intra-class feature matrix and inter-class feature matrix with context prior knowledge;
[0036] In step 3-2, the intra-class feature matrix and the inter-class feature matrix obtained in step 3-1 are enhanced by the self-attention module, and the global dependency is modeled to obtain point cloud features with contextual priors and global semantic associations.
[0037] Wherein, step 3-1 comprises the following steps:
[0038] Step 3-1-1: For the N×512-dimensional feature matrix obtained in step 2, use 1x1 convolution operation to reduce the dimension to N×256 dimensions to obtain a new feature matrix F. By multiplying it with its transposed matrix, we can obtain an N×N-dimensional intra-class feature matrix M and an inter-class feature matrix IM, where I represents the unit matrix; aggregate the intra-class features and inter-class features to obtain the feature matrix F containing context priors. e ,Right now:
[0039] F e =concat(M,(IM)F)
[0040] Among them, concat means concatenating and aggregating the features in the last dimension.
[0041] Step 3-1-2: For the true label G in step 3-1, obtain the N×N dimensional covariance matrix C and calculate the difference between M and C. As part of Loss, the specific calculation formula is as follows:
[0042]
[0043]
[0044]
[0045]
[0046] in, They represent the accuracy within a class, the recall within a class, and the specificity between classes respectively; c ij represents the (i, j) element of the matrix C, m ij represents the (i, j) element of the matrix M, μ is a non-negative minimum value, and in the present invention, μ=0.0001 is set based on experience to control the overflow caused by the divisor being all 0 during the network training process.
[0047] Calculate the learning context matrix, that is, the intra-class feature matrix M (shape is N×N, m n ∈M, n∈[1, N 2 ]) and the matrix C (shape N×N, c n ∈C, n∈[1, N 2 ]) And finally the final context loss is obtained by weighting the two losses The specific calculation formula is as follows:
[0048]
[0049]
[0050] Among them, λ u and λ g Represents their respective weight values. In the present invention, λ u and λ g Set to 1.
[0051] In step 3-2, the self-attention module uses 8-head attention and performs a self-attention on the feature matrix F containing context prior obtained in step 3-1-1. e The dataset is divided into 8 small subsets, and the self-attention matrix is calculated for each subset, and finally summarized into a global attention matrix with overall attention relationship. The global relationship is modeled and strengthened through the self-attention mechanism to obtain the final feature matrix.
[0052] In step 4, the feature matrix obtained in step 3 is passed through a fully connected layer, and finally a Softmax multi-classifier is used to perform multi-label prediction on the input multi-dimensional feature vector to obtain a probability map of semantic segmentation of the point cloud data. The label with the highest predicted probability for each point in the point cloud data is used as the predicted label of the point, and the corresponding true label G i Compare and calculate semantic segmentation loss Same as in step 3-1-2 Add up as the total loss After back propagation, we finally get the trained point cloud segmentation network containing contextual prior knowledge. The specific calculation formula is as follows.
[0053]
[0054]
[0055] Among them, w is the corresponding weight, c is the category, and x is the network output prediction label.
[0056] The method of the present invention is dedicated to solving the problem of segmenting 3D point cloud models into labeled semantic parts. Analyzing and reasoning the model based on the components of the point cloud model is widely used in computer vision, robotics and virtual reality, such as hybrid model analysis, target detection and tracking, 3D reconstruction, style transfer, robot roaming and grasping, etc., which makes this work very meaningful.
[0057] Beneficial effects: The method of the present invention is inspired by first using a relational shape network to extract point cloud features, and then introducing contextual prior knowledge through a context-attention module to constrain the features, and obtain a feature matrix with intra-class and inter-class relationships. Finally, the classifier is used to predict components on the complete feature map to obtain the final semantic segmentation map. In the whole process, this method is embedded in the general point cloud feature extraction backbone network, and integrates prior semantic contextual knowledge to prompt the network to clarify the boundaries of different categories of point cloud components. After being strengthened by the self-attention module, the effect of point cloud semantic segmentation and annotation is further improved. The entire method system is efficient and practical. The method of the present invention optimizes the segmentation effect of the edge area of the components in the general point cloud segmentation process, which not only ensures the overall segmentation accuracy, but also improves the edge details. In addition, the method of the present invention designs a context module that can be easily embedded, which can be widely applied to common point cloud segmentation networks, helping the network to further improve the results of semantic segmentation and annotation of three-dimensional point cloud models. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0059] Figure 1a The original model without segmentation.
[0060] Figure 1b Coloring rendering results for labels after semantic segmentation.
[0061] Figure 2 It is the overall network framework diagram of the method of the present invention.
[0062] Figure 3 This is a framework diagram of the context-attention module in the present invention.
[0063] Figure 4 This is a rendering of the semantic segmentation effect of the method of the present invention on the ShapeNetPart dataset.
[0064] Figure 5 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0065] like Figure 5 As shown, the present invention discloses a context-attention-based 3D point cloud semantic segmentation method, which collects N point cloud data and corresponding component labels in a 3D model to be segmented; inputs the point cloud data and component labels of a model training set into a network model for training, and inputs the point cloud data of a model test set into the trained network model to obtain component prediction labels for all points; the 3D point cloud model is grouped using farthest point sampling in the segmentation network, so that the coverage of the grouped point cloud in the original point cloud data is maximized; a convolution operation based on the Euclidean distance between the point and the sphere center point cloud coordinates is performed on the point cloud data of each group to obtain a point cloud feature matrix with relationship information; the corresponding context feature map is calculated through a context-attention module, and supervised training is performed using a feature map generated by a priori component labels; an attention matrix is calculated for the feature map obtained by the context prior to strengthen the relationship between classes and the global dependency; the probability of the point cloud prediction being each component is calculated by a classifier, and the maximum value is selected as the final predicted component label.
[0066] For a given 3D point cloud model dataset S = {S Train , S Test}, divided into training set S Train ={s1, s2, ...s i , ..., s n} and test set S Test ={s n+1 ,s n+2 , ..., s n+j , ..., s n+m}, where s i represents the i-th model in the training set, s n+jRepresents the jth model in the test set. The present invention completes the semantic segmentation and annotation of the models in the data set through the following steps. The target task is as follows: Figure 1a As shown in the flow chart, Figure 2 and Figure 5 As shown:
[0067] The specific steps include:
[0068] Step 1, collecting data from the input 3D point cloud model data set;
[0069] Step 2: The relational shape network extracts point-by-point features from the point cloud data to obtain point cloud features containing shape information;
[0070] Step 3: Use the context-attention module to constrain and strengthen the extracted point cloud features within and between classes, and obtain point cloud features with contextual priors and global semantic associations;
[0071] Step 4: Use a multi-layer perceptron classifier to classify the point cloud features and obtain the final predicted label for each point in the point cloud data.
[0072] Step 1 includes the following steps:
[0073] Step 1-1: Input 3D point cloud model dataset S = {S Train , S Test} is divided into training set S Train ={s1, s2, ...s i , ..., s n} and test set S Test ={s n+1 ,s n+2 , ..., s n+j , ..., s n+m}, where s i represents the i-th model in the training set, s n+j represents the jth model in the test set;
[0074] Step 1-2, assuming that a single 3D point cloud model s is input i (records the coordinates of all points of the 3D model, which is taken from the ShapeNetPart standard 3D point cloud model semantic segmentation dataset containing 16 types of 3D models) and the label set li for the parts to which all points belong (records the label of the type of part to which each point of the model belongs, and there are 50 types of parts in this dataset), randomly sample N points from all point cloud data as the network input point set P i ={p1, p2, ... p i , ..., p N}, from the label set l i Remove the Pi The corresponding labels form a new label set g i , the data set in step 1-1 is sampled to obtain a new data set P = {P Train , P Test}, so that the characteristic shapes of different point cloud models can be consistent during the network segmentation process. The experiment found that sampling N = 2048 points can effectively take into account the performance of the hardware GPU;
[0075] Step 1-3, for the training set P obtained in step 1-2 Train Random scaling and translation are performed, where the scaling factor u is sampled from a uniform distribution U(0.8, 1.25) and the translation amount is sampled from a uniform distribution U(-0.1, 0.1).
[0076] Wherein, step 1-2 comprises the following steps:
[0077] Step 1-2-1, for a single 3D point cloud model i , whose point cloud set is s i ={s i1 ,s i2 ,...s ij ,..,s in}, where s ij Represents point cloud model s i The j-th point data is sampled with replacement from the index set Q = {1, 2, ..., n}, and repeated N times to obtain the sampled index set Q1 = {q1, q2, ...q k , ..., q N},i k ∈I, where q represents the index of the kth sampling from the set Q;
[0078] Step 1-2-2, collect the point cloud in step 1-2-1 i The point cloud subscript in Q1 and the point cloud corresponding to the element in Q1 are added to the sampling point set P to obtain the new point cloud model data P i ={p1, p2, ... p k , ..., p N}, where p k For step 1-2-1
[0079] Step 1-2-3, repeat steps 1-2-1 and 1-2-2 until all 3D point cloud models in the training set have been sampled.
[0080] In steps 1-3, the coordinates of each point cloud data, that is, the first three dimensions of the point cloud data, are randomly scaled and translated, which can improve the model training effect and robustness.
[0081] Step 2 includes the following steps:
[0082] Step 2-1, for the sampled point cloud training set P Train = {P1, P2, ... P i , ..., P n}, collect the true labels G of each point Train ={G1, G2, ... Gi, ..., G n} and point cloud data are input into the relational shape network for training, and the encoder extracts high-dimensional point cloud features, where P i Refers to the data of the i-th point cloud model, G i Refers to the true label set of each point in the i-th point cloud model;
[0083] Step 2-2, upsample and decode the point cloud features extracted in step 2-1 to obtain point cloud features that conform to the input shape and contain relationship information. Use bilinear interpolation to gradually increase the point cloud cardinality until the input shape N is reached, and finally obtain an N×512-dimensional feature matrix.
[0084] Wherein, step 2-1 comprises the following steps:
[0085] Step 2-1-1, for a single point cloud model data P i , group the point cloud data according to the farthest point sampling strategy, and iteratively select the point with the largest Euclidean distance to all point cloud data as the sphere center to obtain the point cloud grouping PG i ={pg1, pg2, ..., pg i , ..., pg m}, where pg i ={p i1 ,..,p ik ,..p in}, represents the i-th point cloud group, p ik Indicates pg i For the kth point in , the farthest point sampling can maximize the coverage of the sampling point cloud on the original point cloud data;
[0086] Step 2-1-2, PG in step 2-1-1 i After the forward propagation convolution operation, it is extracted into an m×512-dimensional feature matrix f i ;
[0087] Step 2-1-3, repeat steps 2-1-1 and 2-1-2 for a total of 3 times. In each repetition, m is 512, 128, and 1, and n is 32, 32, and 128 respectively. PG is obtained respectively. i -1.PG i -2.PGi -3 and f i -1, f i -2, f i -3, forming a feature matrix under different scale groups.
[0088] Step 3 includes the following steps:
[0089] Step 3-1: For a single point cloud model data P i And its corresponding true label G, the feature matrix obtained in step 2 is passed through the context module to obtain the intra-class feature matrix and inter-class feature matrix with context prior knowledge;
[0090] In step 3-2, the intra-class feature matrix and the inter-class feature matrix obtained in step 3-1 are enhanced by the self-attention module, and the global dependency is modeled to obtain point cloud features with contextual priors and global semantic associations.
[0091] Wherein, step 3-1 comprises the following steps:
[0092] Step 3-1-1: For the N×512-dimensional feature matrix obtained in step 2, use 1x1 convolution operation to reduce the dimension to N×256 dimensions to obtain a new feature matrix F. By multiplying it with its transposed matrix, we can obtain an N×N-dimensional intra-class feature matrix M and an inter-class feature matrix IM, where I represents the unit matrix; aggregate the intra-class features and inter-class features to obtain the feature matrix F containing context priors. e ,Right now:
[0093] F e =concat(M,(IM)F)
[0094] Among them, concat means concatenating and aggregating the features in the last dimension.
[0095] Step 3-1-2: For the true label G in step 3-1, obtain the N×N dimensional covariance matrix C and calculate the difference between M and C. As part of Loss, the specific calculation formula is as follows:
[0096]
[0097]
[0098]
[0099]
[0100] in, They represent the accuracy within the class, the recall within the class, and the specificity between classes respectively; cij represents the (i, j) element of the matrix C, m ij represents the (i, j) element of the matrix M, μ is a non-negative minimum value, and in the present invention, μ=0.0001 is set based on experience to control the overflow caused by the divisor being all 0 during the network training process.
[0101] Calculate the learned context matrix, i.e., the intra-class feature M (shape is N×N, m n ∈M, n∈[1, N 2 ]) and the matrix C (shape N×N, c n ∈C, n∈[1, N 2 ]) And finally the final context loss is obtained by weighting the two losses The specific calculation formula is as follows:
[0102]
[0103]
[0104] Among them, λ u and λ g Represents their respective weight values. In the present invention, λ u and λ g Set to 1.
[0105] In step 3-2, the self-attention mechanism uses 8-head attention and performs a self-attention on the feature matrix F containing context priors obtained in step 3-1-1. e The dataset is divided into 8 small subsets, and the self-attention matrix is calculated for each subset, and finally summarized into a global attention matrix with overall attention relationship. The global relationship is modeled and strengthened through the self-attention mechanism to obtain the final feature matrix.
[0106] In step 4, the feature matrix obtained in step 3 is passed through a fully connected layer, and finally a Softmax multi-classifier is used to perform multi-label prediction on the input multi-dimensional feature vector to obtain a probability map of semantic segmentation of the point cloud data. The label with the highest predicted probability for each point in the point cloud data is used as the predicted label of the point, and the corresponding true label G i Compare and calculate semantic segmentation loss Same as in step 3-1-2 Add up as the total loss After back propagation, we finally get the trained point cloud segmentation network containing contextual prior knowledge. The specific calculation formula is as follows.
[0107]
[0108]
[0109] Among them, w is the corresponding weight, c is the category, and x is the network output prediction label. Test Input into the trained network model to obtain the semantic segmentation annotations of all point clouds in the training set.
[0110] Example
[0111] The objectives and tasks of the present invention are as follows Figure 1a and Figure 1b As shown, Figure 1a is the original model without segmentation, Figure 1b The label coloring rendering result after semantic segmentation, the network structure of the whole method is as follows Figure 2 As shown, Figure 3 It is a detailed schematic diagram of the core context-attention module. The following describes the various steps of the present invention according to the embodiments.
[0112] Step (1) collects data for the input 3D point cloud model dataset S. It is specifically divided into the following steps:
[0113] Step (1.1), the input 3D point cloud model dataset S = {S Train , S Test} is divided into training set S Train ={s1, s2, ...s i , ..., s n} and test set S Test ={s n+1 ,s n+2 , ..., s n+j , ..., s n+m}, where s i represents the i-th model in the training set, s n+j represents the jth model in the test set;
[0114] Step (1.2), input a single 3D point cloud model s i And the label set l of the components to which all points belong i , randomly sample N points from all point cloud data as the network input point set P i ={p1, p2, ... p i , ..., p N}, from the label set l i Remove the P i The corresponding labels form a new label set g i , the data set in step 1-1 is sampled to obtain a new data set P = {P Train , P Test}, so that the characteristic shapes of different point cloud models can be consistent during the network segmentation process; this step can be specifically divided into the following steps:
[0115] Step (1.2.1), for a single 3D point cloud model s i , whose point cloud set is s i ={s i1 ,s i2 ,...s ij ,..,s in}, where s ij Represents point cloud model s i The j-th point data is sampled with replacement from the index set Q = {1, 2, ..., n}, and repeated N times to obtain the sampled index set Q1 = {q1, q2, ...q k , ..., q N},i k ∈I, where q represents the index of the kth sampling from the set Q;
[0116] Step (1.2.2) collects the point cloud s in step (1.2.1) i The point cloud subscript in Q1 and the point cloud corresponding to the element in Q1 are added to the sampling point set P to obtain the new point cloud model data P i ={p1, p2, ... p k , ..., p N}, where p k For step 1-2-1
[0117] Step (1.2.3), repeat steps (1.2.1) and (1.2.2) until all 3D point cloud models in the training set have been sampled.
[0118] Step (1.3), for the training set P obtained in step 1-2 Train Random scaling and translation are performed, where the scaling factor u is sampled from the uniform distribution U(0.8, 1.25) and the translation amount is sampled from the uniform distribution U(-0.1, 0.1). Specifically, the coordinates of each point cloud data are implemented, that is, random scaling and translation are performed on the first 3 dimensions of the point cloud data.
[0119] Step (2), using a relational shape network to extract point-by-point features from the point cloud data to obtain point cloud features containing shape information;
[0120] Step (2.1), for the sampled point cloud training set P Train , collect the true labels G of each point TrainThe point cloud data is input into the relational shape network for training, and the high-dimensional point cloud features are extracted through the encoder; this step can be specifically divided into the following steps:
[0121] Step (2.1.1), for a single point cloud model data P i , group the point cloud data according to the farthest point sampling strategy, and iteratively select the point with the largest Euclidean distance to all point cloud data as the sphere center to obtain the point cloud grouping PG i ={pg1, pg2, ..., pg i , ..., pg m}, where pg i ={p i1 ,..,p ik ,..p in}, represents the i-th point cloud group, p ik Indicates pg i The kth point in
[0122] Step (2.1.2), PG in step (2.1.1) i After the forward propagation convolution operation, it is extracted into an m×512-dimensional feature matrix f i ;
[0123] Step (2.1.3), repeat step (2.1.1) and step (2.1.2) for 3 times, in each repetition, m is 512, 128, 1, n is 32, 32, 128, respectively, and PG is obtained respectively. i -1.PG i -2.PG i -3 and f i -1, f i -2, f i -3.
[0124] Step (2.2), upsample and decode the point cloud features extracted in step (2.1), and use a bilinear interpolation strategy to upsample the point cloud features to N×512 dimensions, that is, point cloud features that conform to the input shape and contain relationship information.
[0125] In step (3), the context-attention module is used to constrain and strengthen the extracted point cloud features within and between classes, so as to obtain point cloud features with contextual priors and global semantic associations.
[0126] Step (3.1), for a single point cloud model data P i And its corresponding true label G, the feature matrix obtained in step 2 is passed through the context module to obtain the intra-class feature matrix and inter-class feature matrix that learn the context prior knowledge; this step can be specifically divided into the following steps:
[0127] Step (3.1.1): For the N×512-dimensional feature matrix obtained in step (2), use a 1x1 convolution operation to reduce the dimension to N×256 dimensions to obtain a new feature matrix F. By multiplying it with its transposed matrix, we obtain an N×N-dimensional intra-class feature matrix M and an inter-class feature matrix IM, where I represents the unit matrix; aggregate the intra-class features and inter-class features to obtain the feature matrix F containing context priors e ,Right now:
[0128] F e =concat(M,(IM)F);
[0129] Among them, concat means concatenating and aggregating the features in the last dimension.
[0130] Step (3.1.2), for the true label G in step (3.1), obtain the N×N dimensional covariance matrix C and calculate the difference between M and C As part of Loss, the specific calculation formula is as follows:
[0131]
[0132]
[0133]
[0134]
[0135] in, They represent the accuracy within the class, the recall within the class, and the specificity between classes respectively; c ij represents the (i, j) element of the matrix C, m ij represents the (i, j) element of the matrix M, μ is a non-negative minimum value, and in the present invention, μ=0.0001 is set based on experience to control the overflow caused by the divisor being all 0 during the network training process.
[0136] Calculate the learned context matrix, i.e., the intra-class feature M (shape is N×N, m n ∈M, n∈[1, N 2 ]) and the matrix C (shape N×N, c n ∈C, n∈[1, N 2 ]) And finally the final context loss is obtained by weighting the two losses The specific calculation formula is as follows:
[0137]
[0138]
[0139] Among them, λ u and λ g Represents their respective weight values. In the present invention, λ u and λ g Set to 1.
[0140] Step (3.2): The intra-class feature matrix and inter-class feature matrix obtained in step 3-1 are enhanced by the self-attention module to model the global dependency relationship and obtain point cloud features with contextual priors and global semantic associations.
[0141] Step (4) uses a multi-layer perceptron classifier to classify the point cloud features and obtain the final predicted label of each point in the point cloud data. The feature matrix obtained in step 3 is passed through a multi-layer perceptron, and finally a Softmax multi-classifier is used to perform multi-label prediction on the input multi-dimensional feature vector to obtain a probability map of semantic segmentation of the point cloud data. The label with the highest predicted probability for each point in the point cloud data is used as the predicted label of the point, and the corresponding true label G i Compare and calculate semantic segmentation loss Same as in step (3.1.2) Add up as the total loss After back propagation, we finally get the trained point cloud segmentation network containing contextual prior knowledge. The specific calculation formula is as follows.
[0142]
[0143]
[0144] Among them, w is the corresponding weight, c is the category, and x is the network output prediction label.
[0145] Results Analysis
[0146] The experimental environment parameters of the method of the present invention are as follows:
[0147] The experimental platform parameters for data collection and training and testing of the point cloud segmentation network with contextual prior integration are Windows 10 64-bit operating system, Intel(R) Core(TM) i7-5820K CPU 3.30GHz, 64GB memory, and a graphics card of Titan X GPU 12GB. The Python programming language is used, and the Pytorch third-party open source library is used to implement it.
[0148] The comparative experimental results of the method of the present invention and the classic point cloud semantic segmentation methods: the method in document 1 (referred to as PointNet), the method in document 2 (referred to as PointNet++), the method in document 4 (referred to as PointCNN), and the method in document 6 (referred to as RSCNN) are analyzed as follows (as shown in Table 1):
[0149] Experiments were conducted on ShapeNetPart, a recognized 3D model point cloud component segmentation dataset. The category names of each category of the dataset are shown in the first column of Table 1, where the category names mean Airplane, Bag, Cap, Car, Chair, Earphone, Guitar, Knife, Lamp, Laptop, Motorbike, Mug, Pistol, Rocket, Skateboard, and Table. The division of the training set and the test set is shown in the second column of Table 1. The comparison of the semantic segmentation and annotation effect renderings is shown in the following figure. Figure 4 As shown; the comparison of semantic segmentation annotation accuracy is shown in Table 1 and Table 2.
[0150] As shown in the comparison of the results in Tables 1 and 2 (Table 1 shows the comparison of the average intersection-over-union ratio of semantic segmentation annotation between the method of the present invention and other methods on the ShapeNetPart dataset, and Table 2 shows the statistical comparison of the average intersection-over-union ratio of semantic segmentation annotation between the method of the present invention and other methods on the ShapeNetPart dataset), the method of the present invention is partially ahead of other methods. Among the 16 object categories, the method of the present invention is ahead of other methods in 10 object categories. The method of the present invention and PointCNN each have their own advantages and disadvantages. As shown in Tables 1 and 2, the method of the present invention exceeds PointCNN in Instance Average IoU (average of instance intersection-over-union ratio of objects) and lags slightly behind in Class Average IoU (average of class intersection-over-union ratio). Specifically for all object categories, the method of the present invention lags behind PointCNN in only 4 categories, and is ahead of PointCNN in the remaining 12 object categories.
[0151] Table 1
[0152]
[0153]
[0154] Table 2
[0155] PointNet PointNet++ PointCNN RSCNN Method of the present invention Class Average IoU 80.4 81.9 84.6 84.0 84.4 Instance Average IoU 83.7 85.1 86.1 86.2 87.1
[0156] In the self-comparison experiment, the context prior module and the self-attention module in the context-attention module are removed respectively, and the accuracy comparison with the final experimental results is shown in Table 3, which shows that the context prior module and the self-attention module can significantly improve the final semantic segmentation annotation accuracy.
[0157] Table 3
[0158]
[0159] The present invention provides a context- and attention-based 3D point cloud semantic segmentation method. There are many methods and approaches to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented using existing technologies.
Claims
1. A context- and attention-based 3D point cloud semantic segmentation method, characterized in that: The following steps are involved: Step 1, collecting data from the input 3D point cloud model data set; Step 2, extracting point-by-point features of the point cloud data through a relational shape network to obtain point cloud features containing shape information; Step 3: Use the context and attention modules to constrain and strengthen the extracted point cloud features within and between classes, and obtain point cloud features with contextual priors and global semantic associations; Step 4: Use a multi-layer perceptron classifier to classify the point cloud features and obtain the final predicted label of each point in the point cloud data; Step 1 includes the following steps: Step 1-1: Input 3D point cloud model dataset S = {S Train ,S Test } is divided into training set S Train ={s1,s2,…s i ,…,s n } and test set S Test ={s n+1 ,s n+2 ,…,s n+j ,…,s n+m }, where s i represents the i-th model in the training set, s n+j represents the jth model in the test set; Step 1-2, set the input single 3D point cloud model s i And the label set l of the components to which all points belong i , randomly sample N points from all point cloud data as the network input point set P i ={p1,p2,…p i ,…,p N }, from the label set l i Take out the point P with the i-th point i The corresponding labels form a new label set g i , i takes values of 1 to N; the data set in step 1-1 is sampled to obtain a new data set P = {P Train ,P Test };P Train represents the sampled point cloud training set, P Test Represents the sampled point cloud test set; Step 1-3, for the training set P obtained in step 1-2 Train Random scaling and translation are performed, where the scaling factor u is sampled from a uniform distribution U(0.8,1.25) and the translation amount is sampled from a uniform distribution U(-0.1,0.1).
2. The method according to claim 1, characterized in that Step 1-2 includes the following steps: Step 1-2-1, for a single 3D point cloud model i , whose point cloud set is s i ={s i1 ,s i2 ,…s ij ,..,s in }, where s ij Represents point cloud model s i The j-th point data, j ranges from 1 to n; sampling with replacement is performed from the index set Q = {1, 2, ..., n}, and repeated N times to obtain the sampled index set Q1 = {q1, q2, ...q k ,…,q N }, i k ∈I, where q k represents the index of the kth sampling from the set Q; Step 1-2-2, collect the point cloud in step 1-2-1 i The point cloud subscript in Q1 and the point cloud corresponding to the element in Q1 are added to the sampling point set P to obtain the new point cloud model data P i ={p1,p2,…p k ,…,p N }, where p k For step 1-2-1 oj j take q k ,Right now Step 1-2-3, repeat steps 1-2-1 and 1-2-2 until all 3D point cloud models in the training set have been sampled.
3. The method according to claim 2, characterized in that In steps 1-3, random scaling and translation are performed on the coordinates of each point cloud data, that is, the first three dimensions of the point cloud data.
4. The method according to claim 3, characterized in that Step 2 includes the following steps: Step 2-1, for the sampled point cloud training set P Train , collect the true label G of each point Train The point cloud data is input into the relational shape network for training, and the high-dimensional point cloud features are extracted through the encoder; Step 2-2, upsample and decode the point cloud features extracted in step 2-1 to obtain point cloud features that conform to the input shape and contain relationship information.
5. The method according to claim 4, characterized in that Step 2-1 includes the following steps: Step 2-1-1, for a single point cloud model data P i , group the point cloud data according to the farthest point sampling strategy, and iteratively select the point with the largest Euclidean distance to all point cloud data as the sphere center to obtain the point cloud grouping PG i ={pg1,pg2,…,pg i ,…,pg m }, where pg i ={p i1 ,..,p ik ,..p in }, pg i represents the i-th point cloud group, p ik Indicates pg i The kth point in Step 2-1-2, PG in step 2-1-1 i After the forward propagation convolution operation, it is extracted into an m×512-dimensional feature matrix f i ; Step 2-1-3, repeat steps 2-1-1 and 2-1-2 for a total of 3 times, in each repetition, m is 512, 128, 1, n is 32, 32, 128, respectively, to obtain the first stage of point cloud grouping PG i -1. Point cloud grouping PG in the second stage i -2. The third stage of point cloud grouping PG i -3 and the point cloud feature matrix f of the first stage i -1. The point cloud feature matrix f of the second stage i -2. The point cloud feature matrix f of the third stage i -3.
6. The method according to claim 5, characterized in that In step 2-2, a bilinear interpolation strategy is used to upsample the point cloud features to N×512 dimensions.
7. The method according to claim 6, characterized in that Step 3 includes the following steps: Step 3-1: For a single point cloud model data P i And its corresponding true label G, the feature matrix obtained in step 2 is passed through the context module to obtain the intra-class feature matrix and inter-class feature matrix that learn the context prior knowledge; In step 3-2, the intra-class feature matrix and the inter-class feature matrix obtained in step 3-1 are enhanced by the self-attention module, and the global dependency is modeled to obtain point cloud features with contextual priors and global semantic associations.
8. The method according to claim 7, characterized in that Step 3-1 includes: Step 3-1-1: For the N×512-dimensional feature matrix obtained in step 2, use 1x1 convolution operation to reduce the dimension to N×256 dimensions to obtain a new feature matrix F. By multiplying it with its transposed matrix, we can obtain an N×N-dimensional intra-class feature matrix M and an inter-class feature matrix IM, where I represents the unit matrix; aggregate the intra-class features and inter-class features to obtain the feature matrix F containing context priors. e ,Right now: F e =concat(M,(I-M)F) Among them, concat means concatenating and aggregating the features in the last dimension; Step 3-1-2: For the true label G in step 3-1, obtain the N×N dimensional covariance matrix C and calculate the difference between m and C. As part of the Loss loss, the specific calculation formula is as follows: in, They represent the accuracy within a class, the recall within a class, and the specificity between classes respectively; c ij represents the (i,j) element of matrix C, m ij represents the (i,j) element of the matrix M, μ is a non-negative minimum value; Calculate the learned context matrix, i.e., the binary cross loss between the intra-class features M and the matrix C And finally the final context loss is obtained by weighting the two losses The specific calculation formula is as follows: Among them, λ u and λ g Represents their respective weight values, and λ u and λ g Set to 1.
9. The method according to claim 8, characterized in that In step 3-2, the self-attention module uses 8-head attention and performs a self-attention on the feature matrix F containing context prior obtained in step 3-1-1. e Perform global relationship modeling and enhancement to obtain the final feature matrix; In step 4, the feature matrix obtained in step 3 is passed through a fully connected layer, and finally a Softmax multi-classifier is used to perform multi-label prediction on the input multi-dimensional feature vector to obtain a probability map of semantic segmentation of the point cloud data. The label with the highest predicted probability for each point in the point cloud data is used as the predicted label of the point, and the corresponding true label G i Compare and calculate semantic segmentation loss Same as in step 3-1-2 Add up as the total loss After back propagation, we finally get the trained point cloud segmentation network containing contextual prior knowledge. The specific calculation formula is as follows: Among them, w is the corresponding weight, c is the category, and x is the network output prediction label.
Citation Information
Patent Citations
Three-dimensional point cloud automatic classification method based on graph convolutional neural network
CN112488210A