3D Point Cloud Task Analysis Method Integrating Local Features and Global Context Information
Through the grouping self-attention mechanism and the global context feature extraction module, combined with spatial shape position coding, the problem of missing local features and global context information in the existing three-dimensional point cloud task analysis is solved, and the analysis accuracy and classification and segmentation effect of point cloud data are improved.
Patent Information
- Application Number
- CN202510038073.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-10
AI Technical Summary
When processing point cloud data, the existing three-dimensional point cloud task analysis methods are difficult to effectively retain the three-dimensional geometric information and fine-grained characteristics of point clouds, and ignore the shape relationship between point clouds, resulting in insufficient analysis accuracy.
Local Group Self-Attention (LGA) is used to capture local interaction information, and inter-region information is transmitted through the Local Group Propagation (LGP), and a global context feature extraction module (GSFM) and spatial shape position encoding (Spatial-Shape Relative Position Encoding (SS-RPE) are introduced to obtain the position relationship between points.
The analysis accuracy of three-dimensional point cloud tasks is improved, the ability to extract local features and capture global context information is enhanced, and the semantic understanding and classification and segmentation effect of point cloud data is improved.
Smart Images

Figure CN119445291B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of three-dimensional point cloud task analysis, and particularly relates to a three-dimensional point cloud task analysis method that fuses local features and global context information. Background Art
[0002] In recent years, three-dimensional point cloud technology has been widely applied in many fields such as autonomous driving and intelligent robots. Different from traditional two-dimensional image data, point cloud data consists of a set of unordered and irregular points. Three-dimensional point cloud task analysis is a key component in the practical application of this technology and is also a research hotspot in environmental intelligent perception. In realizing the understanding and perception of the three-dimensional space, three-dimensional point cloud task analysis plays a crucial role.
[0003] As a perception technology for fine-grained processing of the three-dimensional space, the effect of three-dimensional point cloud task analysis requires high efficiency, stability, and accuracy. With the remarkable stability, reliability, and automatic feature extraction ability shown by deep learning technology in tasks such as image recognition, generation, and segmentation, many studies have gradually applied deep learning to three-dimensional point cloud tasks. However, due to the disorder and sparsity of point cloud data itself, deep learning faces many challenges in processing three-dimensional point clouds.
[0004] Inspired by the successful application of convolutional neural networks in image tasks, researchers have tried to convert irregular point cloud data into regular forms such as images or voxels for applying convolutional neural networks for feature extraction. The main research directions include multi-view-based and voxel-based methods. The multi-view-based method projects the point cloud onto a two-dimensional plane and uses mature two-dimensional convolution for processing; while the voxel-based method first divides the three-dimensional space into regular voxels and then applies three-dimensional convolution to the voxel data. Although these methods perform well, due to the possible loss of position information and the lack of fine-grained features during the voxelization process, they cannot well retain the three-dimensional geometric information of the point cloud.
[0005] To solve these problems, researchers have developed point-based methods that directly use point features and position information to maintain the integrity of position information. This idea has given rise to different feature aggregation methods for learning high-level semantic features. For example, PointNet introduced a method based on a multi-layer perceptron (MLP) to process three-dimensional tasks, but its feature extraction is still limited to local information. For this reason, PointNet++ improved the ability of PointNet by introducing a hierarchical sampling strategy. And DenseKPNet established a complementary relationship between local features and high-level geometric information through multi-scale convolution and density connection to improve the accuracy of point cloud tasks and the ability to extract fine-grained features.
[0006] Recently, Transformer has gradually become the mainstream deep learning model, which solves the problem of long-term dependence and achieves excellent performance in most computer vision tasks. Inspired by the significant achievements of Transformer in two-dimensional vision tasks, researchers have developed a large number of methods for three-dimensional point cloud task analysis. Transformer has a natural advantage in learning the semantic information of geometric features. As a pure Transformer network, PCT first proposed neighborhood embedding to map point clouds to a high-dimensional architecture feature space. PU-Transformer is the first model to introduce Transformer for point cloud upsampling, and a new variant of the multi-head self-attention structure is developed to enhance the point-wise and channel relationships of the feature map. OctForme inherits the order during the octalization process, similar to the z-order, providing scalability and good performance while reducing the computational amount, but it is still restricted by the octree structure itself. On the other hand, FlatFormer uses a window-based sorting strategy to group points into pillars, similar to window partitioning. However, this design lacks scalability in the receptive field. However, these methods only regard the relative position relationship between point clouds as local geometric features, ignoring the shape relationship between point clouds, and the refinement of the rich shape semantic position information of point clouds is still limited. Summary of the Invention
[0007] To overcome the deficiencies of the prior art, the present invention proposes a method for three-dimensional point cloud task analysis that fuses local features and global context information. First, a grouped self-attention mechanism (Local Group Self-Attention, LGA) is designed to capture local interaction information in each region. To capture the separated local region feature relationships, a grouped propagation module (Local Group Propagation, LGP) is proposed to transmit information between different regions through query points, allowing features to spread between neighbors and obtaining more fine-grained feature information. To further expand the effective receptive field, a global context feature extraction module (Global Shape Feature Module, GSFM) is proposed to learn global context information through key shape points. Finally, to solve the position information clues between global contexts, spatial-shape relative position encoding (Spatial-Shape Relative Position Encoding, SS-RPE) is introduced to obtain the position relationship between points.
[0008] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0009] A method for three-dimensional point cloud task analysis that fuses local features and global context information, the method comprising the following steps:
[0010] Step 1. Construct a point cloud data set, and the process is as follows:
[0011] Step 1.1) Define the data set;
[0012] Step 1.2) Preprocess the point cloud data set: For the point cloud data set P, perform preprocessing operations of random rotation, scaling, and adding noise to generate point cloud data sets under various conditions, so as to improve the anti-interference ability of the method to data changes;
[0013] Step 2. Sample the point cloud data set. In order to alleviate the computational burden brought by the huge number of point clouds in the point cloud data set, the number of point clouds in the original point cloud data set is downsampled to a fixed value by the farthest point sampling (FPS) method;
[0014] Step 3. Group the point cloud data set: In three-dimensional solid geometry, a shape can be represented by a set of closely arranged three-dimensional coordinates. For the geometric bodies of the same target, their point cloud data will be similar in spatial features and semantic expressions. Therefore, when performing point cloud tasks, extracting the semantic information of the target geometric body depends on the correlation between each point and its surrounding neighborhood points. KNN (K-Nearest Neighbors) is a commonly used algorithm for determining the neighborhood points of each point; calculating the distance between each point and its K nearest neighbors can better understand the correlation between point cloud data. Grouping the point cloud data through the KNN algorithm can effectively capture this local spatial relationship, which helps to extract the semantic information of the point cloud data;
[0015] Step 4. Extract local features: First construct a neighborhood set, and then extract the features of the neighborhood set;
[0016] Step 5. Neighborhood feature propagation: In order to capture the feature relationship of separated local regions, establish the relationship between neighborhoods through the query points in different regions to allow the point cloud features to spread between neighbors and obtain more fine-grained feature information;
[0017] Step 6. Construct spatial shape position encoding. First construct attention position encoding and azimuth position encoding, and then perform feature mixing;
[0018] Step 7. Construct channel attention;
[0019] Step 8. Enhance local features,
[0020] Step 9. Extract global context features. The global shape feature extraction module (GSFM) is used for global context feature extraction;
[0021] Step 10: 3D point cloud classification and segmentation tasks, the process is as follows:
[0022] Step 10.1): Construct the classification task: The shape feature propagation module (SAPFormer Block) proposed in the present invention consists of steps 1 to 8. For the 3D point cloud classification task, the primary point cloud is obtained through point cloud embedding. The feature encoder consists of a downsampling layer and a SAPFormer module. The classification result is obtained through global average pooling (GAP) and MLP; the downsampling layer (Downsample) is a layer in a deep learning model used to reduce the spatial dimension of the input data; global average pooling is a method used in deep learning models. Different from traditional pooling layers, the purpose of the global average pooling layer is to convert the entire feature map into a single average value, which helps to simplify the model and directly generate class labels after feature extraction; point cloud embedding (Point Embeding) is used to extract the primary features of the point cloud;
[0023] Step 10.2): Construct the segmentation task: For the 3D point cloud segmentation task, the U-Net architecture is adopted. The encoder is the same as the classification task in step 10.1). After each stage is fed to the corresponding upsampling (Upsample), the segmentation result is obtained by accumulation; where U-Net is a deep learning architecture that can effectively process a small amount of training data and at the same time can generate high-resolution segmentation images. It is mainly used to restore the spatial dimension of the feature map, increase its height and width, and upsampling can convert a low-resolution feature map into a high-resolution output.
[0024] Furthermore, in step 1.1), the point cloud data consists of a series of three-dimensional coordinates, and each coordinate point contains the coordinate information of the point cloud in three-dimensional space and the information f reflecting the texture characteristics of the point cloud data. Denote the point cloud data set containing N points as , where is the coordinate information of the i-th point. The texture feature set corresponding to P is denoted as , is the color information of the i-th point.
[0025] The process of step 1.2) is as follows:
[0026] Step 1.2.1): Random rotation, randomly generate a rotation angle , and construct a three-dimensional coordinate rotation matrix according to the angle , and its expression is
[0027] ;
[0028] Multiply the point cloud data set P by the rotation matrix to obtain the point cloud data set after the original point cloud rotates around the z-axis by the angle, and the expression is ,
[0029] ;
[0030] Step 1.2.2), Random scaling, randomly generate a scaling ratio , and multiply the scaling ratio by the rotated point cloud data set to generate the scaled data , and the expression is
[0031] ;
[0032] Step 1.2.3), Add noise, randomly generate a set of three-dimensional coordinates with the same dimension as the point cloud data set , , the value of which is relatively small, add the scaled data set and the noise to add a small amount of noise interference to the original point cloud,
[0033] .
[0034] The process of step 2 is as follows:
[0035] Step 2.1), Determine the sampling initial point, traverse the point cloud data set , and calculate the average value of all points in the three coordinate dimensions, and use the calculated average value as the initial point for farthest point sampling . At the same time, add it to the sampling point set ;
[0036] Step 2.2), Calculate the distance matrix. Farthest point sampling is to select the point with the farthest distance from the point cloud data set . Therefore, in order to obtain the distance between the point cloud data and the sampling point set , record the Euclidean distance between each point and the newly added point in through the distance matrix ;
[0037] Step 2.3), Update the sampling point set, denote the point corresponding to the maximum value in as , and add it to . As updates the sampling points, The Euclidean distance between the newly added sampling points and the point cloud dataset will be recalculated, and at the same time, the point cloud with Euclidean distance less than the original value will be updated; finally, the above operations will be looped to increase the number of point clouds in the sampling point set to a fixed value.
[0038] The process of step 4 is as follows:
[0039] Step 4.1) Construct the neighborhood set:
[0040] According to the point cloud set P constructed in step 1, use KPconv to extract local features. KPconv is a mature point cloud network architecture that proposes kernel point convolution to simulate deformable convolution to obtain local information of the point cloud; then, obtain center points through the FPS algorithm in step 2, that is, obtain local regions, and then divide the point cloud sampled by the KNN algorithm in step 3 into N / K local region blocks, each region contains K points, and group the K nearest neighbor points around the center point of each region, that is where C is the dimension of the point cloud feature, is the number of sampled points;
[0041] Step 4.2) Extract neighborhood set features:
[0042] After obtaining the point cloud after sampling and grouping , use LGA to extract the local neighborhood features after grouping, input the points of each local neighborhood into the MLP (MultiLayer Perceptron) and then generate in sequence, and the feature of each local region is , denoted as is , and the calculation formula is as follows,
[0043] ;
[0044] ;
[0045] ;
[0046] where is the MLP. MLP is a forward propagation neural network with an input layer, hidden layers, and an output layer structure, and generates a description of the input data in the global expression space through multiple calculations of the hidden layers; by calculating the similarity between the query vector and all key vectors , use the Softmax function to obtain a weight distribution for weighted summation of the associated numerical vectors ; The Softmax function is a function that converts data to a value between 0 and 1. C is the dimension of the point cloud features, which prevents gradient explosion. Gradient explosion refers to the situation where, during backpropagation, the gradient value grows exponentially as the number of network layers increases, resulting in a large update of the network weights and making the network unstable; is the attention weight vector, which is the fusion weight ordinal adaptively adjusted according to the saliency of the global and local features. Finally, the numerical vector is multiplied by the weight vector and then passed through the MLP to obtain the grouped point cloud features .
[0047] The process of step 5 is as follows:
[0048] Step 5.1): Construct a neighborhood index. Use the KNN algorithm in step 3 to obtain the index map of the neighbor points for the grouped point cloud features obtained in step 4 and the grouped point cloud set as follows,
[0049] ;
[0050] ;
[0051] Among them, is 's index mapping of the k nearest points, obtained by querying different indices o for the region of the k nearest points ;
[0052] Step 5.2): Neighborhood propagation. Calculate the inverse index mapping through the index mapping , backtrack to another region , query the point through and to obtain all elements in the two sets, and re - establish a set , , where is the coordinate information of the l - th point, is the feature of the l - th point, k_n is the number of all points (including itself) that can be connected to. Given a point , it means that its local region is , and g represents the number of associated regions.
[0053] The process of step 6 is as follows:
[0054] Step 6.1), Attention Position Encoding: Input the query points obtained in Step 5.2), and their associated points into the MLP to increase the feature dimension to . Then, calculate the geometric relative position information and the associated points . , where is the number of distances between points; Finally, pass the point features in Step 5.2), through the MLP in sequence to obtain Q, K, V. The position relationship Pos and the geometric representation G are obtained by the following formulas:
[0055] ;
[0056] ;
[0057] ;
[0058] where is the learnable matrix of geometric representation, is the learnable matrix of position relationship;
[0059] Step 6.2), Azimuth Position Encoding: Input the query points obtained in Step 5.2), and their associated points and the feature into the Azimuth Position Encoding APE for encoding. APE represents the position relationship between point clouds in the polar coordinate system,
[0060] ;
[0061] ;
[0062] ;
[0063] ;
[0064] where is the linear transformation and normalization operation. In the local area constructed in Step 5.2), is the radius between the point and the center of the circle, and represent the angular relationship with the center of the circle to enhance the position information;
[0065] Step 6.3), Feature Mixing. The spatial shape position encoding ss-RPE consists of the attention position encoding in Step 6.1) and the azimuth position encoding in Step 6.2). ss-RPE accumulates the position relationship Pos obtained in Step 6.1) and the features obtained from APE in Step 6.2), as well as the point cloud features to obtain the final feature, which is expressed as follows,
[0066] ;
[0067] ;
[0068] ;
[0069] where, is the feature obtained after encoding, is the position query relationship between points, is the weight of the attention position encoding.
[0070] In Step 7, different channels usually exhibit strong geometric relationships in the point cloud. Therefore, reweighting is performed according to the different semantic information represented between point channels. Specifically, the feature obtained in Step 6.3 is projected to generate and as well as . The weight matrix query embedding is obtained by the matrix product of the key embedding . Then, this matrix is disassembled and repeated operations are performed to obtain channel matrices . The spatial information of the points is propagated into each channel to obtain the information relationship between points, After passing through the softmax normalization function, the weight of a single channel is then multiplied element-wise with the key embedding to maintain channel differences, and then the feature of is updated. This process is expressed as,
[0071] ;
[0072] ;
[0073] ;
[0074] ;
[0075] where, is a linear projection, which is an important concept in linear algebra and mathematical analysis. It maps a vector space to its subspace. is a repeated operation function, which copies D copies and splices them into .
[0076] In step 8, after the point cloud set P passes through steps 4, 5, 6, and 7, it performs local feature extraction again through step 4 to obtain the local final point cloud fine-grained feature ,
[0077] ;
[0078] Among them, LGA(*) represents performing step 4 operation again, is the encoded feature obtained after step 6, is the point after being encoded through step 6.
[0079] In step 9, according to the point cloud set P constructed in step 1, the key shape regions of different objects are captured using a few key shape points (KSP). The positions of KSP points change dynamically according to different object shapes. Therefore, several KSP points are used to represent the specific deformations of such objects. The formula is as follows,
[0080] ;
[0081] ;
[0082] ;
[0083] Among them, is the offset of the point, is the calculated KSP point, is the feature of the input point. The linear layer is a basic layer in a neural network and is also known as the fully connected layer or dense layer. The role of the linear layer in a neural network is to perform a linear transformation on the input data;
[0084] GSFM accepts the coordinate information of the grouped points after step 4.1), as well as the coordinate information P and features of all points. For the convenience of representation, the grouped points are used as the neighborhood boundary sphere, and all points are used as the entire point cloud boundary sphere. The volume ratio of the neighborhood boundary sphere and the point cloud boundary sphere is calculated respectively as the slight deformation of the entire object, and then combined with the specific deformation of the object, more effective global context features are obtained from the shape information in the three-dimensional points.
[0085] ;
[0086] ;
[0087] ;
[0088] Among them, is the volume of the corresponding neighborhood boundary sphere, is the volume of the point cloud boundary sphere, is splicing, is the coordinate of the KSP point, used to represent the position of the local neighborhood, is the original point cloud feature, is the feature of the KSP point, is the final point cloud feature.
[0089] The beneficial effects of the present invention are mainly manifested in:
[0090] 1. The present invention proposes a method for extracting local features of point clouds. This method first extracts features from the local area of the point cloud and reduces the computational amount by grouping the point cloud; then it uses query points to transfer information between different regions to allow features to spread between neighbors, thereby obtaining the context information of the point cloud.
[0091] 2. The present invention designs a new global shape feature extraction module. In order to further expand the effective receptive field, this module uses a few Key Shape point (KSP) points to capture the key shape regions of different objects to expand the receptive field and obtain global context features.
[0092] 3. The present invention designs a brand-new position encoding: Spatial Shape Relative Position Encoding (SS-RPE) captures the position information between points, promotes communication between points, and interacts with the current feature by calculating the difference in spatial shape relative position information. Therefore, it can simultaneously process the unordered features of the point cloud and adaptively enhance the position information.
[0093] 4. The present invention can effectively improve the accuracy of 3D point cloud tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 is the network structure diagram.
[0095] Figure 2 is the SS-RPE structure diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0096] The present invention will be further described below with reference to the accompanying drawings.
[0097] Reference Figure 1 and Figure 2 , a three-dimensional point cloud task analysis method that fuses local features and global context information, the method comprising the following steps:
[0098] Step 1. Construct a point cloud data set, the process is as follows:
[0099] Step 1.1). Define the data set;
[0100] In the said step 1.1), the point cloud data is composed of a series of three-dimensional coordinates, and each coordinate point contains the coordinate information of the point cloud in three-dimensional space and the information f reflecting the texture characteristics of the point cloud data. Denote the point cloud data set containing N points as , where is the coordinate information of the i-th point. The texture feature set corresponding to P is denoted as , is the color information of the i-th point.
[0101] Step 1.2). Preprocess the point cloud data set: For the point cloud data set P, perform preprocessing operations of random rotation, scaling, and adding noise to generate point cloud data sets in various situations, so as to improve the anti-interference ability of the method to data changes. The process is as follows:
[0102] Step 1.2.1). Random rotation, randomly generate a rotation angle , and construct a three-dimensional coordinate rotation matrix according to the angle , and its expression is
[0103] ;
[0104] Multiply the point cloud data set P by the rotation matrix to obtain the point cloud data set after the original point cloud rotates around the z-axis by angle, and the expression is
[0105] ;
[0106] Step 1.2.2). Random scaling, randomly generate a scaling ratio , and multiply the scaling ratio by the rotated point cloud data set to generate the scaled data , and the expression is
[0107] ;
[0108] Step 1.2.3). Add noise, randomly generate a group of data sets that are the same as the point cloud data set Three-dimensional coordinates with the same dimensions , has a smaller value, and the scaled dataset and noise are added together to add slight noise interference to the original point cloud,
[0109] .
[0110] Step 2: Sampling the point cloud dataset. To alleviate the computational burden brought by the huge number of point clouds in the point cloud dataset, the number of point clouds in the original point cloud dataset is downsampled to a fixed value by the farthest point sampling (FPS) method;
[0111] The process of the said Step 2 is as follows:
[0112] Step 2.1): Determine the initial sampling point. Traverse the point cloud dataset and calculate the average value of all points in the three coordinate dimensions. The calculated average value is used as the initial point for the farthest point sampling , and at the same time, add it to the sampling point set ;
[0113] Step 2.2): Calculate the distance matrix. The farthest point sampling is to select the point with the farthest distance from the point cloud dataset. Therefore, in order to obtain the distance between the point cloud data and the sampling point set , record the Euclidean distance between each point and the newly added point in through the distance matrix ; the Euclidean distance of the newly added point;
[0114] Step 2.3): Update the sampling point set. Denote the point corresponding to the maximum value in as , and add it to . As updates the sampling point, will recalculate the Euclidean distance between the newly added sampling point and the point cloud dataset, and at the same time update the point cloud whose Euclidean distance is less than the original value; finally, loop the above operations to increase the number of point clouds in the sampling point set to a fixed value.
[0115] Step 3: Grouping the point cloud dataset: In three-dimensional solid geometry, a shape can be represented by a set of closely arranged three-dimensional coordinates in sequence. For the geometric bodies of the same object, their point cloud data will be similar in spatial characteristics and semantic expressions. Therefore, when performing point cloud tasks, extracting the semantic information of the target geometric body depends on the correlation between each point and its surrounding neighborhood points. KNN (K-Nearest Neighbors) is a commonly used algorithm for determining the neighborhood points of each point. Calculating the distance between each point and its K nearest neighbors can better understand the correlation between point cloud data. By grouping the point cloud data through the KNN algorithm, this local spatial relationship can be effectively captured, which helps to extract the semantic information of the point cloud data;
[0116] Step 4: Extracting local features: First, construct a neighborhood set, and then extract the features of the neighborhood set;
[0117] The process of Step 4 is as follows:
[0118] Step 4.1) Constructing the neighborhood set:
[0119] According to the point cloud set P constructed in Step 1, use KPconv to extract local features. KPconv is a mature point cloud network architecture that proposes kernel point convolution and simulates deformable convolution to obtain the local information of the point cloud; then, through the FPS algorithm in Step 2, center points are obtained, that is, local regions are obtained. Then, use the point cloud sampled by the KNN algorithm in Step 3 to divide it into N / K local region blocks. Each region contains K points. Group the K nearest neighbor points around the center point of each region, that is, , where C is the dimension of the point cloud feature, is the number of points after sampling;
[0120] Step 4.2) Extracting the features of the neighborhood set:
[0121] After obtaining the point cloud after sampling and grouping , use LGA to extract the local neighborhood features after grouping. Input the points of each local neighborhood into the MLP (MultiLayer Perceptron) and then generate in sequence. The feature of each local region is . Denote as . The calculation formula is as follows,
[0122] ;
[0123] ;
[0124] ;
[0125] Among them, is an MLP. The MLP is a forward propagation neural network with an input layer, a hidden layer, and an output layer structure, which generates a description of the input data in the global feature space through multiple calculations in the hidden layer; by calculating the query vector and all key vectors the similarity between them, a weight distribution is obtained using the Softmax function, which is used to weighted sum the associated numerical vectors ; The Softmax function is a function that converts data between 0 and 1. C is the dimension of the point cloud feature to prevent gradient explosion; Gradient explosion refers to the phenomenon that during backpropagation, the gradient value increases exponentially as the number of network layers increases, resulting in a large update of the network weights and making the network unstable; is the attention weight vector, which is an adaptive adjustment fusion weight ordinal according to the saliency of the global feature and the local feature; Finally, the numerical vector is multiplied by the weight vector and then passed through the MLP to obtain the grouped point cloud feature .
[0126] Step 5, Neighborhood Feature Propagation: In order to capture the feature relationship of separated local regions, the relationship between neighborhoods is established through query points in different regions, allowing the point cloud features to propagate between neighbors to obtain more fine-grained feature information;
[0127] The process of step 5 is as follows:
[0128] Step 5.1), Construct a neighborhood index, and use the KNN algorithm in step 3 to obtain the index map of neighbor points for the grouped point cloud feature obtained in step 4 and the grouped point cloud set as follows,
[0129] ;
[0130] ;
[0131] Among them, is of index mapping of the nearest points, and the region of the nearest points is obtained by querying different indexes o ;
[0132] Step 5.2), Neighborhood Propagation, calculate the inverse index mapping through the index mapping to backtrack to another region , query point Through And Obtain all elements in the two sets and re - establish a set , , where is the coordinate information of the l - th point, is the feature of the l - th point, and k_n is the number of all points that can be connected (including itself). Given a point , it means that the local area it belongs to is , and g represents the number of associated regions.
[0133] Step 6: Construct the spatial shape position encoding. First, construct the attention position encoding and azimuth position encoding, and then perform feature mixing;
[0134] The process of step 6 is as follows:
[0135] Step 6.1) Attention position encoding: Input the query point obtained in step 5.2) and its associated points into the MLP to increase the feature dimension to . Then, from and the associated points calculate the geometric relative position information , is the number of distances between points. Finally, the point features in step 5.2) are sequentially passed through the MLP to obtain Q, K, V. The position relationship Pos and geometric representation G are obtained by the following formulas,
[0136] ;
[0137] ;
[0138] ;
[0139] Among them, is the learnable matrix of geometric representation, is the learnable matrix of position relationship;
[0140] Step 6.2) Azimuth position encoding: Input the query point obtained in step 5.2) and its associated points and the feature Input to the azimuth position encoding APE (Azimuth Position Encoding) for encoding. APE represents the positional relationship between point clouds in the polar coordinate system.
[0141] ;
[0142] ;
[0143] ;
[0144] ;
[0145] Among them, is a linear transformation and normalization operation. In the local area constructed in step 5.2), is the radius between the point and the center of the circle, and represent the angular relationship with the center of the circle, enhancing the position information;
[0146] Step 6.3), Feature mixing. The spatial shape position encoding ss - RPE is composed of the attention position encoding in step 6.1) and the azimuth position encoding in step 6.2). ss - RPE accumulates the position relationship Pos obtained in step 6.1), the features obtained from APE in step 6.2) and the point cloud features to obtain the final feature, expressed as follows,
[0147] ;
[0148] ;
[0149] ;
[0150] Among them, is the feature obtained after encoding, is the position query relationship between points, is the weight of the attention position encoding.
[0151] Step 7, Construct channel attention;
[0152] In step 7, different channels usually exhibit strong geometric relationships in the point cloud. Therefore, re - weighting is performed according to the different semantic information represented between point channels. Specifically, the feature obtained in step 6.3) is projected to generate and and , and the query embedding and the key embedding The weight matrix is obtained by the matrix product , and then the matrix is disassembled and repeated operations are performed to obtain channel matrices . The spatial information of the points is propagated to each channel to obtain the information relationship between the points. After passing through the softmax normalization function, the weights of a single channel are multiplied element-wise with the key embedding to maintain channel differences, and then the features of are updated. This process is expressed as
[0153] ;
[0154] ;
[0155] ;
[0156] ;
[0157] Among them, is a linear projection. Linear projection is an important concept in linear algebra and mathematical analysis. It maps a vector space to its subspace. is the repeated operation function. According to , D copies are replicated and spliced into .
[0158] Step 8, Local Feature Enhancement
[0159] In the said Step 8, after the point cloud set P passes through Step 4, Step 5, Step 6, and Step 7, it passes through Step 4 again for a local feature extraction to obtain the local final point cloud fine-grained feature ,
[0160] ;
[0161] Among them, LGA(*) represents performing Step 4 operation again. is the encoded feature obtained after Step 6. is the point after being encoded through Step 6.
[0162] Step 9, Global Context Feature Extraction. The Global Shape Feature Module (GSFM) is used for global context feature extraction.
[0163] In step 9, according to the point cloud set P constructed in step 1, the key shape regions of different objects are captured using a few key shape points (KSP). The positions of the KSP points change dynamically according to the different object shapes. Therefore, several KSP points are used to represent the specific deformations of such objects. The formula is as follows:
[0164] ;
[0165] ;
[0166] ;
[0167] where, is the offset of the point, is the calculated KSP point, is the feature of the input point. The linear layer is a basic layer in a neural network, also known as the fully connected layer or the dense layer. The role of the linear layer in a neural network is to perform a linear transformation on the input data;
[0168] GSFM receives the coordinate information of the grouped points after step 4.1), the coordinate information P of all points, and the features. For the sake of easy representation, the grouped points are used as the neighborhood boundary sphere, and all points are used as the entire point cloud boundary sphere. The volume ratios of the neighborhood boundary sphere and the point cloud boundary sphere are calculated respectively as the slight deformation of the entire object, and then combined with the specific deformation of the object, more effective global context features are obtained from the shape information in the three-dimensional points.
[0169] ;
[0170] ;
[0171] ;
[0172] where, is the volume of the corresponding neighborhood boundary sphere, is the volume of the point cloud boundary sphere, is the concatenation, is the coordinate of the KSP point, is used to represent the position of the local neighborhood, is the original point cloud feature, is the feature of the KSP point, is the final point cloud feature.
[0173] In the Transformer model, the feed-forward neural network (FFN) is one of the core components of the Transformer. It is located after each encoder and decoder layer of the Transformer. The feed-forward neural network is a fully connected feed-forward neural network, which consists of two linear transformations (fully connected layers) and a non-linear activation function.
[0174] Step 10: 3D point cloud classification and segmentation tasks, the process is as follows:
[0175] Step 10.1): Construct the classification task: As Figure 1 Shown on the right, the shape feature propagation module (SAPFormer Block) proposed by the present invention consists of steps 1 to 8. As Figure 1 Shown in the upper half, for the 3D point cloud classification task, the primary point cloud is obtained through point cloud embedding. The feature encoder consists of a downsampling layer and a SAPFormer module. The classification result is obtained through global average pooling (GAP) and MLP. The downsampling layer (Downsample) is a layer in a deep learning model used to reduce the spatial dimensions (such as height and width) of the input data. Global average pooling is a method used in deep learning models. Different from traditional pooling layers (such as max pooling or average pooling), the purpose of the global average pooling layer is to convert the entire feature map into a single average value, which helps to simplify the model and directly generate class labels after feature extraction. Point embedding (Point Embeding) is used to extract the primary features of the point cloud, such as MLP and the Kpconv network architecture;
[0176] Step 10.2): Construct the segmentation task: As Figure 1 Shown in the lower half, for the 3D point cloud segmentation task, the U-Net architecture is adopted. The encoder is the same as the classification task in step 10.1). After each stage is fed to the corresponding upsampling (Upsample), the segmentation result is obtained by accumulation. Among them, U-Net is a deep learning architecture that can effectively process a small amount of training data and can generate high-resolution segmentation images. It is mainly used to restore the spatial dimensions of the feature map and increase its height and width. Upsampling can convert a low-resolution feature map into a high-resolution output.
[0177] In this embodiment, the ShapeNetPart dataset and the ModelNet40 dataset are used to perform performance tests on the present invention. The test process includes the following steps:
[0178] Step 1: Define the comparison method, as follows:
[0179] PointNet: Introduced a method based on MLP to solve 3D point cloud tasks.
[0180] Kpconv: Proposed kernel point convolution, simulating deformable convolution to obtain local information of point clouds.
[0181] PAconv: Proposed position-adaptive convolution to obtain information of point clouds, simulating complex 3D position changes.
[0182] Step 2, Define evaluation metrics: The present invention uses Instance IoU (ins. mIoU) and class IoU (cls.mIoU) as evaluation metrics for the part segmentation task, and uses the mean accuracy (mAcc) of each class and the overall accuracy (OA) of all classes as evaluation metrics for the shape classification.
[0183] ;
[0184] ;
[0185] ;
[0186] ;
[0187] Among them, is the total number of shape categories in the dataset (in the ShapeNet dataset = 16), represents the number of instances of the i-th class, N represents the number of all shape classes, is the number of correct predictions of the i-th class, is the number of samples of the i-th class, is the number of classes in the dataset, is the total number of samples in the dataset;
[0188] Step 3, Dataset introduction, as follows:
[0189] 3.1) Introduction to the part segmentation dataset;
[0190] ShapeNetPart is a dataset for point cloud part segmentation, which consists of 16,880 3D models from 16 different shape categories (such as airplanes and chairs), of which 14,006 models are used for training and 2,874 for testing. The number of parts per class ranges from 2 to 6, and there are a total of 50 different parts.
[0191] 3.2) Introduction to the shape classification dataset;
[0192] ModelNet40 is a shape classification dataset, which consists of 9843 training models and 2468 testing CAD models belonging to 40 categories, and it contains 15,000 objects from 15 classes.
[0193] 3.3) Evaluation of part segmentation experimental results;
[0194] Tables 1 and 2 list the part segmentation results of the present invention on ShapeNetPart. It can be seen from the table that the method is about 0.7% higher than the method based on point convolution paconv and other methods. Moreover, the present invention achieves the best and suboptimal performance in Instance IoU and classIoU respectively. In addition, the part segmentation effect on objects such as car, lamp, laptop, mug, and table is optimal, which verifies the effectiveness of the present invention in point cloud part segmentation tasks.
[0195] Table 1 shows the part segmentation results of the present invention on ShapeNetPart;
[0196] Method Cls.mIoU Ins.mIoU Airplane Bag Cap Car Chair EarPhone Guitar PointNet 80.4 83.7 83.4 78.7 82.5 74.9 89.6 73.0 91.5 paconv 84.6 86.1 84.3 85.0 90.4 79.7 90.6 80.8 92.0 Kpconv 85.1 86.4 84.6 86.3 87.2 81.1 91.1 77.8 92.6 SAPFormer 85.0 86.8 84.9 86.0 89.9 81.3 91.8 75.4 92.1
[0197] Table 2 shows the part segmentation results of the present invention on ShapeNetPart (Table 2 is a continuation of Table 1);
[0198] Method Knife Lamp Laptop MotorBike Mug Pistol Rocket SkateBoard table PointNet 85.9 80.8 95.3 65.2 93.0 81.2 57.9 72.8 80.6 paconv 88.7 82.2 95.9 73.9 94.7 84.7 65.9 81.4 84.0 Kpconv 88.4 82.7 96.2 78.1 95.8 85.4 69.0 82.0 83.6 SAPFormer 87.9 86.4 96.3 76.7 95.8 84.5 68.0 78.7 84.4
[0199] 3.4) Evaluation of shape classification experimental results;
[0200] Table 3 lists the shape classification results of the present invention on ModelNet40. It can be seen from the table that compared with the advanced methods in recent years, the present invention has achieved suboptimal performance in both OA and mAcc, which verifies the effectiveness of the present invention in point cloud classification tasks.
[0201] Table 3 shows the shape classification results of the present invention on ModelNet40;
[0202] Method mAcc OA PointNet 86.2 89.2 paconv - 93.6 kpconv - 92.9 SAPFormer <![CDATA 90.9 > <![CDATA 93.9 >
[0203] The contents described in the embodiments of this specification are merely enumerations of implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.
Claims
1. A three-dimensional point cloud task analysis method that fuses local features and global context information, characterized in that, The method includes the following steps: Step 1: Construct a point cloud data set, the process is as follows: Step 1.1): Define the data set; Step 1.2): Preprocess the point cloud data set: For the point cloud data set P, perform operations of random rotation, scaling, and adding noise for preprocessing to generate point cloud data sets under various conditions; Step 2: Sample the point cloud data set, and downsample the number of points in the original point cloud data set to a fixed value through the farthest point sampling (FPS) method; Step 3: Group the point cloud data set: Group the point cloud data through the KNN algorithm; Step 4: Extract local features: First construct a neighborhood set, and then extract the features of the neighborhood set; the process is as follows: Step 4.1) Construct the neighborhood set: According to the point cloud set P constructed in Step 1, use KPconv to extract local features. KPconv is a mature point cloud network architecture that proposes kernel point convolution and simulates deformable convolution to obtain the local information of the point cloud. Then, obtain M' center points through the FPS algorithm in Step 2, that is, obtain M' local regions, and then use the point cloud sampled by the KNN algorithm in Step 3 Divide it into N / K local region blocks, with K points in each region, and group the K nearest neighbor points around the center point of each region, that is where C is the dimension of the point cloud feature and N' is the number of points after sampling; Step 4.2) Extract the features of the neighborhood set: After obtaining the sampled and grouped point cloud the grouped self-attention mechanism LGA is used to extract the local neighborhood features after grouping, and the points in each local neighborhood are input into the MLP to sequentially generate Q LGA , K LGA , V LGA , and each local region feature is Let M' be N' / K, and the calculation formula is as follows Z local = α(W LGA V LGA )); Among them, φ, δ, α are MLP. MLP is a forward propagation neural network with an input layer, a hidden layer, and an output layer structure. It generates a description of the input data in the global expression space through multiple calculations in the hidden layer; by calculating the query vector Q LGA and all key vectors K LGA to obtain a similarity, and using the Softmax function to get a weight distribution for weighted summation of the associated numerical vectors V LGA ; The Softmax function is a function that converts data between 0 and 1. C is the dimension of the point cloud feature to prevent gradient explosion; Gradient explosion refers to the situation where during backpropagation, the gradient value grows exponentially as the number of network layers increases, resulting in a large update of the network weights and making the network unstable; W LGA is the attention weight vector, which is a fusion weight ordinal adaptively adjusted according to the salience of the global feature and the local feature; Finally, the numerical vector V LGA is multiplied by the weight vector and then passed through the MLP to obtain the grouped point cloud feature Z local ; Step 5: Neighborhood feature propagation: In order to capture the feature relationships of separated local regions, establish the relationships between neighborhoods through query points in M' different regions, so as to allow the point cloud features to propagate between neighbors and obtain more fine-grained feature information; Step 6: Construct a spatial shape position encoding. First construct an attention position encoding and an azimuth position encoding, and then perform feature mixing; the process is as follows: Step 6.1), Attention Position Encoding, input the query point p obtained in Step 5 query and its associated point p l (x i ,y i ) into the MLP, and increase the feature dimension to C”. Then, from p query and the associated point (x i ,y i ) calculate the geometric relative position information N”×N” is the number of distances between points and points; finally, the point features in Step 5 are sequentially passed through the MLP to obtain Q, K, V. The position relationship Pos and the geometric representation G are obtained by the following formulas G = R·W G ; Pos = R·W P ; Among them, W G is a learnable matrix for geometric representation, and W P is a learnable matrix for positional relationship; Step 6.2), azimuth position encoding, input the query point p obtained in step five query and its associated point p l (x i , y i ) and features into the azimuth position encoding APE for encoding. APE represents the positional relationship between point clouds in the polar coordinate system F APE = f p ([ω i , θ i , σ i ) Among them, f p is a linear transformation and a normalization operation. In the local region constructed in Step Five, ω i is the radius between a point and the center of the circle, θ i and σ i represent the angular relationship with the center of the circle, enhancing the position information; Step 6.3), Feature Mixing. The spatial shape position encoding ss-RPE is composed of the attention position encoding in Step 6.1) and the azimuth position encoding in Step 6.2). The ss-RPE accumulates the position relationship Pos obtained in Step 6.1), the feature F obtained from the APE in Step 6.2), APE and the point cloud feature X l to obtain the final feature, which is expressed as follows: P qk = Pos·Q + Pos·K; F = W a ·G + W a ·V + F APE ; Among them, F is the feature obtained after encoding, and P qk is the position query relationship between points, and W a is the weight of the attention position encoding; Step 7: Construct channel attention; Step 8: Enhance local features, Step 9: Extract global context features, and use the global shape feature extraction module (GSFM) for global context feature extraction; Step 10: Three-dimensional point cloud classification and segmentation tasks, the process is as follows: Step 10.1): Construct a classification task: For the three-dimensional point cloud classification task, obtain the primary point cloud through point cloud embedding. The feature encoder consists of a downsampling layer and a SAPFormer module. The SAPFormer module is a shape feature propagation module, and the processing process of this shape feature propagation module consists of Step 1 to Step 8. Obtain the classification result through global average pooling and MLP; Step 10.2): Construct a segmentation task: For the three-dimensional point cloud segmentation task, adopt a U-Net architecture. The encoder is the same as the classification task in Step 10.1). After each stage is fed to the corresponding upsampling, the segmentation result is obtained by accumulation.
2. The three-dimensional point cloud task analysis method for fusing local features and global context information according to claim 1, wherein In the step 1.1), the point cloud data consists of a series of three-dimensional coordinates. Each coordinate point contains the coordinate information (x, y, z) of the point cloud in the three-dimensional space and the information f reflecting the texture characteristics of the point cloud data. The point cloud data set containing N points is denoted as P = {p i | i = 1, …, N}, where p i is the coordinate information of the i-th point, and the texture feature set corresponding to P is denoted as F o = {f i | i = 1, …, N}, and f i is the color information of the i-th point.
3. The three-dimensional point cloud task analysis method for fusing local features and global context information as described in claim 1 or 2, characterized in that The process of Step 1.2) is: Step 1.2.1): Random rotation, randomly generate a rotation angle θ, and construct a three-dimensional coordinate rotation matrix R(θ) according to the angle θ. Its expression is, Multiply the point cloud data set P by the rotation matrix R(θ) to obtain the point cloud data set P′ after the original point cloud rotates by θ degrees around the z-axis. The expression is, Step 1.2.2), Random scaling, randomly generate a scaling ratio λ D , multiply the scaling ratio λ D by the rotated point cloud dataset P′ to generate the scaled data P″, and the expression is P″ = λ D P′; Step 1.2.3): Add noise, randomly generate a set of three-dimensional coordinates J with the same dimension as the point cloud data set P″, and the values of J are relatively small. Add the scaled data set P″ and the noise J to add slight noise interference to the original point cloud, P″ = P″ + J.
4. The three-dimensional point cloud task analysis method for fusing local features and global context information according to claim 3, characterized in that The process of Step 2 is as follows: Step 2.1): Determine the sampling initial point, traverse the point cloud data set P″, and calculate the average value of all points in the three coordinate dimensions. Take the calculated average value as the initial point D_P0 for the farthest point sampling. At the same time, add it to the sampling point set D_S = {D_P0}; Step 2.2), calculate the distance matrix. The farthest point sampling selects the point with the farthest distance from D_S from the point cloud dataset. Therefore, in order to obtain the distance between the point cloud data P″ and the sampling point set D_S, the Euclidean distance between each point and the newly added point in D_S is recorded through the distance matrix D_L; Step 2.3), update the set of sampling points, and denote the point corresponding to the maximum value in D_L as D_P i , and add it to D_S. As D_S updates the sampling points, D_L will recalculate the Euclidean distance between the newly added sampling points and the point cloud dataset, and at the same time update the point cloud with a Euclidean distance less than the original value; finally, loop the above operations to increase the number of point clouds in the sampling point set to a fixed value.
5. The three-dimensional point cloud task analysis method for fusing local features and global context information according to claim 1, wherein The process of the fifth step is as follows: Step 5.1), construct a neighborhood index, and use the grouped point cloud features Z local and the grouped point cloud sets obtained in Step 4 to get an index map of neighbor points using the KNN algorithm in Step 3, as shown below Among them, Map o is the index mapping of the nearest points, obtained by querying different indices o region Ω of the nearest points o ; Step 5.2), neighborhood propagation, through the index mapping Map o to calculate the inverse index mapping Map o -1 , backtrack to another region Ω o+1 , query point p query through Map o and Map o -1 obtain all elements in the two sets and re - establish a set where is the coordinate information of the l - th point, is the feature of the l - th point, k_n is the number of all points that p query can contact. Given a point p query , it means that its local region is {Ω1, Ω2, ·····, Ω g}, and g represents the number of associated regions.
6. The three-dimensional point cloud task analysis method for fusing local features and global context information according to claim 1, wherein In step 7, project the feature F obtained in step 6.3) to generate Q c , K c and V c . Multiply the query embedding Q c and the key embedding K c to obtain the weight matrix Then split this matrix and perform repeated operations to obtain N c channel matrices Propagate the spatial information of the points into each channel to obtain the information relationship between the points Pass through the softmax normalization function, and then multiply the weight of a single channel element-wise with the V of the key embedding c to maintain the channel difference, and then update the features of P group . This process is expressed as Q c , K c , V c = MLP(F); Among them, τ is a linear projection, which is an important concept in linear algebra and mathematical analysis and maps a vector space to its subspace. ρ(·) is a repeated operation function that, according to copies D and splices them together to form 7. The three-dimensional point cloud task analysis method for fusing local features and global context information according to claim 6, characterized in that, In the eighth step, after the point cloud set P undergoes steps four, five, six, and seven, it undergoes local feature extraction again through step four to obtain the final local point cloud fine-grained feature F result , F result = LGA(F; X''); Among them, LGA(*) represents performing the operation of the fourth step again, F is the encoded feature obtained after the sixth step, and X” is the point encoded after the sixth step.
8. The three-dimensional point cloud task analysis method for fusing local features and global context information according to claim 7, wherein In the ninth step, according to the point cloud set P constructed in the first step, the key shape points KSP are used to capture the key shape regions of different objects. The positions of the KSP points change dynamically according to different object shapes. Therefore, several KSP points are used to represent the specific deformations of such objects. The formula is as follows: A KSP = softmax(f k ); Δx k = Linear(f k ); where, Δx k is the offset of the point, ξ ksp is the calculated KSP point, and f k is the feature of the input point; GSFM receives the coordinate information of the grouped points after step 4.1), the coordinate information P of all points, and the features. For the convenience of representation, the grouped points are used as the neighborhood boundary spheres, and all points are used as the entire point cloud boundary sphere. The volume ratio γ of the neighborhood boundary sphere and the point cloud boundary sphere is calculated as the slight deformation of the entire object. Then, combined with the specific deformation of the object, more effective global context features are obtained from the shape information in the three-dimensional points. F res = FFN(f ξ + F + F o ) + f ξ + F + F o ); Among them, v L is the volume of the neighborhood boundary sphere corresponding to p i , v G is the volume of the point cloud boundary sphere, is splicing, (x ξ , y ξ , z ξ ) are the coordinates of the KSP point, p i (x n , y n , z n ) is used to represent the position of the local neighborhood, F o is the original point cloud feature, f ξ is the feature of the KSP point, F res is the final point cloud feature.
Citation Information
Patent Citations
Digital twinborn scene-oriented point cloud automatic semantic modeling method
CN116844004A
Point cloud feature extraction method and system based on graph convolutional neural network, and medium
CN117788990A