A 3D Point Cloud Classification Method Based on Multi-View Attention Convolutional Pooling

By adopting the multi-view attention convolution pooling method in three-dimensional point cloud classification, and using Res2Net and attention mechanism to extract and fuse feature information, the problems of feature information loss and detailed information loss in the existing technology are solved, and higher classification accuracy is achieved.

CN113887385BInactive Publication Date: 2025-06-27UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111150171.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing three-dimensional point cloud classification method has problems such as loss of feature information and loss of detailed information during view dimensionality reduction, resulting in low classification accuracy.

Method used

Using a multi-view attention convolution pooling method, multi-view visual features are extracted through Res2Net, and the feature information and detailed information of the view are extracted using attention mechanism and convolution operations to perform feature fusion and classification.

Benefits of technology

It effectively solves the problems of feature information loss and detailed information loss, improves the accuracy of three-dimensional point cloud classification, and the classification accuracy can reach 93.64%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113887385B_ABST
    Figure CN113887385B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional point cloud classification method based on multi-view attention convolutional pooling, including: voxelizing a three-dimensional point cloud model, and using a set of virtual images under different perspectives to replace the virtual three-dimensional model for the voxelized model; extracting visual features through a deep visual feature extraction module, converting the multi-view visual features extracted from the two-dimensional images by Res2Net into a feature map of size m×n, and its input is denoted as f m×n ; converting the feature vector obtained in the previous step through a visual feature fusion and classification module, converting the feature vector into a C×1 feature vector through a fully connected layer, and applying the SoftMax function to handle the classification problem to obtain the probability distribution of the model to be classified. According to the present invention, the problems of feature information loss caused by feature representation and the loss of detailed information in each view during the dimensionality reduction process are effectively solved, higher classification accuracy is obtained, and superior performance can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three - point cloud classification, and particularly relates to a three - dimensional point cloud classification method based on multi - perspective attention convolutional pooling. Background Art

[0002] With the continuous emergence of three - dimensional imaging and scanning devices such as 3D cameras, Kinects, radars, and depth scanners, the acquisition of point cloud data has become increasingly convenient and accurate. Currently, point cloud data has been widely used in fields such as autonomous driving, intelligent robots, virtual reality, medical diagnosis, and medical imaging. Among the many processes of point cloud data, point cloud classification is the basis for tasks such as target recognition and tracking, scene understanding, and three - dimensional reconstruction in the above - mentioned application fields. Therefore, the problem of three - dimensional point cloud classification has become a current research hotspot and has important research significance.

[0003] Currently, traditional machine learning methods have some limitations, such as long training time and low classification accuracy. In the past decade, due to the rapid development of deep learning technology and the emergence of three - dimensional model datasets (such as ShapeNet, ModelNet, PASCAL3D +, and Stanford Computer Vision and Geometry Laboratory datasets, etc.), deep learning methods have been widely applied to the classification task of three - dimensional point cloud data. According to the different convolution objects, point cloud classification methods based on deep learning methods can be divided into three categories, including voxel - based methods, point - cloud - based methods, and view - based methods. Among them, voxel - based methods convert point clouds into voxels of a fixed size and use convolutional neural networks for classification. For point - cloud - based classification methods, the point cloud is directly input into the neural network to complete classification. View - based methods convert three - dimensional point clouds into two - dimensional images from different angles, making the classification problem of three - dimensional point clouds a two - dimensional image classification problem. The existing methods have the following problems: Effective voxel - based methods are usually limited to small datasets or used for single - object classification, and the computational cost is very high. On large datasets, the classification accuracy of this type of method is generally not high. Therefore, this type of method still has great room for improvement; Point - cloud - based methods, due to the non - uniformity of point cloud density, currently cannot perfectly solve the classification problem of three - dimensional point cloud data adapting to non - uniform point sampling density, and the classification accuracy of this type of method also needs to be improved. In addition, the inability of this type of method to determine the detailed location of discrete objects is also a major bottleneck problem; Effective view - based methods use views in each perspective of multiple perspectives to represent three - dimensional models, and usually show that they can achieve higher classification accuracy requirements with less computational requirements. Compared with voxel - based methods and point - cloud - based methods, this type of method has better classification performance, but some feature information or detailed information is lost during the representation or processing of views, resulting in the accuracy of this type of algorithm not being very high and still having a large room for improvement. Summary of the Invention

[0004] Aiming at the deficiencies existing in the prior art, the purpose of the present invention is to provide a three-dimensional point cloud classification method based on multi-view attention convolutional pooling, which effectively solves the problem of feature information loss caused by feature representation and the problem of losing detailed information in the process of dimensionality reduction for each view, obtains higher classification accuracy, and can achieve superior performance. To achieve the above object and other advantages of the present invention, a three-dimensional point cloud classification method based on multi-view attention convolutional pooling is provided, including:

[0005] S1. Voxelize the three-dimensional point cloud model, and use a set of virtual images under different perspectives to replace the virtual three-dimensional model for the voxelized model;

[0006] S2. Extract visual features through a deep visual feature extraction module, and convert the multi-view visual features extracted from the two-dimensional image by Res2Net into a feature map of size m×n, and its input is denoted as f m×n ;

[0007] S3. Convert the feature vector obtained in step S2 through a visual feature fusion and classification module. The feature vector is converted into a C×1 feature vector through a fully connected layer, and the SoftMax function is applied to handle the classification problem to obtain the probability distribution of the model to be classified.

[0008] Preferably, in step S1, n virtual cameras are placed horizontally on a circular orbit at the same interval angle d, and the capture lenses of the virtual cameras are aligned with the center of the three-dimensional model to simulate the scenario of humans observing the model.

[0009] Preferably, in step S2, a variant Res2Net of the ResNet method is used to extract visual features. The 3×3 convolutional layer is evenly divided into p subsets, which is represented by x={x1, x2, x3, …, x p}, and then each subset (excluding x1) is input into a 3×3 convolution, denoted as Conv p , and then starting from x3, before inputting into Conv p , the output of Conv p-1 is added to increase the possible receptive field within one layer.

[0010] Preferably, the attention-convolutional pooling includes extracting the feature information of the view using the attention mechanism and extracting the detailed information of the view using the convolutional operation. The attention mechanism extracts the feature information of the view to generate three feature maps through four 1×1 convolutional layers, denoted as Query, Key1, Key2, and Value respectively. The feature map Query is transposed into a feature map Q of size n×m T, perform product operations with the feature maps Key1 and Key2 respectively to obtain two feature maps of n×n and Perform product operations on the obtained and Perform another product operation to obtain a feature map of n×n, denoted as f n×n ; then use the softmax activation function as the attention weight; then perform a product operation on Value and the attention weight, and use max pooling to perform dimensionality reduction processing on it to obtain a feature vector of m×1

[0011] Preferably, the product operation for extracting the detailed information of the view includes the original feature map f m×n Generate a feature map of m×1 through a convolutional layer of 1×n, denoted as After the attention mechanism extracts the feature information of the view and the convolutional operation extracts the detailed information of the view, the two feature vectors of m×1 and are concatenated into a feature vector of 2m×1.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0013] (1) By introducing the variant Res2Net of ResNet, feature extraction is performed on a group of multi-view 2D view images, further improving the accuracy in the 3D model classification task.

[0014] (2) Using the attention-convolution mechanism, it can more focus on finding useful information related to the current output in the input data for processing, effectively solving the problem of feature information loss caused by feature representation and the loss of detailed information of each view during the dimensionality reduction process, thereby improving the classification accuracy.

[0015] (3) A large number of experiments are carried out to verify the performance of the proposed method. Compared with several popular methods, the experimental results show that the accuracy of the classification method has been significantly improved, and the classification accuracy can reach 93.64%, proving that the classification framework achieves advanced performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a framework diagram of the 3D point cloud classification method based on multi-view attention convolution pooling according to the present invention;

[0017] Figure 2 is a process diagram of generating 6-view and 12-view data of the 3D model according to the 3D point cloud classification method based on multi-view attention convolution pooling of the present invention;

[0018] Figure 3Schematic diagram of the Res2Net structure used in the 3D point cloud classification method based on multi-view attention convolution pooling of the present invention.

[0019] Figure 4 Flow chart of the attention convolution pooling of the 3D point cloud classification method based on multi-view attention convolution pooling according to the present invention;

[0020] Figure 5 Flow chart of the visual feature fusion classification module of the 3D point cloud classification method based on multi-view attention convolution pooling according to the present invention. Detailed implementation manner

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] Refer to Figures 1-5 , a 3D point cloud classification method based on a multi-view attention-convolution pooling network (MVACPN) includes the following steps: For an original 3D point cloud model, first voxelize it, and then use a set of virtual images V = {V1, V2, V3,..., V n} to replace the virtual 3D model, where V n represents n virtual images generated by a 3D model at n viewpoints. In this process, n virtual cameras are placed horizontally on a circular orbit at the same interval angle d, and the capture lenses of the virtual cameras are aligned with the center of the 3D model to simulate the scenario of humans observing the model. The relationship between the interval angle d and the number n of virtual cameras is In this article, three types of virtual camera arrangements are set. The first is to place 3 virtual cameras and set the interval angle to d = 120°, obtaining 3 views; the second is to place 6 virtual cameras, so the interval angle d = 60° needs to be set, obtaining 6 views, as shown in Figure 2 -a; the third is to place 12 virtual cameras, so the interval angle d = 30° needs to be set, obtaining 12 views, as shown in Figure 2 -b. It can be seen that the method used can also be used for the generation of multi-views at other viewpoints.

[0023] For a set of multi-views V = {V1, V2, V3,..., Vn}, in order to further increase the number of acceptable fields, make the feature extraction ability more powerful, and thus reduce the information loss problem during feature extraction, a variant Res2Net of the ResNet method is used to extract visual features. As Figure 3 shown in the Res2Net module, the basic block in the ResNet structure is replaced by the Res2Net module. First, the 3×3 convolutional layer is evenly divided into p subsets, represented by x = {x1, x2, x3, …, x p}, and then each subset (excluding x1) is input into a 3×3 convolution, denoted as Conv p , and then starting from x3, before inputting Conv p , the output of Conv p-1 is added to increase the possible receptive field within one layer. The process formula of the Res2Net module can be expressed as:

[0024]

[0025] where, y = {y1, y2, y3, …, y p} is the output of the Res2Net module, and then it is connected and passed into a 1×1 convolutional layer to ensure the channel size of the Res2Net residual module.

[0026] The multi-view visual features extracted from the two-dimensional image by Res2Net are transformed into a feature map of size m×n, and its input is denoted as f m×n . The attention-convolutional pooling proposed in this paper can be mainly divided into two parts, as Figure 4 shown. The first part is to use the attention mechanism to extract the feature information of the view; the second part is to use the convolutional operation to extract the detailed information of the view.

[0027] In the first part, for f m×n , three feature maps are generated through four 1×1 convolutional layers, denoted by respectively, and this process can be expressed by formula (2):

[0028]

[0029] where, f m×n represents the input feature map of size m×n, Conv 1×1 is the convolution operation using a 1×1-sized convolution kernel, and Query, Key1, Key2, and Value tables are the feature maps obtained after convolution operations using 1×1-sized convolution kernels.

[0030] Then, the feature map Query is transposed into a feature map Q of size n×mT Perform product operations with the feature maps Key1 and Key2 respectively to obtain two feature maps of n×n and This process can be expressed by formula (3):

[0031]

[0032]

[0033]

[0034] where T represents the transpose of the feature map, represents the product operation between two feature maps.

[0035] Then, perform a product operation on the obtained and again to obtain a feature map of n×n, denoted as f n×n ; then use the softmax activation function as the attention weight; then perform a product operation with the attention weight and use max pooling to reduce its dimension to obtain a feature vector of m×1 This process can be expressed by formula (4):

[0036]

[0037] where softmax represents the activation function softmax, and Max represents max pooling.

[0038] In the second part, generate a feature map of m×1 from the original feature map f m×n through a convolutional layer of 1×n, denoted as This process can be expressed by formula (5):

[0039]

[0040] where f m×n represents the input feature map of size m×n, represents the feature map obtained after convolution operation using a convolutional kernel of size 1×n, and Conv 1×n is the convolution operation using a convolutional kernel of size 1×n.

[0041] After the first part and the second part are completed, concatenate the two feature vectors of m×1 and into a feature vector of 2m×1. This process can be expressed by formula (5):

[0042]

[0043] Among them, Cat represents concatenation, and f A-CP is the feature vector obtained after the attention-convolutional pooling proposed in this paper.

[0044] After obtaining the feature vector.

[0045] It can be sorted out as follows:

[0046]

[0047] Through the previous operations, a 2m×1 feature vector is obtained, which represents the feature information and detail information of each view. Here, a fully connected layer is added, and it is transformed into a C×1 feature vector through the fully connected layer, and the SoftMax function is applied to handle the classification problem to obtain the probability distribution of the model to be classified, as Figure 5 shown. This process can be expressed by formula (8):

[0048]

[0049] Among them, x is the input of the fully connected layer, w is the weight, and b is the bias. is the probability output by SoftMax, and the calculation method of SoftMax is as follows:

[0050]

[0051] Among them, C is the category of the dataset. For example, when using the ModelNet40 dataset, C is set to 40 here.

[0052] In the experiment, the classification performance of the method measured by the overall accuracy OA (Overall Accuracy) of each sample and the average accuracy AA (Average Accuracy) of each category is used. They are defined as

[0053] · The overall accuracy OA (Overall Accuracy) of each sample: It refers to the ratio of the number of samples correctly classified to the total number of samples, and can be expressed by the formula as:

[0054]

[0055] Among them, N is the total number of samples, and x ii is the number of correctly classified samples distributed along the diagonal of the confusion matrix, and C is the category of the dataset.

[0056] · The average accuracy AA (Average Accuracy) of each category: It refers to the ratio between the number correctly predicted for each category and the total number of each category, and finally takes the average of the accuracies of each category, and can be expressed by the formula as:

[0057]

[0058] Among them, recall represents the ratio of the correctly predicted ones in the actual samples, and C represents the number of categories.

[0059] Equipped with 2 NVidia Titan Xp GPUs and 64GB of memory, all experiments were conducted using the PyTorch platform. During the experiments, there were two stages for training. The number of training times in the first stage and the second stage were set to 10 and 20 respectively. In the first stage, only single pictures were classified to fine-tune the model; in the second stage, all views of the model after voxelization of the original 3D point cloud model were trained, so that the entire classification framework could be trained. And in the test stage of the experiment, only the second stage was tested.

[0060] To optimize the overall architecture, Adam was used as the optimizer for both stages. In addition, learning rate decay and weight decay were also set. The initial value of the learning rate (lr) was set to 0.0001, and then the next learning rate was adjusted to half of the previous one. The weight decay used L2 regularization, aiming to accelerate the training of the model and reduce overfitting of the model.

[0061] When selecting the CNN framework for extracting visual features of the views, experiments were conducted and compared with the VGG-11 model proposed by Simonyan et al., the ResNet-50 model proposed by He et al., the variants Res2Net-50 and Res2NeXt-50 of the ResNet model proposed by Gao et al., and the DenseNet-121 model proposed by Huang et al., which were used as the backbone models for the deep visual feature extraction module in the framework. The experimental results are shown in TABLE-Ⅰ. Here, the learning rate (lr) was set to 5×10 -5 , the batch size (bs1) in the first stage and the batch size (bs2) in the second stage were set to 64 and 16 respectively. The most classic max pooling method was used for the feature pooling module, and N was assigned 6 in the N views. It can be seen that both Res2Net-50 and Res2NeXt-50 exceeded 92% and 90% in terms of the performance of OA and AA. Therefore, two models (the sub-optimal Res2Net-50 and the optimal Res2NeXt-50) were selected as the backbone models for subsequent experiments.

[0062] TABLE-Ⅰ. Influence results of different backbone models on classification performance.

[0063]

[0064] Influence of different values of N in the N views on classification performance

[0065] When this paper explores the influence of different values of N in the N - perspective on the classification performance, comparative experiments are conducted by taking N = 3, N = 6, and N = 12 respectively. The experimental results are shown in TABLE - Ⅱ. In the framework model, by adjusting the hyperparameters (learning rate and batchsize), better performance can be achieved. Here, the hyperparameters are set as lr = 1×10 -4 , bs1 = 128, bs2 = 32.

[0066] Through experimental comparison, it is found that in each perspective, the model framework exceeds other methods. And in the 6 - perspective and 12 - perspective, the method of this application exceeds 93% in the performance of OA, reaching a better level. It is worth noting that the classification performance of the method of this application is the best when N = 6. Moreover, as the value of N increases, the training time becomes longer. When N = 6, the training time can be shorter. In the model framework, comparing the backbone model Res2Net - 50 and the backbone model Res2NeXt - 50, the performance of Res2NeXt - 50 in terms of OA and AA is 93.64% and 91.53% respectively when N = 6, both reaching the optimal. Therefore, the backbone model Res2NeXt - 50 and N = 6 are used as the optimal model framework configuration.

[0067] TABLE - Ⅱ. Influence of different values of N in the N - perspective on the classification performance.

[0068]

[0069] Through experiments, the influence of different factors on the classification performance of 3D point clouds is explored, the experimental settings of the optimal model are determined, and the method of this application is compared with the state - of - the - art methods, including voxel - based effective methods, point - cloud - based effective methods, and view - based effective methods. The comparison results are shown in TABLE - Ⅲ.

[0070] The results show that compared with other methods, the algorithm framework proposed in this application has better performance, obtaining a competitive advantage with classification accuracies of 93.64% and 91.53% in terms of performance OA and AA, indicating that the model has high - precision classification ability. This is mainly due to two reasons: (i) MVACPN further increases the number of receptive fields by introducing Res2Net, making the feature extraction ability more powerful, thus reducing the information loss problem in the feature extraction process; (ii) The attention - convolutional feature pooling module in MVACPN includes an attention mechanism and a convolutional operation, using the attention mechanism to extract the feature information of the view and the convolutional operation to extract the detailed information of the view. Compared with traditional methods,

[0071] TABLE - Ⅲ. Comparison of classification results on the ModelNet40 dataset

[0072]

[0073] This method can more focusedly find the useful information related to the current output in the input data, effectively solving the problem of feature information loss caused by feature representation and the loss of detailed information of each view during the dimensionality reduction process, thereby improving the classification accuracy.

[0074] In this application, a multi-view attention-convolutional pooling network framework (MVACPN) is proposed for high-precision 3D point cloud classification tasks. Considering the problem of feature information loss caused by feature representation and the loss of detailed information of each view during the dimensionality reduction process, an attention-convolutional pooling structure is proposed. By using attention-convolution operations, the useful information related to the current output in the input data can be found more focusedly, thereby improving the classification accuracy. A large number of experiments are conducted to obtain the optimal model settings and achieve the best classification accuracy. The model is evaluated on ModelNet 40. The experimental results show that compared with the state-of-the-art methods, the framework can achieve higher classification accuracy, demonstrating its superiority.

[0075] The number of devices and the processing scale described here are used to simplify the description of the present invention. The applications, modifications, and variations of the present invention are obvious to those skilled in the art.

[0076] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. A 3D point cloud classification method based on multi-view attention convolutional pooling, characterized in that Including the following steps: S1. Voxelize the three-dimensional point cloud model, and use a set of virtual images from different perspectives to replace the virtual three-dimensional model for the voxelized model; S2. Extract visual features through the deep visual feature extraction module, convert the multi-view visual features extracted from the two-dimensional image by Res2Net into a feature map of size m×n, and its input is expressed as ; S3. Transform the feature vectors obtained in step S2 through the visual feature fusion classification module. The feature vectors are transformed into C×1 feature vectors through a fully connected layer, and the SoftMax function is applied to handle the classification problem to obtain the probability distribution of the model to be classified; In step S2, the attention mechanism is used to extract the feature information of the view, and the convolution operation is used to extract the detailed information of the view. The attention mechanism extracts the feature information of the view by generating three feature maps through four 1×1 convolutional layers, which are represented by Query, Key1, Key2, and Value respectively. The feature map Query is transposed into a feature map QT of size n×m, and product operations are performed with the feature maps Key1 and Key2 respectively to obtain two feature maps of size n×n and , and the obtained and are subjected to another product operation to obtain a feature map of size n×n, denoted as ; Then use the softmax activation function as the attention weight; Then, multiply Value with the attention weights and perform a dimensionality reduction operation on it using max pooling to obtain an m×1 feature vector ; The product operation to extract the detailed information of the view includes generating an m×1 feature map from the original feature map through a 1×n convolutional layer, denoted as . After the attention mechanism extracts the feature information of the view and the convolutional operation extracts the detailed information of the view, the two m×1 feature vectors and are concatenated into a 2m×1 feature vector.

2. The 3D point cloud classification method based on multi-view attention convolutional pooling according to claim 1, characterized in that In step S1, n virtual cameras are placed horizontally on a circular orbit at the same interval angle d, and the capture lenses of the virtual cameras are aligned with the center of the three-dimensional model to simulate the scenario of human observing the model.

3. The 3D point cloud classification method based on multi-view attention convolutional pooling according to claim 1, wherein, In the step S2, the mutant Res2Net of the ResNet method is used to extract visual features. The 3×3 convolutional layer is evenly divided into p subsets, denoted by . Then, each subset excluding is input into the 3×3 convolution, denoted as . Then, starting from , before inputting , the output of is added to increase the receptive field possible within one layer.