A General and Efficient 3D Point Cloud Analysis Method Based on Distillation Algorithm
By adopting a general and efficient method based on distillation algorithm in 3D point cloud analysis, the problem of designing a general and efficient 3D point cloud analysis system is solved, and excellent performance and model size reduction in segmentation and classification tasks are achieved.
Patent Information
- Application Number
- CN202311121853.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-09-01
AI Technical Summary
The prior art is difficult to design a 3D point cloud analysis system that is both universal and efficient, and lacks a unified perspective in model comparison and analysis, and model size and efficiency are challenges in mobile device deployment.
Using a general and efficient 3D point cloud analysis method based on distillation algorithm, building blocks are divided into neighborhood updaters, neighborhood aggregators, point updaters and space embedding generators through unified building blocks and model compression technology, and their best combination is explored to obtain superior performance and smaller models using knowledge distillation.
Excellent performance in segmentation and classification tasks is achieved, while reducing model size and improving model efficiency, providing a general framework for system comparison and analysis.
Smart Images

Figure CN117237284B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a general and efficient 3D point cloud analysis method based on a distillation algorithm. Background Art
[0002] In the past decade, the popularization of 3D data acquisition technology and the progress of point cloud analysis have promoted the increasing application of 3D point clouds in daily life. 3D point cloud analysis refers to the process of processing and analyzing a large amount of discrete three-dimensional point data obtained by lidar, cameras, or other sensors. These points can represent information such as the shape, size, and position of the object surface.
[0003] In the industrial field, 3D point cloud analysis is widely used in product design, manufacturing, and quality control. For example, in the automotive industry, 3D point cloud analysis can be used to detect quality parameters such as the dimensions and surface flatness of parts and the entire vehicle; in the aerospace field, 3D point cloud analysis can be used to detect defects such as deformation and fatigue cracks in the aircraft hull; in the field of architecture and urban planning, 3D point cloud analysis can be used to establish accurate terrain models to assist in design and construction. In addition, 3D point cloud analysis is also widely used in the fields of digital cultural heritage protection, medical image analysis, virtual reality, and augmented reality.
[0004] Although convolutional neural networks have achieved success in image processing, their ability to process point clouds is limited by their sparsity, disorder, and irregularity. Early methods attempted to apply convolutional neural networks to point clouds through 2D projection or voxelization, but these methods resulted in additional computations and information loss. PointNet (Qi C R, Su H, Mo K, et al. Pointnet: Deep learning on point sets for 3d classification and segmentation[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 652-660.) and PointNet++ (Qi C R, Yi L, Su H, et al. Pointnet++: Deep hierarchical feature learning on point sets in a metric space[J]. Advances in neural information processing systems, 2017, 30.) introduced permutation equivariance and permutation invariance, enabling convolutional neural networks to process unstructured point clouds. However, due to the lack of a unified perspective in these methods, it is difficult to systematically compare and analyze these models. This issue has received extensive attention from researchers, who have continuously tried to find a unified perspective to compare these methods in order to identify the most critical implementation details. PointNeXt (Qian, G., Li, Y., Peng, H., Mai, J., Hammoud, H., Elhoseiny, M., & Ghanem, B. (2022). Pointnext: Revisiting pointnet++ with improved training and scaling strategies. Advances in Neural Information Processing Systems, 35, 23192-23204.) found a comparable perspective. But it did not provide a general framework. In addition, deploying 3D point cloud models on mobile devices usually has high requirements for both the model size and efficiency, so the size of 3D point cloud models also needs attention. Based on the above analysis, it can be concluded that developing a general and efficient framework is a crucial research issue. Summary of the Invention
[0005] The object of the present invention is to provide a general and efficient 3D point cloud analysis method based on the distillation algorithm (abbreviated as MetaPointDistille) for the above problems existing in the prior art. By using unified building blocks and model compression techniques, the building blocks are divided into a neighborhood updater, a neighborhood aggregator, a point updater, and a spatial embedding generator, and their optimal combinations are explored. Knowledge distillation is used to obtain a model with superior performance and a smaller scale.
[0006] The present invention includes the following steps:
[0007] 1) Input 3D point cloud data;
[0008] 2) Train the teacher model;
[0009] 3) The 3D point cloud data is initially feature-extracted through a Multi-Layer Perceptions (MLPs) layer;
[0010] 4) The features extracted in step 3) are further feature-extracted by a Feature Extraction module;
[0011] 5) The features output in step 4) are analyzed by a Meta Combination module;
[0012] 6) Repeat the steps in step 5) K times;
[0013] 7) Repeat the steps in steps 4) to 6) 3 times;
[0014] 8) Perform Feat.Propagation calculations on the information output in steps 6) and 7);
[0015] 9) Repeat the steps in step 8) 3 times;
[0016] 10) Train the teacher model until the loss converges;
[0017] 11) Distill the student model based on the teacher model;
[0018] 12) Given any 3D point cloud data, input it into the student model, and the student model outputs the analysis result.
[0019] The characteristics and effects of the present invention are as follows:
[0020] The general and efficient 3D point cloud analysis method based on the distillation algorithm proposed by the present invention solves a common challenge in 3D point cloud analysis: how to design a general and efficient point cloud system. The present invention decomposes it into two problems, the model universality problem and the model size problem, and uses a framework based on Meta Combination and distillation learning to solve them respectively. The present invention emphasizes the concept of basic building blocks and realizes their optimal combination, called Meta Combination, which not only allows for a more fair comparison of point cloud systems at the macroscopic level but also allows for a more specific design at the microscopic level. Distillation learning is used to further reduce the model size and improve the model performance, which also indirectly proves the universality of the method of the present invention. The present invention achieves excellent results in segmentation and classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a framework diagram of the present invention.
[0022] Figure 2 It is a comparison of component combinations between PointNet++ and the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following embodiments will describe the present invention in detail with reference to the accompanying drawings.
[0024] The latest progress in 3D point cloud analysis has introduced various network architectures into this field. However, due to the lack of a standard framework for describing these networks, it is difficult to conduct a systematic comparison and analysis of them. In addition, due to the emergence of large language models, the network scale of these models is often very large. The present invention designs a general and efficient point cloud analysis framework, provides a general and comparable framework by using basic building blocks, reorganizes the basic building blocks into a neighborhood updater, a neighborhood aggregator, a point updater, and a spatial embedding generator, attempts their best implementation and combination, called Meta Combination, to further implement a general framework. In addition, knowledge distillation is applied on the basis of Meta Combination to further reduce the model size, forming a general and efficient framework.
[0025] The process of the method of the present invention is as Figure 1 shown and includes the following steps:
[0026] 1) Input 3D point cloud data;
[0027] 2) Train a teacher model;
[0028] 3) The 3D point cloud data is initially feature-extracted through a Multi-Layer Perceptions (MLPs) layer;
[0029] 4) Further extract features from the features extracted in step 3) using the Feature Extraction module;
[0030] 41) Subsample the 3D point cloud data;
[0031] 42) Group the subsampled data;
[0032] 43) Use MLPs to extract the point cloud features after grouping;
[0033] 5) Analyze the features output in step 4) through the Meta Combination module;
[0034] 51) Extract features using the MLPs layer;
[0035] 52) Group the features (neighborhood updater);
[0036] 53) Add spatial position information (spatial embedding generator);
[0037] 54) Purify the features through the Max-pooling layer (neighborhood aggregator);
[0038] 55) Extract features using the MLPs layer;
[0039] 56) Perform a residual connection with the most original input features (point updater);
[0040] 6) Repeat the steps in step 5) K times;
[0041] 7) Repeat the steps in steps 3) - 5) 3 times;
[0042] 8) Perform Feat.Propagation calculation on the information output in steps 6) and 7);
[0043] 81) Interpolate the point cloud features output in the previous stage;
[0044] 82) Concatenate with the features output by the previous Meta Combination;
[0045] 83) Extract features using the MLPs layer;
[0046] 9) Repeat the steps in step 8) 3 times;
[0047] 10) Train the teacher model until the loss converges;
[0048] 11) Distill the student model based on the teacher model;
[0049] 111) Calculate the probability p assigned by the teacher model to the i-th class of the sample at temperature T i :
[0050]
[0051] Among them, v i represents the logits of the i-th class in this sample in the teacher model, and T is the temperature used in knowledge distillation.
[0052] 112) Calculate the probability q i assigned by the student model to the i-th class of the sample at temperature T
[0053]
[0054] Among them, z i represents the logits of the i-th class in this sample in the student model, and T is the temperature used in knowledge distillation.
[0055] 113) Calculate the soft label loss L soft between the teacher and student models
[0056]
[0057] Among them, N represents the number of samples, M represents the number of classes, p i and q i correspond to the probabilities assigned by the teacher model and the student model to the i-th class at temperature T, respectively.
[0058] 114) Calculate the hard label loss L hard between the predicted output of the student model and the true label
[0059]
[0060] Among them, cj represents the true value of the logits of the j-th class in this sample.
[0061] 115) The final distillation loss function L distill of the model is
[0062] L distill = αL soft + (1 - α)L hard (5)
[0063] Among them, α is the balance factor between L soft and L hard
[0064] Note: The difference between the student model and the teacher model lies only in the MLP channel size and the number of Meta Combinations in each stage. Their detailed structures are as follows:
[0065] Student model: C = 32, B = (2, 4, 2, 2)
[0066] Teacher model: C = 64, B = (4, 8, 4, 4)
[0067] Among them, C represents the channel size of the MLP, and B represents the number of Meta Combinations in each stage.
[0068] 12) Given any 3D point cloud data, input it into the student model, and the student model outputs the analysis result.
[0069] Figure 1 is the MetaPointDistill framework diagram. The architecture of MetaPointDistill consists of several feature extraction, Meta Combination, and feature propagation blocks. The present invention adopts a two-stage method. First, a larger teacher model is trained, and then knowledge distillation is applied to obtain a more efficient student model. The difference between the teacher and student models lies in the different numbers of Meta Combinations in each stage.
[0070] Figure 2 is a comparison of the component combinations between PointNet++ (Qi C R, Yi L, Su H, et al. Pointnet++: Deep hierarchical feature learning on point sets in a metric space[J]. Advances in neural information processing systems, 2017, 30.) and the method of the present invention. PointNet++ adopts implicit space embedding and ignores the combination with the original features, while the method of the present invention uses explicit space embedding and emphasizes the original features more.
[0071] Table 1 shows the performance comparison between the present invention and other methods on the S3DIS Area-5 dataset, where MetaPointDistll is the present invention.
[0072] Table 1
[0073]
[0074] As can be seen from Table 1, the present invention achieves excellent results in segmentation and classification tasks.
[0075] The latest advancements in 3D point cloud analysis have introduced various network architectures into this field. However, due to the lack of a standard framework for describing these networks, it is difficult to conduct systematic comparisons and analyses on them. Additionally, due to the emergence of large language models, the network scale of these models tends to be very large. This invention proposes a method called MetaPointDistille, which employs unified building blocks and model compression techniques. When comparing systems, a unified framework structure is needed for fair and convenient comparison. At the same time, the scale of this framework cannot be too large, so model compression techniques need to be used. This invention divides the building blocks into neighborhood updaters, neighborhood aggregators, point updaters, and spatial embedding generators, and explores their optimal combinations. Furthermore, this invention uses knowledge distillation to obtain models with superior performance and smaller scales. In summary, the method of this invention demonstrates great potential in addressing the challenges encountered in the actual production process.
Claims
1. A general and efficient 3D point cloud analysis method based on a distillation algorithm, characterized in that It includes the following steps: 1) Input 3D point cloud data; 2) Train the teacher model; 3) The 3D point cloud data is initially feature-extracted through the MLPs layer; 4) The features extracted in step 3) are further used to extract 3D point cloud features by the Feature Extraction module; 5) The features output in step 4) are further analyzed through the Meta Combination module; 6) Repeat the steps in step 5) K times; 7) Repeat the steps in steps 4) - 6) 3 times; 8) Perform Feat.Propagation calculation on the information output in steps 6) and 7): 81) Interpolate the point cloud features output in the previous stage; 82) Concatenate with the features output by the previous Meta Combination; 83) The MLPs layer extracts features; 9) Repeat the steps in step 8) 3 times; 10) Train the teacher model until the loss converges; 11) Distill the student model based on the teacher model, specifically as follows: 111) Calculate the probability p assigned by the teacher model to the i-th class of the sample at temperature T i : where, v i represents the logits of the i-th class in the teacher model for this sample, and T is the temperature used in knowledge distillation; 112) Calculate the probability q assigned by the student model to the i-th class of the sample at temperature T i : Among them, z i represents the logits of the i-th class in this sample in the student model, and T is the temperature used in knowledge distillation; 113) Calculate the soft label loss L between the teacher and student models soft : where N represents the number of samples, M represents the number of classes, and p i and q i correspond to the probabilities assigned to the i-th class by the teacher model and the student model at temperature T, respectively; 114) Calculate the hard-label loss L between the predicted output of the student model and the true label hard : Among them, cj represents the true value of the logits of the j-th class in this sample; 115) The final distillation loss function L of the model distill is as follows: L distill = αL soft +(1 - α)L hard (5) where α is the balance factor of L soft and L hard balance factor; The differences between the student model and the teacher model are only in the MLP channel size and the number of Meta Combinations in each stage, and their detailed structures are as follows: Student model: C = 32, B = (2, 4, 2, 2) Teacher model: C = 64, B = (4, 8, 4, 4) Among them, C represents the MLP channel size, and B represents the number of Meta Combinations in each stage; 12) Given any 3D point cloud data, input it into the student model, and the student model outputs the analysis result.
2. The general and efficient 3D point cloud analysis method based on the distillation algorithm according to claim 1, wherein In step 4), the extraction of 3D point cloud features includes: 41) Subsample the 3D point cloud data; 42) Group the subsampled data; 43) Use MLPs to extract the point cloud features of the grouped data.
3. The general and efficient 3D point cloud analysis method based on the distillation algorithm according to claim 1, characterized in that In step 5), further extract 3D point cloud features by combining the point cloud neighborhood: 51) The MLPs layer extracts features; 52) Group the features; 53) Add spatial position information; 54) Through the max-pooling layer Max-pooling, purify the features; 55) The MLPs layer extracts features; 56) Perform a residual connection with the most original input features.