Three-dimensional point cloud completion method and system based on multi-scale structured knowledge distillation
By employing a multi-scale structured knowledge distillation method, a 3D point cloud completion network model was constructed, which solved the problem of loss of high-order shape details in point cloud data completion and achieved more complete and robust point cloud reconstruction.
Patent Information
- Application Number
- CN202511395403.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Existing technologies often lead to the loss of high-order shape details during point cloud data completion, especially in the encoding of incomplete point clouds, making it impossible to accurately understand or reconstruct them. Furthermore, existing methods neglect the understanding of semantic information during the encoding process.
A multi-scale structured knowledge distillation method is adopted. By constructing a multi-scale hierarchical knowledge self-distillation encoding module and a reconstruction decoding module, the differences between different local features are learned by maximizing mutual information. The distribution of the deepest module is designated as supervision to construct a 3D point cloud completion network model and achieve complete point cloud completion.
It improves the completeness of point cloud data completion, enabling more effective recovery of high-order shape details of point clouds and enhancing the robustness and accuracy of point cloud reconstruction.
Smart Images

Figure CN120876323A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of point cloud data completion technology, and in particular to a three-dimensional point cloud completion method and system based on multi-scale structured knowledge distillation. Background Technology
[0002] Current research on point cloud data mainly falls into two categories: reconstruction and understanding. PointNet and PointNet++ establish the basic models for point cloud understanding, while PCN and PU-Net define the process for point cloud reconstruction (including completion and upsampling). Their common characteristic is that, before subsequent understanding or reconstruction modules, point cloud data, whether partial or complete, should always be encoded as an implicit encoding with high-order information. Furthermore, these basic models primarily employ simple aggregation operations for encoding, as this is a prerequisite for encoding unordered point clouds. Aggregation typically requires symmetric functions, such as max pooling, summation, and global set abstraction. Therefore, due to information loss, these methods inevitably cannot accurately understand or reconstruct shapes. At the same time, simply encoding point clouds, especially incomplete ones, can lead to the loss of high-order shape details. In real-world scanning, the quality of point clouds varies greatly; they can be sparse, discontinuous, or incomplete. This significantly increases the need to encode consistent and robust representations from point clouds with missing regions.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main objective of this application is to propose a 3D point cloud completion method and system based on multi-scale structured knowledge distillation, which can learn the differences between different local features by maximizing the mutual information between them, and specify the distribution of the deepest module as the supervision of the previous modules, thereby improving the completeness of data completion.
[0005] To achieve the above objectives, one aspect of this application proposes a 3D point cloud completion method based on multi-scale structured knowledge distillation, the method comprising: Constructing an incomplete 3D point cloud dataset; A multi-scale hierarchical knowledge self-distillation encoding module and a reconstruction decoding module are introduced to construct a 3D point cloud data completion network model. The incomplete 3D point cloud dataset is completed by using the 3D point cloud data completion network model to obtain a complete 3D point cloud dataset.
[0006] In some embodiments, the 3D point cloud data completion network model includes a multi-scale hierarchical knowledge self-distillation encoding module, a multilayer perceptron module, and a reconstruction decoding module, wherein the multi-scale hierarchical knowledge self-distillation encoding module, the multilayer perceptron module, and the reconstruction decoding module are connected sequentially. The specific expression of the loss function of the 3D point cloud data completion network model is as follows: ; In the above formula, This represents the loss function for a 3D point cloud data completion network model. Represents the empirical coefficient. Represents the reconstruction loss function. This represents the self-distillation loss function. This represents the corresponding completion result for the target set. Representing a complete truth point cloud, This represents the predicted point cloud output by the reconstruction decoding module. This indicates the actual resolution value output by the reconstruction decoding module.
[0007] In some embodiments, the multi-scale hierarchical knowledge self-distillation encoding module includes several self-distillation modules, each of which includes a k-nearest neighbor layer, a farthest point sampling layer, a multilayer perceptron layer, and a max pooling layer. The reconstruction decoding module includes a PointNet layer, a local attention mechanism layer, a one-dimensional deconvolution layer, a first multilayer perceptron layer, a second multilayer perceptron layer, and an upsampling layer.
[0008] In some embodiments, the step of performing 3D point cloud completion on the incomplete 3D point cloud dataset based on the 3D point cloud data completion network model to obtain a complete 3D point cloud dataset includes: The incomplete 3D point cloud dataset is input into the 3D point cloud data completion network model; Based on the multi-scale hierarchical knowledge self-distillation coding module of the 3D point cloud data completion network model, the incomplete 3D point cloud dataset is subjected to self-distillation loss coding to obtain sparse point cloud data and neighborhood features. The multilayer perceptron module based on the three-dimensional point cloud data completion network model performs data dimensionality reduction processing on the sparse point cloud data and the neighborhood features to obtain a coarse point cloud dataset. The reconstruction and decoding module based on the 3D point cloud data completion network model reconstructs the coarse point cloud dataset to obtain the complete 3D point cloud dataset.
[0009] In some embodiments, the multi-scale hierarchical knowledge self-distillation encoding module based on the 3D point cloud data completion network model performs self-distillation loss encoding on the incomplete 3D point cloud dataset to obtain sparse point cloud data and neighborhood features, including: The incomplete 3D point cloud dataset is input into the multi-scale hierarchical knowledge self-distillation encoding module; Based on the self-distillation module of the multi-scale hierarchical knowledge self-distillation coding module, the self-distillation module extracts features from the incomplete 3D point cloud dataset to obtain several preliminary neighborhood features. The preliminary neighborhood features are calculated using the softmax activation function to obtain the probability distributions corresponding to several preliminary neighborhood features; Using the probability distribution corresponding to the last preliminary neighborhood feature as a supervision signal, knowledge feedback is achieved through KL divergence, and the sparse point cloud data and the neighborhood feature are output.
[0010] In some embodiments, the self-distillation module based on the multi-scale hierarchical knowledge self-distillation encoding module performs feature extraction on the incomplete 3D point cloud dataset to obtain several preliminary neighborhood features, including: The incomplete 3D point cloud dataset is input into the self-distillation module of the multi-scale hierarchical knowledge self-distillation encoding module; Based on the k-nearest neighbor layer of the self-distillation module, a k-nearest neighbor search is performed on the incomplete 3D point cloud dataset to construct local nearest neighbor features of the point cloud data. Based on the farthest point sampling layer of the self-distillation module, the center point of the local nearest neighbor feature of the point cloud data is sampled to obtain the center point of the local nearest neighbor feature. Based on the multilayer perceptron layer of the self-distillation module, feature mapping is performed on the center point of the local nearest neighbor feature to obtain the mapped local nearest neighbor feature; Based on the max pooling layer of the self-distillation module, the mapped local nearest neighbor features are aggregated to obtain several preliminary neighborhood features.
[0011] In some embodiments, the loss function of the self-distillation module is specifically as follows: ; In the above formula, This represents the loss function of the self-distillation module. Denotes KL divergence, This represents the predicted label distribution. Represents the true label distribution. Indicates the distribution of the last layer. .
[0012] In some embodiments, the reconstruction decoding module based on the 3D point cloud data completion network model reconstructs the coarse point cloud dataset to obtain the complete 3D point cloud dataset, including: The coarse point cloud dataset is input into the reconstruction and decoding module of the 3D point cloud data completion network model; Based on the PointNet layer of the reconstruction decoding module, point cloud registration processing is performed on the coarse point cloud dataset to obtain the current point cloud neighborhood features; Based on the local attention mechanism layer of the reconstruction decoding module, local feature extraction processing is performed on the current point cloud neighborhood features to obtain local features; Obtain the global features of the coarse point cloud dataset and combine them with the local features to obtain the point cloud association features; Based on the one-dimensional deconvolution layer of the reconstruction decoding module, the point cloud association features are copied upwards to obtain the copied point cloud association features. Based on the first and second multilayer perceptron layers of the reconstruction decoding module, feature mapping is performed on the copied point cloud associated features to obtain the point-by-point displacement of the point cloud features. Based on the upsampling layer of the reconstruction decoding module, the coarse point cloud dataset is upsampled to obtain the upsampled coarse point cloud dataset. The point-by-point displacement of the point cloud features is added to the upsampled coarse point cloud dataset to obtain the complete 3D point cloud dataset.
[0013] In some embodiments, the loss function of the reconstruction decoding module is specifically as follows: ; In the above formula, This represents the loss function of the reconstruction decoding module. This represents a non-complete 3D point cloud dataset. Represents a complete 3D point cloud dataset. This represents point cloud data points in an incomplete 3D point cloud dataset. This represents the point cloud data points in a complete 3D point cloud dataset. This represents the L2 norm.
[0014] To achieve the above objectives, another aspect of this application proposes a 3D point cloud completion system based on multi-scale structured knowledge distillation, the system comprising: The first module is used to construct an incomplete 3D point cloud dataset; The second module is used to introduce a multi-scale hierarchical knowledge self-distillation encoding module and a reconstruction decoding module to construct a three-dimensional point cloud data completion network model. The third module is used to perform 3D point cloud completion on the incomplete 3D point cloud dataset based on the 3D point cloud data completion network model to obtain a complete 3D point cloud dataset.
[0015] The embodiments of this application include at least the following beneficial effects: This application provides a 3D point cloud completion method and system based on multi-scale structured knowledge distillation. This scheme constructs an incomplete 3D point cloud dataset, further introduces a multi-scale hierarchical knowledge self-distillation encoding module and a reconstruction decoding module, constructs a 3D point cloud data completion network model, proposes a hierarchical structure, focuses on more effective encoding of point clouds, utilizes multi-scale hierarchical knowledge distillation, learns the differences between different local features by maximizing the mutual information between them, specifies the distribution of the deepest module as the supervision of previous modules, and the plug-and-play module can be integrated into most point cloud processing networks as an encoder, recovers point clouds from coarse to fine through the reconstruction decoding module, and finally performs 3D point cloud completion on the incomplete 3D point cloud dataset based on the 3D point cloud data completion network model, improving the completeness of data completion. Attached Figure Description
[0016] Figure 1 This is a flowchart of the 3D point cloud completion method based on multi-scale structured knowledge distillation provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the 3D point cloud completion system based on multi-scale structured knowledge distillation provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the three-dimensional point cloud data completion network model provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the multi-scale hierarchical knowledge self-distillation encoding module provided in the embodiments of this application; Figure 5 This is a schematic diagram of the reconstruction decoding module provided in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the completion effects of different models on different objects in the PCN dataset provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Before providing a detailed description of the embodiments of this application, some of the nouns and terms used in the embodiments of this application will be explained first. Point cloud, as a 3D data representation with precise geographic coordinates and rich semantic information, plays a crucial role in many fields. Point cloud applications are wide-ranging, including but not limited to virtual reality, autonomous driving systems, medical modeling, and building maintenance. In practical applications, point cloud data either needs to be processed in real time or stored in large storage devices for subsequent analysis. When the cost of data processing and storage exceeds the transmission cost, point cloud reconstruction technology becomes particularly important for economic efficiency in practical applications. Especially in scenarios with time constraints or spanning large distances, point cloud reconstruction becomes the most feasible method.
[0022] Some shortcomings exist in related technologies. For example, existing completion methods such as SnowflakeNet and SeedFormer focus more on the decoding process of missing parts, i.e. restoration and upsampling, while ignoring the function of the encoding process in enriching the understanding of the model's intrinsic semantic information.
[0023] In view of this, this application provides a 3D point cloud completion method based on multi-scale structured knowledge distillation. By incorporating a plug-and-play hierarchical self-distillation (HSD) framework, it aims to maximize mutual information from multiple coverage scales and transfer the most unique implicit features in the deepest hyperspace to other scales. The last layer classifier in the series at different scales is designated as the teacher network (deepest layer) to supervise the classifiers in the previous network layers, thereby transferring important knowledge to the hyperspace at different perceptual scales.
[0024] Reference Figure 1 , Figure 1 A flowchart of a 3D point cloud completion method based on multi-scale structured knowledge distillation provided in this embodiment of the invention is shown below. Figure 1 The method includes the following steps: S100, Construct an incomplete 3D point cloud dataset; In some specific embodiments, any 3D point cloud dataset can be used. This embodiment uses the common PCN completion dataset, a publicly available dataset derived from ShapeNet, which has already been paired with incomplete inputs and complete outputs. By default, each object uses 2048 points as input and 16384 points as output.
[0025] S200 introduces a multi-scale hierarchical knowledge self-distillation encoding module and a reconstruction decoding module to construct a three-dimensional point cloud data completion network model. It should be noted that in some embodiments, such as Figure 3 As shown, the 3D point cloud data completion network model includes a multi-scale hierarchical knowledge self-distillation encoding module, a multi-layer perceptron module, and a reconstruction decoding module. The multi-scale hierarchical knowledge self-distillation encoding module, the multi-layer perceptron module, and the reconstruction decoding module are connected sequentially. The multi-scale hierarchical knowledge self-distillation encoding module includes several self-distillation modules, each of which includes a k-nearest neighbor layer, a farthest point sampling layer, a multi-layer perceptron layer, and a max pooling layer. The reconstruction decoding module includes a PointNet layer, a local attention mechanism layer, a one-dimensional deconvolution layer, a first multi-layer perceptron layer, a second multi-layer perceptron layer, and an upsampling layer.
[0026] In some specific embodiments, the expression for the loss function of the 3D point cloud data completion network model is as follows: ; In the above formula, This represents the loss function for a 3D point cloud data completion network model. Represents the empirical coefficient. Represents the reconstruction loss function. This represents the self-distillation loss function. This represents the corresponding completion result of the target set. Representing a complete truth point cloud, This represents the predicted point cloud output by the reconstruction decoding module. This indicates the actual resolution value output by the reconstruction decoding module.
[0027] S300. Based on the three-dimensional point cloud data completion network model, perform three-dimensional point cloud completion on the incomplete three-dimensional point cloud dataset to obtain a complete three-dimensional point cloud dataset. It should be noted that in some embodiments, step S300 may include: S310, inputting the incomplete 3D point cloud dataset into the 3D point cloud data completion network model; S320. Based on the multi-scale hierarchical knowledge self-distillation coding module of the three-dimensional point cloud data completion network model, self-distillation loss coding is performed on the incomplete three-dimensional point cloud dataset to obtain sparse point cloud data and neighborhood features. Specifically, the incomplete 3D point cloud dataset is input into the multi-scale hierarchical knowledge self-distillation encoding module; based on the self-distillation module of the multi-scale hierarchical knowledge self-distillation encoding module, feature extraction is performed on the incomplete 3D point cloud dataset to obtain several preliminary neighborhood features; the preliminary neighborhood features are calculated using the softmax activation function to obtain the probability distributions corresponding to the preliminary neighborhood features; the probability distribution corresponding to the last preliminary neighborhood feature is used as a supervision signal, and knowledge backflow is achieved through KL divergence to output the sparse point cloud data and the neighborhood features.
[0028] More specifically, the incomplete 3D point cloud dataset is input into the self-distillation module of the multi-scale hierarchical knowledge self-distillation encoding module; based on the k-nearest neighbor layer of the self-distillation module, a k-nearest neighbor search is performed on the incomplete 3D point cloud dataset to construct local nearest neighbor features of the point cloud data; based on the farthest point sampling layer of the self-distillation module, center point sampling is performed on the local nearest neighbor features of the point cloud data to obtain the center points of the local nearest neighbor features; based on the multilayer perceptron layer of the self-distillation module, feature mapping is performed on the center points of the local nearest neighbor features to obtain the mapped local nearest neighbor features; based on the max pooling layer of the self-distillation module, the mapped local nearest neighbor features are aggregated to obtain several preliminary neighborhood features.
[0029] In some specific embodiments, such as Figure 4As shown, it should first be noted that traditionally, various aggregation operations are used to encode global or local shapes, such as max pooling and set abstraction. Therefore, these methods inevitably ignore specific low-order geometric features. This phenomenon can be further amplified, especially when the input is incomplete. To mitigate the feature performance degradation caused by aggregation, this invention utilizes multi-scale hierarchical knowledge distillation to learn the differences between different local features by maximizing their mutual information. Since they can be integrated into any hierarchical model, they can replace the encoder in various point cloud completion methods. Specifically, this invention applies self-distillation modules to SnowflakeNet and PMP-Net++ to replace their set abstraction modules. Specifically, with incomplete point clouds... As input, the encoder integrates three ensemble abstraction modules (SA, including k-nearest neighbor kNN, farthest point sampling FPS, MLP, and max pooling AgG) to cloud-compute the neighborhood features of the input points. The probability distributions y of the three SAs are obtained by applying softmax activation. The last SA is then subjected to self-distillation loss. The signal is passed back to the two students in front, SA ( and Although no real probability distribution is available, only the geometric prior is, its implicit encoding can still be used to represent a particular object and its virtual probability distribution. Therefore, we map the encodings from different aggregation operations to the probability space, using softmax activation to make the different distributions act as virtual objective functions. We specify the distribution of the deepest module as supervision over previous modules. Let... For KL divergence, the self-distillation loss is expressed as: ; In the above formula, This represents the loss function of the self-distillation module. Denotes KL divergence, This represents the predicted label distribution. Represents the true label distribution. Indicates the distribution of the last layer. .
[0030] in For predicted values, Note the predicted label distribution in the self-distillation module. Also related to the true value distribution Comparisons can be made. Knowledge can be fed back through this formula during the forward training steps to provide stronger supervision.
[0031] S330. A multilayer perceptron module based on the three-dimensional point cloud data completion network model performs data dimensionality reduction processing on the sparse point cloud data and the neighborhood features to obtain a coarse point cloud dataset. S340. The reconstruction and decoding module based on the three-dimensional point cloud data completion network model performs data reconstruction on the coarse point cloud dataset to obtain the complete three-dimensional point cloud dataset.
[0032] Specifically, the coarse point cloud dataset is input into the reconstruction decoding module of the 3D point cloud data completion network model; based on the PointNet layer of the reconstruction decoding module, point cloud registration processing is performed on the coarse point cloud dataset to obtain the current point cloud neighborhood features; based on the local attention mechanism layer of the reconstruction decoding module, local feature extraction processing is performed on the current point cloud neighborhood features to obtain local features; global features of the coarse point cloud dataset are obtained and combined with the local features to obtain point cloud association features; based on the one-dimensional deconvolution layer of the reconstruction decoding module, the point cloud association features are copied upwards to obtain copied point cloud association features; based on the first multilayer perceptron layer and the second multilayer perceptron layer of the reconstruction decoding module, feature mapping is performed on the copied point cloud association features to obtain the point-by-point displacement of the point cloud features; based on the upsampling layer of the reconstruction decoding module, the coarse point cloud dataset is upsampled to obtain an upsampled coarse point cloud dataset; the point-by-point displacement of the point cloud features is added to the upsampled coarse point cloud dataset to obtain the complete 3D point cloud dataset.
[0033] In some specific embodiments, such as Figure 5 As shown, based on the sparse points sampled by SA in the encoder and its neighborhood features A coarse point cloud can be obtained through dimensionality reduction mapping using an MLP (Multilayer Perceptron). This invention uses a decoder to... A complete reconstruction is performed. During the decoding process, this invention utilizes the max-pooled neighborhood features obtained from the encoder. This is considered as a global feature G. Subsequently, this invention integrates three split-based deconvolution (SPD) modules to recover the point cloud from coarse to fine. Each SPD inputs the current coarse point cloud from the previous step into PointNet to obtain the current point cloud's neighborhood features. By combining global features G and local attention, the correlation between neighborhood features and their surrounding points and global features G is obtained. And this association is replicated upwards using one-dimensional deconvolution. Finally, based on MLP, these features are mapped into point-by-point displacements. Simultaneously, based on the point cloud obtained in the previous SPD step (here, a rough point cloud input to the first SPD), For example, by simply upsampling (coordinate copying), the number of points is doubled, and finally, these upsampled points are added to the point-by-point displacement to obtain the complete point cloud output. ,in This is the final output of the SPD. This process is equivalent to rearranging the generated points based on the upsampled original coordinates, rather than simply shuffling the higher-order representations like PU-Net, or performing grid stitching like PCN. In general, the network aims to learn from incomplete inputs. Reconstructing the complete point cloud ,in and These represent the number of points in the ground truth point cloud and the incomplete input point cloud, respectively. Specifically, the reconstruction loss can be expressed as: ; In the above formula, This represents the loss function of the reconstruction decoding module. This represents a non-complete 3D point cloud dataset. Represents a complete 3D point cloud dataset. This represents point cloud data points in an incomplete 3D point cloud dataset. This represents the point cloud data points in a complete 3D point cloud dataset. This represents the L2 norm.
[0034] Finally, as Figure 6 As shown, the completion error of the model in this embodiment of the invention and other models for the PCN dataset is compared ( Through more efficient encoding (MSG and HSD), the completion accuracy has been improved across the board, as detailed in Table 1 below.
[0035] Table 1. Comparison of completion errors for each model
[0036] Please see Figure 2 This application also provides a 3D point cloud completion system based on multi-scale structured knowledge distillation, which can implement the above-mentioned 3D point cloud completion method based on multi-scale structured knowledge distillation. The system includes: The first module 201 is used to construct an incomplete 3D point cloud dataset; The second module 202 is used to introduce a multi-scale hierarchical knowledge self-distillation encoding module and a reconstruction decoding module to construct a three-dimensional point cloud data completion network model. The third module 203 is used to perform 3D point cloud completion on the incomplete 3D point cloud dataset based on the 3D point cloud data completion network model to obtain a complete 3D point cloud dataset.
[0037] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0038] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A 3D point cloud completion method based on multi-scale structured knowledge distillation, characterized in that, The method includes the following steps: Constructing an incomplete 3D point cloud dataset; A multi-scale hierarchical knowledge self-distillation encoding module and a reconstruction decoding module are introduced to construct a 3D point cloud data completion network model. Based on the aforementioned 3D point cloud data completion network model, the incomplete 3D point cloud dataset is completed to obtain a complete 3D point cloud dataset. The step of performing 3D point cloud completion on the incomplete 3D point cloud dataset based on the 3D point cloud data completion network model to obtain a complete 3D point cloud dataset includes: The incomplete 3D point cloud dataset is input into the 3D point cloud data completion network model; Based on the multi-scale hierarchical knowledge self-distillation coding module of the 3D point cloud data completion network model, the incomplete 3D point cloud dataset is subjected to self-distillation loss coding to obtain sparse point cloud data and neighborhood features. The multilayer perceptron module based on the three-dimensional point cloud data completion network model performs data dimensionality reduction processing on the sparse point cloud data and the neighborhood features to obtain a coarse point cloud dataset. The reconstruction and decoding module based on the 3D point cloud data completion network model reconstructs the coarse point cloud dataset to obtain the complete 3D point cloud dataset.
2. The method according to claim 1, characterized in that, The 3D point cloud data completion network model includes a multi-scale hierarchical knowledge self-distillation encoding module, a multilayer perceptron module, and a reconstruction decoding module. The multi-scale hierarchical knowledge self-distillation encoding module, the multilayer perceptron module, and the reconstruction decoding module are connected sequentially. The specific expression of the loss function of the 3D point cloud data completion network model is as follows: ; In the above formula, This represents the loss function for a 3D point cloud data completion network model. Represents the empirical coefficient. Represents the reconstruction loss function. This represents the self-distillation loss function. This represents the corresponding completion result for the target set. Representing a complete truth point cloud, This represents the predicted point cloud output by the reconstruction decoding module. This indicates the actual resolution value output by the reconstruction decoding module.
3. The method according to claim 1, characterized in that, The multi-scale hierarchical knowledge self-distillation encoding module includes several self-distillation modules, each of which includes a k-nearest neighbor layer, a farthest point sampling layer, a multilayer perceptron layer, and a max pooling layer. The reconstruction decoding module includes a PointNet layer, a local attention mechanism layer, a one-dimensional deconvolution layer, a first multilayer perceptron layer, a second multilayer perceptron layer, and an upsampling layer.
4. The method according to claim 1, characterized in that, The multi-scale hierarchical knowledge self-distillation encoding module based on the 3D point cloud data completion network model performs self-distillation loss encoding on the incomplete 3D point cloud dataset to obtain sparse point cloud data and neighborhood features, including: The incomplete 3D point cloud dataset is input into the multi-scale hierarchical knowledge self-distillation encoding module; Based on the self-distillation module of the multi-scale hierarchical knowledge self-distillation coding module, the self-distillation module extracts features from the incomplete 3D point cloud dataset to obtain several preliminary neighborhood features. The preliminary neighborhood features are calculated using the softmax activation function to obtain the probability distributions corresponding to several preliminary neighborhood features; Using the probability distribution corresponding to the last preliminary neighborhood feature as a supervision signal, knowledge feedback is achieved through KL divergence, and the sparse point cloud data and the neighborhood feature are output.
5. The method according to claim 4, characterized in that, The self-distillation module based on the multi-scale hierarchical knowledge self-distillation encoding module extracts features from the incomplete 3D point cloud dataset to obtain several preliminary neighborhood features, including: The incomplete 3D point cloud dataset is input into the self-distillation module of the multi-scale hierarchical knowledge self-distillation encoding module; Based on the k-nearest neighbor layer of the self-distillation module, a k-nearest neighbor search is performed on the incomplete 3D point cloud dataset to construct local nearest neighbor features of the point cloud data. Based on the farthest point sampling layer of the self-distillation module, the center point of the local nearest neighbor feature of the point cloud data is sampled to obtain the center point of the local nearest neighbor feature. Based on the multilayer perceptron layer of the self-distillation module, feature mapping is performed on the center point of the local nearest neighbor feature to obtain the mapped local nearest neighbor feature; Based on the max pooling layer of the self-distillation module, the mapped local nearest neighbor features are aggregated to obtain several preliminary neighborhood features.
6. The method according to claim 3, characterized in that, The loss function of the self-distillation module is as follows: ; In the above formula, This represents the loss function of the self-distillation module. Denotes KL divergence, This represents the predicted label distribution. Represents the true label distribution. Indicates the distribution of the last layer. .
7. The method according to claim 1, characterized in that, The reconstruction and decoding module based on the 3D point cloud data completion network model reconstructs the coarse point cloud dataset to obtain the complete 3D point cloud dataset, including: The coarse point cloud dataset is input into the reconstruction and decoding module of the 3D point cloud data completion network model; Based on the PointNet layer of the reconstruction decoding module, point cloud registration processing is performed on the coarse point cloud dataset to obtain the current point cloud neighborhood features; Based on the local attention mechanism layer of the reconstruction decoding module, local feature extraction processing is performed on the current point cloud neighborhood features to obtain local features; Obtain the global features of the coarse point cloud dataset and combine them with the local features to obtain the point cloud association features; Based on the one-dimensional deconvolution layer of the reconstruction decoding module, the point cloud association features are copied upwards to obtain the copied point cloud association features. Based on the first and second multilayer perceptron layers of the reconstruction decoding module, feature mapping is performed on the copied point cloud associated features to obtain the point-by-point displacement of the point cloud features. Based on the upsampling layer of the reconstruction decoding module, the coarse point cloud dataset is upsampled to obtain the upsampled coarse point cloud dataset. The point-by-point displacement of the point cloud features is added to the upsampled coarse point cloud dataset to obtain the complete 3D point cloud dataset.
8. The method according to claim 7, characterized in that, The loss function of the reconstruction decoding module is as follows: ; In the above formula, This represents the loss function of the reconstruction decoding module. This represents a non-complete 3D point cloud dataset. Represents a complete 3D point cloud dataset. This represents point cloud data points in an incomplete 3D point cloud dataset. This represents the point cloud data points in a complete 3D point cloud dataset. This represents the L2 norm.
9. A 3D point cloud completion system based on multi-scale structured knowledge distillation, characterized in that, The system includes: The first module is used to construct an incomplete 3D point cloud dataset; The second module is used to introduce a multi-scale hierarchical knowledge self-distillation encoding module and a reconstruction decoding module to construct a three-dimensional point cloud data completion network model. The third module is used to perform 3D point cloud completion on the incomplete 3D point cloud dataset based on the 3D point cloud data completion network model to obtain a complete 3D point cloud dataset. The step of performing 3D point cloud completion on the incomplete 3D point cloud dataset based on the 3D point cloud data completion network model to obtain a complete 3D point cloud dataset includes: The incomplete 3D point cloud dataset is input into the 3D point cloud data completion network model; Based on the multi-scale hierarchical knowledge self-distillation coding module of the 3D point cloud data completion network model, the incomplete 3D point cloud dataset is subjected to self-distillation loss coding to obtain sparse point cloud data and neighborhood features. The multilayer perceptron module based on the three-dimensional point cloud data completion network model performs data dimensionality reduction processing on the sparse point cloud data and the neighborhood features to obtain a coarse point cloud dataset. The reconstruction and decoding module based on the 3D point cloud data completion network model reconstructs the coarse point cloud dataset to obtain the complete 3D point cloud dataset.
Citation Information
Patent Citations
Point cloud completion method based on teacher-student network and course learning
CN117710255A
Self-supervised three-dimensional point cloud completion method for underwater target object
CN118470515A
Sparse point cloud understanding method based on joint learning and hierarchical self-distillation
CN119832538A
Cross-modal distillation point cloud up-sampling method guided by multi-view depth map
CN120071074A
3D anomaly detection method and device based on self-distillation
CN120411951A