Point cloud up-sampling method based on lightweight neural network
Through the lightweight neural network framework and the point cloud upsampling method optimized by knowledge distillation loss function, the problem of unbalanced performance and computing efficiency in the existing technology is solved, efficient point cloud upsampling performance and low computing resource requirements are achieved, and the processing capability of point cloud data is improved.
Patent Information
- Application Number
- CN202510274179.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-08
AI Technical Summary
The existing point cloud upsampling methods are difficult to balance between performance and computing efficiency, especially in large data scenarios with high computing requirements, and the existing methods fail to effectively utilize the sparseness of point clouds, resulting in increased computing complexity.
Using a lightweight neural network framework, through multiple rounds of training of teacher networks and student networks, combined with global context holders, globally perceived Transformer upsamplers, local geometry holders and locally perceived Mamba enhancers, a double-aligned knowledge distillation loss function is designed to optimize network parameters to achieve efficient point cloud upsampling.
It realizes high upsampling performance on PU-GAN datasets, reduces computing resource requirements, network parameters and calculation floating-point numbers, takes into account point cloud upsampling performance and computing costs, and improves the understanding ability of point cloud data.
Smart Images

Figure CN120279353A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a point cloud upsampling technology, in particular to a point cloud upsampling method based on a lightweight neural network. Background Art
[0002] With the rapid development of lidar sensor technology, point cloud data plays an important role in representing the structure of complex three-dimensional objects. The compact storage form of point cloud data makes it favored in many applications, and these applications highly rely on the quality of the point cloud. Different from image data regularly arranged on a pixel grid, point cloud data usually contains disordered, sparse, and noisy data points due to the limitations of acquisition devices, making it difficult to directly apply traditional image processing techniques to point cloud processing. The point cloud upsampling technology emerged to improve the quality of point clouds by increasing point cloud density, reducing noise, and enhancing geometric fidelity.
[0003] Currently, point cloud upsampling methods are mainly divided into two categories: optimization-based point cloud upsampling methods and deep learning-based point cloud upsampling methods. Optimization-based point cloud upsampling methods utilize shape prior information such as normal vectors, density, and curvature to guide the generation process from sparse point clouds to dense point clouds. Deep learning-based point cloud upsampling methods, on the other hand, learn the mapping relationship between sparse point sets and dense point sets through a large amount of data.
[0004] However, the existing point cloud upsampling technologies still face the following challenges:
[0005] First, although existing methods can maintain the overall geometric shape of the original point cloud when generating dense point clouds, it is difficult to achieve a balance between upsampling performance and computational efficiency. This limits the performance of point cloud upsampling methods in establishing global dependencies and local geometric structures, especially in scenarios such as autonomous driving where a large amount of point cloud data is required. This defect is particularly prominent in such scenarios. Existing work is mostly limited to relatively low upsampling quality, lacking effective exploration of the balance between network parameters and performance, and it is difficult to meet the actual application requirements.
[0006] Second, some existing methods have good performance, but due to the excessive number of parameters, the computational requirements increase rapidly as the input data volume increases. To avoid clustering effects and maintain a uniform distribution of points, a large number of parameters need to be optimized, which further increases the computational burden. Some methods attempt to convert point clouds into regular 3D voxel grids, but this method has extremely high computational complexity and memory costs, and fails to fully utilize the sparsity of point clouds. Although the attention mechanism of the Transformer network is suitable for processing disordered point cloud data, the complexity of the attention mechanism grows quadratically with the increase in input data volume. Although the global receptive field of the attention mechanism can provide better global information aggregation ability, it also significantly increases the computational complexity. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a point cloud upsampling method based on a lightweight neural network, which can effectively improve the performance of point cloud upsampling and save computing resources.
[0008] The technical solution adopted by the present invention to solve the above technical problem is: a point cloud upsampling method based on a lightweight neural network, characterized in that: First, construct a training set containing pairs of sparse point cloud blocks and ground-truth dense point cloud blocks, and build a lightweight neural network composed of a teacher network and a student network; Second, input the pairs of sparse point cloud blocks and ground-truth dense point cloud blocks in the training set into the lightweight neural network for multiple rounds of network training, and after the network training is completed, obtain the student network training model; Third, use the student network training model to upsample each sparse point cloud block in the test set, and after upsampling, obtain the student dense point cloud block corresponding to each sparse point cloud block in the test set; where:
[0009] The teacher network is composed of a global context keeper and a globally-aware Transformer upsampler, and the student network is composed of a local geometry keeper and a locally-aware Mamba enhancer; First, input the sparse point cloud block into the global context keeper and the local geometry keeper at the same time, the global context keeper outputs the teacher feature map, and the local geometry keeper outputs the student feature map; Second, input the sparse point cloud block and the teacher feature Figure 1 map into the globally-aware Transformer upsampler to obtain the teacher dense point cloud block; and input the sparse point cloud block and the student feature Figure 1 map into the locally-aware Mamba enhancer to obtain the student dense point cloud block;
[0010] When training the lightweight neural network, first train the teacher network, and then train the student network after freezing the parameters of the teacher network; when training the teacher network, calculate the teacher network loss before the end of each round of network training to optimize the parameters of the teacher network to constrain the network training, stop the training of the teacher network after multiple rounds of network training, and freeze the parameters of the teacher network; when training the student network, calculate the student network loss before the end of each round of network training to optimize the parameters of the student network to constrain the network training, stop the training of the student network after multiple rounds of network training, and obtain the student network training model, where the student network loss includes a dual-alignment knowledge distillation loss for aligning features and points.
[0011] The global context retainer includes a first multi-layer perceptron layer, a first max pooling layer, a first broadcasting layer, a first distribution feature calculation layer, a second broadcasting layer, a first convolutional layer, a first batch normalization layer, a first non-linear activation layer, a second max pooling layer, a second distribution feature calculation layer, a third broadcasting layer, a second convolutional layer, a second batch normalization layer, a second non-linear activation layer, a third distribution feature calculation layer, and a first affine transformation layer; the implementation process of the global context retainer is as follows: simultaneously input the sparse point cloud block into the first multi-layer perceptron layer, the second distribution feature calculation layer, and the second max pooling layer, pass the feature map output by the first multi-layer perceptron layer through the first max pooling layer and the first broadcasting layer in sequence, simultaneously pass the feature map output by the first multi-layer perceptron layer through the first distribution feature calculation layer and the second broadcasting layer, perform a channel connection operation on the feature map output by the first broadcasting layer and the feature map output by the second broadcasting layer to obtain a first channel connection feature map, the first channel connection feature map passes through the first convolutional layer, the first batch normalization layer, and the first non-linear activation layer in sequence, perform a channel connection operation on the sparse point cloud block, the feature map obtained by passing the feature map output by the second distribution feature calculation layer through the third broadcasting layer, and the feature map output by the second max pooling layer to obtain a second channel connection feature map, the second channel connection feature map passes through the second convolutional layer, the second batch normalization layer, and the second non-linear activation layer in sequence, perform a channel connection operation on the feature map output by the first non-linear activation layer and the feature map output by the second non-linear activation layer to obtain a third channel connection feature map, the third channel connection feature map passes through the third distribution feature calculation layer and the first affine transformation layer in sequence, and perform a residual connection operation on the feature map output by the first multi-layer perceptron layer and the feature map output by the first affine transformation layer, and use the obtained feature map as the teacher feature map output by the global context retainer.
[0012] The global perception Transformer upsampler includes a first nearest neighbor interpolation layer, a second multi-layer perceptron layer, a third max pooling layer, a fourth broadcast layer, a third multi-layer perceptron layer, a PT layer, a first deconvolution layer, a second nearest neighbor interpolation layer, a third non-linear activation layer, a fourth multi-layer perceptron layer, and a fourth non-linear activation layer; the implementation process of the global perception Transformer upsampler is as follows: The sparse point cloud block is simultaneously input into the first nearest neighbor interpolation layer and the second multi-layer perceptron layer. The feature map output by the second multi-layer perceptron layer and the teacher feature map output by the global context keeper are subjected to a residual connection operation, and the feature map obtained by the residual connection is channel-connected with the feature map obtained after passing through the third max pooling layer and the fourth broadcast layer in sequence to obtain a fourth channel-connected feature map. The fourth channel-connected feature map passes through the third multi-layer perceptron layer and the PT layer in sequence. The feature map output by the PT layer is simultaneously input into the first deconvolution layer and the second nearest neighbor interpolation layer. The feature map output by the first deconvolution layer and the feature map output by the second nearest neighbor interpolation layer are channel-connected to obtain a fifth channel-connected feature map. The fifth channel-connected feature map passes through the third non-linear activation layer, the fourth multi-layer perceptron layer, and the fourth non-linear activation layer in sequence. The feature map obtained by performing a residual connection operation on the feature map output by the first nearest neighbor interpolation layer and the feature map output by the fourth non-linear activation layer is used as the teacher dense point cloud block output by the global perception Transformer upsampler.
[0013] The local geometry keeper includes a fifth multi-layer perceptron layer, a first K-nearest neighbor grouping layer, a fourth distribution feature calculation layer, a third convolution layer, a third batch normalization layer, a fifth non-linear activation layer, a fifth distribution feature calculation layer, a fourth convolution layer, a fourth batch normalization layer, a sixth non-linear activation layer, a sixth distribution feature calculation layer, and a second affine transformation layer; the implementation process of the local geometry keeper is as follows: The sparse point cloud block is input into the fifth multi-layer perceptron layer. The feature map output by the fifth multi-layer perceptron layer and the sparse point cloud block are together input into the first K-nearest neighbor grouping layer. One of the feature maps output by the first K-nearest neighbor grouping layer passes through the fourth distribution feature calculation layer, the third convolution layer, the third batch normalization layer, and the fifth non-linear activation layer in sequence. The other feature map output by the first K-nearest neighbor grouping layer passes through the fifth distribution feature calculation layer, the fourth convolution layer, the fourth batch normalization layer, and the sixth non-linear activation layer in sequence. The feature map output by the fifth non-linear activation layer and the feature map output by the sixth non-linear activation layer are channel-connected to obtain a sixth channel-connected feature map. The sixth channel-connected feature map passes through the sixth distribution feature calculation layer and the second affine transformation layer in sequence. The feature map obtained by performing a residual connection operation on the feature map output by the fifth multi-layer perceptron layer and the feature map output by the second affine transformation layer is used as the student feature map output by the local geometry keeper.
[0014] The locally perceptive Mamba enhancer includes a second K-nearest neighbor grouping layer, a sixth multi-layer perceptron layer, a fourth max pooling layer, a serialization layer, a Mamba layer, a second deconvolution layer, a third nearest neighbor interpolation layer, a seventh non-linear activation layer, a seventh multi-layer perceptron layer, an eighth non-linear activation layer, and a fourth nearest neighbor interpolation layer; the implementation process of the locally perceptive Mamba enhancer is as follows: the sparse point cloud block and the student features output by the local geometry retainer Figure 1 are input into the second K-nearest neighbor grouping layer, and the channel connection operation is performed on the two feature maps output by the second K-nearest neighbor grouping layer to obtain the seventh channel-connected feature map. The seventh channel-connected feature map passes through the sixth multi-layer perceptron layer, the fourth max pooling layer, the serialization layer, and the Mamba layer in sequence. The feature map output by the Mamba layer is simultaneously input into the second deconvolution layer and the third nearest neighbor interpolation layer. The channel connection operation is performed on the feature map output by the second deconvolution layer and the feature map output by the third nearest neighbor interpolation layer to obtain the eighth channel-connected feature map. The eighth channel-connected feature map passes through the seventh non-linear activation layer, the seventh multi-layer perceptron layer, and the eighth non-linear activation layer. The feature map obtained by performing the residual connection operation on the feature map obtained by passing the sparse point cloud block through the fourth nearest neighbor interpolation layer and the feature map output by the eighth non-linear activation layer is used as the student dense point cloud block output by the locally perceptive Mamba enhancer.
[0015] The construction process of the training set is as follows: First, multiple original mesh models are selected as the source of point cloud data; then, Poisson disk sampling operations are performed on each original mesh model to generate sparse point clouds and dense point clouds respectively. Among them, the number of points in the sparse point cloud is significantly less than the number of points in the dense point cloud; then, the farthest point sampling algorithm is used for the sparse point cloud to find multiple seed points; then, for each seed point, the nearest neighbor algorithm is used to extract sparse point cloud blocks from the sparse point cloud and extract dense point cloud blocks from the dense point cloud. The dense point cloud blocks are used as the true dense point cloud blocks, and the sparse point cloud blocks and the true dense point cloud blocks are paired to form a training pair. Among them, the number of points in the sparse point cloud block is significantly less than the number of points in the dense point cloud block; the number of training pairs generated by each original mesh model is the same as the number of seed points, and the final number of generated training pairs is equal to the product of the number of original mesh models and the number of seed points.
[0016] Denote the teacher network loss as L T , Denote the student network loss as L S ,L S =L CD +L DAD , L DAD =L KL (FG1,FG4)+L MSE (FG1,FG4)+L KL (FG3,FG5)+L MSE(FG3, FG5), where L T is obtained by calculating the chamfer distance loss between the ground-truth dense point cloud patch and the teacher dense point cloud patch output by the globally-aware Transformer upsampler. FG3 represents the teacher dense point cloud patch output by the globally-aware Transformer upsampler. represents the number of points in FG3, p O represents a point in FG3, min(·) is the function to take the minimum value, P HR represents the ground-truth dense point cloud patch, p T represents a point in P HR in represents the number of points in P HR ||·||2 is the symbol for the two-norm operation, L CD represents the chamfer distance loss between the ground-truth dense point cloud patch and the student dense point cloud patch output by the locally-aware Mamba enhancer, L DAD represents the dual-alignment knowledge distillation loss, FG5 represents the student dense point cloud patch output by the locally-aware Mamba enhancer. represents the number of points in FG5, represents a point in FG5, L KL L(·) represents the KL divergence loss, L MSE L(·) represents the mean squared error loss, FG1 represents the teacher feature map output by the global context retainer, FG4 represents the student feature map output by the local geometry retainer.
[0017] The process of obtaining the test set is as follows: First, randomly select M test sparse point clouds; then, for each test sparse point cloud, use the farthest point sampling algorithm to find r×S / (3×N) seed points; then, for each seed point, use the nearest neighbor algorithm to extract a sparse point cloud patch containing N points from the test sparse point cloud. For one test sparse point cloud, a total of r×S / (3×N) sparse point cloud patches are extracted; finally, the M×r×S / (3×N) sparse point cloud patches form the test set; where M≥1, r represents the upsampling rate, S represents the number of points in the sparse point cloud, and 3×N represents the size of the sparse point cloud patch.
[0018] After upsampling to obtain the student dense point cloud patch corresponding to each sparse point cloud patch in the test set, the corresponding student dense point cloud patches obtained by upsampling the r×S / (3×N) sparse point cloud patches from the same test sparse point cloud are spliced together, and then the farthest point sampling algorithm is used to sample r×S points, finally completing the r-fold upsampling of the test sparse point cloud.
[0019] Compared with the prior art, the advantages of the present invention are:
[0020] 1) The method of the present invention proposes a point cloud upsampling framework composed of lightweight neural networks, aiming to achieve the best trade-off between performance and computational cost and ensure meeting the requirements of practical applications. The point cloud upsampling combines lightweight neural networks and knowledge distillation, which can effectively improve the performance of point cloud upsampling and save computational resources, as reflected by achieving high performance on the test set of PU-GAN: Chamfer distance (0.243), Hausdorff distance (1.521), Earth Mover's distance (2.306), average distance from point to plane (1.675), and keeping the number of network parameters of the lightweight neural network within the lightweight range: number of network parameters (0.832MB), floating-point operations (19.808GB).
[0021] 2) The point cloud upsampling framework aligns the student feature map and the teacher feature map, as well as the student dense point cloud block and the teacher dense point cloud block, by designing a dual-aligned knowledge distillation loss function, considering the feature similarity of features and three-dimensional coordinates, point similarity, feature distribution similarity, and point distribution similarity, thereby constraining the learning process of the network for point cloud data. In addition, a global context keeper and a globally-aware Transformer upsampler are designed to learn the global geometric information of the point cloud, and a local geometry keeper and a locally-aware Mamba enhancer are designed to learn the local detail information of the point cloud. These two designs enable the point cloud upsampling framework to not only have the global modeling ability of Transformer to learn geometric contour information but also be able to take into account the local modeling ability to learn edge detail information, thus enhancing the understanding ability of the point cloud upsampling framework for point cloud data. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the overall implementation block diagram of the method of the present invention;
[0023] Figure 2 is the schematic diagram of the composition structure of the global context keeper in the lightweight neural network built by the method of the present invention;
[0024] Figure 3 is the schematic diagram of the composition structure of the globally-aware Transformer upsampler in the lightweight neural network built by the method of the present invention;
[0025] Figure 4 is the schematic diagram of the composition structure of the local geometry keeper in the lightweight neural network built by the method of the present invention;
[0026] Figure 5 is the schematic diagram of the composition structure of the locally-aware Mamba enhancer in the lightweight neural network built by the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] The present invention will be further described in detail below with reference to the embodiments of the drawings.
[0028] A point cloud upsampling method based on a lightweight neural network proposed by the present invention, the overall implementation block diagram of which is as Figure 1 shown. First, a training set containing pairs of sparse point cloud blocks and ground truth dense point cloud blocks is constructed, and a lightweight neural network composed of a teacher network and a student network is built; second, the pairs of sparse point cloud blocks and ground truth dense point cloud blocks in the training set are input into the lightweight neural network for multiple rounds of network training. After the network training is completed, a student network training model is obtained; third, the student network training model is used to upsample each sparse point cloud block in the test set, and the student dense point cloud block corresponding to each sparse point cloud block in the test set is obtained after upsampling; where:
[0029] The teacher network is composed of a global context maintainer and a globally-aware Transformer upsampler, and the student network is composed of a local geometry maintainer and a locally-aware Mamba enhancer; first, the sparse point cloud block is input into the global context maintainer and the local geometry maintainer at the same time. The global context maintainer outputs a teacher feature map, and the local geometry maintainer outputs a student feature map; second, the sparse point cloud block and the teacher feature Figure 1 map are input into the globally-aware Transformer upsampler to obtain a teacher dense point cloud block; and the sparse point cloud block and the student feature Figure 1 map are input into the locally-aware Mamba enhancer to obtain a student dense point cloud block.
[0030] When training the lightweight neural network, the teacher network is trained first, and after freezing the parameters of the teacher network, the student network is trained; when training the teacher network, the teacher network loss is calculated before the end of each round of network training to optimize the parameters of the teacher network to constrain the network training. After multiple rounds (such as 100 rounds) of network training are completed, the training of the teacher network is stopped, and the parameters of the teacher network are frozen; when training the student network, the student network loss is calculated before the end of each round of network training to optimize the parameters of the student network to constrain the network training. After multiple rounds (such as 100 rounds) of network training are completed, the training of the student network is stopped, and a student network training model is obtained. Among them, the student network loss includes a dual-alignment knowledge distillation loss used to align features and points.
[0031] In a specific embodiment, the teacher network loss is denoted as L T , the student network loss is denoted as L S , L S = L CD + L DAD , L DAD = L KL (FG1, FG4) + L MSE(FG1,FG4)+L KL (FG3,FG5)+L MSE (FG3,FG5), where L T Obtained by calculating the chamfer distance loss between the ground-truth dense point cloud block and the teacher dense point cloud block output by the globally-aware Transformer upsampler. FG3 represents the teacher dense point cloud block output by the globally-aware Transformer upsampler, represents the number of points in FG3, p O represents a point in FG3, min(·) is the minimum function, P HR represents the ground-truth dense point cloud block, p T represents P HR a point in represents the number of points in P HR ||·||2 is the symbol for the two-norm operation, L CD represents the chamfer distance loss between the ground-truth dense point cloud block and the student dense point cloud block output by the locally-aware Mamba enhancer, L DAD represents the double-alignment knowledge distillation loss. FG5 represents the student dense point cloud block output by the locally-aware Mamba enhancer, represents the number of points in FG5, represents a point in FG5, L KL L(·) represents the KL divergence loss, L MSE L(·) represents the mean squared error loss. FG1 represents the teacher feature map output by the global context retainer, and FG4 represents the student feature map output by the local geometry retainer.
[0032] Further defined as, e.g., Figure 2As shown, the global context retainer includes a first multi-layer perceptron layer, a first max pooling layer, a first broadcast layer, a first distribution feature calculation layer, a second broadcast layer, a first convolutional layer, a first batch normalization layer, a first non-linear activation layer, a second max pooling layer, a second distribution feature calculation layer, a third broadcast layer, a second convolutional layer, a second batch normalization layer, a second non-linear activation layer, a third distribution feature calculation layer, and a first affine transformation layer; the implementation process of the global context retainer is as follows: the sparse point cloud block is simultaneously input into the first multi-layer perceptron layer, the second distribution feature calculation layer, and the second max pooling layer. The feature map output by the first multi-layer perceptron layer is sequentially passed through the first max pooling layer and the first broadcast layer. At the same time, the feature map output by the first multi-layer perceptron layer is sequentially passed through the first distribution feature calculation layer and the second broadcast layer. The feature map output by the first broadcast layer and the feature map output by the second broadcast layer are subjected to a channel connection operation (Concatenation) to obtain a first channel connection feature map. The first channel connection feature map is sequentially passed through the first convolutional layer, the first batch normalization layer, and the first non-linear activation layer. The sparse point cloud block, the feature map obtained by passing the feature map output by the second distribution feature calculation layer through the third broadcast layer, and the feature map output by the second max pooling layer are subjected to a channel connection operation (Concatenation) to obtain a second channel connection feature map. The second channel connection feature map is sequentially passed through the second convolutional layer, the second batch normalization layer, and the second non-linear activation layer. The feature map output by the first non-linear activation layer and the feature map output by the second non-linear activation layer are subjected to a channel connection operation (Concatenation) to obtain a third channel connection feature map. The third channel connection feature map is sequentially passed through the third distribution feature calculation layer and the first affine transformation layer. The feature map output by the first multi-layer perceptron layer and the feature map output by the first affine transformation layer are subjected to a residual connection operation (Residual Connection), and the obtained feature map is used as the teacher feature map output by the global context retainer.
[0033] In this embodiment, the size of the sparse point cloud block is 3×N, where N is taken as 256. The input size of the first multi-layer perceptron layer is 3×N, and the output size is 128×N. The input size of the first max pooling layer is 128×N, and the output size is 128×1. The input size of the first broadcast layer is 128×1, and the output size is 128×N. The input size of the first distribution feature calculation layer is 128×N, and the output size is 128×1. The input size of the second broadcast layer is 128×1, and the output size is 128×N. The input size of the first convolutional layer is 256×N, and the output size is 64×N. The input size of the first batch normalization layer is 64×N, and the output size is 64×N. The input size of the first non-linear activation layer is 64×N, and the output size is 64×N. The input size of the second max pooling layer is 3×N, and the output size is 1×N. The input size of the second distribution feature calculation layer is 3×N, and the output size is 3×1. The input size of the third broadcast layer is 3×1, and the output size is 3×N. The input size of the second convolutional layer is 7×N, and the output size is 64×N. The input size of the second batch normalization layer is 64×N, and the output size is 64×N. The input size of the second non-linear activation layer is 64×N, and the output size is 64×N. The input size of the third distribution feature calculation layer is 128×N, and the output size is 128×1. The input size of the first affine transformation layer is 128×1, and the output size is 128×N. See Figure 2, specifically: input the sparse point cloud block into the first multi-layer perceptron layer, and denote the output feature map as FM1; input FM1 into the first max pooling layer, and denote the output feature map as FP1; input FP1 into the first broadcast layer, and denote the output feature map as FB1; input FM1 into the first distribution feature calculation layer, and denote the output feature map as FD1; input FD1 into the second broadcast layer, and denote the output feature map as FB2; perform a channel connection operation on FB1 and FB2, and denote the obtained first channel connection feature map as FC1; input FC1 into the first convolutional layer, and denote the output feature map as FV1; input FV1 into the first batch normalization layer, and denote the output feature map as FN1; input FN1 into the first non-linear activation layer, and denote the output feature map as FR1; input the sparse point cloud block into the second max pooling layer, and denote the output feature map as FP2; input the sparse point cloud block into the second distribution feature calculation layer, and denote the output feature map as FD2; input FD2 into the third broadcast layer, and denote the output feature map as FB3; perform a channel connection operation on the sparse point cloud block, FP2 and FB3, and denote the obtained second channel connection feature map as FC2; input FC2 into the second convolutional layer, and denote the output feature map as FV2; input FV2 into the second batch normalization layer, and denote the output feature map as FN2; input FN2 into the second non-linear activation layer, and denote the output feature map as FR2; perform a channel connection operation on FR1 and FR2, and denote the obtained third channel connection feature map as FC3; input FC3 into the third distribution feature calculation layer, and denote the output feature map as FD3; input FD3 into the first affine transformation layer, and denote the output feature map as FA1; perform a residual connection operation on FA1 and FM1, and denote the obtained feature map as FG1, and FG1 is used as the teacher feature map output by the global context keeper; where the size of FM1 is 128×N, the sizes of FP1 and FP2 correspond to 128×1 and 1×N respectively, the sizes of FB1, FB2 and FB3 correspond to 128×N, 128×N and 3×N respectively, the sizes of FD1, FD2 and FD3 correspond to 128×1, 3×1 and 128×1 respectively, the sizes of FC1, FC2 and FC3 correspond to 256×N, 7×N and 128×N respectively, the sizes of FV1, FV2, FN1, FN2, FR1 and FR2 are all 64×N, and the sizes of FA1 and FG1 are all 128×N.
[0034] Further limited, such as Figure 3As shown in the figure, the global perception Transformer upsampler includes a first nearest neighbor interpolation layer, a second multi-layer perceptron layer, a third max pooling layer, a fourth broadcast layer, a third multi-layer perceptron layer, a PT (PointTransformer network) layer, a first deconvolution layer, a second nearest neighbor interpolation layer, a third non-linear activation layer, a fourth multi-layer perceptron layer, and a fourth non-linear activation layer; the implementation process of the global perception Transformer upsampler is as follows: the sparse point cloud block is input into the first nearest neighbor interpolation layer and the second multi-layer perceptron layer at the same time, the feature map output by the second multi-layer perceptron layer and the teacher feature map output by the global context keeper are subjected to a residual connection operation (ResidualConnection), and the feature map obtained by the residual connection is channel-connected (Concatenation) with the feature map obtained after passing through the third max pooling layer and the fourth broadcast layer in sequence to obtain a fourth channel-connected feature map. The fourth channel-connected feature map passes through the third multi-layer perceptron layer and the PT layer in sequence, and the feature map output by the PT layer is input into the first deconvolution layer and the second nearest neighbor interpolation layer at the same time. The feature map output by the first deconvolution layer and the feature map output by the second nearest neighbor interpolation layer are channel-connected (Concatenation) to obtain a fifth channel-connected feature map. The fifth channel-connected feature map passes through the third non-linear activation layer, the fourth multi-layer perceptron layer, and the fourth non-linear activation layer in sequence. The feature map obtained by performing a residual connection operation (ResidualConnection) on the feature map output by the first nearest neighbor interpolation layer and the feature map output by the fourth non-linear activation layer is used as the teacher dense point cloud block output by the global perception Transformer upsampler.
[0035] In this embodiment, the input size of the first nearest neighbor interpolation layer is 3×N, the output size is 3×4N, the input size of the second multi-layer perceptron layer is 3×N, the output size is 128×N, the input size of the third max pooling layer is 128×N, the output size is 128×1, the input size of the fourth broadcast layer is 128×1, the output size is 128×N, the input size of the third multi-layer perceptron layer is 256×N, the output size is 128×N, the input size of the PT layer is 128×N, the output size is 128×N, the input size of the first deconvolution layer is 128×N, the output size is 128×4N, the input size of the second nearest neighbor interpolation layer is 128×N, the output size is 128×4N, the input size of the third non-linear activation layer is 256×4N, the output size is 256×4N, the input size of the fourth multi-layer perceptron layer is 256×4N, the output size is 3×4N, and the input size of the fourth non-linear activation layer is 3×4N, the output size is 3×4N. See Figure 3, specifically: input the sparse point cloud block into the first nearest neighbor interpolation layer, and denote the output feature map as FK1; input the sparse point cloud block into the second multi-layer perceptron layer, and denote the output feature map as FM2; perform a residual connection operation on FM2 and the teacher feature map output by the global context retainer, and denote the obtained feature map as FG2; input FG2 into the third max pooling layer, and denote the output feature map as FP3; input FP3 into the fourth broadcast layer, and denote the output feature map as FB4; perform a channel connection operation on FG2 and FB4, and denote the obtained fourth channel connection feature map as FC4; input FC4 into the third multi-layer perceptron layer, and denote the output feature map as FM3; input FM3 into the PT layer, and denote the output feature map as FPT; input FPT into the first deconvolution layer, and denote the output feature map as FDC1; input FPT into the second nearest neighbor interpolation layer, and denote the output feature map as FK2; perform a channel connection operation on FDC1 and FK2, and denote the obtained fifth channel connection feature map as FC5; input FC5 into the third non-linear activation layer, and denote the output feature map as FR3; input FR3 into the fourth multi-layer perceptron layer, and denote the output feature map as FM4; input FM4 into the fourth non-linear activation layer, and denote the output feature map as FR4; perform a residual connection operation on FR4 and FK1, and denote the obtained feature map as FG3, and FG3 is the teacher dense point cloud block output by the global perception Transformer upsampler; among them, the sizes of FM2, FM3, and FM4 are 128×N, 128×N, and 3×4N respectively, the size of FPT is 128×N, the sizes of FG2 and FG3 are 128×N and 3×4N respectively, the size of FP3 is 128×1, the size of FB4 is 128×N, the sizes of FC4 and FC5 are 256×N and 256×4N respectively, the size of FDC1 is 128×4N, the sizes of FK1 and FK2 are 3×4N and 128×4N respectively, and the sizes of FR3 and FR4 are 256×4N and 3×4N respectively.
[0036] Further defined as Figure 4As shown, the local geometry retainer includes a fifth multi-layer perceptron layer, a first K-nearest neighbor grouping layer, a fourth distribution feature calculation layer, a third convolutional layer, a third batch normalization layer, a fifth non-linear activation layer, a fifth distribution feature calculation layer, a fourth convolutional layer, a fourth batch normalization layer, a sixth non-linear activation layer, a sixth distribution feature calculation layer, and a second affine transformation layer; the implementation process of the local geometry retainer is as follows: the sparse point cloud block is input into the fifth multi-layer perceptron layer, the feature map output by the fifth multi-layer perceptron layer and the sparse point cloud block are input into the first K-nearest neighbor grouping layer together, one of the feature maps output by the first K-nearest neighbor grouping layer passes through the fourth distribution feature calculation layer, the third convolutional layer, the third batch normalization layer, and the fifth non-linear activation layer in sequence, the other feature map output by the first K-nearest neighbor grouping layer passes through the fifth distribution feature calculation layer, the fourth convolutional layer, the fourth batch normalization layer, and the sixth non-linear activation layer in sequence, the feature map output by the fifth non-linear activation layer and the feature map output by the sixth non-linear activation layer are subjected to a channel connection operation to obtain a sixth channel connection feature map, the sixth channel connection feature map passes through the sixth distribution feature calculation layer and the second affine transformation layer in sequence, and the feature map obtained by performing a residual connection operation on the feature map output by the fifth multi-layer perceptron layer and the feature map output by the second affine transformation layer is used as the student feature map output by the local geometry retainer.
[0037] In this embodiment, the input size of the fifth multi-layer perceptron layer is 3×N, and the output size is 128×N. The two input sizes of the first K-nearest neighbor grouping layer are 128×N and 3×N respectively, and the two output sizes are 128×N×12 and 3×N×12 respectively. The input size of the fourth distribution feature calculation layer is 128×N×12, and the output size is 128×N. The input size of the third convolutional layer is 128×N, and the output size is 64×N. The input size of the third batch normalization layer is 64×N, and the output size is 64×N. The input size of the fifth non-linear activation layer is 64×N, and the output size is 64×N. The input size of the fifth distribution feature calculation layer is 3×N×12, and the output size is 3×N. The input size of the fourth convolutional layer is 3×N, and the output size is 64×N. The input size of the fourth batch normalization layer is 64×N, and the output size is 64×N. The input size of the sixth non-linear activation layer is 64×N, and the output size is 64×N. The input size of the sixth distribution feature calculation layer is 128×N, and the output size is 128×1. The input size of the second affine transformation layer is 128×1, and the output size is 128×N. See Figure 4, specifically: input the sparse point cloud block into the fifth multi-layer perceptron layer, and denote the output feature map as FM5; input FM5 and the sparse point cloud block into the first K-nearest neighbor grouping layer, and denote the two output feature maps as FKM1 and FKP1 respectively; input FKM1 into the fourth distribution feature calculation layer, and denote the output feature map as FD4; input FD4 into the third convolutional layer, and denote the output feature map as FV3; input FV3 into the third batch normalization layer, and denote the output feature map as FN3; input FN3 into the fifth non-linear activation layer, and denote the output feature map as FR5; input FKP1 into the fifth distribution feature calculation layer, and denote the output feature map as FD5; input FD5 into the fourth convolutional layer, and denote the output feature map as FV4; input FV4 into the fourth batch normalization layer, and denote the output feature map as FN4; input FN4 into the sixth non-linear activation layer, and denote the output feature map as FR6; perform a channel connection operation on FR5 and FR6, and denote the obtained sixth channel connection feature map as FC6; input FC6 into the sixth distribution feature calculation layer, and denote the output feature map as FD6; input FD6 into the second affine transformation layer, and denote the output feature map as FA2; perform a residual connection operation on FM5 and FA2, and denote the obtained feature map as FG4, and FG4 is used as the student feature map output by the local geometry retainer; where the size of FM5 is 128×N, the sizes of FKM1 and FKP1 correspond to 128×N×12 and 3×N×12 respectively, the sizes of FD4 and FD5 correspond to 128×N and 3×N respectively, the sizes of FV3, FV4, FN3, FN4, FR5 and FR6 are all 64×N, the size of FC6 is 128×N, the size of FD6 is 128×1, and the sizes of FA2 and FG4 are both 128×N.
[0038] Further defined, as Figure 5 shown, the locally perceptive Mamba enhancer includes a second K-nearest neighbor grouping layer, a sixth multi-layer perceptron layer, a fourth max pooling layer, a serialization layer, a Mamba layer, a second deconvolution layer, a third nearest neighbor interpolation layer, a seventh non-linear activation layer, a seventh multi-layer perceptron layer, an eighth non-linear activation layer, a fourth nearest neighbor interpolation layer; the implementation process of the locally perceptive Mamba enhancer is: input the sparse point cloud block and the student feature Figure 1It is input into the second K-nearest neighbor grouping layer, and channel connection operations are performed on the two feature maps output by the second K-nearest neighbor grouping layer to obtain the seventh channel-connected feature map. The seventh channel-connected feature map successively passes through the sixth multi-layer perceptron layer, the fourth max pooling layer, the serialization layer, and the Mamba layer. The feature map output by the Mamba layer is simultaneously input into the second deconvolution layer and the third nearest neighbor interpolation layer. Channel connection operations are performed on the feature map output by the second deconvolution layer and the feature map output by the third nearest neighbor interpolation layer to obtain the eighth channel-connected feature map. The eighth channel-connected feature map passes through the seventh non-linear activation layer, the seventh multi-layer perceptron layer, and the eighth non-linear activation layer. The feature map obtained by performing a residual connection operation on the feature map obtained after passing the sparse point cloud block through the fourth nearest neighbor interpolation layer and the feature map output by the eighth non-linear activation layer is used as the student dense point cloud block output by the Mamba enhancer for local perception.
[0039] In this embodiment, the two input sizes of the second K-nearest neighbor grouping layer are 128×N and 3×N respectively, and the two output sizes are 128×N×12 and 3×N×12 respectively. The input size of the sixth multi-layer perceptron layer is 131×N×12, and the output size is 128×N×12. The input size of the fourth max pooling layer is 128×N×12, and the output size is 128×N. The input size of the serialization layer is 128×N, and the output size is 128N×1. The input size of the Mamba layer is 128N×1, and the output size is 128×N. The input size of the second deconvolution layer is 128×N, and the output size is 128×4N. The input size of the third nearest neighbor interpolation layer is 128×N, and the output size is 128×4N. The input size of the seventh non-linear activation layer is 256×4N, and the output size is 256×4N. The input size of the seventh multi-layer perceptron layer is 256×4N, and the output size is 3×4N. The input size of the eighth non-linear activation layer is 3×4N, and the output size is 3×4N. The input size of the fourth nearest neighbor interpolation layer is 3×N, and the output size is 3×4N. See Figure 5 Specifically: the sparse point cloud block and the student features output by the local geometry retainer Figure 1It is input into the second K-nearest neighbor grouping layer, and the two output feature maps are denoted as FKM2 and FKP2 respectively; a channel connection operation is performed on FKM2 and FKP2, and the resulting seventh channel-connected feature map is denoted as FC7; FC7 is input into the sixth multi-layer perceptron layer, and the output feature map is denoted as FM6; FM6 is input into the fourth max pooling layer, and the output feature map is denoted as FP4; FP4 is input into the serialization layer, and the output feature map is denoted as FS; FS is input into the Mamba layer, and the output feature map is denoted as FB; FB is input into the second deconvolution layer, and the output feature map is denoted as FDC2; FB is input into the third nearest neighbor interpolation layer, and the output feature map is denoted as FK3; a channel connection operation is performed on FDC2 and FK3, and the resulting eighth channel-connected feature map is denoted as FC8; FC8 is input into the seventh non-linear activation layer, and the output feature map is denoted as FR7; FR7 is input into the seventh multi-layer perceptron layer, and the output feature map is denoted as FM7; FM7 is input into the eighth non-linear activation layer, and the output feature map is denoted as FR8; the sparse point cloud block is input into the fourth nearest neighbor interpolation layer, and the output feature map is denoted as FK4; a residual connection operation is performed on FK4 and FR8, and the resulting feature map is denoted as FG5, and FG5 is the student dense point cloud block output by the local perception Mamba enhancer; where the sizes of FKM2 and FKP2 are 128×N×12 and 3×N×12 respectively, the size of FC7 is 131×N×12, the size of FM6 is 128×N×12, the size of FP4 is 128×N, the size of FS is 128N×1, the size of FB is 128×N, the size of FDC2 is 128×4N, the size of FK3 is 128×4N, the size of FC8 is 256×4N, the size of FR7 is 256×4N, the size of FM7 is 3×4N, the size of FR8 is 3×4N, the size of FK4 is 3×4N, and the size of FG5 is 3×4N.
[0040] Here, in the constructed lightweight neural network, the channel concatenation operation and the residual connection operation are both conventional operations in the neural network; all multi-layer perceptron layers are in the form of three layers, each layer consists of one-dimensional convolution, batch normalization, and a linear activation function (ReLU) and has the same channel dimension. The channel dimension of the first layer is the same as the input dimension of the multi-layer perceptron layer where it is located, and the channel dimension of the third layer is the same as the output dimension of the multi-layer perceptron layer where it is located. The channel dimension of the second layer, i.e., the middle layer, is 128; the broadcast layer replicates and expands the channel dimension, which is a conventional operation in the neural network; the distribution feature calculation layer calculates the variance along the point dimension; the non-linear activation layer (ReLU) is a conventional operation in the neural network; the affine transformation layer is used for affine transformation and is documented in L. Dinh, D. Krueger, and Y. Bengio, “NICE: Non-linear Independent Components Estimation,” arXiv: Learning, 2014. [Online]. Available: https: / / api.semanticscholar.org / CorpusID:13995862. (Non-linear Independent Components Estimation); the PT (Point Transformer network) layer is an existing network module and is documented in H. Zhao, L. Jiang, J. Jia, P. Torr, and V. Koltun, “Point Transformer,” in 2021 IEEE / CVF International Conference on Computer Vision (ICCV), 2021, pp. 16239-16248. (Point Transformer); the nearest neighbor interpolation layer is a conventional operation in the neural network; the K-nearest neighbor grouping layer is documented in R.Q. Charles, L. Yi, H. Su, and L.J. Guibas, “PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space,” in Proceedings of the 31st International Conference on Neural Information Processing Systems, ser. NIPS’17. Red Hook, NY, USA: Curran Associates Inc., 2017, p. 5105-5114.(Learning the Depth Hierarchical Features of Point Sets in Metric Spaces) is recorded; the serialization layer serializes the three-dimensional point cloud into one dimension, as recorded in T. Zhang, X. Li, H. Yuan, S. Ji, and S. Yan, “Point Cloud Mamba: Point Cloud Learning via State Space Model,” ArXiv, vol. abs / 2403.00762, 2024. [Online]. Available: https: / / api.semanticscholar.org / CorpusID: 268230692. (Point Cloud Mamba); the Mamba layer is an existing network module, as recorded in A. Gu and T. Dao, “Mamba: Linear-Time Sequence Modeling with Selective State Spaces,” ArXiv, vol. abs / 2312.00752, 2023. [Online]. Available: https: / / api.semanticscholar.org / CorpusID: 265551773. (Mamba: Linear-Time Sequence Modeling with Selective State Spaces).
[0041] In a specific embodiment, the construction process of the training set is as follows: First, select a plurality of, such as at least 120, original mesh models (in mesh format) as the source of point cloud data; then perform Poisson disk sampling operations on each original mesh model to generate sparse point clouds and dense point clouds respectively. Among them, the number of points in the sparse point cloud is significantly less than that in the dense point cloud. For example, the sparse point cloud contains 2048 points and the dense point cloud contains 8192 points; then use the farthest point sampling algorithm for the sparse point cloud to find a plurality of, such as 200, seed points; then use the nearest neighbor algorithm for each seed point to extract sparse point cloud patches from the sparse point cloud and extract dense point cloud patches from the dense point cloud. The dense point cloud patches are used as the true dense point cloud patches, and the sparse point cloud patches and the true dense point cloud patches are paired to form a training pair. Among them, the number of points in the sparse point cloud patch is significantly less than that in the dense point cloud patch. For example, the sparse point cloud patch contains 256 points and the dense point cloud patch contains 1024 points; the number of training pairs generated by each original mesh model is the same as the number of seed points, that is, 200, and the final number of generated training pairs is equal to the product of the number of original mesh models and the number of seed points, that is, at least 24000. The process of obtaining the test set is as follows: First, arbitrarily select M test sparse point clouds; then use the farthest point sampling algorithm for each test sparse point cloud to find r×S / (3×N) seed points; then use the nearest neighbor algorithm for each seed point to extract sparse point cloud patches containing N points from the test sparse point cloud. For one test sparse point cloud, a total of r×S / (3×N) sparse point cloud patches are extracted; finally, M×r×S / (3×N) sparse point cloud patches form the test set; where M≥1, r represents the upsampling rate, S represents the number of points in the sparse point cloud, and 3×N represents the size of the sparse point cloud patch. In this embodiment, r = 4, S = 2048, and N = 256 are taken. After upsampling to obtain the student dense point cloud patches corresponding to each sparse point cloud patch in the test set, the student dense point cloud patches corresponding to the r×S / (3×N) sparse point cloud patches from the same test sparse point cloud after upsampling are spliced together, and then the farthest point sampling algorithm is used to sample r×S points, and finally the r-fold upsampling of the test sparse point cloud is completed.
[0042] To further illustrate the feasibility and effectiveness of the method of the present invention, experiments are carried out on the method of the present invention.
[0043] In the experiment, the PU-GAN dataset is selected. The training set in the PU-GAN dataset includes 24000 training pairs, and the test set includes 27 sparse point cloud patches. The method of the present invention is used to test on the test set in the PU-GAN dataset.
[0044] In the experiment, six commonly used objective parameters were selected to evaluate the performance of the method of the present invention, which are Chamfer Distance (CD), Hausdorff Distance (HD), Earth Mover's Distance (EMD), Average Point-to-Face Distance (P2F), number of network parameters (Parms), and floating-point operations (19.808). Table 1 shows the upsampling performance tested using the method of the present invention on the test set of the PU-GAN dataset.
[0045] Table 1 Upsampling performance tested using the method of the present invention on the test set of the PU-GAN dataset
[0046] CD HD EMD P2F Parms(MB) Calculate floating point numbers (GB) The method of the present invention 0.243 1.521 2.306 1.675 0.832 19.808
[0047] From the results given in Table 1, it can be found that the method of the present invention achieves lower CD, HD, EMD, and P2F on the existing dataset, which indicates that the student dense point cloud blocks obtained by the method of the present invention are very close to the ground truth dense point cloud blocks. Moreover, the method of the present invention can effectively complete the upsampling of sparse point cloud blocks, presenting a lightweight number of network parameters (0.832MB) and floating-point operations (19.808GB), taking into account both performance and computational overhead.
Claims
1. A point cloud upsampling method based on a lightweight neural network, characterized in that: First, construct a training set containing pairs of sparse point cloud patches and ground-truth dense point cloud patches, and build a lightweight neural network consisting of a teacher network and a student network; second, input the pairs of sparse point cloud patches and ground-truth dense point cloud patches in the training set into the lightweight neural network for multiple rounds of network training, and after the network training is completed, obtain the student network training model; Third, use the student network training model to upsample each sparse point cloud patch in the test set, and after upsampling, obtain the student dense point cloud patch corresponding to each sparse point cloud patch in the test set; where: The teacher network consists of a global context retainer and a globally-aware Transformer upsampler, and the student network consists of a local geometry retainer and a locally-aware Mamba enhancer; first, input the sparse point cloud patch into the global context retainer and the local geometry retainer at the same time, the global context retainer outputs the teacher feature map, and the local geometry retainer outputs the student feature map; second, input the sparse point cloud patch and the teacher feature map into the globally-aware Transformer upsampler together to obtain the teacher dense point cloud patch; and input the sparse point cloud patch and the student feature map into the locally-aware Mamba enhancer together to obtain the student dense point cloud patch; When training the lightweight neural network, first train the teacher network, and then train the student network after freezing the parameters of the teacher network; when training the teacher network, calculate the teacher network loss before the end of each round of network training to optimize the parameters of the teacher network to constrain the network training, stop the training of the teacher network after multiple rounds of network training, and freeze the parameters of the teacher network; when training the student network, calculate the student network loss before the end of each round of network training to optimize the parameters of the student network to constrain the network training, stop the training of the student network after multiple rounds of network training, and obtain the student network training model, where the student network loss includes the dual-alignment knowledge distillation loss used to align features and points.
2. The method for upsampling point clouds based on a lightweight neural network according to claim 1, wherein The global context keeper includes a first multi-layer perceptron layer, a first max pooling layer, a first broadcast layer, a first distribution feature calculation layer, a second broadcast layer, a first convolutional layer, a first batch normalization layer, a first non-linear activation layer, a second max pooling layer, a second distribution feature calculation layer, a third broadcast layer, a second convolutional layer, a second batch normalization layer, a second non-linear activation layer, a third distribution feature calculation layer, and a first affine transformation layer; the implementation process of the global context keeper is as follows: simultaneously input the sparse point cloud block into the first multi-layer perceptron layer, the second distribution feature calculation layer, and the second max pooling layer, pass the feature map output by the first multi-layer perceptron layer through the first max pooling layer and the first broadcast layer in sequence, simultaneously pass the feature map output by the first multi-layer perceptron layer through the first distribution feature calculation layer and the second broadcast layer, perform a channel connection operation on the feature map output by the first broadcast layer and the feature map output by the second broadcast layer to obtain a first channel connection feature map, the first channel connection feature map passes through the first convolutional layer, the first batch normalization layer, and the first non-linear activation layer in sequence, perform a channel connection operation on the sparse point cloud block, the feature map obtained by passing the feature map output by the second distribution feature calculation layer through the third broadcast layer, and the feature map output by the second max pooling layer to obtain a second channel connection feature map, the second channel connection feature map passes through the second convolutional layer, the second batch normalization layer, and the second non-linear activation layer in sequence, perform a channel connection operation on the feature map output by the first non-linear activation layer and the feature map output by the second non-linear activation layer to obtain a third channel connection feature map, the third channel connection feature map passes through the third distribution feature calculation layer and the first affine transformation layer in sequence, and perform a residual connection operation on the feature map output by the first multi-layer perceptron layer and the feature map output by the first affine transformation layer, and use the obtained feature map as the teacher feature map output by the global context keeper.
3. A point cloud upsampling method based on a lightweight neural network according to claim 1 or 2, characterized in that The globally-aware Transformer upsampler includes a first nearest neighbor interpolation layer, a second multi-layer perceptron layer, a third max pooling layer, a fourth broadcast layer, a third multi-layer perceptron layer, a PT layer, a first deconvolution layer, a second nearest neighbor interpolation layer, a third non-linear activation layer, a fourth multi-layer perceptron layer, and a fourth non-linear activation layer. The implementation process of the globally-aware Transformer upsampler is as follows: The sparse point cloud block is simultaneously input into the first nearest neighbor interpolation layer and the second multi-layer perceptron layer. The feature map output by the second multi-layer perceptron layer and the teacher feature map output by the global context keeper are subjected to a residual connection operation. The feature map obtained by the residual connection is channel-connected with the feature map obtained after passing through the third max pooling layer and the fourth broadcast layer in sequence to obtain a fourth channel-connected feature map. The fourth channel-connected feature map passes through the third multi-layer perceptron layer and the PT layer in sequence. The feature map output by the PT layer is simultaneously input into the first deconvolution layer and the second nearest neighbor interpolation layer. The feature map output by the first deconvolution layer and the feature map output by the second nearest neighbor interpolation layer are channel-connected to obtain a fifth channel-connected feature map. The fifth channel-connected feature map passes through the third non-linear activation layer, the fourth multi-layer perceptron layer, and the fourth non-linear activation layer in sequence. The feature map obtained by performing a residual connection operation on the feature map output by the first nearest neighbor interpolation layer and the feature map output by the fourth non-linear activation layer is used as the teacher dense point cloud block output by the globally-aware Transformer upsampler.
4. The method for upsampling point clouds based on a lightweight neural network according to claim 3, wherein The local geometry retainer includes a fifth multi-layer perceptron layer, a first K-nearest neighbor grouping layer, a fourth distribution feature calculation layer, a third convolution layer, a third batch normalization layer, a fifth non-linear activation layer, a fifth distribution feature calculation layer, a fourth convolution layer, a fourth batch normalization layer, a sixth non-linear activation layer, a sixth distribution feature calculation layer, and a second affine transformation layer. The implementation process of the local geometry retainer is as follows: The sparse point cloud block is input into the fifth multi-layer perceptron layer. The feature map output by the fifth multi-layer perceptron layer and the sparse point cloud block are input into the first K-nearest neighbor grouping layer together. One of the feature maps output by the first K-nearest neighbor grouping layer passes through the fourth distribution feature calculation layer, the third convolution layer, the third batch normalization layer, and the fifth non-linear activation layer in sequence. The other feature map output by the first K-nearest neighbor grouping layer passes through the fifth distribution feature calculation layer, the fourth convolution layer, the fourth batch normalization layer, and the sixth non-linear activation layer in sequence. The feature map output by the fifth non-linear activation layer and the feature map output by the sixth non-linear activation layer are channel-connected to obtain a sixth channel-connected feature map. The sixth channel-connected feature map passes through the sixth distribution feature calculation layer and the second affine transformation layer in sequence. The feature map obtained by performing a residual connection operation on the feature map output by the fifth multi-layer perceptron layer and the feature map output by the second affine transformation layer is used as the student feature map output by the local geometry retainer.
5. The method for upsampling point clouds based on a lightweight neural network according to claim 4, characterized in that The local perception Mamba enhancer includes a second K-nearest neighbor grouping layer, a sixth multi-layer perceptron layer, a fourth max pooling layer, a serialization layer, a Mamba layer, a second deconvolution layer, a third nearest neighbor interpolation layer, a seventh non-linear activation layer, a seventh multi-layer perceptron layer, an eighth non-linear activation layer, and a fourth nearest neighbor interpolation layer; the implementation process of the local perception Mamba enhancer is as follows: the sparse point cloud block and the student feature map output by the local geometry retainer are input into the second K-nearest neighbor grouping layer together, and channel connection operations are performed on the two feature maps output by the second K-nearest neighbor grouping layer to obtain a seventh channel-connected feature map. The seventh channel-connected feature map passes through the sixth multi-layer perceptron layer, the fourth max pooling layer, the serialization layer, and the Mamba layer in sequence. The feature map output by the Mamba layer is input into the second deconvolution layer and the third nearest neighbor interpolation layer at the same time. Channel connection operations are performed on the feature map output by the second deconvolution layer and the feature map output by the third nearest neighbor interpolation layer to obtain an eighth channel-connected feature map. The eighth channel-connected feature map passes through the seventh non-linear activation layer, the seventh multi-layer perceptron layer, and the eighth non-linear activation layer. The feature map obtained by performing residual connection operations on the feature map obtained by passing the sparse point cloud block through the fourth nearest neighbor interpolation layer and the feature map output by the eighth non-linear activation layer is used as the student dense point cloud block output by the local perception Mamba enhancer.
6. The method for upsampling point clouds based on a lightweight neural network according to claim 1, characterized in that The construction process of the training set is as follows: First, select multiple original mesh models as the source of point cloud data; then perform Poisson disk sampling operations on each original mesh model to generate sparse point clouds and dense point clouds respectively. Among them, the number of points in the sparse point cloud is significantly less than the number of points in the dense point cloud; then use the farthest point sampling algorithm for the sparse point cloud to find multiple seed points; then use the nearest neighbor algorithm for each seed point to extract sparse point cloud blocks from the sparse point cloud and extract dense point cloud blocks from the dense point cloud. The dense point cloud block is used as the true dense point cloud block, and the sparse point cloud block and the true dense point cloud block are paired to form a training pair. Among them, the number of points in the sparse point cloud block is significantly less than the number of points in the dense point cloud block; the number of training pairs generated by each original mesh model is the same as the number of seed points, and the final number of generated training pairs is equal to the product of the number of original mesh models and the number of seed points.
7. A point cloud upsampling method based on a lightweight neural network according to claim 1, characterized in that Denote the teacher network loss as \(L\) T , Denote the student network loss as \(L\) S , \(L\) S = \(L\) CD + \(L\) DAD , \(L\) DAD = \(L\) KL (FG1, FG4)+ \(L\) MSE (FG1, FG4)+ \(L\) KL (FG3, FG5)+ \(L\) MSE (FG3, FG5), where \(L\) T is obtained by calculating the chamfer distance loss between the ground-truth dense point cloud patch and the teacher dense point cloud patch output by the globally-aware Transformer upsampler. FG3 represents the teacher dense point cloud patch output by the globally-aware Transformer upsampler, represents the number of points in FG3, \(p\) O represents a point in FG3, min(·) is the minimum function, \(P\) HR represents the ground-truth dense point cloud patch, \(p\) T represents a point in \(P\) HR . represents the number of points in \(P\) HR , \(\|\cdot\|_2\) is the symbol of the two-norm operation, \(L\) CD represents the chamfer distance loss between the ground-truth dense point cloud patch and the student dense point cloud patch output by the locally-aware Mamba enhancer, \(L\) DAD represents the dual-alignment knowledge distillation loss, FG5 represents the student dense point cloud patch output by the locally-aware Mamba enhancer, represents the number of points in FG5, represents a point in FG5, \(L\) KL (·) represents the KL divergence loss, \(L\) MSE (·) represents the mean squared error loss, FG1 represents the teacher feature map output by the global context retainer, and FG4 represents the student feature map output by the local geometry retainer.
8. A method for upsampling point clouds based on a lightweight neural network according to claim 1, characterized in that The acquisition process of the test set is as follows: First, arbitrarily select M test sparse point clouds; then use the farthest point sampling algorithm for each test sparse point cloud to find r×S / (3×N) seed points; then use the nearest neighbor algorithm for each seed point to extract sparse point cloud blocks containing N points from the test sparse point cloud. A total of r×S / (3×N) sparse point cloud blocks are extracted for one test sparse point cloud; finally, M×r×S / (3×N) sparse point cloud blocks form the test set; where M≥1, r represents the upsampling rate, S represents the number of points in the sparse point cloud, and 3×N represents the size of the sparse point cloud block.
9. A method for point cloud upsampling based on a lightweight neural network according to claim 8, characterized in that After upsampling to obtain the student dense point cloud blocks corresponding to each sparse point cloud block in the test set, the corresponding student dense point cloud blocks obtained by upsampling the r×S / (3×N) sparse point cloud blocks from the same test sparse point cloud are spliced together, and then r×S points are sampled using the farthest point sampling algorithm, finally completing the r-fold upsampling of the test sparse point cloud.
Citation Information
Cited By
Point cloud up-sampling method and system based on Mama model
CN121032785A