Cross-modal distillation point cloud up-sampling method guided by multi-view depth map
Through the multi-view depth map-guided cross-modal distillation method, the dual-branch feature extraction module and knowledge distillation structure are used to solve the problems of insufficient characterization and lack of details in the point cloud upsampling method in the existing technology, and high-quality point cloud density and refinement are achieved.
Patent Information
- Application Number
- CN202411941584.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing point cloud up-sampling methods only utilize the three-dimensional coordinates of a single mode, resulting in insufficient characterization and lack of geometric details, making it difficult to effectively dense and refine sparse point clouds.
The multi-view depth map-guided cross-modal distillation point cloud upsampling method is adopted. Through the dual-branch cross-modal feature extraction module and multi-view to point feature fusion module, combined with the knowledge distillation structure of the teacher network and the student network, the features of the point cloud and multi-view depth map are fully extracted and the fine geometric structure is generated.
The quality of upsampled point clouds is significantly improved, and the dense point clouds generated have better geometric details and distribution uniformity, solving the problems of insufficient characterization and missing details in the existing methods.
Smart Images

Figure CN120071074A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of point cloud upsampling, and in particular, to a point cloud upsampling method based on cross-modal distillation guided by multi-view depth maps in deep learning. Background Art
[0002] Point cloud data can concisely and efficiently represent three-dimensional objects and three-dimensional scenes, and is widely used in fields such as autonomous driving, three-dimensional reconstruction, and virtual reality. Due to hardware and computing power limitations, the original point cloud data directly obtained by three-dimensional sensors is often sparse, unevenly distributed, and contains noise. This low-quality point cloud data greatly affects the performance of downstream tasks.
[0003] Point cloud upsampling is a strategy that can densify point clouds, process sparse point clouds, and obtain dense, evenly distributed, and noise-free point clouds. Currently, although existing methods can achieve the densification of sparse point clouds, they still face two challenges: existing methods only use the single modality of the three-dimensional coordinates of sparse point clouds, resulting in insufficient single-modality representations; the dense point clouds generated by existing methods lack geometric details. Therefore, how to fully utilize other available modality data to extract sufficient cross-modal representations of sparse point clouds and generate fine geometric details in upsampled point clouds remains a challenge.
[0004] In view of this, the purpose of this application is to provide a point cloud upsampling method based on cross-modal distillation guided by multi-view depth maps. This method proposes a dual-branch cross-modal feature extraction module to fully extract point cloud features and multi-view depth map features. At the same time, the cross-modal feature extraction module also includes a cross-modal feature fusion module that fuses pixel-level multi-view depth map features into the features of each point in the point cloud. In addition, in order to further generate fine geometric structures and make full use of the real dense point cloud data during the training process, this method also proposes a detail estimation and distillation structure, which includes a teacher network and a student network. The input of the teacher network is the initial point cloud data, the multi-view depth maps of the paired initial point cloud data, and the multi-view depth maps of the real dense point cloud data. First, the teacher network is trained, and then the student network only takes the initial point cloud data and the multi-view depth maps of the initial point cloud data as input and is trained under the guidance of the pre-trained teacher network to learn the upsampling method of the teacher network through knowledge distillation.
[0005] The embodiments of this application provide a point cloud upsampling method based on cross-modal distillation guided by multi-view depth maps. The point cloud upsampling network includes a cross-modal feature extraction module and a detail estimation and distillation structure, and the cross-modal feature extraction module also includes a cross-modal feature fusion module; the determination method includes:
[0006] Obtain the initial point cloud data and the real dense point cloud data of the target three-dimensional object or three-dimensional scene, where the real dense point cloud data is only used for training, and only the initial point cloud data is used during the testing process;
[0007] Input the initial point cloud data into the multi-view depth map rendering module to obtain the multi-view depth map of the initial point cloud data;
[0008] Input the real dense point cloud data into the multi-view depth map rendering module to obtain the multi-view depth map of the real dense point cloud data;
[0009] Input the initial point cloud data into the point cloud branch in the cross-modal feature extraction module of the upsampling teacher network, and splice the multi-view depth map of the initial point cloud data and the multi-view depth map of the real dense point cloud data along the channel dimension under the same view and then input them into the depth map branch in the cross-modal feature extraction module of the upsampling teacher network to obtain the cross-modal representation of the initial point cloud data;
[0010] Input the cross-modal representation of the initial point cloud data into the upsampling tail module of the teacher network to obtain the rough upsampled point cloud predicted by the upsampling teacher network and the refined upsampled point cloud predicted by the upsampling teacher network;
[0011] Optimize the upsampling teacher network based on the target values to obtain the optimized upsampling teacher network; wherein, the target values include the CD distance between the rough upsampled point cloud predicted by the upsampling teacher network and the real dense point cloud, the CD distance between the refined dense upsampled point cloud predicted by the upsampling teacher network and the real dense point cloud, the L 1 loss between the multi-view depth map of the rough upsampled point cloud predicted by the upsampling teacher network and the multi-view depth map of the real dense point cloud, and the L 1 loss of the multi-view depth map of the refined upsampled point cloud predicted by the upsampling teacher network and the multi-view depth map of the real dense point cloud;
[0012] Load the optimized upsampling teacher network, freeze the parameters of the optimized upsampling teacher network, and initialize the parameters of the upsampling student network; wherein, the upsampling student network has the same network structure as the upsampling teacher network;
[0013] Input the initial point cloud data into the point cloud branch of the upsampling student network, and input the multi-view depth map of the initial point cloud data into the depth map branch of the upsampling student network to obtain the cross-modal representation of the initial point cloud data;
[0014] Input the cross-modal representation of the initial point cloud data extracted by the upsampling student network into the upsampling tail module of the upsampling student network to obtain the rough upsampled point cloud predicted by the upsampling student network and the refined upsampled point cloud predicted by the upsampling student network;
[0015] Optimize the upsampling student network based on the objective function value to obtain the optimized upsampling student network; where the objective value includes the reconstruction loss and the knowledge distillation loss; among them, the reconstruction loss includes the CD distance between the rough upsampled point cloud predicted by the upsampling student network and the real dense point cloud, the CD distance between the refined dense upsampled point cloud predicted by the upsampling student network and the real dense point cloud, the L 1 loss between the multi-view depth map of the rough upsampled point cloud predicted by the upsampling student network and the multi-view depth map of the real dense point cloud, and the L 1 loss between the multi-view depth map of the refined upsampled point cloud predicted by the upsampling student network and the multi-view depth map of the real dense point cloud; the knowledge distillation loss includes two parts, one part is the knowledge distillation loss based on the response, and the other part is the knowledge distillation loss based on the feature; among them, the knowledge distillation loss based on the response includes the L 2 loss between the rough upsampled point cloud predicted by the upsampling student network and the rough upsampled point cloud predicted by the upsampling teacher network, and the L 2 loss between the rough upsampled point cloud predicted by the upsampling student network and the rough upsampled point cloud predicted by the upsampling teacher network; the knowledge distillation loss based on the feature is the L 2 loss between the multi-scale representation of the dense point cloud output by the depth map branch of the upsampling teacher network and the multi-scale representation of the dense point cloud predicted by the depth map branch of the upsampling student network;
[0016] Furthermore, the cross-modal feature extraction module includes a point cloud branch, a depth map branch, and a multi-view depth map to point feature fusion module, which are used to extract point cloud features and the multi-view depth map features corresponding to the point cloud, and to fuse the multi-view depth map features into the point cloud features respectively;
[0017] The depth map branch in the cross-modal feature extraction module includes multiple ResNet structures; each ResNet structure includes a residual downsampling module and a residual convolution module; where the residual downsampling module reduces the resolution of the depth map to half of the resolution of the input feature map, and the number of channels becomes twice that of the input feature map;
[0018] The point cloud branch in the cross-modal feature extraction module includes multiple densely connected DenseNet structures; the input of each DenseNet structure is the concatenation of the point cloud features output by all previous DenseNet structures and the point cloud features after fusing with the multi-view depth map features output by the ResNet structure of the same layer; each DenseNet consists of a dynamic graph construction module and multiple densely connected EdgeConv; among them, the input of each EdgeConv is the concatenation of the dynamic graph and the outputs of all previous EdgeConv;
[0019] The multi-view to point feature fusion module in the cross-modal feature extraction module takes the point cloud features output by the DenseNet structure and the multi-view depth map features output by the ResNet of the corresponding layer as inputs; this module includes a multi-view cross-attention structure and a feature fusion structure for attention; among them, the multi-view cross-attention aligns the point cloud features and the multi-view depth map features in the feature space through different linear layers, and the calculation formula is as follows:
[0020]
[0021] Among them, represents the point cloud features output by the l-th layer DenseNet structure, represents the multi-view depth map features output by the l-th layer ResNet structure, is the number of channels, H and W represent the original depth map spatial resolution, N V represents the number of viewpoints of the multi-view depth map, i represents the index of the depth map viewpoint, represents the set of projection matrices of the point cloud features, is the number of channels of, projects the point cloud features into a query set under different viewpoints and is the projection matrix of the depth map, projects the depth map features into a key set under different viewpoints and the value set d k is the number of channels of the query and the key, d v is the number of channels of the value;
[0022] Among them, is the point cloud feature, is the depth map feature, N v is the number of viewpoints of the multi-view depth map;
[0023] Subsequently, cross-attention is calculated between Q i 、K i and V i at the same viewpoint, aggregates useful depth map features for each point in the point cloud, and the calculation formula is as follows:
[0024]
[0025] Among them, is the softmax activation function, is the depth map features from different aggregated perspectives;
[0026] Subsequently, is concatenated along the channel dimension and passed through a linear layer to integrate the depth map features from different aggregated perspectives. The calculation formula is as follows:
[0027]
[0028] Among them, and represent the weight parameters and bias parameters of the linear layer, and Concat represents concatenation along the channel dimension;
[0029] Furthermore, through the attention fusion module, the is further fused with the point cloud feature . The formula for this process is as follows:
[0030]
[0031] Among them, is the point cloud feature after fusion, and β is the weighting parameter predicted by the sub-network ;
[0032] The weight prediction sub-network first adds the corresponding elements of and , then passes through a batch normalization layer to reduce the difference in data distribution between different modalities. The normalized features will be sent to the global branch and the local branch. Both the global branch and the local branch are composed of two multi-layer perceptrons. In addition to the multi-layer perceptron mentioned above, the global branch also performs global pooling on the input features first. Subsequently, the features output by the two branches are added element-wise and then sent to the sigmoid activation function to obtain the above-mentioned weight β. The formula for this process is as follows:
[0033]
[0034] Among them, MLP G and MLP L are the multi-layer perceptrons of the global branch and the local branch respectively, GAP is the global pooling operation, is the batch normalization operation, E[·] and Var[·] respectively represent calculating the expectation and standard deviation, ω and μ are two learnable parameters, ∈ = 10-6 is a very small constant, ReLU(x)=max{0,x} is the ReLU activation function;
[0035] A computer system provided by the present invention comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the cross-modal distillation point cloud upsampling method based on multi-view depth map guidance is implemented;
[0036] Beneficial effects: Compared with the existing point cloud upsampling methods, this method has the following advantages:
[0037] 1. This paper designs a cross-modal point cloud upsampling deep neural network framework, which renders point cloud data into depth maps and significantly improves the quality of upsampled point clouds without introducing additional information.
[0038] 2. The present invention designs a point cloud cross-modal feature extraction module to obtain the cross-modal representation of the initial point cloud data. The cross-modal feature extraction module includes a point cloud branch and a depth map branch. Considering the depth map branch, it is mainly composed of the ResNet structure, and the multi-scale representation of the multi-view depth map is fully extracted through residual downsampling and residual convolution modules. Considering the point cloud branch, it is mainly composed of the DenseNet structure, and the point cloud features at different levels are effectively organized through dense connections within blocks and dense connections between blocks.
[0039] 3. The present invention designs a multi-view point feature fusion module, which acts on different levels of the cross-modal feature extraction module, and fuses the multi-view depth map features output by ResNet into the point cloud features output by the same layer DenseNet. The multi-view point feature fusion module includes a multi-view cross-attention and attention feature fusion module, which can effectively aggregate beneficial depth map features;
[0040] 4. The present invention designs a detail estimation structure based on knowledge distillation. Specifically, this structure includes an upsampling teacher network and an upsampling student network. Considering the upsampling teacher network, this network takes the initial point cloud data, the multi-view depth maps of the paired initial point cloud data, and the multi-view depth maps of the real dense point cloud as inputs to predict the upsampled point cloud and generate multi-level detail representations of the dense point cloud. Considering the upsampling student network, this network only takes the initial point cloud data and the multi-view depth maps of the initial point cloud data as inputs to predict the upsampled point cloud. The detail estimation structure based on knowledge distillation first trains the upsampling teacher network. During the training of the upsampling student network, the teacher network only performs inference and freezes the parameters of the pre-trained teacher network. The upsampling student network learns the multi-level detail representations of the dense point cloud extracted from the upsampling teacher network through a response-based knowledge distillation loss and a feature-based knowledge distillation loss;
[0041] 5. The experimental results in the following specific embodiments confirm the effectiveness and superiority of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic structural diagram of a cross-modal distillation point cloud upsampling method guided by multi-view depth maps according to the present invention.
[0043] Figure 2 It is a schematic structural diagram of DenseNet and ResNet according to the present invention.
[0044] Figure 3 It is a schematic diagram of a multi-view to point (MVP) feature fusion module according to the present invention.
[0045] Figure 4 It is an effect diagram of the upsampled point cloud according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The present invention will be further described in detail below with reference to the drawings and embodiments. The embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] A cross-modal distillation point cloud upsampling method guided by multi-view depth maps provided by an embodiment of the present invention proposes a novel cross-modal point cloud upsampling network architecture, which renders point cloud data in different directions to obtain multi-view depth maps of the point cloud data. A dual-branch cross-modal feature extraction module for point clouds is designed. This module includes a point cloud branch and a multi-view depth map branch, which are respectively used to extract point cloud features and multi-view depth map features. Among them, the point cloud branch is composed of a DenseNet structure, while the multi-view depth map branch is composed of a ResNet structure. These two structures are very different and can respectively process irregular point cloud attributes and regular multi-view depth maps. In addition, in order to fully fuse the features of these two modalities, the present invention designs a multi-view to point (MVP) feature fusion module, which is composed of a multi-view cross-attention and an attention-based feature fusion module, and can adaptively fuse beneficial multi-view depth map features into point cloud features. In addition, the present invention also proposes a detail estimation structure based on knowledge distillation. This structure includes an upsampling teacher network and an upsampling student network. This structure enables the upsampling network to improve its reconstruction ability for geometric details and complex structures in dense point clouds without introducing additional information. The structural flow diagram of the network based on the embodiment of the present invention is as shown in Figure 1 shown.
[0048] Specifically, for a cross-modal distillation point cloud upsampling method guided by multi-view depth maps according to an embodiment of the present invention, first, the initial point cloud data is rendered to obtain point cloud depth maps from different perspectives; then, the initial point cloud data and the multi-view depth maps of the initial point cloud data are input into the cross-modal feature extraction module to obtain the cross-modal representation of the initial point cloud data; then, the cross-modal representation of the initial point cloud data is input into the upsampling tail to obtain a rough upsampled point cloud and a refined upsampled point cloud;
[0049] In addition, in order to enable the network to generate more refined geometric structures and local details, the present invention also proposes a detail estimation structure based on knowledge distillation; the detail estimation structure includes an upsampling teacher network and an upsampling student network, and the upsampling teacher network and the upsampling student network have the same network structure; among them, the upsampling teacher network is first trained. The upsampling teacher network takes not only the initial point cloud data as input, but also the multi-view depth maps of the aligned initial point cloud data and the multi-view depth maps of the real dense point cloud data as input, so as to generate multi-level detail representations of the dense point cloud; after the upsampling teacher network is fully trained, freeze the parameters of the upsampling teacher network, and perform inference to obtain the multi-level detail representations of the dense point cloud extracted by the teacher network; the upsampling student network only takes the initial point cloud data and the multi-view depth maps of the initial point cloud data as input to obtain the predicted multi-level detail representations of the dense point cloud; then, based on feature-based knowledge distillation and response-based knowledge distillation, guide the upsampling student network to learn the multi-level detail representations of the dense point cloud extracted by the upsampling teacher network;
[0050] The following will combine Figures 1 to 4 to detail the specific steps of the above method:
[0051] Step 1: Obtain the initial point cloud data and the real dense point cloud data of the target object, which are used as the input data and the supervision data of the upsampling network respectively, and normalize the initial point cloud data and the real dense point cloud data to obtain the normalized initial point cloud data N×3 and the normalized real dense point cloud data rN×3, where N is the number of points in the initial point cloud data and r is the upsampling rate;
[0052] Step 2: Project the normalized initial point cloud data and the real dense point cloud data to obtain the multi-view depth maps N v ×H×W×1 of the initial point cloud data, the multi-view depth maps N v ×H×W×1 of the real dense point cloud data, and the depth map of the initial point cloud data and the depth map of the real dense data N v ×H×W×2 after splicing, where N v represents the number of viewpoints of the multi-view depth map, and the number of channels of each depth map is 1;
[0053] Step 3: Initialize the parameters of the upsampling teacher network, and input the initial point cloud data N×3 and the depth map of the initial point cloud data and the depth map of the real dense data N v ×H×W×2 after splicing into the cross-modal feature extraction module in the upsampling teacher network to extract the cross-modal representation of the initial point cloud data and the multi-level detail representation of the dense point cloud data
[0054] Step 4: Input the cross-modal representation of the initial point cloud data extracted by the upsampling teacher network into the upsampling tail module of the upsampling teacher network to obtain the rough dense point cloud predicted by the teacher network and the refined dense point cloud
[0055] Step 5: Optimize the upsampling teacher network based on the target value of the upsampling teacher network to obtain the parameters of the optimized upsampling teacher network;
[0056] Step 6: Initialize the upsampling student network, and load and freeze the optimized upsampling teacher network obtained in Step 5;
[0057] Step 7: Input the initial point cloud data N×3 and the multi-view depth map N v ×H×W×1 of the initial point cloud data into the cross-modal feature extraction module of the upsampling student network to extract the cross-modal representation of the initial point cloud data and the multi-level detailed representation of the dense point cloud data predicted by the upsampling student network
[0058] Step 8: Input the cross-modal representation of the initial point cloud extracted by the upsampling student network into the upsampling tail of the upsampling student network to obtain the rough dense point cloud predicted by the upsampling student network and the refined dense point cloud
[0059] Step 9: The loaded upsampling teacher network only performs inference to obtain the multi-level detailed representation of the dense point cloud data extracted by the teacher network as well as the rough dense point cloud predicted by the teacher network and the refined dense point cloud
[0060] Step 10: Optimize the upsampling student network with the target value of the upsampling student network to obtain the optimized upsampling student network;
[0061] In this example, in Step 1, the initial point cloud data and the real dense point cloud data of the target object are obtained. It should be noted that the point cloud data refers to a set of discrete point data in a three-dimensional coordinate system. The data in the data set includes at least the position data of the point, such as the three-dimensional coordinate vector (x, y, z). Among them, in the real environment, the real dense point cloud data of the target object can be obtained through three-dimensional sensors such as lidar. In the laboratory environment, the corresponding real dense point cloud data can be generated from the three-dimensional mesh model through algorithms such as Poisson disk sampling. The initial point cloud data of the corresponding target object can be obtained through downsampling algorithms such as farthest point sampling or random downsampling.
[0062] In step 2, the normalized initial point cloud data and the real dense point cloud data are projected to obtain the multi-view depth map N of the initial point cloud data v ×H×W×1, the multi-view depth map N of the real dense point cloud data v ×H×W×1, and the depth map of the spliced initial point cloud data and the depth map N of the real dense data v ×H×W×2. The specific steps are as follows:
[0063] 2-1. To save computational effort and reduce the loss of point cloud information during the projection process, in this example, the coordinates of three orthogonal projection viewpoints (-v, 0, 0), (0, -v, 0), and (0, 0, -v) are used, where v represents the distance from the three projection points to the origin of the coordinates. Then, the projection planes corresponding to (-v, 0, 0), (0, -v, 0), and (0, 0, -v) are x = -v, y = -v, and z = -v. In this example, v is set to 1.4. Since the point cloud is normalized to the unit sphere, if v is too small, a large number of points will be lost after projection, and if v is too small, a large number of points will be projected together, resulting in aggregation. Therefore, it is recommended that v take a real number near 1.5;
[0064] 2-2. Calculate the distance from each point in the normalized point cloud data to the projection plane, and this distance is the depth value corresponding to each point's position in the projected depth map. For a point (x, y, z) in the normalized point cloud, the distances to the three projection planes corresponding to the projection viewpoints (-v, 0, 0), (0, -v, 0), and (0, 0, -v) are d x = |x + v|, d y = |y + v|, d z = |z + v|;
[0065] 2-3. Calculate the coordinates of the depth map obtained by projecting the point (x, y, z) in the point cloud under the viewpoints (-v, 0, 0), (0, -v, 0), and (0, 0, -v) and where, and should be integers between 1 and H, and should be integers between 1 and W, where H and W are the preset resolutions of the depth map; Take the calculation of as, where is rounded up to ensure that the coordinates on the depth map are integers. If exceeds the preset depth map resolution range, the point is directly discarded; and The calculation can rotate the point (x, y, z) around the y-axis and x-axis respectively, convert the cases where the projection points are (-v, 0, 0) and (0, -v, 0) into the case where the projection point is (0, 0, -v), and then perform the calculation. In this example, H and W are set to 128. In principle, integer values around 128 can be taken. However, if H and W are too small, the spatial resolution of the depth map will be too small, and information is likely to be lost after projection. If H and W are too large, the projected point cloud will be too sparse and not compact, making it difficult to extract features;
[0066] 2-4. The rounding step in step 2-3 may cause points with different depths to be mapped to the same 2D depth map coordinates. To solve this problem, this method calculates the harmonic mean of multiple depth values at the same position as the final depth value, making the final depth value close to the projected depth value corresponding to the point closest to the projection point in the point cloud;
[0067] 2-5. Perform the steps 2-1 to 2-4 on the normalized initial point cloud data N×3 and the normalized real dense point cloud data rN×3 respectively, and obtain the depth map N of the initial point cloud data v ×H×W×1, and the depth map N of the real dense point cloud data v ×H×W×1, and splice these two groups of depth maps along the channel dimension to obtain the spliced depth map of the initial point cloud data and the depth map of the real dense data N v ×H×W×2;
[0068] In this example, the cross-modal feature extraction module of the upsampling teacher network in step 3 and the cross-modal feature extraction module of the upsampling student network in step 7 have the same structure, both including a point cloud branch and a multi-view depth map branch (see Figure 1 ), where the point cloud branch is composed of a DenseNet structure, and the depth map branch is composed of a ResNet structure. These two structures are as Figure 2 shown. The specific steps of cross-modal feature extraction are as follows:
[0069] 3-1. Transform the initial point cloud data N×3 through a 1D convolutional layer with a kernel size of 1 to obtain N×g. Referring to PU-GAN and Dis-PU, g is set to 24 in this example;
[0070] 3-2. Transform the multi-view depth map of the point cloud data through a 2D convolution with a kernel size of 3 to obtain the feature map H×W×C of the multi-view depth map 0 . In this example, H and W are the spatial resolution of the depth map, set to 128, and C 0 is set to 16. In principle, C 0It can be set to other positive integer values, but it is recommended to set it to a smaller value because during the process of feature extraction using the ResNet structure for the depth map, the spatial resolution will decrease and the number of channels will increase. A larger C 0 will bring a relatively large computational complexity;
[0071] 3 - 3. Input the feature map H×W×C of the multi-view depth map 0 into multiple cascaded ResNet structures, where each ResNet structure is as Figure 2 shown, to obtain multi-level detailed representations In this example, L is set to 4. Since each ResNet structure will reduce the spatial resolution of the original depth map to half of the input, when L is set to 4, it is equivalent to performing 16-fold downsampling. At this time, the spatial resolution of the feature map of the depth map is already very small. It is found in the experiment that 4 times of downsampling is sufficient, but L can be tried to be set to 5 or even 6.
[0072] 3 - 4. Input the point cloud feature N×g into multiple densely connected DenseNet structures. This structure includes constructing a dynamic graph N×K×2C and densely connected EdgeConv (edge convolution) modules. The input of each EdgeConv is the concatenation of all the previous EdgeConvs and the dynamic graph. The output of the DenseNet structure is the max pooling N×(2C + ng) of the concatenation of all EdgeConvs and the dynamic graph. Denote the multiple point cloud features extracted by the DenseNet structure as Referring to PU-GAN and Dis-PU, in this example, K is set to 16 and n is set to 3; it should be noted that except for the input of the first DenseNet structure being only the point cloud feature N×g, the input of the remaining DenseNet structures is the point cloud feature N×g and the point cloud features extracted by all the previous DenseNet structures and the features after fusion with the multi-view depth map features along the channel dimension, and use a 1D convolution with a kernel size of 1 to reduce the number of channels (see Figure 1 );
[0073] 3 - 5. Use the multi-view to point (MVP) feature fusion module to fuse the output of each DenseNet structure with the output of the corresponding layer ResNet structure ;
[0074] 3 - 6. Concatenate the outputs of all MVP feature fusion modules along the channel dimension and compress the number of channels through a 1D convolution with a kernel of 1 to obtain the cross-modal representation N×C of the initial point cloud data;
[0075] In this example, the structure of the multi-view to point feature fusion module in steps 3-5 is as Figure 3 shown. This module fuses the point cloud features and depth map features of the same layer and the specific steps are as follows: The specific steps are as follows:
[0076] 3-5-1. First, through different linear layers, align the point cloud features and multi-view depth map features in the feature space. The calculation formula is as follows:
[0077]
[0078] Among them, i represents the index of the depth map view point, represents the set of projection matrices of the point cloud features, which projects the point cloud features into a query set under different views and are the projection matrices of the depth map, which project the depth map features into a key set and a value set under different views and value set d k is the number of channels of the query and the key, d v is the number of channels of the value. In this example, both d k and d v are set to 56, and can also be set to other integer values. It is recommended to set them to integer values near 64, which can ensure the model performance while having a relatively small computational cost;
[0079] 3-5-2. Then, calculate the cross-attention between Q i , K i and V i of the same view to aggregate useful depth map features for each point in the point cloud. The calculation formula is as follows:
[0080]
[0081] Among them, is the softmax activation function, is the aggregated depth map features under different views;
[0082] 3-5-3. Then, splice along the channel dimension and pass through a linear layer to integrate the aggregated depth map features under different views. The calculation formula is as follows:
[0083]
[0084] Among them, and denote the weight parameter and bias parameter of the linear layer, and Concat denotes concatenation along the channel dimension;
[0085] 3-5-4. Integrate the integrated depth map features obtained in step 3-5-3 with the point cloud features by means of attention feature fusion, and the calculation formula is as follows:
[0086]
[0087] where sigmoid(x) = 1 / (1 + e -x ) is the sigmoid activation function, denotes the point cloud features that fuse the depth map features, which are the finally extracted point cloud features of each layer, β denotes the weighting weight, and is calculated by a two-branch weight prediction sub-network obtained;
[0088] The weight prediction sub-network first adds and element-wise, then uses a batch normalization layer to reduce the difference in data distribution between different modalities. The normalized features will be fed into the global branch and the local branch. Both the global branch and the local branch are composed of two multi-layer perceptrons. In addition to the aforementioned multi-layer perceptron, the global branch also performs global pooling on the input features first, and then adds the features output by the two branches element-wise and sends them into the sigmoid activation function to obtain the above-mentioned weight β. The formula for this process is as follows:
[0089]
[0090] where MLP G and MLP L are the multi-layer perceptrons of the global branch and the local branch respectively, GAP is the global pooling operation, is the batch normalization operation, E[·] and Var[·] denote calculating the expectation and standard deviation respectively, ω and μ are two learnable parameters, ∈ = 10 -6 is a very small constant, ReLU(x) = max{0, x} is the ReLU activation function. In this example, MLP G and MLP L are both composed of two linear layers with the same structure but different parameters, batch normalization, and ReLU activation functions (see Figure 3 );
[0091] The objective numerical function of the upsampling teacher network in step 5 of this example includes: the CD distance between the rough dense point cloud predicted by the upsampling teacher network and the real dense point cloud, the CD distance between the refined dense point cloud predicted by the upsampling teacher network and the real dense point cloud, the mean square error between the rough dense point cloud predicted by the upsampling teacher network and the depth map rendered from the real dense point cloud at the same viewing angle, and the mean square error between the refined dense point cloud predicted by the upsampling teacher network and the depth map rendered from the real dense point cloud at the same viewing angle, only including the reconstruction loss terms, and the formula of the total objective numerical function of the upsampling teacher network is expressed as follows:
[0092]
[0093] Among them, α is a hyperparameter used to balance the reconstruction quality of the rough upsampled point cloud and the refined upsampled point cloud, represents the CD distance loss function, and its calculation formula is as follows:
[0094]
[0095] S 1 and S 2 represent two point clouds, |S 1 | and |S 2 | represent the number of points in the two point clouds, Φ represents a differentiable point cloud depth map renderer, which is used to measure the difference between the multi-view depth map of the reconstructed upsampled point cloud and the multi-view depth map of the real dense point cloud. In this method, the differentiable point cloud depth map renderer proposed in SpareNet is directly used as the implementation of Φ, and represent the rough upsampled point cloud and the refined upsampled point cloud output by the upsampling module of the teacher network, represents the real dense point cloud, α is a hyperparameter used to balance the reconstruction quality of the rough upsampled point cloud and the refined upsampled point cloud, ||·|| 1 and ||·|| 2 respectively represent the L 1 norm and the L 2 norm;
[0096] The process of step 7 is basically the same as that of step 3, but there are two differences. First, the input of step 7 includes the initial point cloud data N×3 and the concatenated depth map of the initial point cloud data and the depth map of the real dense data N v ×H×W×2, while the input of the upsampling teacher network in step 3 includes the point cloud data N×3 and the aligned depth map of the initial point cloud data and the depth map of the dense point cloud N v×H×W×2. Additionally, the parameters in step 7 are the parameters of the upsampling student network, and the parameters in step 3 are the parameters of the upsampling teacher network; the multi-layer depth map features extracted by the student network
[0097] The process of step 8 is basically the same as that of step 4. Both use the dense generation and multi-scale spatial correction modules in SSPU-Net as the upsampling tail module, but the parameters are different;
[0098] The upsampling network loaded in step 9 only performs inference. It should be emphasized that during the training process of the student network, every time there is an upsampling, the upsampling teacher network uses the corresponding input data for inference to obtain multi-level detailed representations of the dense point cloud data and the rough dense point cloud predicted by the teacher network and the refined dense point cloud and will be used for the calculation of the target values of the upsampling student network;
[0099] In step 10, the target values of the upsampling student network include the reconstruction loss and the knowledge distillation loss. Among them, the reconstruction loss part includes: the CD distance between the rough dense point cloud predicted by the upsampling student network and the real dense point cloud, the CD distance between the refined dense point cloud predicted by the upsampling student network and the real dense point cloud, the mean square error of the depth maps rendered from the rough dense point cloud predicted by the upsampling student network and the real dense point cloud from the same perspective, and the mean square error of the depth maps rendered from the refined dense point cloud predicted by the upsampling student network and the real dense point cloud from the same perspective. The formula for this part of the loss is as follows:
[0100]
[0101] where α is a hyperparameter used to balance the reconstruction quality of the rough upsampled point cloud and the refined upsampled point cloud, and its value is the same as that of α in the teacher network loss function, is the CD distance loss function, represents the rough upsampled point cloud predicted by the student network, represents the refined upsampled point cloud predicted by the student network;
[0102] The knowledge distillation loss function of the upsampling student network includes: the feature-based knowledge distillation loss function and the response-based knowledge distillation loss. Among them, the feature-based knowledge distillation loss includes the error loss between the multi-level detailed representations predicted by the upsampling student network and the multi-level representations of the real dense point cloud extracted by the upsampling teacher network, and the error loss between the spatial attention maps. The formula for this part of the loss function is as follows:
[0103]
[0104] Among them, is the loss between the spatial attention maps of multi-level detail representations, is the error loss between multi-level detail representations, Operator H×W×C→H×W sums the squares of all channel values of each pixel to obtain the spatial attention map, that is, C A is the number of channels of the input feature map A;
[0105] Therefore, the calculation formula of the total loss of feature-based knowledge distillation is as follows:
[0106]
[0107] Among them, λ AT and λ mimic are both hyperparameters used to balance and sizes;
[0108] The calculation formula of the knowledge distillation loss term based on response is as follows:
[0109]
[0110] Among them, and are the rough upsampled point cloud and the refined upsampled point cloud output by the upsampled student network respectively, and are the rough upsampled point cloud and the refined upsampled point cloud output by the upsampled teacher network;
[0111] The calculation formula of the total loss of the knowledge distillation term is as follows:
[0112]
[0113] Among them, λ response is a hyperparameter used to balance and sizes;
[0114] The formula of the total loss function of the upsampled student network is expressed as follows:
[0115]
[0116] Among them, λ distill is a hyperparameter used to balance and sizes;
[0117] In an experiment, the upsampling teacher network and the upsampling student network are optimized according to the above loss functions respectively, where α in the loss function is set to 0.75, and λ mimic is set to 0.1, and λ AT is set to 1, and λ response is set to 500, and λ distill is set to 0.15. Among them, the parameter α can also be set to other decimals near 0.75, and it is recommended to be greater than 0.5 and less than 1.0, which helps the network improve the reconstruction quality of the rough upsampled point cloud. The parameter λ mimic , λ AT and λ response can also be set to other small values near the given values. The setting principle is to adjust λ mimic , λ AT and λ response to make adjusted to be of the same order of magnitude as . Similarly, λ distill can also be set to other small values near 0.15, but do not set it too large, and try to make less than The upsampling teacher network and the upsampling student network adopt the same training strategy. They are trained for a total of 200 rounds, and the learning rate is multiplied by 0.7 for decay every 35 rounds. The initial learning rate is set to 0.001, and the batchsize is set to 32. In this example, the PyTorch platform is used for code implementation, and the Adam optimizer is used to optimize the model. Among them, the number of training rounds, the initial learning rate, and the decay strategy of the learning rate are all related to the batchsize. Since the upsampling teacher network and the upsampling student network in this example are both trained on a single RTX 3090 graphics card, it is found through experiments that setting the batchsize to 32 can make better use of the graphics card. Therefore, the initial learning rate is set to 0.001 to accelerate the initial optimization of the model parameters in the early stage of training with a relatively large initial learning rate. As the training progresses, the learning rate decays continuously, and finally the model parameters reach an optimal value with a relatively small learning rate.
[0118] This application conducts experiments on two classic datasets, PU-GAN and PUGeo-Net, and compares the DCD loss, CD loss, and HD loss between the upsampled target dense point cloud data and the real dense point cloud data on the corresponding test sets, as well as the mean P2FM and standard deviation P2FS of the point-to-surface (P2F) of the target dense point cloud. In the experiment, experiments are conducted for upsampling rates of 2 times, 4 times, and 8 times respectively. For the test point cloud of the PU-GAN dataset, the number of points N is 2048, and for the test point cloud of the PUGeo-Net dataset, the number of points N is 5000.
[0119] As shown in Table 1 below are the test results on the PU-GAN dataset, and Table 2 are the test results on the PUGeo-Net dataset. The present invention (CD) and the present invention (DCD) respectively represent the results obtained by the point cloud upsampling network proposed in the present invention under the CD loss function and the DCD loss function. The smaller all the metrics are, the closer they are to the true dense point cloud.
[0120]
[0121] Table 1 Experimental data table of point cloud upsampling on the PU-GAN dataset
[0122]
[0123] Table 2 Experimental data table of point cloud upsampling on the PUGeo-Net dataset
[0124] It can be seen that the point cloud upsampling network in the present invention has achieved an advanced point cloud upsampling effect. In addition to making the input sparse point cloud denser, the upsampled point cloud is evenly distributed and can maintain geometric structures and detailed information. The specific subjective upsampling effect can be seen Figure 4 .
[0125] Based on the same inventive concept, an embodiment of the present invention provides a structure-aware point cloud upsampling system based on deep learning, including an input module, a point cloud upsampling network model, and an output module; the input module is used to obtain a three-dimensional point cloud containing coordinate information and input it into the point cloud upsampling network model; the point cloud upsampling network model is used to upsample the input point cloud to obtain a point cloud that meets the upsampling magnification requirements; the output module is used to reconstruct the object coordinates based on the point cloud output by the point cloud upsampling network model; the upsampling network model includes: a point cloud depth map rendering module for rendering the input point cloud into a depth map; a cross-modal feature extraction module for extracting the cross-modal representation of the input point cloud data; an MVP cross-modal feature fusion module acting on different layers of the cross-modal feature extraction module for fusing point cloud features and multi-view depth map features; a detail estimation structure based on knowledge distillation, including a teacher network and a student network, further improving the ability of the upsampling network to reconstruct complex geometric structures and fine local details in the case of introducing additional information. The specific implementation of each module of the network model can be referred to the above method embodiment and will not be elaborated here. Those not detailed in the present invention are all prior arts.
[0126] Based on the same inventive concept, a computer system provided by an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the above-mentioned structure-aware point cloud upsampling method based on deep learning are implemented.
Claims
1. A multi-view depth map guided cross-modal distillation point cloud upsampling method, characterized by: First, the input initial point cloud is projected from multiple perspectives to obtain depth maps of the initial point cloud at different perspectives. Then, the initial point cloud and its multi-perspective depth map are input into the cross-modal feature extraction module to extract the cross-modal representation of the initial point cloud data. Then, the extracted cross-modal representation of the initial point cloud is input into the upsampling module to obtain the upsampled dense point cloud. The cross-modal feature extraction module includes a depth map branch and a point cloud branch, which are used to extract multi-view depth map features and initial point cloud features respectively, and multiple cross-modal feature fusion modules, which are used to fuse multi-view depth map features of different layers with features of initial point cloud; in the depth map branch, multiple ResNet modules are used to extract features of the multi-view depth map of the initial point cloud and gradually downsample to obtain multi-layer depth map features; in the point cloud branch, the irregular initial point cloud data is regarded as a graph signal, and the initial point cloud features of different layers are organized by graph convolution and internal dense connection and external dense connection to obtain initial point cloud features of different layers; the depth map features of the same layer are fused into the point cloud features by using the cross-modal feature fusion module; The cross-modal fusion module includes a cross-attention module and an attention feature fusion module. First, the point cloud features and multi-view depth map features are aligned in different feature spaces through linear projection. Then, cross attention is calculated in the aligned feature space to obtain the aggregated depth map features useful for each point in the initial point cloud; then the input point cloud features are fused with the aggregated depth map features through the attention feature fusion module to obtain the final fused point cloud features; Secondly, the point cloud upsampling method is a detail estimation and distillation structure composed of a teacher network and a student network, wherein the teacher network and the student network have the same structure as above; the detail estimation and distillation structure first completes the training of the teacher network, and inputs the initial point cloud data and the multi-view depth map of the spliced initial point cloud data and the multi-view depth map of the real dense point cloud data into the cross-modal feature extraction module of the upsampling teacher network to obtain the multi-level detail representation of the real dense point cloud and the upsampled point cloud of the teacher network; then, guided by the fully pre-trained teacher network, the student network with the same structure is trained, and the initial point cloud data and the multi-view depth map of the initial point cloud data are input into the cross-modal feature extraction module of the student network to obtain the multi-level detail representation of the real dense point cloud predicted by the student network, and the upsampling result of the student network; then, through feature knowledge distillation and response knowledge distillation, the multi-level detail representation of the real dense point cloud predicted by the student network is constrained to be close to the multi-level detail representation of the real dense point cloud extracted by the teacher network, and the upsampling result predicted by the student network is close to the upsampling result predicted by the teacher network.
2. The method according to claim 1, characterized in that The following steps are involved: Step 1: Obtain the initial point cloud data and the real dense point cloud data of the target object as the input data and supervision data of the upsampling network respectively, and normalize the initial point cloud data and the real dense point cloud data to obtain the normalized initial point cloud data N×3 and the normalized real dense point cloud data rN×3, where N is the number of points in the initial point cloud data and r is the upsampling rate; Step 2: Project the normalized initial point cloud data and the real dense point cloud data to obtain a multi-view depth map N of the initial point cloud data v ×H×W×1, multi-view depth map of real dense point cloud data N v ×H×W×1, as well as the depth map of the spliced initial point cloud data and the depth map of the real dense data N v ×H×W×2, where N v Represents the number of viewpoints of the multi-view depth map. Since this method uses three orthogonal projection viewpoints, N V =3, W and H represent the spatial resolution of the depth map, and the number of channels of each depth map is 1; Step 3: Initialize the parameters of the upsampling teacher network, and convert the initial point cloud data N×3 and the depth map of the spliced initial point cloud data and the depth map N of the real dense data v ×H×W×2, which is input into the cross-modal feature extraction module in the upsampling teacher network to extract the cross-modal representation of the initial point cloud data and the multi-level detail representation of the dense point cloud data Step 4: Input the cross-modal representation of the initial point cloud data extracted by the upsampling teacher network into the upsampling tail module of the upsampling teacher network to obtain the rough and dense point cloud predicted by the teacher network. and refine dense point clouds Step 5: Optimize the upsampling teacher network based on the target value of the upsampling teacher network to obtain the parameters of the optimized upsampling teacher network; Step 6: Initialize the upsampled student network, load and freeze the optimized upsampled teacher network obtained in step 5; Step 7: Initial point cloud data N×3 and multi-view depth map N of initial point cloud data v ×H×W×1 is input into the cross-modal feature extraction module of the upsampling student network to extract the cross-modal representation of the initial point cloud data and the multi-level detail representation of the dense point cloud data predicted by the upsampling student network Step 8: The cross-modal representation of the initial point cloud extracted by the upsampling student network is input into the upsampling tail of the upsampling student network to obtain the rough and dense point cloud predicted by the upsampling student network. and refine dense point clouds Step 9: The loaded upsampled teacher network is only inferred to obtain a multi-level detailed representation of the dense point cloud data extracted by the teacher network And the rough dense point cloud predicted by the teacher network and refine dense point clouds Step 10: The target value of the up-sampled student network optimizes the up-sampled student network to obtain an optimized up-sampled student network; Step 1: Obtain the initial point cloud data N×3 and the real dense point cloud data rN×3 of the target object. Point cloud data refers to a collection of a series of discrete points in three-dimensional space, where N is the number of points in the initial point cloud data and r is the upsampling rate. The data summarized in the data set includes at least the position information of each point, that is, the three-dimensional coordinates. Normalization refers to normalizing the point cloud data to the unit sphere. Specifically, first calculate the average value of the xyz coordinates of the point cloud data as the center of the sphere, and then calculate the maximum value from all points in the point cloud to the center of the sphere as the radius. Subtract the center coordinates of the point cloud data and divide by the radius to complete the normalization operation of the point cloud. Step 2 projects the normalized initial point cloud data and the real dense point cloud data to obtain a multi-view depth map N of the initial point cloud data. v ×H×W×1, multi-view depth map of real dense point cloud data N v ×H×W×1, as well as the depth map of the spliced initial point cloud data and the depth map of the real dense data N v ×H×W×2, the specific steps are as follows: 2-1. Determine the coordinates of three orthogonal projection viewpoints (-v, 0, 0), (0, -v, 0) and (0, 0, -v), where v represents the distance from the three projection points to the origin of the coordinate system. The projection planes corresponding to (-v, 0, 0), (0, -v, 0) and (0, 0, -v) are x = -v, y = -v and z = -v. Since the point cloud is normalized to the unit sphere, in order to be able to completely project the point cloud onto the projection plane, the minimum value of v should be no less than 1. 2-2. Calculate the distance from each point in the normalized point cloud data to the projection plane. This distance is the depth value of each point in the depth map after projection. The distances from the point (x, y, z) in the normalized point cloud data to the three projection planes x = -v, y = -v and z = -v are d respectively. x =|x+v|, d y =|y+v|,d z =|z+v|; 2-3. Calculate the coordinates of the depth map obtained by projecting the point (x, y, z) in the point cloud at (-v, 0, 0), (0, -v, 0) and (0, 0, -v) viewpoints as well as in, and Should be an integer between 1 and H. and It should be an integer between 1 and W, where H and W are the spatial resolutions of the preset depth map. As an example, in To round up, ensure that the coordinates on the depth map are integers. If If the point exceeds the preset depth map resolution range, the point will be discarded directly; and The calculation can be performed by rotating the point (x, y, z) around the y axis and x axis respectively, converting the projection points (-v, 0, 0) and (0, -v, 0) to the projection point (0, 0, -v) before performing the calculation; 2-4. The rounding step in step 2-3 may cause points of different depths to be mapped to the same 2D depth map coordinates. To solve this problem, this method calculates the harmonic mean of multiple depth values at the same position as the final depth value, so that the final depth value is close to the projected depth value corresponding to the point closest to the projection point in the point cloud; 2-5. Perform steps 2-1 to 2-4 on the normalized initial point cloud data N×3 and the normalized real dense point cloud data rN×3 to obtain the depth map N of the initial point cloud data. v ×H×W×1, the depth map of the real dense point cloud data N v ×H×W×1, and concatenate these two sets of depth maps along the channel dimension to obtain the depth map of the concatenated initial point cloud data and the depth map N of the real dense data v ×H×W×2; In step 3, the initial point cloud data N×3 and the depth map of the spliced initial point cloud data and the depth map N of the real dense data are combined. v ×H×W×2, which is input into the cross-modal feature extraction module in the upsampling teacher network to extract the cross-modal representation of the initial point cloud data and the multi-level detail representation of the dense point cloud data The following steps are involved: 3-1. Input the initial point cloud data into the point cloud branch of the cross-modal feature extraction module of the upsampling teacher network, and concatenate the depth map of the initial point cloud data and the depth map N of the real dense data. v ×H×W×2, input to the deep map branch of the cross-modal feature extraction module of the teacher network; 3-2. First, the depth map branch in the dual-branch cross-modal feature extraction module uses a ResNet module composed of a residual structure to extract multi-layer and multi-view depth map features. in is the feature map of the depth map output by the lth layer of the cross-modal feature extraction module, is the number of channels of the depth map feature map output by the l-th layer ResNet module; then, the point cloud branch uses the DenseNet module composed of densely connected dynamic graph convolution to extract multi-layer point cloud features in is the feature map of the point cloud output by the lth layer of the cross-modal feature extraction module, is the number of channels of the point cloud feature map output by the lth layer, and L is the number of layers in the cross-modal feature extraction module; 3-3. The input of DenseNet in step 3-2 is the concatenation of all the previous point cloud features along the channel dimension, and firstly a linear layer is used to compress the number of channels of the input point cloud features, and then the KNN algorithm is used to find K neighboring points with similar features in the feature space of the point cloud, and the edge features are calculated by subtracting the center point features from the neighboring point features, and the calculated edge features are concatenated with the center point features in the channel dimension to obtain a constructed dynamic graph N×K×2C, where N is the number of points in the point cloud, K is the number of neighboring points, and C is the number of channels of the point cloud feature map after the number of channels is reduced by a one-dimensional convolution; 3-4. Use densely connected graph convolution to transform the dynamic graph N×K×2C obtained in step 3-3. The output of each graph convolution is an incremental feature N×K×g, where g is the number of channels of the output incremental feature. The input of each graph convolution is the concatenation of the dynamic graph N×K×2C and the incremental features output by all previous graph convolutions. The final output of the DenseNet module is the concatenation of the dynamic graph N×K×2C and the incremental features N×K×g output by each graph convolution in the channel dimension N×K×(2C+ng), and the result N×(2C+ng) after the maximum pooling in the K adjacent dimensions, where n is the number of graph convolutions in DenseNet. 3-5. Multi-view depth map features of each layer in step 3-1 It is necessary to fuse the point cloud features of the same layer through multi-view point (MVP) features. In , the point cloud features and multi-view depth map features are first aligned in the feature space through different linear layers, that is, the point cloud features and multi-view depth map features are cross-attention calculated, and the calculation formula is as follows: in, i represents the index of the depth map viewpoint, A set of projection matrices representing point cloud features, projecting point cloud features into query sets at different perspectives and is the projection matrix of the depth map, which projects the depth map features into a key set under different perspectives With value collection d k is the number of channels between the query and the key, d v is the number of channels with the value; Then at the same perspective Q i , K i With V i The cross attention is calculated between them to aggregate useful depth map features for each point in the point cloud. The calculation formula is as follows: in, is the softmax activation function, It is the depth map features under different perspectives of aggregation; Afterwards, The depth map features of different perspectives are integrated by splicing along the channel dimension and passing through a linear layer. The calculation formula is as follows: in, and Represents the weight parameters and bias parameters of the linear layer, and Concat represents concatenation along the channel dimension; 3-6. The integrated depth map features obtained in step 3-5 Point cloud features The fusion is performed by attention feature fusion, and the calculation formula is as follows: Where sigmoid(x)=1 / (1+e -x ) is the sigmoid activation function, Represents the point cloud features fused with the depth map features, as the final point cloud features extracted at each layer, β represents the weighted weight, and a two-branch weight prediction subnetwork Calculated; Weight prediction subnetwork First, and The corresponding elements are added, and then a batch normalization layer is used to mitigate the difference in data distribution between different modalities. The normalized features will be sent to the global branch and the local branch. The global branch and the local branch are composed of two multi-layer perceptrons. In addition to the multi-layer perceptron mentioned above, the global branch must first perform global pooling on the input features, and then add the corresponding elements of the features output by the two branches and send them to the sigmoid activation function to obtain the above weight β. The formula of this process is as follows: Among them, MLP G With MLP L is a multi-layer perceptron of global branch and a multi-layer perceptron of local branch, GAP is a global pooling operation, is a batch normalization operation, E[·] and Var[·] represent the expected and standard deviation respectively, ω and μ are two learnable parameters, ∈=10 -6 is a very small constant, ReLU(x)=max{0,x} is the ReLU activation function; 3-7. Fuse the point cloud features obtained by steps 3-5 and 3-6 Splicing is performed along the channel dimension, and the number of channels is reduced through a one-dimensional convolution to obtain a cross-modal representation of the initial point cloud data extracted by the teacher network. This cross-modal representation will be used as the input of the upsampling tail module on the teacher network to obtain the dense point cloud reconstructed by the teacher network, and the multi-layer depth map features extracted by the teacher network are It is called the multi-layer detail representation of dense point cloud data, denoted as In step 5, the upsampling teacher network is optimized based on the target value of the upsampling teacher network to obtain the parameters of the optimized upsampling teacher network, including the following steps: 5-1. Calculate the loss function between the upsampled point cloud reconstructed by the teacher network and the real dense point cloud. The calculation formula is as follows: in, Indicates the chamfer distance, and the calculation formula is as follows: S1 and S2 represent two point clouds, |S1| and |S2| represent the number of points in the two point clouds, Φ represents a differentiable point cloud depth map renderer, which is used to measure the difference between the multi-view depth map of the reconstructed upsampled point cloud and the multi-view depth map of the real dense point cloud. This method directly uses the differentiable point cloud depth map renderer proposed in SpareNet as the implementation of Φ. and represents the coarse upsampled point cloud and the refined upsampled point cloud output by the upsampling module on the teacher network, represents the real dense point cloud, α is a hyperparameter used to balance the reconstruction quality of the coarse upsampled point cloud and the refined upsampled point cloud, ||·||1 and ||·||2 represent the L1 norm and L2 norm respectively; 5-2. Use the Adam optimizer to optimize the parameters of the teacher network until The function curve is basically stable, fluctuating up and down, and the fluctuation range is within 0.7%. convergence; The process of step 7 is basically the same as that of step 3, but there are two differences. First, the input of step 7 includes the initial point cloud data N×3 and the depth map of the spliced initial point cloud data and the depth map N of the real dense data. v ×H×W×2, and the input of the up-sampled teacher network in step 3 includes point cloud data N×3 as well as the depth map of the aligned initial point cloud data and the depth map of the dense point cloud N v ×H×W×2, secondly, the parameters in step 7 are the parameters of the upsampled student network, and the parameters in step 3 are the parameters of the upsampled teacher network; the multi-layer depth map features extracted from the student network Recorded as The process of step 8 is basically the same as that of step 4. Both use the dense generation and multi-scale spatial correction modules in SSPU-Net as upsampling modules, but the parameters are different. The upsampling network loaded in step 9 is only used for inference. It should be emphasized that every time the student network is upsampled during the training process, the upsampling teacher network must use the corresponding input data for inference to obtain a multi-level detailed representation of the dense point cloud data. And the rough dense point cloud predicted by the teacher network and refine dense point clouds as well as Will be used to calculate the target value of the upsampling student network; Step 10: Optimize the upsampled student network by the target value of the upsampled student network to obtain the optimized upsampled student network. The specific steps are as follows: 10-1. Calculate the loss function of the student network, including the reconstruction loss term And the knowledge distillation loss term The calculation formula of the reconstruction loss term is as follows: Among them, α is a hyperparameter used to balance the reconstruction quality of the coarse upsampled point cloud and the refined upsampled point cloud, and its value is the same as α in the teacher network loss function. is the CD distance loss function, represents the coarse upsampled point cloud predicted by the student network, Refined upsampled point cloud representing the student network predictions; The knowledge distillation loss function of the upsampled student network includes: feature-based knowledge distillation loss function and response-based knowledge distillation loss. The feature-based knowledge distillation loss includes the error loss between the multi-level detail representation predicted by the upsampled student network and the multi-level representation of the real dense point cloud extracted by the upsampled teacher network, as well as the error loss between the spatial attention maps. The formula of this part of the loss function is as follows: in, Operator The squares of all channel values of each pixel are summed to obtain the spatial attention map, i.e., C A is the number of channels of the input feature map A; Therefore, the calculation formula for the total loss of feature-based knowledge distillation is as follows: Among them, λ AT With λ mimic These are all hyperparameters used to adjust and size; The response-based knowledge distillation loss term is calculated as follows: in, and The rough upsampled point cloud and the refined upsampled point cloud output by the upsampled student network are respectively and The coarse upsampled point cloud and the refined upsampled point cloud output by the upsampled teacher network; Then the total loss of the knowledge distillation term is: Among them, λ response is a hyperparameter used to balance and size; Then the total loss function of the student network is: Among them, λ distill is a hyperparameter used to balance and size; 10-2. Use the Adam optimizer to optimize the parameters of the student network. The training strategy is the same as that of the teacher network until Convergence, reflected in The function curve is basically stable, fluctuating up and down, and the fluctuation range is within 0.7%.
Citation Information
Cited By
Three-dimensional point cloud completion method and system based on multi-scale structured knowledge distillation
CN120876323A