A Self-Supervised Point Cloud Completion Method Guided by Image Features

The method addresses the limitations of existing point cloud completion methods by using image feature guidance for self-supervised point cloud completion, enhancing accuracy and density without ground truth data.

CN120070224BActive Publication Date: 2025-07-15ZHEJIANG UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510541242.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-15
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing point cloud completion method combining image information fails to make full use of point cloud spatial features, directly using nearest point depth instead of inaccurate, insufficient feature fusion, and requires real and complete ground truth point cloud supervision.

Method used

The self-supervised completion method based on image feature guidance is adopted, image features are extracted through the convolution layer and the ResBlock module, point cloud feature encoding is combined with EdgeConv and GuideConv layers, and self-attention graph pooling and feature expansion are used to generate accurate dense point clouds using image features, and trained through the chamfer distance loss function.

Benefits of technology

It realizes that there is no need for real ground truth point cloud supervision, and makes full use of image guidance to generate more accurate and dense point cloud completion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070224B_ABST
    Figure CN120070224B_ABST
Patent Text Reader

Abstract

The present invention discloses a point cloud self-supervised completion method based on image feature guidance, step 1: acquiring a camera image Image Perform encoding and decoding operations to extract features. Step 2: Obtain the original sparse point cloud Points , random point sampling is performed on the original sparse point cloud to obtain N point, and for this N Points construct feature tensors, denoted as Points 1; Step 3: For the feature tensor constructed in step 2 Points 1. Perform feature encoding to obtain features; Step 4: Expand the features obtained in step 3 to obtain point cloud features; Step 5: Point cloud The feature uses the fully connected layer to reconstruct the point cloud and obtain the completed point cloud feature, which is recorded as Points 2; Step 6: When training the network, the original sparse point cloud needs to be Points Sampling the farthest point N point, and for this N Points construct feature tensors, denoted as Points 3. Calculation Points 2 and Points The chamfer distance between 3 is used as the loss function during network training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and specifically to a deep learning method for guiding point cloud self-supervised completion using image features. Background Art

[0002] With the rapid development of fields such as autonomous driving and mobile robots, cameras and lidars have become the most widely used sensors. The point cloud data obtained by lidars can accurately represent depth information, but the resolution of point clouds is often low, and there are cases of missing parts of point clouds due to occlusion and other situations, and point clouds need to be completed in terms of precise measurement and the like. The images obtained by cameras, on the other hand, have rich information such as color and texture, and have a high resolution. Therefore, point cloud completion combined with image information has become one of the current research hotspots.

[0003] Currently, many scholars have proposed point cloud completion methods that combine image information. Among them, the technical solution relatively close to the present invention is as follows: The literature (Jie Tang, FeiPeng Tian, Wei Feng, et al. Learning Guided Convolutional Network for Depth Completion[J]. IEEE Transactions on Image Processing, 2021, 30:1116-1129.) proposed to fuse image features into different stages of an encoder with sparse depth features to guide sparse depth upsampling. This method is implemented based on the depth map corresponding to the point cloud rather than directly based on the point cloud, resulting in the loss of point cloud spatial features; The literature (Jiarong Wang, Ming Zhu, Bo Wang, et al. KDA3D: Key-Point Densification and Multi-Attention Guidance for 3D Object Detection[J]. Remote Sensing, 2020, 12(11):1895-1921.) proposed a key-point densification operation, projecting each point cloud in the foreground area as a key-point cloud into the image to obtain the key pixels corresponding to each key-point cloud. Then, the neighboring pixels of each key pixel are clustered to search for extended pixels close to the key pixel. The extended pixels are back-projected into the 3D space to obtain a pseudo point cloud, where the depth information of the pseudo point cloud is directly replaced by the depth information of its corresponding key-point cloud; Patent Application No.: 202210425289.6, Title: Point Cloud Upsampling Method, Device, Computer Device and Storage Medium, uses a two-dimensional panoramic image semantic segmentation map and coordinate transformation relationship to obtain a point cloud data semantic segmentation map, and performs feature extraction and feature extension on the point cloud data to generate extended points corresponding to each point cloud to obtain a dense point cloud. This method only utilizes the semantic information of the image and does not fully play the role of image-guided point cloud completion; The literature (He Xing, Zhu Zhe, Yan Xuefeng, et al. Cross-Modal Transformer Point Cloud Completion Algorithm Integrating Image Information[J]. Journal of Computer-Aided Design & Computer Graphics, 2024, 36:1-9.) proposed to extract point cloud features and image features using a point cloud branch and an image branch respectively, fuse the point cloud features and image features through a feature fusion module to generate a full-resolution point cloud, and then use the real and complete point cloud to supervise the completed point cloud. In this method, the feature fusion module directly concatenates and combines the point cloud features and image features. Due to the heterogeneity of image and point cloud data, this simple feature fusion method is also difficult to fully exploit the potential of image-guided point cloud completion.

[0004] In summary, the current point cloud completion methods that combine image information have the following deficiencies: (1) Completion based on depth maps does not fully utilize the spatial features of the point cloud; (2) The depth information of the completed point cloud is directly replaced by the depth of neighboring points, which is inaccurate; (3) The point cloud features and image features are directly concatenated and fused, without fully exploiting the potential of image-guided point cloud completion; (4) Real and complete ground truth point clouds are required to supervise the completion. Summary of the Invention

[0005] In view of the above problems existing in the existing point cloud completion methods that combine image information, the present invention proposes a self-supervised point cloud completion method guided by image features.

[0006] A self-supervised point cloud completion method guided by image features includes the following steps:

[0007] Step 1: Perform encoding and decoding operations on the acquired camera image Image to extract image features. The encoding part uses 1 convolutional layer and L ResBlock modules in series, and the decoding part uses L -1 transposed convolutional layers in series. Add the output features of the l th ResBlock module in the encoding part and the output features of the L - l th transposed convolutional layer in the decoding part to obtain the feature F l , where l = 1, 2, …, L -1;

[0008] Step 2: Obtain the original sparse point cloud Points , randomly sample N points from the original sparse point cloud, and construct a N -dimensional feature tensor for these points, denoted as Points 1, where N represents the number of points, and C represents the feature dimension of each point;

[0009] Step 3: Perform feature encoding on the feature tensor Points 1 constructed in Step 2 to obtain -dimensional features, where represents the feature dimension of each point after feature encoding;

[0010] Step 4: Expand the -dimensional features obtained in Step 3 to obtain -dimensional point cloud features, where r represents a preset feature expansion rate. Denote the feature dimension of each point after feature expansion;

[0011] Step 5: Use a fully connected layer to reconstruct the point cloud with the -dimensional features obtained in Step 4, and obtain the -dimensional completed point cloud features, denoted as Points 2, where rN denotes the number of points after completion;

[0012] Step 6: During network training, it is necessary to perform farthest point sampling on the original sparse point cloud Points to obtain N points, and construct a N -dimensional feature tensor for these points, denoted as Points 3, then calculate the chamfer distance between Points 2 and Points 3 according to Equation (1) as the loss function during network training, where denotes Points the chamfer distance value between 3 and Points 2, denotes a preset weight coefficient, p denotes Points the points in 2, denotes Points the points in 3;

[0013] .

[0014] Furthermore, the process of feature encoding in Step 3 is as follows:

[0015] 3.1): The network for feature encoding uses 1 EdgeConv layer and L -1 GuideConv layers in series. Whether to use a self-attention graph pooling layer for pooling operations can be selected after each GuideConv layer;

[0016] 3.2): Concatenate the output features of the EdgeConv layer and each GuideConv layer in the dimension of the features, and perform feature dimensionality reduction processing on the concatenated features through a SharedMLP layer to obtain -dimensional features.

[0017] In 3.1), the GuideConv layer is a depthwise separable convolution form of the EdgeConv layer, and both its depth convolution kernel parameters and pointwise convolution kernel parameters are guided and generated by the image features in Step 1, where F l guides the lThe convolution kernel parameters of the GuideConv layer are generated. The specific convolution kernel parameter generation process is as follows:

[0018] 3.1.1): Using standard convolutional layer processing F l get The characteristic of the dimension is l The depth convolution kernel parameters of the GuideConv layer, where N and C l Respectively represent l The number of points to be processed in the GuideConv layer and the feature dimensions of the points. When performing deep convolution on each feature dimension N Parameters are not shared between points;

[0019] 3.1.2) Using average pooling F l get dimensional features, where M l express F l The number of channels;

[0020] 3.1.3) For the results obtained in 3.1.2) The dimension features are obtained by using the fully connected layer The characteristic of the dimension is l The point-by-point convolution kernel parameters of the GuideConv layer, where C l+1 Indicates l The feature dimension of the points output by the GuideConv layer.

[0021] Furthermore, the feature expansion process in step 4 is as follows:

[0022] 4.1): Replace the Feature copy of dimension r For copies r The characteristics of the two r The SharedMLP network branches process the r share dimensional perturbation coefficient, where each SharedMLP network branch is composed of L 1 SharedMLP layer in series;

[0023] 4.2): r share The perturbation coefficients of the dimension are spliced into 4.1) r share In the copy feature of the dimension, we get r share dimensional features;

[0024] 4.3): For the r copies dimensional features, respectively use r EdgeConv network branches for processing to obtain r copies dimensional point cloud features. Among them, each EdgeConv network branch consists of L 2 cascaded EdgeConv layers.

[0025] The beneficial effects of the present invention are as follows:

[0026] By using the point cloud completion method based on image feature guidance of the present invention, the spatial features of the point cloud can be fully utilized, and the potential of image-guided point cloud completion can be fully exerted. It is not necessary to use the ground truth point cloud for supervised completion, and a more accurate and dense point cloud completion result can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Schematic diagram of the camera image used in this embodiment;

[0028] Figure 2 Schematic diagram of the original sparse point cloud used in this embodiment;

[0029] Figure 3 Schematic diagram of the model structure for encoding point cloud features guided by image features used in this embodiment;

[0030] Figure 4 Schematic diagram of the feature expansion model structure used in this embodiment;

[0031] Figure 5 Schematic diagram of the point cloud reconstruction model structure in this embodiment;

[0032] Figure 6 Schematic diagram of the point cloud reconstruction effect example in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The following will elaborate on the specific implementation manners of the point cloud completion method based on image feature guidance of the present invention in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] A point cloud completion method based on image feature guidance of the present invention includes the following steps:

[0035] Step 1: Perform encoding and decoding operations on the acquired camera image Image to extract image features as Figure 1 shown. The encoding part uses 1 convolutional layer and LA series of L consecutive deconvolution layers are used in the decoding part. The l output features of the L - l th ResBlock module in the encoding part and the output features of the F l th deconvolution layer in the decoding part are added together to obtain the feature l where L = 1, 2, …, L -1. In this embodiment, Image is set to 5, and the camera image used is Figure 1 as shown;

[0036] Step 2: Obtain the original sparse point cloud Points As shown Figure 2 , random point sampling is performed on the original sparse point cloud to obtain N points, and a N -dimensional feature tensor is constructed for these points, denoted as Points 1, where N represents the number of points, and C represents the feature dimension of each point. In this embodiment, the coordinate values of each point ( x , y, z ) are used as its features, and its feature dimension C is 3, N is set to 2048, and the original sparse point cloud obtained is Points as shown Figure 2 ;

[0037] Step 3: Perform feature encoding on the feature tensor Points 1 constructed in Step 2 to obtain -dimensional features, where represents the feature dimension of each point after feature encoding;

[0038] The process of feature encoding in Step 3 is as follows:

[0039] 3.1): The network for feature encoding uses 1 EdgeConv layer and L -1 consecutive GuideConv layers. Whether to use a self-attention graph pooling layer (SAGpool) for pooling operation can be selected after each GuideConv layer;

[0040] In 3.1), the GuideConv layer is a depthwise separable convolution form of the EdgeConv layer, and both its depth convolution kernel parameters and pointwise convolution kernel parameters are guided and generated by the image features in Step 1. Among them, Fl Generate the convolution kernel parameters of the l th GuideConv layer. The specific process of generating the convolution kernel parameters is as follows:

[0041] 3.1.1): Process using a standard convolutional layer F l Obtain -dimensional features as the depth convolution kernel parameters of the l th GuideConv layer. Among them, N and C l respectively represent the number of points to be processed and the feature dimension of the points in the l th GuideConv layer. When performing depth convolution on each feature dimension, N parameters are not shared among

[0042] points; F l 3.1.2) Process using average pooling Obtain M l -dimensional features, where F l represents the number of channels of

[0043] 3.1.3) For the -dimensional features obtained in 3.1.2), process using a fully connected layer to obtain -dimensional features as the pointwise convolution kernel parameters of the l th GuideConv layer. Among them, C l+1 represents the feature dimension of the points output by the l th GuideConv layer.

[0044] 3.2): Concatenate the output features of the EdgeConv layer and each GuideConv layer in the feature dimension, and perform feature dimensionality reduction processing on the concatenated features through a SharedMLP layer to obtain -dimensional features. The model structure using image features to guide point cloud feature encoding in this embodiment is as Figure 3 shown.

[0045] Step 4: Expand the -dimensional features obtained in Step 3 to obtain -dimensional point cloud features, where r represents the preset feature expansion rate, represents the feature dimension of each point after feature expansion. In this embodiment, r is set to 1. The feature expansion model structure is as Figure 4 shown;

[0046] The process of feature expansion in step 4 is as follows:

[0047] 4.1): Copy the -dimensional features in step 3 r times. For each copy of the r features, use r SharedMLP network branches to process and obtain r copies of -dimensional perturbation coefficients. Among them, each SharedMLP network branch is composed of L 1 SharedMLP layer in series. In this embodiment, L 1 is set to 3;

[0048] 4.2): Concatenate the r copies of -dimensional perturbation coefficients to the r copies of -dimensional copied features in 4.1) respectively to obtain r copies of -dimensional features;

[0049] 4.3): For the r copies of -dimensional features in 4.2), use r EdgeConv network branches to process and obtain r copies of -dimensional point cloud features. Among them, each EdgeConv network branch is composed of L 2 EdgeConv layers in series. In this embodiment, L 2 is set to 4.

[0050] Step 5: Use a fully connected layer to perform point cloud reconstruction on the -dimensional features obtained in step 4 to obtain -dimensional completed point cloud features, denoted as Points 2, where rN represents the number of points after completion. The point cloud reconstruction model structure in this embodiment is as shown in Figure 5 and the effect of point cloud reconstruction is as shown in Figure 6 ;

[0051] Step 6: During network training, it is necessary to perform farthest point sampling on the original sparse point cloud Points to obtain N points, and construct a N -dimensional feature tensor for these points, denoted as Points 3, and then calculate Points 2 and PointsThe chamfer distance between 3 is used as the loss function during network training:

[0052]

[0053] where denotes Points the chamfer distance value between 3 and Points 2, denotes a preset weight coefficient, p denotes Points a point in 2, denotes Points a point in 3. In this embodiment, the public dataset ShapeNet-ViPC is used for training and testing. It is set to 0.25. The experimental data comparing the test results with existing point cloud completion methods are shown in Table 1;

[0054] Table 1 Comparison results of the L2 chamfer distance values between the completion results of the point cloud completion method and the real point cloud (unit: x10 3 ):

[0055] ;

[0056] The experimental data results in Table 1 show that the self-supervised point cloud method proposed in this invention patent achieves the minimum L2 chamfer distance value, far superior to other self-supervised and weakly supervised methods.

[0057] The above embodiments are only the preferred embodiments of the present invention and do not limit the technical solutions of the present invention. Any technical solutions that can be achieved on the basis of the above embodiments without creative labor shall be regarded as falling within the scope of the patent protection of the present invention.

Claims

1. A self-supervised point cloud completion method guided by image features, characterized in that, It includes the following steps: Step 1: Perform encoding and decoding operations on the acquired camera image Image to extract image features. The encoding part consists of 1 convolutional layer and L ResBlock modules connected in series, and the decoding part consists of L-1 transposed convolutional layers connected in series. Add the output features of the l-th ResBlock module in the encoding part and the output features of the (L-l)-th transposed convolutional layer in the decoding part to obtain the feature F l , where l = 1, 2, …, L-1; Step 2: Obtain the original sparse point cloud Points, randomly sample N points from the original sparse point cloud, and construct an N×C-dimensional feature tensor for these N points, denoted as Points1, where N represents the number of points and C represents the feature dimension of each point; Step 3: Perform feature encoding on the feature tensor Points1 constructed in Step 2 to obtain an N×C'-dimensional feature, where C' represents the feature dimension of each point after feature encoding; The process of feature encoding in Step 3 is as follows: 3.1): The network for feature encoding consists of 1 EdgeConv layer and L - 1 GuideConv layers connected in series; 3.1) The GuideConv layer in it is the depthwise separable convolution form of the EdgeConv layer, and both its depth convolution kernel parameters and pointwise convolution kernel parameters are generated guided by the image features in step 1, where F l Guide the generation of the convolution kernel parameters of the l-th GuideConv layer. The specific process of generating the convolution kernel parameters is as follows: 3.1.1): Process F using a standard convolutional layer l Obtain a 1×N×C l -dimensional feature as the depth convolution kernel parameter of the l-th GuideConv layer, where N and C l respectively represent the number of points to be processed in the l-th GuideConv layer and the feature dimension of the points. When performing depth convolution on each feature dimension, the parameters are not shared among the N points; 3.1.2) Average pooling is used to process F l to obtain features of 1×1×M l dimensions, where M l represents the number of channels of F l ; 3.1.3) For the 1×1×M l -dimensional features obtained in 3.1.2), a fully connected layer is used for processing to obtain 1×1×C l ×C l+1 -dimensional features as the pointwise convolution kernel parameters of the l-th GuideConv layer, where C l+1 represents the feature dimension of the points output by the l-th GuideConv layer; 3.2): Concatenate the output features of the EdgeConv layer and each GuideConv layer in the dimension of the features, and perform feature dimensionality reduction on the concatenated features through a SharedMLP layer to obtain an N×C'-dimensional feature; Step 4: Expand the N×C'-dimensional feature obtained in Step 3 to obtain an rN×C''-dimensional point cloud feature, where r represents a preset feature expansion rate and C'' represents the feature dimension of each point after feature expansion; The process of feature expansion in Step 4 is as follows: 4.1): Copy the N×C'-dimensional feature in Step 3 r times, and use r SharedMLP network branches to process the r copied features respectively to obtain r N×2-dimensional perturbation coefficients, where each SharedMLP network branch consists of L1 SharedMLP layers connected in series; 4.2): Concatenate the r N×2-dimensional perturbation coefficients to the r N×C'-dimensional copied features in 4.1) respectively to obtain r N×(C'+2)-dimensional features; 4.3): Use r EdgeConv network branches to process the r N×(C'+2)-dimensional features in 4.2) respectively to obtain r N×C''-dimensional point cloud features, where each EdgeConv network branch consists of L2 EdgeConv layers connected in series; Step 5: Use a fully connected layer to perform point cloud reconstruction on the rN×C''-dimensional feature obtained in Step 4 to obtain an rN×C-dimensional completed point cloud feature, denoted as Points2, where rN represents the number of completed points; Step 6: During network training, it is necessary to perform farthest point sampling on the original sparse point cloud Points to obtain N points, and construct an N×C-dimensional feature tensor for these N points, denoted as Points3, and calculate the chamfer distance between Points2 and Points3 as the loss function during network training; The chamfer distance between Points2 and Points3 is calculated according to Equation (1) in Step 6: where CD(Point3,Point2) represents the chamfer distance value between Points3 and Points2, λ represents a preset weight coefficient, p represents a point in Points2, and represents a point in Points3.

Citation Information

Patent Citations

  • Point cloud up-sampling method and device, computer equipment and storage medium

    CN114898094A

  • Deep three-dimensional point cloud classification network construction method based on competitive attention fusion

    CN112990336A

  • Denoising method based on multiscale distribution score for point cloud

    US20240296528A1