Point cloud self-supervision completion method based on image feature guidance

By introducing the technology of image feature encoding and point cloud feature expansion into the point cloud completion method, combined with the chamfer distance loss function learned by self-supervised learning, the problem of inaccurate point cloud spatial feature loss and depth information in the existing method is solved, and more accurate and dense point cloud completion is achieved.

CN120070224AActive Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510541242.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing point cloud completion method combining image information has problems such as point cloud spatial feature loss, inaccurate depth information, insufficient feature fusion, and the need for real and complete ground truth point clouds for supervision and completion.

Method used

A point cloud self-supervised completion method based on image feature guidance is proposed. Image features are extracted by encoding and decoding the camera image, and point cloud feature encoding and expansion is combined with EdgeConv and GuideConv layers, point cloud reconstruction is used to use the full connection layer, and network training is performed by chamfering distance as a loss function.

Benefits of technology

This method can make full use of the spatial characteristics of point clouds and fully utilize the potential of image-guided point cloud completion, without the need for ground truth point cloud supervision, and obtain more accurate and dense point cloud completion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070224A_ABST
    Figure CN120070224A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud self-supervision completion method based on image feature guidance, which comprises the following steps: step 1, carrying out encoding and decoding operation on an acquired camera image to extract features, step 2, acquiring an original sparse point cloud Point, carrying out random point sampling on the original sparse point cloud to obtain N points, constructing a feature tensor for the N points, and recording the feature tensor as Points1; step 3, carrying out feature coding on the feature tensor Points1 constructed in the step 2 to obtain features; 4, performing feature expansion on the features obtained in the step 3 to obtain point cloud features; 5, performing point cloud reconstruction on the point cloud features obtained in the step 4 by using a full connection layer to obtain complemented point cloud features, and recording the complemented point cloud features as Points2; and step 6, during network training, carrying out farthest point sampling on the original sparse point cloud Points to obtain N points, constructing a feature tensor for the N points, recording the feature tensor as Points3, and calculating a chamfering distance between the Points2 and the Points3 as a loss function during network training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and specifically to a deep learning method for guiding point cloud self-supervised completion using image features. Background Art

[0002] With the rapid development of fields such as autonomous driving and mobile robots, cameras and lidars have become the most widely used sensors. The point cloud data obtained by lidars can accurately represent depth information, but the resolution of point clouds is often low, and there are cases of missing partial point clouds due to occlusion and other situations, and point clouds need to be completed in terms of precise measurement. The images obtained by cameras, on the other hand, have rich information such as color and texture, and have a high resolution. Therefore, combining image information for point cloud completion has become one of the current research hotspots.

[0003] Currently, many scholars have proposed point cloud completion methods that combine image information. Among them, the technical solutions relatively close to the present invention are as follows: The literature (Jie Tang, FeiPeng Tian, Wei Feng, et al. Learning Guided Convolutional Network for Depth Completion[J]. IEEE Transactions on Image Processing, 2021, 30:1116-1129.) proposed to fuse image features into different stages of an encoder with sparse depth features to guide sparse depth upsampling. This method is implemented based on the depth map corresponding to the point cloud rather than directly based on the point cloud, resulting in the loss of point cloud spatial features; The literature (Jiarong Wang, Ming Zhu, Bo Wang, et al. KDA3D: Key-Point Densification and Multi-Attention Guidance for 3D Object Detection[J]. Remote Sensing, 2020, 12(11):1895-1921.) proposed a key-point densification operation, where each point cloud in the foreground area is projected into the image as a key-point cloud to obtain the key pixels corresponding to each key-point cloud. Then, the neighboring pixels of each key pixel are clustered to search for extended pixels close to the key pixel. The extended pixels are back-projected into the 3D space to obtain a pseudo point cloud, where the depth information of the pseudo point cloud is directly replaced by the depth information of its corresponding key-point cloud; Patent Application No.: 202210425289.6, Title: Point Cloud Upsampling Method, Device, Computer Equipment and Storage Medium, uses a two-dimensional panoramic image semantic segmentation map and coordinate transformation relationship to obtain a point cloud data semantic segmentation map, and performs feature extraction and feature expansion on the point cloud data to generate extended points corresponding to each point cloud to obtain a dense point cloud. This method only utilizes the semantic information of the image and does not fully play the role of image-guided point cloud completion; The literature (He Xing, Zhu Zhe, Yan Xuefeng, et al. Cross-Modal Transformer Point Cloud Completion Algorithm Incorporating Image Information[J]. Journal of Computer-Aided Design & Computer Graphics, 2024, 36:1-9.) proposed to separately extract point cloud features and image features using a point cloud branch and an image branch, fuse the point cloud features and image features through a feature fusion module to generate a full-resolution point cloud, and then use the real and complete point cloud to supervise the completed point cloud. In this method, the feature fusion module directly concatenates and combines the point cloud features and image features. Due to the heterogeneity of image and point cloud data, this simple feature fusion method is also difficult to fully exploit the potential of image-guided point cloud completion.

[0004] In summary, the current point cloud completion methods that combine image information have the following deficiencies: (1) Completion based on depth maps does not fully utilize the spatial features of the point cloud; (2) The depth information of the completed point cloud is directly replaced by the depth of neighboring points, which is inaccurate; (3) The point cloud features and image features are directly concatenated and fused, without fully exploiting the potential of image-guided point cloud completion; (4) Real and complete ground truth point clouds are required to supervise the completion. Summary of the Invention

[0005] In view of the above problems existing in the existing point cloud completion methods that combine image information, the present invention proposes a self-supervised point cloud completion method guided by image features.

[0006] A self-supervised point cloud completion method guided by image features includes the following steps: Step 1: Perform encoding and decoding operations on the acquired camera image Image to extract image features. The encoding part uses 1 convolutional layer and L ResBlock modules in series, and the decoding part uses L -1 transposed convolutional layers in series. Add the output features of the l th ResBlock module in the encoding part and the output features of the L - l th transposed convolutional layer to obtain the feature F l , where l = 1, 2, …, L -1; Step 2: Obtain the original sparse point cloud Points , randomly sample N points from the original sparse point cloud, and construct a N -dimensional feature tensor for these points, denoted as Points 1 , where N represents the number of points, and C represents the feature dimension of each point; Step 3: Perform feature encoding on the feature tensor Points 1 constructed in Step 2 to obtain -dimensional features, where represents the feature dimension of each point after feature encoding; Step 4: Expand the -dimensional features obtained in Step 3 to obtain -dimensional point cloud features, where r represents a preset feature expansion rate, and represents the feature dimension of each point after feature expansion; Step 5: Use a fully connected layer to perform point cloud reconstruction on the -dimensional features to obtain -dimensional completed point cloud features, denoted as Points 2 , where rN represents the number of completed points; Step 6: During network training, it is necessary to perform farthest point sampling on the original sparse point cloud Points to obtain N points, and construct a N -dimensional feature tensor for these points, denoted as Points 3 , and then calculate the chamfer distance between Points 2 and Points 3 according to Equation (1) as the loss function during network training, where represents Points 3 and Points 2 the chamfer distance value between, represents a preset weight coefficient, p represents Points 2 the points in, represents Points 3 the points in; .

[0007] Furthermore, the process of feature encoding in Step 3 is as follows: 3.1): The network for feature encoding uses 1 EdgeConv layer and L -1 GuideConv layers in series, and whether to use a self-attention graph pooling layer for pooling operation can be selected after each GuideConv layer; 3.2): Concatenate the output features of the EdgeConv layer and each GuideConv layer in the feature dimension, and perform feature dimensionality reduction processing on the concatenated features through a SharedMLP layer to obtain -dimensional features.

[0008] The GuideConv layer in 3.1) is a depthwise separable convolution form of the EdgeConv layer, and its depth convolution kernel parameters and pointwise convolution kernel parameters are both guided and generated by the image features in Step 1, where F l guides the lGeneration of the convolution kernel parameters of a GuideConv layer. The specific process of generating the convolution kernel parameters is as follows: 3.1.1): Process using a standard convolution layer F l Obtain -dimensional features as the depth convolution kernel parameters of the l -th GuideConv layer. Among them, N and C l respectively represent the number of points to be processed and the feature dimension of the points in the l -th GuideConv layer. When performing depth convolution on each feature dimension, N parameters are not shared among points; 3.1.2) Process using average pooling F l Obtain -dimensional features, where M l represents F l the number of channels; 3.1.3) For the -dimensional features obtained in 3.1.2), process using a fully connected layer to obtain -dimensional features as the pointwise convolution kernel parameters of the l -th GuideConv layer. Among them, C l+1 represents the feature dimension of the points output by the l -th GuideConv layer.

[0009] Furthermore, the process of feature expansion in step 4 is as follows: 4.1): Copy the -dimensional features in step 3 r times. For the r copies of features, process them respectively using r SharedMLP network branches to obtain r copies of -dimensional perturbation coefficients. Among them, each SharedMLP network branch is composed of L 1 SharedMLP layers in series; 4.2): Concatenate the r copies of -dimensional perturbation coefficients to the r copies of -dimensional copied features in 4.1) respectively to obtain r copies of -dimensional features; 4.3): For the r copies in 4.2) The features of the r dimensions are processed by r copies of dimensional point cloud features are obtained. Among them, each EdgeConv network branch consists of L 2 EdgeConv layers in series.

[0010] The beneficial effects of the present invention are as follows: By using the point cloud completion method based on image feature guidance of the present invention, the spatial features of the point cloud can be fully utilized, and the potential of image-guided point cloud completion can be fully exerted. It is not necessary to use the ground truth point cloud for supervised completion, and a more accurate and dense point cloud completion result can be obtained. Brief Description of the Drawings

[0011] Figure 1 Schematic diagram of the camera image used in this embodiment; Figure 2 Schematic diagram of the original sparse point cloud used in this embodiment; Figure 3 Schematic diagram of the model structure for encoding point cloud features guided by image features used in this embodiment; Figure 4 Schematic diagram of the feature expansion model structure used in this embodiment; Figure 5 Schematic diagram of the point cloud reconstruction model structure used in this embodiment; Figure 6 Schematic diagram of the point cloud reconstruction effect example used in this embodiment. Detailed Embodiment

[0012] The following combines the embodiments to elaborate in detail the specific implementation manner of the point cloud completion method based on image feature guidance of the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0013] A point cloud completion method based on image feature guidance of the present invention includes the following steps: Step 1: Perform encoding and decoding operations on the acquired camera image Image to extract image features as Figure 1 shown. The encoding part uses 1 convolutional layer and L ResBlock modules in series. The decoding part uses L -1 transposed convolutional layers in series. Add the output features of the l th ResBlock module in the encoding part and the output features of the L - l th transposed convolutional layer in the decoding part to obtain the feature Fl , where l = 1, 2, …, L -1. In this embodiment, L is set to 5, and the camera images used Image are as Figure 1 shown; Step 2: Obtain the original sparse point cloud Points As Figure 2 shown, perform random point sampling on the original sparse point cloud to obtain N points, and construct a N -dimensional feature tensor for these points, denoted as Points 1 , where N represents the number of points, C represents the feature dimension of each point. In this embodiment, the coordinate values of each point ( x , y, z ) are used as its features, and its feature dimension C is 3, N is set to 2048, and the original sparse point cloud obtained Points is as Figure 2 shown; Step 3: Perform feature encoding on the feature tensor Points 1 constructed in Step 2 to obtain -dimensional features, where represents the feature dimension of each point after feature encoding; The process of feature encoding in Step 3 is as follows: 3.1): The network for feature encoding uses 1 EdgeConv layer and L -1 GuideConv layers in series. Whether to use a self-attention graph pooling layer (SAGpool) for pooling operation can be selected after each GuideConv layer; The GuideConv layer in 3.1) is a depthwise separable convolution form of the EdgeConv layer, and both its depth convolution kernel parameters and pointwise convolution kernel parameters are guided and generated by the image features in Step 1. Among them, F l guides the generation of the convolution kernel parameters of the l th GuideConv layer. The specific process of generating convolution kernel parameters is as follows: 3.1.1): Use a standard convolution layer to process F l to obtain -dimensional features as the depth convolution kernel parameters of the l th GuideConv layer. Among them, N andC l respectively represent the number of points to be processed and the feature dimension of the points in the l th GuideConv layer. When performing depth convolution on each feature dimension, N parameters are not shared among points; 3.1.2) Use average pooling to process F l to obtain -dimensional features, where M l represents F l the number of channels; 3.1.3) Use a fully connected layer to process the -dimensional features obtained in 3.1.2) to obtain -dimensional features as the pointwise convolution kernel parameters of the l th GuideConv layer, where C l+1 represents the l th GuideConv layer's output point feature dimension.

[0014] 3.2): Concatenate the output features of the EdgeConv layer and each GuideConv layer in the feature dimension, and perform feature dimensionality reduction on the concatenated features through a SharedMLP layer to obtain -dimensional features. In this embodiment, the model structure using image features to guide point cloud feature encoding is as Figure 3 shown.

[0015] Step 4: Expand the -dimensional features obtained in Step 3 to obtain -dimensional point cloud features, where r represents the preset feature expansion rate, represents the feature dimension of each point after feature expansion. In this embodiment, r is set to 1, and the feature expansion model structure is as Figure 4 shown; The process of feature expansion in Step 4 is as follows: 4.1): Copy the -dimensional features in Step 3 r times, and respectively process the copied r times of features through r SharedMLP network branches to obtain r times of -dimensional perturbation coefficients. Among them, each SharedMLP network branch consists of L 1A series of SharedMLP layers are cascaded. In this embodiment, L 1 it is set to 3; 4.2): Concatenate r copies of the perturbation coefficients of dimension r copies of the copied features of dimension r copies to obtain features of dimension 4.3): For the r copies of the features of dimension r process them respectively using r copies of EdgeConv network branches to obtain L 2 copies L 2 of point cloud features of dimension

[0016] Step 5: Use a fully connected layer to reconstruct the point cloud from the -dimensional features obtained in Step 4 to obtain -dimensional completed point cloud features, denoted as Points 2 , where rN represents the number of points after completion. In this embodiment, the point cloud reconstruction model structure is as shown in Figure 5 , and the effect of point cloud reconstruction is as shown in Figure 6 ; Step 6: During network training, it is necessary to perform farthest point sampling on the original sparse point cloud Points to obtain N points, and construct a N -dimensional feature tensor for these points, denoted as Points 3 , and then calculate the chamfer distance between Points 2 and Points 3 according to Equation (1) as the loss function during network training:

[0017] where represents Points 3 the chamfer distance value between Points 2 and represents the preset weight coefficient, p representsPoints 2 The points in represent Points 3 the points in. In this embodiment, the public dataset ShapeNet-ViPC is used for training and testing. It is set to 0.25. The experimental data comparing the test results with existing point cloud completion methods are shown in Table 1; Table 1 Comparison results of the L2 chamfer distance values between the completion results of the point cloud completion method and the real point cloud (unit: x10 3 ): ; The experimental data results in Table 1 show that the self-supervised point cloud method proposed in this invention patent achieves the minimum L2 chamfer distance value, far superior to other self-supervised and weakly supervised methods.

[0018] The above embodiments are only the preferred embodiments of the present invention and do not limit the technical solutions of the present invention. Any technical solutions that can be achieved on the basis of the above embodiments without creative labor shall be regarded as falling within the scope of the patent rights of this invention patent.

Claims

1. A point cloud self-supervised completion method based on image feature guidance, characterized in that: The steps include: Step 1: Acquire the camera image Image Perform encoding and decoding operations to extract image features. The encoding part uses 1 convolutional layer and L ResBlock modules are connected in series, and the decoding part uses L -1 deconvolution layer in series, the encoding part l The output features of the ResBlock module and the decoding part L - l The output features of the deconvolution layers are added to obtain the feature F l ,in l =1, 2, …, L -1; Step 2: Get the original sparse point cloud Points , random point sampling is performed on the original sparse point cloud to obtain N point, and for this N Point construction The feature tensor of dimension is denoted as Points 1, where N represents the number of points, C Represents the feature dimension of each point; Step 3: For the feature tensor constructed in step 2 Points 1 Perform feature encoding and obtain dimensional features, where Represents the feature dimension of each point after feature encoding; Step 4: Take the data obtained in step 3 and The feature expansion of the dimension is carried out to obtain dimensional point cloud features, where r represents the preset feature expansion rate, Represents the feature dimension of each point after feature expansion; Step 5: Take the data obtained in step 4 and The feature of the dimension is used to reconstruct the point cloud using the fully connected layer, and we get The dimension of the completed point cloud feature is recorded as Points 2, where rN Indicates the number of completed points; Step 6: When training the network, the original sparse point cloud needs to be Points Sampling the farthest point N point, and for this N Point construction The feature tensor of dimension is denoted as Points 3. Calculation Points 2 and Points The chamfer distance between 3 is used as the loss function during network training.

2. The point cloud completion method based on image feature guidance according to claim 1, characterized in that The process of feature encoding in step 3 is as follows: 3.1): The feature encoding network uses 1 EdgeConv layer and L -1 GuideConv layer in series; 3.2): The output features of the EdgeConv layer and each GuideConv layer are concatenated in the feature dimension, and the concatenated features are processed through a SharedMLP layer for feature dimensionality reduction to obtain Dimensional characteristics.

3. According to the method of claim 2, the point cloud completion method based on image feature guidance is characterized in that the GuideConv layer in 3.1) is a depth-separable convolution form of the EdgeConv layer, and its depth convolution kernel parameters and point-by-point convolution kernel parameters are generated by the image feature guidance in step 1, wherein, F l Guide l The convolution kernel parameters of the GuideConv layer are generated. The specific convolution kernel parameter generation process is as follows: 3.1.1): Using standard convolutional layer processing F l get The characteristic of the dimension is l The depth convolution kernel parameters of the GuideConv layer, where N and C l Respectively represent l The number of points to be processed in the GuideConv layer and the feature dimensions of the points. When performing deep convolution on each feature dimension N Parameters are not shared between points; 3.1.2) Using average pooling F l get dimensional features, where M l express F l The number of channels; 3.1.3) For the results obtained in 3.1.2) The dimension features are obtained by using the fully connected layer The characteristic of the dimension is l The point-by-point convolution kernel parameters of the GuideConv layer, where C l+1 Indicates l The feature dimension of the points output by the GuideConv layer.

4. The point cloud completion method based on image feature guidance according to claim 1, characterized in that: The process of feature expansion in step 4 is as follows: 4.1): Replace the Feature copy of dimension r For copies r The characteristics of the two r The SharedMLP network branches process the r share dimensional perturbation coefficient, where each SharedMLP network branch is composed of L 1 SharedMLP layer in series; 4.2): r share The perturbation coefficients of the dimension are spliced ​​into 4.1) r share In the copy feature of the dimension, we get r share Characteristics of dimensions; 4.3): For 4.2) r share The characteristics of the dimensions are respectively r EdgeConv network branches are processed to obtain r share dimensional point cloud features, where each EdgeConv network branch consists of L 2 EdgeConv layers are connected in series.

5. The point cloud completion method based on image feature guidance according to claim 1, characterized in that: In step 6, according to formula (1), Points 2 and Points Chamfer distance between 3: ; in express Points 3 and Points 2. represents the pre-set weight coefficient, p express Points The point in 2, express Points The point in 3.

Citation Information

Patent Citations

  • System and method for completing three-dimensional point cloud target of laser radar

    CN109613557A

  • Point cloud geometric compression method based on deep convolutional network

    CN110691243A

  • Multi-scale and folding structure-based terracotta warrior point cloud shape completion method and system

    CN112837420A

  • Deep three-dimensional point cloud classification network construction method based on competitive attention fusion

    CN112990336A

  • Three-dimensional point cloud up-sampling method and system, equipment and medium

    CN113674403A

Cited By

  • Ship point cloud completion method and system based on multi-modal data fusion

    CN120278874A