Point cloud completion method and device

Through the feature extraction and fusion of the CLIP model and VIT encoder, combined with the cross attention mechanism, efficient and accurate completion of point cloud data is achieved, structural correlation modeling and multimodal fusion problems in point cloud completion are solved, and the integrity and quality of point cloud data are improved.

CN120339433APending Publication Date: 2025-07-18NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510439991.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing technology has structural correlation modeling problems, geometric-semantic collaborative optimization dilemma and multimodal fusion technology obstacles in point cloud completion. Especially under object occlusion, differences in surface material reflection characteristics and sensor resolution limitations, it is difficult to effectively capture cross-region structural continuity and perform accurate multimodal data fusion.

Method used

The CLIP model is used to predict and feature extraction of images and point cloud data, combine VIT encoder and position perception module for feature fusion, and use the cross attention mechanism to achieve accurate point cloud completion through the point cloud completion network.

Benefits of technology

It improves the accuracy and efficiency of point cloud completion, reduces the dependence on high-precision devices and complex operations, provides richer context information, enhances the model's understanding of point cloud data, and solves the problems of structural correlation modeling and multimodal fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339433A_ABST
    Figure CN120339433A_ABST
Patent Text Reader

Abstract

The invention relates to a point cloud completion method and device, and belongs to the technical field of computer application. Generating a two-dimensional orthographic projection drawing of the preprocessed point cloud data, performing normalization processing, copying the projection drawing into two parts, complementing one part, and performing block processing on the other part; predicting the category of the complemented projection drawing through a CLIP model to obtain text features and image features; inputting the complemented image into a VIT code to obtain a local feature of the point cloud; inputting the partitioned images into a VIT encoder and a position sensing module PA to obtain position sensing features; fusing the text features, the image features and the position perception features to form fused features; using a cross attention mechanism to interact the local features with the fusion features to obtain enhanced point cloud features, and inputting the enhanced point cloud features into a point cloud completion network to obtain a completed point cloud; according to the method, the complementation precision and effect are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer application technology, and specifically relates to a point cloud completion method and device. Background Art

[0002] The core challenge of point cloud completion technology stems from the complex acquisition conditions in actual application scenarios. Factors such as object occlusion, differences in surface material reflection characteristics, sensor resolution limitations, and field of view constraints together lead to partial missing of the acquired point cloud data. This data incompleteness not only restricts the performance of downstream applications, but also puts forward deep technical requirements for the completion algorithm. The existing technology has the following technical problems:

[0003] (1) Difficulty in modeling structural correlation. The defective area and the existing point cloud have significant correlation at the level of potential structural features. However, when mining this structural correlation, existing algorithms are limited by their inability to extract local features and are unable to effectively capture cross-regional structural continuity.

[0004] (2) The dilemma of geometric-semantic co-optimization. Although existing methods can generate visually coherent geometric shapes, there is a common problem of separation between geometric reconstruction and semantic understanding. This lack of semantic information directly leads to significant limitations in existing models in terms of overall morphological cognition and defective area localization.

[0005] (3) Technical barriers to multimodal fusion. Multimodal data fusion (such as RGB images and point cloud data) can provide complementary texture information and spatial clues for completion tasks. However, combining these images with point cloud data requires accurate camera calibration and complex spatial alignment, which is very challenging in practical applications. Summary of the invention

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] To achieve the above object, the present invention specifically adopts the following technical solutions:

[0008] A point cloud completion method comprises the following steps:

[0009] Step 1, data acquisition and preprocessing: obtain raw point cloud data from the sensor and preprocess it;

[0010] Step 2, generate two-dimensional orthographic projection images: use the projection method to generate six two-dimensional orthographic projection images of the preprocessed point cloud data, and normalize each projection image;

[0011] Step 3: Projection image completion and block processing: Copy the generated projection image twice, one for completion and the other for block processing ;

[0012] Step 4, CLIP Model Prediction and Feature Extraction: Input the completed image into the CLIP model to predict its category and obtain the corresponding text label , and acquire the text features and image features ;

[0013] Step 5, VIT Encoder Feature Extraction and Position Perception: Input the completed image into the VIT encoder to obtain the local features of the point cloud ; then input the segmented image into the VIT encoder to get the features of each block , and input into the position perception module PA to obtain the position perception features ;

[0014] Step 6, Multimodal Feature Fusion: Fuse the text features , image features and position perception features to form the fused features ;

[0015] Step 7, Cross-Attention Operation: Use the cross-attention mechanism to interact the local features of the point cloud obtained in Step 5 with the fused features obtained in Step 6 to obtain the enhanced point cloud features ;

[0016] Step 8, Input to the Point Cloud Completion Network and Completion: Input the enhanced point cloud features into the point cloud completion network to obtain the finally completed point cloud .

[0017] A point cloud completion device includes the following modules:

[0018] Data Acquisition and Preprocessing Module, which acquires the original point cloud data from the sensor and performs preprocessing;

[0019] Generate 2D Orthographic Projection Map Module, which uses the projection method to generate six 2D orthographic projection maps of the preprocessed point cloud data and normalizes each projection map;

[0020] Projection Map Completion and Block Processing Module, which duplicates the generated projection map, one for completion and the other for block processing to obtain ;

[0021] CLIP Model Prediction and Feature Extraction Module, which inputs the completed image Predict its category through the CLIP model to obtain the corresponding text label , and obtain the text features and the image features ;

[0022] VIT encoder feature extraction and position perception module, input the complemented image into the VIT encoder to obtain the local features of the point cloud ; then input the segmented image into the VIT encoder to obtain the features of each block , and input into the position perception module PA to obtain the position perception features ;

[0023] Multimodal feature fusion module, fuse the text features , the image features and the position perception features to form the fused features ;

[0024] Cross-attention operation module, using the cross-attention mechanism, interact the local features of the point cloud obtained in step 5 with the fused features obtained in step 6 to obtain the enhanced point cloud features ;

[0025] Input the point cloud completion network and the completion module, input the enhanced point cloud features into the point cloud completion network to obtain the finally completed point cloud .

[0026] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the steps of the point cloud completion method when executing the program.

[0027] A non-transitory computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of the point cloud completion method when executed by a processor.

[0028] The present invention has the following beneficial effects:

[0029] The present invention utilizes the position perception module to accurately identify the position information of the missing part in the point cloud. This ability makes the completion process more accurate and generates more accurate completion results. The position perception module analyzes the local and global features of the projection map to learn the spatial information of the missing part in each block, thereby guiding the point cloud completion network to more accurately locate and complete the missing area, and solving the problem of structural correlation modeling.

[0030] The present invention utilizes the pre-trained vision-language model CLIP, enabling the model to learn rich feature representations from a large number of image and text pairs, establishing the connection between semantics and geometry, and solving the geometric-semantic co-optimization dilemma.

[0031] The present invention integrates multi-modal data by fusing point cloud, image, and text information. This integration not only provides richer context information, enhancing the model's understanding ability of point cloud data, but also helps the model learn more comprehensive feature representations from different perspectives and information sources. In addition, the present invention also uses the cross-attention mechanism to further enhance the correlation between point cloud features and multi-modal features. This mechanism enables the model to pay more attention to the features important for the completion task during the completion process, improving the accuracy and efficiency of completion and solving the multi-modal fusion technical obstacle.

[0032] Based on solving the three technical problems, the present invention uses the pre-trained model of the CLIP model, greatly reducing the time and resources required for training the model from scratch, while improving the performance of the model in the point cloud completion task. Compared with traditional methods, this technology avoids complex preprocessing steps, such as precise camera calibration and spatial alignment. This simplification makes the pairing process of point cloud and image data more efficient, reducing the dependence on high-precision equipment and complex operations. Through this simplification, the present invention can adapt to different data sets faster and provide faster completion results. Brief Description of the Drawings

[0033] Figure 1 It is a flowchart of the point cloud completion method of the present invention. Detailed Embodiment

[0034] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0035] The point cloud completion method of the present invention has been widely experimentally verified on multiple data sets. The results show that the method has achieved excellent completion effects on different types of point cloud data, demonstrating good robustness and generalization ability.

[0036] The execution environment of the present invention is a Core 4-core computer with a 4.0 GHZ central processing unit and 128 G bytes of memory, and is trained using the Adam optimizer with a batch size of 32 on two RTX 4090 GPUs through the deep learning framework PyTorch. On the premise that the computer memory and video memory permit, the present invention can also be based on other execution environments.

[0037] Such asFigure 1 As shown in Figure 1 , the point cloud completion method of the present invention includes the following steps:

[0038] Step 1: Data acquisition and preprocessing. Use a depth camera or lidar device (sensor) to scan the target scene to obtain the original point cloud data , and perform preprocessing on the original point cloud data , including operations such as noise reduction, filtering, point cloud resampling, and outlier removal, to improve the data quality. The preprocessed point cloud data is converted into a unified coordinate system and scale for subsequent processing.

[0039] Step 2: Generate two-dimensional orthographic projection maps. Orthogonally project the preprocessed point cloud data along the X, Y, and Z axes to generate two-dimensional orthographic projection maps of six main views , - which are respectively the front view, rear view, top view, bottom view, left view, and right view of the point cloud data . The two-dimensional orthographic projection maps are used for subsequent feature extraction and completion operations. Normalize each projection map to eliminate the influence of different scales on subsequent processing, and save the depth information of each projection map for use in the completion process.

[0040] Step 3: Projection map completion and block processing. Duplicate the two-dimensional orthographic projection maps. For each two-dimensional projection map ( taking values from 1 to 6, that is, the two-dimensional projection maps of six main views) in one copy of the two-dimensional orthographic projection maps, use an image inpainting algorithm to perform image completion based on the VGG deep learning model to fill in the missing parts, making the finally output point cloud data more complete and reliable, and obtaining the completed image . Divide the other copy of the two-dimensional projection maps into several blocks, such as 2×2 blocks, denoted as image , to prepare for the processing of the position perception module in the subsequent steps.

[0041] Step 4: CLIP model prediction and feature extraction. Input the completed image into the pre-trained vision-language model CLIP to predict its category and obtain the corresponding text label . At the same time, extract the image features and text features from the CLIP model. These features contain the visual information of the image and the semantic information of the text. Feature extraction can be represented by the following formula:

[0042] ;

[0043] Among them, CLIP (Contrastive Language-Image Pretraining) is a vision-language model that uses a pre-trained vision-language model to process text and image data. Through contrastive learning of images and text, it achieves cross-modal feature understanding and generation. It can associate images with text, enabling the model to understand the semantic content of images, thereby providing context information for point cloud completion.

[0044] Step 5: VIT Encoder Feature Extraction and Position Awareness. Input the completed image into the VIT encoder of the Vision Transformer model to obtain the local features of the point cloud . Input the segmented image into the VIT encoder of the Vision Transformer model to obtain the features of each block . Then input into the Position Awareness Module PA. The Position Awareness Module PA understands the spatial information of the missing parts in each block and outputs the position awareness features . The extraction of the position awareness features can be expressed by the following formula:

[0045] ;

[0046] ;

[0047] ;

[0048] Among them, the VIT (Vision Transformer) encoder is a deep learning model based on the Transformer architecture, specifically designed to process projected image data. This module first divides the input image into multiple small blocks and learns independent parameters for each block to capture local features. This process enables the VIT encoder to effectively identify details in the image. Especially when dealing with point clouds with complex structures, the capture of local features is crucial. In this way, the VIT encoder can not only extract the nuances of the image but also provide key information for point cloud completion, helping the model identify the specific location of the missing parts in the image.

[0049] The main objective of the Position Awareness Module PA is to learn the position information of the missing parts in the point cloud. This module divides the projected image into multiple blocks and assigns weights to each block to identify the position of the missing parts. This position awareness ability enables the network to more accurately locate the areas in the point cloud that need to be completed, thereby improving the accuracy of completion. Specifically, the Position Awareness Module can analyze the features of each image block and adjust the weights according to its relative position in the overall image, which can effectively improve the recognition rate of the missing areas.

[0050] Step 6: Multimodal feature fusion. Combine the text features , image features and location-aware features for fusion, which can be achieved through concatenation, weighted averaging, or a more complex feature fusion network. This process integrates features from different modalities to form multimodal fusion features . Consider the point cloud features and the fusion features as a point cloud triple to complete the construction of the dataset for input to the point cloud completion network. Feature fusion is represented by the following formula:

[0051] ;

[0052] where is the concatenation operation.

[0053] This step combines the features obtained from the CLIP encoder and the VIT encoder with the features of the point cloud data itself, integrates the local features and global features through the feature fusion layer, and outputs the final image features. These fused features not only retain the detailed information of the image but also incorporate the context semantics, providing richer information for the subsequent point cloud decoder. Through this feature fusion mechanism, it can be ensured that the network fully utilizes the local and global information of the image during the completion process, thereby achieving a more accurate completion effect.

[0054] Step 7: Cross-attention operation. Utilize the cross-attention mechanism to interact the local features of the point cloud obtained in Step 5 with the fusion features obtained in Step 6 , enabling the point cloud features to capture the context information from the image and text. The cross-attention operation is represented by the following formula:

[0055] ;

[0056] where is the cross-attention mechanism, is the enhanced point cloud features, which can focus on relevant information by weighting the input features. In multimodal learning, the cross-attention mechanism is particularly effective as it can establish connections between different modalities, enabling the information of each modality to enhance each other.

[0057] Step 8: Input to the point cloud completion network and completion. Input the enhanced point cloud features into the point cloud completion network to obtain the finally completed point cloud .

[0058] The point cloud completion network utilizes this information to predict and generate the missing parts of the point cloud. The point cloud completion network consists of multiple convolutional layers and upsampling layers, gradually restoring the geometric structure and detailed information of the point cloud. After a series of processes, the network finally outputs complete point cloud data, which not only contains the information of the original point cloud but also includes the newly added points after completion, significantly improving the integrity and quality of the point cloud. Point cloud completion can be represented by the following formula:

[0059] ;

[0060] Among them, are four advanced point cloud completion networks, namely SnowflakeNet, PoinTr, AdaPoinTr, and CRA-PCN.

[0061] Meanwhile, in order to evaluate the completion effect, the CDL1 loss function is adopted for optimization, and its calculation formula is as follows:

[0062] ;

[0063] Among them, is the loss function, is the total number of points in the point cloud, is the completed points predicted by the network, is the corresponding point in the original point cloud. The CDL1 loss function can minimize the error between the completion result and the original data by calculating the absolute difference between the predicted points and the real points, thereby improving the accuracy and reliability of the completion. In this way, the point cloud completion network can not only effectively fill in the missing parts but also optimize the quality of the output point cloud, making the finally generated point cloud reach a relatively high level in terms of geometric shape and detail performance.

[0064] As described above, the present invention can effectively complete incomplete point cloud data, improve the integrity and quality of the point cloud, and is applicable to multiple fields such as autonomous driving, robot navigation, and virtual reality. The present invention can be applied to three-dimensional object reconstruction, especially in the point cloud completion task. By effectively completing the scanned defective point cloud, it can effectively assist three-dimensional reconstruction. Compared with some existing classic point cloud completion networks, the point cloud completion method of the present invention has higher measurement accuracy.

[0065] As described above, it is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments according to the technical essence of the present invention still belong to the technology of the present invention.

[0066] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code. The solutions in the embodiments of the present invention can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript, etc.

[0067] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0068] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0069] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0070] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0071] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A point cloud completion method, characterized in that, It includes the following steps: Step 1, data acquisition and preprocessing: Obtain the original point cloud data from the sensor and perform preprocessing; Step 2, generate two-dimensional orthographic projection maps: Use the projection method to generate six two-dimensional orthographic projection maps of the preprocessed point cloud data, and perform normalization processing on each projection map; Step 3, Projection Map Completion and Block Processing: Duplicate the generated projection map, complete one copy, and perform block processing on the other copy to obtain ; Step 4, CLIP model prediction and feature extraction: For the completed image predict its category through the CLIP model to obtain the corresponding text label , and obtain the text features and the image features ; Step 5, VIT encoder feature extraction and position perception: Input the completed image into the VIT encoder to obtain the local features of the point cloud ; Then, the segmented images are input into the VIT encoder to obtain the features of each block , and are input into the position-aware module PA to obtain the position-aware features ; Step 6, Multimodal Feature Fusion: Fuse the text features , image features and location awareness features to form fused features ; Step 7, cross-attention operation: Using the cross-attention mechanism, the local features of the point cloud obtained in step 5 are interacted with the fused features obtained in step 6 to obtain enhanced point cloud features ; Step 8, Input to the point cloud completion network and completion: Input the enhanced point cloud features into the point cloud completion network to obtain the finally completed point cloud .

2. The point cloud completion method according to claim 1, wherein In the said Step 1, the preprocessing includes noise reduction, filtering, point cloud resampling, and removal of outliers.

3. The point cloud completion method according to claim 1, wherein In step 2, the preprocessed point cloud data is orthogonally projected along the X, Y, and Z axes to generate two-dimensional orthographic projection diagrams of six main views , - which are respectively the front view, rear view, top view, bottom view, left view, and right view of the point cloud data ; and each projection diagram is normalized.

4. The point cloud completion method according to claim 1, characterized in that In the said step 3, the completion includes: for each two-dimensional projection diagram in one of the two-dimensional orthographic projection diagrams, using an image inpainting algorithm to perform image completion based on a VGG deep learning model to fill in the missing parts and obtain a completed image .

5. The point cloud completion method according to claim 1, wherein In step 4, the completed projection image is input into the CLIP model to predict its category and obtain the corresponding text label ; meanwhile, image features and text features are extracted from the CLIP model, and the feature extraction is expressed by the following formula: 。 6. The point cloud completion method according to claim 1, wherein In step 6, the text feature , the image feature and the location perception feature are fused to form a multi-modal fusion feature ; the point cloud feature and the fusion feature are regarded as a point cloud triple to complete the construction of the dataset; the feature fusion is represented by the following formula: ; Among them, is a splicing operation; the fusion is achieved by splicing, weighted average or other feature fusion networks.

7. The point cloud completion method according to claim 1, characterized in that In Step 8, the point cloud completion is expressed by the following formula: ; Among them, are four advanced point cloud completion networks, namely SnowflakeNet, PoinTr, AdaPoinTr, and CRA-PCN; Meanwhile, the CDL1 loss function is adopted for optimization, and the calculation formula is as follows: ; Among them, is the total number of points in the point cloud, is the completed points predicted by the network, is the corresponding point in the original point cloud.

8. A point cloud completion device, characterized in that, It includes the following modules: Data acquisition and preprocessing module, which obtains the original point cloud data from the sensor and performs preprocessing; Generate two-dimensional orthographic projection map module, which uses the projection method to generate six two-dimensional orthographic projection maps of the preprocessed point cloud data, and performs normalization processing on each projection map; Projection map completion and block processing module; Make two copies of the generated projection diagram, complete one copy, and perform block processing on the other copy to obtain ; CLIP model prediction and feature extraction module, for the completed projection image Predict its category through the CLIP model to obtain the corresponding text label , and obtain the text features and image features ; The VIT encoder feature extraction and position perception module inputs the completed image into the VIT encoder to obtain the local features of the point cloud ; Then, the segmented images are input into the VIT encoder to obtain the features of each block , and the features are input into the position-aware module PA to obtain the position-aware features ; Multimodal feature fusion module, which fuses text features , image features and location features to form fused features ; Cross-attention operation module, using the cross-attention mechanism, to use the cross-attention mechanism to obtain the local features of the point cloud obtained in step 5 and the fused features obtained in step 6 are interacted to obtain enhanced point cloud features ; The input point cloud completion network and the completion module input the enhanced point cloud features into the point cloud completion network to obtain the finally completed point cloud .

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the said program, it realizes the steps of the point cloud completion method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it realizes the steps of the point cloud completion method as described in any one of claims 1 to 7.