Cross-modal point cloud completion method based on retrieval assistance
By constructing a 3D point cloud asset library and extracting structural features using rendering images and contrastive language images pre-trained models, and combining information fusion with a dual-gating mechanism, the problems of structural generalization and detail loss in point cloud completion methods are solved, achieving more accurate and stable point cloud generation.
Patent Information
- Application Number
- CN202511049168.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
Existing point cloud completion methods have shortcomings in handling structural generalization and loss of detail information. In particular, when real-world data contains arbitrary rotation angles, unseen categories, or sparse displays, the generation effect is poor, and the differences between different modalities affect the generation of fine-grained point clouds.
A retrieval-assisted cross-modal point cloud completion method is adopted. By constructing a 3D point cloud asset library, structural features are extracted using rendering images and contrastive language image pre-trained models, and information is fused by combining a dual-gating mechanism to generate a complete point cloud.
It improves the accuracy and robustness of point cloud completion, and can maintain stable generation performance under conditions of high missingness and no class samples, demonstrating stronger generalization ability and robustness.
Smart Images

Figure CN120976067A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a point cloud processing method, and more particularly to a cross-modal point cloud completion method based on retrieval assistance. Background Technology
[0002] With the development of 3D computer vision technology, point cloud data is increasingly being used in fields such as autonomous driving and 3D scene understanding. Point cloud data is a collection of points in three-dimensional space, each containing X, Y, and Z coordinate information, and sometimes attributes such as color and reflectivity. However, due to inherent limitations in scanning conditions, such as viewpoint occlusion and surface reflectivity, raw point cloud data often exhibits incompleteness. Recovering complete, high-fidelity 3D point clouds is crucial for many downstream tasks.
[0003] Existing point cloud completion methods suffer from two main problems: (1) Structural generalization limitations: Structural features rely on data-driven training methods. When real-world data contains arbitrary rotation angles, unseen categories, or sparse displays, these features may not be generated based on limited input information. (2) Loss of detail: For detailed targets, inferring the details of missing structures from partial input is extremely challenging. Therefore, some methods have introduced instance images captured by RGB cameras on 3D scanners to guide generation. However, the inherent differences between different modalities affect the generation of fine-grained point clouds. Summary of the Invention
[0004] The purpose of this invention is to provide a retrieval-assisted cross-modal point cloud completion method. By progressively obtaining information from the reference point cloud through steps such as structural feature extraction and dual-gating mechanism processing, the method improves the ability to generate a complete and dense point cloud from a defective point cloud, and achieves more accurate and robust completion of the defective point cloud.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] A retrieval-assisted cross-modal point cloud completion method is proposed. Based on the rendered image, a cross-modal 3D point cloud asset library containing feature encoding is constructed. By retrieving similar reference point clouds in the 3D point cloud asset library, structural feature encoding and dual-gating filtering features are applied to the input point cloud, the image, and the reference point cloud. On this basis, progressive fusion of the input point cloud and the reference point cloud is performed to generate a seed for the complete point cloud and complete the refinement.
[0007] Specifically, the steps include the following:
[0008] Step 1: First, construct rendering images of 12 views for the 3D point cloud asset library. For each view, use the Contrastive Language Image Pretraining (CLIP) model to encode and average the results to build a retrieval-enabled 3D point cloud asset library.
[0009] Step 2: Given an input image, use CLIP encoding similarity to find the most similar reference point cloud in the 3D point cloud asset library for the incomplete input point cloud;
[0010] Step 3: Perform structural feature extraction on the image, input point cloud, and reference point cloud using convolutional neural network and local sphere query neighborhood vector feature aggregation respectively, to mine structural feature encoding;
[0011] Step 4: The reference point cloud features are processed through a dual-gating mechanism. The dual-gating mechanism simultaneously calculates the similarity between the reference point cloud features and the input point cloud features, as well as encodes the concatenated vector of the reference point cloud features and the global features of the input point cloud.
[0012] Step 5: Gradually fuse the input point cloud and reference point cloud information step by step, and generate a sparse and coarse seed through global variables;
[0013] Step 6: Based on the local features of the sparse and coarse seed, progressively learn relevant structural features from the input point cloud and the reference point cloud, and complete local refinement to obtain the complete point cloud.
[0014] The specific methods for constructing and retrieving the 3D point cloud asset library are as follows:
[0015] Step 1: In order to obtain reference point clouds, a 3D point cloud asset library containing object and human body categories needs to be built first, and 12 images are rendered for each object from different perspectives.
[0016] Step 2: Using the Contrastive Language Image Pretraining (CLIP) coding model, all rendered images in each 3D point cloud are encoded and averaged to obtain a vector. This vector is then used as the encoded feature of the 3D point cloud.
[0017] Step 3: When using the model, the corresponding 3D point cloud model is retrieved by image CLIP embedding or text retrieval. This means that 3D point cloud retrieval supports both image-based and text-based methods. When retrieval by rendering an image is not possible, a reference point cloud is obtained by generating a 3D point cloud from an image or text. This shows that the acquisition of reference point clouds supports multiple input methods to adapt to different needs.
[0018] Structural feature encoding and dual gating mechanism, the specific methods are as follows:
[0019] Step 1: Feature extraction from point cloud and image: For image, a convolutional neural network is used; for point cloud, a local sphere query neighborhood vector feature aggregation method is used to extract structural features.
[0020] Step 2: Shared structure encoding: Design a shared structure encoder that uses the self-attention mechanism in Vision-Transformer to capture long-range unified information between different modalities and between different objects in the same space; apply absolute position encoding only to the input point cloud and not to the reference point cloud to learn structural information that is not affected by pose.
[0021] Step 3: Feature Fusion: After the shared structure encoding, the alignment features of the cross-modal input are fused to retain the global structural features in the input image, which helps to learn global information from the incomplete structure in the incomplete input and avoids the interference caused by the absolute position difference in the reference point cloud;
[0022] Step 4: Similarity and Missing Components Control Gate: Design a dual-channel control gate to focus on relevant and missing components of the reference sample; the first control gate is used to encode feature correlation, enhance relevant information and mask irrelevant components; the second control gate is used to detect missing components, combining similar features with the global input point cloud for encoding and amplifying the impact of missing components.
[0023] Step 5: Information Integration: Simultaneously use the two control gates mentioned above to extract useful information from the reference point cloud, increasing the representational ability of similar and missing features in the reference point cloud, and reconstructing the reference prior features. A progressive fusion method for generating a complete point cloud from the input and reference point clouds first generates a seed, then refines the local structure of the seed to obtain the complete point cloud. The specific method is as follows:
[0024] Step 1: Initialize the seed point cloud: Combine the global variables of the input point cloud and the reference point cloud to generate a complete but sparse seed point cloud, which will serve as the basis for subsequent refinement.
[0025] Step 2: Utilize retrieval features to assist in generation; use structurally encoded retrieval features as an auxiliary tool to infer missing parts based on existing shape structures and restore geometric details while maintaining data fidelity;
[0026] Step 3: Stepwise refinement of feature learning: Using the seed point cloud as an intermediary variable, further learn details from the input point cloud and the reference point cloud to decode the local neighborhood structure of the seed point cloud; this process effectively processes the reference point cloud through control gates to ensure the relevance and accuracy of the information;
[0027] Step 4: Progressive Local Generation: In the progressive local generation process, features are progressively learned from the aligned input point cloud and the reference point cloud features processed by the control gate, from global to local, thereby optimizing the quality of the final output point cloud.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] This invention employs a cross-modal point cloud completion framework. This framework introduces a retrieval assistance mechanism, combining structural feature encoding methods and a dual-gating mechanism to effectively filter and select features, thereby achieving a progressive fusion of information between the input and reference point clouds. During the fusion process, it gradually transitions from global structural features to local detailed features, fully mining and utilizing the complementary information between the input and reference point clouds. This coarse-to-fine, global-to-local fusion strategy not only enhances the modeling ability for the rationality and continuity of the missing region structure but also effectively improves the completeness and accuracy of the generated results. Experimental results show that when processing point cloud data with severely missing structures, this method can generate more realistic and semantically consistent geometric shapes. Especially when facing high levels of missing data and cases where no category samples are found, it maintains stable completion performance, demonstrating stronger generalization ability and robustness. Attached Figure Description
[0030] Figure 1 This is a flowchart of the algorithm for the cross-modal point cloud completion method based on retrieval assistance proposed in this invention.
[0031] Figure 2(a) is a defect cloud image; Figure 2(b) is the completion result of the traditional method EGIINet; Figure 2(c) is the completion result of the method proposed in this invention; Figure 2(d) is the experimental ground truth image. Detailed Implementation
[0032] The specific details of each step of the present invention will be described in detail below with reference to the accompanying drawings.
[0033] like Figure 1 As shown, the present invention provides a cross-modal point cloud completion method based on retrieval assistance, comprising the following steps:
[0034] Step 1: Construction of Cross-Modal Asset Library and Retrieval of Reference Samples
[0035] First, it is necessary to construct 12 rendering images from different perspectives of a 3D point cloud asset library and obtain their encoding. In order to obtain reference point clouds, a 3D point cloud asset library containing object and human body categories needs to be constructed, and 12 images are rendered for each object from different perspectives; this asset library is accomplished by expanding the existing 3D point cloud asset library and generating corresponding single-view rendering images.
[0036] Step 2: Given an input image, the most similar reference point cloud in the 3D point cloud asset library is found using CLIP encoding similarity for the incomplete input point cloud. This invention first encodes the 3D object rendering images from 12 viewpoints using the multimodal pre-trained neural network model CLIP. Then, based on the input single-view image I∈R... H×W×C Or, a text encoder can quickly retrieve semantically similar reference 3D samples P. r The core of the retrieval process lies in extracting feature representations from images and point clouds, and finding the most relevant reference point cloud through semantic matching.
[0037] Step 3: Perform structural feature extraction on the image, input point cloud, and reference point cloud using convolutional neural network and local sphere query neighborhood vector feature aggregation respectively, to mine structural feature encoding;
[0038] After obtaining the reference point cloud, a structure-shared feature encoding module is used to perform joint feature extraction on the input point cloud P, the reference point cloud P_r, and the image I. This module is implemented in the following way:
[0039] 1. Local surrogate coding: For image I, a block-based coding technique is used to divide the image into multiple regions and convert them into feature vectors F_i through 2D convolution.
[0040] F i =Conv2D(Patch(I))
[0041] 2. Neighborhood Aggregation Encoding: For the input point cloud P and the reference point cloud P_r, a region proxy encoding method is used, where each point represents the local structural relationship through aggregation of its neighborhood information. Neighboring points are identified through ball query to better capture key structural information.
[0042] F p =GraphConv(F p -ballQuery(P i ,F p ))
[0043] Step 4: Feature Alignment and Blending
[0044] The reference point cloud features are processed through a dual-gating mechanism, which simultaneously calculates the similarity between the reference and input point cloud features, and encodes the concatenated vector of the reference and input point cloud global features. After encoding by a shared feature module, features from different modalities are aligned and fused. Specifically:
[0045] A dual-gating mechanism was designed: a similarity control gate and a missing information perception control gate were designed to enhance relevant features and suppress irrelevant interference, respectively, while highlighting information in the reference sample that is related to the missing part of the input point cloud.
[0046] First, image features F i It provides global structural information, helping the model learn the overall distribution of missing structures from the image. Input point cloud features F p and reference point cloud features F' p The influence of absolute position embedding is avoided during the encoding process, thereby alleviating the spatial misalignment problem caused by changes in the pose of the reference sample.
[0047] The fused features retain global structural information from the input image and incorporate geometric priors from the reference point cloud, laying the foundation for subsequent generation tasks.
[0048] Step 5: Gradually fuse the input point cloud and reference point cloud information step by step, and generate a sparse and coarse seed through global variables;
[0049] To progressively generate a complete point cloud, a progressive fusion of the input and reference point clouds is introduced. The main implementation steps of this module include:
[0050] 1. Seed point cloud generation: By coupling the global variables of the input point cloud and the reference sample, a sparse but complete initial point cloud (called "seed") is generated.
[0051] 2. Hierarchical Feature Fusion: The geometric details of the seed point cloud are learned layer by layer from global to local. In this process, the progressive generator uses reference features processed by the shared feature encoding module and input features to progressively optimize the shape and structure of the seed point cloud.
[0052] 3. Simultaneously, the structured coding features of the input point cloud and the processed reference point cloud are used as auxiliary tools to infer the missing parts and restore geometric details, while maintaining the authenticity of the data.
[0053] Step 6: Based on the local features of the sparse and coarse seed, progressively learn relevant structural features from the input point cloud and the reference point cloud, and complete local refinement to obtain the complete point cloud.
[0054] To further improve the quality of the seed point cloud, the progressive generator focuses on restoring local details. The specific operation is as follows:
[0055] First, a local query generation is performed, using global variables and location information to generate a local query Q for the seed point cloud. i This is to re-represent seed features and reflect their local neighborhood information.
[0056] Then, feature interaction and reconstruction are performed. By combining global input and retrieval priors with a multilayer perceptron, the local structure of the seed point cloud is gradually refined to generate more accurate geometric details.
[0057] After completing the above steps, the progressive generator module outputs a complete point cloud result. The final generated point cloud not only matches the input data in overall contour but also fills in the missing details by retrieving geometric priors from the samples. Furthermore, by comparing the original input point cloud and the generated result, the model's performance in cross-modal completion tasks can be evaluated, verifying its performance in terms of geometric fidelity and detail richness.
[0058] The beneficial effects of the present invention are verified using the following embodiments:
[0059] Example 1:
[0060] This embodiment of a retrieval-assisted cross-modal point cloud completion method is prepared according to the following steps:
[0061] The data used in the experiment are the test defect cloud and images of ShapeNet-ViPC. The original image and ground truth image are shown in Figure 2(a) and Figure 2(d). Figure 2(b) is the point cloud completion result of the traditional method, and Figure 2(c) is the point cloud completion result of the method of the present invention. The experimental results verify the effectiveness of the cross-modal point cloud completion method based on retrieval assistance proposed in this invention.
[0062] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A cross-modal point cloud completion method based on retrieval assistance, characterized in that: A cross-modal 3D point cloud asset library with feature encoding is constructed based on the rendered image. By searching for similar reference point clouds in the 3D point cloud asset library, structural feature encoding and dual-gating filtering features are applied to the input point cloud, image, and reference point cloud. On this basis, progressive fusion of the input point cloud and reference point cloud is performed to generate the seed of the complete point cloud and complete the refinement. Specifically, the steps include the following: Step 1: First, construct rendering images of 12 views for the 3D point cloud asset library. For each view, use the Contrastive Language Image Pretraining (CLIP) model to encode and average the results to build a retrieval-enabled 3D point cloud asset library. Step 2: Given an input image, use CLIP encoding similarity to find the most similar reference point cloud in the 3D point cloud asset library for the incomplete input point cloud; Step 3: Perform structural feature extraction of the image, input point cloud, and reference point cloud using convolutional neural network and local sphere query neighborhood vector feature aggregation respectively, and mine structural feature encoding; Step 4: The reference point cloud features are processed through a dual-gating mechanism. The dual-gating mechanism simultaneously calculates the similarity between the reference point cloud features and the input point cloud features, as well as encodes the concatenated vector of the reference point cloud features and the global features of the input point cloud. Step 5: Gradually fuse the input point cloud and reference point cloud information step by step, and generate a sparse and coarse seed through global variables; Step 6: Based on the local features of the sparse and coarse seed, progressively learn relevant structural features from the input point cloud and the reference point cloud, and complete local refinement to obtain the complete point cloud.
2. The cross-modal point cloud completion method based on retrieval assistance according to claim 1, characterized in that: The specific methods for constructing and retrieving the 3D point cloud asset library are as follows: Step 1: In order to obtain reference point clouds, a 3D point cloud asset library containing object and human body categories needs to be built first, and 12 images are rendered for each object from different perspectives. Step 2: Using the Contrastive Language Image Pretraining (CLIP) coding model, all rendered images in each 3D point cloud are encoded and averaged to obtain a vector. This vector is then used as the encoded feature of the 3D point cloud. Step 3: When using the model, the corresponding 3D point cloud model is retrieved by image CLIP embedding or text retrieval. This means that 3D point cloud retrieval supports both image-based and text-based methods. When retrieval by rendering an image is not possible, a reference point cloud is obtained by generating a 3D point cloud from an image or text. This shows that the acquisition of reference point clouds supports multiple input methods to adapt to different needs.
3. The cross-modal point cloud completion method based on retrieval assistance according to claim 1, characterized in that: Structural feature encoding and dual gating mechanism, the specific methods are as follows: Step 1: Feature extraction from point cloud and image: For image, a convolutional neural network is used; for point cloud, a local sphere query neighborhood vector feature aggregation method is used to extract structural features. Step 2: Shared structure encoding: Design a shared structure encoder that uses the self-attention mechanism in Vision-Transformer to capture long-range unified information between different modalities and between different objects in the same space; apply absolute position encoding only to the input point cloud and not to the reference point cloud to learn structural information that is not affected by pose. Step 3: Feature Fusion: After the shared structure encoding, the alignment features of the cross-modal input are fused to retain the global structural features in the input image, which helps to learn global information from the incomplete structure in the incomplete input and avoids the interference caused by the absolute position difference in the reference point cloud; Step 4: Similarity and Missing Components Control Gate: Design a dual-channel control gate to focus on relevant and missing components of the reference sample; the first control gate is used to encode feature correlation, enhance relevant information and mask irrelevant components; the second control gate is used to detect missing components, combining similar features with the global input point cloud for encoding and amplifying the impact of missing components. Step 5: Information Integration: Simultaneously use the two control gates mentioned above to obtain useful information from the reference point cloud, increase the representation ability of similar and missing features in the reference point cloud, and reconstruct the reference prior features.
4. The cross-modal point cloud completion method based on retrieval assistance according to claim 1, characterized in that: A method for progressively fusing input and reference point clouds to generate a complete point cloud first generates a seed, and then refines the local structure of the seed to obtain the complete point cloud. The specific method is as follows: Step 1: Initialize the seed point cloud: Combine the global variables of the input point cloud and the reference point cloud to generate a complete but sparse seed point cloud, which will serve as the basis for subsequent refinement. Step 2: Utilize retrieval features to assist in generation; use structurally encoded retrieval features as an auxiliary tool to infer missing parts based on existing shape structures and restore geometric details while maintaining data fidelity; Step 3: Stepwise refinement of feature learning: Using the seed point cloud as an intermediary variable, further learn details from the input point cloud and the reference point cloud to decode the local neighborhood structure of the seed point cloud; this process effectively processes the reference point cloud through control gates to ensure the relevance and accuracy of the information; Step 4: Progressive Local Generation: In the progressive local generation process, features are progressively learned from the aligned input point cloud and the reference point cloud features processed by the control gate, from global to local, thereby optimizing the quality of the final output point cloud.