Point cloud restoration method based on multi-modal information guidance
By building a 3D shape corpus and multimodal point cloud repair network model, and using the CLIP model to extract multimodal features for point cloud repair, the problems of point cloud data sparsity and semantic information are solved, and more efficient point cloud repair performance is achieved.
Patent Information
- Application Number
- CN202510036842.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art faces the problems of inherent sparsity of point cloud data and lack of semantic information in point cloud repair, resulting in limited performance of 3D point cloud reconstruction, detection and classification.
Using a point cloud repair method based on multimodal information guidance, by building a 3D shape corpus, acquiring point cloud, text description and rendering images, constructing a ternary data pair, and designing a multimodal point cloud repair network model, using the CLIP model to extract multimodal features, and performing multi-stage feature fusion to achieve point cloud repair.
It significantly improves the performance of point cloud repair, can more accurately predict the semantic and geometric features of incomplete three-dimensional shapes, and improves the generalization ability and reliability of the model.
Smart Images

Figure CN120107115A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of fully automated products and relates to a point cloud repair method based on multi-modal information guidance. Background Art
[0002] With the increasing popularity of 3D scanning equipment, it has become more convenient to obtain point cloud data. However, due to factors such as occlusion, reflection, transparency and low resolution during the scanning process, point cloud data is often incomplete, sparse and missing. These incomplete data may lead to semantic ambiguity and significantly affect the performance of downstream tasks such as 3D point cloud reconstruction, detection and classification. Therefore, point cloud restoration has gradually become a research hotspot in the field of 3D vision in recent years. Although deep learning-based methods have achieved excellent performance, the inherent sparsity of scanned point cloud data and the lack of semantic information still hinder the further development of 3D point cloud restoration. Fortunately, the recent rapid development of multimodal technology has provided new possibilities and effective ways to address this challenge.
[0003] Previous multimodal restoration methods
[20]
[21]
[23]
[26]
[28]
[43]
[42]
[44] Due to the small size of training data, the training of additional modal branches is insufficient.
[13] et al. demonstrated that pre-trained 2D models can improve the performance of 3D point cloud learning tasks.
[0004] Existing text generation techniques, such as D3Net
[15] Methods lack the description of model geometry details.
[16] and Text2Shape
[17] Such algorithms can be used for various 3D objects (especially from ShapeNet-ViPC [5] However, these methods face challenges in building large-scale, multi-category 3D object text datasets due to the associated cost. Summary of the invention
[0005] In order to solve the above problems, the technical solution adopted by the present invention is: a point cloud repair method based on multimodal information guidance, comprising the following steps:
[0006] Build a 3D shape corpus to describe the geometry of objects;
[0007] The point cloud data of an object with missing point cloud data, the text describing the geometric shape of the object with missing point cloud data using a 3D shape corpus, and the rendered images of the object randomly selected from 24 perspectives are constructed into a data set in the form of three-dimensional data pairs;
[0008] Construct a multimodal point cloud repair network model for repairing the complete geometric shape of an object;
[0009] The multimodal point cloud restoration network model is trained based on the training set data to obtain a trained multimodal point cloud restoration network model;
[0010] The test set data is input into the trained multimodal point cloud repair network model to repair the geometric shape of objects that lack point cloud data.
[0011] Further: the multimodal point cloud restoration network model includes:
[0012] Point cloud encoding module: used to encode the missing point cloud data of the object to obtain the fine-grained features of the point cloud;
[0013] Image encoding module: used to encode the randomly selected rendered images of the object under 24 viewing angles to obtain fine-grained features of the image;
[0014] CLIP module: used to encode the randomly selected rendering images of the object under 24 viewing angles to obtain the global features of the image and to encode the text description information of the geometric shape description of the object lacking point cloud data to obtain the global features of the text;
[0015] Multi-stage feature fusion module: used to fuse the global features of the image and the global features of the text with the fine-grained features of the point cloud, and then perform fine-grained feature fusion with the fine-grained features of the image under the local attention mechanism; perform multi-stage fine-grained multi-head attention feature fusion under the guidance of the global features of the image and the global features of the text, and predict the data of the complete point cloud object corresponding to the input incomplete point cloud data;
[0016] Loss function module: It is used to use the chamfer distance as the loss function to train the data of the same shape as the complete point cloud data, and optimize the multimodal point cloud restoration network model.
[0017] Furthermore: the image encoder module adopts the image encoder of the CLIP model, and the CLIP module adopts the text encoder of the CLIP model.
[0018] Further: after fusing the global features of the image and the global features of the text with the fine-grained features of the point cloud, fine-grained features are fused with the fine-grained features of the image under the local attention mechanism; multi-stage fine-grained multi-head attention feature fusion is performed under the guidance of the global features of the image and the global features of the text to obtain the complete point cloud data with the same shape data as follows:
[0019] In the first stage of fusion, the global text feature Gt and the global visual feature G i After concatenation, these multimodal features are repeated and combined with the fine-grained point cloud features To splice;
[0020] The fusion of multimodal features is completed through the shared MLP operation, and the output is recorded as After the first multimodal fusion, the fine-grained image features F i With the multimodal feature y 1 In this operation, the cross attention mechanism is used to complete the fusion of fine-grained features, and the output result is recorded as
[0021] Then, using the fine-grained features y 2 With the multimodal global feature G i and G t The second fusion process is the same as the first stage operation process. The output of the second stage is recorded as
[0022] Further: The formula of the loss function is as follows:
[0023]
[0024] Among them: The first item encourages the predicted point cloud Y to be as close as possible to the real point cloud Y gt , and the second term ensures that Y can cover the real point cloud Y gt .
[0025] Further: the construction of the text corpus of the geometric shape of the object adopts BLIP-2 to receive a single-view rendered image and a text prompt question as input, and generate a text prompt answer as output, wherein the text prompt question adopts a series of language prompts as follows:
[0026] Category question answering: used to give a coarse-grained description of the shape appearance;
[0027] Existence question answering: used to determine whether a specific component exists in a given 3D shape;
[0028] Quantity question answering: used to accurately describe the quantity of a specific component in a given 3D shape;
[0029] Appearance question answering: used to give a fine-grained description of the component's appearance, constructing a question related to the component.
[0030] A point cloud repair device based on multimodal information guidance, comprising:
[0031] Building Module I: used to build a 3D shape corpus for describing the geometric shapes of objects;
[0032] Acquisition module: used to acquire the missing point cloud data of an object, use the 3D shape corpus to describe the geometric shape of the object that lacks point cloud data, and construct a dataset in the form of three-dimensional data pairs using randomly selected rendering images of the object under 24 viewing angles;
[0033] Building Module II: Multimodal point cloud restoration network model for restoring the complete geometric shape of an object;
[0034] The training module trains the multimodal point cloud restoration network model based on the training set data to obtain a trained multimodal point cloud restoration network model;
[0035] Implementation module: used to input the test set data into the trained multimodal point cloud repair network model to repair the geometric shape of the object that lacks point cloud data. The present invention provides a point cloud repair method based on multimodal information guidance, and proposes a novel multimodal fusion network for point cloud repair, which can simultaneously fuse point cloud, visual and text information to more generalize and accurately predict the semantic and geometric features of incomplete three-dimensional shapes. Specifically, in order to deal with the problem of insufficient prior information brought by small-scale data sets, a visual-language model pre-trained on a large-scale image-text pairing data set, namely the CLIP model, is used to obtain the global feature encoding (global feature encoding of the image and global feature encoding of the text) used to describe the object. As a result, the text encoder and visual encoder of the model show stronger generalization ability. Subsequently, we designed a multi-stage feature fusion strategy to gradually fuse text and global visual features into point cloud features for guiding point cloud features. In addition, to further explore the effectiveness of fine-grained text descriptions, we constructed a text corpus containing rich geometric details to provide more detailed information for three-dimensional shape repair. These fine-grained text descriptions are used for network training and evaluation. Sufficient quantitative and qualitative experiments show that the proposed method has significant advantages over existing point cloud restoration methods.
[0036] In summary, this application has the following advantages:
[0037] In order to improve the generalization ability and reliability, the CLIP model is adopted to migrate the multimodal features of images and texts into the restoration framework of this application.
[0038] A large-scale fine-grained 3D shape corpus, ViPC-Text, is proposed to explore the complementarity between textual descriptions and the missing semantics and structures of incomplete shapes.
[0039] A novel point cloud inpainting network is designed, which can simultaneously utilize three modal information: point cloud, image and text. Extensive experiments show that the proposed method outperforms the existing state-of-the-art methods in performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0041] Figure 1 It is the point cloud inpainting method proposed in this application, which predicts reliable complete shapes by leveraging rich complementary information from corresponding rendered images and detailed text descriptions;
[0042] Figure 2 This application proposes a CLIP-based
[14] Model's fine-grained text and image-guided point cloud inpainting architecture;
[0043] Figure 3 is a sample instance from the ViPC-Text dataset;
[0044] Figure 4 It is a 3D component corpus assisted by a Large-Scale Language Model (LLM);
[0045] Figure 5 is a qualitative comparison of different methods, including single-modal methods (PCN
[21] , ECG
[43] , VRCNet
[44] , SeedFormer
[27] ) and multimodal methods (XMFNet [6] and methods of the present application;
[0046] Figure 6 In ShapeNet-ViPC [5] Visual comparison of recent point cloud restoration methods and the proposed method on unknown categories;
[0047] Figure 7 Ablation experiments on different text information. DETAILED DESCRIPTION
[0048] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0049] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] In this application, we propose a novel multimodal point cloud inpainting network to enhance shape perception by combining visual and textual information (e.g. Figure 1 As shown in Figure 2). Unlike previous multimodal 3D inpainting methods or dense deep inpainting tasks that only focus on the fusion of point cloud and image pairing, we introduce additional text descriptions to enrich the model's understanding of shape semantics. These text descriptions provide rich semantic information for point clouds, and their effectiveness has been fully verified in the fields of language modeling, text-driven image generation, and text-driven 3D shape generation.
[0051] The point cloud repair method based on multimodal information guidance includes the following steps:
[0052] S1: Build a 3D shape corpus to describe the geometric shapes of objects;
[0053] S2: Obtain missing point cloud data of an object, use a 3D shape corpus to describe the geometric shape of the object that lacks point cloud data, and randomly select rendered images of the object from 24 perspectives to construct a data set in the form of three-dimensional data pairs;
[0054] S3: construct a multimodal point cloud repair network model for repairing the complete geometric shape of an object;
[0055] S4: training the multimodal point cloud restoration network model based on the training set data to obtain a trained multimodal point cloud restoration network model;
[0056] S5: Input the test set data into the trained multimodal point cloud repair network model to repair the geometric shape of objects that lack point cloud data.
[0057] The steps S1 / S2 / S3 / S4 / S5 are performed sequentially;
[0058] The multimodal point cloud restoration network model includes:
[0059] Point cloud encoding module: used to encode the missing point cloud data of the object to obtain the fine-grained features of the point cloud;
[0060] Image encoding module: used to encode the rendered images randomly selected from 24 perspectives of the object to obtain fine-grained features and global features of the image;
[0061] Text encoding module: used to encode the text description information of the geometric shape of the object lacking point cloud data to obtain the fine-grained features and global features of the text;
[0062] Multi-stage feature fusion module: used to fuse the global features of the image and the global features of the text with the fine-grained features of the point cloud, and then perform fine-grained feature fusion with the fine-grained features of the image under the local attention mechanism; perform multi-stage fine-grained multi-head attention feature fusion under the guidance of the global features of the image and the global features of the text, and predict the data of the complete point cloud object corresponding to the input incomplete point cloud data;
[0063] Loss function module: It is used to use the chamfer distance as the loss function to train the data of the same shape as the complete point cloud data, and optimize the multimodal point cloud restoration network model.
[0064] CD: Chamfer Distance is a distance measurement method commonly used to calculate the similarity between two sets of point clouds. It is widely used in computer vision and 3D point cloud processing, such as point cloud matching, shape reconstruction, and generative model evaluation.
[0065] The calculation method of Chamfer Distance is given two sets of point clouds S1S_1S1 and S2S_2S2, each containing a point set:
[0066] ·S1={p1,p2,...,pm}S_1=\{p_1,p_2,...,p_m\}S1={p1,p2,...,pm}
[0067] ·S2={q1,q2,...,qn}S_2=\{q_1,q_2,...,q_n\}S2={q1,q2,...,qn}
[0068] The Chamfer Distance is defined as the sum of the two distances:
[0069]
[0070] in:
[0071] 1. For each point ppp (from S1S_1S1), find the point qqq in S2S_2S2 that is closest to ppp and calculate the square of the Euclidean distance between them.
[0072] 2. For each point qqq (from S2S_2S2), find the point ppp in S1S_1S1 that is closest to qqq and calculate the square of the Euclidean distance between the two.
[0073] Add the two distances together to get the Chamfer Distance.
[0074] When training the network for multimodal point cloud restoration, the selected images can be images from any of the 24 perspectives.
[0075] 24 angle division method: The spherical coordinate system is divided into 24 perspectives, and the angle division is performed using a non-uniform division method. Each perspective is represented by a pair of (theta, phi) angles.
[0076] 1) Horizontal rotation angle (theta): From 0 to 360 degrees, 24 angles are selected as sampling points for horizontal rotation. These sampling points are determined according to specific rules (for example, sampling at certain intervals) and are not evenly distributed.
[0077] 2) Vertical rotation angle (phi): From 0 to 180 degrees, choose the appropriate vertical angle to achieve different viewing angles. The distribution of phi is combined with the distribution of theta angles to determine the position of each viewing angle.
[0078] This application proposes to use pre-trained CLIP
[14] Model, transfers its multimodal knowledge (including visual and textual modalities) to our point cloud restoration network. Compared with ShapeNet-ViPC [5] The number of training samples in the dataset is about 608k, and the pre-trained CLIP
[14] The model uses approximately 400 million training samples, which ensures that more effective text and global visual features can be extracted.
[0079] In order to ensure consistent features are extracted from images, texts, and point clouds, a triple database containing a rich text corpus, i.e., a dataset in the form of triple data pairs, is constructed;
[0080] We use large-scale modeling techniques to construct text-image-point cloud triplets, thereby constructing the ViPC-Text dataset, which focuses on capturing the description information of the geometric parts of 3D shapes. It is worth noting that our method can generate large-scale and concise text descriptions, thus facilitating the generation of text descriptions for arbitrary 3D shapes.
[0081] The goal of this method is to explore the effectiveness of rich multimodal information in point cloud restoration. Our network architecture is as follows Figure 2As shown. [6] Different from this, our network input consists of an incomplete point cloud A text description and a randomly selected rendered image from 24 viewpoints The final prediction is a complete point cloud First, we use CLIP to exploit the strong generalization ability of large-scale models.
[14] Model to extract multimodal information features. Global geometric perception features and Obtained by visual encoder and text encoder respectively (see Figure 2 Bottom). Secondly, we adopt a multi-stage feature fusion strategy in the backbone network to effectively fuse multimodal information. Finally, Figure 2 As shown, we propose an efficient text generation algorithm and construct a corpus containing fine-grained geometric descriptions to further improve the model's understanding of semantics and geometric structures.
[0082] Furthermore, the image encoder module adopts CLIP
[14] The image encoder of the Contrastive Language–Image Pretraining model (which uses large-scale image and text pairing data to train a general visual model through contrastive learning, achieving strong migration capabilities on multiple tasks) and the text encoder module uses CLIP
[14] The text encoder of the model.
[0083] CLIP
[14] It is a pre-trained model consisting of a visual encoder and a text encoder, which is trained by matching about 400 million image-text pairs. The large scale of the training dataset ensures the generalization ability of the model, and the alignment of multimodal features enables the two encoders to extract latent features with unified semantics and visual consistency. These aligned features play an important role in promoting point cloud restoration tasks.
[0084] In this application, we integrate this pre-trained model into our restoration architecture. First, from ShapeNet-ViPC [5] The point cloud rendered image of the dataset is passed as input to CLIP
[14] Meanwhile, the text encoder can receive a specific text prompt, which includes a definition template: “This is a [category]”, where the word “category” should match the ShapeNet-ViPC [5]In addition, richer descriptions such as “This is a classic style, white metal table” can also be used as semantic extensions following the above prompts. The additional descriptions constitute the ViPC-Text dataset we proposed. The corresponding point cloud, image and text examples are as follows: Figure 3 As shown in Figure 2, the extracted latent features can be applied in multiple stages of the restoration model to enhance the fusion of multimodal information. Subsequent quantitative experiments and visual comparison results demonstrate the effectiveness of the extracted latent features.
[0085] The 3D shape corpus constructed to describe the geometric shape of an object is based on a 3D component corpus assisted by a large-scale language model (LLM);
[0086] As we all know, incomplete point cloud data often lacks specific components, so accurate geometric description is required for effective restoration. Although existing work can generate text descriptions for 3D shapes, it is extremely challenging to build a rich, accurate and detailed text corpus to describe a large number of 3D shapes through manual annotation. Therefore, this paper proposes a method that can leverage the existing ShapeNet-ViPC [5] Datasets and BLIP-2
[39] This large-scale language model (LLM) automatically and efficiently builds a fine-grained geometric text corpus for describing 3D shape components. It is worth emphasizing that the entire construction process is automated and efficient.
[0087] like Figure 4 As shown in Figure 2, we take the chair model as an example to show the process of generating a fine-grained geometric corpus.
[39] It receives a single-view rendered image and a textual prompt question as input and generates a textual prompt answer as output.
[39] To be able to describe the fine-grained geometric appearance of 3D shapes, we employ a range of linguistic cues:
[0088] Category Question Answering: To provide a coarse-grained description of shape appearance, we construct a category-dependent question. For example, input: "Please describe the geometric appearance of [chair]?"; output: "This is a brown office [chair]".
[0089] Existence Question Answering: To determine whether a particular component exists in a given 3D shape, we formulate an existence question, which helps us efficiently combine the output descriptions of multiple components. For example, input: "Does [chair] have [legs]?"; output: "yes" or "no".
[0090] Quantity Question Answering: To more accurately describe the number of specific components in a given 3D shape, we designed a question about the number of specific components. For example, input: "How many [legs] does [chair] have?"; output: "This [chair] has four [legs]".
[0091] Appearance question answering: To provide a fine-grained description of component appearance, we construct a component-related question. For example, input: "Please provide some rich geometric description of the [seat] of [chair]?"; output: "The [seat] of this [chair] has a rectangular appearance".
[0092] Due to BLIP-2
[39] The length of the generated text description is usually more than 258 tokens, so we need to use the natural language denoising network Bart
[40] The length of the text description is compressed to 50-58 tokens. In addition, in order to enhance the overall expressiveness of the text, we also manually added a symmetry description at the beginning of each sample. Figure 5 In the figure, the components of the model for the 13 categories in the dataset are shown. This information can be used to generate text-point cloud pairs corresponding to each category.
[0093]
[0094] Furthermore, after the global features of the image and the global features of the text are fused with the fine-grained features of the point cloud, fine-grained features are fused with the fine-grained features of the image under the local attention mechanism; under the guidance of the global features of the image and the global features of the text, multi-stage fine-grained multi-head attention feature fusion is performed to obtain the complete point cloud data with the same shape data as follows:
[0095] Given the global features and extracted from the visual encoder and text encoder of the multimodal feature extraction module, our goal is to fuse these multimodal features into the base network to enhance the base network's understanding of the incomplete 3D shape structure and semantics. Figure 2 As shown in , we adopt a two-stage multimodal feature fusion strategy in the basic network, where the first fusion operation is performed before the fine-grained cross-attention fusion, and the second fusion operation is performed after the attention fusion. Specifically, in the first stage of fusion, we use the global text feature G i and visual features G t These multimodal features are then repeated and combined with the fine-grained point cloud features. The fusion of multimodal features is completed through the shared MLP operation, and the output is recorded as After the first multimodal fusion, we further transform the fine-grained features y 2With the multimodal feature G i and G t In this operation, the cross attention mechanism is used to complete the fusion of fine-grained features, and the output result is recorded as Then, we use the fine-grained features y 2 With the multimodal feature G i and G t The second fusion is performed. Since the second fusion process is similar to the first one, it will not be repeated here. The output of the second stage is recorded as
[0096] Similar to previous work, we use a symmetric version of Chamfer Distance (CD) as the loss function, which is used to measure the difference between the predicted point cloud and the true point cloud. The formula of the loss function is as follows:
[0097]
[0098] Among them: The first item encourages the predicted point cloud Y to be as close as possible to the real point cloud Y gt , and the second term ensures that Y can cover the real point cloud Y gt .
[0099] Embodiment 1:
[0100] 1. Database and Experimental Setup
[0101] In our experiments, we use the widely used ShapeNet-ViPC [5] The proposed method is trained and tested on the dataset, which contains 13 categories of 3D shapes. Among these categories, 8 categories (airplane, lamp, cabinet, chair, table, sofa, boat, and car) are used for training and evaluation, while the remaining 5 categories (bench, display, speaker, firearm, and mobile phone) are reserved exclusively for the final zero-shot evaluation. It is worth noting that the training and evaluation dataset division is the same as XMFNet [6] same.
[0102] To explore the effectiveness of rich text descriptions in our network, we created a new text corpus called ViPC-Text, which is derived from ShapeNet-ViPC [5] The dataset contains 38,328 triplets, each of which contains a text description, a 3D object, and a set of rendered images from 24 viewpoints. Figure 3 The length of the text descriptions in our corpus ranges from 50 to 58 tokens. These concise texts can be directly used as CLIP
[14] The input of the model.
[0103] Corpus Generation: The corpus generation algorithm was implemented on a single NVIDIA RTX A6000. In practice, the algorithm consumes approximately 26GB of GPU memory and takes an average of 1 to 4 seconds to generate a text description.
[0104] Network training: Throughout the training process, we follow the same [6] The same training settings as described in . Our method is implemented using the PyTorch framework, and all experiments are performed on a 4-GPU NVIDIA A800-SXM4-80G cluster with 320GB of GPU memory. The algorithm consumes approximately 160GB of GPU memory on the cluster when running in parallel. We use the Adam optimizer. The multimodal point cloud repair network model is trained for 400 epochs with a batch size of 560 and an initial learning rate of 0.00209.
[0105] Complexity Analysis: We present the complexity analysis in Table 1, which lists the inference time and number of parameters on a single NVIDIA A6000 GPU. Our proposed method is a multimodal point cloud inpainting method. Compared with the other two point cloud inpainting methods, our method utilizes CLIP
[14] This results in a relatively large number of parameters. However, it demonstrates better computational efficiency and requires less processing time. The comparison results show that our method strikes a balance between cost and performance.
[0106] Table 1. Comparison of the complexity of the method in this paper. This application uses THOP (PyTorch-OpCounter) to compare the number of multiplication and addition operations (MACs) and the number of parameters (Params) of the method in this application and the two classic methods.
[0107] Table 1. Comparison of the complexity of our methods.
[0108] Methods <![CDATA[PCN
[21] ]]> <![CDATA[SVDFormer
[56] ]]> Ours MACs(G) 14.708 25.680 22.499 Params(M) 6.864 19.620 121.930
[0109] 2. Quantitative results of point cloud reconstruction
[0110] 2.1 Quantitative comparison on known categories
[0111] The proposed method is compared with other single-modal and multi-modal methods. In order to quantitatively evaluate the performance of various methods, we use Chamfer Distance (CD) and F-score indicators to measure the ShapeNet-ViPC [5]Reconstruction quality on the dataset. Tables 2 and 3 report the quantitative results of various methods. (Considering the limitation of computing resources, we only conducted comparative experiments with the most advanced method SVDFormer on four categories. The experimental results are shown in Table 4, which shows that the method of this application is still highly competitive.) The method of this application shows significant advantages in both single-modal and multi-modal methods, especially when compared with the most advanced multi-modal fusion method. In addition, we observed that XMFNet [6] The performance of is significantly affected by the tuning of training parameters, especially when training by category, and the effect of the “lamp” category is particularly obvious. When multiple categories use the same parameter set, our method consistently outperforms XMFNet. [6] .
[0112] We then present a qualitative comparison, see Figure 6 . With XMFNet [6] Similarly, the proposed method is compared with several single-modality repair methods, including PCN
[21] , ECG
[43] 、VRCNet
[44] In the multimodal comparison, due to ViPC [5] Incomplete code and CSDN [7] Code is not available, we only work with XMFNet [6] Compared with other methods, our method is able to generate more complete shapes and reduce noise. Visual results show that our network is more robust in predicting the semantic information of the missing parts, thus achieving more accurate reconstruction. Specifically, Figure 6 As shown in , our method reconstructs a more complete tail and lampshade for the aircraft and lamp models respectively. Figure 6 For the car model in the paper, our method achieves more accurate wheel reconstruction. More importantly, the rich text information in our work enables us to predict the geometry of the shape more accurately than other methods. Figure 6 It clearly demonstrates the improvement of our thin and long reconstruction compared with other methods.
[0113] Table 2 ShapeNet-ViPC [5] The performance of the state-of-the-art methods is quantitatively compared using the average chamfer distance per point and the average of 2048 points on the eight known categories of the dataset. The best results are marked in bold. * indicates that the code is not available or has not been completed.
[0114] Table 2 ShapeNet-ViPC [5] The performance of the dataset is evaluated against the existing state-of-the-art methods.
[0115]
[0116]
[0117] Table 3 ShapeNet-ViPC [5] The performance of the state-of-the-art methods is quantitatively compared using the mean of the F-Score@0.001 of 2048 points on the eight known categories of the dataset. The best results are marked in bold. * indicates that the code is not available or has not been completed.
[0118] Table 3 ShapeNet-ViPC [5] Dataset evaluation and performance of existing state-of-the-art methods
[0119]
[0120] Table 4 Using ShapeNet-ViPC on four known categories [5] The quantitative comparison of the data set is performed, and the evaluation indicator is the mean Chamfer Distance (×10 -3 ) and mean F-Score @ 0.001 over 2048 points. The best result is highlighted in bold. * indicates code is unavailable or incomplete. “Chamfer Distance / F-Score”.
[0121] Table 4 Using ShapeNet-ViPC on four known categories [5] Quantitative comparison of datasets
[0122] Methods Mean Airplane Chair Sofa SVDFormer 1.439 / 0.778 0.735 / 0.932 2.100 / 0.610 1.484 / 0.793 Ours 1.081 / 0.861 0.539 / 0.971 1.193 / 0.844 1.512 / 0.767
[0123] 2.2 Quantitative comparison on unknown categories
[0124] To verify the generalization ability of the proposed method, we also show the [5] Quantitative and qualitative results on five unknown categories on the dataset. Specifically, for all methods that need to be compared, we first perform [5] We train more general models on the 8 known categories in the dataset, and then evaluate these models on the 4 unknown categories (actually the same as CSDN [7] Similarly, we only tested four unknown categories of objects). Still using the same metrics, Table 5 reports the quantitative results of CD and F-Score respectively. Our method still outperforms the state-of-the-art single-modal methods such as PointAttN and PoinTr
[26] , and the multimodal method XMFNet[6] In addition, we show a visual comparison of our method with other methods, including the Transformer-based method PoinTr
[26] And multi-modal method XMFNet [6] ,like Figure 7 As shown. We can observe that although PoinTr
[26] It performs well on quantitative results, but it does not work well in recovering missing shapes of unknown categories. Then, XMFNet [6] The quality of the repaired shape generated only with the aid of image information is poor. In contrast, the method of the present application can generate clearer results with stronger structural details. This comparison further proves that the method of the present application successfully utilizes the complementary information provided by images and text to complete point cloud repair.
[0125] Table 5 ShapeNet-ViPC [5] For the five unknown categories of the data set, the mean ChamferDistance (×10 -3 ) and mean F-Score@0.001 (2048 points). The best result is highlighted in bold.
[0126] Table 5 ShapeNet-ViPC [5] Quantitative comparison on five unknown categories of the dataset
[0127]
[0128] 2.3 Ablation Experiment
[0129] 2.3.1 Performance Analysis of Multimodal Information
[0130] G i As shown in the first and second rows of Tables 6 and 7, we conducted ablation experiments to evaluate the performance of CLIP
[14] Effectiveness of visual encoder. In this experiment, only CLIP
[14] The extracted visual global feature G i It is important to note that CLIP
[14] The input images used in the benchmark model are the same as those used in the baseline model. [6] ) compared to the introduction of CLIP
[14] The visual global features in ShapeNet-ViPC [5] On the known categories of the dataset, the CD and F-Score indicators are improved by 16.42% and 4.90% respectively. This experiment verifies that CLIP is more efficient than the fine-grained image features used in the baseline model.
[14] The visual module can provide more extensive visual information, thus achieving significant performance improvement. Figure 5 It can be seen that the normal vectors predicted by our method perform better in surface reconstruction, and the reconstructed surface is more accurate and complete, which shows the robustness of our proposed normal estimation method.
[0131] G t Performance analysis of the proposed method. Compared with visual information, the semantic description text of point clouds is easier to obtain in practical applications. In order to verify the effectiveness of introducing text information, we use simple text to guide the process of point cloud completion task. In fact, we only use a simple prompt "This is a [category]" as the input of our multimodal feature extraction module. The experimental results in the first and third rows of Tables 6 and 7 show that this simple global text template is also effective for guiding the point cloud completion process. As shown in the experiments, the CD and F-Score indicators are improved by 13.31% and 3.39% respectively. Therefore, the experimental results show that only a simple text related to the object category can provide effective semantic information.
[0132] G i +G t Performance analysis. Considering the complementarity between multimodal information, we also conducted ablation experiments to verify the effectiveness of the fusion of text and visual information. As shown in the first and sixth rows of Tables 6 and 7, when both visual information and simple text information are fused into our completion network, the CD and F-Score indicators are improved by 18.16% and 5.28%, respectively. This shows that there is indeed complementarity between the three modalities of point cloud, image and text description, and the use of these three modal information can better improve the quality of 3D shape reconstruction. In addition, we also verified the effectiveness of fine-grained geometric description text information. As shown in the last two rows of Tables 6 and 7, compared with simple text, the text description used in our ViPC-Text dataset brought 1.86% and 0.48% improvement in CD and F-Score indicators, respectively. Finally, compared with XMFNet [6] Compared with the previous method, our method improves CD and F-Score by 19.68% and 5.78% respectively. Figure 2 As shown in the red dashed box in , our method has more advantages in structure and detail reconstruction. Our reconstruction results are clearer and sharper, and the point cloud distribution is more uniform.
[0133] Table 6 ShapeNet-ViPC [5]On the dataset, we conducted an ablation experiment on the multimodal information used in the network, and used Chamfer Distance as an indicator to evaluate the performance improvement brought by different multimodal information and different fusion strategies.
[0134] Table 6: Ablation experiments using multimodal information on the ShapeNet-ViPC[5] dataset
[0135]
[0136] Table 7 ShapeNet-ViPC [5] On the dataset, we conducted an ablation experiment on the multimodal information used in the network, and used the mean F-score@0.001 as an indicator to evaluate the performance improvement brought by different multimodal information and different fusion strategies.
[0137] Table 7 shows the ablation experiments using multimodal information on the ShapeNet-ViPC [5] dataset.
[0138]
[0139] 2.3.2 Multi-stage fusion strategy
[0140] As shown in the fourth to sixth rows in Tables 6 and 7, we conducted an ablation experiment to explore multimodal feature fusion strategies. Specifically, we compared three fusion strategies:
[0141] Multimodal feature fusion is performed before fine-grained fusion based on baseline models.
[0142] Multimodal feature fusion is performed after fine-grained fusion based on baseline models.
[0143] Two-stage multimodal feature fusion is performed before and after the fine-grained fusion based on the baseline model.
[0144] Through systematic experimental comparison, we found that the most effective point cloud restoration fusion strategy is to adopt the third method, that is, to perform two-stage multimodal feature fusion before and after fine-grained fusion. This strategy can more fully extract and utilize the complementarity between multimodal information, thereby significantly improving the performance of point cloud restoration.
[0145] Current multimodal point cloud inpainting methods face two major challenges. First, the alignment of cross-modal features needs to be further optimized to enhance the correlation between multimodal features, thereby fully leveraging the complementary advantages of multimodal information. Second, most existing point cloud inpainting methods focus on coarse-grained, large-scale region inpainting, while having difficulties in fine-grained, small-scale region restoration. These challenges highlight the key bottlenecks in multimodal point cloud inpainting research and provide valuable inspiration for future research directions.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
[0147] [1] Z.Ma, S.Liu, A review of 3d reconstruction techniques in civil engineering 575and their applications, Advanced Engineering Informatics 37(2018)163–174.
[0148] [2] P. Mandikal, V B Radhakrishnan, Dense 3d point cloud reconstruction using a deep pyramid network, in: 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE, 2019, pp. 1052–1060.580
[0149] [3]Y.Guo,H.Wang,Q.Hu,H.Liu,L.Liu,M.Bennamoun,Deep learning for 3dpoint clouds:A survey,IEEE transactions on pattern analysis and machineintelligence 43(12)(2020)4338–4364.
[0150] [4]H.Wang,L.Ding,S.Dong,S.Shi,A.Li,J.Li,Z.Li,L.Wang,CAGroup3d:Class-aware grouping for 3d object detection on point clouds,in:A.H.Oh,585A.Agarwal,D.Belgrave,K.Cho(Eds.),Advances in Neural Information ProcessingSystems,2022.URL https: / / openreview.net / forum?id=nLKkHwYP4Au
[0151] [5]X.Zhang,Y.Feng,S.Li,C.Zou,H.Wan,X.Zhao,Y.Guo,Y.Gao,View guidedpoint cloud completion,in:Proceedings ofthe IEEE / CVF Conference 590onComputer Vision and Pattern Recognition,2021,pp.15890–15899.
[0152] [6]E.Aiello,D.Valsesia,E.Magli,Cross-modal learning for image-guidedpoint cloud shape completion,Advances in Neural Information ProcessingSystems 35(2022)37349–37362.
[0153] [7]Z.Zhu,L.Nan,H.Xie,H.Chen,J.Wang,M.Wei,J.Qin,CSDN:595Cross-ModalShape-Transfer Dual-Refinement Network for Point Cloud Completion,IEEETransactions on Visualization&Computer Graphics 30(07)(2024)3545–3563.doi:10.1109 / TVCG.2023.3236061.URL https: / / doi.ieeecomputersociety.org / 10.1109 / TVCG.2023.3236061600
[0154] [8]G.Chen,J.Lin,H.Qin,Uamd-net:A unified adaptive multimodal neuralnetwork for dense depth completion,IEEE Trans.Cir.and Sys.for VideoTechnol.33(10)(2023)5406–5419.doi:10.1109 / TCSVT.2023.3254650.URL https: / / doi.org / 10.1109 / TCSVT.2023.3254650605
[0155] [9]L.Zhao,W.Zheng,Y.Duan,J.Zhou,J.Lu,Sptr:Structure-preservingtransformer for unsupervised indoor depth completion,IEEE Transactions onCircuits and Systems for Video Technology.
[0156]
[10] T.Brown,B.Mann,N.Ryder,M.Subbiah,J.D.Kaplan,P.Dhariwal,A.Neelakantan,P.Shyam,G.Sastry,A.Askell,et al.,Language models are few-shot610learners,Advances in neural information processing systems 33(2020)1877–1901.
[0157]
[11] A.Ramesh,P.Dhariwal,A.Nichol,C.Chu,M.Chen,Hierarchical text-conditional image generation with clip latents,arXiv preprint arXiv:2204.06125.615
[0158]
[12] B.Poole,A.Jain,J.T.Barron,B.Mildenhall,Dreamfusion:Text-to-3dusing 2d diffusion,arXiv preprint arXiv:2209.14988.
[0159]
[13] H.Peng,B.Li,B.Zhang,X.Chen,T.Chen,H.Zhu,Multi-view vision fusionnetwork:Can 2d pre-trained model boost 3d point cloud data-scarce learning?,IEEE Transactions on Circuits and Systems for Video Technology.620
[0160]
[14] A.Radford,J.W.Kim,C.Hallacy,A.Ramesh,G.Goh,S.Agarwal,G.Sastry,A.Askell,P.Mishkin,J.Clark,et al.,Learning transferable visual models fromnatural language supervision,in:International conference on machine learning,PMLR,2021,pp.8748–8763.
[0161]
[15] D.Z.Chen,Q.Wu,M.Nieβner,A.X.Chang,D3net:A speaker-listener625architecture for semi-supervised dense captioning and visual grounding inrgb-d scans(2021).arXiv:2112.01551.
[0162]
[16] Z.Han,C.Chen,Y.-S.Liu,M.Zwicker,Shapecaptioner:Generative captionnetwork for 3d shapes by learning a mapping from parts detected in multipleviews to sentences,2020,pp.1018–1027.doi:10.1145 / 3394171.6303413889.
[0163]
[17] K.Chen,C.B.Choy,M.Savva,A.X.Chang,T.Funkhouser,S.Savarese,Text2shape:Generating shapes from natural language by learningjointembeddings,in:Computer Vision–ACCV 2018:14th Asian Conference on ComputerVision,Perth,Australia,December 2–6,2018,Revised Selected 635 Papers,Part III14,Springer,2019,pp.100–116.
[0164]
[18] Y.LeCun,Y.Bengio,G.Hinton,Deep learning,Nature(2015)436–444doi:10.1038 / nature14539.URL http: / / dx.doi.org / 10.1038 / nature14539
[0165]
[19] Y.Guo,Y.Liu,A.Oerlemans,S.Lao,S.Wu,M.S.Lew,Deep learning for640visual understanding:A review,Neurocomputing 187(2016)27–48.
[0166]
[20] B.Fei,W.Yang,W.-M.Chen,Z.Li,Y.Li,T.Ma,X.Hu,L.Ma,Comprehensivereview ofdeep learning-based 3d point cloud completion processing andanalysis,IEEE Transactions on Intelligent Transportation Systems.
[0167]
[21] W.Yuan,T.Khot,D.Held,C.Mertz,M.Hebert,Pcn:Point completion 645network,in:2018 International Conference on 3D Vision(3DV),IEEE,2018,pp.728–737.
[0168]
[22] L.P.Tchapmi,V.Kosaraju,H.Rezatofighi,I.Reid,S.Savarese,Topnet:Structural point cloud decoder,in:Proceedings ofthe IEEE / CVF Conference onComputer Vision and Pattern Recognition,2019,pp.383–392.650
[0169]
[23] M.Liu,L.Sheng,S.Yang,J.Shao,S.-M.Hu,Morphing and sampling networkfor dense point cloud completion,in:Proceedings of the AAAI conference onartificial intelligence,Vol.34,2020,pp.11596–11603.
[0170]
[24] M.-H.Guo,J.-X.Cai,Z.-N.Liu,T.-J.Mu,R.R.Martin,S.-M.Hu,Pct:Pointcloud transformer,Computational Visual Media(2021)187–199doi:10.6551007 / s41095-021-0229-5.URL http: / / dx.doi.org / 10.1007 / s41095-021-0229-5
[0171]
[25] X.Pan,Z.Xia,S.Song,L.E.Li,G.Huang,3d object detection withpointformer,in:2021 IEEE / CVF Conference on Computer Vision and PatternRecognition(CVPR),2021.doi:10.1109 / cvpr46437.2021.00738.660 URL http: / / dx.doi.org / 10.1109 / cvpr46437.2021.00738
[0172]
[26] X.Yu,Y.Rao,Z.Wang,Z.Liu,J.Lu,J.Zhou,Pointr:Diverse point cloudcompletion with geometry-aware transformers,in:Proceedings of the IEEE / CVFinternational conference on computer vision,2021,pp.12498–12507.665
[0173]
[27] P.Xiang,X.Wen,Y.-S.Liu,Y.-P.Cao,P.Wan,W.Zheng,Z.Han,Snowflakenet:Point cloud completion by snowflake point deconvolution with skiptransformer,in:Proceedings of the IEEE / CVF international conference on computer vision,2021,pp.5499–5509.
[0174]
[28] H.Zhou,Y.Cao,W.Chu,J.Zhu,T.Lu,Y.Tai,C.Wang,Seedformer:Patch670seeds based point cloud completion with upsample transformer,in:EuropeanConference on ComputerVision,Springer,2022,pp.416–432.
[0175]
[29] Z.Chen,F.Long,Z.Qiu,T.Yao,W.Zhou,J.Luo,T.Mei,Anchorformer:Pointcloud completion from discriminative nodes.
[0176]
[30] S.Li,P.Gao,X.Tan,M.Wei,G.Pointr,Proxyformer:Proxy alignment 675assisted point cloud completion with missing part sensitive transformer
[0177]
[31] L.Tan,X.Lin,D.Niu,D.Wang,M.Yin,X.Zhao,Projected generative ad13versarial network for point cloud completion,IEEE Transactions on Circuitsand Systems for Video Technology 33(2)(2022)771–781.
[0178]
[32] H.Xiao,Y.Li,W.Kang,Q.Wu,Distinguishing and matching-aware unsu680pervised point cloud completion,IEEE Transactions on Circuits and Systems forVideo Technology.
[0179]
[33] A.Mao,Y.Tang,J.Huang,Y.He,Dmf-net:Image-guided point cloudcompletion with dual-channel modality fusion and shape-aware upsamplingtransformer,arXiv preprint arXiv:2406.17319.685
[0180]
[34] X.Zhu,R.Zhang,B.He,Z.Zeng,S.Zhang,P.Gao,Pointclip v2:Adaptingclip for powerful 3d open-world learning,arXiv preprint arXiv:2211.11682.
[0181]
[35] D.Hegde,J.M.J.Valanarasu,V.M.Patel,Clip goes 3d:Leveraging prompttuning for language grounded 3d recognition,arXiv preprint arXiv:2303.11313.690
[0182]
[36] L.Xue,N.Yu,S.Zhang,J.Li,R. J.Wu,C.Xiong,R.Xu,J.C.Niebles,S.Savarese,Ulip-2:Towards scalable multimodal pre-training for 3dunderstanding,arXiv preprint arXiv:2305.08275.
[0183]
[37] Z.Han,C.Chen,Y.-S.Liu,M.Zwicker,Shapecaptioner:Generative captionnetwork for 3d shapes by learning a mapping from parts detected in multiple695 views to sentences,in:Proceedings of the 28th ACM InternationalConference on Multimedia,2020,pp.1018–1027.
[0184]
[38] D.Zhenyu Chen,Q.Wu,M.Nieβner,A.X.Chang,D3net:A unifiedspeakerlistener architecture for 3d dense captioning and visual grounding,arXiv eprints(2021)arXiv–2112.700
[0185]
[39] J.Li,D.Li,C.Xiong,S.Hoi,Blip:Bootstrapping language-imagepretraining for unified vision-language understanding and generation.
[0186]
[40] M.Lewis,Y.Liu,N.Goyal,M.Ghazvininejad,A.Mohamed,O.Levy,V.Stoyanov,L.Zettlemoyer,Bart:Denoising sequence-to-sequence pretraining fornatural language generation,translation,and comprehension,705 arXiv preprintarXiv:1910.13461.
[0187]
[41] T.Groueix,M.Fisher,V.G.Kim,B.C.Russell,M.Aubry,A papier-mach^e′approach to learning 3d surface generation,in:Proceedings ofthe IEEEconference on computer vision and pattern recognition,2018,pp.216–224.
[0188]
[42] Y.Yang,C.Feng,Y.Shen,D.Tian,Foldingnet:Point cloud auto-encodervia 710 deep grid deformation,in:Proceedings ofthe IEEE conference oncomputer vision and pattern recognition,2018,pp.206–215.
[0189]
[43] L.Pan,Ecg:Edge-aware point cloud completion with graphconvolution,IEEE Robotics and Automation Letters 5(3)(2020)4392–4398.
[0190]
[44] L.Pan,X.Chen,Z.Cai,J.Zhang,H.Zhao,S.Yi,Z.Liu,Variationalrelational 715 point completion network,in:Proceedings of the IEEE / CVFconference on computer vision and pattern recognition,2021,pp.8524–8533.
[0191]
[45] Z.Huang,Y.Yu,J.Xu,F.Ni,X.Le,Pf-net:Point fractal network for 3dpoint cloud completion,in:Proceedings of the IEEE / CVF conference on computervision and pattern recognition,2020,pp.7662–7670.720
[0192]
[46] H.Xie,H.Yao,S.Zhou,J.Mao,S.Zhang,W.Sun,Grnet:Gridding residualnetwork for dense point cloud completion,in:Computer Vision–ECCV 2020:16thEuropean Conference,Glasgow,UK,August 23–28,2020,Proceedings,Part IX,Springer,2020,pp.365–381.
[0193]
[47] J.Wang,Y.Cui,D.Guo,J.Li,Q.Liu,C.Shen,Pointattn:You only need725attention for point cloud completion,arXiv preprint arXiv:2203.08485.
[0194]
[48] W.Zhang,Z.Dong,J.Liu,Q.Yan,C.Xiao,et al.,Point cloud completionvia skeleton-detail transformer,IEEE Transactions on Visualization andComputer Graphics.
[0195]
[49] A.Paszke,S.Gross,F.Massa,A.Lerer,J.Bradbury,G.Chanan,T.Killeen,730 Z.Lin,N.Gimelshein,L.Antiga,et al.,Pytorch:An imperative style,highperformance deep learning library,Advances in neuralinformationprocessing systems 32.
[0196]
[50] D.P.Kingma,J.Ba,Adam:A method for stochastic optimization,arXivpreprintarXiv:1412.6980.735
[0197]
[51] Z.Zhu,H.Chen,X.He,W.Wang,J.Qin,M.Wei,Svdformer:Complementingpoint cloud via self-view augmentation and self-structure dual-generator,in:Proceedings of the IEEE / CVF International Conference on ComputerVision,2023,pp.14508–14518.
[0198]
[52] C.-H.Shen,H.Fu,K.Chen,S.-M.Hu,Structure recovery by partassembly,740ACM Transactions on Graphics(Proceedings of ACM SIGGRAPH Asia201231(6)(2012)180:1–180:11.
[0199]
[53] S.Choi,Q.-Y.Zhou,S.Miller,V.Koltun,A large dataset of objectscans,arXiv:1602.02481.
[0200]
[54] H.Jun,A.Nichol,Shap-e:Generating conditional 3d implicitfunctions,745 arXiv preprint arXiv:2305.02463.
[0201]
[55] A.Nichol,H.Jun,P.Dhariwal,P.Mishkin,M.Chen,Point-e:A system forgenerating 3d point clouds from complex prompts,arXiv preprint arXiv:2212.08751.
[0202]
[56] Zhu,Zhe,et al."Svdformer:Complementing point cloud via self-viewaugmentation and self-structure dual-generator."Proceedings of the IEEE / CVFInternational Conference on ComputerVision.2023。
Claims
1. A point cloud restoration method based on multimodal information guidance, characterized in that: The following steps are involved: Build a 3D shape corpus to describe the geometry of objects; Obtaining point cloud data of objects that lack point cloud data, constructing a 3D shape corpus that describes the geometric shapes of objects that lack point cloud data, and constructing a dataset in the form of triple data pairs using randomly selected rendered images of the object under 24 viewing angles; Construct a multimodal point cloud repair network model for complete geometric repair of an object; The multimodal point cloud restoration network model is trained based on the training set data to obtain a trained multimodal point cloud restoration network model; The test set data is input into the trained multimodal point cloud repair network model to repair the geometric shape of objects that lack point cloud data.
2. The point cloud restoration method based on multimodal information guidance according to claim 1, characterized in that: The multimodal point cloud restoration network model includes: Point cloud encoding module: used to encode the point cloud data input by the network to obtain fine-grained features of the point cloud; Image encoding module: used to encode the randomly selected rendered images of the object under 24 viewing angles to obtain fine-grained features of the image; CLIP module: used to encode the randomly selected rendering images of the object under 24 viewing angles to obtain the global features of the image and to encode the text description information of the geometric shape description of the object lacking point cloud data to obtain the global features of the text; Multi-stage feature fusion module: used to fuse the global features of the image and the global features of the text with the fine-grained features of the point cloud, and then perform fine-grained feature fusion with the fine-grained features of the image under the local attention mechanism; perform multi-stage fine-grained multi-head attention feature fusion under the guidance of the global features of the image and the global features of the text, and predict the object of the complete point cloud data corresponding to the object of the input incomplete point cloud data; Loss function module: The chamfer distance is used as the loss function to calculate the difference between the predicted point cloud and the complete point cloud in the dataset to optimize the multimodal point cloud restoration network model.
3. The point cloud repair method based on multimodal information guidance according to claim 2, characterized in that: The image encoder module adopts the image encoder of the CLIP model, and the CLIP module adopts the text encoder of the CLIP model.
4. The point cloud restoration method based on multimodal information guidance according to claim 1, characterized in that: After fusing the global features of the image and the global features of the text with the fine-grained features of the point cloud, the fine-grained features are fused with the fine-grained features of the image under the local attention mechanism; under the guidance of the global features of the image and the global features of the text, multi-stage fine-grained multi-head attention feature fusion is performed to obtain the complete point cloud data with the same shape data as follows: In the first stage of fusion, the global text feature G t and the global visual feature G i After concatenation, these multimodal features are repeated and combined with the fine-grained point cloud features To splice; The fusion of multimodal features is completed through the shared MLP operation, and the output is recorded as After the first multimodal fusion, the fine-grained image features F i It is fused with the multimodal feature y1; in this operation, the cross attention mechanism is used to complete the fusion of fine-grained features, and the output result is recorded as Then, the fine-grained feature y2 is used together with the multimodal global feature G i and G t The second fusion process is the same as the first stage operation process. The output of the second stage is recorded as 5. The point cloud repair method based on multimodal information guidance according to claim 2, characterized in that: The formula of the loss function is as follows: Among them: The first item encourages the predicted point cloud Y to be as close as possible to the real point cloud Y gt , and the second term ensures that Y can cover the real point cloud Y gt .
6. The point cloud restoration method based on multimodal information guidance according to claim 1, characterized in that: The construction of the text corpus of object geometry adopts BLIP-2 to receive a single-view rendered image and a text prompt question as input, and generate a text prompt answer as output. The text prompt question adopts a series of language prompts as follows: Category question answering: used to give a coarse-grained description of the shape appearance; Existence question answering: used to determine whether a specific component exists in a given 3D shape; Quantity question answering: used to accurately describe the quantity of a specific component in a given 3D shape; Appearance question answering: used to give a fine-grained description of the component's appearance, constructing a question related to the component.
7. A point cloud repair device based on multimodal information guidance, characterized in that: include: Building Module I: used to build a 3D shape corpus for describing the geometric shapes of objects; Acquisition module: used to acquire the missing point cloud data of an object, use the 3D shape corpus to describe the geometric shape of the object that lacks point cloud data, and construct a dataset in the form of three-dimensional data pairs using randomly selected rendering images of the object from 24 perspectives; Building Module II: Multimodal point cloud restoration network model for restoring the complete geometric shape of an object; The training module trains the multimodal point cloud restoration network model based on the training set data to obtain a trained multimodal point cloud restoration network model; Implementation module: used to input the test set data into the trained multimodal point cloud repair network model to repair the geometric shape of objects that lack point cloud data.