Bidirectional cross-modal image guide point cloud restoration method with multi-scale progressive refinement
Through a multi-scale progressively refined bidirectional cross-modal image-guided point cloud repair method, the two-way interaction between images and point cloud data is solved, and the traditional method is difficult to restore details in complex scenarios, achieving a more refined and natural point cloud repair effect.
Patent Information
- Application Number
- CN202510542201.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Traditional point cloud repair methods are difficult to effectively restore details in complex scenarios, especially when there are a large number of missing data or sparse areas, the repair effect is poor.
A two-way cross-modal image-guided point cloud repair method with multi-scale progressive refinement is adopted. Through the two-way interaction compensation module and the point cloud refinement module, a more refined and natural point cloud model is gradually generated through the two-way interaction between images and point cloud data.
It significantly improves the accuracy and nature of point cloud repair, and can generate more accurate and complete point clouds in complex scenarios and large-scale absences, improving data quality and application performance.
Smart Images

Figure CN120070269A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision for artificial intelligence drones or embodied robots. Specifically, it relates to a bidirectional cross-modal image-guided point cloud repair method with multi-scale progressive refinement. Background Art
[0002] During the flight of a drone, it is necessary to obtain three-dimensional information of the surrounding environment in real time to build an environmental model and perform autonomous navigation. This method can obtain image and point cloud data through sensors such as cameras and lidar carried by the drone, and use a bidirectional interaction compensation module to fuse the texture and color information in the image with the geometric depth information in the point cloud to generate an environmental point cloud model. When an embodied robot interacts with humans or the environment, it needs to accurately perceive the surrounding space information. This method can generate a refined and high-quality point cloud model, enabling the robot to more accurately identify and understand information such as the shape, position, and size of surrounding objects. An embodied robot needs to understand the scene it is in to make reasonable decisions and actions. This method creates a global representation of rich semantic information by fusing image and point cloud data. During the autonomous navigation process, an embodied robot needs to perceive the surrounding environment in real time and plan a path to avoid obstacles. The point cloud model generated by this method can provide three-dimensional space information for the robot.
[0003] Point cloud repair technology aims to restore missing or sparse point cloud data caused by occlusion, sensor noise, or resolution limitations. With the wide application of three-dimensional scanning technology, the demand for point cloud data in fields such as autonomous driving, robot vision, three-dimensional reconstruction, and virtual reality is increasing. However, the inevitable missing problems in the point cloud acquisition process affect the integrity and accuracy of the data. Therefore, point cloud repair has become a key technology to improve data quality and application performance. Point cloud repair can not only improve the integrity and accuracy of the data, enhance three-dimensional modeling and environmental perception capabilities, but also be of great significance for improving the safety of autonomous driving systems, optimizing the immersion in virtual reality, and supporting the processing of large-scale data sets. By repairing the missing point cloud data, the accuracy and reliability of subsequent applications can be guaranteed, especially in complex environments. In the future, point cloud repair will combine emerging technologies such as deep learning to continuously improve the repair accuracy and computational efficiency, promoting the further development of computer vision, artificial intelligence, and robotics. But there are still some problems with point cloud repair at present: (1) Traditional point cloud repair methods (such as algorithms based on interpolation, surface fitting, or neighborhood reconstruction) mainly rely on the local structure or spatial neighborhood information of the point cloud. Although these methods are simple and have high computational efficiency, they often cannot effectively restore details in complex scenes, especially when there is a large amount of missing data or sparse regions, and the repair effect is poor.
[0004] (2) Traditional methods mostly rely on fixed-scale repair strategies, which may perform well at a certain scale but cannot effectively handle details at different levels. Especially in the case of complex geometric structures or large-scale missing parts, conventional methods may not be able to restore global and local details simultaneously.
[0005] (3) Traditional repair methods may cause detail loss or generate unnatural repaired areas when dealing with large-scale missing regions. For example, simple interpolation methods may make the repaired area too smooth, losing details and structures in the real scene. Summary of the Invention
[0006] The present invention aims to provide a two-way cross-modal image-guided point cloud repair method with multi-scale progressive refinement, aiming to use two-way cross-modal and image information to guide point cloud repair, and use the multi-scale progressive refinement strategy to greatly improve the repair effect. Especially when facing large-scale missing parts, complex scenes or dynamic targets, the repair results are more refined and natural.
[0007] To solve the above problems, the technical solutions adopted by the present invention are as follows: The present invention provides a two-way cross-modal image-guided point cloud repair method with multi-scale progressive refinement, including the following steps: Step 1: Obtain view images and incomplete point clouds , and convert the incomplete point clouds into depth maps through perspective projection; Step 2: Based on the two-way interactive compensation module with multi-scale progressive refinement, compensate for the missing geometric depth information in the view image modality according to the depth map, and compensate for the missing texture and color information in the point cloud modality by the view image modality to obtain the final enhanced feature representation ; Step 3: Decode the final enhanced feature representation based on the coarse point cloud generation module , and generate a coarse point cloud representing the approximate shape of the completed object ; Step 4: Use the point cloud refinement module to refine the coarse point cloud , and gradually generate the predicted refined point cloud through the fusion of global shape and local structure information guided by two-stream features .
[0008] As a further description of the above technical solution, in step 1, a set of points need to be projected onto six orthogonal planes to generate a sparse depth image; for each point in the point cloud, perspective projection is used to derive its 2D coordinates on the image plane; the coordinates are discretized, and the depth value at each image position is determined by the weighted average of nearby pixels.
[0009] As a further description of the above technical solution, in step 2, the bidirectional interaction compensation module includes a point encoder, a depth encoder, and an image encoder pre-trained on the ImageNet image dataset; among them, the point encoder consists of three layers of set abstraction layers, and extracts point cloud features from point clouds at different levels and scales through hierarchical downsampling ; The depth encoder uses the ResNet-18 model to extract depth map features from the depth map ; The image encoder is used to extract perspective image features from the view images .
[0010] As a further description of the above technical solution, the bidirectional interaction compensation module includes two cross-modal Transformers; the first cross-modal Transformer establishes a semantic association between the depth map features and the perspective image features , and the specific implementation is represented by the following formula: Where: represents the fused feature obtained by concatenating and ; represents the position embedding; represents the projection viewpoint; , , represents a linear layer; , , represents the result after linear transformation, which is used for the operation of the self-attention mechanism; represents layer normalization; represents a feed-forward network composed of two linear layers; The second cross-modal Transformer establishes a semantic association between the enhanced view feature representation and the point cloud feature , and the specific implementation formula is as follows: Where: Represents the enhanced view feature and the point cloud feature are connected to obtain the fused feature; Represents the finally obtained enhanced point cloud feature representation; and respectively represent the fused features after primary and secondary normalization; , , represent linear layers; , , , , , represent the results after linear transformation and are used for the operations of the self-attention mechanism.
[0011] As a further description of the above technical solution, the enhanced point cloud feature representation is connected to the initial point cloud feature to obtain the final enhanced feature representation , and the final enhanced feature representation is used as the input of the coarse point cloud generation module and the point cloud refinement module.
[0012] As a further description of the above technical solution, in step 3, the coarse point cloud generation module uses transposed convolution to decode the features of each point from the final enhanced feature representation .
[0013] As a further description of the above technical solution, the coarse point cloud generation module adopts a two-branch method. One branch uses the NodeShuffle module to generate new points from the latent space by combining the spatial information provided by adjacent points, and the other branch uses the self-attention mechanism to simulate the long-range dependencies between single-point features and capture the correlations between points to generate new points.
[0014] As a further description of the above technical solution, the new points from the two branches are combined together through element-wise addition to obtain the newly generated single points. The newly generated single points are connected to the final enhanced feature representation and then input into a multi-layer perceptron to generate the transitional coarse point cloud . The transitional coarse point cloud is connected to the incomplete point cloud , and the farthest point sampling algorithm is used to obtain the coarse point cloud .
[0015] As a further description of the above technical solution, the point cloud refinement module includes a global shape prediction sub-module, a local structure optimization sub-module, a DFGF sub-module, and an up-sampling sub-module; the global shape prediction sub-module is used to generate the transitional coarse point cloud according to the feature representation and the coarse point cloud Generate a feature representation representing the global shape ; The local structure optimization sub-module is used to iteratively aggregate and refine the local features of the incomplete point cloud based on edge convolution and farthest point sampling to obtain a feature representation suitable for enhancing the local structure ; The DFGF sub-module is used to obtain local structure features and based on the feature representation and global shape features ; The upsampling sub-module is used to rely on the fused feature representation of the local structure features and global shape features to generate a coordinate offset through MLP and reshaping operations, and use the coordinate offset to finely adjust the coarse point cloud and finally obtain a complete point cloud .
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The bidirectional interaction compensation module is introduced, which effectively solves the cross-modal data challenge in image-guided point cloud completion. By enhancing the feature representation, this module bridges the information gap between the point cloud and image modalities.
[0017] (2) The point cloud refinement module is proposed. This module adopts a two-stream structure to optimize the point cloud generation process, significantly enhancing the generation accuracy of the overall shape and local details of the point cloud, and obtaining a more accurate and complete point cloud.
[0018] (3) The proposed bidirectional cross-modal image-guided point cloud repair method with multi-scale progressive refinement successfully alleviates the sensitivity of previous models to input views, and can reasonably generate a complete point cloud with a fine-grained semantic structure, achieving the best performance on the benchmark dataset.
[0019] To make the above objects, features, and advantages of the present invention more obvious and understandable, specific embodiments of the present invention are hereinafter given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 is the overall framework diagram of the method of the present invention; Figure 2It is the diagram of the bidirectional interaction compensation module of the method of the present invention; Figure 3 It is the diagram of the coarse point cloud generation module of the method of the present invention; Figure 4 It is the diagram of the point cloud refinement module of the method of the present invention; Figure 5 It is the diagram of the experimental result of the method of the present invention. Detailed implementation manners
[0022] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention.
[0023] The embodiments of the present invention provide a bidirectional cross-modal image-guided point cloud repair network BiMPR-Net with multi-scale progressive refinement as Figure 1 shown. The BiMPR-Net model includes a bidirectional interaction compensation module, a coarse point cloud generation module (CPGM) and a point cloud refinement module. The whole completion process follows the strategy of gradually generating the predicted point cloud from coarse to fine.
[0024] The embodiments of the present invention provide a bidirectional cross-modal image-guided point cloud repair method with multi-scale progressive refinement, which is implemented based on the Figure 1 shown BiMPR-Net network and includes the following steps: Step 1, obtain the view image and the incomplete point cloud , and convert the incomplete point cloud into a depth map through perspective projection.
[0025] In this embodiment, a set of points needs to be projected onto six orthogonal planes to generate a sparse depth image. For each point in the point cloud, perspective projection is used to derive its 2D coordinates on the image plane. The coordinates are discretized, and the depth value at each image position is determined by the weighted average of nearby pixels, which helps to reduce noise by averaging adjacent pixel values.
[0026] Step 2, use the bidirectional interaction compensation module to compensate for the missing information in the point cloud modality and the depth map modality, where the depth map obtained by projecting the point cloud is used to compensate for the missing geometric depth information in the view image modality, and conversely, the view image modality is used to compensate for the missing texture and color information in the point cloud modality. This process effectively fuses cross-modal data by enhancing the feature representation and creates a more comprehensive and richer global representation of semantic information.
[0027] In this embodiment, the left part of the bidirectional interaction compensation module is as Figure 1 shown andFigure 2 As shown, it includes three encoders for extracting key feature information from different modal data, namely a point encoder for processing incomplete point cloud data, a depth encoder for processing depth maps, and an image encoder for processing view images.
[0028] Among them, inspired by PointNet++, the point cloud encoder consists of three layers of grouped sampling and a point cloud network, and refines and extracts key information from point clouds at different levels and scales through hierarchical downsampling to obtain point cloud features. .
[0029] The depth encoder uses the classic ResNet-18 model, and the number of input channels of the first convolutional layer is adjusted to 1 to obtain the final depth map features with a dimension of 128. .
[0030] Similarly, the image encoder uses ResNet-18 as the backbone to extract view image features from view images. , and the dimension is also 128.
[0031] To address the potential limitations of model training and generalization performance due to limited data, the present invention pre-trains the image encoder on the ImageNet image dataset to transfer visual modal knowledge from a large-scale dataset to the BiMPR-Net model proposed by the present invention, enhancing the generalization performance of the image encoder and making it better at capturing precise semantic information in the view image feature representation.
[0032] The present invention uses a cross-modal Transformer to establish semantic associations between depth map features and view image features , and uses the depth map obtained by point cloud projection to make up for the lack of spatial position and distance information in a single view image. The specific implementation of this cross-modal Transformer is represented by the following formulas: (1) (2) (3) (4) Where: represents the fused features obtained by concatenating and ; represents the position embedding; represents the projection view point; , , represents the linear layer; , , represents the result after linear transformation and is used for the operation of the self-attention mechanism; denotes layer normalization; denotes a feed-forward network composed of two linear layers. The present invention utilizes self-attention to learn fused features in the long-distance context information. During this process, the present invention conveys information about the projection viewpoints to the model through positional encoding to help the model understand the spatial relationships and viewpoint differences between features. By obtaining attention weights, the model selectively focuses on different parts of the fused features to obtain more discriminative fused features . To further enhance the feature representation ability, the present invention introduces a feed-forward network to enable the model to learn more complex and abstract feature representations . can be regarded as an enhanced view feature representation after compensating for depth image information.
[0033] Subsequently, the enhanced view feature representation obtained after passing through the first cross-modal Transformer is fed into the second cross-modal Transformer to establish semantic associations with the point cloud features . This interaction process compensates for the missing appearance and visual attribute information in the incomplete point cloud. Considering the increased complexity of dimensions and feature interactions, compared with the previous cross-modal Transformer, additional self-attention layers are introduced here. This enhancement aims to better learn the correlations between cross-modal features. Since the point cloud features inherently contain position information, the present invention omits position embeddings in this context. The formula is expressed as follows: (5) (6) (7) (8) (9) (10) Where: represents the enhanced view feature connected with the point cloud feature to obtain the fused feature; and respectively represent the fused features after primary and secondary normalization; , , represent linear layers; , , , , , represents the result after linear transformation and is used for the operation of the self-attention mechanism; denotes a feed-forward network composed of two linear layers.
[0034] After passing through the second cross-modal Transformer, an enhanced point cloud feature representation is obtained .
[0035] Finally, the present invention connects the enhanced point cloud feature representation with the initial point cloud feature to obtain the final enhanced feature representation of the model , which is the global representation. This feature representation is obtained by seamlessly fusing the initial point cloud feature with the features supplemented and enriched through cross-modal interaction and model learning, so as to more accurately capture and express the essential features and key information in the point cloud data.
[0036] Step 3, decoding the final enhanced feature representation based on the Coarse Point Cloud Generation Module (CPGM) to generate a coarse point cloud representing the approximate shape of the completed object ( = 1024).
[0037] As Figure 1 in the middle part of Figure 3 and shown, different from the fully connected decoder that focuses on capturing the global geometry or the folding-based decoder that helps represent the local geometry, the embodiment of the present invention initially uses transposed convolution (deconvolution) to decode the feature of each point from the final enhanced feature representation .
[0038] For each point feature , the present invention connects it with the repeated , and then integrates it through a convolutional layer. The feature representation is connected with the intermediate feature map, and the feature representation of each point is enriched by merging the global shape information and local details. This fusion improves the accuracy and structural integrity of the generated coarse point cloud, thereby promoting more effective feature learning. Therefore, the generated point cloud not only retains the overall shape but also captures the detailed information. In addition, this feature fusion guides the decoding process, enabling the decoder to simultaneously focus on local features and the broader object structure. This dual focus helps the model generate a fine-grained point cloud consistent with the overall object shape, ensuring a balanced representation of details and structure. This method can generate a more accurate and structurally reasonable point cloud.
[0039] To increase the number of single-point features for generating new points, the present invention adopts a dual-branch method. One branch utilizes NodeShuffle (a module for point cloud upsampling, whose core idea is to encode the spatial information of point neighborhoods using a graph convolutional network (GCN) to generate new points) to generate new points from the latent space by combining the spatial information provided by adjacent points. The other branch employs a self-attention mechanism to effectively model the long-range dependencies between single-point features, capturing the correlations between points to generate new points.
[0040] By combining the advantages of the two branches through element-wise addition, that is, combining the new points of the two branches to obtain the newly generated single points, it promotes the integration of local spatial information and global correlation learning, thereby promoting the generation of new point features.
[0041] Subsequently, the newly generated single points are connected to the final enhanced feature representation .
[0042] To further preserve the structural information of the input point cloud, the transitional coarse point cloud is connected to the incomplete point cloud , and the farthest point sampling algorithm is used to obtain the final coarse point cloud of size . .
[0043] Step 4: Use the point cloud refinement module to refine the coarse point cloud generated in Step 3 , and gradually generate the predicted refined point cloud by fusing the global shape and local structure information guided by the dual-stream features, so as to improve the quality and accuracy of the finally completed point cloud.
[0044] As Figure 1 shown in the right part of Figure 4 , the point cloud refinement module includes four key sub-modules: a global shape prediction sub-module, a local structure optimization sub-module, two stacked DFGF (dual-stream feature-guided fusion) sub-modules, and an upsampling sub-module.
[0045] The input of the global shape prediction sub-module consists of the enhanced feature representation and the coarse point cloud . Features are extracted from using an MLP (multi-layer perceptron). Then the enhanced feature representation is input into another MLP to reduce the feature dimension and capture the most important feature information. After that, the outputs of these two MLPs are connected to obtain the feature representation representing the global shape.
[0046] The local structure optimization sub-module takes the incomplete point cloud As the input, through two layers of EdgeConv (Edge Convolution) and one layer of FPS (Farthest Point Sampling), the local features of the incomplete point cloud are iteratively aggregated and refined to obtain a feature representation suitable for enhancing the local structure .
[0047] Feature representation and are both input into the DFGF sub-module to facilitate the interaction and fusion of global and local information.
[0048] As Figure 4 shown in Figure (e) of: First, the present invention adopts a self-attention mechanism to learn the context information in the global shape and local structure feature representations, and the outputs are represented as and ; Subsequently, group convolution is used to perform position encoding on , and the results before and after convolution are directly added as the source of query ( ) in the cross-attention mechanism, while serves as the source of key ( ) and value ( ) in the cross-attention mechanism. The model uses the cross-attention mechanism to achieve the guided fusion of two-stream features and semantic correlations. By using the global shape information as , the local shape information as , the guided fusion of two-stream features is realized. The structural information serves as and , and this process guides the model to focus on the local information related to the global shape, thereby extracting more relevant and meaningful features from the local structure. Further processing involves enhancement through cross-attention, as well as the locally structured features of global information fusion, to obtain . Finally, self-attention is used again to learn the relationship in the feature representation , and the output is further enhanced through to obtain the final global shape feature representation, denoted as .
[0049] Connect the local structure features output by the second DFGF sub-module and the global shape features to form a fused feature representation , which can capture both the global shape and local details.
[0050] The upsampling sub-module relies on the feature information of the global shape and local structure, that is, the fused feature representation , to generate a set of coordinate offsets through MLP and reshaping operations, and uses the coordinate offsets for the coarse point cloud Perform fine adjustment to obtain a finer and more accurate complete point cloud , that is, the entire repair process is completed.
[0051] The method of the present invention is different from the traditional shape completion methods that directly infer the complete shape from the incomplete input point cloud. The present invention utilizes the additional modal information provided by the view images to help obtain a more accurate shape of the incomplete point cloud. Specifically, given an incomplete point cloud and any view image as inputs, utilize the information from both modalities to infer the missing parts of the incomplete point cloud and generate a complete point cloud . It should be noted that the number of points in the input and output point clouds of the present invention is the same.
[0052] A large number of quantitative and qualitative experiments conducted on the benchmark dataset show that the BiMPR-Net model achieves state-of-the-art performance. The repair results are good, as Figure 5 shown, demonstrating the excellence and necessity of the present invention.
[0053] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A bidirectional cross-modal image-guided point cloud inpainting method with multi-scale progressive refinement, characterized in that: The following steps are involved: Step 1: Get the view image and incomplete point cloud , the incomplete point cloud The depth map is obtained through perspective projection transformation; Step 2: Based on the bidirectional interactive compensation module with multi-scale progressive refinement, the missing geometric depth information in the view image modality is compensated according to the depth map, and the missing texture and color information in the point cloud modality is compensated by the view image modality to obtain the final enhanced feature representation. ; Step 3: Decode the final enhanced feature representation based on the coarse point cloud generation module , generating a coarse point cloud representing the approximate shape of the completed object ; Step 4: Use the point cloud refinement module to refine the coarse point cloud Perform refinement processing, guide the fusion of global shape and local structure information through dual-stream features, and gradually generate predicted refined point clouds .
2. The bidirectional cross-modal image-guided point cloud restoration method according to claim 1, characterized in that: In step 1, a set of points needs to be projected onto six orthogonal planes to generate a sparse depth image; for each point in the point cloud, perspective projection is used to derive its 2D coordinates on the image plane; the coordinates are discretized, and the depth value of each image position is determined by the weighted average of nearby pixels.
3. The bidirectional cross-modal image-guided point cloud restoration method according to claim 1, characterized in that: In step 2, the bidirectional interaction compensation module includes a point encoder, a depth encoder, and an image encoder pre-trained on the ImageNet image dataset; the point encoder consists of three layers of set abstraction layers, which extract point cloud features from point clouds of different levels and scales through hierarchical downsampling. ; The depth encoder uses the ResNet-18 model to extract depth map features from the depth map ; The image encoder is used to extract the view image features from the view image .
4. The bidirectional cross-modal image-guided point cloud restoration method according to claim 3, characterized in that: The bidirectional interaction compensation module includes two cross-modal Transformers; the first cross-modal Transformer is based on the depth map feature and view image features A semantic association is established between them, and the specific implementation is expressed by the following formula: in: Indicates that by and The fused features obtained by concatenation; Represents positional embedding; Represents the projection viewpoint; , , represents the linear layer; , , Represents the result after linear transformation, which is used for the operation of the self-attention mechanism; Representation layer normalization; represents a feed-forward network consisting of two linear layers; The second cross-modal Transformer enhances view feature representation Point cloud features A semantic association is established between them. The specific implementation formula is as follows: in: Represents enhanced view features Point cloud features Concatenate the fused features; represents the final enhanced point cloud feature representation; and Represent the fusion features after primary and secondary normalization respectively; , , represents the linear layer; , , , , , Represents the result after linear transformation and is used for the operation of the self-attention mechanism.
5. The bidirectional cross-modal image-guided point cloud restoration method according to claim 4, characterized in that: Enhanced point cloud feature representation With the initial point cloud features Connect them to get the final enhanced feature representation , the final enhanced feature representation As the input of the coarse point cloud generation module and the point cloud refinement module.
6. The bidirectional cross-modal image-guided point cloud restoration method according to claim 5, characterized in that: In step 3, the coarse point cloud generation module uses transposed convolution to extract the final enhanced feature representation Decode the features of each point.
7. The bidirectional cross-modal image-guided point cloud restoration method according to claim 6, characterized in that: The coarse point cloud generation module adopts a dual-branch approach. One branch uses the NodeShuffle module to generate new points from the latent space by combining the spatial information provided by adjacent points. The other branch uses the self-attention mechanism to simulate the long-distance dependencies between single-point features and capture the correlation between points to generate new points.
8. The bidirectional cross-modal image-guided point cloud restoration method according to claim 7, characterized in that: The new points of the two branches are combined by element-by-element addition to obtain a newly generated single point. The newly generated single point is combined with the final enhanced feature representation After connection, input into the multi-layer perceptron to generate a transitional coarse point cloud , the transition coarse point cloud With incomplete point cloud Connect and use the farthest point sampling algorithm to get a coarse point cloud .
9. The bidirectional cross-modal image-guided point cloud restoration method according to claim 8, characterized in that: The point cloud refinement module includes a global shape prediction submodule, a local structure optimization submodule, a DFGF submodule, and an upsampling submodule; the global shape prediction submodule is used to represent the and coarse point cloud Generate a feature representation that represents the global shape ; The local structure optimization submodule is used to optimize the incomplete point cloud based on edge convolution and farthest point sampling Iteratively aggregate and refine the local features to obtain a feature representation suitable for enhancing the local structure ; The DFGF submodule is used to represent and Get local structural features and global shape features ; The upsampling submodule is used to rely on local structural features and global shape features Fusion feature representation , generate coordinate offsets through MLP and reshaping operations, and use coordinate offsets to correct the coarse point cloud Make fine adjustments to finally get a complete point cloud .
Citation Information
Patent Citations
Three-dimensional reconstruction method and device for urban aerial image, electronic equipment and medium
CN116563465A
Dataset generation method for self-supervised learning scene point cloud completion based on panoramas
US20230094308A1
Deep learning-based high-precision point cloud completion method and apparatus
WO2024060395A1
Cited By
Residential building three-dimensional diagram enhancement processing method
CN120259102A
A method for enhancing a three-dimensional map of a residential building
CN120259102B
Panoramic depth estimation method integrating geometric and semantic optimization and related device
CN121095311A
A panoramic depth estimation method integrated with geometry and semantic optimization and related devices
CN121095311B
Three-dimensional point cloud repairing method and device guided by single view
CN121788397A