Controllable point cloud completion method and system based on large model prior

By introducing deep image redrawing and multi-view generation of text and image regulation, combined with the feedforward three-dimensional Gaussian sputtering generation model, the problem of patching missing parts in point cloud data is solved, and point cloud completion is achieved in the absence of truth value, which is highly robust and universal.

CN120236016AActive Publication Date: 2025-07-01WUHAN UNIV

Patent Information

Application Number
CN202510704342.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively repair the missing parts of the point cloud data acquired by lidar, especially in the absence of true value data. Traditional methods are limited by training data and deep learning methods are difficult to achieve results in practical applications.

Method used

Introduce deep image redrawing and multi-view generation with text and image regulation, combined with the feedforward three-dimensional Gaussian sputtering generation model, feature-level control and fusion are performed through three-plane expression to achieve point cloud completion based on large-models.

Benefits of technology

It achieves robust completion of point clouds in the absence of truth value, and is suitable for shape completion of a variety of three-dimensional point cloud assets, with strong general applicability and controllability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236016A_ABST
    Figure CN120236016A_ABST
Patent Text Reader

Abstract

The invention discloses a controllable point cloud completion method and system based on large model prior. The method comprises the steps that incomplete point cloud orthographic depth projection is carried out; redrawing point cloud depth projection based on an image generation model; generating three-dimensional representation guidance based on a multi-view diffusion model and a feedforward Gaussian sputtering generation model; and feature level fusion control and result sampling based on three-plane expression. According to the method, text and image regulation and control are introduced on the basis of a point cloud completion framework, and estimation of an unknown part is realized through text-controlled depth image redrawing and multi-view generation; the result of the feedforward three-dimensional Gaussian sputtering generation model and the incomplete point cloud are unified to three-plane expression for feature level control and fusion, and large model reasoning result-based point cloud completion is realized. According to the method, a point cloud completion framework with a result capable of being manually adjusted, high robustness and high universality is realized, and the method can be applied to an application scene which has a point cloud completion requirement and lacks a truth value for targeted training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of computer vision and laser scanning data processing, and particularly relates to point cloud data processing, point cloud quality enhancement, reconstruction content, and an automated technology for enhancing laser point cloud measurement data. Background Art

[0002] Laser point cloud is a kind of data expressing the surface position information of an object obtained by lidar, which consists of a series of three - dimensional coordinates. As an easily - obtained 3D discrete representation, laser point cloud can be used to model urban objects and provide an accurate description of the physical world, and is one of the main data sources for building a high - precision digital base. Point cloud has also received extensive attention in many fields of computer vision, such as 3D reconstruction, scene understanding, and autonomous driving. However, due to the inherent limitations of the scanning principle, the original point cloud obtained by lidar inevitably has missing parts caused by occlusion, reflection, etc. These missing parts will seriously affect the accurate progress of downstream tasks such as segmentation and reconstruction. Therefore, repairing the missing parts is a key step in geometric quality enhancement.

[0003] Traditional point cloud completion methods can only repair the missing parts of specific surface shapes based on artificial priors. Deep - learning - based point cloud completion methods are limited by training data, and are over - designed for uniformly - distributed synthetic data, making it difficult to handle situations outside the dataset distribution, uneven point cloud distribution, and the presence of noise. In actual use, due to the lack of ground - truth data for training, deep - learning - based methods are difficult to achieve the same effect as on synthetic data. The above problems limit the practical application of various current point cloud completion methods.

[0004] Based on the requirements for the generation and reconstruction of various 3D assets, in recent years, 3D generation models based on image generation models or large 3D datasets have emerged continuously, providing important prior knowledge for point clouds. However, the optimization algorithm based on the image generation model consumes too much computing power, which is unacceptable for the point cloud completion task that needs to process objects in batches. The utilization of the prior of the image generation model in the field of point cloud completion only stays at the result level, and the utilization of the prior of the feed - forward 3D model is still in the blue ocean of development. The feed - forward 3D generation model is superior to the optimization - based method in terms of generation speed, and its generation ability for the target level is gradually increasing. The development and utilization of its prior knowledge are of great significance for the point cloud completion task that needs to infer unknown parts. Summary of the Invention

[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a controllable point cloud completion method and system based on the prior knowledge of large models. By introducing text and image regulation on the basis of the point cloud completion framework, the speculation of unknown parts is realized through text-controlled depth image redrawing and multi-view generation; by unifying the results of the feed-forward three-dimensional Gaussian sputtering generation model and the defective point cloud into a tri-planar representation for feature-level control and fusion, point cloud completion based on the inference results of large models is realized.

[0006] According to one aspect of the specification of the present invention, a controllable point cloud completion method based on the prior knowledge of large models is provided, including: Performing orthographic depth projection on the obtained target-level defective point cloud to obtain an orthographic projection depth map; Performing large model inference based on the orthographic projection depth map to obtain end features; Inputting the end features and the target-level defective point cloud into a trained point cloud completion model, and outputting the completed complete point cloud; wherein, the training of the point cloud completion model includes: Constructing a target-level defective point cloud dataset; Performing orthographic depth projection and large model inference on the defective point clouds in the target-level cascaded defective point cloud dataset in sequence to obtain end features corresponding to the defective point clouds; Inputting the end features and the defective point cloud into the point cloud completion model, and the point cloud completion model encodes the end features and the defective point cloud into a tri-planar representation of the complete point cloud, samples the complete point cloud based on the tri-planar representation of the complete point cloud, and obtains the complete point cloud; Constructing a loss function to supervise the training process of the point cloud completion model, and outputting the trained point cloud completion model.

[0007] As a further technical solution, performing large model inference based on the orthographic projection depth map includes: Based on the orthographic projection depth map, using a text-conditioned image editing large model to generate a false color image; Based on the false color image, using a multi-view diffusion model to generate a panoramic image; Based on the panoramic image, using a feed-forward three-dimensional generation model to generate a 3D Gaussian result.

[0008] As a further technical solution, after generating the 3D Gaussian result, it further includes: Using the end features output by the feed-forward three-dimensional generation model as the prior knowledge of the large model, and screening the 3D Gaussian point coordinates output by the feed-forward three-dimensional generation model according to the occupancy rate as the point cloud representation of the inference result of the large model.

[0009] As a further technical solution, after inputting the end features and the defective point cloud into the point cloud completion model, it further includes: Deform the end feature into the form of a multi-channel image and input it into a neural network based on the Transformer architecture for feature extraction to obtain a three-plane representation of the end feature; Input the defective point cloud into a neural network based on the PointNet and residual connection structure to extract point-by-point features and project them onto three coordinate planes. Use a neural network based on U-Net to extract features from the coordinate planes projected with point-by-point features to obtain a three-plane representation of the defective point cloud; Input the three-plane representation of the end feature and the three-plane representation of the defective point cloud into a feature fusion network based on cross-attention to obtain a three-plane representation of the complete point cloud after fusion.

[0010] As a further technical solution, after obtaining the three-plane representation of the complete point cloud after fusion, it further includes: Use a neural network based on a multi-layer perceptron to sample the complete point cloud from the three-plane representation.

[0011] As a further technical solution, construct a loss function, including: Use the one-way chamfer distance loss between the defective point cloud and the output complete point cloud, use the chamfer distance loss between the Gaussian point cloud generated by the feed-forward three-dimensional generation model and the output complete point cloud, and calculate the density penalty loss of the output point cloud itself. Weight-sum the three losses to construct the loss function.

[0012] According to one aspect of the specification of the present invention, a controllable point cloud completion system based on a large model prior is provided for implementing the controllable point cloud completion method based on a large model prior.

[0013] As a further technical solution, the system includes: A first main module for performing orthographic depth projection on the acquired target-level defective point cloud to obtain an orthographic projection depth map; A second main module for performing large model inference based on the orthographic projection depth map to obtain an end feature; A third main module for inputting the end feature and the target-level defective point cloud into a trained point cloud completion model to output the completed complete point cloud; wherein, the training of the point cloud completion model includes: Construct a target-level defective point cloud dataset; Perform orthographic depth projection and large model inference on the defective point clouds in the target-level cascaded defective point cloud dataset in sequence to obtain end features corresponding to the defective point clouds; Input the end feature and the defective point cloud into the point cloud completion model. The point cloud completion model encodes the three-plane representation of the end feature and the defective point cloud to obtain a three-plane representation of the complete point cloud, and samples the complete point cloud based on the three-plane representation of the complete point cloud to obtain the complete point cloud; Construct a loss function to supervise the training process of the point cloud completion model, and output the trained point cloud completion model.

[0014] According to one aspect of the specification of the present invention, there is provided a controllable point cloud completion system based on large model priors, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the described controllable point cloud completion method based on large model priors.

[0015] According to one aspect of the specification of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the described controllable point cloud completion method based on large model priors is implemented.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By introducing text and image regulation on the basis of the point cloud completion framework, the present invention realizes the speculation of unknown parts through text-controlled depth image redrawing and multi-view generation; by unifying the results of the feed-forward three-dimensional Gaussian sputtering generation model and the defective point cloud into a tri-planar representation for feature-level control and fusion, point cloud completion based on the inference results of the large model is realized.

[0017] 2. The present invention can combine large model priors and the original geometric structure of defective point clouds to complete the point cloud of defective point clouds, and has strong robustness to the target object type, point cloud density, and uneven point cloud distribution. Good results can also be obtained in the case of lack of ground truth and type characteristics, and it can be applied to the enhancement of original point cloud data and the shape completion of various three-dimensional point cloud assets. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0019] Figure 1 is the overall technical process provided by the embodiment of the present invention; Figure 2 is a schematic diagram of a defective point cloud provided by the embodiment of the present invention; Figure 3 In (a) is a schematic diagram of the method for obtaining dense ortho-depth provided by the embodiment of the present invention, Figure 3 and (b) in is a schematic diagram of the large model inference pipeline provided by the embodiment of the present invention; Figure 4It is the point cloud completion result diagram provided by the embodiments of the present invention. Detailed implementation manners

[0020] For the target-level defective point cloud, the present invention proposes a controllable point cloud completion method based on large model prior, including projecting the defective point cloud into a dense orthographic depth image; generating a false color image, a panoramic image, and a 3D Gaussian representation according to the depth image; encoding the defective point cloud and the 3D Gaussian representation obtained by large model inference into a triplane space and fusing them; and sampling a complete point cloud from the fused triplane representation.

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or individual embodiment provided by the present invention can be combined with each other arbitrarily to form a new technical solution. Such combination is not restricted by the order of steps and / or the mode of structural composition, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions conflicts with each other or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0022] See Figure 1 , for the controllable point cloud completion method based on large model prior proposed by the embodiments of the present invention, first project the defective point cloud into a dense orthographic depth image, then use an image generation large model conditioned on the depth image and text to generate a false color image from the dense orthographic depth image, use a multi-view diffusion model to generate four panoramic images from a single false color image, and use a feed-forward three-dimensional generation large model to generate a three-dimensional Gaussian representation of the complete shape obtained by large model prior inference from the four panoramic images. The present invention uses the end feature of the large model inference process as prior guidance, encodes the end feature and the defective point cloud into the triplane representation space, and performs fusion on the triplane representation. Finally, the present invention samples a complete point cloud from the fused triplane representation containing the three-dimensional representation of the complete shape. The specific implementation process includes the following steps: Step 1, obtain its depth image through orthographic depth projection based on the target-level defective point cloud data.

[0023] The embodiments of the present invention further propose the implementation manner of Step 1 as follows: First, rotate the object target-level defective point cloud to the axis-aligned direction and normalize it into the unit sphere space so that its geometric center is located at the origin of the three-dimensional coordinate system. Then, perform an orthographic depth projection from its front side to obtain the point cloud depth image. Use the fill-in-fast algorithm based on dilated convolution to fill the sparse point cloud depth image into a dense point cloud depth image.

[0024] The specific implementation of step 1 in the embodiment is as follows: 1.1) In the embodiment, rotate to axis-aligned and normalize: Let the defective point cloud be , representing the coordinates of each point in the defective point cloud, N representing the number of points, representing the set of points . First, construct a three-dimensional rectangular coordinate system with the positive direction of the z-axis vertically upward. Calculate its geometric center point coordinates as shown in Equation (1) :

[0025]

[0026] Move the defective point cloud to its geometric center at the origin as shown in Equation (2):

[0027] Calculate the distance from the point farthest from the origin in the defective point cloud to the origin as shown in Equation (3), and normalize the defective point cloud to the unit sphere space based on this:

[0028] Manually rotate the defective point cloud to make it in the axis-aligned direction, and the projection obtained from the y-axis direction is relatively complete.

[0029] 1.2) In the embodiment, orthographic depth projection: Perform orthographic depth projection from the y-axis direction alignment. The mapping relationship between the point cloud coordinates and the depth image pixel coordinates and grayscale is shown in Equations (4), (5), and (6). Among them, H and W are the height and width of the depth image, and the depth image is composed of three channels of RGB, and the color is assigned as eight-bit integer. In the embodiment, the same value is assigned to all channels to implement a grayscale image.

[0030]

[0031] Among them, the subscript represents the pixel coordinates projected onto the image, the subscript part represents the xyz three-dimensional coordinates of the point cloud, and R, G, and B represent the three color channels of the image. When projecting, it is necessary to judge whether the current point is blocked by other points according to the depth size.

[0032] 1.3) In the embodiment, orthographic depth projection void filling: After projecting into a depth image, a void filling algorithm is used to fill the voids in the depth image. The void filling algorithm is the fill-in-fast algorithm, which is based on dilated convolution.

[0033] The process and intermediate result examples of step 1 are as Figure 3 shown in (a) of

[0034] Step 2, use a text-conditioned image editing large model (such as ControlNet) to draw the depth image into a false color image. The dense orthographic depth projection obtained in step 1 and the text prompt are used as the inputs of the text-conditioned image editing large model based on the depth map, and a false color image is obtained, and then its background is removed.

[0035] The specific implementation of step 2 in the embodiment is as follows: The orthographic depth image obtained in step 1 is used as the conditional input of the image editing large model. In the embodiment, the image editing large model is ControlNet based on the depth map. The control text is an artificial conditional input, which should contain a detailed description of the target point cloud, including descriptions of type, name, features, etc.

[0036] Step 3, based on the false color image, use a multi-view diffusion model (MVDream) to generate four panoramic images from a single false color image.

[0037] The specific implementation of step 3 in the embodiment is as follows: The false color image obtained in step 3 is used as the input of the multi-view image generation large model. In the embodiment, the multi-view image generation large model is MVDream. The control text is an artificial conditional input, which should contain a detailed description of the target point cloud, including descriptions of type, name, features, etc.

[0038] Step 4, use a feed-forward 3D generation model (LGM) to generate corresponding 3D Gaussian results from the four panoramic images. The four panoramic images obtained in step 3 are used as the inputs of the feed-forward 3D generation model (LGM), and the output 3D generation results (Gaussian sputtering) are obtained. The position attributes of the Gaussian sputtering are used as 3D coordinates, and they are screened according to the occupancy rate to obtain Gaussian point clouds for subsequent supervision. The output (terminal feature) of the last layer of the main neural network module (U-Net) in the 3D generation model is used as the prior of the large model inference for subsequent unified 3D representation and feature-level fusion control.

[0039] The specific implementation of step 4 in the embodiment is as follows: Use the four omnidirectional images obtained in step 4 as the input of the feed-forward three-dimensional Gaussian generation large model. In the embodiment, the feed-forward three-dimensional Gaussian generation large model is LGM. In LGM, the four omnidirectional images After passing through the U-Net, the end features are obtained . After the end features pass through the projection layer and the corresponding Gaussian attribute activation function, they are deformed into 3D Gaussian expressions . By setting a threshold for the Gaussian occupancy attribute opacity for screening, and retaining its position attribute as the point cloud coordinates, a Gaussian point cloud is obtained .

[0040] Steps 2, 3, and 4 are the processes of the large model inference pipeline provided by the embodiment of the present invention, that is, the process of the second main module. The process schematic diagram and intermediate result examples are as Figure 3 shown in (b) of

[0041] In summary, in steps 1 to 4 of the embodiment of the present invention, first, the object target-level defective point cloud is rotated to the axis-aligned direction and normalized to the unit sphere space so that its geometric center is located at the origin of the three-dimensional coordinate system. Then, an orthographic depth projection is performed from its relatively complete observation surface to obtain a point cloud depth image. The sparse point cloud depth image is filled into a dense point cloud depth image using the fill-in-fast algorithm. The dense orthographic depth projection and the text prompt are used as the input of the image editing large model (ControlNet) based on the depth map condition to obtain a false color image. Then, its background is removed. The background-free false color image is used as the input of the multi-view generation model (MVDream) to obtain four multi-view consistent omnidirectional images. The four omnidirectional images are used as the input of the feed-forward three-dimensional generation model (LGM) to obtain the output three-dimensional generation result (Gaussian sputtering). The position attribute of the Gaussian sputtering is used as the three-dimensional coordinate, and it is screened according to the occupancy rate to obtain a Gaussian point cloud for subsequent supervision. The output (end features) of the last layer of the main neural network module (U-Net) in the three-dimensional generation model is used as the large model inference prior for subsequent unified three-dimensional representation and feature-level fusion control.

[0042] Step 5, use the Transformer-based triplane encoder to encode the large model prior into a triplane expression. Deform the end features obtained in step 4 into the form of a multi-channel image, input it into the neural network based on the Transformer architecture for feature extraction, and then deform it into the format of triplane features through the upsampling layer.

[0043] The specific implementation of step 5 in the embodiment is as follows: The specific implementation of the end feature three - plane encoder can refer to the image three - plane encoder in LRM. This encoder is based on the Transformer architecture and converts the initial three - plane embedding into a three - plane representation containing image geometric information conditional on the input image. In the embodiment, the end feature After deformation and one - layer convolution, it obtains the features input to the Transformer . The output features of the Transformer obtain a Gaussian three - plane representation after deformation and a de - convolution layer .

[0044] Step 6: Use a point - cloud three - plane encoder based on PointNet and projection relationships to encode the defective point cloud into a three - plane representation. First, input the defective point cloud into a neural network based on PointNet and residual connection structure to extract point - by - point features. Then project the defective point cloud onto three coordinate planes according to the coordinates, and assign the point - by - point features of the corresponding points to the corresponding pixels on the coordinate planes. Integrate the features of the three planes to obtain the projection of the point cloud in the three - plane space. Finally, use a neural network based on U - Net to extract features from this three - plane representation to obtain the final point - cloud three - plane representation.

[0045] The specific implementation of Step 6 in the embodiment is as follows: The architecture of the point - cloud three - plane encoder can refer to the point - cloud three - plane encoder in GenSDF during specific implementation. This encoder extracts the point - by - point features of the input point cloud and projects them onto three coordinate planes, and then extracts features from the coordinate planes with projected features to obtain a three - plane representation containing the geometric information of the point cloud. The point - cloud three - plane encoder encodes the defective point cloud into a point - cloud three - plane representation .

[0046] Step 7: Use a neural network based on cross - attention to fuse the three - plane representation of the large - model prior with the three - plane representation of the defective point cloud to obtain a three - plane representation of the final complete shape. Integrate the three - plane representation of the end feature obtained in Step 5 and the three - plane representation of the point cloud in Step 6 to the same size through a projection layer, and then input them into a feature fusion network based on cross - attention to obtain the fused three - plane representation.

[0047] The specific implementation of Step 7 in the embodiment is as follows: First, the three - plane representation of the end feature and the three - plane representation of the point cloud are transformed into , and then these two features are fused through stacked cross-attention and self-attention layers. Among them, the fusion method of the cross-attention layer is shown in Equation (7), the calculation process of the self-attention layer is shown in Equation (8), and the overall fusion process is shown in Equations (9) and (10). Among them, MHA is the abbreviation of MultiHeadAttention, and the three inputs of this network layer are query, key, and value in sequence.

[0048]

[0049] The fused features After deformation and deconvolution layers, it is transformed into a fused three-plane representation .

[0050] Step 8, use a neural network based on a multi-layer perceptron to sample a complete point cloud from the three-plane representation. First, randomly initialize an initial point cloud, query the corresponding three-plane features of each point in the initial point cloud one by one in the three-plane representation obtained in Step 7, input the queried features into the neural network based on a multi-layer perceptron, calculate the coordinate offset, and add the coordinate offset to the coordinates of the query point to obtain the coordinates of the points on the final complete point cloud. After looping through all the points in the initial point cloud, the complete point cloud can be obtained.

[0051] The specific implementation of Step 8 in the embodiment is as follows: First, randomly initialize the point cloud , for each point in, use its three-dimensional coordinates to query the corresponding features on the three corresponding coordinate planes in to obtain the features of three corresponding pixels. Combine the queried features to obtain the point-by-point features of N points , and input into the point cloud three-plane feature decoder based on a multi-layer perceptron to decode the corresponding point-by-point offset . Add the point-by-point offset to the initial random point cloud to obtain the point-by-point coordinates of the complete point cloud, that is, Equation (11).

[0052]

[0053] Step 9, training phase. First, set the number of points used for sampling the complete point cloud in Step 8 to be the same as the number of points in the input defective point cloud. Calculate the mean of the point-to-point chamfer distances from the defective point cloud to the complete point cloud as Loss 1. Then, downsample the Gaussian point cloud obtained in Step 4 to the same number of points as the complete point cloud, and calculate the mean of the point-to-point chamfer distances from the Gaussian point cloud to the complete point cloud and from the complete point cloud to the Gaussian point cloud as Loss 2. Also, calculate the density penalty loss of the complete point cloud as Loss 3. The total loss is the weighted sum of Loss 1, Loss 2, and Loss 3, which is used to supervise the training of the neural network involved in Steps 5 to 8.

[0054] The specific implementation of Step 9 in the embodiment is as follows: Unidirectional chamfer distance loss is calculated as shown in Equation (12), and the chamfer distance loss is calculated as shown in Equation (13). Here, X and Y are two point clouds, and N(X) represents the number of points in point cloud X. The density penalty loss is calculated as shown in Equation (14), where STD is the standard deviation function. When calculating the density penalty loss, first calculate the point-to-point distance , select the largest 16 values among them, and calculate their standard deviation as the loss. The overall loss of the present invention is calculated as shown in Equation (15).

[0055]

[0056] In the embodiment, the controllable point cloud completion framework based on the large model prior constructed by the present invention was trained for 160 rounds on eight known categories of the ShapeNet-ViPC dataset without relying on the ground truth.

[0057] The point cloud completion results obtained through the above steps can be applied to the augmentation of the original point cloud data and the shape completion of various 3D point cloud assets. Schematic diagrams before and after point cloud completion are as shown in Figure 2 and Figure 4 shown.

[0058] In summary, in the implementation of steps 5 to 9 in the embodiments of the present invention, first, the end features obtained in step 4 are transformed into the form of a multi-channel image and input into a neural network based on the Transformer architecture for feature extraction. Then, it is transformed into the format of three-plane features through an upsampling layer. The defective point cloud is input into a neural network based on the PointNet and residual connection structure to extract point-by-point features. Then, the defective point cloud is projected onto three coordinate planes according to the coordinates respectively, and the point-by-point features of the corresponding points are assigned to the corresponding pixels on the coordinate planes. The features of the three planes are integrated to obtain the projection of the point cloud in the three-plane space. Finally, a neural network based on U-Net is used to extract features from the three-plane representation to obtain the final three-plane representation of the point cloud. The three-plane representations of the end features and the point cloud are integrated to the same size through a projection layer. Then, it is input into a feature fusion network based on cross-attention to obtain the fused three-plane representation. An initial point cloud is randomly initialized, and the points in the initial point cloud are queried one by one for the corresponding three-plane features in the three-plane representation obtained in step 7. The queried features are input into a neural network based on a multi-layer perceptron to calculate the coordinate offset, and the coordinate offset is added to the coordinates of the query point to obtain the coordinates of the points on the final complete point cloud. After querying all the points in the initial point cloud in a loop, the complete point cloud is obtained. The average of the point-by-point chamfer distances from the defective point cloud to the complete point cloud is calculated as loss 1. Then, the Gaussian point cloud obtained in step 4 is downsampled to the same number of points as the complete point cloud, and the average of the point-by-point chamfer distances from the scattered Gaussian point cloud to the complete point cloud and from the complete point cloud to the Gaussian point cloud is calculated as loss 2. In addition, the density penalty loss of the complete point cloud is calculated as loss 3. The total loss is the weighted sum of loss 1, loss 2, and loss 3, which is used to supervise the training of the neural networks involved in steps 5 to 8.

[0059] In specific implementation, the method proposed by the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. The system device for implementing the method, such as a computer-readable storage medium storing the corresponding computer program of the technical solution of the present invention and a computer device including running the corresponding computer program, should also be within the protection scope of the present invention.

[0060] In some possible embodiments, a controllable point cloud completion system based on the prior of a large model is provided, including the following modules: The first main module, the orthographic depth projection module, is used to perform orthographic depth projection on the obtained target-level defective point cloud to obtain an orthographic projection depth map; this module includes point cloud axis alignment and normalization, orthographic depth projection, and void filling, and is used to project the defective point cloud into a dense orthographic depth image.

[0061] The second main module, the large model inference pipeline module, is used to perform large model inference based on the orthographic projection depth map to obtain end features; this module includes an image editing large model, a multi-view generation large model, and a 3D generation large model, which are used to obtain a reasonable complete 3D shape from the prior knowledge of the large model.

[0062] The third main module, the tri-planar encoding network, is used to input the end features and the target-level residual point cloud into the trained point cloud completion model and output the completed complete point cloud; this module includes a point cloud tri-planar encoder, a Gaussian end feature tri-planar encoder, a tri-planar feature fusion network, and a point cloud sampling network, which are used to unify the inference results of the large model and the residual point cloud into the same 3D expression, realize feature-level fusion control, and obtain a point cloud expression that conforms to both the original shape and the large model's speculation from the fusion results.

[0063] In some possible embodiments, a controllable point cloud completion system based on large model prior is provided, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a controllable point cloud completion method based on large model prior as described above.

[0064] In some possible embodiments, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium, and when the computer program is executed, the controllable point cloud completion method based on large model prior as described above is implemented.

[0065] In summary of the above embodiments, the present invention provides a controllable point cloud completion method based on a generative large model, including orthographic depth projection of the residual point cloud; point cloud depth projection redrawing based on an image generation model; 3D representation guidance based on a multi-view diffusion model and a feed-forward Gaussian sputtering generation model; feature-level fusion control and result sampling based on tri-planar expression. The present invention introduces text and image regulation on the basis of the point cloud completion framework, and realizes the speculation of unknown parts through text-controlled depth image redrawing and multi-view generation; by unifying the results of the feed-forward 3D Gaussian sputtering generation model and the residual point cloud into the tri-planar expression for feature-level control and fusion, the point cloud completion based on the inference results of the large model is realized. In the model training stage, the present invention can be trained without using the ground truth, relying only on the residual point cloud and the inference results of the large model. The present invention realizes a point cloud completion framework with adjustable results, strong robustness, and strong versatility by combining the speculation of the large model on the incomplete part of the point cloud and the shape of the original residual point cloud at the implicit expression level, and can be applied to application scenarios where there is a need for point cloud completion and there is a lack of ground truth for targeted training, such as 3D reconstruction of various infrastructures and assets.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A controllable point cloud completion method based on the prior of large models, characterized in that, Including: Performing orthographic depth projection on the obtained target-level defective point cloud to obtain an orthographic projection depth map; Performing large model inference based on the orthographic projection depth map to obtain end features; Inputting the end features and the target-level defective point cloud into a trained point cloud completion model to output a completed full point cloud; wherein, the training of the point cloud completion model includes: Constructing a target-level defective point cloud dataset; Performing orthographic depth projection and large model inference on the defective point clouds in the target-level cascaded defective point cloud dataset in sequence to obtain end features corresponding to the defective point clouds; Inputting the end features and the defective point cloud into a point cloud completion model, and the point cloud completion model performs triple-plane expression encoding on the end features and the defective point cloud to obtain a full point cloud triple-plane expression, and performs full point cloud sampling based on the full point cloud triple-plane expression to obtain a full point cloud; Constructing a loss function to supervise the training process of the point cloud completion model and outputting a trained point cloud completion model.

2. The controllable point cloud completion method based on the prior of the large model according to claim 1, wherein, Performing large model inference based on the orthographic projection depth map includes: Based on the orthographic projection depth map, using a text-conditioned image editing large model to generate a false color image; Based on the false color image, using a multi-view diffusion model to generate a panoramic image; Based on the panoramic image, using a feed-forward three-dimensional generation model to generate a 3D Gaussian result.

3. The controllable point cloud completion method based on the prior of the large model according to claim 2, wherein, After generating the 3D Gaussian result, it further includes: Using the end features output by the feed-forward three-dimensional generation model as the large model prior, and using the 3D Gaussian point coordinates output by the feed-forward three-dimensional generation model screened by occupancy as the point cloud representation of the inference result of the large model.

4. The controllable point cloud completion method based on the prior of the large model according to claim 1, wherein, After inputting the end features and the defective point cloud into the point cloud completion model, it further includes: Deforming the end features into the form of a multi-channel image and inputting them into a neural network based on the Transformer architecture for feature extraction to obtain the triple-plane expression of the end features; Inputting the defective point cloud into a neural network based on the PointNet and residual connection structure to extract point-by-point features and project them onto three coordinate planes, and using a neural network based on the U-Net to perform feature extraction on the coordinate planes projected with point-by-point features to obtain the triple-plane expression of the defective point cloud; Inputting the triple-plane expression of the end features and the triple-plane expression of the defective point cloud into a feature fusion network based on cross-attention to obtain the triple-plane expression of the fused full point cloud.

5. The controllable point cloud completion method based on large model prior according to claim 4, characterized in that After obtaining the triple-plane expression of the fused full point cloud, it further includes: Using a neural network based on a multi-layer perceptron to sample a full point cloud from the triple-plane expression.

6. The controllable point cloud completion method based on large model prior according to claim 3, wherein, Constructing a loss function includes: Taking the one-way chamfer distance loss between the defective point cloud and the output full point cloud, taking the chamfer distance loss between the Gaussian point cloud generated by the feed-forward three-dimensional generation model and the output full point cloud, and calculating the density penalty loss of the output point cloud itself, and weighted summing the three losses to construct a loss function.

7. A controllable point cloud completion system based on the prior of large models, characterized in that, For implementing a controllable point cloud completion method according to any one of claims 1-6.

8. The controllable point cloud completion system based on the prior of the large model according to claim 7, characterized in that, The system includes: A first main module for performing orthographic depth projection on the obtained target-level defective point cloud to obtain an orthographic projection depth map; The second main module is used to perform large model inference based on the orthographic projection depth map to obtain the end - point features; The third main module is used to input the end - point features and the target - level defective point cloud into the trained point cloud completion model and output the completed full point cloud; wherein, the training of the point cloud completion model includes: Construct a target - level defective point cloud dataset; Perform orthographic depth projection and large model inference on the defective point clouds in the target - level cascaded defective point cloud dataset in sequence to obtain the end - point features corresponding to the defective point clouds; Input the end - point features and the defective point cloud into the point cloud completion model, and the point cloud completion model performs tri - plane expression encoding on the end - point features and the defective point cloud to obtain the tri - plane expression of the full point cloud, and perform full point cloud sampling based on the tri - plane expression of the full point cloud to obtain the full point cloud; Construct a loss function to supervise the training process of the point cloud completion model and output the trained point cloud completion model.

9. A controllable point cloud completion system based on the prior of a large model, characterized in that, It includes a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a controllable point cloud completion method based on large model prior as described in any one of claims 1 - 6.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer - readable storage medium, and when the computer program is executed, it implements a controllable point cloud completion method based on large model prior as described in any one of claims 1 - 6.

Citation Information

Patent Citations

  • Incomplete point cloud completion method based on hidden space topological structure constraint

    CN113205466A

  • Point cloud dense completion method based on deep learning

    CN113628140A

  • Point cloud completion method based on attention mechanism

    CN115131245A

  • Self-supervised three-dimensional point cloud completion method for underwater target object

    CN118470515A

  • Sparse visual angle three-dimensional reconstruction method based on depth prior information

    CN118657888A

Cited By

  • Point cloud completion method and device based on text prompt, equipment and storage medium

    CN121860895A