Text prompt-based point cloud completion method and device, equipment and storage medium

By using a text-based point cloud completion method, combining a diffusion model and a completion model with a large language model and a 3D point cloud encoder, a complete point cloud of the target object is generated. This solves the problem of poor point cloud completion effect in existing technologies and improves completion efficiency and reliability.

CN121860895BActive Publication Date: 2026-05-29湖南工商大学

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
湖南工商大学
Filing Date
2026-03-17
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing point cloud completion methods struggle to handle complex structures and large-scale missing scenes, leading to geometric distortions, loss of details, and discontinuous contours in 3D reconstruction models, severely reducing the accuracy of target object recognition.

Method used

A text-based point cloud completion method is adopted. The complete point cloud of the target object is generated by using a diffusion model and a completion model. The large language model and 3D point cloud encoder are combined with a cross-attention mechanism to generate point cloud features through cross-modal feature fusion, and the complete point cloud of the target object is automatically generated.

Benefits of technology

It improves the efficiency and reliability of point cloud completion for target objects, reduces manual intervention, and generates point clouds that better match the actual physical spatial distribution, thereby improving the accuracy of 3D reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860895B_ABST
    Figure CN121860895B_ABST
Patent Text Reader

Abstract

The application relates to the visual technology field and the artificial intelligence technology field, and discloses a point cloud completion method and device based on a text prompt, equipment and a storage medium. The method comprises the following steps: generating a diffusion loss of a diffusion model according to a first loss model, and generating a completion loss of a completion model according to a second loss model; generating a total loss of an overall model according to the diffusion loss of the diffusion model, the completion loss of the completion model and a total loss model, updating model parameters of the overall model, and saving the updated overall model; determining text semantic features of a preset object based on text prompt information corresponding to a target object, determining fusion features of the target object based on a plurality of non-overlapping local regions in a current incomplete point cloud and the text semantic features of the preset object, and completing the current incomplete point cloud based on the fusion features of the target object by using the updated overall model, so as to generate a complete point cloud of the target object. The application can improve the completion efficiency of the complete point cloud of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of visual technology and artificial intelligence technology, and in particular to point cloud completion methods, apparatus, devices and storage media based on text prompts. Background Technology

[0002] In practical applications of 3D vision and 3D reconstruction, 3D scanning equipment is an important means of acquiring geometric information of object surfaces. During actual acquisition, scanning results are easily affected by factors such as object occlusion, single-view or limited-view acquisition limitations, complex environmental lighting and noise interference, as well as the accuracy of the equipment itself and the acquisition distance, resulting in incomplete structure of the point cloud data of the target object.

[0003] Incomplete point cloud data directly causes geometric distortion, loss of detail, and discontinuous contours in 3D reconstruction models, severely reducing the accuracy of subsequent target object recognition and failing to meet practical needs. Traditional point cloud completion methods mostly rely on simple interpolation or manual completion, which are difficult to handle complex structures and large-scale missing scenes, resulting in limited repair effects and weak generalization ability. Therefore, how to generate complete point clouds of target objects is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This application provides a point cloud completion method, apparatus, device, and storage medium based on text prompts to solve the aforementioned technical problem of how to generate a complete point cloud of a target object.

[0005] In a first aspect, embodiments of this application provide a point cloud completion method based on text prompts, applied to electronic devices, the point cloud completion method comprising:

[0006] Obtain a preset defect cloud, and determine the fusion features of the preset object based on the preset defect cloud;

[0007] The fusion features of the preset object are input into the diffusion model, and the diffusion model generates the predicted point cloud of each component within the preset object.

[0008] Based on the predicted point cloud of each component within the preset object, the real point cloud of each component within the preset object, and the first loss model, the diffusion loss of the diffusion model is generated, and the completion loss of the completion model is generated based on the second loss model.

[0009] Based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, the total loss of the overall model is generated, and the model parameters of the overall model are updated until the total loss of the overall model meets the preset conditions, at which point the updated overall model is saved.

[0010] The current residual point cloud and text prompt information corresponding to the target object are obtained. Based on the text prompt information corresponding to the target object, the text semantic features of the preset object are determined. Based on multiple non-overlapping local regions in the current residual point cloud, the text semantic features of the preset object and the predefined method, the fusion features of the target object are determined. The updated overall model is based on the fusion features of the target object to complete the current residual point cloud and generate the complete point cloud of the target object.

[0011] In one possible implementation of the first aspect, obtaining a preset defect cloud and determining the fusion features of a preset object based on the preset defect cloud includes:

[0012] Obtain a sample set, and obtain training samples from the sample set. The training samples include the preset defect cloud corresponding to the preset object, the text prompt information corresponding to the preset object, and the real point cloud of each component in the preset object.

[0013] Extract the nouns of each component within the preset object from the text prompts corresponding to the preset object, input the nouns of each component within the preset object into the large language model, and generate an extended description of the nouns of each component within the preset object through the large language model.

[0014] A text encoder is used to encode the textual prompts for the preset object and the extended descriptions of the nouns for each component within the preset object, generating the textual semantic features of the preset object. The total number of components within the preset object is calculated. A K-means clustering algorithm is used to divide the preset defect cloud based on the total number of components, generating multiple non-overlapping local regions within the preset defect cloud. These non-overlapping local regions are then input into a 3D point cloud encoder, which extracts features from them, generating local geometric features of the preset defect cloud. The entire region of the preset defect cloud is also input into the 3D point cloud encoder, which extracts features from it, generating global geometric features. A cross-attention mechanism is then used to fuse the local geometric features, the corresponding global geometric features, and the textual semantic features of the preset object across modalities, generating the fused features of the preset object.

[0015] In one possible implementation of the first aspect, the step of generating the total loss of the overall model based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, updating the model parameters of the overall model, and saving the updated overall model only after the total loss of the overall model meets a preset condition, includes:

[0016] Based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, the total loss of the overall model is generated. The model parameters of the overall model are updated using backpropagation until the total loss of the overall model is less than the preset loss value. Only then is the updating of the model parameters of the overall model stopped and the updated overall model is saved.

[0017] In one possible implementation of the first aspect, the steps of obtaining the current residual point cloud corresponding to the target object and the text prompt information corresponding to the target object, determining the text semantic features of a preset object based on the text prompt information, determining the fusion features of the target object based on multiple non-overlapping local regions within the current residual point cloud, the text semantic features of the preset object, and a predefined method, and then completing the current residual point cloud based on the fusion features of the target object using the updated overall model to generate a complete point cloud of the target object, include:

[0018] Obtain the current defect cloud and text prompt information corresponding to the target object. Extract the nouns of each component within the target object from the text prompt information. Input the nouns of each component within the target object into the large language model. Generate an extended description of the nouns of each component in the current defect cloud through the large language model.

[0019] The text encoder encodes the text prompts corresponding to the target object and the extended descriptions of the nouns of each component within the target object, generating the text semantic features of the target object. It also counts the number of components within the target object and uses the K-means clustering algorithm to divide the current defect cloud based on the total number of components, generating multiple non-overlapping local regions within the current defect cloud. These non-overlapping local regions are then input into a 3D point cloud encoder, which extracts features from them to generate the local geometric features of the current defect cloud. Finally, the entire region of the current defect cloud is input into the 3D point cloud encoder, which extracts features from it to generate the global geometric features of the current defect cloud.

[0020] A cross-attention mechanism is employed to fuse the local geometric features of the current incomplete point cloud, the corresponding global geometric features, and the textual semantic features of the target object across modalities, generating a fused feature of the target object. This fused feature is then input into an updated diffusion model, which generates a point cloud of the missing component. Affine transformation parameters are used to perform rotation, scaling, shearing, and translation operations on the point cloud of the missing component, generating an adjusted point cloud of the missing component. The adjusted point cloud of the component is then used to complete the current incomplete point cloud, generating a complete point cloud of the target object.

[0021] In one possible implementation of the first aspect, the first loss model is defined as follows:

[0022] ;

[0023] This represents the diffusion loss in the diffusion model;

[0024] This represents the KL divergence of the diffusion model;

[0025] This represents the total number of sub-parts after an object has been decomposed.

[0026] Indicates the sequence number of the sub-components within the preset object;

[0027] This represents the actual point cloud of the i-th sub-component within a predefined object;

[0028] This represents the predicted point cloud of the i-th sub-component within the preset object at the i-th step of the diffusion process;

[0029] This represents the predicted point cloud of the i-th sub-component within the preset object in the step preceding the i-th step of the diffusion process;

[0030] An extended description of the noun representing the i-th sub-component within a predefined object;

[0031] Represents the true distribution;

[0032] This represents the expectation under the true distribution;

[0033] These represent the model parameters of the diffusion model;

[0034] This represents the predicted probability distribution generated by the diffusion model based on the model parameters;

[0035] In one possible implementation of the first aspect, the second loss model is defined as follows:

[0036] ;

[0037] This represents the completion loss of the completion model;

[0038] This represents the predicted point cloud of a predefined object. Represents the true point cloud of a predefined object;

[0039] This indicates the total number of points for the preset object; This represents the total number of points in the actual point cloud of the preset object;

[0040] To represent any point in the predicted point cloud; Represents any point in a real point cloud;

[0041] express and The Euclidean distance between them;

[0042] express and The Euclidean distance between them.

[0043] In one possible implementation of the first aspect, the total loss model is defined as follows:

[0044] ;

[0045] This represents the total loss of the entire model;

[0046] This represents the diffusion loss in the diffusion model;

[0047] This represents the completion loss of the completion model; It is a parameter used to adjust the proportion of diffusion loss; It is a parameter used to adjust the proportion of compensation loss; .

[0048] Secondly, embodiments of this application provide a point cloud completion device based on text prompts, applied to electronic devices, including:

[0049] The acquisition module is used to acquire a preset defect cloud and determine the fusion features of a preset object based on the preset defect cloud;

[0050] The input module is used to input the fusion features of the preset object into the diffusion model, and the diffusion model generates the predicted point cloud of each component within the preset object.

[0051] The first generation module is used to generate the diffusion loss of the diffusion model based on the predicted point cloud of each component in the preset object, the real point cloud of each component in the preset object and the first loss model, and to generate the completion loss of the completion model based on the second loss model.

[0052] The update module is used to generate the total loss of the overall model based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, and update the model parameters of the overall model until the total loss of the overall model meets the preset conditions before saving the updated overall model.

[0053] The second generation module is used to obtain the current residual point cloud corresponding to the target object and the text prompt information corresponding to the target object. Based on the text prompt information corresponding to the target object, the text semantic features of the preset object are determined. Based on multiple non-overlapping local regions in the current residual point cloud, the text semantic features of the preset object and the predefined method, the fusion features of the target object are determined. The updated overall model is based on the fusion features of the target object to complete the current residual point cloud and generate the complete point cloud of the target object.

[0054] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the point cloud completion method described in the first aspect above.

[0055] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the point cloud completion method described in the first aspect above.

[0056] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the point cloud completion method described in the first aspect.

[0057] The beneficial effects of the embodiments of this application are as follows:

[0058] Firstly, the current residual point cloud corresponding to the target object and the text prompt information corresponding to the target object are obtained. Based on the text prompt information corresponding to the target object, the text semantic features of the preset object are determined. Based on multiple non-overlapping local regions in the current residual point cloud, the text semantic features of the preset object and the predefined method, the fusion features of the target object are determined. The updated overall model completes the current residual point cloud based on the fusion features of the target object to generate the complete point cloud of the target object. Since no manual completion is required, the completion time of the complete point cloud of the target object is reduced, which is conducive to improving the completion efficiency of the complete point cloud of the target object.

[0059] Secondly, since the complete point cloud of the target object is automatically generated without being affected by human intervention, it helps to improve the reliability of the complete point cloud of the target object. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is an application scenario diagram of the point cloud completion method provided in the embodiments of this application;

[0062] Figure 2 This is a flowchart illustrating the point cloud completion method provided in the embodiments of this application;

[0063] Figure 3 A flowchart illustrating the implementation of S204 provided in this application embodiment;

[0064] Figure 4 A schematic block diagram of the point cloud completion device provided in the embodiments of this application;

[0065] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0066] Figure 6 This is a flowchart of the point cloud completion method provided in the embodiments of this application. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0068] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0069] It should be understood that in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0070] Furthermore, the technical solutions of the various embodiments can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0071] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0072] The point cloud completion method provided in this application can be applied to electronic devices, including but not limited to servers, mobile phones, tablets, wearable devices, vehicle-mounted devices, and laptops. This application does not impose any restrictions on the specific type of electronic device.

[0073] Please see Figure 1 , Figure 1 The application scenario diagram of the point cloud completion method provided in the embodiments of this application is described in detail below:

[0074] The electronic device connects to the database, retrieves a sample set from the database, and retrieves training samples from the sample set. The training samples include the preset defect cloud corresponding to the preset object, the text prompt information corresponding to the preset object, and the real point cloud of each component in the preset object.

[0075] Extract the nouns of each component within the preset object from the text prompts corresponding to the preset object, input the nouns of each component within the preset object into the large language model, and generate an extended description of the nouns of each component within the preset object through the large language model.

[0076] A text encoder is used to encode the textual prompts for the preset object and the extended descriptions of the nouns for each component within the preset object, generating the textual semantic features of the preset object. The total number of components within the preset object is calculated. A K-means clustering algorithm is used to divide the preset defect cloud based on the total number of components, generating multiple non-overlapping local regions within the preset defect cloud. These non-overlapping local regions are then input into a 3D point cloud encoder, which extracts features from them, generating local geometric features of the preset defect cloud. The entire region of the preset defect cloud is also input into the 3D point cloud encoder, which extracts features from it, generating global geometric features. A cross-attention mechanism is then used to fuse the local geometric features, the corresponding global geometric features, and the textual semantic features of the preset object across modalities, generating the fused features of the preset object.

[0077] K-means clustering is a distance-based unsupervised learning algorithm whose core logic is to divide the data into K compact and independent clusters through iterative optimization.

[0078] In this embodiment, a cross-attention mechanism is used to fuse the local geometric features of the preset defect cloud, the global geometric features corresponding to the preset defect cloud, and the textual semantic features of the preset object across modal features to generate the fused features of the preset object. This approach enables the fused features of the preset object to have richer representation capabilities and stronger robustness, which can effectively improve the overall model's ability to understand incomplete, noisy, and point clouds with different poses.

[0079] Please see Figure 2 , Figure 2 This is a flowchart illustrating the point cloud completion method provided in this application embodiment, which can be applied to electronic devices.

[0080] like Figure 2 As shown, the point cloud completion method provided in this application includes the following steps, detailed below:

[0081] S201, Obtain a preset defect cloud, and determine the fusion features of a preset object based on the preset defect cloud;

[0082] The step of obtaining a preset defect cloud and determining the fusion features of a preset object based on the preset defect cloud includes:

[0083] Obtain a sample set, and obtain training samples from the sample set. The training samples include the preset defect cloud corresponding to the preset object, the text prompt information corresponding to the preset object, and the real point cloud of each component in the preset object.

[0084] Extract the nouns of each component within the preset object from the text prompts corresponding to the preset object, input the nouns of each component within the preset object into the large language model, and generate an extended description of the nouns of each component within the preset object through the large language model.

[0085] A text encoder is used to encode the textual prompts for the preset object and the extended descriptions of the nouns for each component within the preset object, generating the textual semantic features of the preset object. The total number of components within the preset object is calculated. A K-means clustering algorithm is used to divide the preset defect cloud based on the total number of components, generating multiple non-overlapping local regions within the preset defect cloud. These non-overlapping local regions are then input into a 3D point cloud encoder, which extracts features from them, generating local geometric features of the preset defect cloud. The entire region of the preset defect cloud is also input into the 3D point cloud encoder, which extracts features from it, generating global geometric features. A cross-attention mechanism is then used to fuse the local geometric features, the corresponding global geometric features, and the textual semantic features of the preset object across modalities, generating the fused features of the preset object.

[0086] The total number of components of a preset object is the total number of all components within the preset object.

[0087] For ease of explanation, the following example is provided:

[0088] For example, if the total number of parts of a preset object is 4, the K-means clustering algorithm is used to divide the preset defect cloud and generate 4 non-overlapping local regions within the preset defect cloud.

[0089] For example, if the total number of parts of a preset object is 5, the K-means clustering algorithm is used to divide the preset defect cloud and generate 5 non-overlapping local regions within the preset defect cloud.

[0090] The preset object is a preset airplane, and the target object is the airplane to be processed.

[0091] For ease of explanation, the following example is provided:

[0092] For example, the preset airplane has four parts, and the names of the four parts of the preset airplane are: preset nose, preset fuselage, preset wings, and preset tail.

[0093] The default extended description of the machine head is: a machine head with a sharp front end and a streamlined shape;

[0094] The default description of the fuselage extension is: a long cylindrical main fuselage used to support the overall structure;

[0095] The default description of the wings is: wings that are symmetrically deployed on the left and right sides of the fuselage;

[0096] The default description of the tail fin extension is: the tail section of the fuselage, used to stabilize the flight attitude.

[0097] S202, input the fusion features of the preset object into the diffusion model, and generate the predicted point cloud of each component within the preset object through the diffusion model;

[0098] S203, Based on the predicted point cloud of each component within the preset object, the real point cloud of each component within the preset object, and the first loss model, generate the diffusion loss of the diffusion model, and based on the second loss model, generate the completion loss of the completion model.

[0099] The first loss model is defined as follows:

[0100] ;

[0101] This represents the diffusion loss in the diffusion model;

[0102] This represents the KL divergence of the diffusion model;

[0103] This represents the total number of sub-parts after an object has been decomposed.

[0104] Indicates the sequence number of the sub-components within the preset object;

[0105] This represents the actual point cloud of the i-th sub-component within a predefined object;

[0106] This represents the predicted point cloud of the i-th sub-component within the preset object at the i-th step of the diffusion process;

[0107] This represents the predicted point cloud of the i-th sub-component within the preset object in the step preceding the i-th step of the diffusion process;

[0108] This represents the condition vector of the i-th sub-component within the preset object;

[0109] Represents the true distribution;

[0110] This represents the expectation under the true distribution;

[0111] These represent the model parameters of the diffusion model;

[0112] This represents the predicted probability distribution generated by the diffusion model based on the model parameters.

[0113] The full Chinese name for KL divergence is Kullback Leibler Divergence, and its full English name is Kullback Leibler Divergence. The second loss model is defined as follows:

[0114] ;

[0115] This represents the completion loss of the completion model;

[0116] This represents the predicted point cloud of a predefined object. Represents the true point cloud of a predefined object;

[0117] This indicates the total number of points for the preset object; This represents the total number of points in the actual point cloud of the preset object;

[0118] To represent any point in the predicted point cloud; Represents any point in a real point cloud;

[0119] express and The Euclidean distance between them;

[0120] express and The Euclidean distance between them.

[0121] S204. Based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, generate the total loss of the overall model, update the model parameters of the overall model, and save the updated overall model only when the total loss of the overall model meets the preset conditions.

[0122] The overall model consists of a diffusion model and a completion model. The diffusion model generates the global content, ensuring the overall texture, style, and semantic consistency. The completion model completes the image by filling in missing regions in the global content output by the diffusion model.

[0123] The total loss model is defined as follows:

[0124] ;

[0125] This represents the total loss of the entire model;

[0126] This represents the diffusion loss in the diffusion model;

[0127] This represents the completion loss of the completion model; It is a parameter used to adjust the proportion of diffusion loss; It is a parameter used to adjust the proportion of compensation loss; .

[0128] S205: Obtain the current defect cloud corresponding to the target object and the text prompt information corresponding to the target object. Based on the text prompt information corresponding to the target object, determine the text semantic features of the preset object. Based on multiple non-overlapping local regions in the current defect cloud, the text semantic features of the preset object, and the predefined method, determine the fusion features of the target object. The updated overall model completes the current defect cloud based on the fusion features of the target object, generating the complete point cloud of the target object.

[0129] The preset object is a preset airplane, and the target object is the airplane to be processed.

[0130] Among these methods, utilizing the updated overall model based on the fusion features of the aircraft to be processed to complete the current incomplete point cloud and generate a complete point cloud of the aircraft can play an important role in practical aerospace engineering. The overall model can accurately reconstruct the complete 3D geometry of missing areas such as the nose, fuselage, wings, and tail of the aircraft to be processed, based on learned prior knowledge of the aircraft structure, thus obtaining a complete point cloud of the aircraft. This complete point cloud of the aircraft can be directly used for aircraft digital modeling, assembly accuracy inspection, component damage assessment, flight simulation modeling, and digital twin construction. It can effectively reduce the time spent on manual repairs, decrease reliance on high-precision complete scans, improve the efficiency and reliability of 3D model reconstruction, and provide data support for aircraft design.

[0131] The process involves: acquiring the current residual point cloud corresponding to the target object and the corresponding text prompt information; determining the text semantic features of the preset object based on the text prompt information; determining the fusion features of the target object based on multiple non-overlapping local regions within the current residual point cloud, the text semantic features of the preset object, and a predefined method; and finally, completing the current residual point cloud based on the fusion features of the target object to generate a complete point cloud of the target object, including:

[0132] Obtain the current defect cloud and text prompt information corresponding to the target object. Extract the nouns of each component within the target object from the text prompt information. Input the nouns of each component within the target object into the large language model. Generate an extended description of the nouns of each component in the current defect cloud through the large language model.

[0133] The text encoder encodes the text prompts corresponding to the target object and the extended descriptions of the nouns of each component within the target object, generating the text semantic features of the target object. It also counts the number of components within the target object and uses the K-means clustering algorithm to divide the current defect cloud based on the total number of components, generating multiple non-overlapping local regions within the current defect cloud. These non-overlapping local regions are then input into a 3D point cloud encoder, which extracts features from them to generate the local geometric features of the current defect cloud. Finally, the entire region of the current defect cloud is input into the 3D point cloud encoder, which extracts features from it to generate the global geometric features of the current defect cloud.

[0134] A cross-attention mechanism is employed to fuse the local geometric features of the current incomplete point cloud, the corresponding global geometric features, and the textual semantic features of the target object across modalities, generating a fused feature of the target object. This fused feature is then input into an updated diffusion model, which generates a point cloud of the missing component. Affine transformation parameters are used to perform rotation, scaling, shearing, and translation operations on the point cloud of the missing component, generating an adjusted point cloud of the missing component. The adjusted point cloud of the component is then used to complete the current incomplete point cloud, generating a complete point cloud of the target object.

[0135] The preset object is a preset airplane, and the target object is the airplane to be processed.

[0136] For ease of explanation, the following example is provided:

[0137] For example, the aircraft to be processed has four parts, and the names of the four parts of the aircraft to be processed are: the nose to be processed, the fuselage to be processed, the wings to be processed, and the tail to be processed.

[0138] The extended description of the machine head to be processed is: a machine head with an ellipsoidal front end, an overall smooth, blunt-tipped ellipsoidal outline, and no sharp edges;

[0139] The extended description of the fuselage to be processed is: adopting a wide-body, short and stout load-bearing fuselage;

[0140] The extended description of the two wings to be processed is: wings that are symmetrically deployed on the left and right sides of the fuselage;

[0141] The extended description of the tail fin to be processed is as follows: It is composed of two symmetrically inclined wing surfaces, with an overall V-shaped structure. The wing surfaces are smooth and thin, with regular outlines and streamlined transitions at the edges.

[0142] The total number of parts of the current object is the total number of all parts within the current object.

[0143] For ease of explanation, the following example is provided:

[0144] For example, if the current object has a total of 4 parts, the K-means clustering algorithm is used to divide the current defect cloud and generate 4 non-overlapping local regions within the current defect cloud.

[0145] For example, if the current object has a total of 5 parts, the K-means clustering algorithm is used to divide the current defect cloud and generate 5 non-overlapping local regions within the current defect cloud.

[0146] Specifically, by performing rotation, scaling, shearing, and translation operations on the point cloud of the missing component through affine transformation parameters, an adjusted point cloud of the missing component is generated. The adjusted point cloud of the component is then used to complete the current incomplete point cloud, generating a complete point cloud of the target object. This method can accurately fill in the missing area while preserving the current incomplete point cloud, making the complete point cloud of the target object more consistent with the actual physical spatial distribution.

[0147] For clarity, please refer to Figure 6 , Figure 6 This is a flowchart of the point cloud completion method provided in the embodiments of this application, which is described in detail below:

[0148] The text prompt for the aircraft to be processed is: This is a symmetrical aircraft with a nose, fuselage, wings and tail.

[0149] The extended description of the machine head to be processed is as follows:

[0150] 1. The head is a smooth, conical dome that blends seamlessly with the cylindrical body.

[0151] 2. The fuselage is slender and cylindrical, narrowing at both ends, with a smooth surface.

[0152] 3. The two wings are flat airfoil structures that extend laterally from the middle of the fuselage.

[0153] 4. The entire tail section consists of a vertical tail section that rises from the rear of the fuselage and two horizontally symmetrical tail sections on both sides.

[0154] Let represent the local geometric features of the current defect cloud, let represent the global geometric features corresponding to the current defect cloud, and let C represent the textual semantic features of the target object.

[0155] A cross-attention mechanism is employed to fuse the cross-modal features of the target object with C, generating a fused feature of the target object. The fused feature of the target object is then input into the updated diffusion model, which generates a point cloud of the missing part. Affine transformation parameters are used to perform rotation, scaling, shearing, and translation operations on the point cloud of the missing part to generate an adjusted point cloud of the missing part. The adjusted point cloud of the part is then used to complete the current incomplete point cloud, generating a complete point cloud of the target object.

[0156] The beneficial effects of the embodiments of this application are as follows:

[0157] Firstly, the current residual point cloud corresponding to the target object and the text prompt information corresponding to the target object are obtained. Based on the text prompt information corresponding to the target object, the text semantic features of the preset object are determined. Based on multiple non-overlapping local regions in the current residual point cloud, the text semantic features of the preset object and the predefined method, the fusion features of the target object are determined. The updated overall model completes the current residual point cloud based on the fusion features of the target object to generate the complete point cloud of the target object. Since no manual completion is required, the completion time of the complete point cloud of the target object is reduced, which is conducive to improving the completion efficiency of the complete point cloud of the target object.

[0158] Secondly, since the complete point cloud of the target object is automatically generated without being affected by human intervention, it helps to improve the reliability of the complete point cloud of the target object.

[0159] Please see Figure 3 , Figure 3 The implementation flowchart of S204 provided in the embodiments of this application is described in detail below:

[0160] S301, Based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, generate the total loss of the overall model;

[0161] S302 uses backpropagation to update the model parameters of the overall model until the total loss of the overall model is less than the preset loss value, then stops updating the model parameters of the overall model and saves the updated overall model.

[0162] In this embodiment, the model parameters of the overall model are not updated until the total loss of the overall model is less than a preset loss value. The updated overall model is then saved. This effectively filters out model parameters that meet the accuracy requirements and avoids the accidental saving of intermediate models with poor generalization ability, thereby improving the reliability of the updated overall model.

[0163] For the point cloud completion method described in the above embodiments, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic block diagram of the point cloud completion device provided in the embodiments of this application. Figure 4 The point cloud completion device 400 shown can be applied to, for example... Figure 1 The application scenario diagram shows electronic devices. The following section uses electronic devices as an example to illustrate this. Figure 4 The point cloud completion device 400 shown will be described in detail. The point cloud completion device 400 may include an acquisition module 401, an input module 402, a first generation module 403, an update module 404, and a second generation module 405.

[0164] The acquisition module 401 is used to acquire a preset defect cloud and determine the fusion features of a preset object based on the preset defect cloud;

[0165] The input module 402 is used to input the fusion features of the preset object into the diffusion model, and generate the predicted point cloud of each component in the preset object through the diffusion model.

[0166] The first generation module 403 is used to generate the diffusion loss of the diffusion model based on the predicted point cloud of each component in the preset object, the real point cloud of each component in the preset object and the first loss model, and to generate the completion loss of the completion model based on the second loss model.

[0167] The update module 404 is used to generate the total loss of the overall model based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, and update the model parameters of the overall model until the total loss of the overall model meets the preset conditions before saving the updated overall model.

[0168] The second generation module 405 is used to obtain the current defect cloud corresponding to the target object and the text prompt information corresponding to the target object. Based on the text prompt information corresponding to the target object, the text semantic features of the preset object are determined. Based on multiple non-overlapping local regions in the current defect cloud, the text semantic features of the preset object and the predefined method, the fusion features of the target object are determined. The updated overall model completes the current defect cloud based on the fusion features of the target object to generate the complete point cloud of the target object.

[0169] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0170] The beneficial effects of the embodiments of this application are as follows:

[0171] Firstly, the current residual point cloud corresponding to the target object and the text prompt information corresponding to the target object are obtained. Based on the text prompt information corresponding to the target object, the text semantic features of the preset object are determined. Based on multiple non-overlapping local regions in the current residual point cloud, the text semantic features of the preset object and the predefined method, the fusion features of the target object are determined. The updated overall model completes the current residual point cloud based on the fusion features of the target object to generate the complete point cloud of the target object. Since no manual completion is required, the completion time of the complete point cloud of the target object is reduced, which is conducive to improving the completion efficiency of the complete point cloud of the target object.

[0172] Secondly, since the complete point cloud of the target object is automatically generated without being affected by human intervention, it helps to improve the reliability of the complete point cloud of the target object.

[0173] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0174] like Figure 5 As shown, Figure 5 The electronic device includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 executes the computer program 22 to implement the steps in any of the above method embodiments.

[0175] The electronic device may include, but is not limited to, processor 20 and memory 21. Those skilled in the art will understand that... Figure 5 This is merely an example of an electronic device and does not constitute a limitation on electronic devices. It may include more or fewer components than shown in the illustration, or combinations of certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0176] The processor 20 is used to run a computer program 22 stored in the memory 21, and performs the following steps when executing the computer program 22:

[0177] Obtain a preset defect cloud, and determine the fusion features of the preset object based on the preset defect cloud;

[0178] The fusion features of the preset object are input into the diffusion model, and the diffusion model generates the predicted point cloud of each component within the preset object.

[0179] Based on the predicted point cloud of each component within the preset object, the real point cloud of each component within the preset object, and the first loss model, the diffusion loss of the diffusion model is generated, and the completion loss of the completion model is generated based on the second loss model.

[0180] Based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, the total loss of the overall model is generated, and the model parameters of the overall model are updated until the total loss of the overall model meets the preset conditions, at which point the updated overall model is saved.

[0181] The current residual point cloud and text prompt information corresponding to the target object are obtained. Based on the text prompt information corresponding to the target object, the text semantic features of the preset object are determined. Based on multiple non-overlapping local regions in the current residual point cloud, the text semantic features of the preset object and the predefined method, the fusion features of the target object are determined. The updated overall model is based on the fusion features of the target object to complete the current residual point cloud and generate the complete point cloud of the target object.

[0182] The processor 20 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors, field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0183] In some embodiments, the memory 21 may be an internal storage unit of the electronic device, such as a hard disk or memory of the electronic device. In other embodiments, the memory 21 may also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device.

[0184] Furthermore, the memory 21 may include both internal storage units and external storage devices of the electronic device. The memory 21 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 21 can also be used to temporarily store data that has been output or will be output.

[0185] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0186] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0187] The computer-readable storage medium may also be an external storage device of the point cloud completion device or electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, or non-transitory computer-readable storage medium equipped on the point cloud completion device or electronic device.

[0188] Since the computer program stored in the computer-readable storage medium can execute any of the text-based point cloud completion methods provided in the embodiments of this application, the computer-readable storage medium can achieve the beneficial effects that any of the text-based point cloud completion methods provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0189] This application provides a computer program product that, when run on an electronic device, causes the electronic device to execute the aforementioned point cloud completion method.

[0190] When a computer program is loaded into an electronic device, it can perform the following steps:

[0191] Obtain a preset defect cloud, and determine the fusion features of the preset object based on the preset defect cloud;

[0192] The fusion features of the preset object are input into the diffusion model, and the diffusion model generates the predicted point cloud of each component within the preset object.

[0193] Based on the predicted point cloud of each component within the preset object, the real point cloud of each component within the preset object, and the first loss model, the diffusion loss of the diffusion model is generated, and the completion loss of the completion model is generated based on the second loss model.

[0194] Based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, the total loss of the overall model is generated, and the model parameters of the overall model are updated until the total loss of the overall model meets the preset conditions, at which point the updated overall model is saved.

[0195] The current residual point cloud and text prompt information corresponding to the target object are obtained. Based on the text prompt information corresponding to the target object, the text semantic features of the preset object are determined. Based on multiple non-overlapping local regions in the current residual point cloud, the text semantic features of the preset object and the predefined method, the fusion features of the target object are determined. The updated overall model is based on the fusion features of the target object to complete the current residual point cloud and generate the complete point cloud of the target object.

[0196] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0197] Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium includes: an entity or device for carrying computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium.

[0198] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0199] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A point cloud completion method based on text prompts, characterized in that, The point cloud completion method, applied to electronic devices, includes: Obtain a preset defect cloud, and determine the fusion features of the preset object based on the preset defect cloud; The fusion features of the preset object are input into the diffusion model, and the diffusion model generates the predicted point cloud of each component within the preset object. Based on the predicted point cloud of each component within the preset object, the real point cloud of each component within the preset object, and the first loss model, the diffusion loss of the diffusion model is generated, and the completion loss of the completion model is generated based on the second loss model. Based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, the total loss of the overall model is generated, and the model parameters of the overall model are updated until the total loss of the overall model meets the preset conditions, at which point the updated overall model is saved. The process involves acquiring the current defect cloud and textual prompts corresponding to the target object. From the textual prompts, the nouns for each component within the target object are extracted. These nouns are then input into a large language model, which generates extended descriptions for each component noun in the current defect cloud. A text encoder encodes the textual prompts and extended descriptions for each component, generating semantic features of the target object. The total number of components is then calculated. K-means clustering is used to divide the current defect cloud into multiple non-overlapping local regions. These non-overlapping local regions are input into a 3D point cloud encoder, which then processes the current defect cloud. Feature extraction is performed on multiple non-overlapping local regions within the point cloud to generate local geometric features of the current incomplete point cloud. The entire region of the current incomplete point cloud is then input into a 3D point cloud encoder, which extracts features from the entire region to generate global geometric features. A cross-attention mechanism is used to fuse the local geometric features, the corresponding global geometric features, and the textual semantic features of the target object across modalities to generate fused features of the target object. These fused features are then input into an updated diffusion model to generate point clouds of the missing components. Affine transformation parameters are used to perform rotation, scaling, shearing, and translation operations on the point clouds of the missing components to generate adjusted point clouds of the missing components. The adjusted point clouds of the components are then used to complete the current incomplete point cloud, generating the complete point cloud of the target object.

2. The point cloud completion method according to claim 1, characterized in that, The process of obtaining a preset residual cloud and determining the fusion features of a preset object based on the preset residual cloud includes: Obtain a sample set, and obtain training samples from the sample set. The training samples include the preset defect cloud corresponding to the preset object, the text prompt information corresponding to the preset object, and the real point cloud of each component in the preset object. Extract the nouns of each component within the preset object from the text prompts corresponding to the preset object, input the nouns of each component within the preset object into the large language model, and generate an extended description of the nouns of each component within the preset object through the large language model. A text encoder is used to encode the textual prompts for the preset object and the extended descriptions of the nouns for each component within the preset object, generating the textual semantic features of the preset object. The total number of components within the preset object is calculated. A K-means clustering algorithm is used to divide the preset defect cloud based on the total number of components, generating multiple non-overlapping local regions within the preset defect cloud. These non-overlapping local regions are then input into a 3D point cloud encoder, which extracts features from them, generating local geometric features of the preset defect cloud. The entire region of the preset defect cloud is also input into the 3D point cloud encoder, which extracts features from it, generating global geometric features. A cross-attention mechanism is then used to fuse the local geometric features, the corresponding global geometric features, and the textual semantic features of the preset object across modalities, generating the fused features of the preset object.

3. The point cloud completion method according to claim 1, characterized in that, The process involves generating the total loss of the overall model based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model; updating the model parameters of the overall model; and saving the updated overall model only after the total loss of the overall model meets a preset condition. This includes: Based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, the total loss of the overall model is generated. The model parameters of the overall model are updated using backpropagation until the total loss of the overall model is less than the preset loss value. Only then is the update of the overall model parameters stopped and the updated overall model saved.

4. The point cloud completion method according to claim 1, characterized in that, The first loss model is defined as follows: ; This represents the diffusion loss in the diffusion model; This represents the KL divergence of the diffusion model; This represents the total number of sub-parts after an object has been decomposed. Indicates the sequence number of the sub-components within the preset object; Indicates the first [object] within the preset object The actual point cloud of each sub-component; Indicates the first [object] within the preset object The sub-components in the diffusion process Predicted point cloud of the step; Indicates the first [object] within the preset object The sub-components in the diffusion process The predicted point cloud of the previous step; Indicates the first [object] within the preset object The corresponding condition vectors of each sub-component; Represents the true distribution; This represents the expectation under the true distribution; These represent the model parameters of the diffusion model; This represents the predicted probability distribution generated by the diffusion model based on the model parameters.

5. The point cloud completion method according to claim 1, characterized in that, The second loss model is defined as follows: ; This represents the completion loss of the completion model; This represents the predicted point cloud of a predefined object. Represents the true point cloud of a predefined object; This indicates the total number of points for the preset object; This represents the total number of points in the actual point cloud of the preset object. To represent any point in the predicted point cloud; Represents any point in a real point cloud; express and The Euclidean distance between them; express and The Euclidean distance between them.

6. The point cloud completion method according to claim 1, characterized in that, The total loss model is defined as follows: ; This represents the total loss of the entire model; This represents the diffusion loss in the diffusion model; This represents the completion loss of the completion model; It is a parameter used to adjust the proportion of diffusion loss; It is a parameter used to adjust the proportion of compensation loss; .

7. A point cloud completion device based on text prompts, characterized in that, Applied to electronic devices, including: The acquisition module is used to acquire a preset defect cloud and determine the fusion features of a preset object based on the preset defect cloud; The input module is used to input the fusion features of the preset object into the diffusion model, and the diffusion model generates the predicted point cloud of each component within the preset object. The first generation module is used to generate the diffusion loss of the diffusion model based on the predicted point cloud of each component in the preset object, the real point cloud of each component in the preset object and the first loss model, and to generate the completion loss of the completion model based on the second loss model. The update module is used to generate the total loss of the overall model based on the diffusion loss of the diffusion model, the completion loss of the completion model, and the total loss model, and update the model parameters of the overall model until the total loss of the overall model meets the preset conditions before saving the updated overall model. The second generation module is used to acquire the current defect cloud corresponding to the target object and the text prompt information corresponding to the target object. From the text prompt information, it extracts the nouns of each component within the target object. These nouns are input into a large language model, which generates extended descriptions of each component noun in the current defect cloud. A text encoder encodes the text prompt information and the extended descriptions of each component noun to generate the text semantic features of the target object. It then counts the number of components within the target object and uses a K-means clustering algorithm to divide the current defect cloud based on the total number of components, generating multiple non-overlapping local regions within the current defect cloud. These non-overlapping local regions are input into a 3D point cloud encoder. Feature extraction is performed on multiple non-overlapping local regions within the current incomplete point cloud to generate local geometric features. The overall region of the current incomplete point cloud is then input into a 3D point cloud encoder, which extracts features from the overall region to generate global geometric features. A cross-attention mechanism is used to fuse the local geometric features, the corresponding global geometric features, and the textual semantic features of the target object across modalities to generate fused features of the target object. These fused features are then input into an updated diffusion model to generate point clouds of the missing components. Affine transformation parameters are used to perform rotation, scaling, shearing, and translation operations on the point clouds of the missing components to generate adjusted point clouds of the missing components. The adjusted point clouds of the components are then used to complete the current incomplete point cloud, generating the complete point cloud of the target object.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the point cloud completion method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the point cloud completion method as described in any one of claims 1 to 6.