A 3D Pre-training Method and System Based on Diffusion Model

By adopting a diffusion model-based method in 3D point cloud pre-training, the problems of performance degradation and increase in computational volume caused by negative samples and large batch sizes in the prior art are solved, and more efficient point cloud pre-training is achieved, and richer geometric information and better model performance are obtained.

CN117115588BActive Publication Date: 2025-06-17SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311085634.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-25
Publication Date
2025-06-17
Estimated Expiration
2043-08-25

AI Technical Summary

Technical Problem

The existing 3D point cloud pre-training methods have problems of performance degradation and training instability when designing negative samples and large batches, and mask modeling increases the network computing volume.

Method used

Using a 3D pre-training method based on diffusion model, the visible point cloud word segmentation features are encoded through an encoder, combined with the masked point cloud word segmentation features are spliced, and the diffusion model is generated through the conditional embedding module. The loss function is determined according to the multi-stage uniform sampling strategy, and the decoder is guided to decode to complete the training.

Benefits of technology

Without changing the network architecture and increasing the amount of computing, the pre-training method is improved, and the multi-level boot characteristics of the diffusion model can be used to achieve a comprehensive understanding of the point cloud, obtain richer geometric information, improve the effect and performance of model training, and save the cost of point cloud pre-training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117115588B_ABST
    Figure CN117115588B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a 3D pre-training method and system based on a diffusion model. The method encodes the visible point cloud token features through an encoder to obtain deep visible point cloud token features, wherein the visible point cloud token features are generated from the acquired point cloud blocks; the deep visible point cloud token features are concatenated with the masked point cloud token features to obtain concatenated features, wherein the masked point cloud token features are generated from the acquired point cloud blocks; the concatenated features are encoded through a conditional embedding module to obtain the guiding conditions of the diffusion model; a loss function is determined according to the guiding conditions of the diffusion model and a multi-stage uniform sampling strategy, and the loss function is used to guide the decoder to perform decoding to complete the training. The solution of the present invention realizes a comprehensive understanding of point clouds without changing the network architecture and without increasing the computational amount, improves the training effect and performance of the model, and further saves the pre-training cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer vision technology, and in particular, to a 3D pre-training method and system based on a diffusion model. Background Art

[0002] With the rapid development of the 3D computer vision field, point clouds have become an important data form for describing the shape and spatial position of three-dimensional objects. The application scope of point cloud data is very wide, including fields such as three-dimensional scene reconstruction, autonomous driving, and intelligent robots. The research on 3D point cloud pre-training technology can provide important support for these tasks, thus greatly reducing the time and resource costs of artificial intelligence model development and deployment.

[0003] The current mainstream 3D pre-training methods can be divided into two categories: contrastive learning and masked modeling. Contrastive learning performs pre-training by comparing the similarities and differences between samples. For example, in the field of 3D point clouds, two point clouds are constructed from different perspectives, and the similarity of point features between the two point clouds is compared to achieve the pre-training of the point cloud network; or only using a single-view depth map for feature contrast learning can also learn effective feature representations; or according to the cross-modal contrastive learning method, comparing the point cloud with the rendered 2D image for contrastive learning helps to learn generalizable point cloud representations. However, 3D pre-training methods based on contrastive learning often require designing negative samples and a large batch size (batchsize). If the negative samples are selected improperly, the model may learn incorrect feature representations, resulting in performance degradation; a large batch size may lead to a more unstable training process, poor training effects, and affect the model performance, resulting in the risk of model collapse. Masked modeling does not require designing negative samples and a large batch size, but current masked modeling mostly learns multi-level geometric knowledge by modifying the hierarchical network architecture, resulting in an increase in network computational complexity. Summary of the Invention

[0004] In view of this, the embodiments of the present invention provide a 3D pre-training method and system, an electronic device, and a computer storage medium based on a diffusion model to at least partially solve the above problems.

[0005] According to the first aspect of the embodiments of the present invention, a 3D pre-training method based on a diffusion model is provided, including encoding the visible point cloud token features through an encoder to obtain deep visible point cloud token features, where the visible point cloud token features are generated from the obtained point cloud blocks; splicing the deep visible point cloud token features with the masked point cloud token features to obtain a spliced feature, where the masked point cloud token features are generated from the obtained point cloud blocks; encoding the spliced feature through a conditional embedding module to obtain the guiding condition of the diffusion model; determining a loss function according to the guiding condition of the diffusion model and a multi-stage uniform sampling strategy, and the loss function is used to guide the decoder to perform decoding to complete the training.

[0006] In one implementation, before encoding the visible point cloud tokenization features through an encoder to obtain deep visible point cloud tokenization features, it includes obtaining point cloud patches through the farthest point sampling algorithm and the K-nearest neighbor algorithm; centering the point cloud patches to update the point cloud patches; embedding the updated point cloud patches into the tokenization features to obtain point cloud tokenization features; randomly masking the point cloud tokenization features to generate visible point cloud tokenization features and masked point cloud tokenization features, where the visible point cloud tokenization features are represented as The masked point cloud tokenization features are represented as is the number of masked point cloud tokenization features, g = s - r is the number of visible point cloud tokenization features and m is the ratio of random masking.

[0007] In another implementation, obtaining point cloud patches through the farthest point sampling algorithm and the K-nearest neighbor algorithm includes sampling the preset point cloud data through the farthest point sampling algorithm to obtain s center points, where the preset point cloud data contains n points and the preset point cloud data is represented as X ∈ R n×3 , and the s center points are represented as Based on each center point C of the s center points i , k points closest to each center point are selected as point cloud patches through the K-nearest neighbor algorithm, where each center point is represented as C i , and the point cloud patches are represented as Then the point cloud patch The calculation formula is:

[0008]

[0009] In another implementation, embedding the updated point cloud patches into the tokenization features to obtain point cloud tokenization features includes using the simplified PointNetξ φ (·) with parameter φ to embed the updated point cloud patches into the tokenization features , which is represented as:

[0010]

[0011] In another implementation, encoding the visible point cloud tokenization features through an encoder to obtain deep visible point cloud tokenization features includes inputting the position information of the visible point cloud tokenization features into the position embedding formula ψ τ (·) with parameter τ, outputting the position features of the visible point cloud tokenization features, and the position features of the visible point cloud tokenization features are represented as Outputting the position features of the visible point cloud tokenization features and the visible point cloud tokenization features Connect to obtain combined features; through the encoder Φ ρ (·) of the Transformer block extracts features from the combined features to obtain deep visible point cloud tokenization features, and the deep visible point cloud tokenization features are represented as Among them, the encoder Φ ρ (·) consists of standard Transformer blocks with parameters ρ. Then, the visible point cloud tokenization features are encoded through the encoder, and the formula for obtaining the deep visible point cloud tokenization features is:

[0012]

[0013] In another implementation, the concatenated features are encoded through a conditional embedding module to obtain the guidance conditions for the diffusion model, including encoding the concatenated features through the conditional embedding module f ω (·) with parameters ω to obtain the guidance conditions c for the diffusion model. Among them, the formula for the guidance conditions c of the diffusion model is:

[0014]

[0015] In another implementation, the loss function is determined according to the guidance conditions of the diffusion model and the multi-stage uniform sampling strategy, including defining the training objective of the pre-trained model according to the guidance conditions c of the diffusion model, and the formula is:

[0016]

[0017] The time step interval [1, T] for each preset point cloud data is divided into h stages:

[0018] Among them

[0019] Randomly sample a time step t from the h stages, calculate the loss h times and perform an average operation on the h losses to obtain the loss function:

[0020]

[0021] According to the second aspect of the embodiments of the present invention, a 3D pre-training system based on a diffusion model is provided, including a first encoding module for encoding the visible point cloud token features through an encoder to obtain deep visible point cloud token features, where the visible point cloud token features are generated from the acquired point cloud blocks; a splicing module for splicing the deep visible point cloud token features with the masked point cloud token features to obtain spliced features, where the masked point cloud token features are generated from the acquired point cloud blocks; a second encoding module for encoding the spliced features through a conditional embedding module to obtain the guiding conditions of the diffusion model; and a determination module for determining a loss function according to the guiding conditions of the diffusion model and a multi-stage uniform sampling strategy, where the loss function is used to guide the decoder to perform decoding to complete the training.

[0022] According to the third aspect of the embodiments of the present invention, an electronic device is provided, including a processor and a memory storing a program. Wherein, the program includes instructions that, when executed by the processor, cause the processor to execute the method as in the first aspect.

[0023] According to the fourth aspect of the embodiments of the present invention, a computer storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method as in the first aspect is implemented.

[0024] In summary, the solution of the embodiments of the present invention, without changing the network architecture and without increasing the computational amount, through improving the pre-training method and utilizing the multi-stage guiding characteristics of the diffusion model, realizes a comprehensive understanding of the point cloud, obtains richer geometric information, improves the effect and performance of model training, and further saves the cost of point cloud pre-training. Description of the Drawings

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0026] Figure 1 It is a flowchart of the steps of the 3D pre-training method based on the diffusion model according to the embodiments of the present invention.

[0027] Figure 2 It is an overall framework diagram of the 3D pre-training method based on the diffusion model according to the embodiments of the present invention.

[0028] Figure 3A It is a visualization result diagram of the 3D pre-training method based on the diffusion model according to the embodiments of the present invention.

[0029] Figure 3B It is a comparison diagram of semantic segmentation results according to another embodiment of the present invention.

[0030] Figure 4 To correspond to Figure 1 The structural block diagram of the 3D pre-training system based on the diffusion model corresponding to the embodiment.

[0031] Figure 5 The structural schematic diagram of an electronic device according to another embodiment of the present invention. Detailed implementation manners

[0032] In order to have a clearer understanding of the technical features, objectives, and effects of the embodiments of the present application, the specific implementation manners of the embodiments of the present application will now be described with reference to the accompanying drawings.

[0033] In this document, "schematic" means "serving as an instance, example, or illustration", and any illustration or implementation manner described as "schematic" in this document should not be construed as a more preferred or more advantageous technical solution.

[0034] To simplify the drawings, only the parts related to the present application are schematically shown in each figure, and they do not represent their actual structures as products. Additionally, to simplify the drawings for easy understanding, in some figures, for components with the same structure or function, only one or more of them are schematically shown, or only one or more of them are labeled.

[0035] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present invention.

[0036] For ease of understanding, before describing the specific embodiments of the present invention in detail, an exemplary description of the prior art of the present invention will be given first.

[0037] Many studies in the existing technology, including VisualChatGPT[1]Wu C, Yin S, Qi W, et al. Visual chatgpt: Talking, drawing and editing with visual foundation models[J]. arXiv preprint arXiv:2303.04671, 2023. and SAM[2]Kirillov A, Mintun E, RaviN, et al. Segment anything[J]. arXiv preprint arXiv:2304.02643, 2023., have demonstrated the excellent performance of pre-trained models in a wide range of downstream tasks. With the rapid development of the 3D computer vision field, point clouds have become an important data form for describing the shape and spatial position of three-dimensional objects. The application scope of point cloud data is very wide, including fields such as three-dimensional scene reconstruction, autonomous driving, and intelligent robots. The research on 3D point cloud pre-training technology can provide important support for these tasks, thus greatly reducing the time and resource costs of artificial intelligence model development and deployment. The current mainstream 3D pre-training methods can be divided into two categories: contrastive learning and masked modeling.

[0038] Contrastive learning performs pre-training by comparing the similarities and differences between samples. For example:

[0039] In the 3D point cloud field, PointContrast[3]Xie S, Gu J, Guo D, et al. Pointcontrast: Unsupervised pre-training for 3d point cloud understanding[C] / / ComputerVision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16. Springer International Publishing, 2020: 574-591. constructs two point clouds from different perspectives and pre-trains the point cloud network by comparing the similarity of point features between the two point clouds.

[0040] DepthContrast[4]Zhang Z,Girdhar R,Joulin A,et al.Self-supervised pretraining of 3d features on any point-cloud[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2021:10252-10263.Feature contrast learning using only the depth map of a single view can also learn effective feature representations.

[0041] CrossPoint[5]Afham M,Dissanayake I,Dissanayake D,et al.Crosspoint:Self-supervised cross-modal contrastive learning for 3d point cloud understanding[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2022:9902-9912.A cross-modal contrastive learning method is proposed, which performs contrastive learning between point clouds and rendered 2D images, helping to learn generalizable point cloud representations.

[0042] However, the above contrastive learning-based methods often require the design of negative samples and large batch sizes, and at the same time cannot completely avoid model collapse, thus affecting performance.

[0043] Different from contrastive learning, masked modeling pre-trains the encoder by recovering masked information, without the need to design negative samples and large batch sizes, where:

[0044] Point-BERT[6] Yu X, Tang L, Rao Y, et al. Point-bert: Pre-training 3d pointcloud transformers with masked point modeling[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022: 19313-19322. Convert point cloud patches into discrete tokens through dVAE[7] Rolfe J T. Discrete variational autoencoders[J]. arXiv preprint arXiv:1609.02200, 2016. and predict the tokens of the masked part.

[0045] Furthermore, Point-MAE[8] Pang Y, Wang W, Tay F E H, et al. Masked autoencoders for point cloud self-supervised learning[C] / / Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part II. Cham: Springer Nature Switzerland, 2022: 604-621. directly masks and recovers point cloud patches to pre-train the encoder.

[0046] Point-M2AE[9] Zhang R, Guo Z, Gao P, et al. Point-M2AE: multi-scale masked autoencoders for hierarchical point cloud pre-training[J]. arXiv preprint arXiv:2205.14401, 2022. constructs a hierarchical network structure that can gradually model geometric information and capture feature information, achieving good results in downstream tasks.

[0047] Joint-MAE

[10] Guo Z, Li X, Heng P A. Joint-mae: 2d-3d joint masked autoencoders for 3d point cloud pre-training[J]. arXiv preprint arXiv:2302.14007, 2023. Focus on the correlation between 2D images and 3D point clouds, and propose a hierarchical module for cross-modal interaction to reconstruct the masked information of both modalities.

[0048] However, Point-M2AE and Joint-MAE learn multi-level geometric knowledge by modifying the hierarchical network architecture, which greatly increases the network computation.

[0049] Therefore, the present invention analyzes the above problems and proposes a solution. Without changing the network architecture and increasing the computation, by improving the training method and utilizing the gradually guiding characteristic of the diffusion model, the network learns multi-level geometric knowledge by restoring the noisy point clouds at different stages, so as to obtain better performance in various downstream tasks.

[0050] The following further illustrates the specific implementation of the embodiments of the present invention with reference to the accompanying drawings of the embodiments of the present invention.

[0051] Referring to FIG. 1, it is a flowchart of the steps of the 3D pre-training method based on the diffusion model according to the embodiment of the present invention. Figure 2 It is an overall framework diagram of the 3D pre-training method based on the diffusion model according to the embodiment of the present invention.

[0052] According to Figure 1 and Figure 2 , the embodiments of the present invention mainly include the following steps:

[0053] Step S110, encode the visible point cloud token features through an encoder to obtain deep visible point cloud token features, where the visible point cloud token features are generated from the obtained point cloud blocks.

[0054] It should be understood that the encoder is a module or network structure that converts input data into a latent representation or feature vector, responsible for extracting latent geometric features, and is composed of standard Transformer blocks.

[0055] It should also be understood that encoding the visible point cloud token features through the encoder to obtain deep visible point cloud token features enables subsequent processing steps to better understand and utilize these features.

[0056] Step S120: Concatenate the deep visible point cloud tokenized features and the occluded point cloud tokenized features to obtain concatenated features, where the occluded point cloud tokenized features are generated from the acquired point cloud patches.

[0057] It should be understood that the purpose of concatenating the deep visible point cloud tokenized features and the occluded point cloud tokenized features to obtain the concatenated features is to combine these two types of features to obtain a more rich and diverse feature representation, providing more comprehensive information for subsequent processing.

[0058] Step S130: Encode the concatenated features through a conditional embedding module to obtain the guidance conditions for the diffusion model.

[0059] It should be understood that the conditional embedding module takes the concatenated features as input and encodes them through a series of operations and network layers. The purpose of conditional embedding is to extract the key information in the concatenated features and generate the guidance conditions for the diffusion model, which can help the decoder perform decoding operations better.

[0060] Step S140: Determine the loss function according to the guidance conditions of the diffusion model and the multi-stage uniform sampling strategy. The loss function is used to guide the decoder to perform decoding to complete the training.

[0061] In summary, the solution of the embodiment of the present invention, without changing the network architecture and without increasing the computational amount, through improving the pre-training method and utilizing the multi-level guidance characteristics of the diffusion model, realizes a comprehensive understanding of the point cloud, obtains richer geometric information, improves the effect and performance of model training, and further saves the cost of point cloud pre-training. In addition, the 3D pre-training method of the embodiment of the present invention has strong flexibility and can be extended to various network models to improve their performance.

[0062] In one implementation, before encoding the visible point cloud tokenized features through an encoder to obtain the deep visible point cloud tokenized features, it includes obtaining point cloud patches through the farthest point sampling algorithm and the K-nearest neighbor algorithm; centering the point cloud patches to update the point cloud patches; embedding the updated point cloud patches into the tokenized features to obtain the point cloud tokenized features; randomly occluding the point cloud tokenized features to generate the visible point cloud tokenized features and the occluded point cloud tokenized features, where the visible point cloud tokenized features are represented as The occluded point cloud tokenized features are represented as is the number of occluded point cloud tokenized features the number of visible point cloud tokenized features, and m is the random occlusion ratio. g = s - r is

[0063] It should be understood that generating visible point cloud tokenization features and occluded point cloud tokenization features can help the model learn features with different importance. For visible point cloud tokenization features, the model can focus on learning key and obvious features. While occluded point cloud tokenization features can help the model learn the ability to infer and fill in occlusions, improving the understanding ability and modeling ability of the complete point cloud data.

[0064] By randomly occluding some point cloud tokenization features, the diversity of the dataset can be increased, thereby improving the generalization ability of the model. Occluding different point cloud tokenization features can simulate occlusion situations in real scenarios, enabling the model to better adapt to processing under various occlusion conditions.

[0065] Through centering processing, the position of the point cloud block can be normalized to the origin or a specified position, eliminating the influence of position on subsequent processing, and improving the consistency, stability, and description ability of the data; by embedding the updated point cloud block into the tokenization feature, the feature representation can be enriched, and a better feature basis can be provided for subsequent processing tasks.

[0066] In another implementation, point cloud blocks are obtained through the farthest point sampling algorithm and the K-nearest neighbor algorithm, including sampling the preset point cloud data through the farthest point sampling algorithm to obtain s center points. Among them, the preset point cloud data contains n points, and the preset point cloud data is represented as X ∈ R n×3 , and the s center points are represented as Based on each center point C of the s center points i , k points closest to each center point are selected as point cloud blocks through the K-nearest neighbor algorithm. Among them, each center point is represented as C i , and the point cloud block is represented as Then the point cloud block The calculation formula is:

[0067]

[0068] In another implementation, the updated point cloud block is embedded into the tokenization feature to obtain the point cloud tokenization feature, including using the simplified version of PointNetξ φ (·) to embed the updated point cloud block into the tokenization feature , which is expressed as:

[0069]

[0070] In another implementation, the visible point cloud tokenization feature is encoded through an encoder to obtain the deep visible point cloud tokenization feature, including inputting the position information of the visible point cloud tokenization feature into the position embedding formula ψ with parameter τ τ(·), output the position feature of the visible point cloud tokenization feature, and the position feature of the visible point cloud tokenization feature is represented as The output position feature of the visible point cloud tokenization feature is connected to the visible point cloud tokenization feature to obtain a combined feature; through the encoder Φ ρ (·)'s Transformer block extracts features from the combined feature to obtain a deep visible point cloud tokenization feature, and the deep visible point cloud tokenization feature is represented as Among them, the encoder Φ ρ (·) consists of standard Transformer blocks with parameters ρ. Then, the formula for encoding the visible point cloud tokenization feature through the encoder to obtain the deep visible point cloud tokenization feature is:

[0071]

[0072] In another implementation, the concatenated feature is encoded through a conditional embedding module to obtain the guiding condition of the diffusion model, including encoding the concatenated feature through the conditional embedding module f ω (·) to obtain the guiding condition c of the diffusion model. Among them, the formula for the guiding condition c of the diffusion model is:

[0073]

[0074] In another implementation, the loss function is determined according to the guiding condition of the diffusion model and the multi-stage uniform sampling strategy, including defining the training objective of the pre-trained model according to the guiding condition c of the diffusion model, and the formula is:

[0075]

[0076] The time step interval [1, T] for each preset point cloud data is divided into h stages:

[0077] Among them

[0078] Randomly sample a time step t from the h stages, calculate the loss h times and perform an average operation on the h losses to obtain the loss function:

[0079]

[0080] It should be understood that in the fine-tuning stage, the encoder is separated from the diffusion model, and a specific decoder and task head are introduced for training to adapt to various different 3D downstream tasks.

[0081] It should also be understood that in the pre-training process of the present invention, the original point cloud is restored, and pre-training can also be carried out by restoring the masked point cloud patches. During the pre-training process, the denoising network of the diffusion model can be replaced by a transformer network structure, and the pre-training network Encoder can be replaced by any 2D / 3D feature network.

[0082] The novel pre-training model proposed by the present invention does not require a hierarchical network structure. By improving the 3D pre-training method and utilizing the multi-level guidance characteristics of the diffusion model, it can achieve a comprehensive understanding of the original input, enabling the network to obtain richer geometric information, thus performing better in downstream tasks and effectively saving the cost of point cloud pre-training. In addition, the pre-training paradigm proposed by the invention can be flexibly extended to various network models and improve their performance.

[0083] Figure 3A This is a visualization result diagram of the 3D pre-training method based on the diffusion model according to an embodiment of the present invention. Each row respectively shows the original input point cloud, i.e., the preset point cloud data, the masked point cloud, and the generated point clouds at different stages during the inverse diffusion process. According to Figure 3A It can be seen that even when 80% of the points are masked, the solution of the present invention can still restore high-quality point clouds, proving the effectiveness of the solution of the present invention.

[0084] Figure 3B This is a comparison diagram of semantic segmentation results according to another embodiment of the present invention, showing a qualitative comparison of semantic segmentation results on the S3DIS dataset. The first column is the original input point cloud, and the second to fourth columns are the detection results of PointNeXt, Point-MAE, and the solution of the present invention respectively. The last column is the ground truth label. The segmentation result of the solution of the present invention is closer to the ground truth label, has a better segmentation effect on the blackboard and the table, has fewer misclassifications compared to PointNeXt and Point-MAE trained from scratch, learns more hierarchical geometric knowledge, and can perform better in downstream tasks.

[0085] See Figure 4 For Figure 1 This is a structural block diagram of a 3D pre-training system 400 based on the diffusion model corresponding to the embodiment. The 3D pre-training system 400 based on the diffusion model includes:

[0086] A first encoding module 410, configured to encode the visible point cloud token features through an encoder to obtain deep visible point cloud token features, where the visible point cloud token features are generated from the obtained point cloud patches.

[0087] A splicing module 420, configured to splice the deep visible point cloud token features and the masked point cloud token features to obtain splicing features, where the masked point cloud token features are generated from the obtained point cloud patches.

[0088] A second encoding module 430, configured to encode the spliced features through a conditional embedding module to obtain guiding conditions for the diffusion model.

[0089] A determination module 440, configured to determine a loss function according to the guiding conditions of the diffusion model and a multi-stage uniform sampling strategy, where the loss function is used to guide the decoder to perform decoding to complete training.

[0090] In summary, the solution of the embodiment of the present invention, without changing the network architecture and without increasing the computational amount, realizes a comprehensive understanding of point clouds, obtains richer geometric information, improves the training effect and performance of the model, and further saves the cost of point cloud pre-training by improving the pre-training method and utilizing the multi-level guiding characteristics of the diffusion model.

[0091] The system of this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein. In addition, the function implementation of each module in the system of this embodiment can refer to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated herein either.

[0092] According to a third aspect of the embodiments of the present invention, an electronic device is provided. Refer to Figure 5 , and now the structural block diagram of an electronic device 500 that can be used as a server or a client of the present application will be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0093] The electronic device 500 may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.

[0094] The processor 502, the communication interface 504, and the memory 506 communicate with each other through the communication bus 508. The communication interface 404 is used to communicate with other electronic devices or servers.

[0095] The processor 502 is configured to execute a program 510, and specifically may execute the relevant steps in the foregoing method embodiments.

[0096] Specifically, the program 510 may include program code that includes computer operation instructions.

[0097] The processor 502 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0098] The memory 506 is used to store the program 510. The memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0099] Specifically, the program 510 can be used to cause the processor 502 to perform the following operations: encoding the visible point cloud tokenization features through an encoder to obtain deep visible point cloud tokenization features; splicing the deep visible point cloud tokenization features with the occluded point cloud tokenization features to obtain spliced features; encoding the spliced features through a conditional embedding module to obtain the guidance conditions of the diffusion model; determining the loss function according to the guidance conditions of the diffusion model and the multi-stage uniform sampling strategy.

[0100] In addition, for the specific implementation of each step in the program 510, reference can be made to the corresponding steps and descriptions in the corresponding units in the above method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.

[0101] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0102] An exemplary embodiment of the present invention also provides a computer storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the methods of the embodiments of the present invention.

[0103] The method according to an embodiment of the present invention can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and to be downloaded through a network and stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or an FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0104] It should be understood that although this specification is described according to various embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0105] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application, and the patent protection scope of the embodiments of the present application shall be defined by the claims.

Claims

1. A 3D pre-training method based on a diffusion model, characterized in that, Including: Encoding the visible point cloud tokenization features through an encoder to obtain deep visible point cloud tokenization features, where the visible point cloud tokenization features are generated from the acquired point cloud blocks; Concatenating the deep visible point cloud tokenization features with the occluded point cloud tokenization features to obtain concatenated features, where the occluded point cloud tokenization features are generated from the acquired point cloud blocks; Encoding the concatenated features through a conditional embedding module to obtain the guidance conditions for the diffusion model; Determining a loss function according to the guidance conditions of the diffusion model and a multi-stage uniform sampling strategy, including: Defining the training objective of the pre-trained model according to the guidance condition c of the diffusion model, with the formula: Dividing the time step interval [1, T] for each preset point cloud data into h stages: wherein Randomly sampling a time step t from the h stages, calculating the loss h times and performing an average operation on the h losses to obtain the loss function: Q i = [d×i + 1, d×(i + 1)]; Wherein, the loss function is used to guide the decoder to perform decoding to complete the training.

2. The method according to claim 1, characterized in that, Before encoding the visible point cloud tokenization features through the encoder to obtain deep visible point cloud tokenization features, it includes: Obtaining point cloud blocks through the farthest point sampling algorithm and the K-nearest neighbor algorithm; Centering the point cloud blocks to update the point cloud blocks; Embedding the updated point cloud blocks into the tokenization features to obtain point cloud tokenization features; Randomly mask the point cloud tokenization features to generate the visible point cloud tokenization features and the masked point cloud tokenization features, where the visible point cloud tokenization features are represented as The masked point cloud tokenization features are represented as For the number of the masked point cloud tokenization features , g = s - r is the number of the visible point cloud tokenization features , and m is the random masking ratio.

3. The method according to claim 2, characterized in that, The obtaining of point cloud blocks through the farthest point sampling algorithm and the K-nearest neighbor algorithm includes: Sampling the preset point cloud data through the farthest point sampling algorithm to obtain s center points, where the preset point cloud data contains n points, and the preset point cloud data is represented as X∈R n×3 , and the s center points are represented as For each center point C among s center points i , k points closest to each center point are selected as the point cloud blocks through the K-nearest neighbor algorithm, where each center point is denoted as C i , the point cloud blocks are denoted as Then the point cloud block has the following calculation formula:

4. The method according to claim 3, characterized in that, The embedding of the updated point cloud blocks into the tokenization features to obtain point cloud tokenization features includes: Through a simplified version of PointNet ξ with parameter φ φ (·) will update the point cloud block and embed it into the tokenized features which is expressed as:

5. The method according to claim 4, characterized in that, The encoding of the visible point cloud tokenization features through the encoder to obtain deep visible point cloud tokenization features includes: Input the position information of the visible point cloud tokenization feature into the position embedding formula ψ with parameter τ τ (·), and output the position feature of the visible point cloud tokenization feature, which is represented as The position feature of the output visible point cloud word segmentation feature is connected to the visible point cloud word segmentation feature to obtain a combined feature; Through the encoder Φ ρ The Transformer block of (·) extracts features from the combined features to obtain the deep visible point cloud tokenized features, which are represented as where the encoder Φ ρ is composed of standard Transformer blocks with parameters ρ. Then, the formula for encoding the visible point cloud tokenized features through the encoder to obtain the deep visible point cloud tokenized features is:

6. The method according to claim 5, characterized in that, The encoding of the concatenated features through the conditional embedding module to obtain the guidance conditions for the diffusion model includes: Condition embedding module \(f\) with parameter \(\omega\) ω (·) encodes the concatenated features to obtain the guidance condition \(c\) of the diffusion model, where the formula for the guidance condition \(c\) of the diffusion model is:

7. A 3D pre-training system based on a diffusion model, characterized in that, Including: A first encoding module for encoding the visible point cloud tokenization features through an encoder to obtain deep visible point cloud tokenization features, where the visible point cloud tokenization features are generated from the acquired point cloud blocks; A concatenation module for concatenating the deep visible point cloud tokenization features with the occluded point cloud tokenization features to obtain concatenated features, where the occluded point cloud tokenization features are generated from the acquired point cloud blocks; A second encoding module for encoding the concatenated features through a conditional embedding module to obtain the guidance conditions for the diffusion model; A determination module for determining a loss function according to the guidance conditions of the diffusion model and a multi-stage uniform sampling strategy, including: Defining the training objective of the pre-trained model according to the guidance condition c of the diffusion model, with the formula: Dividing the time step interval [1, T] for each preset point cloud data into h stages: wherein Randomly sampling a time step t from the h stages, calculating the loss h times and performing an average operation on the h losses to obtain the loss function: Q i = [d×i + 1, d×(i + 1)]; Wherein, the loss function is used to guide the decoder to perform decoding to complete the training.

8. An electronic device, characterized in that, Including: A processor; A memory storing a program; Wherein, the program includes instructions that, when executed by the processor, cause the processor to execute the steps of the method according to any one of claims 1-6.

9. A computer storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, the method described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Point cloud pre-training method based on shielding modeling

    CN116109763A

  • Three-dimensional reconstruction method and device based on binocular vision and diffusion model

    CN116503553A