A 3D Point Cloud Reconstruction Method Based on Self-Training Conditional Diffusion Model

Through the three-dimensional point cloud reconstruction method of self-training conditional diffusion model, self-training is used to self-train, which solves the dependence problem of existing technology on large-scale annotation data, and realizes high-quality three-dimensional point cloud reconstruction in small sample scenarios, improving the robustness and applicability of the model.

CN118691742BActive Publication Date: 2025-07-11HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410736167.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-07-11
Estimated Expiration
2044-06-07

AI Technical Summary

Technical Problem

The existing three-dimensional point cloud reconstruction method relies on large-scale annotation data, which makes it difficult to obtain annotation data in actual applications and affects the performance of model reconstruction.

Method used

A three-dimensional point cloud reconstruction method based on self-training conditional diffusion model is adopted. By introducing a conditional aggregation module, feature consistency loss and shape nature module, combined with reconstruction consistency loss, self-training is used to reduce the dependence on large-scale annotation data, and improve the model's ability to utilize image information.

Benefits of technology

在小样本场景下保持良好的重建能力,提升三维点云重建的准确度和细节表现力,降低对大规模标注数据的依赖,增强模型的鲁棒性和适用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118691742B_ABST
    Figure CN118691742B_ABST
Patent Text Reader

Abstract

The present invention relates to a three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model, comprising: obtaining an image to be reconstructed and Gaussian noise; constructing a three-dimensional point cloud reconstruction network model, introducing a conditional aggregation module and a feature consistency loss, obtaining a three-dimensional point cloud reconstruction network model with conditional diffusion, using the three-dimensional point cloud reconstruction network model with conditional diffusion as a teacher sub-model and a student sub-module, combining a shape naturalness module and a reconstruction consistency loss, obtaining a three-dimensional point cloud reconstruction network model with self-training conditional diffusion; inputting the image to be reconstructed and the Gaussian noise into the three-dimensional point cloud reconstruction network model with self-training conditional diffusion to obtain a point cloud reconstruction result. The present invention can effectively utilize image information, improve the performance of three-dimensional point cloud reconstruction, and at the same time reduce the dependence on large-scale labeled data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly to a three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model. Background Art

[0002] In the context of the current digital age, the progress of technology has given rise to an urgent need for a more accurate and comprehensive understanding of the world. Using images for three-dimensional point cloud reconstruction has become a key technology connecting the real physical world and the digital world. The reconstructed three-dimensional point cloud can provide a comprehensive view including shape, size, position, and mutual spatial relationships, making it possible to simulate and analyze complex scenes, and has broad application value in fields such as virtual reality, autonomous driving, and cultural heritage protection. Therefore, studying three-dimensional point cloud reconstruction has important theoretical value and practical significance.

[0003] Scholars at home and abroad have conducted in-depth research on three-dimensional point cloud reconstruction and achieved important results. Among the existing literature, the most famous and effective three-dimensional point cloud reconstruction methods mainly include: 1. Point set generation network for reconstructing 3D objects from a single image: A point cloud generation network based on a variational autoencoder is proposed. The network consists of two encoder branches, a predictor branch, and a decoder branch. The encoder part extracts global and local features from the input image, the decoder decodes the global features to generate the global shape of the object, the predictor uses the local features to generate the detailed information of the object, and finally the global shape and detailed information are combined as the reconstruction result. 2. Latent embedding matching for accurately and diversely reconstructing 3D point clouds from a single image: It is proposed to map the image and the point cloud into the latent feature space through the encoder for matching, and then the decoder reconstructs various possible point clouds according to the image features. 3. Attention-based dense point cloud reconstruction from a single image: An attention-based method is proposed to reconstruct a dense point cloud from the image, and the attention mechanism enables the encoder part to learn specific details of the shape structure. 4. Using a deep pyramid network for dense 3D point cloud reconstruction: It is proposed to adopt a deep pyramid network to improve the resolution of the reconstructed point cloud by integrating local and global features of the point cloud. 5. Cascade generation network for dense point cloud reconstruction from a single image: A dense point cloud generation method called CGNet is proposed, which realizes staged point cloud reconstruction through a pre-reconstruction network and an upsampling network, and while effectively reducing the spatial complexity, reconstructs a point cloud with a clearer geometric contour. 6. Single-image three-dimensional point cloud reconstruction based on visual enhancement: A point cloud reconstruction network is proposed to reconstruct a finer point cloud from the image by paying more attention to edge and corner information. 7. Single-view three-dimensional point cloud reconstruction based on a single encoder and multiple decoder deep networks: It is proposed to use a single encoder and multiple decoder deep network architecture for three-dimensional point cloud reconstruction, where each decoder reconstructs certain fixed viewpoints. Then all viewpoints are fused to reconstruct a dense point cloud. Figure 3 dimensional point cloud reconstruction: It is proposed to use a single encoder and multiple decoder deep network architecture for three-dimensional point cloud reconstruction, where each decoder reconstructs certain fixed viewpoints. Then all viewpoints are fused to reconstruct a dense point cloud. Summary of the Invention

[0004] The object of the present invention is to propose a three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model to solve the problems existing in the above-mentioned prior art, which can effectively utilize image information, improve the performance of three-dimensional point cloud reconstruction, and at the same time reduce the dependence on large-scale labeled data.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model, comprising:

[0007] Obtain an image to be reconstructed and Gaussian noise;

[0008] Construct a three-dimensional point cloud reconstruction network model, introduce a conditional aggregation module and a feature consistency loss, obtain a conditional diffusion three-dimensional point cloud reconstruction network model, use the conditional diffusion three-dimensional point cloud reconstruction network model as a teacher sub-model and a student sub-module, and combine a shape naturalness module and a reconstruction consistency loss to obtain a self-training conditional diffusion three-dimensional point cloud reconstruction network model;

[0009] Input the image to be reconstructed and the Gaussian noise into the self-training conditional diffusion three-dimensional point cloud reconstruction network model to obtain a point cloud reconstruction result.

[0010] Optionally, the feature consistency loss is:

[0011]

[0012] Wherein, is the feature consistency loss, C is the number of channels of the feature map, and are the values of the feature maps Z o and Z a at the i-th element position and the c-th channel respectively, and M is the total number of elements on the feature map.

[0013] Optionally, the self-training conditional diffusion three-dimensional point cloud reconstruction network model includes:

[0014] A teacher sub-model, which is used to obtain unlabeled image data, perform three-dimensional point cloud reconstruction on the unlabeled image data, and obtain a reconstructed point cloud;

[0015] A shape naturalness module, which is used to explore the internal connection between the reconstructed point cloud and the unlabeled image data, obtain a pseudo-labeled point cloud with a natural structure of the point cloud and consistent with the unlabeled image data to expand the three-dimensional point cloud reconstruction training set, and train the student sub-module in combination with a denoising loss function and the reconstruction consistency loss;

[0016] A student sub-module, which is used to dynamically update the parameters of the teacher sub-model, and obtain the point cloud reconstruction result according to the image to be reconstructed and the Gaussian noise.

[0017] Optionally, the shape naturalness module includes:

[0018] An image feature extraction sub-module, which is used to extract features from the unlabeled image data to obtain first image features;

[0019] A point cloud feature extraction sub-module, which is used to extract the geometric structure information of the reconstructed point cloud to obtain point cloud features;

[0020] A multi-scale convolution sub-module, which is used to splice the first image features and the point cloud features to obtain fused features;

[0021] A self-attention sub-module, which is used to recalibrate the fused features to obtain recalibrated features, comprehensively analyze the recalibrated features to obtain a point cloud quality assessment result, judge whether the structure of the point cloud is natural and whether it is consistent with the unlabeled image data according to the point cloud quality assessment result, obtain a pseudo-labeled point cloud with a natural structure of the point cloud and consistent with the unlabeled image data to expand the three-dimensional point cloud reconstruction training set, and train the student sub-module in combination with a denoising loss function and the reconstruction consistency loss function.

[0022] Optionally, the image feature extraction sub-module includes:

[0023] A two-dimensional convolution unit, which is used to perform preliminary convolution operations on the unlabeled image data;

[0024] A normalization unit, which is used to normalize the convolution operation result to obtain shallow features;

[0025] A number of residual units, which are used to perform deep processing on the shallow features to obtain deep feature maps;

[0026] A max-pooling unit, which is used to reduce the dimension of the deep feature map to obtain first image features.

[0027] Optionally, the point cloud feature extraction sub-module includes:

[0028] A number of SA units, which are used to downsample the reconstructed point cloud to obtain key point information, aggregate the point information around the key points in the key point information to obtain a local point set, extract features for each point in the local point set, and perform max-pooling on the features of all points in the feature dimension to obtain hierarchical features;

[0029] An MLP unit, which is used to process the hierarchical features to obtain point cloud features.

[0030] Optionally, the multi-scale convolution sub-module includes:

[0031] A reshaping unit for reshaping the size of the point cloud feature into a first image feature size;

[0032] A splicing unit for splicing and cascading the reshaped point cloud feature and the first image feature;

[0033] Several groups of convolution units for performing convolution processing on the spliced and cascaded feature map, splicing and cascading multi-scale information, and obtaining a fused feature.

[0034] Optionally, the method for obtaining the denoising loss function is:

[0035]

[0036] Wherein, is the denoising loss function, ∈ θ (x t , t, Z Il ) is the predicted noise, ∈ is Gaussian noise, x t is the noisy point cloud, t is the time step, is the second image feature;

[0037] The method for obtaining the reconstruction consistency loss is:

[0038]

[0039] Wherein, is the reconstruction consistency loss function, P r is the reconstructed point cloud of the enhanced image reconstructed by the student network, P p is the pseudo-label point cloud corresponding to the unlabeled image data, d(·) is the similarity metric function, is the weight function related to the time step t.

[0040] Optionally, the student sub-model includes:

[0041] A feature extraction model for extracting features from the image to be reconstructed to obtain a second image feature;

[0042] A first reverse denoising module for reverse denoising the Gaussian noise to obtain a noisy point cloud;

[0043] A conditional aggregation module for splicing the image feature and the noisy point cloud to obtain an input point cloud;

[0044] A noise estimation module for predicting the input point cloud to obtain a predicted noise;

[0045] The second reverse denoising module is used to perform reverse denoising on the predicted noise to obtain the point cloud reconstruction result.

[0046] Optionally, obtaining the input point cloud includes:

[0047] Perform bilinear interpolation calculation on the D dimensions of the second image feature to obtain an interpolated feature map. Combine the internal and external parameter matrices of the camera parameters to convert the coordinates of any point in the noisy point cloud into the pixel coordinates of the interpolated feature map, obtain the D-dimensional vector corresponding to the pixel coordinates of the interpolated feature map, splice the D-dimensional vector and the coordinates of the noisy point cloud, and obtain the input point cloud.

[0048] The beneficial effects of the present invention are:

[0049] Compared with the prior art, the advantages of the present invention are as follows: Traditional three-dimensional point cloud reconstruction methods usually rely on large-scale labeled data for training. However, in actual application scenarios, it is difficult to obtain labeled data. To reduce the dependence of the model reconstruction performance on large-scale labeled data and enable the model to maintain good reconstruction ability in small-sample scenarios, the present invention proposes a three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model; to improve the model's understanding of image data and thus improve the accuracy of the model's reconstruction of three-dimensional point clouds, the present invention adds a feature consistency loss to the image feature module to enable the model to fully capture the essential information in the image; to improve the detail expressiveness of the three-dimensional point cloud reconstruction result, the present invention designs a conditional aggregation method to aggregate image information and point clouds through projection, enabling the denoising diffusion probability model to focus on the reconstruction process of each point and enhancing the reconstruction accuracy of the model; the present invention proposes a shape naturalness module to enhance the pseudo-labeled point clouds generated during the self-training process, retain the pseudo-labeled point clouds with higher quality, effectively augment the training data, and improve the model's utilization ability for unlabeled images; to further enhance the model's feature learning ability for pseudo-labeled point clouds, the present invention proposes a reconstruction consistency loss, adopting the idea of consistency regularization to improve the robustness and reconstruction performance of the model.

[0050] The three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model proposed by the present invention has good three-dimensional point cloud reconstruction performance, can effectively utilize image information to reconstruct high-quality three-dimensional point clouds. At the same time, it can effectively reduce the model's dependence on large-scale labeled data and maintain good reconstruction ability in small-sample data scenarios, having good effectiveness and applicability. Brief Description of the Drawings

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0052] Figure 1 Flowchart of the 3D point cloud reconstruction method based on the self-training conditional diffusion model for the embodiments of the present invention;

[0053] Figure 2 Schematic structural diagram of the 3D point cloud reconstruction network model based on the diffusion model for the embodiments of the present invention;

[0054] Figure 3 Schematic detailed structural diagram of the image feature extraction module in the 3D point cloud reconstruction network model based on the diffusion model for the embodiments of the present invention;

[0055] Figure 4 Schematic detailed structural diagram of the noise estimation module in the 3D point cloud reconstruction network model based on the diffusion model for the embodiments of the present invention;

[0056] Figure 5 Schematic diagram of the point body convolution process of the noise estimation module in the 3D point cloud reconstruction network model based on the diffusion model for the embodiments of the present invention;

[0057] Figure 6 Schematic structural diagram of the conditional diffusion 3D point cloud reconstruction network model for the embodiments of the present invention;

[0058] Figure 7 Schematic diagram of the bilinear interpolation calculation used for conditional aggregation in the conditional diffusion 3D point cloud reconstruction network model for the embodiments of the present invention;

[0059] Figure 8 Comparison chart of the visualization results of the reconstruction results between the conditional diffusion 3D point cloud reconstruction network model for the embodiments of the present invention and other advanced 3D point cloud reconstruction models;

[0060] Figure 9 Schematic structural diagram of the self-training conditional diffusion 3D point cloud reconstruction network model for the embodiments of the present invention;

[0061] Figure 10 Schematic structural diagram of the shape naturalness module in the self-training conditional diffusion 3D point cloud reconstruction network model for the embodiments of the present invention;

[0062] Figure 11 Schematic structural diagram of the image feature extraction part used in the shape naturalness module of the self-training conditional diffusion 3D point cloud reconstruction network model for the embodiments of the present invention;

[0063] Figure 12 Schematic diagram of the detailed structure of the residual block used in the image feature extraction part of the shape naturalness module in the self-training conditional diffusion three-dimensional point cloud reconstruction network model of the embodiment of the present invention;

[0064] Figure 13 Schematic diagram of the structure of the point cloud feature extraction part used in the shape naturalness module in the self-training conditional diffusion three-dimensional point cloud reconstruction network model of the embodiment of the present invention;

[0065] Figure 14 Schematic diagram of the structure of the multi-scale convolution part used in the shape naturalness module in the self-training conditional diffusion three-dimensional point cloud reconstruction network model of the embodiment of the present invention;

[0066] Figure 15 Training process diagram of the student network in the self-training conditional diffusion three-dimensional point cloud reconstruction network model of the embodiment of the present invention after introducing the reconstruction consistency loss;

[0067] Figure 16 Visualization result comparison diagram of the reconstruction results between the self-training conditional diffusion three-dimensional point cloud reconstruction network model of the embodiment of the present invention and other advanced three-dimensional point cloud reconstruction models. Detailed implementation manners

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0070] This embodiment discloses a three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model, including: obtaining an image to be reconstructed and Gaussian noise; constructing a three-dimensional point cloud reconstruction network model, introducing a conditional aggregation module and a feature consistency loss, obtaining a conditional diffusion three-dimensional point cloud reconstruction network model, using the conditional diffusion three-dimensional point cloud reconstruction network model as a teacher sub-model and a student sub-model, and combining a shape naturalness module and a reconstruction consistency loss to obtain a self-training conditional diffusion three-dimensional point cloud reconstruction network model; inputting the image to be reconstructed and the Gaussian noise into the self-training conditional diffusion three-dimensional point cloud reconstruction network model to obtain a point cloud reconstruction result, specifically:

[0071] Construct a 3D point cloud reconstruction network model based on a diffusion model, introduce a conditional aggregation method into the 3D point cloud reconstruction network model based on the diffusion model to construct a conditional diffusion 3D point cloud reconstruction network model; based on the conditional diffusion 3D point cloud reconstruction network model, combine a shape naturalness module and a reconstruction consistency loss to construct a self-training conditional diffusion 3D point cloud reconstruction network model; obtain an image and Gaussian noise, input the image and Gaussian noise into the self-training conditional diffusion 3D point cloud reconstruction network model, and output a 3D point cloud reconstruction result for 3D point cloud reconstruction.

[0072] Based on the conditional diffusion 3D point cloud reconstruction network model, combining a shape naturalness module and a reconstruction consistency loss, constructing the self-training conditional diffusion 3D point cloud reconstruction network model includes: using the conditional diffusion 3D point cloud reconstruction network model as both the teacher network and the student network, simultaneously constructing and pre-training the shape naturalness module, using the teacher network to reconstruct pseudo-labeled point clouds for unlabeled images, using the pre-trained shape naturalness module to explore the internal connections between the image and the point cloud at multiple scales, retaining the pseudo-labeled point clouds with natural shapes, effectively augmenting the training data; introducing the reconstruction consistency loss into the student network to enhance the student network's feature learning ability for the pseudo-labeled point clouds and improve the reconstruction accuracy; on this basis, constructing a self-training conditional diffusion 3D point cloud reconstruction network model.

[0073] Constructing and pre-training the shape naturalness module includes: using image data and point cloud data as the input of the shape naturalness module, respectively extracting features from the image data and point cloud data by an image feature extraction part and a point cloud feature extraction part to obtain image features and point cloud features, then fusing the image features and point cloud features by a multi-scale convolution part to obtain fused features, recalibrating the fused features through a self-attention layer to obtain calibrated features, and processing the calibrated features by two fully connected layers to output whether the shape of the point cloud data is natural.

[0074] Using the cross-entropy loss as the loss function of the shape naturalness module, using labeled images and their point cloud data as samples of natural shapes, training the conditional diffusion 3D point cloud reconstruction network model with a small amount of data and training rounds to obtain a pre-trained model, using it to reconstruct point clouds for some image data, and using the reconstructed point clouds and the images as samples of unnatural shapes to construct a training set to pre-train the shape naturalness module.

[0075] Introducing the reconstruction consistency loss into the student network includes: adding Gaussian noise to an unlabeled image to obtain an enhanced image, and defining the reconstruction consistency loss calculation formula as follows: P r is the 3D point cloud reconstructed by the student network for the enhanced image, P pis the pseudo-label corresponding to the unlabeled image data, and d(·) is a similarity metric function. is a weight function related to time step t.

[0076] The three-dimensional point cloud reconstruction includes: inputting the image and Gaussian noise into the student network in the three-dimensional point cloud reconstruction network model of self-training conditional diffusion, and the trained student network outputs the three-dimensional point cloud reconstruction result.

[0077] The shape naturalness module includes: an image feature extraction sub-module for extracting features from the unlabeled image data to obtain the first image feature; a point cloud feature extraction sub-module for extracting the geometric structure information of the reconstructed point cloud to obtain the point cloud feature; a multi-scale convolution sub-module for splicing the first image feature and the point cloud feature to obtain a fused feature; a self-attention sub-module for recalibrating the fused feature to obtain a recalibrated feature, comprehensively analyzing the recalibrated feature to obtain the point cloud quality evaluation result, judging whether the structure of the point cloud is natural and consistent with the unlabeled image data according to the point cloud quality evaluation result, obtaining a pseudo-labeled point cloud with a natural structure and consistent with the unlabeled image data to expand the three-dimensional point cloud reconstruction training set, and training the student sub-model in combination with the denoising loss function and the reconstruction consistency loss function.

[0078] The image feature extraction sub-module includes: a two-dimensional convolution unit for performing preliminary convolution operations on the unlabeled image data; a normalization unit for normalizing the convolution operation result to obtain a shallow feature; several residual units for performing deep processing on the shallow feature to obtain a deep feature map; a max pooling unit for reducing the dimension of the deep feature map to obtain the first image feature.

[0079] The point cloud feature extraction sub-module includes: several SA units for downsampling the reconstructed point cloud to obtain key point information, aggregating the point information around the key points in the key point information to obtain a local point set, extracting features for each point in the local point set, and performing max pooling on the features of all points in the feature dimension to obtain a hierarchical feature; an MLP unit for processing the hierarchical feature to obtain the point cloud feature.

[0080] The multi-scale convolution sub-module includes: a reshaping unit for reshaping the size of the point cloud feature to the size of the first image feature; a splicing unit for splicing and cascading the reshaped point cloud feature and the first image feature; several groups of convolution units for performing convolution processing on the spliced and cascaded feature map, splicing and cascading multi-scale information to obtain a fused feature.

[0081] The student sub-model includes: a feature extraction model for extracting features from the image to be reconstructed to obtain second image features; a first reverse denoising module for performing reverse denoising on the Gaussian noise to obtain a noise point cloud; a conditional aggregation module for splicing the image features and the noise point cloud to obtain an input point cloud; a noise estimation module for predicting the input point cloud to obtain predicted noise; and a second reverse denoising module for performing reverse denoising on the predicted noise to obtain the point cloud reconstruction result, specifically:

[0082] Building the three-dimensional point cloud reconstruction network model based on the diffusion model includes: building an image feature extraction module and a noise estimation module, building a denoising loss function in the noise estimation module, and building the final three-dimensional point cloud reconstruction network model on this basis.

[0083] Building the image feature extraction module includes: taking the image as input, encoding it through four encoding layers, and then decoding it through five decoding layers to build the image feature extraction module. Among them, the four encoding layers take the image as input and encode it to obtain a feature map as the image features output by the image feature extraction module.

[0084] Building the noise estimation module includes: splicing the image features and the noise point cloud as the input point cloud of the noise estimation module. Each point in the input point cloud can be expressed as: f F (p)=f I (Z)|f p (p), where p is any point in the noise point cloud, f F (p) is the feature of point p after splicing, Z is the image feature, f I (·) converts the image feature into a one-dimensional vector, f p (p) is the coordinate vector of point p, and | represents the splicing operation. Input the input point cloud into the noise estimation module for noise prediction and output the predicted noise.

[0085] Building the denoising loss function includes: adding Gaussian noise to the real point cloud as the noise point cloud, splicing it with the image feature to obtain the input point cloud, inputting the input point cloud into the noise estimation module to obtain the predicted noise, and calculating the mean square error between the Gaussian noise and the predicted noise as the denoising loss function, which can be expressed as: where ∈ is the Gaussian noise, ∈ θ is the predicted noise.

[0086] Introduce a conditional aggregation method into the three-dimensional point cloud reconstruction network model based on the diffusion model. The construction of the three-dimensional point cloud reconstruction network model with conditional diffusion includes: introducing the conditional aggregation method into the noise estimation module to replace the method of splicing the image features and the noisy point cloud to generate the input point cloud, and constructing the three-dimensional point cloud reconstruction network model with conditional diffusion.

[0087] A three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model provided in this embodiment is as Figure 1 shown, and includes:

[0088] S1. Construct a three-dimensional point cloud reconstruction network model based on the diffusion model, and use a single image to complete the three-dimensional point cloud reconstruction. Use the three-dimensional point cloud reconstruction data set as the input of the model. The structure of the three-dimensional point cloud reconstruction network model based on the diffusion model is as Figure 2 shown. The specific method for constructing the three-dimensional point cloud reconstruction network model based on the diffusion model is as follows:

[0089] Construct and pre-train an image feature extraction module using the image data in the three-dimensional point cloud reconstruction data set. The structure of the image feature extraction module is as Figure 3 shown. The image feature extraction module consists of an encoder and a decoder part. The encoder is used to extract the feature map of the input image, which consists of four layers, and each layer contains convolution calculation, BN, and ReLU activation calculation operations. The decoder then decodes the feature map to reconstruct the image, which consists of five layers, and each layer contains deconvolution calculation, BN, and ReLU activation calculation operations. The overall goal of the pre-trained image feature extraction module is to make the original image X and the output image data as consistent as possible. To prevent the reconstruction process from becoming an identity mapping, a constraint is added to the encoding process: an enhanced image is obtained by adding Gaussian noise to the original image as the module input, and the module reconstructs the image data that is consistent with the original image.

[0090] The original image is represented as X, and the image reconstructed by the encoder is represented as Define the reconstruction loss calculation formula as follows:

[0091]

[0092] In the formula, C represents the number of color channels of the image, X i,c and respectively represent the component values of the original image and the reconstructed image at the i-th pixel position and the c-th color channel.

[0093] By inputting the enhanced image into the encoder for encoding, the latent space feature map Z aMeanwhile, the original image X is also input into the encoder to extract features, obtaining the feature map Z o Calculate Z a and Z o The difference is used as the feature consistency loss, which encourages the two to be as close as possible, ensuring that the model does not produce large output changes for small changes in the input, improving the network's learning ability for the essential features of the image, and enhancing the robustness of the network. Denote the feature map obtained by inputting the enhanced image into the encoder as Z a and the feature map obtained by inputting the original image into the encoder as Z o Define the feature consistency loss as follows:

[0094]

[0095] In the formula, C represents the number of channels of the feature map, and represent the values of the feature maps Z o and Z a at the i-th element position and the c-th channel respectively.

[0096] Based on the above design, the target loss function of the module consists of the following two parts, namely the reconstruction loss and the feature consistency loss. The overall loss function is Use the overall loss function to pre-train the image feature extraction module. After the pre-training is completed, the 3D point cloud reconstruction network model based on the diffusion model only uses the encoder part to extract image features and obtains the feature map.

[0097] Furthermore, construct a noise estimation module. The structure of the noise estimation module is as Figure 4 shown. Add Gaussian noise to the real point cloud in the 3D point cloud reconstruction dataset according to the noise addition process of the denoising diffusion probability model to obtain the noisy point cloud. Use the image feature extraction module to extract image features to obtain the feature map as the image feature. Concatenate the image feature and the noisy point cloud as the input point cloud of the noise estimation module. Each point in the input point cloud can be expressed as: f F (p) = f I (Z)|f p (p), where p is any point in the noisy point cloud, f F (p) is the feature of point p after concatenation, Z is the image feature, f I (·) converts the image feature into a one-dimensional vector, f p (p) is the coordinate vector of point p, and | represents the concatenation operation. Input the input point cloud into the noise estimation module, and then perform feature learning through set abstraction, global attention layer, feature propagation layer, and MLP to predict the noise to perform the reverse denoising process of the denoising diffusion probability model.

[0098] InFigure 4 In this process, as the set abstraction layer continuously processes the input, the obtained feature granularity gradually increases, and the information receptive field in the feature map gradually expands. To enable the network to fully exploit the important local and global feature information therein, a global attention layer is added. Through the self-attention mechanism, the global features are integrated to highlight the important information among them. That is, after the last set abstraction layer extracts the point cloud features, the resulting feature map is [N′, D C , where N′ is the number of finally downsampled points, and D C is the feature vector of each downsampled point, with a length of C. Through the self-attention mechanism, a convolution operation with a stride of 1 and a kernel size of 1 is used to convolve D C to obtain query, key, and value vectors with a length of C. The similarity between the query and the key is calculated through dot product operation, and the result is normalized by softmax to obtain a weight vector. The weight vector and the value vector are dot product-operated to obtain the feature vector D′ C of each downsampled point, so as to enhance the feature extraction and representation ability of the network, allow the network to pay more attention to important features, and comprehensively reflect the characteristics of the point cloud data.

[0099] In addition, the noise estimation module uses point-body convolution to extract features from the point cloud data in the set abstraction and feature propagation layers. Point-body convolution uses a shared FC to perform point-level feature extraction on the input data to obtain point features. At the same time, the input is voxelized, and voxel features are extracted through a three-dimensional convolutional layer to obtain a feature map. The self-attention layer and the channel attention layer focus on the important information therein, and interpolation is used to de-voxelize to obtain voxel-level features. Finally, the features extracted by the two methods are fused to effectively capture spatial information by utilizing the regularity of the voxel grid while retaining the flexibility of the point cloud data representation, thereby realizing the fine modeling of complex geometric structures. Figure 5 Show the process diagram of point-body convolution. In Figure 5 it is Denotes element-wise addition operation. The voxelization process of the point cloud will result in information loss. The number of points contained in each voxel is different, and the points in some voxels may be relatively sparse, and their importance at the feature level is also weak. To further enhance the feature representation ability of the 3D convolutional layer, the result after 3D convolution is processed through the self-attention mechanism. Through the self-attention layer, the module focuses on the feature information of voxels containing more points and reduces the attention to the voxel features at sparse point cloud locations. The calculation process is as follows: The feature X extracted by 3D convolution has a shape of [N, C, D, H, W], where N is the batch size, C is the number of channels, and D, H, and W are the depth, height, and width respectively. Apply a 3D convolution operation with a stride of 1 and a convolution kernel of 1 to X to obtain the corresponding Q, K, and V. Calculate the similarity between Q and K through the dot product operation, convert it to a weight using softmax, and operate the result with V to obtain the final output result X'. The result is element-wise added to the original input through a residual connection to ensure that the network neither loses feature information nor improves the attention to key information.

[0100] In the point body convolution operation, a channel attention mechanism is additionally adopted to dynamically adjust the relationship between channels so that it can focus on more important feature channels. First, perform global average pooling on the output feature map X' of the self-attention layer to obtain the global average value Z of each channel. The calculation formula is as follows:

[0101]

[0102] where, is the dimension of the global average value Z.

[0103] On this basis, perform "Squeeze" and "Excitation" operations on the features through two linear layers and an activation function to calculate the weight of each channel, that is, perform dimensionality reduction (Suqeeze) through the first linear transformation to obtain U = W1Z + b1, is the weight of the first linear layer, r is the dimensionality reduction ratio, C / r represents the number of output channels, C is the number of input channels, and b1 is the bias term. is the result after dimensionality reduction. Pass U through an activation function to obtain the activated result V = f(U), where f(·) is the activation function. For the activated result V, use the second linear layer to perform dimensionality increase (Excitation): S = σ(W2V + b2), where is the weight of this layer, b2 is the bias term, and σ is the sigmoid function used to normalize the weight between 0 and 1. is the weight of each channel. Finally, the obtained weights are used to perform weighted calculation on the three-dimensional convolution result X to recalibrate the feature information, that is, X′ = X ⊙ S′, where ⊙ represents element-wise multiplication, and S′ is reshaped from [N, C] to [N, C, 1, 1, 1] to match the dimension of X, thereby ensuring the compatibility of the multiplication operation. X′ is the final output feature, and its dimension is the same as that of X. Through the SE operation, it is ensured that the three-dimensional convolution network can learn the dependencies between different channels, enabling it to more effectively utilize data features during the training process, reducing the dependence on unimportant features, and improving the performance and generalization ability of the model.

[0104] In the way of adding time encoding embeddings, the specific stages in the generation process are guided, so that the three-dimensional point cloud reconstruction network model based on the diffusion model can accurately adjust the noise prediction strategy at each time step. For the sampled time step t, it is initially encoded through sine and cosine functions to obtain the periodicity and continuity of time. The encoding calculation formula is as follows:

[0105]

[0106] In the formula, t is the time step, d is the encoded dimension, i = {0, 1,..., (d - 2) / 2} is the index of the dimension. After encoding, a d-dimensional embedding vector is obtained. To further refine the time encoding and enable the model to better understand and utilize time information, the obtained embedding vector T E is input into two fully connected layers and one activation function to calculate the final encoded information T E ′.

[0107] T E ′ is concatenated with the input of each layer of SA and FP to embed time information, enhancing the model's ability to capture and recover features at various scales from coarse to fine details of the data, and improving the stability of the training process and the quality of the final reconstructed point cloud.

[0108] The mean square error between the Gaussian noise and the predicted noise is calculated as the denoising loss function, which can be expressed as: where ∈ is the Gaussian noise, ∈ θ is the predicted noise, and the three-dimensional point cloud reconstruction network model based on the diffusion model is trained with this denoising loss function.

[0109] S2. Introduce a conditional aggregation method in the three-dimensional point cloud reconstruction network model based on the diffusion model to construct a conditional diffusion three-dimensional point cloud reconstruction network model, enhance the model's ability to utilize image information, and improve the accuracy of the three-dimensional point cloud reconstruction result.

[0110] The conditional diffusion three-dimensional point cloud reconstruction network model is as Figure 6 shown. The specific method for constructing the conditional diffusion three-dimensional point cloud reconstruction network model is as follows:

[0111] Introduce a conditional aggregation method to replace the method of splicing image features and noisy point clouds to generate input point clouds in the 3D point cloud reconstruction network model based on the diffusion model, and construct the conditional diffusion 3D point cloud reconstruction network model.

[0112] The conditional aggregation method is as follows: After inputting the image into the pre-trained image feature extraction module, obtain the feature map Z(H, W, D) corresponding to the image as the image feature, where H, W, and D are the height, width, and depth of the feature map respectively. Since the feature map obtained through the feature extraction network is different in size from the original image. First, use the bilinear interpolation method to expand the image feature so that it is finally the same size as the input image. Figure 7 Show the bilinear interpolation calculation process.

[0113] In Figure 7 , for the eigenvalue f(P) at the unknown point P=(x, y) in the feature map, take the eigenvalues f(Q 11 =(x1, y1), Q 12 =(x1, y2), Q 21 =(x2, y1), Q 22 =(x2, y2) of its adjacent four feature points, and calculate the eigenvalue of point P using the bilinear interpolation method according to the distance weighted information. The calculation formula is as follows: 11 )、f(Q 12 )、f(Q 21 )、f(Q 22 ), where x2, x1, y2, and y1 are the indices of the corresponding feature points on the feature map.

[0114]

[0115] Perform the above interpolation calculation on each of the D dimensions of the feature map to obtain the interpolated feature map Z

[0116] (H I (H o ,W o ,D), where H o 、W o are the original height and width of the input image. Then, using the camera parameters in the 3D point cloud reconstruction dataset, including the camera intrinsics and extrinsics, adopt the idea of coordinate transformation to transform each point in the 3D point cloud into the 2D space, construct a point-level feature mapping relationship, and through cascading operations, add the corresponding image information to each point to guide the subsequent diffusion process of the point in the diffusion model. The overall process is as follows: First, use the projective transformation parameters in the camera intrinsics, that is, the distance f (focal length) from the camera focus to the image plane, and the scaling and origin translation coefficients α, β, c x 、cy , let f x = α·f, f y = β·f, calculate the intrinsic matrix K as follows:

[0117]

[0118] After that, use the rotation matrix R and translation vector t in the camera extrinsic parameters. Assume P c (x c , y c , z c ) is the coordinate in the camera coordinate system, and P w (x, y, z) is its coordinate in the world coordinate system. Through the camera extrinsic parameters, the transformation relationship between Pc and Pw can be obtained as P c = RP w + t. According to this process, the calculation formula for the extrinsic matrix T of the camera is derived as follows:

[0119]

[0120] In the case of obtaining K and T, for any point p = (x, y, z) in the point cloud, from the transformation function z c [u v 1] T = K'T[x y z 1] T calculate the corresponding pixel coordinates (u v), where the calculation formula for K' is as follows:

[0121]

[0122] On this basis, construct the point-level alignment function of the image feature and the point cloud where f(p) is the coordinate feature of p, and Z D (u, v) is the D-dimensional vector corresponding to the (u, v) coordinates in the image feature, represents the concatenation operation. Through the transformation function and the point-level alignment function, each point in the noisy point cloud can accurately obtain its corresponding image feature, and then obtain the overall feature of the point, realizing conditional aggregation. Use the noisy point cloud after aggregating the image information as the input of the noise estimation module to complete the construction of the 3D point cloud reconstruction network model for conditional diffusion.

[0123] To verify the effectiveness of the conditional aggregation method proposed in the present invention, in this embodiment, the Chamfer Distance (CD) and Earth Mover's Distance (EMD) are used as evaluation metrics for quantitative analysis in the ShapeNet validation set. The reconstruction capabilities of the 3D point cloud reconstruction network model PRN based on the diffusion model, the conditional diffusion 3D point cloud reconstruction network model CDPRN, and other advanced 3D point cloud reconstruction models are compared and analyzed. To make the comparison results more intuitive, without special instructions, the CD values and EMD values of the experimental results of all models are enlarged by 100 times.

[0124] Table 1 is a CD comparison table of the reconstruction quality of each model on the ShapeNet dataset. The smaller the CD value, the smaller the gap between the reconstructed point cloud and the original point cloud.

[0125] Table 1

[0126]

[0127] Analyzing Table 1, it can be seen that CDPRN has better reconstruction effects in each category than PSGN and 3D-LMNet. Compared with CGNet and 3D-VENet, CDPRN has better performance in the reconstruction of 10 categories, and only has slightly lower reconstruction effects in the categories of cars, sofas, and mobile phones. In addition, the average CD value of the reconstruction quality of CDPRN is 0.32 higher than that of 3D-VENet. At the same time, it has a relatively large improvement compared with PSGN, 3D-LMNet, and CGNet. At the same time, compared with PRN, the CD values of CDPRN reconstructed in each category are lower, indicating that the proposed conditional aggregation method can effectively utilize image information and the reconstructed 3D point cloud has higher accuracy.

[0128] Table 2 is an EMD comparison table of the reconstruction quality of each model on the ShapeNet dataset. The smaller the EMD value, the smaller the gap between the reconstructed point cloud and the original point cloud.

[0129] Table 2

[0130]

[0131]

[0132] As can be seen from Table 2, for the proposed CDPRN in the bench and gun categories, the EMD value of the reconstruction quality is slightly higher than that of CGNet, and the EMD values of the reconstruction quality in all other categories are lower than those of the other five models. This shows that in terms of reconstruction accuracy, CDPRN is generally better than the other five models. Especially in the reconstruction of categories such as airplanes, cars, chairs, table lamps, and sofas, CDPRN performs relatively better, indicating that CDPRN is more accurate in reconstructing the geometry of these categories. At the same time, the average EMD value of the reconstruction quality of CDPRN is 4.39, which is lower than the lowest value of 4.52 among the other five models, further confirming that CDPRN is superior to the other five models in overall reconstruction performance.

[0133] To further analyze the reconstruction ability of the conditional diffusion-based 3D point cloud reconstruction network model, in this embodiment, visual analysis is performed on the reconstruction results of CPDRN, 3D-LMNet, and PSGN. Figure 8 is a comparison chart of the point cloud visualization results. As can be seen from Figure 8 it, the point cloud reconstructed by CDPRN is more accurate in terms of shape details. Compared with PSGN and 3D-LMNet, it can better capture and reconstruct the geometric features of objects. For example, in the reconstruction of an airplane, CDPRN can better depict the edges of the wings and the shape of the tail fin. For the reconstruction of chairs and sofas, it can more accurately reconstruct the curves of the legs and backs. When reconstructing complex structures such as chair legs, CDPRN can better maintain the original structural characteristics without obvious breaks or deformations in the structure, indicating that the model has strong capabilities in reconstructing slender or detail-rich structures. In terms of noise point suppression, the point cloud reconstructed by CDPRN is relatively clean. In contrast, PSGN and 3D-LMNet have more discrete points in the reconstruction of some objects such as sofas and monitors, indicating that CDPRN can effectively reduce the generation of noise points during the reconstruction process and improve the point cloud quality. In addition, the point cloud reconstructed by CDPRN shows a more uniform density distribution without obvious sparse or dense areas. When reconstructing objects containing small-scale features such as speakers or bench handles, compared with PSGN and 3D-LMNet, CDPRN shows higher reconstruction accuracy, indicating its excellent performance in dealing with small-scale and high-precision features. Based on the above comparison results, CDPRN shows advantages in multiple aspects such as the quality of point cloud reconstruction, detail retention, noise suppression, density consistency, and perception of small-scale features. Compared with PSGN and 3D-LMNet, the reconstruction accuracy and performance have been significantly improved.

[0134] S3. Based on the conditional diffusion-based 3D point cloud reconstruction network model, combined with the shape naturalness module and the reconstruction consistency loss, construct a self-training conditional diffusion-based 3D point cloud reconstruction network model to enhance the model's reconstruction ability under small-sample data and improve the model's applicability.

[0135] The network model structure of three-dimensional point cloud reconstruction with self-training conditional diffusion is as Figure 9 shown. The specific method for constructing the three-dimensional point cloud reconstruction network model with self-training conditional diffusion is as follows:

[0136] First, based on the conditional diffusion three-dimensional point cloud reconstruction network model CDPRN, the teacher and student networks are constructed. Subsequently, the teacher network is pre-trained using the three-dimensional point cloud reconstruction dataset to ensure that it can effectively capture key geometric structures and feature information. Then, the teacher network performs three-dimensional point cloud reconstruction on a large amount of unlabeled image data to generate reconstructed point clouds. On this basis, the pre-trained shape naturalness module mines the internal connection between the reconstructed point clouds and the unlabeled images through multi-scale convolutional feature fusion, and retains the higher-quality reconstructed point clouds as pseudo-label point clouds corresponding to the unlabeled images, further expanding the three-dimensional point cloud reconstruction training set. The student network is trained on the expanded three-dimensional point cloud reconstruction training set to learn a wider data distribution and improve the reconstruction ability and generalization ability. At the same time, by means of parameter update, the parameters of the teacher network are dynamically updated by the parameters of the student network to improve the model stability. Repeat the above reconstruction process of the teacher network and the training process of the student network to achieve the purpose of self-training. In addition, a reconstruction consistency loss is introduced into the student network to enhance the feature learning ability of the student network for pseudo-label point clouds and improve the reconstruction performance of the model.

[0137] The design and pre-training method of the further shape naturalness module are as follows:

[0138] Figure 10 is the structure diagram of the shape naturalness module. In Figure 10 it, the shape naturalness module mainly consists of an image feature extraction part, a point cloud feature extraction part, a multi-scale convolutional part, and a self-attention part.

[0139] The image feature extraction part is used to extract the shape and texture features contained in the image. Figure 11 is the structure diagram of the image feature extraction part. In Figure 11 it, for the input image data I, the image feature extraction part first performs a preliminary convolution operation through a two-dimensional convolutional layer, normalizes the obtained feature map to obtain shallow features, uses 3 residual block structures to perform in-depth processing on the shallow features to obtain deep feature maps, and finally reduces the dimension of the deep feature maps through a max-pooling layer with a stride of 2 to obtain the final image features, denoted as F I , whose shape is: [w I ,h I ,d I , where w I , h I and d I are the width, height, and depth of the feature map respectively. Figure 12It is the structure diagram of the residual block. In Figure 12 for each residual block, the Conv2d therein is a two-dimensional convolutional layer, and a convolutional kernel with a size of 3×3 is used. The convolutional stride is set to 1, and the padding is 1 to ensure that the length and width of the input data do not change before and after convolution. BN is the batch normalization layer, which normalizes each small batch of data so that its output mean is close to 0 and the variance is close to 1 to improve the stability of the training process and accelerate the convergence speed. At the same time, the ReLU activation function is used to increase the non-linear expression ability of the residual block so that it can learn more complex data representations. For the input data x, the calculation process by the residual block can be expressed as:

[0140]

[0141] In the formula, BN(·) is the batch normalization function, C(·) represents the convolution operation, ReLU(·) is the activation function calculation, represents the element-wise addition operation.

[0142] The point cloud feature extraction part is used to extract the geometric structure information of the point cloud data, Figure 13 which is the structure diagram of the point cloud feature extraction part. In Figure 13 the point cloud feature extraction part first extracts the features of the point cloud through three set abstraction (SA) layers. Each SA layer includes three parts: sampling, grouping, and point feature extraction. For the sampling process, the farthest point sampling (FPS) method is used to downsample the input point cloud data to obtain the key point information. By the method of ball query grouping, the point information around the key points is aggregated to obtain the local point set. The features of each point in the local point set are extracted through a multi-layer perceptron (MLP) with shared weights. Finally, the features of all points on each local point set are max-pooled in the feature dimension as the features of the key points corresponding to the local point set. Through the three SAs, the hierarchical features of the point cloud from local to global are extracted to better capture the subtle structural difference information in the point cloud data. The hierarchical features are further processed by an MLP containing three connection layers to obtain the final point cloud features, denoted as F p , whose shape is: [N′, D P , where N′ is the number of points in the point cloud features, and D p is the length of the feature vector of each point in the point cloud features.

[0143] After obtaining the image features and point cloud features, to solve the problem that there are significant differences in the internal structure and scale properties between the image and point cloud data, a multi-scale convolutional feature fusion strategy is designed to deeply explore and strengthen the internal connection between the two modalities, improve the shape naturalness module's understanding ability of the local context relationship between the features of the two modalities, enhance the effect of information fusion, and further enhance the discriminant ability of the shape naturalness module. Figure 14It is a structural diagram of multi-scale convolution feature fusion. In Figure 14 , the multi-scale convolution feature fusion process is as follows. First, reshape the point cloud feature size to the image feature map size, that is, convert [N′, D P to [w I , h I , d P , where w I and h I are the width and height of the feature map of the image feature respectively, and d p is the depth of the reshaped point cloud feature map. Concatenate the adjusted point cloud feature and the image feature to obtain the fused feature map [w I , h I , d F , where d F = d I + d p , d I is the depth of the image feature map, and d p is the depth of the point cloud feature map. Then, perform convolution operations on the fused feature map through three convolutional layers with different kernel sizes to obtain multi-scale feature information. To enable the model to capture different sizes of dependence ranges, use convolutional kernels of small, medium, and large sizes, that is, perform convolution processing on the fused feature map through convolutional kernels of sizes 3, 5, and 9 respectively. Finally, concatenate the multi-scale information to obtain the fused feature, and its calculation process is where X is the original fused feature map, C i (·) represents the convolution operation of three convolutional layers with a kernel size of i, represents the element-wise addition operation. To ensure that the size of the feature map remains unchanged after passing through the convolutional layer, zero-padding is performed on the input during convolution.

[0144] The result of multi-scale feature fusion To further mine the key features therein, a self-attention method is used to recalibrate the fusion result. Multiply with the weight matrices W Q , W K and W V to perform a linear transformation, and then generate the corresponding query key and value Use and to calculate the score, and scale it using the square root of the dimension d of k . The calculation formula is as follows:

[0145]

[0146] On this basis, the softmax function is used to normalize the Score to obtain the attention weights, and the attention weights are weighted and summed with to obtain the final output as the recalibrated features. Since all elements are considered, the important parts in the features are further extracted. Finally, the recalibrated features are comprehensively analyzed through a fully connected layer and an activation function layer, and the point cloud quality evaluation result is output, that is, whether the point cloud data can be used as the three-dimensional rendering of the corresponding image and whether its shape performance is natural.

[0147] The training objective of the shape naturalness module is to judge whether the point cloud is natural relative to the image, that is, whether there is a strong consistency between the two. It outputs the naturalness result through the fully connected layer of the last single neuron and the sigmoid activation function. Indicates that the structure of the point cloud is natural, the quality is good, and it is consistent with the image at the same time, and can be used as the pseudo-point cloud data of the image; Indicates that the quality of the point cloud is average, the consistency relationship with the image is poor, and it cannot be used as the pseudo-point cloud data corresponding to the image. Based on the above training objectives, the training and test data of the shape naturalness module are labeled with the label y ∈ {0, 1} to indicate whether the structure of the point cloud is natural and whether it is consistent with the image representation. For the sample pairs with natural shape and consistent with the image representation, they are marked as y = 1, and for the sample pairs with unnatural shape or inconsistent with the image representation, they are marked as y = 0. Cross-entropy is used as the loss, and the loss function is defined as follows:

[0148]

[0149] In the formula, N is the batch size, y i is the true label of the i-th sample, is the result predicted by the network.

[0150] When training the shape naturalness module, the labeled images and their point cloud data are used as the sample pairs with natural shape. The three-dimensional point cloud reconstruction network model CDPRN of conditional diffusion is trained with less data and fewer training rounds to obtain the training model. It reconstructs the point cloud from part of the image data, and the reconstructed point cloud and the image are used as the sample pairs with unnatural shape. By constructing a dataset of sample pairs with natural shape and unnatural shape, the shape naturalness module is fully trained. To maintain sample balance, the ratio of the sample pairs with natural shape to the sample pairs with unnatural shape in the constructed dataset is 1:1.

[0151] Furthermore, the loss function and training process of the student network are optimized to make it fully utilize the internal information of the unlabeled data. The specific implementation is as follows:

[0152] Using the idea of consistency regularization in semi-supervised learning, a reconstruction consistency loss is introduced. That is, noise perturbation is added to the unlabeled images to obtain enhanced images, and the results output by the model are constrained to be consistent with the pseudo-labels. Further leveraging the role of the pseudo-labeled point cloud data, the feature learning ability of the student network for unlabeled image data is enhanced. The calculation formula for the reconstruction consistency loss is as follows:

[0153]

[0154] In the formula, P r is the reconstructed point cloud of the enhanced image by the student network, and P p is the pseudo-labeled point cloud corresponding to this unlabeled image data. d(·) is the similarity metric function, is the weight function related to the time step t.

[0155] Taking the chamfer distance between P r and P p as the similarity evaluation metric d(·), the chamfer distance CD is the L2 distance calculated point by point between two point clouds. It is differentiable everywhere and the point-by-point calculation is independent. At the same time, compared with EMD, its computational complexity is lower, and it can be used as an efficient loss function to evaluate the overall similarity between two point clouds.

[0156] The denoising loss calculation formula of the original CDPRN is as follows:

[0157]

[0158] The calculation formula for the overall loss function of the improved CDPRN is as follows:

[0159]

[0160] In the formula, ρ is a hyperparameter that controls the influence degree of the reconstruction consistency loss on the network.

[0161] Using the improved overall loss function as the training objective of the student network.

[0162] At the same time, the training process of the student network is adjusted to adapt to the expanded dataset. Figure 15 is the training process diagram of the student network. In Figure 15 , when training with labeled data, the training method of the student network is the same as that of the conditional diffusion three-dimensional point cloud reconstruction network model CDPRN. That is, first, using the image feature extraction module, the input labeled image data I label is used to extract features to obtain the condition Randomly sample the time step t, and then add noise to the input labeled point cloud data to obtain is the labeled point cloud, is the weight coefficient of noise accumulation, and then the noise is predicted by the noise estimation module Finally, calculate the denoising loss Use the gradient Update the network parameters

[0163] When training with unlabeled data, since the reconstruction consistency loss is added as the unsupervised loss, for the input unlabeled image data I unlabel and the pseudo-labeled point cloud The training process of the student network is as follows

[0164] (1) Add noise perturbation to I unlabel by adding noise to obtain the enhanced result I augment ;

[0165] (2) Feed I augment into the feature extraction network to extract features and obtain the feature map

[0166] (3) Sample the time step t

[0167] (4) Add noise to to obtain

[0168] (5) Predict the noise by the noise estimation module Calculate the denoising loss

[0169] (6) The noise estimation module performs the reverse denoising process to reconstruct the point cloud The calculation formula is as follows

[0170]

[0171] (7) Calculate the reconstruction consistency loss, that is, calculate

[0172] (8) Calculate the total loss

[0173] (9) Update the parameters of the student network by the gradient Update the parameters of the student network

[0174] In the above training process of the student network, due to the different sampling time steps t, the noise level of x t is different. When t is close to the total time step T, x t contains a relatively high level of noise, resulting in the reconstructed point cloud also containing more noise. At this time, reduce the impact of the reconstruction consistency loss on the network. When t is close to 1, x t contains a relatively low level of noise. At this time, the reconstructed point cloud It contains less noise, which can increase the impact of the reconstruction consistency loss between it and the pseudo-labeled point cloud on the network. According to the above idea, the function expression is as follows:

[0175]

[0176] where T is the total number of time steps.

[0177] Furthermore, in the way of Exponential Moving Average (EMA), the weight parameters of the teacher network are dynamically updated by the student network parameters. The teacher network can continuously generate more refined pseudo-labeled point cloud data synchronized with the training stage, thereby expanding richer training data for the training set and improving the overall stability and adaptability of the network at the same time. The EMA calculation formula is as follows:

[0178] θ′ s = αθ′ s-1 +(1 - α)θ s

[0179] In the formula, θ′ s is the weight parameter of the teacher network when the training step is s, θ s is the weight parameter of the student network when the training step is s, and α is the decay factor.

[0180] During the training process of the self-training conditional diffusion three-dimensional point cloud reconstruction network model, only by dynamically adjusting the teacher network parameters in the above way can the stability of the model during training be ensured. When the teacher network parameters are updated, it uses the reverse denoising process to perform point cloud reconstruction on the unlabeled image data. After being evaluated by the shape naturalness module, the training set is further expanded. The student network, as the main part of the self-training conditional diffusion three-dimensional point cloud reconstruction network model training, is trained using the expanded training set.

[0181] To verify the performance of the model in the small-sample data scenario, this embodiment divides the training set in the three-dimensional point cloud reconstruction dataset, that is, divides the ShapeNet dataset according to the ratios of 25%, 45%, and 60%, conducts three groups of experiments on the self-training conditional diffusion three-dimensional point cloud reconstruction network model, and conducts a comparative analysis of the performance of other models at the 100% ratio. To facilitate the analysis of the experimental results, this embodiment names the results obtained by training the self-training conditional diffusion three-dimensional point cloud reconstruction network model ST-CDPRN with 25%, 45%, and 60% ratio training sets as ST@25%, ST@45%, and ST@60% respectively. Table 3 is the CD comparison table of the reconstruction quality of each model on the ShapeNet dataset.

[0182] Table 3

[0183]

[0184] Analysis of Table 3 shows that although under the training with 25% proportion of labeled data, the overall CD between the reconstructed point cloud and the real point cloud of ST-CDPRN is higher than that of other comparison models, its reconstruction performance is slightly lower than that of other models. However, when trained with 45% proportion of labeled data, in the reconstruction of some shapes such as monitors and mobile phones, its performance is better than that of PSGN and 3D-LMNet, indicating that when dealing with relatively simple three-dimensional structures, ST-CDPRN can achieve better reconstruction results with less labeled data. Under the training with 60% proportion of labeled data, the average CD of ST-CDPRN is 5.51, and the average CD of PSGN is 5.64, indicating that the overall reconstruction performance of ST-CDPRN is better than that of PSGN. At the same time, by comparing the average CD of 3D-LMNet, CGNet, and 3D-VENet, it can be seen that ST-CDPRN achieves comparable performance with less data training. This shows that ST-CDPRN can effectively utilize unlabeled data to improve the reconstruction quality and effectively reduce the dependence on a large amount of labeled data.

[0185] Table 4 is a comparison table of EMD of the reconstruction quality of each model on the ShapeNet dataset.

[0186] Table 4

[0187]

[0188] Analysis of Table 4 shows that although the reconstruction performance of the model trained by ST-CDPRN is slightly insufficient under the conditions of 25% and 45% proportion of labeled data, under the training with 60% proportion of labeled data, the EMD value between the reconstructed point cloud and the real point cloud is smaller, indicating that the reconstructed point cloud is more similar to the real point cloud, and it is better than the performance of the PSGN and 3D-LMNet models. This shows that the model can effectively utilize unlabeled data to improve performance and maintain good reconstruction performance in small-sample data scenarios.

[0189] Furthermore, in this embodiment, from a qualitative perspective, the visual reconstruction results of PSGN, 3D-LMNet, ST@25%, ST@45%, and ST@60% are compared. Figure 16This is a comparison chart of visualization results. Analyzing the results, it can be seen that ST-CDPRN can achieve good reconstruction results under different proportions of labeled data. When trained with 25% proportion of labeled data, although there are problems such as rough contour reconstruction, uneven point distribution density, or missing details for each object, it can still maintain good overall structural consistency. With the appropriate increase in the proportion of labeled data during training, the reconstruction results of the model are improved, and the accuracy and detail integrity are enhanced. For example, when trained with 45% proportion of labeled data, the reconstruction accuracy is improved, the point distribution is more uniform, and the shape of the object can be captured better; when trained with 60% proportion of labeled data, the wings and tails of the airplane, the legs and backs of the chair, and the leg structures of the bench all have obvious improvements in details and shapes. These visual results prove that ST-CDPRN can make full use of the information of unlabeled data and can also achieve high-quality 3D point cloud reconstruction with limited labeled data training. Although too little data will limit the model's ability to capture details and shapes, as the available data volume increases, ST-CDPRN can capture details more accurately and maintain the geometric consistency of the object, showing its good adaptability and reconstruction ability for objects with different complexity levels. At the same time, when trained with 60% proportion of labeled data, the reconstruction results of ST-CDPRN are comparable to those of PSGN and 3D-LMNet in terms of detail performance. These qualitative results are consistent with the previous quantitative analysis, further verifying the practicality and effectiveness of ST-CDPRN in 3D point cloud reconstruction in the scenario of a small amount of labeled data.

[0190] In addition, to verify the effectiveness of the shape naturalness module, reconstruction consistency loss, and self-training method proposed in the present invention, the following ablation experiments are carried out on 60% of the data in the training set of the 3D reconstruction dataset in this embodiment:

[0191] (1) Ablate the shape naturalness module. To ensure the normal operation of the model, all the point cloud data reconstructed by the teacher network are directly expanded into the dataset, and the impact on the model performance is analyzed. The model after ablation is named w / o-SN.

[0192] (2) Ablate the reconstruction consistency loss and analyze its impact on the model performance. The model after ablation is named w / o-RL.

[0193] (3) Ablate the shape naturalness module and the reconstruction consistency loss, and analyze the impact of only using the self-training method on the model performance. The model after ablation is named w / o-SN-RL.

[0194] (4) Ablate the self-training method and only train the conditional diffusion 3D point cloud reconstruction network model CDPRN under the condition of a small amount of labeled data.

[0195] Table 5 is a CD comparison table of the reconstruction quality of each model on the ShapeNet dataset.

[0196] Table 5

[0197]

[0198]

[0199] Analysis of Table 5 shows that compared with CDPRN, the average CD value of the reconstruction quality of ST - CDPRN decreases by 1.22, and the reconstruction performance of each category is better than that of CDPRN. This indicates that ST - CDPRN can effectively utilize unlabeled data to improve the reconstruction performance of the model. At the same time, in terms of the reconstruction performance of each category, compared with ST - CDPRN, the average CD of the reconstruction quality of w / o - RL increases by 0.47, indicating that the reconstruction consistency loss plays a certain role in improving the reconstruction ability of the model. w / o - SN is inferior to ST - CDPRN and CDPRN, indicating that the shape naturalness module plays a huge role in enhancing the internal correlation between images and point clouds, improving the quality of the training set, and improving the model performance. At the same time, the results of w / o - SN - RL are similar to those of w / o - SN, indicating that in the absence of the shape naturalness module to evaluate the quality of pseudo - labels, the reconstruction consistency loss cannot fully utilize the pseudo - label information with higher quality, and the help for improving the model performance is small. Table 6 is a comparison table of the EMD of the reconstruction quality of each model on the ShapeNet dataset.

[0200] Table 6

[0201]

[0202]

[0203] Analysis of Table 6 shows that the EMD between the reconstructed point cloud of ST - CDPRN and the real point cloud is better than that of CDPRN in each category. Especially for the reconstruction of table lamps and tables, the addition of pseudo - labels enables the model to learn more extensive features and achieve better reconstruction results. Compared with ST - CDPRN, the reconstruction ability of w / o - SN and w / o - SN - RL has regressed significantly, even far lower than that of CDPRN, further indicating that the shape naturalness module can effectively reduce the noise introduced during the training process, enhance the quality of the training set, and improve the model performance.

[0204] S4. Obtain image data from the 3D point cloud reconstruction dataset, randomly sample Gaussian noise, input the image and Gaussian noise into the self - training conditional diffusion 3D point cloud reconstruction model, and the student network in it performs the reverse denoising process to output the 3D point cloud reconstruction result for 3D point cloud reconstruction.

[0205] From the above experimental results, it can be concluded that the three-dimensional point cloud reconstruction network model with conditional diffusion constructed in this embodiment can effectively improve the accuracy of the three-dimensional point cloud reconstruction results. At the same time, constructing a self-training three-dimensional point cloud reconstruction network model with conditional diffusion can effectively handle the three-dimensional point cloud reconstruction task in the scenario of small sample data and maintain good reconstruction performance.

[0206] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model, characterized in that, Including: Obtain the image to be reconstructed and Gaussian noise; Construct a 3D point cloud reconstruction network model, introduce a conditional aggregation module and a feature consistency loss, obtain a conditional diffusion 3D point cloud reconstruction network model, use the conditional diffusion 3D point cloud reconstruction network model as the teacher sub-model and the student sub-module, and combine a shape naturalness module and a reconstruction consistency loss to obtain a self-training conditional diffusion 3D point cloud reconstruction network model; The shape naturalness module is used to explore the internal connection between the reconstructed point cloud and the unlabeled image data, obtain a pseudo-labeled point cloud with a natural structure of the point cloud and consistent with the unlabeled image data to expand the 3D point cloud reconstruction training set, and train the student sub-module in combination with a denoising loss function and a reconstruction consistency loss; The shape naturalness module includes: An image feature extraction sub-module for extracting features from the unlabeled image data to obtain first image features; A point cloud feature extraction sub-module for extracting geometric structure information of the reconstructed point cloud to obtain point cloud features; A multi-scale convolution sub-module for splicing the first image features and the point cloud features to obtain fused features; A self-attention sub-module for recalibrating the fused features to obtain recalibrated features, comprehensively analyzing the recalibrated features to obtain a point cloud quality assessment result, judging whether the structure of the point cloud is natural and consistent with the unlabeled image data according to the point cloud quality assessment result, obtaining a pseudo-labeled point cloud with a natural structure of the point cloud and consistent with the unlabeled image data to expand the 3D point cloud reconstruction training set, and training the student sub-module in combination with a denoising loss function and the reconstruction consistency loss; The student sub-module includes: A feature extraction model for extracting features from the image to be reconstructed to obtain second image features; A conditional aggregation module for splicing the second image features and a noisy point cloud to obtain an input point cloud; Input the image to be reconstructed and the Gaussian noise into the self-training conditional diffusion 3D point cloud reconstruction network model to obtain a point cloud reconstruction result.

2. The three-dimensional point cloud reconstruction method based on the self-training conditional diffusion model according to claim 1, wherein The feature consistency loss is: Among them, is the feature consistency loss, C is the number of channels of the feature map, and are the values of the feature maps Z o and Z a at the i-th element position and the c-th channel respectively, and M is the total number of elements on the feature map.

3. The three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model according to claim 1, wherein The self-training conditional diffusion 3D point cloud reconstruction network model includes: A teacher sub-model for obtaining unlabeled image data, performing 3D point cloud reconstruction on the unlabeled image data to obtain a reconstructed point cloud; A student sub-module for dynamically updating the parameters of the teacher sub-model and obtaining the point cloud reconstruction result according to the image to be reconstructed and the Gaussian noise.

4. The three-dimensional point cloud reconstruction method based on the self-training conditional diffusion model according to claim 1, wherein The image feature extraction sub-module includes: A two-dimensional convolution unit for performing preliminary convolution operations on the unlabeled image data; A normalization unit for normalizing the convolution operation result to obtain shallow features; A number of residual units for performing deep processing on the shallow features to obtain a deep feature map; A max-pooling unit for reducing the dimension of the deep feature map to obtain first image features.

5. The 3D point cloud reconstruction method based on self-training conditional diffusion model according to claim 1, characterized in that, The point cloud feature extraction sub-module includes: A number of SA units are used to downsample the reconstructed point cloud, obtain key point information, aggregate the point information around the key points in the key point information, obtain a local point set, extract features for each point in the local point set, and perform max pooling on the features of all points in the feature dimension to obtain hierarchical features; An MLP unit is used to process the hierarchical features to obtain point cloud features.

6. The three-dimensional point cloud reconstruction method based on the self-training conditional diffusion model according to claim 4, wherein The multi-scale convolution sub-module includes: A reshaping unit is used to reshape the size of the point cloud features into the first image feature size; A splicing unit is used to splice and concatenate the reshaped point cloud features with the first image features; A number of groups of convolution units are used to perform convolution processing on the concatenated feature maps, splice and concatenate multi-scale information, and obtain fused features.

7. The 3D point cloud reconstruction method based on the self-training conditional diffusion model according to claim 1, characterized in that The method for obtaining the denoising loss function is: Among them, is the denoising loss function, is the predicted noise, ∈ is Gaussian noise, and x t is the noisy point cloud, t is the time step, is the second image feature; The method for obtaining the reconstruction consistency loss is: Among them, is the reconstruction consistency loss function, and P r is the reconstructed point cloud of the student network for the enhanced image reconstruction, and P p is the pseudo-label point cloud corresponding to the unlabeled image data, and d(·) is the similarity metric function. is the weight function related to the time step t.

8. The three-dimensional point cloud reconstruction method based on the self-training conditional diffusion model according to claim 1, wherein The student sub-module includes: A first reverse denoising module is used to perform reverse denoising on the Gaussian noise to obtain a noisy point cloud; A noise estimation module is used to predict the input point cloud to obtain predicted noise; A second reverse denoising module is used to perform reverse denoising on the predicted noise to obtain the point cloud reconstruction result.

9. The three-dimensional point cloud reconstruction method based on a self-training conditional diffusion model according to claim 3, wherein Obtaining the input point cloud includes: Performing bilinear interpolation calculation on the D dimension of the second image feature to obtain an interpolated feature map, combining the internal and external parameter matrices of the camera parameters, converting the coordinates of any point in the noisy point cloud into the pixel coordinates of the interpolated feature map, obtaining the D-dimensional vector corresponding to the pixel coordinates of the interpolated feature map, splicing the D-dimensional vector with the coordinates of the noisy point cloud, and obtaining the input point cloud.