Polarization three-dimensional reconstruction method and system based on prior guide diffusion model

By guiding the diffusion model with priori knowledge, a conditional U-Net diffusion model is constructed using an encoder, vector quantizer, and decoder. This solves the problems of existing methods such as sensitivity to noise and insufficient generalization ability, and achieves high-precision and robust polarization 3D reconstruction.

CN120612424APending Publication Date: 2025-09-09WUHAN UNIV
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510678302.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing deep learning-based polarization 3D reconstruction methods rely too much on labeled data, have limited generalization capabilities, are sensitive to noise, resulting in distortion of details, and the generated models are not robust enough in noisy environments.

Method used

A prior-guided diffusion model is adopted, and through multi-level collaborative optimization, a conditional U-Net diffusion model is constructed using the encoder, vector quantizer and decoder. Combined with polarization prior information, the three-dimensional normal vector is gradually denoised and restored. A staged training strategy is adopted, first pre-training the VQGAN encoder and decoder, then training the diffusion model, and using Markov chains and attention mechanisms to achieve high-fidelity reconstruction.

Benefits of technology

It improves the 3D reconstruction accuracy and robustness against noise interference in complex scenes, the physical authenticity and stability of the generation process, and can generate surface normals with rich details in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612424A_ABST
    Figure CN120612424A_ABST
Patent Text Reader

Abstract

The invention provides a polarization three-dimensional reconstruction method and a polarization three-dimensional reconstruction system based on a prior guide diffusion model, which apply a diffusion model in the field of polarization three-dimensional reconstruction and improve the recovery capability of complex details and the robustness of noise interference resistance. According to the method, a two-stage training mode is adopted, and the generation quality and the calculation efficiency are balanced through step-by-step optimization. In the first stage, a VQGAN codec is independently trained to learn high-efficiency low-dimensional potential representation of an image, and direct high-cost calculation in a pixel space is avoided; in the second stage, learnable parameters of the VQGAN are frozen, a diffusion model is trained on a trained potential space, gradual denoising is guided through a priori condition, and potential features are generated and mapped back to an image space. The diffusion model effectively fuses the physical constraint of the polarization clue and the data prior in the gradual denoising process, and the surface normal with rich details can still be stably generated in the case of noise interference or information loss. Experimental results show that the method provided by the invention is excellent in surface normal reconstruction in a plurality of complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and deep learning technology, and relates to a polarization three-dimensional reconstruction method and system based on a priori guided diffusion model, which is suitable for three-dimensional reconstruction application scenarios with high precision requirements. Background Art

[0002] Polarization 3D reconstruction, an important passive 3D reconstruction method, has made significant progress in theoretical research and practical applications in recent years. Polarization 3D reconstruction estimates the target surface normal based on the target's reflective properties of polarized light, enabling 3D reconstruction of the target surface. Compared to traditional 3D reconstruction techniques based on stereo vision or structured light projection, this technology offers advantages such as non-contact measurement, immunity to ambient stray light interference, and long-range detection. It provides an effective solution for low-texture surfaces and highly reflective materials, where traditional methods fail.

[0003] In recent years, deep learning has been introduced into the field of polarization 3D reconstruction. In 2020, Kondo et al. proposed a new polarization bidirectional reflectance distribution function model. [1] , used to create a synthetic dataset of polarization images and estimate surface normals based on convolutional neural networks. Ba et al. proposed a U-Net based [2] Subsequently, Lei et al. constructed a new complex scene dataset [3] , and proposed a deep learning-based polarization 3D reconstruction method. Existing deep learning-based polarization 3D reconstruction methods mostly use discriminative architectures such as CNN and Transformer. These models directly map the polarization image and its prior information to the target surface normal vector, without explicitly modeling the data distribution. While the training process is relatively efficient, they rely too much on labeled data, limiting their generalization capabilities for unknown scenes. Furthermore, discriminative models are sensitive to input noise, especially in low-signal-to-noise ratio polarization images. Noise can be directly propagated to the output through end-to-end mapping, resulting in distorted details.

[0004] Unlike discriminative models, generative models such as diffusion models do not directly learn deterministic mappings, but instead achieve reconstruction by modeling the joint probability distribution of polarization data and target three-dimensional parameters. In the forward process, the diffusion model gradually adds Gaussian noise to erode the true value of the three-dimensional normal vector. In the reverse process, it uses priors such as the polarization azimuth and degree of polarization to guide the back diffusion, learns a denoising function that recovers the three-dimensional structure from the noise, and ultimately generates the target three-dimensional normal vector that conforms to physical laws. The diffusion model implicitly learns the manifold structure of the polarization data during training, and exhibits greater robustness to polarization images with noise, partial occlusion, and material changes. In addition, the model adopts a progressive generation mechanism, which can gradually refine high-frequency detail information during the generation process, restore finer geometric features, and significantly improve the accuracy and realism of three-dimensional reconstruction. Therefore, the application of diffusion models to polarization three-dimensional reconstruction tasks has great potential.

[0005] Related references are as follows: [1]Y. Kondo, T. Ono, L. Sun, Y. Hirasawa, and J. Murayama, “Accuratepolarimetric BRDF for real polarization scene rendering,” in Proc.Eur.Conf. Comput. Vis. (ECCV), 2020, pp. 220–236. [2]Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutionalnetworks for biomedical image segmentation,” in Proc. MICCAI, 2015, pp. 234–241. [3] C. Lei, C. Qi, J. Xie, N. Fan, V. Koltun, and Q. Chen, “Shapefrompolarization for complex scenes in the wild,” in Proc. IEEE / CVFConf.Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 12 632–12 641. Summary of the Invention To address the shortcomings of existing technologies, this paper proposes a novel polarization-based 3D reconstruction method and system based on a priori-guided diffusion model. An encoder module encodes and decodes the input 3D normal vector, mapping the high-dimensional continuous image space into a structured discrete latent space, achieving efficient data compression while preserving key features. A conditional U-Net diffusion model is constructed in the encoded latent space to learn the mapping distribution from polarization modes to 3D normal vectors, accurately recovering the target surface normal field and thus reconstructing high-quality 3D objects.

[0006] The technical solution adopted by the present invention is: a polarization 3D reconstruction method based on a priori guided diffusion model, which realizes high-quality 3D reconstruction through multi-level collaborative optimization. First, a polarization prior map group is obtained from polarization images at different polarization angles, including prior information such as polarization degree, non-polarized image, re-represented azimuth angle and field of view encoding, as conditional constraint information to guide the diffusion model, so as to strengthen the correlation between 3D geometric features and optical properties. Then, VQGAN is introduced as the core module of encoding and decoding, in which the encoder compresses high-dimensional image information into a low-dimensional discrete latent space through convolutional layers and vector quantization technology, and the decoder reconstructs image details through adversarial training, while retaining high-frequency textures and reducing data dimensions, providing efficient and fidelity intermediate representation for subsequent diffusion models. Finally, a diffusion architecture based on polarization prior guidance is adopted to embed the U-Net backbone into the diffusion framework. The forward diffusion process starts from the 3D normal vector and gradually adds noise to make it gradually degenerate into a pure noise state; the reverse process uses the polarization prior as a guide, gradually removes the noise through the model, and accurately restores the original 3D normal vector. This diffusion model based on polarization prior guidance uses a Markov chain to gradually learn the conditional probability mapping relationship between the three-dimensional normal vector and the noise, ensuring the physical authenticity and stability of the generation process. A phased training strategy is adopted. The VQGAN encoder and decoder are first pre-trained on the scene dataset SPW, and a composite loss function suitable for the network is designed to establish a high-fidelity low-dimensional latent space representation. In the second stage, the learnable parameters of the pre-trained VQGAN encoder are frozen, and the diffusion model is trained exclusively. The model combines time embedding, residual blocks, and attention mechanisms to achieve efficient modeling of the data, and generates outputs that conform to physical constraints through a priori guidance mechanism. Finally, we perform three-dimensional reconstruction based on the trained model. The method includes the following steps: Step 1: Obtain the three-dimensional normal vector of the target surface and the polarization prior input map group, use the polarization prior input map group as the original prior condition, and obtain the true value of the three-dimensional normal vector at the same time; Step 2: construct a codec network, which consists of four parts: encoder, vector quantizer, decoder and discriminator; The encoder downsamples and extracts features from the input 3D normal vector, mapping it to a continuous latent space. The vector quantizer maps the continuous features output by the encoder to corresponding discrete codebook vectors using a nearest neighbor lookup table. The decoder reconstructs the image from the discrete codebook vectors, and the discriminator is used for adversarial training. Step 3: Construct a diffusion model based on prior guidance. The forward process of the diffusion model starts from the distribution of discrete codebook vectors and converts the data into pure noise by gradually adding Gaussian noise. The reverse process concatenates the prior condition information encoded by the encoder and vector quantizer with the noisy discrete vector at the current diffusion step t in the channel dimension. The noise is gradually removed by predicting the noise through the U-Net network, and the original latent variables are restored. Finally, the latent variables are converted into the reconstruction results in the pixel space through the decoder.

[0007] Furthermore, the encoder consists of a series of convolutional layers, residual blocks, and attention blocks. The specific processing process is as follows: The encoder first maps the input image from 3 channels to basic channels, keeping the spatial resolution unchanged, , H and W represent the height and width of the image respectively; then through three resolution layers, each layer executes two residual blocks ResnetBlock, and performs downsampling operations after the first two layers to reduce the spatial resolution by half; in each resolution layer, the number of output channels is 、 and ; Then the features are further extracted and fused through the middle part containing two residual blocks; Finally, after normalization and Swish activation function processing, a 3×3 convolution layer is used to map the feature map to 3 channels, and the output spatial resolution is ( × )’s latent variable feature map.

[0008] Furthermore, the vector quantizer converts the continuous latent features extracted by the encoder into is mapped into a discrete vector space, which is represented by an embedding layer ( ,3) constitutes, used to define a The specific processing process of the quantizer is as follows.

[0009] First, the encoded continuous latent features , the shape is ) is rearranged to , and flattened to , where H and W represent the height and width of the feature respectively, and B represents the number of samples contained in a batch. Then calculate each input vector With codebook vector The Euclidean distance is calculated by the formula:

[0010] To select the nearest codebook vector Replace the original input to complete the discretization; the discretization result is gradient preserved through the direct technology, and finally the codebook vector Restore to the input The same shape is achieved by replacing the continuous vector with the best discrete approximation in the codebook.

[0011] Furthermore, the loss function of the vector quantizer consists of two parts: embedding loss and commitment loss. The embedding loss only produces gradients for the codebook vector, which makes the codebook vector Close to the encoder output , the commitment loss is only applied to the encoder output The purpose of generating gradients is to make the encoder output close to the codebook vector; the loss function formula is as follows:

[0012]

[0013]

[0014] in, represents the quantization loss, and is the intermediate amount, , is the hyperparameter for adjusting the loss, z is the output of the encoder, is the nearest neighbor vector found in the codebook, sg( ) means stopping the gradient, that is, the gradient of this variable is not updated during back propagation.

[0015] Furthermore, the decoder receives 3 channels and has a resolution of × The quantized tensor is mapped to channels, the resolution remains unchanged; then enters the intermediate module, and passes through two Resnet residual block of the channel and a The global attention block of the channel is used for deep extraction and fusion of features; then enter the upsampling stage, execute 3 residual blocks in the first level, and increase the number of channels from becomes , by upsampling the resolution from × Restore to × ; In the second level, 3 residual blocks are also executed, and the number of channels is changed from becomes , and upsample again to increase the resolution to ; 3 residual blocks are executed at the third level, the number of channels remains unchanged, and the resolution remains ; Finally, the output The channel feature map is normalized and Swish activated, and then mapped to the final 3-channel image through 3×3 convolution, completing the gradual reconstruction from the latent space to the high-resolution image.

[0016] Furthermore, in step 2, the following total loss is used as a guide for optimizing the generator network, which consists of an encoder, a vector quantizer, and a decoder;

[0017] in, is the total loss, is the normal vector reconstruction loss, is the perceptual loss, To combat losses, is the quantization loss, 、 are hyperparameters for adjusting perceptual loss and quantization loss, To combat loss dynamics, It is an adaptive weight, which balances the reconstruction loss and the adversarial loss by the ratio of the gradient magnitude; 、 、 The expression is as follows:

[0018]

[0019]

[0020] in represents the dot product, represents the reconstructed surface normal, is the true value corresponding to the surface normal, is the pixel position ( ), is the true value of the surface normal at that location, Represents the first layer feature extractor, Represents the discriminative feature map output by the discriminator.

[0021] Furthermore, the discriminator distinguishes real images from generated images through hinge loss, which is expressed as:

[0022] in, Indicates hinge loss, Express expectations, represents the reconstructed surface normal, is the true value corresponding to the surface normal, Represents the discriminative feature map output by the discriminator.

[0023] Furthermore, in step 3, the forward process of the diffusion model is the low-dimensional latent variable That is, the discrete codebook vector Gradually add Gaussian noise, and finally pass Step 1: transform the data into Gaussian noise , Indicates that the mean is 0 and the covariance matrix is ​​the identity matrix I Normal distribution; The reverse process from Gaussian noise Departure, pass conditions Guide the adjustment of denoising direction and gradually denoise and reconstruct the latent variables .

[0024] Furthermore, in step 3, the learnable parameters of the codec are frozen, and the diffusion model is trained on the dataset using the variational lower bound loss function. Supervised optimization of this stage of training aims to minimize the noise of the U-Net network predictions With real noise The difference between , enables the diffusion model to learn to effectively denoise under conditional guidance:

[0025] in, Express expectations, is the standard normally distributed noise, Indicates the Step data, Indicates conditional information. It is a U-Net network used to predict the added noise; Finally, the latent variables reconstructed by denoising the diffusion model The decoder converts it back to the original data space to obtain the final reconstruction result.

[0026] The present invention also provides a polarization three-dimensional reconstruction system based on a priori guided diffusion model, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, a polarization three-dimensional reconstruction method based on a priori guided diffusion model as described in the above technical solution is implemented.

[0027] Compared with existing technologies, the present invention offers the following advantages and benefits: The present invention proposes a 3D reconstruction method based on a prior-guided diffusion model. This generative model, a diffusion model, is applied to the field of polarization 3D reconstruction, improving the ability to recover complex details and robustness against noise interference. This method employs a two-stage training approach, balancing generation quality and computational efficiency through step-by-step optimization. In the first stage, the VQGAN is independently trained to learn an efficient, low-dimensional latent representation of the image, avoiding costly computations directly in pixel space. In the second stage, the learnable parameters of the VQGAN are frozen, and a diffusion model is trained on the trained latent space. This model generates latent features through step-by-step denoising guided by prior conditions and maps them back to the image space. The diffusion model effectively integrates the physical constraints of polarization cues with data priors during the step-by-step denoising process, enabling the stable generation of detailed surface normals despite noise interference or information loss. Experimental results demonstrate that the proposed method performs well in reconstructing surface normals in multiple complex scenarios, demonstrating the potential of prior-guided diffusion models for polarization 3D reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the polarization prior map group.

[0029] Figure 2 This is the VQGAN encoding and decoding network structure diagram.

[0030] Figure 3 It is a schematic diagram of the prior guided diffusion model.

[0031] Figure 4 It is a structural diagram of two network cascades.

[0032] Figure 5 This is the result diagram of the forward process of the diffusion model.

[0033] Figure 6 This is the result diagram of the reverse process of the diffusion model.

[0034] Figure 7 The scene-level dataset SPW of the embodiment is a result of 3D reconstruction of the target surface. DETAILED DESCRIPTION

[0035] In order to facilitate ordinary technicians in this field to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0036] This invention mainly targets the application requirements of three-dimensional reconstruction with high precision, and proposes a new polarization three-dimensional reconstruction method based on deep learning. The encoder module is used to encode and decode the input three-dimensional normal vector, mapping the high-dimensional continuous image space to a structured discrete latent space, achieving efficient data compression and key feature retention. In the encoded latent space, a conditional U-Net diffusion model is constructed to learn the mapping distribution from polarization mode to three-dimensional normal vector, accurately restore the normal vector field of the target surface, and thus reconstruct a high-quality three-dimensional target. The method specifically includes the following steps: Step 1: Obtain the three-dimensional normal vector of the target surface and the polarization prior input map group, and use the polarization prior input map group as the original prior condition: obtain the polarization images of the target object and the target scene at different polarization angles, and calculate the non-polarized images based on the multi-angle polarization images ( ), polarization image ( ), polarization angle image ( ), polarization angle coded image ( ) and field of view encoding ( ) image and other polarization prior input maps. At the same time, the polarization prior target surface reconstruction result, that is, the true value of the 3D normal vector, is obtained.

[0037] In this embodiment, the specific implementation of step 1 includes the following sub-steps: Step 1.1: The intensity of each pixel in the polarization image at different polarization angles is expressed as:

[0038] in is the unpolarized intensity, is the polarization angle, is the degree of polarization, from the polarization angle Four polarization images of ( ) are calculated separately , and ; Step 1.2, surface normal azimuth and zenith angle From the polarization angle and polarization degree Calculated, expressed as:

[0039] When specular reflection dominates, and The relationship is:

[0040] in is the target surface reflectivity; When diffuse reflection dominates, and The relationship is:

[0041] Further get the target surface normal:

[0042] in, is the three-dimensional normal vector, 、 、 are the component values ​​of the target surface normal vector in three directions respectively; Step 1.3, get the phase angle Re-expression of polarization angle encoding and field of view encoding , where the field of view encoding It can be obtained from the intrinsic parameters of the camera, and further obtain the polarization prior map group ,in The reverse process used to guide the diffusion model:

[0043]

[0044] The final polarization prior map group is as shown in the attached Figure 1 shown.

[0045] Step 2: Construct an encoder-decoder network to provide a high-fidelity, low-dimensional operating space for the diffusion model. The encoder-decoder network consists of four parts: an encoder, a vector quantizer, a decoder, and a discriminator. The encoder consists of a series of convolutional layers, residual blocks, and attention blocks. It downsamples and extracts features from the input target three-dimensional normal vector and maps it to a continuous latent space. The vector quantizer maps the continuous features output by the encoder to corresponding discrete codebook vectors through a nearest neighbor lookup table, constrains the latent space dimension, and forces the network to learn a limited and semantically rich set of representations. The decoder structure also consists of residual blocks and attention blocks. It gradually restores the spatial resolution through upsampling and reconstructs the image from the discrete codebook vector. The discriminator is used for adversarial training to improve the quality and authenticity of the generated image.

[0046] Then, the VQGAN encoder and decoder are independently trained based on the scene dataset SPW. The input of the encoder is the surface normal vector. The composite loss functions such as cosine similarity loss, perceptual loss, quantization loss and adversarial loss are used to calculate the gap between the reconstruction result and the true value of the surface normal vector, which is used to improve the encoder's perception of the surface normal vector.

[0047] In this embodiment, the specific implementation of step 2 includes the following sub-steps: Step 2.1, the VQGAN encoder-decoder consists of four parts: encoder, vector quantizer, decoder and discriminator. The encoder first transforms the input image (3D normal vector) through a 3×3 convolutional layer. ) is mapped from 3 channels to base channels, keeping the spatial resolution unchanged ( , H and W represent the height and width of the image respectively). Then, through three resolution layers, two residual blocks (ResnetBlock) are executed in each layer, and downsampling operations are performed after the first two layers to halve the spatial resolution. In each resolution layer, the number of output channels is 、 and Then, the features are further extracted and fused through the middle part containing two residual blocks. Finally, after normalization and Swish activation function processing, a 3×3 convolution layer is used to map the feature map to 3 channels, and the output spatial resolution is ( × ). The entire encoder structure performs two downsampling operations, each time halving the spatial resolution, and the final image resolution is reduced to one-quarter of the original.

[0048] Step 2.2, the quantizer converts the continuous latent features extracted by the encoder into Mapped into a discrete vector space, its main structure consists of an embedding layer ( ,3) constitutes, used to define a The specific processing of the quantizer is as follows.

[0049] First, the encoded continuous latent features (Shape is ) is rearranged to , and flattened to , where B represents the number of samples contained in a batch, and then calculate each input vector With codebook vector (common The Euclidean distance of the two numbers is calculated by the formula:

[0050] To select the nearest codebook vector to replace the original input, thus completing the discretization. The discretized result is gradient preserved through the straight-through technique to ensure that the model can be trained end-to-end. Finally, the codebook vector (This is for each input Calculate it and all codebook vectors Euclidean distance, replace the codebook vector with the one with the smallest distance obtained) is restored to the same value as the input The same shape is achieved by replacing the continuous vector with the best discrete approximation in the codebook.

[0051] The loss function of the quantizer consists of two parts: embedding loss and commitment loss. The embedding loss only produces gradients for the codebook vector, which makes the codebook vector Close to the encoder output , the commitment loss is only applied to the encoder output The purpose of generating gradients is to make the encoder output close to the codebook vector. The loss function formula is as follows:

[0052]

[0053]

[0054] in, , is the hyperparameter for adjusting the loss, z is the output of the encoder, is the nearest neighbor vector found in the codebook, sg( ) means stopping the gradient, that is, the gradient of this variable is not updated during back propagation.

[0055] Step 2.3, the decoder module receives 3 channels and a resolution of × The quantized tensor is mapped to channels, the resolution remains unchanged; then enters the intermediate module, and passes through two Resnet residual block of the channel and a The global attention block of the channel is used for deep extraction and fusion of features; then enter the upsampling stage, execute 3 residual blocks in the first level, and increase the number of channels from becomes , by upsampling the resolution from × Restore to × ; In the second level, 3 residual blocks are also executed, and the number of channels is changed from becomes , and upsample again to increase the resolution to ; 3 residual blocks are executed at the third level, the number of channels remains unchanged, and the resolution remains ; Finally, the output The channel feature map is normalized and Swish activated, and then mapped to the final 3-channel image through 3×3 convolution, completing the gradual reconstruction from the latent space to the high-resolution image.

[0056] Furthermore, in step 2, the following loss function is used as a guide for optimizing the generator network (composed of the encoder, vector quantizer, and decoder):

[0057] in, is the total loss, is the normal vector reconstruction loss, is the perceptual loss, To combat losses, is the quantization loss mentioned in step 2.2, 、 are hyperparameters for adjusting perceptual loss and quantization loss, To combat the loss dynamic factor, it is set to 0 at the beginning of training, allowing the network to focus on learning pixel-level and perceptual reconstruction first, and then adding the adversarial loss after the basic reconstruction quality stabilizes. It is an adaptive weight that balances the reconstruction loss and the adversarial loss by the ratio of the gradient magnitude. 、 、 The expression is as follows:

[0058]

[0059]

[0060] in represents the dot product, represents the reconstructed surface normal, is the true value corresponding to the surface normal, is the pixel position ( ), is the true value of the surface normal at that location, Represents the first layer feature extractor, Represents the discriminant feature map output by the discriminator. The larger the value, the more "real" it is.

[0061] The discriminator distinguishes real images from generated images through hinge loss, forcing the generator to produce more realistic images. The expression of this loss function is:

[0062] Step 3: Construct a diffusion architecture based on prior guidance, the core of which is the U-Net network. The forward process of the diffusion model starts from the distribution of the encoded data (i.e., the discrete codebook vector) and gradually adds Gaussian noise to convert the data into pure noise. The reverse process combines the prior condition information encoded by the encoder and quantizer with the input data (the noised discrete vector at the current diffusion step t). ) Splicing is performed on the channel dimension, and the noise is gradually removed by predicting the noise through the U-Net network to restore the original latent variable Finally, the VQGAN decoder converts the latent variables into pixel-space reconstruction results. We freeze the learnable parameters of the VQGAN encoder and decoder and continue to train the diffusion model on the dataset, using the variational lower bound loss function to supervise and optimize the training of this stage.

[0063] In this embodiment, the specific implementation of step 3 includes the following sub-steps: Step 3.1, in the latent space, the forward process of the diffusion model is to calculate the low-dimensional latent variables (i.e., discrete codebook vector ) gradually add Gaussian noise, and finally pass Step 1: transform the data into Gaussian noise , Indicates that the mean is 0 and the covariance matrix is ​​the identity matrix I Specifically, the forward process is defined as a Markov chain, which is expressed as follows

[0064] Each step The conditional probability is:

[0065] in, Indicates the The data after adding noise, Indicates the Step data, is a predefined noise scheduling parameter that controls the amount of noise added.

[0066] Step 3.2, reverse process of diffusion model from Gaussian noise Departure, pass conditions Guide the adjustment of denoising direction and gradually denoise and reconstruct the latent variables The reverse process is defined as a parameterized Markov chain:

[0067] The conditional distribution at each step is:

[0068] in, express The data after denoising, Indicates the Step data, Indicates conditional information. It is a mean prediction network, which is used to predict the mean of the denoised data. It is a covariance prediction network, which is used to predict the covariance of denoised data.

[0069] Furthermore, in step 3, the following variational lower bound loss function is used as a guide for network optimization, with the goal of minimizing the noise of the U-Net network prediction With real noise The difference enables the model to learn to effectively denoise under conditional guidance:

[0070] in, is the standard normally distributed noise, It is a U-Net network used to predict the added noise.

[0071] Finally, we denoise the reconstructed latent representation by the diffusion model The decoder converts it back to the original data space to obtain the final reconstruction result.

[0072] We use the proposed prior-guided diffusion model to reconstruct the target surface normal on the test set of the scene-level dataset SPW, and the results are shown in the attached figure. Figure 7 shown.

[0073] The results show that the reconstruction results obtained by the method proposed in this invention can reconstruct the target surface information with high quality on scene-level data, have stronger ability to reconstruct detail information, and are more generalizable.

[0074] On the other hand, an embodiment of the present invention also provides a polarization three-dimensional reconstruction system based on a priori guided diffusion model, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements a polarization three-dimensional reconstruction method based on a priori guided diffusion model as described in the above technical solution.

[0075] It should be understood that parts not elaborated in detail in this specification belong to the prior art.

[0076] It should be understood that the above description of the embodiments is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.

Claims

1. A polarization 3D reconstruction method based on a priori guided diffusion model, characterized in that: The steps include: Step 1: Obtain the three-dimensional normal vector of the target surface and the polarization prior input map group, use the polarization prior input map group as the original prior condition, and obtain the true value of the three-dimensional normal vector at the same time; Step 2: construct a codec network, which consists of four parts: encoder, vector quantizer, decoder and discriminator; The encoder downsamples and extracts features from the input 3D normal vector and maps it to a continuous latent space; The vector quantizer maps the continuous features output by the encoder to corresponding discrete codebook vectors through a nearest neighbor lookup table. The decoder reconstructs the image from the discrete codebook vectors, and the discriminator is used for adversarial training. Step 3: Construct a diffusion model based on prior guidance. The forward process of the diffusion model starts from the distribution of discrete codebook vectors and converts the data into pure noise by gradually adding Gaussian noise. The reverse process concatenates the prior condition information encoded by the encoder and vector quantizer with the noisy discrete vector at the current diffusion step t in the channel dimension. The noise is gradually removed by predicting the noise through the U-Net network, and the original latent variables are restored. Finally, the latent variables are converted into the reconstruction results in the pixel space through the decoder.

2. The polarization 3D reconstruction method based on a priori guided diffusion model according to claim 1, characterized in that: The encoder consists of a series of convolutional layers, residual blocks, and attention blocks. The specific processing process is as follows: The encoder first maps the input image from 3 channels to Basic channels, keeping the spatial resolution unchanged, , H and W represent the height and width of the image respectively; then through three resolution layers, each layer executes two residual blocks ResnetBlock, and performs downsampling operations after the first two layers to reduce the spatial resolution by half; in each resolution layer, the number of output channels is 、 and ; Then the features are further extracted and fused through the middle part containing two residual blocks; Finally, after normalization and Swish activation function processing, a 3×3 convolution layer is used to map the feature map to 3 channels, and the output spatial resolution is ( × )’s latent variable feature map.

3. The polarization 3D reconstruction method based on a priori guided diffusion model according to claim 1, characterized in that: The vector quantizer converts the continuous latent features extracted by the encoder into is mapped into a discrete vector space, which is represented by an embedding layer ( ,3) constitutes, used to define a The specific processing process of the quantizer is as follows. First, the encoded continuous latent features , the shape is ) is rearranged to , and flattened to , where H and W represent the height and width of the feature respectively, and B represents the number of samples contained in a batch. Then calculate each input vector With codebook vector The Euclidean distance is calculated by the formula: To select the nearest codebook vector Replace the original input to complete the discretization; the discretization result is gradient preserved through the direct technology, and finally the codebook vector Restore to the input The same shape is achieved by replacing the continuous vector with the best discrete approximation in the codebook.

4. The polarization 3D reconstruction method based on a priori guided diffusion model according to claim 3, characterized in that: The loss function of the vector quantizer consists of two parts: embedding loss and commitment loss. The embedding loss only produces gradients for the codebook vector, which makes the codebook vector Close to the encoder output , the commitment loss is only applied to the encoder output The purpose of generating gradients is to make the encoder output close to the codebook vector; the loss function formula is as follows: in, represents the quantization loss, and is the intermediate amount, , is the hyperparameter for adjusting the loss, z is the output of the encoder, is the nearest neighbor vector found in the codebook, sg( ) means stopping the gradient, that is, the gradient of this variable is not updated during back propagation.

5. The polarization 3D reconstruction method based on a priori guided diffusion model according to claim 1, characterized in that: The decoder receives 3 channels and has a resolution of × The quantized tensor is mapped to channels, the resolution remains unchanged; Then enter the middle module, go through two Resnet residual block of the channel and a The global attention block of the channel is used for deep extraction and fusion of features; then enter the upsampling stage, execute 3 residual blocks in the first level, and increase the number of channels from becomes , by upsampling the resolution from × Restore to × ; In the second level, three residual blocks are also executed, and the number of channels is increased from becomes , and upsample again to increase the resolution to ; 3 residual blocks are executed at the third level, the number of channels remains unchanged, and the resolution remains ; Finally, the output The channel feature map is normalized and Swish activated, and then mapped to the final 3-channel image through 3×3 convolution, completing the gradual reconstruction from the latent space to the high-resolution image.

6. The polarization 3D reconstruction method based on a priori guided diffusion model according to claim 1, characterized in that: In step 2, the following total loss is used as a guide for optimizing the generator network, which consists of an encoder, a vector quantizer, and a decoder. in, is the total loss, is the normal vector reconstruction loss, is the perceptual loss, To combat losses, is the quantization loss, 、 are hyperparameters for adjusting perceptual loss and quantization loss, To combat the loss dynamics, It is an adaptive weight, which balances the reconstruction loss and the adversarial loss by the ratio of the gradient magnitude; 、 、 The expression is as follows: in represents the dot product, represents the reconstructed surface normal, is the true value corresponding to the surface normal, is the pixel position ( ), is the true value of the surface normal at that location, Represents the first layer feature extractor, Represents the discriminative feature map output by the discriminator.

7. The polarization 3D reconstruction method based on a priori guided diffusion model according to claim 1, characterized in that: The discriminator distinguishes real images from generated images through hinge loss, which is expressed as: in, Indicates hinge loss, Express expectations, represents the reconstructed surface normal, is the true value corresponding to the surface normal, Represents the discriminative feature map output by the discriminator.

8. The polarization 3D reconstruction method based on a priori guided diffusion model according to claim 1, characterized in that: In step 3, the forward process of the diffusion model is the low-dimensional latent variable That is, the discrete codebook vector Gradually add Gaussian noise, and finally pass Step 1: transform the data into Gaussian noise , Indicates that the mean is 0 and the covariance matrix is ​​the identity matrix I Normal distribution; The reverse process from Gaussian noise Departure, pass conditions Guide the adjustment of denoising direction and gradually denoise and reconstruct the latent variables .

9. The polarization 3D reconstruction method based on a priori guided diffusion model according to claim 8, characterized in that: In step 3, the learnable parameters of the codec are frozen, and the diffusion model is trained on the dataset using the variational lower bound loss function. Supervised optimization of this stage of training aims to minimize the noise of the U-Net network predictions With real noise The difference between , enables the diffusion model to learn to effectively denoise under conditional guidance: in, Express expectations, is the standard normally distributed noise, Indicates the Step data, Indicates conditional information. It is a U-Net network used to predict the added noise; Finally, the latent variables reconstructed by denoising the diffusion model The decoder converts it back to the original data space to obtain the final reconstruction result.

10. A polarization 3D reconstruction system based on a priori guided diffusion model, characterized by: The invention comprises a memory, a processor and a computer program stored in the memory and capable of running on the processor, wherein when the processor executes the program, a polarization three-dimensional reconstruction method based on a priori guided diffusion model as claimed in any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Three-dimensional CT skeleton diffusion reconstruction method and system based on two-dimensional X-ray

    CN120894458A

  • Method and system for three-dimensional ct bone diffusion reconstruction based on two-dimensional x-rays

    CN120894458B

  • Anisotropic MRI super-resolution method based on one-step diffusion model

    CN121120897A

  • Three-dimensional shape generation method

    CN121170200A

  • Discrete feature modeling method based on vector quantization variational auto-encoder and visual information generation system

    CN121259517A