High-resolution three-dimensional reconstruction method of fusion diffusion model

By employing a fusion diffusion model for 3D reconstruction, utilizing a text encoder and a conditional diffusion model, combined with TSDF weighted fusion, the problems of geometric boundary misalignment and color inconsistency during multi-view fusion are solved, achieving high-precision 3D modeling, which is particularly suitable for real-time digital twin applications.

CN120976443AActive Publication Date: 2025-11-18SHENZHEN SENSING DATA TECH CO LTD +1

Patent Information

Application Number
CN202511495376.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing 3D reconstruction methods suffer from geometric boundary misalignment and color inconsistency when fusing multiple views, resulting in obvious stitching artifacts. They are difficult to guide with high-level semantic information, and in complex scenes, the textures are blurred and the structures are inconsistent, leading to poor stability of the generated results and failing to meet the requirements of real-time or high-resolution reconstruction.

Method used

A high-resolution 3D reconstruction method using a fusion diffusion model is proposed. This method constructs a 3D reconstruction network, which includes a text encoder, a renderer, a VAE encoder, a conditional diffusion model, and an MVS module. The training process is divided into three stages. Semantic guidance and conditional diffusion sampling are introduced, and TSDF weighted fusion and local super-resolution fine-tuning are combined to optimize color consistency and depth consistency loss, thereby improving reconstruction quality.

Benefits of technology

It significantly improves the semantic consistency and detail fidelity of the reconstruction results, reduces stitching artifacts, enhances the stability of real-time rendering in complex scenes, and achieves high-precision 3D modeling, making it particularly suitable for real-time digital twin applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976443A_ABST
    Figure CN120976443A_ABST
Patent Text Reader

Abstract

The invention discloses a high-resolution three-dimensional reconstruction method of a fusion diffusion model, which belongs to the technical field of image data processing, and comprises the following steps: constructing an original data set D; constructing an enhanced training set; constructing a three-dimensional reconstruction network which comprises a text encoder, a renderer, a VAE encoder, a conditional diffusion model, a VAE decoder and an MVS module; training and fine-tuning the conditional diffusion model in three stages to obtain a three-dimensional reconstruction model, acquiring an image sequence and a text instruction of a scene to be reconstructed, and performing reconstruction by using the three-dimensional reconstruction model. According to the method, highly consistent geometric and color reduction can be kept under the multi-view condition, and splicing artifacts are remarkably reduced. Through semantic guidance optimization, texture details and structural consistency of the reconstruction model are greatly improved. Conditional diffusion sampling enables the model to accurately restore local details in a complex scene, and the stability of real-time rendering is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer modeling technology, and in particular to a high-resolution three-dimensional reconstruction method using a fusion diffusion model. Background Technology

[0002] With the rapid development of deep learning and diffusion models, diffusion-based 3D reconstruction optimization methods have shown great potential in fields such as 3D modeling, computer vision, and virtual reality. Ensuring point cloud accuracy is crucial during reconstruction, and real-time rendering and simulation also require high resolution and consistency to meet the needs of building high-precision environment models. However, existing technologies still have some drawbacks, such as: existing methods often exhibit geometric boundary misalignment and color inconsistency during multi-view fusion, leading to obvious stitching artifacts, especially in complex scenes. Existing methods struggle to utilize high-level semantic information for guidance, relying solely on pixel or feature alignment, which can easily result in blurred local textures or inconsistencies with the overall structure. Traditional rendering-based or MVS-based methods are prone to losing details under changes in lighting, occlusion, or noise, and the generated results have poor stability, making them unable to meet the demands of real-time or high-resolution reconstruction.

[0003] Definitions: The MVS (Multi-View Stereo) algorithm is a method for reconstructing 3D scenes using images from multiple different perspectives. By combining image data from multiple perspectives, the MVS algorithm can construct highly detailed 3D models.

[0004] A VAE (Variational Auto-Encoder) consists of an encoder and a decoder. The encoder probabilistically encodes the original input (image, speech, text sequence, etc.) into a latent space. At this point, the encoder output is not a single latent representation, but rather the probability distribution parameters (usually the mean and variance) in the latent space. This allows the model to sample in the latent space to generate new data points. Then, based on the reparameter technique, the latent variable z (also called the latent variable) is generated using the mean and variance. The decoder is symmetric to the encoder and is used to restore the sampled latent variable z (latent space) into data with the same shape as the original input (original data space).

[0005] Diffusion probabilistic models (Diffusion models) are a class of generative models based on a probabilistic denoising process. Their principle can be decomposed into forward diffusion and backward diffusion processes. Forward diffusion is a fixed, predefined Markov chain. It starts with clean raw data and gradually adds small amounts of Gaussian noise over a series of T time steps until the original data is almost completely submerged in noise, ultimately approximating a pure Gaussian noise distribution. The backward process involves training a neural network (conditional denoising network) to reverse the forward noise addition process. That is, given a noisy sample x... t Given the current time step t and condition c, the conditional denoising network needs to predict the corresponding previous sample x with slightly less noise. t−1 Or predict the movement from x in the forward process t−1 Change to x t The noise added during the process. Conditions can be introduced in the reverse process to guide the denoising network to denoise.

[0006] CLIP (Contrastive Language-Image Pretraining) is a cross-modal neural network model that maps images and text to the same semantic space through contrastive learning, achieving semantic association between images and text. It employs a dual-encoder architecture, where the image encoder converts the image into a vector representation, and the text encoder encodes the text description into a vector representation. These two vector representations are then mapped to a shared semantic space through a linear projection layer, and cosine similarity is used to measure the degree of image-text matching.

[0007] TSDF (Truncated Signed Distance Function) is a widely used representation method in computer vision and 3D reconstruction, especially when using volumetric representation. TSDF can efficiently represent scenes in 3D space, particularly in applications requiring large-scale scenes and real-time interaction. Summary of the Invention

[0008] The purpose of this invention is to provide a high-resolution 3D reconstruction method for a fusion diffusion model that solves problems such as significant model distortion caused by semantic consistency and loss of detail, geometric boundary misalignment and color inconsistency during multi-view fusion, and stitching artifacts.

[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows: a high-resolution three-dimensional reconstruction method for a fusion diffusion model, comprising the following steps; S1, construct the original dataset D; Collect multiple sets of samples, each set of samples including a real 3D model and text instructions, and combine all samples to form the original dataset D; S2, constructing an enhanced training set; Divide the training set D proportionally from D. train Data augmentation is performed on samples within the training set to obtain an augmented training set. , Where K is Total number of internal samples T (k) They are respectively Enhanced realistic 3D model and enhanced text commands for the k-th sample group; S3 constructs a 3D reconstruction network, including a text encoder, renderer, VAE encoder, conditional diffusion model, VAE decoder, and MVS module; The text encoder is used to input T. (k) Generate the corresponding semantic vector e T ; The renderer is used to... Rendered as a color image sequence I (k) and depth map D (k) sequence, , I (k) Composed of color images from N perspectives, For I (k) A color image from a mid-angle perspective n. for The corresponding depth map; The VAE encoder is used to press I (k) sampling weight w k from The corresponding latent variable z0 is generated by sampling in the middle; The conditional diffusion model is used to take the latent variable z0 as input and e T As a condition, a denoising latent variable is generated through a forward noise addition and reverse denoising process. ; The VAE decoder is used to... generate Reconstruction graph And reconstruct the graph from N perspectives. Constructing a sequence of reconstructed graphs ; The MVS module is used to... Generate reconstructed 3D model ; S4, training the 3D reconstruction network; Noise prediction loss based on conditional diffusion model, and The conditional diffusion model is trained using reconstruction loss, renderer color consistency loss, and depth consistency loss. Then, image quality and 3D reconstruction constraints are used to fine-tune the conditional diffusion model, resulting in the 3D reconstruction model M. ref ; S5, acquire the image sequence and text instructions of the scene to be reconstructed, and process them via M... ref The conditional diffusion model and VAE decoder obtain the corresponding reconstructed graph sequence, which is then processed by the MVS module to obtain the reconstructed 3D model.

[0010] Preferably, the renderer is a differentiable renderer. When rendering as a 2D image, first determine the projection plane based on the perspective of the 2D image. Find the vertex p, obtain its surface normal n(p), and perform Laplacian smoothing according to the following formula to obtain the smoothed surface normal n'(p). Replace n(p) with n'(p). , In the formula, N(p) is The set of vertices adjacent to the midpoint p, n(y) is the surface normal of vertex y in N(p), and β is the smoothing coefficient.

[0011] Preferably, the sampling weights are obtained according to the following formula: , In the formula, Color image sequence I (k) and reconstructed graph sequence Reconstruction loss between, Indicates to I (k) Find the gradient, where ε is a minimum value to prevent the denominator from being zero.

[0012] Preferably, the conditional diffusion model includes a forward diffusion process and a reverse diffusion process; The forward diffusion process is used to add noise to z0 step by step according to the following formula, and generate a noisy data at each time step, where the noisy data at time step t is z t ; , In the formula, α t ϵ is the noise scheduling coefficient that decreases with time step t, and ϵ is the standard Gaussian noise that conforms to the standard normal distribution N(0,1). The reverse diffusion process is used to obtain z according to the following formula. t The noise ϵ added during the forward diffusion process from time step t-1 to t is predicted. θ (z t ,t,e T And predict the noisy data z at time step t-1. t-1 ; , In the formula, The noise standard deviation coefficient for the reverse diffusion process. This is the second type of noise sampling.

[0013] As a preferred option, S4 is divided into three phases: Phase 1, Phase 2, and Phase 3. Phase 1: Training the Conditional Diffusion Model; Noise prediction loss L based on conditional diffusion model diff , and Reconstruction loss L recon Color consistency loss of the renderer L color and depth consistency loss L depth Construct a first loss function L1 to minimize the network parameters of the conditional diffusion model until the 3D reconstruction network converges. Label the converged 3D reconstruction network as M. S1 ; Phase Two: Fine-tuning of the image quality-guided conditional diffusion model; Based on L1, a local loss L is introduced for color images and reconstructed images. local and VGG perceived loss L perc Construct a second loss function L2 to minimize the L2 adjustment of the conditional diffusion model network parameters to M. S1 Convergence, and the converged M S1 Marked as M S2 ; Phase 3: Fine-tuning of the 3D reconstruction constraint-guided diffusion model; Based on L2, a photometric consistency error E is introduced for MVS module reconstruction. photo Construct a second loss function L3 to minimize the L3-adjusted conditional diffusion model network parameters to M. S2 Convergence yields the 3D reconstruction model M. ref .

[0014] Preferably, in S4, stage one includes steps Sa1~Sa2: Sa1, construct the first loss function L1; , In the formula, L diff For the noise prediction loss of the conditional diffusion model, L recon L represents the reconstruction loss between the color image sequence and the reconstructed image sequence. color For the color consistency loss of the renderer, L depth λ1, λ2, and λ3 are the depth consistency loss of the renderer, and the weights of the corresponding multiplication terms in L1 are respectively. Sa2, from training set Dtrain A batch of samples is fed into a 3D reconstruction network. L1 is calculated, and the network parameters of the conditional diffusion model are adjusted by minimizing L1 until the 3D reconstruction network converges. The converged 3D reconstruction network is then labeled as M. S1 .

[0015] Preferably, the color consistency loss L color and depth consistency loss L depth Calculate according to the following formulas respectively: , In the formula, g is The total number of mid vertices, for The g-th vertex p g I i (π i (p g )) is p g Projection position π in the color map of viewpoint i i (p g The color of ) I j (π j (p g )) is p g The projection position π in the color image at viewpoint j j (p g The color of ) , D i (u,v) represents the depth value at coordinates (u,v) in the depth map of viewpoint i. Let (u,v) be the reprojected coordinates of the depth map at viewpoint j. Let j be the coordinates in the depth map. The depth value.

[0016] Preferably, in S4, stage two includes steps Sb1~Sb2: Sb1, construct the second loss function L2; , In the formula L local For localized loss, L perc For VGG sensing loss, ~ These are the weights of the corresponding multiplication terms in L2; Sb2, using training set D train Training M S1 And adjust M to minimize L1 S1 The network parameters of the conditional diffusion model converge, yielding M. S2 .

[0017] As a preferred option, local loss Llocal VGG perceived loss L perc We obtain them respectively from the following formulas: , for Detailed areas, I n ( )for medium pixel pixel values, for medium pixel Pixel values; , In the formula, ϕ(⋅) is the feature extraction function, and G SR (⋅) represents a super-resolution network. It is the square of the L2 norm.

[0018] As a preferred option, the photometric consistency error E photo Calculate according to the following formula; , In the formula, for M new Inner vertex p, I i (π i (p) represents the projection position π of p in the color image at viewpoint i. i The color of (p), I j (π j (p) represents the projection position π of p in the color image at viewpoint j. j The color of (p); It is an L1 norm; The MVS module is used to... Generate reconstructed 3D model The method is as follows; The MVS module generates the initial reconstruction model M new Then M new and The TSDF values ​​are fused according to the fusion weights to form the reconstructed 3D model. ; , , In the formula, for M new Let p be an interior vertex, and x be the nearest observed point corresponding to p. for The TSDF value of vertex p. , They are respectively and M new The TSDF value of vertex p, w gt wnew They are respectively and M new The weight of M, δ is the control M new The velocity decay factor of the weight, exp(∙) is the exp function.

[0019] Regarding the dataset: After constructing the original dataset D, this invention divides it into a training set and performs data augmentation to obtain the augmented training set. Used for subsequent model training to improve the robustness of the 3D reconstruction network.

[0020] Regarding renderers: Renderers are used to... When rendering a 2D image, The surface normal n(p) of vertex p is smoothed using Laplacian smoothing to obtain the smoothed surface normal n'(p) which replaces n(p), thereby reducing artifacts introduced by discrete meshes and rendering, and evenly distributing the normal direction to make the lighting transition more natural.

[0021] Regarding the conditional diffusion model: The latent variable z0 is used as input, and e... T As a condition, and with the present invention focusing on training the conditional diffusion model, the training is divided into three stages.

[0022] Phase 1 is used to train the conditional diffusion model. At this stage, the MVS module in the 3D reconstruction network does not participate, and the text instructions T in the samples... (k) The semantic vectors generated by the text encoder will be rendered by the renderer. Rendered as a color image sequence I (k) and depth map D (k) The sequence is then processed by a VAE encoder for I. (k) The process involves color image sampling, conditional diffusion model denoising, denoising, and VAE decoder decoding to generate a reconstructed image corresponding to the color image. In the loss function part, stage one omits the noise prediction loss L of the conditional diffusion model itself. diff Also considering and Reconstruction loss L recon and cross-view color consistency loss L color and depth consistency loss L depth ,get Then, the network parameters of the conditional diffusion model are updated using an optimization algorithm until the conditional diffusion model converges. This stage mainly involves learning to generate multi-view reconstruction graph sequences that meet the text conditions from noise.

[0023] Phase two involves fine-tuning the conditional diffusion model based on image quality guidance. Building upon Phase one, a local loss L is also introduced. local and VGG perceived loss L perc L localPreserve details in small areas (such as edges and textures); L perc High-level features are extracted using pre-trained networks (such as VGG19) to ensure overall semantics and clarity. The MVS module does not participate in this stage and mainly fine-tunes the quality of the reconstructed graph to obtain a higher quality reconstructed graph sequence.

[0024] Phase 3 involves fine-tuning the conditional diffusion model based on 3D reconstruction constraints. Building upon Phase 2, the MVS module then participates, performing 3D reconstruction on the reconstructed image sequence output from Phase 2 and outputting an initial reconstruction model M. new In the process of 3D reconstruction, a photometric consistency error E is introduced. photo To constrain and ensure the consistency of reprojection from different viewpoints, thereby ensuring M new The accuracy of geometric reconstruction, and in M new and A fusion weight is introduced between them to smooth the transition and retain the advantages of both.

[0025] Compared with the prior art, the advantages of the present invention are as follows: (1) This invention introduces semantic guidance in multi-view conditional diffusion to naturally fill in the missing geometric and texture areas, significantly improving the semantic consistency and detail fidelity of the reconstruction results. As a result, the texture details and structural consistency of the reconstructed 3D model are greatly improved.

[0026] (2) Conditional diffusion sampling enables the model to accurately reproduce local details in complex scenes, improving the stability of real-time rendering. Based on the conditional diffusion model, a color consistency loss L is introduced for the renderer. color Deep consistency loss L depth For both the color image and the reconstructed image, a local loss L is introduced. local and VGG perceived loss L perc This allows for highly consistent geometry and color reproduction under multi-view conditions, significantly reducing stitching artifacts.

[0027] (3) The present invention divides the training into three stages. Through the three-level reconstruction strategy of coarse, medium and fine, combined with TSDF weighted fusion and local super-resolution fine adjustment, it not only achieves smooth and continuous overall geometry, but also accurately corrects early errors and restores high-frequency details.

[0028] In summary, this invention significantly improves the quality and application value of multi-view 3D reconstruction by introducing semantic guidance and conditional diffusion sampling, providing an innovative solution for computer vision and virtual reality. By addressing the issues of semantic consistency and loss of detail, this method offers a new technical path for high-precision 3D modeling, with particularly broad application prospects in real-time digital twins. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the three-dimensional reconstruction network of the present invention; Figure 2 This is a flowchart of the present invention. Detailed Implementation

[0030] The present invention will be further described below with reference to the embodiments and accompanying drawings.

[0031] Example 1: See Figure 1 and Figure 2 A high-resolution 3D reconstruction method based on a fusion diffusion model includes the following steps; S1, construct the original dataset D; Collect multiple sets of samples, each set of samples including a real 3D model and text instructions, and combine all samples to form the original dataset D; S2, constructing an enhanced training set; Divide the training set D proportionally from D. train Data augmentation is performed on samples within the training set to obtain an augmented training set. , Where K is Total number of internal samples T (k) They are respectively Enhanced realistic 3D model and enhanced text commands for the k-th sample group; S3 constructs a 3D reconstruction network, including a text encoder, renderer, VAE encoder, conditional diffusion model, VAE decoder, and MVS module; The text encoder is used to input T. (k) Generate the corresponding semantic vector e T ; The renderer is used to... Rendered as a color image sequence I (k) and depth map D (k) sequence, , I (k) Composed of color images from N perspectives, For I (k) A color image from a mid-angle perspective n. for The corresponding depth map; The VAE encoder is used to press I (k) sampling weight w k from The corresponding latent variable z0 is generated by sampling in the middle; The conditional diffusion model is used to take the latent variable z0 as input and e T As a condition, a denoising latent variable is generated through a forward noise addition and reverse denoising process. ; The VAE decoder is used to... generate Reconstruction graph And reconstruct the graph from N perspectives. Constructing a sequence of reconstructed graphs ; The MVS module is used to... Generate reconstructed 3D model ; S4, training the 3D reconstruction network; Phase 1: Training the Conditional Diffusion Model; Noise prediction loss L based on conditional diffusion model diff , and Reconstruction loss L recon Color consistency loss of the renderer L color and depth consistency loss L depth Construct a first loss function L1 to minimize the network parameters of the conditional diffusion model until the 3D reconstruction network converges. Label the converged 3D reconstruction network as M. S1 ; Phase Two: Fine-tuning of the image quality-guided conditional diffusion model; Based on L1, a local loss L is introduced for color images and reconstructed images. local and VGG perceived loss L perc Construct a second loss function L2 to minimize the L2 adjustment of the conditional diffusion model network parameters to M. S1 Convergence, and the converged M S1 Marked as M S2 ; Phase 3: Fine-tuning of the 3D reconstruction constraint-guided diffusion model; Based on L2, a photometric consistency error E is introduced for MVS module reconstruction. photo Construct a second loss function L3 to minimize the L3-adjusted conditional diffusion model network parameters to M. S2 Convergence yields the 3D reconstruction model M. ref ; S5, acquire the image sequence and text instructions of the scene to be reconstructed, and process them via M... ref The conditional diffusion model and VAE decoder obtain the corresponding reconstructed graph sequence, which is then processed by the MVS module to obtain the reconstructed 3D model.

[0032] Regarding text commands, they can be set to describe the actual 3D model, for example: Text instruction 1: A wooden chair with a backrest, brown in color; Text instruction 2: A red two-door sports car with smooth body lines; Text instruction 3: A white, modern office building with a glass curtain wall facade.

[0033] Based on the method of this embodiment, during training, press I. (k) sampling weight w k from The intermediate sample is used as the latent variable z0, and noise is added step by step to obtain z. t The conditional diffusion model aims to learn how to diffuse under given conditions e. T Next, gradually restore the noise to its original state. The resulting image is a reconstructed image.

[0034] During inference, the conditional diffusion model starts from a pure noise vector. Initially, a reverse denoising process is performed step-by-step to obtain the reconstructed image. At this point, noise sampling is independent of... Instead, it is free sampling, and the guidance of the generation process depends on T. (k) The corresponding semantic vector e T .

[0035] Example 2: See Figures 1 to 2 More specifically, based on Example 1, the renderer is a differentiable renderer, which will... When rendering as a 2D image, first determine the projection plane based on the perspective of the 2D image. Find the vertex p, obtain its surface normal n(p), and perform Laplacian smoothing according to the following formula to obtain the smoothed surface normal n'(p). Replace n(p) with n'(p). , In the formula, N(p) is The set of vertices adjacent to the midpoint p, n(y) is the surface normal of vertex y in N(p), and β is the smoothing coefficient.

[0036] The sampling weights are obtained according to the following formula: , In the formula, Color image sequence I (k) and reconstructed graph sequence Reconstruction loss between, Indicates to I (k) Find the gradient, where ε is a minimum value to prevent the denominator from being zero.

[0037] The conditional diffusion model includes a forward diffusion process and a reverse diffusion process; The forward diffusion process is used to add noise to z0 step by step according to the following formula, and generate a noisy data at each time step, where the noisy data at time step t is z t; , In the formula, α t ϵ is the noise scheduling coefficient that decreases with time step t, and ϵ is the standard Gaussian noise that conforms to the standard normal distribution N(0,1). The reverse diffusion process is used to obtain z according to the following formula. t The noise ϵ added during the forward diffusion process from time step t-1 to t is predicted. θ (z t ,t,e T And predict the noisy data z at time step t-1. t-1 ; , In the formula, The noise standard deviation coefficient for the reverse diffusion process. This is the second type of noise sampling.

[0038] Example 3: See Figure 1 , Figure 2 Based on Example 1, we provide the specific implementation steps of S4. The rest is the same as in Example 1 or Example 2. In S4, stage one includes steps Sa1 to Sa2: Sa1, construct the first loss function L1; , In the formula, L diff For the noise prediction loss of the conditional diffusion model, L recon L represents the reconstruction loss between the color image sequence and the reconstructed image sequence. color For the color consistency loss of the renderer, L depth λ1, λ2, and λ3 are the depth consistency loss of the renderer, and the weights of the corresponding multiplication terms in L1 are respectively. The color consistency loss L color and depth consistency loss L depth Calculate according to the following formulas respectively: , In the formula, g is The total number of mid vertices, for The g-th vertex p g I i (π i (p g )) is p g Projection position π in the color map of viewpoint i i (p g The color of ) I j (π j (p g )) is p gThe projection position π in the color image at viewpoint j j (p g The color of ) , D i (u,v) represents the depth value at coordinates (u,v) in the depth map of viewpoint i. Let (u,v) be the reprojected coordinates of the depth map at viewpoint j. Let j be the coordinates in the depth map. The depth value; Sa2, from training set D train A batch of samples is fed into a 3D reconstruction network. L1 is calculated, and the network parameters of the conditional diffusion model are adjusted by minimizing L1 until the 3D reconstruction network converges. The converged 3D reconstruction network is then labeled as M. S1 .

[0039] Phase 2 includes steps Sb1~Sb2: Sb1, construct the second loss function L2; , In the formula L local For localized loss, L perc For VGG sensing loss, ~ These are the weights of the corresponding multiplication terms in L2; Local loss L local VGG perceived loss L perc We obtain them respectively from the following formulas: , for Detailed areas, I n ( )for medium pixel pixel values, for medium pixel Pixel values; , In the formula, ϕ(⋅) is the feature extraction function, and G SR (⋅) represents a super-resolution network. The square of the L2 norm; Sb2, using training set D train Training M S1 And adjust M to minimize L1 S1 The network parameters of the conditional diffusion model converge, yielding M. S2 .

[0040] Phase three specifically refers to L2 already containing L diff L recon L depth L local L perc Based on this, increase the photometric consistency error E photo We obtain L3, and then redistribute the weights of these 6 parts within L3. During training, we adjust the parameters of the conditional diffusion model network to M by minimizing L3. S2 Convergence yields the 3D reconstruction model M. ref .

[0041] Photometric consistency error E photo Calculate according to the following formula; , In the formula, for M new Inner vertex p, I i (π i (p) represents the projection position π of p in the color image at viewpoint i. i The color of (p), I j (π j (p) represents the projection position π of p in the color image at viewpoint j. j The color of (p); It is an L1 norm; Additionally, during model training, the MVS module is used to... Generate reconstructed 3D model The specific method is as follows: The MVS module generates the initial reconstruction model M new Then M new and The TSDF values ​​are fused according to the fusion weights to form the reconstructed 3D model. ; , , In the formula, for M new Let p be an interior vertex, and x be the nearest observed point corresponding to p. for The TSDF value of vertex p. , They are respectively and M new The TSDF value of vertex p, w gt w new They are respectively and M new The weight of M, δ is the control M new The velocity decay factor of the weight, exp(∙) is the exp function.

[0042] During model inference, no Then there is no need to merge.

[0043] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A high-resolution 3D reconstruction method using a fusion diffusion model, characterized in that, Includes the following steps; S1, construct the original dataset D; Collect multiple sets of samples, each set of samples including a real 3D model and text instructions, and combine all samples to form the original dataset D; S2, constructing an enhanced training set; Divide the training set D proportionally from D. train Data augmentation is performed on samples within the training set to obtain an augmented training set. , Where K is Total number of internal samples T (k) They are respectively Enhanced realistic 3D model and enhanced text commands for the k-th sample group; S3 constructs a 3D reconstruction network, including a text encoder, renderer, VAE encoder, conditional diffusion model, VAE decoder, and MVS module; The text encoder is used to input T. (k) Generate the corresponding semantic vector e T ; The renderer is used to... Rendered as a color image sequence I (k) and depth map D (k) sequence, , I (k) Composed of color images from N perspectives, For I (k) A color image from a mid-angle perspective n. for The corresponding depth map; The VAE encoder is used to press I (k) sampling weight w k from The corresponding latent variable z0 is generated by sampling in the middle; The conditional diffusion model is used to take the latent variable z0 as input and e T As a condition, a denoising latent variable is generated through a forward noise addition and reverse denoising process. ; The VAE decoder is used to... generate Reconstruction graph And reconstruct the graph from N perspectives. Constructing a sequence of reconstructed graphs ; The MVS module is used to... Generate reconstructed 3D model ; S4, training the 3D reconstruction network; Noise prediction loss based on conditional diffusion model, and The conditional diffusion model is trained using reconstruction loss, renderer color consistency loss, and depth consistency loss. Then, image quality and 3D reconstruction constraints are used to fine-tune the conditional diffusion model, resulting in the 3D reconstruction model M. ref ; S5, acquire the image sequence and text instructions of the scene to be reconstructed, and process them via M... ref The conditional diffusion model and VAE decoder obtain the corresponding reconstructed graph sequence, which is then processed by the MVS module to obtain the reconstructed 3D model.

2. The high-resolution three-dimensional reconstruction method based on a fusion diffusion model according to claim 1, characterized in that, The renderer is a differentiable renderer. When rendering as a 2D image, first determine the projection plane based on the perspective of the 2D image. Find the vertex p, obtain its surface normal n(p), and perform Laplacian smoothing according to the following formula to obtain the smoothed surface normal n'(p). Replace n(p) with n'(p). , In the formula, N(p) is The set of vertices adjacent to the midpoint p, n(y) is the surface normal of vertex y in N(p), and β is the smoothing coefficient.

3. The high-resolution three-dimensional reconstruction method based on a fusion diffusion model according to claim 1, characterized in that, The sampling weights are obtained according to the following formula: , In the formula, Color image sequence I (k) and reconstructed graph sequence Reconstruction loss between, Indicates to I (k) Find the gradient, where ε is a minimum value to prevent the denominator from being zero.

4. The high-resolution three-dimensional reconstruction method based on a fusion diffusion model according to claim 1, characterized in that, The conditional diffusion model includes a forward diffusion process and a reverse diffusion process; The forward diffusion process is used to add noise to z0 step by step according to the following formula, and generate a noisy data at each time step, where the noisy data at time step t is z t ; , In the formula, α t ϵ is the noise scheduling coefficient that decreases with time step t, and ϵ is the standard Gaussian noise that conforms to the standard normal distribution N(0,1). The reverse diffusion process is used to obtain z according to the following formula. t The noise ϵ added during the forward diffusion process from time step t-1 to t is predicted. θ (z t ,t,e T And predict the noisy data z at time step t-1. t-1 ; , In the formula, The noise standard deviation coefficient for the reverse diffusion process. This is the second type of noise sampling.

5. The high-resolution three-dimensional reconstruction method of the fusion diffusion model according to claim 1, characterized in that, S4 is divided into three phases: Phase 1, Phase 2, and Phase 3. Phase 1: Training the Conditional Diffusion Model; Noise prediction loss L based on conditional diffusion model diff , and Reconstruction loss L recon Color consistency loss of the renderer L color and depth consistency loss L depth Construct a first loss function L1 to minimize the network parameters of the conditional diffusion model until the 3D reconstruction network converges. Label the converged 3D reconstruction network as M. S1 ; Phase Two: Fine-tuning of the image quality-guided conditional diffusion model; Based on L1, a local loss L is introduced for color images and reconstructed images. local and VGG perceived loss L perc Construct a second loss function L2 to minimize the L2 adjustment of the conditional diffusion model network parameters to M. S1 Convergence, and the converged M S1 Marked as M S2 ; Phase 3: Fine-tuning of the 3D reconstruction constraint-guided diffusion model; Based on L2, a photometric consistency error E is introduced for MVS module reconstruction. photo Construct a second loss function L3 to minimize the L3-adjusted conditional diffusion model network parameters to M. S2 Convergence yields the 3D reconstruction model M. ref .

6. The high-resolution three-dimensional reconstruction method of the fusion diffusion model according to claim 5, characterized in that, In S4, stage one includes steps Sa1~Sa2: Sa1, construct the first loss function L1; , In the formula, L diff For the noise prediction loss of the conditional diffusion model, L recon L represents the reconstruction loss between the color image sequence and the reconstructed image sequence. color For the color consistency loss of the renderer, L depth λ1, λ2, and λ3 are the depth consistency loss of the renderer, and the weights of the corresponding multiplication terms in L1 are respectively. Sa2, from training set D train A batch of samples is fed into a 3D reconstruction network. L1 is calculated, and the network parameters of the conditional diffusion model are adjusted by minimizing L1 until the 3D reconstruction network converges. The converged 3D reconstruction network is then labeled as M. S1 .

7. The high-resolution three-dimensional reconstruction method based on a fusion diffusion model according to claim 6, characterized in that, The color consistency loss L color and depth consistency loss L depth Calculate according to the following formulas respectively: , In the formula, g is The total number of mid vertices, for The g-th vertex p g I i (π i (p g )) is p g Projection position π in the color map of viewpoint i i (p g The color of ) I j (π j (p g )) is p g The projection position π in the color image at viewpoint j j (p g The color of ) , D i (u,v) represents the depth value at coordinates (u,v) in the depth map of viewpoint i. Let (u,v) be the reprojected coordinates of the depth map at viewpoint j. Let j be the coordinates in the depth map. The depth value.

8. The high-resolution three-dimensional reconstruction method of the fusion diffusion model according to claim 6, characterized in that, In S4, phase two includes steps Sb1~Sb2: Sb1, construct the second loss function L2; , In the formula L local For local loss, L perc For VGG sensing loss, ~ These are the weights of the corresponding multiplication terms in L2; Sb2, using training set D train Training M S1 And adjust M to minimize L1 S1 The network parameters of the conditional diffusion model converge, yielding M. S2 .

9. The high-resolution three-dimensional reconstruction method of the fusion diffusion model according to claim 8, characterized in that, Local loss L local VGG perceived loss L perc We obtain them respectively from the following formulas: , for Detailed areas, I n ( )for medium pixel pixel values, for medium pixel Pixel values; , In the formula, ϕ(⋅) is the feature extraction function, and G SR (⋅) represents a super-resolution network. It is the square of the L2 norm.

10. The high-resolution three-dimensional reconstruction method of the fusion diffusion model according to claim 5, characterized in that, Photometric consistency error E photo Calculate according to the following formula; , In the formula, for M new Inner vertex p, I i (π i (p) represents the projection position π of p in the color image at viewpoint i. i The color of (p), I j (π j (p) represents the projection position π of p in the color image at viewpoint j. j The color of (p); It is an L1 norm; The MVS module is used to... Generate reconstructed 3D model The method is as follows; The MVS module generates the initial reconstruction model M new Then M new and The TSDF values ​​are fused according to the fusion weights to form the reconstructed 3D model. ; , , In the formula, for M new Let p be an interior vertex, and x be the nearest observed point corresponding to p. for The TSDF value of vertex p. , They are respectively and M new The TSDF value of vertex p, w gt w new They are respectively and M new The weight of M, δ is the control M new The velocity decay factor of the weight, exp(∙) is the exp function.

Citation Information

Patent Citations

  • Diffusion-guided three-dimensional reconstruction

    CN118918241A

  • Diffusion-guided three-dimensional reconstruction

    US20240412458A1

  • Diffusion model and generative adversarial network-based image segmentation method and apparatus

    WO2025050542A1

Cited By

  • Semantic-driven image reconstruction method and device based on edge and color assistance

    CN121239857A

  • Image-text guided multi-modal feature driven three-dimensional model generation method

    CN121600194A

  • A method for generating a three-dimensional model driven by multi-modal features guided by text

    CN121600194B

  • Multi-view three-dimensional reconstruction method based on video diffusion model

    CN122066871A

  • A Multi-View 3D Reconstruction Method Based on Video Diffusion Model

    CN122066871B