Diffusion model driven robust three-dimensional shape recovery method

By using a diffusion model-driven approach, combined with massive generative priors and a self-attention mechanism, the robustness and efficiency issues of 3D topography restoration in existing technologies are addressed, achieving efficient 3D topography reconstruction in low-contrast and noisy environments.

CN122049271APending Publication Date: 2026-05-15SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANXI UNIV
Filing Date
2026-02-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing focused topography restoration techniques show reduced robustness when dealing with low contrast, lighting variations, and noise interference, and data-driven methods have limited generalization ability in real-world scenarios, making it difficult to accurately reconstruct 3D topography.

Method used

A diffusion model-driven approach is adopted, which introduces massive generative priors and image-depth joint latent guidance mechanism, combined with cross-focal plane self-attention mechanism, and utilizes the learning rate warm-up and adaptive decay mechanism in the diffusion model to quickly capture global physical distribution and refine continuous gradient of morphology, thereby realizing 3D morphology reconstruction.

Benefits of technology

It effectively solves the problems of depth jumps and detail loss in low-contrast regions in traditional methods, improves the spatial coherence and computational efficiency of 3D topography restoration, and realizes efficient real-time processing of large-scale focal stack data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049271A_ABST
    Figure CN122049271A_ABST
Patent Text Reader

Abstract

The invention relates to a robust three-dimensional shape recovery method based on a diffusion model, and belongs to the field of computer vision. The method comprises the following steps: firstly, carrying out normalization preprocessing on a collected multi-focus image sequence, and mapping the multi-focus image sequence to a potential space through an encoder; then, the potential feature sequence and the deep potential representation noise are spliced, and a comprehensive guide condition is constructed in combination with focusing prior and time embedding; then, a denoising network integrated with a cross-focal plane self-attention module is input, and clear focal plane information is adaptively aggregated to predict a diffusion perturbation term. And finally, obtaining the denoised depth potential representation through inverse diffusion iterative sampling, and obtaining an absolute depth map through decoding and inverse normalization processing. According to the method, by fusing diffusion model generation prior and focusing imaging physical characteristics, the problems of spatial ambiguity and morphology discontinuity are effectively relieved, the precision and robustness of microscopic scene three-dimensional reconstruction are improved, and the method is suitable for applications such as handheld camera three-dimensional rendering and precision part surface detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to focused shape restoration in the field of computer vision, specifically involving a robust 3D shape restoration method driven by a diffusion model. Background Technology

[0002] Focused topography restoration is a technique for 3D reconstruction using acquired multi-focus image sequences. Its basic principle involves analyzing multi-focus image sequences acquired by an imaging system at different focal plane positions, and using focus evaluation operators to assess the focusing behavior of pixels as a function of depth, thereby recovering the surface topography of the object under test. This method offers advantages such as multi-scene adaptability, low hardware cost, and ease of achieving high-resolution imaging, making it widely applicable in fields such as 3D rendering with handheld cameras and surface inspection of precision components.

[0003] The rapid development of intelligent manufacturing technology has placed increasingly higher demands on measurement accuracy in the manufacturing field. In shape restoration methods, modeling-based methods mainly rely on manually designed focus evaluation operators. These operators assess the sharpness level of local regions and then obtain depth information based on the location index of the sharpest region. However, these methods are sensitive to environmental noise and often struggle to accurately determine focus information when dealing with low-contrast areas or complex textured surfaces, thus affecting the reliability of the restoration results. Restoring 3D depth information from a 2D focused image sequence is a typical ill-posed problem, with multiple solutions exhibiting ambiguity in its solution space. Especially when facing imaging scenarios such as low-contrast textures and varying exposure conditions, the local features provided by the image are limited, leading to the possibility of multiple solutions satisfying physical imaging constraints. This further increases the difficulty of accurately restoring the 3D shape. To address this ill-posed problem, introducing rich visual prior knowledge to effectively constrain the solution space has become an important approach to improving the robustness and accuracy of the method. While existing methods have attempted to leverage data-driven strategies to train models with large datasets to learn deep mapping relationships, their generalization ability is often limited by the scale and diversity of labeled data. The learned mapping relationships are often relatively simple and fail to cover complex and varied real-world scenarios. In recent years, models based on probabilistic generation have shown significant advantages in signal denoising and data generation. By simulating the noise removal process in data distribution, they can concentrate energy on effective regions on high-dimensional manifolds, making them theoretically particularly suitable for handling noisy open-scene 3D reconstruction tasks. However, there are currently no precedents for the direct application of such models to solving 3D reconstruction problems, and their specific implementation and performance optimization in shape restoration still require further exploration.

[0004] In summary, the main challenges facing existing focused topography reconstruction techniques are: modeling methods cannot effectively overcome the reduced robustness caused by low contrast, lighting variations, and noise interference in real-world scenes; data-driven methods can only directly learn the mapping relationship between multi-focus images and 3D topography results, and cannot effectively reconstruct high-noise or depth-of-field missing conditions in real-world scenes; therefore, how to utilize the learning rate warm-up and adaptive decay mechanism in diffusion models to quickly capture the global physical distribution correlation between sequences in the early training stage, achieve accurate fitting of continuous topography gradients in the later refinement stage, and ultimately achieve efficient fine-tuning and adaptation of pre-trained knowledge to the focused topography reconstruction task remains a technical problem to be solved. Summary of the Invention

[0005] To address the technical shortcomings of well-posed topography restoration, such as solution space ambiguity and topography discontinuities caused by ill-posedness, this invention provides a robust 3D topography restoration method driven by a diffusion model, comprising the following steps:

[0006] Step 1: Acquire multi-focus image sequences of microscopic scenes ,in Indicates the total number of image sequences. This represents the position of an image sequence under a certain focal plane, and the size of a single image is [value missing]. , Determine the physical focal length range based on the image's height, width, and number of channels. ;

[0007] Step 2: The multifocus image sequence obtained in Step 1... The data is input into the data preprocessing module via equation (1) and used in conjunction with preset distribution adjustment parameters. Perform normalization processing to obtain a normalized image sequence. ,

[0008]

[0009] in It is a pixel offset constant. The pixel scaling factor. This is a zero-centralized bias term;

[0010] Step 3: Normalize the image sequence obtained in Step 2. Input to the preset encoder module via equation (2) Perform latent spatial feature encoding to obtain latent feature sequences. ,

[0011]

[0012] in For predefined potential spatial distribution adjustment factors, These represent the height, width, and number of channels of the potential feature, respectively.

[0013] Step 4: Calculate the latent feature sequence obtained in Step 3. The input is fed into the feature stitching module via equation (3) and combined with the depth latent characterization noise generated at the current moment. Execution channel dimension concatenation to construct a joint guidance feature sequence ,

[0014]

[0015] Meanwhile, the multifocus sequence recorded in step 1 Obtaining focused prior condition features With equation (4) and diffusion time step The temporal conditional features obtained by sinusoidal position encoding transformation To achieve interactive integration and build comprehensive guidance conditions ,

[0016]

[0017] in This represents the concatenation, weighted fusion, and interactive attention operations of conditional information.

[0018] Step 5: Combine the joint guiding feature sequence obtained in Step 4 With comprehensive guidance conditions The input is fed into the preset denoising network via equation (5). It is used to predict perturbation terms in the latent characterization of injection depth.

[0019]

[0020] Feature enhancement is performed using its internally integrated cross-focal-plane self-attention module, and then the features are reorganized into a vertical feature sequence centered on the pixel spatial location using Equation (6). Then, the correlation weights between different focal planes are calculated to achieve adaptive aggregation of sharp features.

[0021]

[0022] in, They are the characteristic sequences. The query, key, and value matrix obtained through linear transformation The proportional factor for the head of attention;

[0023] Step 6: Calculate the denoising network prediction perturbation term obtained in Step 5. The equation (7) is input into the inverse diffusion iterative sampling module, and is used in conjunction with a preset noise scheduling accumulation coefficient. Perform sampling update, from to Update the deep latent representation frame by frame until a fully denoised deep latent representation is obtained. ,

[0024]

[0025] Step 7: Apply the fully denoised deep latent representation obtained in Step 6 The input is given by equation (8) to the preset decoder module. Perform a reconstruction from the latent space to the physical pixel space to obtain the initial depth prediction results. ,

[0026]

[0027] Step 8: Calculate the prediction results obtained in Step 7. The equation (9) is input into the numerical inverse regression normalization module, and used in conjunction with the physical focal length range recorded in step 1. Perform a linear affine transformation, and then map the depth map from the normalized numerical range back to the true physical range to obtain the absolute depth map of the physical space. ,

[0028]

[0029] Clip is the clipping function. Finally, range truncation is performed on the output to ensure that the depth values ​​of all pixels are within the preset physical effective range, thus obtaining the final 3D shape reconstruction result. .

[0030] Compared with the prior art, the present invention has the following advantages:

[0031] (1) This invention provides a strong geometric constraint for the typical ill-posed problem of focusing on topography restoration by introducing a large amount of generative priors implied by the diffusion model. In addition, it effectively narrows the range of solution space that satisfies physical imaging constraints by using the image-depth joint potential guidance mechanism and the optimal solution search process in the probability space. Combined with the cross-focal plane self-attention mechanism, it can extract long-range dependencies from the pixel-level longitudinal evolution trajectory, so that the generated depth field has excellent spatial coherence. It effectively solves the problem of low-contrast region depth jump and detail loss caused by the reliance on manual evaluation operators in traditional methods.

[0032] (2) The present invention adopts a generative reconstruction framework based on latent space, which maps the complex high-dimensional multi-focus sequence calculation to a compact low-dimensional manifold space, greatly reducing the redundant computational load and system resource overhead in the data processing process. In addition, a fast denoising sampling strategy is adopted, which achieves cross-order-of-magnitude compression of the inference cycle while maintaining the accuracy of morphology restoration, and solves the technical bottleneck of traditional diffusion models in processing large-scale focus stack data in real time. Attached Figure Description

[0033] Figure 1 A flowchart of a robust 3D topography recovery method driven by a diffusion model;

[0034] Figure 2 A network diagram illustrating a robust 3D topography recovery method driven by a diffusion model;

[0035] Figure 3 This refers to a sequence of 30 multi-focus images with different focal planes collected in step 1 of Embodiment 1 of the present invention.

[0036] Figure 4 This is the depth map finally calculated in step 8 of embodiment 1 of the present invention;

[0037] Figure 5 This is a three-dimensional structural diagram obtained in step 8 of embodiment 1 of the present invention;

[0038] Figure 6 This is a 3D structure diagram obtained by other methods. Detailed Implementation

[0040] Example 1

[0041] like Figure 1 , Figure 2 As shown, a robust 3D topography restoration method driven by a diffusion model includes the following steps:

[0042] Step 1: Acquire multi-focus image sequences of microscopic scenes ,in Indicates the total number of image sequences. This represents the position of an image sequence under a certain focal plane, and the size of a single image is [value missing]. , Determine the physical focal length range based on the image's height, width, and number of channels. In the embodiments, such as Figure 3 As shown, the sequence number Image high Image width and image channels ;

[0043] Step 2: The multifocus image sequence obtained in Step 1... The data is input into the data preprocessing module via equation (1) and used in conjunction with preset distribution adjustment parameters. Perform normalization processing to obtain a normalized image sequence. ,

[0044]

[0045] in It is a pixel offset constant. The pixel scaling factor. This is a zero-centralized bias term;

[0046] Step 3: Normalize the image sequence obtained in Step 2. Input to the preset encoder module via equation (2) Perform latent spatial feature encoding to obtain latent feature sequences. ,

[0047]

[0048] in For predefined potential spatial distribution adjustment factors, These represent the height, width, and number of channels of the potential feature, respectively.

[0049] Step 4: Calculate the latent feature sequence obtained in Step 3. The input is fed into the feature stitching module via equation (3) and combined with the depth latent characterization noise generated at the current moment. Execution channel dimension concatenation to construct a joint guidance feature sequence ,

[0050]

[0051] Meanwhile, the multifocus sequence recorded in step 1 Obtaining focused prior condition features With equation (4) and diffusion time step The temporal conditional features obtained by sinusoidal position encoding transformation To achieve interactive integration and build comprehensive guidance conditions ,

[0052]

[0053] in This represents the concatenation, weighted fusion, and interactive attention operations of conditional information.

[0054] Step 5: Combine the joint guiding feature sequence obtained in Step 4 With comprehensive guidance conditions The input is fed into the preset denoising network via equation (5). It is used to predict perturbation terms in the latent characterization of injection depth.

[0055]

[0056] Feature enhancement is performed using its internally integrated cross-focal-plane self-attention module, and then the features are reorganized into a vertical feature sequence centered on the pixel spatial location using Equation (6). Then, the correlation weights between different focal planes are calculated to achieve adaptive aggregation of sharp features.

[0057]

[0058] in, They are the characteristic sequences. The query, key, and value matrix obtained through linear transformation The proportional factor for the head of attention;

[0059] Step 6: Calculate the denoising network prediction perturbation term obtained in Step 5. The equation (7) is input into the inverse diffusion iterative sampling module, and is used in conjunction with a preset noise scheduling accumulation coefficient. Perform sampling update, from to Update the deep latent representation frame by frame until a fully denoised deep latent representation is obtained. ,

[0060]

[0061] Step 7: Apply the fully denoised deep latent representation obtained in Step 6 The input is given by equation (8) to the preset decoder module. Perform a reconstruction from the latent space to the physical pixel space to obtain the initial depth prediction results. ,

[0062]

[0063] Step 8: Calculate the prediction results obtained in Step 7. The equation (9) is input into the numerical inverse regression normalization module, and used in conjunction with the physical focal length range recorded in step 1. Perform a linear affine transformation, then map the depth map from the normalized numerical range back to the true physical range to obtain... Figure 4 Absolute depth map of physical space ,

[0064]

[0065] Clip is the clipping function. Finally, range truncation is performed on the output to ensure that the depth values ​​of all pixels are within the preset physical valid range, resulting in the following: Figure 5 The final three-dimensional topography reconstruction result .

Claims

1. A diffusion model-driven robust 3D topography restoration method, comprising the following steps: Step 1: Acquire multi-focus image sequences of microscopic scenes ,in Indicates the total number of image sequences. This represents the position of an image sequence under a certain focal plane, and the size of a single image is [value missing]. , Determine the physical focal length range based on the image's height, width, and number of channels. ; Step 2: The multifocus image sequence obtained in Step 1... The data is input into the data preprocessing module via equation (1) and used in conjunction with preset distribution adjustment parameters. Perform normalization processing to obtain a normalized image sequence. , ; in It is a pixel offset constant. The pixel scaling factor. This is a zero-centralized bias term; Step 3: Normalize the image sequence obtained in Step 2. Input to the preset encoder module via equation (2) Perform latent spatial feature encoding to obtain latent feature sequences. , ; in For predefined potential spatial distribution adjustment factors, These represent the height, width, and number of channels of the potential feature, respectively. Step 4: Calculate the latent feature sequence obtained in Step 3. The input is fed into the feature stitching module via equation (3) and combined with the depth latent characterization noise generated at the current moment. Execution channel dimension concatenation to construct a joint guidance feature sequence , ; Meanwhile, the multifocus sequence recorded in step 1 Obtaining focused prior condition features With equation (4) and diffusion time step The temporal conditional features obtained by sinusoidal position encoding transformation To achieve interactive integration and build comprehensive guidance conditions , ; in This represents the concatenation, weighted fusion, and interactive attention operations of conditional information. Step 5: Combine the joint guiding feature sequence obtained in Step 4 With comprehensive guidance conditions The input is fed into the preset denoising network via equation (5). It is used to predict perturbation terms in the latent characterization of injection depth. ; Feature enhancement is performed using its internally integrated cross-focal-plane self-attention module, and then the features are reorganized into a vertical feature sequence centered on the pixel spatial location using Equation (6). Then, the correlation weights between different focal planes are calculated to achieve adaptive aggregation of sharp features. ; in, They are the characteristic sequences. The query, key, and value matrix obtained through linear transformation The proportional factor for the head of attention; Step 6: Calculate the denoising network prediction perturbation term obtained in Step 5. The equation (7) is input into the inverse diffusion iterative sampling module, and is used in conjunction with a preset noise scheduling accumulation coefficient. Perform sampling update, from to Update the deep latent representation frame by frame until a fully denoised deep latent representation is obtained. , ; Step 7: Apply the fully denoised deep latent representation obtained in Step 6 The input is given by equation (8) to the preset decoder module. Perform a reconstruction from the latent space to the physical pixel space to obtain the initial depth prediction results. , ; Step 8: Calculate the prediction results obtained in Step 7. The equation (9) is input into the numerical inverse regression normalization module, and used in conjunction with the physical focal length range recorded in step 1. Perform a linear affine transformation, and then map the depth map from the normalized numerical range back to the true physical range to obtain the absolute depth map of the physical space. , ; Clip is the clipping function, and the output is trunculated to ensure that the depth values ​​of all pixels are within the preset physical effective range, thus obtaining the final 3D shape reconstruction result. .