Monocular depth estimation method for polarization-guided diffusion model

By employing a gated fusion strategy that integrates polarization and RGB images, the problem of insufficient accuracy in monocular depth estimation in complex regions is solved, achieving high-precision and robust depth perception.

CN120894408APending Publication Date: 2025-11-04NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511008514.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing monocular depth estimation methods are inadequate in complex regions such as areas lacking texture, transparent surfaces, or highly reflective surfaces, and the lack of polarization information leads to algorithm degradation.

Method used

By fusing polarized images and monocular RGB images, gated fusion is performed in the latent space using a variational autoencoder and a lightweight confidence predictor. Taking advantage of the complementary advantages of polarization characteristics and visual cues, a lightweight confidence predictor is designed to perform information-weighted fusion.

Benefits of technology

It achieves more accurate depth estimation in complex scenes, and improves the estimation accuracy and robustness of the model in areas such as textureless, highly reflective, and transparent surfaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894408A_ABST
    Figure CN120894408A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and particularly relates to a polarization-guided diffusion model monocular depth estimation method. According to the method, the polarization image is integrated into a monocular depth estimation framework for the first time, a learnable lightweight confidence predictor is designed, and gating fusion is performed on the RGB image and the polarization image on the level of a potential space based on the predicted confidence, so that the information of the RGB image and the polarization image is reasonably integrated, and the accuracy of depth estimation is improved. And more accurate depth estimation of complex scenes such as texture-free, high-reflection and transparent surfaces can be realized. Compared with the current most advanced method, the scheme provided by the invention guides the network to carry out dynamic information weighting according to the regional confidence coefficient through explicit modeling of the complementarity of the polarization mode and the RGB mode in the potential feature space. On the premise that the inference cost is not obviously increased, more excellent performance and visual quality are shown, and particularly, better structural integrity is shown in complex areas such as texture-free, high-reflection and transparent surfaces.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and particularly relates to a polarization-guided diffusion model monocular depth estimation method. BACKGROUND

[0002] With the wide application of computer vision technology in automatic driving, robot navigation, virtual reality and other fields, the demand for high-precision depth information is increasing. In the field of depth estimation, monocular depth estimation is a highly valuable technology. It is based on the two-dimensional image collected by a single camera, and through the analysis of visual clues such as texture, geometric structure, shadow, etc. in the image, combined with deep learning algorithm or traditional geometric model, the three-dimensional depth information of objects in the scene is inferred from two-dimensional information. Due to the low cost, small volume and easy deployment of monocular camera, monocular depth estimation has a wide application prospect in automatic driving, robot navigation, augmented reality and many other fields. At present, diffusion models have gradually become the mainstream direction of research in the field of monocular depth estimation technology due to their strong pre-training prior.

[0003] In recent years, polarization-guided three-dimensional perception has also made significant progress. When light encounters the surface of an object and is reflected or refracted, its polarization state will change, and polarization imaging technology can capture these changes in polarization state. By analyzing the polarization light information at different angles and different positions, the details of the object surface material, shape, roughness, etc. can be obtained, and then a three-dimensional model of the object can be constructed to realize depth perception of the scene. This three-dimensional perception method using polarization information can provide a unique perspective different from traditional visual perception.

[0004] Chinese patent "CN202311689601.3 A monocular depth estimation method based on diffusion model" discloses a monocular depth estimation method based on diffusion model. This patent improves the use of conditional prompts, making the diffusion model more flexible and diverse in monocular depth estimation tasks. The patent proposes a preprocessing method that uses depth distribution knowledge in monocular depth estimation to reasonably preprocess and optimize the conditional prompts to improve the model's generalization ability in different scenarios, and introduces a data enhancement method to further optimize the model performance and depth map prediction effect. However, the method only estimates depth through RGB images without providing additional polarization information, which may cause the algorithm to degrade in complex areas such as texture missing areas, transparent surfaces or highly reflective surfaces.

[0005] The Chinese patent "CN202311620723.7 Polarization three-dimensional reconstruction method and system based on priori guided fusion network" proposes a polarization three-dimensional reconstruction method and system based on a priori guided fusion network. This patent is implemented based on a double-branch architecture, a feature correction module is designed to correct defects in the channel dimension and the spatial dimension respectively. In addition, a feature fusion module based on an effective cross-attention mechanism is proposed to fuse polarization and shadow prior features, achieving high-precision surface normal vector estimation and thus reconstructing high-quality three-dimensional targets. However, the proposed method is based on a discriminative model, which does not have the prior modeling capability compared to generative models, so there is still a bottleneck in model performance. SUMMARY

[0006] The purpose of the present application is to propose a polarization-guided diffusion model monocular depth estimation method that fuses polarization images and monocular RGB images. By exploiting the complementary advantages of the polarization characteristics of light and the visual cues of monocular images, high-precision and robust scene depth perception is achieved.

[0007] The technical solution of the present application is as follows: a polarization-guided diffusion model monocular depth estimation method, comprising the following steps:

[0008] Step one: input RGB images, polarization images, and real depth values, and map them to [-1, 1] respectively; the purpose of mapping is to convert the three into the input form of the variational autoencoder;

[0009] Step two: encode the mapped RGB images, mapped polarization images, and mapped real depth values through the variational autoencoder to obtain the RGB input in the latent space, the polarization input in the latent space, and the real depth value in the latent space; the variational autoencoder compresses the mapped RGB images, polarization images, and real depth values into a low-dimensional space for diffusion;

[0010] Step three: gate fusion of the RGB input in the latent space and the polarization input in the latent space;

[0011] Step four: train the lightweight confidence predictor and the U-Net to optimize the network parameters;

[0012] Step five: freeze the trained parameters to estimate the depth of the scene.

[0013] The polarization image includes scene polarization degree p and scene polarization angle f; the mapping process of the polarization image is as follows: the scene polarization degree is linearly mapped to p norm ∈[-1,1]; the original value range of the scene polarization angle f is [0, p], which is mapped to a two-dimensional vector form: f norm =[cos2f,sin2f], finally, the mapping form of the polarization image is:

[0014] x p = [ρ norm , cos2φ, sin2φ].

[0015] The RGB image is a conventional color image of a scene; the real depth value is an actual depth value of the scene, used for supervision during training.

[0016] The variational autoencoder adopts frozen pre-training parameters and does not participate in training throughout the whole process.

[0017] The step three is to obtain a confidence map through latent space confidence prediction; a gated guided latent space fusion strategy is provided to realize adaptive joint reconstruction between the RGB mode and the polarization mode, and obtain a fusion signal of the latent space.

[0018] The latent space confidence prediction specifically comprises: a lightweight confidence predictor f c (z pol ,z rgb ) is used to estimate the reliability of the polarization image in each latent space position; the lightweight confidence predictor is a shallow convolutional neural network, which receives the RGB image of the latent space and the polarization image of the latent space as input and outputs a single-channel confidence map; the last layer of the lightweight confidence predictor is activated by a Sigmoid function, so as to ensure that the output alpha value range is limited to [0, 1].

[0019] The gated guided latent space fusion strategy specifically comprises the following steps:

[0020] z f = α * z pol + (1-α) * z rgb

[0021] In the above formula, alpha represents the predicted confidence, which is a single-channel matrix, * represents element-wise multiplication, 1 represents a unit matrix, z rgb and z pol respectively represent the RGB input of the latent space and the polarization input of the latent space.

[0022] (4.1) Noise addition: during the training process, noise is added to the real depth value of the latent space to obtain a noisy latent space real depth value:

[0023]

[0024] Wherein ε ~ N (0, I) is a standard normal distribution; {β1,...,β T} are pre-set noise variance coefficients; z0 is the latent space real depth value without adding noise.

[0025] (4.2) Parameter training: splice the added noise potential depth value and the fusion signal obtained in step three into the U-Net of the diffusion model, since the confidence predictor and the U-Net form a strong coupling relationship in the task of final depth estimation, the U-Net is jointly optimized with the lightweight confidence predictor in step three, and the loss function is,

[0026]

[0027] Wherein, epsilon θ (z t ,t,z f ) is the noise predicted by the U-Net, and epsilon is the generated initial noise.

[0028] The step five reasoning process is as follows:

[0029] Generate initial random noise, input monocular RGB image and polarization image;

[0030] The monocular RGB image and the polarization image are encoded and fused in the latent space gate by the variational autoencoder;

[0031] Splice the initial random noise and the fusion result into the U-Net network;

[0032] Subtract the initial random noise from the U-Net network output result, and denoise;

[0033] Input the denoising result into the decoder of the variational autoencoder to obtain the estimated depth value.

[0034] The beneficial effects of the present application: the present application first integrates the polarization image into the monocular depth estimation framework, designs a learnable lightweight confidence predictor, and performs gate fusion of the RGB image and the polarization image based on the predicted confidence in the latent space, thereby reasonably integrating the information of the two, and realizing more accurate depth estimation of complex scenes such as textureless, high reflection and transparent surfaces. BRIEF DESCRIPTION OF DRAWINGS

[0035] Figure 1 The depth estimation effect diagram of the present application; (a)-(h) are respectively the depth estimation results and the corresponding original images of two data set samples under different methods: (a)-(d) are a sample in the HyperPol data set, which respectively represent the real depth map, the depth map predicted based on the pure RGB method, the depth map predicted by the method of the present application and the corresponding original RGB image; (e)-(h) are a sample in the HAMMER data set, which respectively represent the real depth map, the depth map predicted based on the pure RGB method, the depth map predicted by the method of the present application and the corresponding original RGB image.

[0036] Figure 2 Training flowchart of the present application;

[0037] Figure 3 Inference flowchart of the present application;

[0038] Figure 4 Lightweight confidence predictor schematic diagram. DETAILED DESCRIPTION

[0039] As Figures 2-4 shown, a polarization-guided diffusion model monocular depth estimation method includes the following steps:

[0040] Step one: input RGB image, polarization image and real depth value, map the three to [-1, 1] respectively; the purpose of mapping is to convert the three into the input form of variational autoencoder;

[0041] Step two: encode the mapped RGB image, mapped polarization image and mapped real depth value respectively through the variational autoencoder to obtain the RGB input of latent space, polarization input of latent space and real depth value of latent space; the variational autoencoder compresses the mapped RGB image, polarization image and real depth value into low-dimensional space for diffusion;

[0042] Step three: gate fusion of the RGB input of latent space and the polarization input of latent space;

[0043] Step four: train the lightweight confidence predictor and U-Net to optimize the network parameters;

[0044] Step five: freeze the trained parameters to estimate the depth of the scene.

[0045] The polarization image includes scene polarization degree ρ and scene polarization angle φ; the mapping process of the polarization image is as follows: the scene polarization degree is linearly mapped to ρ norm ∈[-1,1]; the scene polarization angle φ originally takes the value range of [0, π], which is mapped into a two-dimensional vector form: φ norm =[cos2φ,sin2φ], finally, the form of the polarization image mapping is:

[0046] x p =[ρ norm ,cos2φ,sin2φ].

[0047] The RGB image is a regular color image of the scene; the real depth value is the actual depth value of the scene, which is used for supervision when training.

[0048] The variational autoencoder adopts frozen pre-training parameters and does not participate in training throughout.

[0049] The step three is to obtain a confidence map through latent space confidence prediction; a gated guided latent space fusion strategy is provided to realize adaptive joint reconstruction between the RGB mode and the polarization mode, and a fusion signal of the latent space is obtained.

[0050] The latent space confidence prediction is specifically: a lightweight confidence predictor f c (z pol ,z rgb ) is used to estimate the reliability of the polarization image in each latent space position; the lightweight confidence predictor is a shallow convolutional neural network, which receives the RGB image of the latent space and the polarization image of the latent space as input, and outputs a single-channel confidence map; the last layer of the lightweight confidence predictor is activated by a Sigmoid function, so as to ensure that the output alpha value range is limited to [0, 1].

[0051] The gated guided latent space fusion strategy is specifically as follows:

[0052] z f =α*z pol +(1-α)*z rgb

[0053] In the above formula, alpha represents the predicted confidence, which is a single-channel matrix, * represents element-wise multiplication, 1 represents a unit matrix, z rgb and z pol respectively represent the RGB input of the latent space and the polarization input of the latent space.

[0054] (4.1) Noise addition: during the training process, noise is added to the latent space real depth value to obtain a noisy latent space real depth value:

[0055]

[0056] Wherein ε ~ N (0, I) is a standard normal distribution; {β1,...,β T} is a pre-set noise variance coefficient; z0 is the latent space real depth value without adding noise;

[0057] (4.2) Parameter training: the noisy latent depth value and the fusion signal obtained in step three are spliced and input into the U-Net of the diffusion model; since the confidence predictor and the U-Net form a strong coupling relationship in the task of final depth estimation, the U-Net and the lightweight confidence predictor in step three are optimized jointly, and the loss function is,

[0058]

[0059] wherein ε θ (z t ,t,z f ) is the U-Net predicted noise, and epsilon is the generated initial noise.

[0060] The step five reasoning process is as follows:

[0061] Generate an initial random noise, input monocular RGB image and polarization image;

[0062] The monocular RGB image and polarization image are encoded by the variational autoencoder and the latent space gating fusion is performed;

[0063] The initial random noise is spliced with the fusion result and input into the U-Net network;

[0064] The initial random noise is subtracted from the U-Net network output result, and denoising is performed;

[0065] The denoising result is input into the decoder of the variational autoencoder to obtain the estimated depth value.

[0066] Taking augmented reality (AR) as an example, the application introduces a polarization image on the basis of a traditional RGB input, and provides more abundant information about surface material and geometric structure. The system simultaneously collects an RGB image and a polarization image through a camera equipped with a polarization filter, and inputs the two as a joint input into a pre-trained depth estimation model. Compared with a method using only an RGB, the introduction of polarization information enables the model to have stronger geometric perception ability when facing difficult scenes such as reflective surfaces and texture missing areas. The high-quality depth map output by the model can be used for positioning, occlusion processing and interaction effect of virtual objects in a real scene, thereby significantly improving the immersion and stability of AR applications.

[0067] Compared with the current most advanced method, the scheme proposed by the application models the complementarity of the polarization and RGB modalities in the latent feature space explicitly, and guides the network to perform dynamic information weighting according to the regional confidence. Without significantly increasing the reasoning cost, the scheme exhibits more excellent performance and visual quality, especially in complex regions such as textureless, highly reflective and transparent surfaces, and exhibits better structural integrity. The results of testing 500 samples on the simulation dataset HyperPol and 200 samples on the real dataset HAMMER show that the effect of the scheme can reach the optimal performance.

Claims

1. A polarization-guided diffusion model monocular depth estimation method, characterized in that, The steps include the following: Step 1: Input RGB image, polarization image and true depth value, and map each of them to the range [-1, 1]; the purpose of mapping is to convert the three into the input form of variational autoencoder; Step 2: Encode the mapped RGB image, the mapped polarization image, and the mapped true depth value using a variational autoencoder to obtain the RGB input, polarization input, and true depth value of the latent space. The variational autoencoder compresses the mapped RGB image, polarization image, and true depth value into a low-dimensional space for diffusion. Step 3: Perform gated fusion on the RGB input and polarization input of the latent space; Step 4: Train the lightweight confidence predictor and U-Net, and optimize the network parameters; Step 5: Freeze the trained parameters and perform depth estimation on the scene.

2. The polarization-guided diffusion model monocular depth estimation method according to claim 1, characterized in that, The polarization image includes the scene polarization degree ρ and the scene polarization angle φ; the mapping process of the polarization image is as follows: the scene polarization degree is linearly mapped to ρ. norm ∈[-1,1]; the original value range of the scene polarization angle φ is [0,π], which is mapped to a two-dimensional vector form: φ norm =[cos2φ,sin2φ], and finally, the polarization image mapping takes the form of: x p =[ρ norm ,cos2φ,sin2φ].

3. The polarization-guided diffusion model monocular depth estimation method according to claim 1, characterized in that, The RGB image is a regular color image of the scene; the true depth value is the actual depth value of the scene, used for supervision during training.

4. The polarization-guided diffusion model monocular depth estimation method according to claim 1, characterized in that, The variational autoencoder uses frozen pre-trained parameters and does not participate in the training process.

5. The polarization-guided diffusion model monocular depth estimation method according to claim 1, characterized in that, Step three involves obtaining a confidence map through latent space confidence prediction; and providing a gated latent space fusion strategy to achieve adaptive joint reconstruction between RGB modes and polarization modes to obtain a fused signal in the latent space.

6. The polarization-guided diffusion model monocular depth estimation method according to claim 5, characterized in that, The latent space confidence prediction specifically involves proposing a lightweight confidence predictor f. c (z pol ,z rgb The lightweight confidence predictor is used to estimate the reliability of the polarization image of the latent space at each latent space location. The lightweight confidence predictor is a shallow convolutional neural network that takes the RGB image of the latent space and the polarization image of the latent space as input and outputs a single-channel confidence map. The last layer of the lightweight confidence predictor is activated by the Sigmoid function to ensure that the output α value range is limited to [0,1].

7. The polarization-guided diffusion model monocular depth estimation method according to claim 6, characterized in that, The specific gating-guided potential spatial fusion strategy is as follows: With f =α*z pol +(1-α)*z rgb In the above formula, α represents the prediction confidence level, which is a single-channel matrix; * indicates element-wise multiplication; 1 represents the identity matrix; and z rgb and z pol These represent the RGB input and the polarization input through the latent space, respectively.

8. The polarization-guided diffusion model monocular depth estimation method according to claim 1, characterized in that, Step four is as follows: (4.1) Noise addition: During the training process, noise is added to the true depth value of the latent space to obtain the noisy true depth value of the latent space: Where ε ~ N(0,I) is a standard normal distribution; {β1,...,β T } is the pre-defined noise variance coefficient; z0 is the true depth value of the potential space without added noise; (4.2) Parameter Training: The latent depth values ​​after adding noise are concatenated with the fused signal obtained in step three and then input into the U-Net of the diffusion model. Since the confidence predictor and U-Net form a strong coupling relationship in the final depth estimation task, U-Net and the lightweight confidence predictor in step three are jointly optimized. The loss function is: Where, ε θ (z t ,t,z f ) is the noise predicted by U-Net, and ε is the generated initial noise.

9. The polarization-guided diffusion model monocular depth estimation method according to claim 1, characterized in that, The reasoning process for step five is as follows: Generate initial random noise, and input a monocular RGB image and a polarization image; Variational autoencoder encoding and latent spatial gating fusion are performed on monocular RGB images and polarization images; The initial random noise and the fusion result are concatenated and then input into the U-Net network; The initial random noise is subtracted from the output of the U-Net network to perform noise reduction. The denoising result is input into the decoder of the variational autoencoder to obtain the estimated depth value.

Citation Information

Patent Citations

  • Polarization three-dimensional reconstruction method and system based on prior guidance fusion network

    CN117671142A

  • Monocular image depth estimation method based on diffusion model

    CN117689704A