Monocular Depth Estimation Method and Device Based on Improved Depth Distribution Compensation and Frequency Domain Feature Fusion
Through the adaptive depth bias compensation mechanism and frequency-aware fusion module, combined with the stable diffusion model, the shortcomings of the monocular depth estimation method in long-distance depth prediction and detail recovery are solved, and efficient depth feature generation and accurate depth estimation are achieved.
Patent Information
- Application Number
- CN202510116744.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing monocular depth estimation method has shortcomings in long-distance depth prediction and detail recovery. It is mainly due to the long-tail effect of the depth distribution of the training data set, which reduces the prediction accuracy of the model for long-distance areas. The existing method ignores the ability of frequency domain features to portray depth edge details.
The adaptive depth bias compensation mechanism is designed, and the depth features are decomposed high-frequency and low-frequency through the frequency-aware fusion module, and the adaptive weighting method is used to fuse the high-low-frequency components, and the stable diffusion model is combined for deep feature extraction and denoising training, and the loss weight allocation of the depth estimation model is dynamically optimized.
It significantly improves the capture effect of edge details and the consistency of depth estimation, especially in long-distance areas and complex scenarios, solving the prediction error problem caused by sparse samples in long-distance areas.
Smart Images

Figure CN120031934B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a monocular depth estimation method and device based on improved depth distribution compensation and frequency domain feature fusion. Background Art
[0002] Depth estimation is a fundamental task in computer vision, aiming to predict the depth information of each pixel from a monocular RGB image. Monocular depth estimation has wide application value in multiple vision tasks such as 3D reconstruction, autonomous driving, and augmented reality. However, existing monocular depth estimation methods have significant deficiencies in long-distance depth prediction and detail restoration. The main reason is that there is a significant depth distribution long-tail effect in the training dataset, that is, the number of samples in the near-distance area is much larger than that in the long-distance area, resulting in the depth estimation model tending to focus more on the near-distance area, thus reducing the prediction accuracy of the long-distance area.
[0003] To solve this problem, in recent years, some improvement strategies have been introduced in depth estimation research, including methods such as distance-aware loss functions and multi-expert models. These methods can alleviate the long-tail distribution effect to a certain extent, but due to their dependence on specific training data distributions, the model often shows a decline in generalization performance when dealing with out-of-distribution data. In addition, most existing depth estimation algorithms rely on spatial domain features and ignore the ability of frequency domain features to depict depth edge details. Frequency domain analysis has been proven to have advantages in enhancing edge features and suppressing noise, but directly introducing frequency information into the depth estimation network still has problems such as unreasonable fusion methods and insufficient feature expression capabilities.
[0004] In recent years, depth estimation methods based on the Stable Diffusion Model have received extensive attention. The Stable Diffusion Model can generate depth maps with high semantic consistency by gradually adding random noise to the image and learning the reverse denoising process. However, existing Stable Diffusion depth estimation models also suffer from the long-tail distribution effect and still have significant biases in long-distance area depth prediction. At the same time, existing methods rely on the U-Net network in the spatial domain during the feature extraction process and ignore the potential of frequency domain information, resulting in insufficient edge detail restoration capabilities. Summary of the Invention
[0005] To solve the deficiencies of existing monocular depth estimation methods in long-distance depth prediction and detail restoration, the primary objective of the present invention is to provide a monocular depth estimation method based on improved depth distribution compensation and frequency domain feature fusion, which solves the prediction error problem caused by sparse samples in the long-distance area by designing an adaptive depth bias compensation mechanism, and adopts an adaptive fusion strategy to enhance the feature expression ability, significantly improving the capture effect of edge details and the consistency of depth estimation.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A monocular depth estimation method based on improved depth distribution compensation and frequency-domain feature fusion, the method comprising the following steps in sequence:
[0007] (1) Construct a monocular depth estimation model: Introduce a frequency-aware fusion module based on the U-Net network model. The U-Net network model and the frequency-aware fusion module form a monocular depth estimation model. The input depth feature map is decomposed into high-frequency components and low-frequency components through the fast Fourier transform, and the high-frequency components and low-frequency components are fused through an adaptive weighting method;
[0008] (2) Design the loss function of the monocular depth estimation model: Construct an adaptive depth bias compensation mechanism based on depth distribution weights and edge gradient weights, dynamically allocate the training loss weights of each depth region according to the adaptive depth bias compensation mechanism, and calculate the total loss function to balance the estimation accuracy of the far and near depth regions;
[0009] (3) Train the monocular depth estimation model: Use the stable diffusion model as the generation network, and combine it with the monocular depth estimation model for depth feature extraction and denoising training to obtain the trained monocular depth estimation model and noise
[0010] (4) Perform monocular depth estimation: Input the monocular RGB image to be estimated and the initial noise into the trained monocular depth estimation model, and generate a high-precision depth map through multi-step denoising.
[0011] Step (1) specifically refers to: Introduce a frequency-aware fusion module, and the frequency-aware fusion module uses the fast Fourier transform to decompose the depth feature map to enhance the ability to capture depth edge details; First, perform the fast Fourier transform on the depth feature map z d ∈R H×W×C in the input latent space to convert the spatial domain features to the frequency domain representation:
[0012]
[0013] where, Z freq represents the depth feature in the frequency domain, u and v are the frequency components in the horizontal and vertical directions respectively, W represents the width of the depth feature map in the latent space, H represents the height of the depth feature map in the latent space, C represents the number of channels of the depth feature map in the latent space; dx represents the position of the depth feature map in the latent space in the horizontal direction; dy represents the position of the depth feature map in the latent space in the vertical direction;
[0014] To further separate high-frequency information and low-frequency information, decomposition is performed based on the normalized frequency radius ρ(k), where ρ(k) is defined as:
[0015]
[0016] In the formula, k x is the frequency component along the horizontal direction in the frequency domain, and k y is the frequency component along the vertical direction in the frequency domain;
[0017] According to the normalized frequency radius ρ(k), the high-frequency mask M high and the low-frequency mask M low are defined as:
[0018]
[0019] M low (k) = 1 - M high (k)
[0020] where τ is the set frequency threshold; based on the high-frequency mask M high and the low-frequency mask M low masks, the high-frequency component and the low-frequency component are respectively extracted:
[0021] Z high = F -1 (M high ·Z freq )
[0022] Z low = F -1 (M low ·Z freq )
[0023] where F -1 represents the inverse Fourier transform, which is used to convert the high- and low-frequency components from the frequency domain back to the spatial domain; the high-frequency component Z high contains the edge and detail information in the depth feature map, and the low-frequency component Z low retains the global depth structure of the depth feature map;
[0024] To optimize the utilization of the high- and low-frequency components, an adaptive weighting method is used for feature fusion, and the fused feature Z FAF is:
[0025] Z FAF = α1Z high + α2Z low
[0026] α1 + α2 = 1
[0027] Among them, α1 is a learnable parameter used to control the weight of high-frequency components in the depth feature map; α2 is a learnable parameter used to control the weight of low-frequency components in the depth feature map.
[0028] Step (2) specifically refers to: designing an adaptive depth bias compensation mechanism, which consists of a depth distribution weight map w distance and an edge gradient weight map w edge to address the long-tailed effect problem of depth distribution;
[0029] The depth distribution weight map w distance is achieved by calculating the distribution density of pixel depths:
[0030]
[0031] where d is the original depth map, β is a smoothing parameter, z mean represents the mean depth of each pixel point in the original depth map d, and P is the percentile of the depth distribution; the edge gradient weight map w edge then calculates the edge gradient magnitude of the original depth map d through the Sobel operator:
[0032]
[0033] where G x and G y are the gradients in the horizontal and vertical directions respectively;
[0034] The total loss function L total is:
[0035]
[0036] where N represents the number of training samples, λL var is the variance regularization term used to enhance the diversity of latent features; ∈ θ,i is the predicted noise of the i-th sample, and ∈ t,i is the noise added at time t for the i-th sample; w combined = λ base w distance +(1 - λ base )w edge where λ base is a hyperparameter used to balance the contribution ratio of the depth distribution weight and the edge gradient weight; the dynamic adjustment of w combined can effectively adapt to changes in different depth regions and edge characteristics, thereby optimizing the depth estimation performance.
[0037] Step (3) specifically refers to: during the training process, using a stable diffusion model and combining it with a monocular depth estimation model for depth feature extraction and denoising training;
[0038] First, encode the input RGB image x and the original depth map d into the latent space respectively:
[0039] z x = E(x)
[0040] z d = E(d)
[0041] where z x is the RGB feature map in the latent space, z d is the depth feature map in the latent space, E(x) represents the encoder used to encode the RGB image, and E(d) represents the encoder used to encode the depth map;
[0042] In the forward diffusion process, by gradually adding Gaussian noise to the depth feature map z d in the latent space, depth latent representations containing different levels of noise are generated
[0043]
[0044] where t ∈ {1, …, T} is the current time step, representing the stage of the diffusion process; T is the maximum number of time steps; is the noise attenuation coefficient at time step t, controlling the balance between the original depth map d and the noise;
[0045] represents Gaussian noise of the standard normal distribution; through forward diffusion, depth latent representations containing more noise are gradually generated to simulate the noise distribution that may appear in the actual input;
[0046] In the reverse denoising process, it works in cooperation with a monocular depth estimation model to gradually remove the noise and restore a high-quality depth latent representation. The mathematical expression of the denoising process is:
[0047]
[0048] where represents the noise predicted by the U-Net network, with the input being the depth latent representation at the current time step t and the RGB image representation z x in the latent space; is the frequency-domain enhancement result of the frequency-aware fusion module for the depth features; is the depth latent representation generated by reverse diffusion; through the method of gradually denoising, the depth latent representation gradually recovers from a high-noise state to a low-noise state, and finally generates a high-quality denoised depth feature map in the latent space After the reverse diffusion is completed, the decoder D of the variational autoencoder VAE is used to decode the denoised depth feature map of the latent space into a predicted depth map:
[0049]
[0050] wherein, is the predicted depth map.
[0051] Step (4) specifically includes the following steps in sequence:
[0052] (4a) Latent space encoding: The input monocular RGB image to be estimated is first encoded by the encoder E of the frozen variational autoencoder VAE to generate a latent space representation;
[0053] (4b) Initial noise generation: To initiate the reverse diffusion process, initial noise is sampled from a standard normal distribution
[0054]
[0055] The initial noise represents the initial state of the depth latent representation, and will be gradually restored to a high-quality depth feature through multi-step denoising;
[0056] (4c) Reverse diffusion denoising process: The noise is removed through multiple iterations to generate a gradually optimized depth latent representation;
[0057] (4d) Dynamic optimization through an adaptive depth bias compensation mechanism: An adaptive depth bias compensation mechanism is introduced to dynamically optimize the depth latent features;
[0058] (4e) After the depth decoding iterative denoising process is completed, the finally generated depth feature map of the denoised latent space is passed to the frozen VAE decoder D and reconstructed into a predicted depth map
[0059] Another object of the present invention is to provide an electronic device, including:
[0060] a processor; and
[0061] a memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor executes the above-mentioned monocular depth estimation method based on improved depth distribution compensation and frequency domain feature fusion.
[0062] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the above-mentioned monocular depth estimation method based on improved depth distribution compensation and frequency-domain feature fusion.
[0063] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, by systematically analyzing the long-tail distribution characteristics of depth data, the present invention designs an adaptive depth bias compensation mechanism to dynamically adjust the weight distribution of the near and far regions, fundamentally solving the prediction error problem caused by sparse samples in the far region; Second, the present invention uses a frequency-aware fusion module to decompose depth features into high-frequency and low-frequency components through fast Fourier transform, and adopts an adaptive fusion strategy to enhance the feature expression ability, significantly improving the capture effect of edge details and the consistency of depth estimation; Third, in the training stage, the present invention combines a stable diffusion model to achieve efficient depth feature generation through multi-step noise addition and denoising optimization; in the inference stage, through the adaptive depth bias compensation mechanism and frequency enhancement strategy, the performance of depth estimation in the far region and complex scenes is further improved; Fourth, experimental results show that the present invention shows significant advantages on multiple datasets, especially showing excellent accuracy and robustness in the far region and detail recovery, providing an advanced solution for computer vision applications such as 3D reconstruction, autonomous driving, and augmented reality. Description of the Drawings
[0064] Figure 1 is the flowchart of the method of the present invention;
[0065] Figure 2 is a diagram showing the depth distribution characteristics of the present invention in exploring multiple mainstream depth estimation datasets, and the performance of the present invention in each depth region;
[0066] Figure 3 is a diagram showing the results of ablation experiments on thresholds of the present invention on the KITTI dataset;
[0067] Figure 4 is a diagram showing the results of ablation experiments on thresholds of the present invention on the NYUv2 dataset;
[0068] Figure 5 is a qualitative comparison diagram of the monocular depth estimation performance of the present invention on multiple datasets;
[0069] Figure 6 is a qualitative comparison diagram of the present invention on natural scene samples;
[0070] Figure 7 、 Figure 8 、 Figure 9 、 Figure 10They are all qualitative comparison charts of the present invention on multiple datasets. Detailed implementation manners
[0071] As Figure 1 shown, a monocular depth estimation method based on improved depth distribution compensation and frequency domain feature fusion, the method includes the following steps in sequence:
[0072] (1) Construct a monocular depth estimation model: Introduce a frequency-aware fusion module on the basis of the U-Net network model. The U-Net network model and the frequency-aware fusion module form a monocular depth estimation model. The input depth feature map is decomposed into high-frequency components and low-frequency components through the fast Fourier transform, and the high-frequency components and low-frequency components are fused through an adaptive weighting method;
[0073] (2) Design the loss function of the monocular depth estimation model: Construct an adaptive depth bias compensation mechanism based on depth distribution weights and edge gradient weights, dynamically allocate the training loss weights of each depth region according to the adaptive depth bias compensation mechanism, and calculate the total loss function to balance the estimation accuracy of the far and near distance depth regions;
[0074] (3) Train the monocular depth estimation model: Use the stable diffusion model as the generation network, and combine the monocular depth estimation model for depth feature extraction and denoising training to obtain the trained monocular depth estimation model and noise
[0075] (4) Perform monocular depth estimation: Input the monocular RGB image to be estimated and the initial noise into the trained monocular depth estimation model, and generate a high-precision depth map through multi-step denoising.
[0076] Statistical analysis is carried out on multiple depth datasets (such as DIODE, KITTI, NYUv2, ScanNet, etc.), the depth data distribution is calculated, and it is found that the training samples show a significant long-tailed distribution characteristic in the depth space, that is, the number of samples in the near distance region (such as the 0-20% depth interval) is significantly more than the number of samples in the far distance region (more than the 60% depth interval), resulting in a large depth prediction deviation in the far distance region.
[0077] In the DIODE and KITTI datasets, the depth range is large, and the number of samples in the far distance region (for example, the depth exceeds 80 meters) is significantly insufficient. In the NYUv2 and ScanNet datasets, although the depth range is relatively small, the proportion of the number of samples in the near distance region (for example, the depth is less than 2 meters) is extremely high, thus limiting the prediction accuracy of the model in the far distance region. The present invention further verifies this problem through visual analysis of the sample distribution in different depth regions and finds that:
[0078] 1. The depth estimation error in the near region is usually low, while the error in the far region increases significantly;
[0079] 2. The performance degradation of the model in the far region is not only manifested in the deviation of the predicted value, but also in the blurring of edge details and the lack of depth continuity.
[0080] To more intuitively show the impact of the long-tail distribution on the model performance, the present invention further compares the changing trends of depth estimation errors in the near and far regions. The experimental results are as Figure 2 shown. In the near region (for example, within a range less than 2 meters in NYUv2), the average absolute relative error (AbsRel) is less than 0.1, while in the far region (for example, within a range greater than 60 meters in KITTI), the error increases to more than 0.2. This significant performance gap indicates that the imbalance of the depth distribution is the main factor affecting the prediction accuracy of the model in the far region. Based on the above analysis, the present invention designs an adaptive depth bias compensation mechanism (BiasMap), which can dynamically adjust the loss weights in the near and far regions to balance the optimization objectives in different depth regions. The present invention provides a systematic solution idea through statistical analysis and distribution quantization, providing a scientific basis for alleviating the long-tail distribution effect.
[0081] Step (1) specifically refers to: introducing a frequency-aware fusion module, which uses the fast Fourier transform to decompose the depth feature map to enhance the ability to capture depth edge details; first, perform the fast Fourier transform on the depth feature map z d ∈R H×W×C in the input latent space to convert the spatial domain features to the frequency domain representation:
[0082]
[0083] where Z freq represents the depth feature in the frequency domain, u and v are the frequency components in the horizontal and vertical directions respectively, W represents the width of the depth feature map in the latent space, H represents the height of the depth feature map in the latent space, and C represents the number of channels of the depth feature map in the latent space; dx represents the position of the depth feature map in the latent space in the horizontal direction; dy represents the position of the depth feature map in the latent space in the vertical direction;
[0084] To further separate the high-frequency information and the low-frequency information, decompose based on the normalized frequency radius ρ(k), where ρ(k) is defined as:
[0085]
[0086] In the formula, k x is the frequency component along the horizontal direction in the frequency domain, k yis the frequency component along the vertical direction in the frequency domain;
[0087] According to the normalized frequency radius ρ(k), the high-frequency mask M is defined high and the low-frequency mask M low :
[0088]
[0089] M low (k) = 1 - M high (k)
[0090] where τ is the set frequency threshold; Based on the high-frequency mask M high and the low-frequency mask M low masks, the high-frequency component and the low-frequency component are extracted respectively:
[0091] Z high = F -1 (M high ·Z freq )
[0092] Z low = F -1 (M low ·Z freq )
[0093] where F -1 represents the inverse Fourier transform, which is used to convert the high- and low-frequency components from the frequency domain back to the spatial domain; The high-frequency component Z high contains the edge and detail information in the depth feature map, and the low-frequency component Z low retains the global depth structure of the depth feature map;
[0094] To optimize the utilization of the high- and low-frequency components, an adaptive weighted method is used for feature fusion, and the fused feature Z FAF is:
[0095] Z FAF = α1Z high + α2Z low
[0096] α1 + α2 = 1
[0097] where α1 is a learnable parameter used to control the weight of the high frequency in the depth feature map; α2 is a learnable parameter used to control the weight of the low-frequency component in the depth feature map. The adaptive fusion strategy can dynamically adjust the weight allocation according to the specific scenario, so as to balance detail capture and global consistency.
[0098] In addition, the present invention further conducts an experimental analysis on the enhancement effects of high and low frequency components. The results show that using only high frequency components can significantly improve the edge sharpness in the depth map, but may introduce texture noise; while using only low frequency components helps to maintain the overall smoothness of the depth map, but may lead to loss of details. The high-low frequency fusion achieved through the frequency-aware fusion module, i.e., the FAF module, can enhance edge details while maintaining the consistency of the global depth, thus significantly improving the accuracy and robustness of depth estimation. By introducing the frequency-aware fusion module, efficient decomposition and fusion of depth features in the frequency domain are realized, resulting in a significant improvement in edge detail recovery and global consistency, providing strong technical support for the depth estimation method proposed by the present invention. The introduction of the FAF module significantly improves the model's ability to process depth edge details and distant regions.
[0099] Step (2) specifically refers to: designing an adaptive depth bias compensation mechanism, which consists of a depth distribution weight map w distance and an edge gradient weight map w edge to solve the problem of the long-tailed effect of depth distribution;
[0100] The depth distribution weight map w distance is realized by calculating the distribution density of pixel depths:
[0101]
[0102] where d is the original depth map, β is a smoothing parameter, z mean represents the mean of the depths of each pixel point in the original depth map d, and P is the percentile of the depth distribution; the edge gradient weight map w edge calculates the edge gradient magnitude of the original depth map d through the Sobel operator:
[0103]
[0104] where G x and G y are the gradients in the horizontal and vertical directions respectively;
[0105] The total loss function L total is:
[0106]
[0107] where N represents the number of training samples, λL var is the variance regularization term, used to enhance the diversity of latent features; ∈ θ,i is the prediction noise of the i-th sample, ∈ t,i is the noise added to the i-th sample at time t; w combined =λbase w distance +(1 - λ base )w edge , where λ base is a hyperparameter used to balance the contribution ratio of the depth distribution weight and the edge gradient weight; the dynamic adjustment of w combined can effectively adapt to changes in different depth regions and edge characteristics, thereby optimizing the depth estimation performance.
[0108] Step (3) specifically refers to: during the training process, a stable diffusion model is adopted, combined with a monocular depth estimation model for depth feature extraction and denoising training;
[0109] First, the input RGB image x and the original depth map d are encoded into the latent space respectively:
[0110] z x = E(x)
[0111] z d = E(d)
[0112] In the formula, z x is the RGB feature map in the latent space, z d is the depth feature map in the latent space, E(x) represents the encoder used to encode the RGB image, and E(d) represents the encoder used to encode the depth map;
[0113] In the forward diffusion process, Gaussian noise is gradually added to the depth feature map z d in the latent space to generate depth latent representations containing different levels of noise
[0114]
[0115] where t ∈ {1,..., T} is the current time step, representing the stage of the diffusion process; T is the maximum number of time steps; is the noise attenuation coefficient at time step t, controlling the balance between the original depth map d and the noise;
[0116] represents Gaussian noise of the standard normal distribution; through forward diffusion, depth latent representations containing more noise are gradually generated to simulate the noise distribution that may appear in the actual input;
[0117] In the reverse denoising process, the monocular depth estimation model is used to work together to gradually remove the noise and restore high-quality depth latent representations. The mathematical expression of the denoising process is:
[0118]
[0119] wherein, represents the noise predicted by the U-Net network, and the input is the depth latent representation at the current time step t and the RGB image representation z of the latent space x ; is the frequency-domain enhancement result of the frequency-aware fusion module for the depth features; is the depth latent representation generated by reverse diffusion; through the step-by-step denoising method, the depth latent representation gradually recovers from the high-noise state to the low-noise state, and finally generates a high-quality denoised depth feature map zd0 of the latent space; after the reverse diffusion is completed, the decoder D of the variational autoencoder VAE is used to decode the denoised depth feature map zd0 of the latent space into a predicted depth map:
[0120]
[0121] wherein, is the predicted depth map.
[0122] Step (4) specifically includes the following steps in sequence:
[0123] (4a) Latent space encoding: The input monocular RGB image to be estimated is first encoded by the encoder E of the frozen variational autoencoder VAE to generate a latent space representation;
[0124] (4b) Initial noise generation: To initiate the reverse diffusion process, initial noise is sampled from the standard normal distribution
[0125]
[0126] The initial noise represents the initial state of the depth latent representation, and will be gradually restored to high-quality depth features through multi-step denoising;
[0127] (4c) Reverse diffusion denoising process: The noise is removed through multiple iterations to generate a gradually optimized depth latent representation;
[0128] (4d) Dynamic optimization through the adaptive depth bias compensation mechanism: The adaptive depth bias compensation mechanism is introduced to dynamically optimize the depth latent features;
[0129] (4e) After the depth decoding iterative denoising process is completed, a denoised depth feature map of the latent space is finally generated and is passed to the frozen VAE decoder D to be reconstructed into a predicted depth map
[0130] An embodiment of the present application may also be an electronic device, including:
[0131] A processor; and
[0132] A memory in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the present method.
[0133] An embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the above-mentioned present method.
[0134] As Figure 2 shown, the present invention explores the depth distribution characteristics of multiple mainstream depth estimation datasets (including DIODE, KITTI, NYUv2, and ScanNet), and the performance of the present invention in each depth region. The purple curve in the figure represents the joint average depth distribution, intuitively reflecting the imbalance of the depth distribution. The close-range region (0-20% depth range) occupies the vast majority of the sample quantity, while the samples in the far-range region (more than 60% depth range) are significantly sparse. This long-tail distribution problem is one of the main challenges faced by the current monocular depth estimation task. Figure 2 The performance of the VistaDepth and Marigold methods are respectively marked with red stars and blue stars. The horizontal axis represents the depth interval, and the vertical axis is the δ1 metric. The value of δ1 is the average of the results on the NYUv2 and KITTI test sets, and is used to measure the prediction accuracy of the model in different depth intervals. The results show that the performance of VistaDepth (red stars) is better than that of Marigold (blue stars) in each depth interval, especially showing a significant advantage in the far-range region (more than 60% depth range). The role of the frequency-aware fusion module in enhancing edge details and maintaining global consistency further improves the prediction accuracy of VistaDepth in the far-range region.
[0135] As Figure 3 、 Figure 4 shown, it shows the evaluation of the influence of different percentile thresholds, that is, the percentile P of the depth distribution, on the depth estimation performance when the present invention uses the BiasMap mechanism on the KITTI dataset and the NYUv2 dataset. Figure 3 is the performance evaluation result on the KITTI dataset, Figure 3 showing the trade-off relationship between the accuracy of depth estimation in the close-range region (evaluated by δ1) and the prediction error in the far-range region (evaluated by the absolute relative error AbsRel) as the percentile threshold changes. It can be observed from Figure 3 that when P is set to the 70th percentile, the model achieves the best balance between the close-range region and the far-range region, that is, while ensuring high accuracy in the close-range region, effectively reducing the prediction error in the far-range region.Figure 4 For similar evaluation results on the NYUv2 dataset, the performance of the BiasMap mechanism in indoor scenes was analyzed. Figure 4 Also shown in [reference] is the impact of different percentile thresholds on the accuracy (δ1) in the close - range area and the error (AbsRel) in the far - range area. The results indicate that the 70th percentile threshold can also achieve the best balance on the indoor dataset, making the optimization objectives of the model more balanced between different depth regions.
[0136] As Figure 5 shown, it is a qualitative comparison graph of the monocular depth estimation (MDE) performance of the present invention on multiple datasets. Figure 5 Shows the comparison results of VistaDepth proposed by the present invention with other existing monocular depth estimation algorithms on different datasets, including datasets such as NYUv2, KITTI, ETH3D, DIODE, and ScanNet. VistaDepth shows obvious advantages in capturing distant details and maintaining scene consistency. In the NYUv2 dataset, VistaDepth can accurately capture the depth details of the door edge, while the prediction results of other methods are blurred in the edge area. In the KITTI dataset, VistaDepth is more accurate in estimating the depth of distant vehicles, with clear vehicle boundaries, avoiding the common blurring phenomenon in traditional methods. In the ETH3D and DIODE datasets, VistaDepth significantly outperforms other methods in predicting the depth details of building windows, fully demonstrating its ability to capture high - frequency details. In the ScanNet dataset, the depth estimation results of VistaDepth in complex indoor scenes show high global consistency. For example, VistaDepth can correctly reflect the depth relationship of indoor furniture, while the results of other methods usually have obvious mistakes or depth discontinuities in the detail areas.
[0137] As Figure 6 shown, it is a qualitative comparison graph of the present invention on natural scene samples. Figure 6 Compares the performance of VistaDepth with other existing depth estimation algorithms (including DPT, Depth Anything, and Marigold), using a consistent color coding to represent the depth range, where red represents the near - plane area and blue represents the far - plane area. Figure 6 Intuitively demonstrates the significant advantages of VistaDepth in distant depth reconstruction and detail preservation.
[0138] As Figure 7 、 Figure 8 、 Figure 9 、 Figure 10 shown, Figure 7 、 Figure 8, Figure 9 , Figure 10 The performance of VistaDepth of the present invention is compared with other existing depth estimation algorithms (including DPT, Depth Anything, and Marigold), and a consistent color coding is used to represent the depth range, where red represents the near-plane area and blue represents the far-plane area. Figure 7 , Figure 8 , Figure 9 , Figure 10 Intuitively demonstrates the significant advantages of VistaDepth in long-distance depth reconstruction and detail retention.
[0139] In summary, through systematic analysis of the long-tail distribution characteristics of depth data, the present invention designs an adaptive depth bias compensation mechanism to dynamically adjust the weight distribution of the near and far regions, fundamentally solving the prediction error problem caused by sparse samples in the long-distance region; uses a frequency-aware fusion module to decompose depth features into high-frequency and low-frequency components through fast Fourier transform, and adopts an adaptive fusion strategy to enhance the feature expression ability, significantly improving the capture effect of edge details and the consistency of depth estimation; in the training stage, the present invention combines a stable diffusion model to achieve efficient depth feature generation through multi-step noise addition and denoising optimization; in the inference stage, through the dynamic optimization of the adaptive depth bias compensation mechanism and the frequency enhancement strategy, the performance of depth estimation in the long-distance region and complex scenes is further improved; experimental results show that the present invention shows significant advantages on multiple datasets, especially demonstrating excellent accuracy and robustness in the long-distance region and detail recovery, providing an advanced solution for computer vision applications such as 3D reconstruction, autonomous driving, and augmented reality.
Claims
1. A monocular depth estimation method based on improved depth distribution compensation and frequency domain feature fusion, characterized in that: The method includes the following steps in sequence: (1) Construct a monocular depth estimation model: Introduce a frequency-aware fusion module based on the U-Net network model. The U-Net network model and the frequency-aware fusion module form a monocular depth estimation model. The input depth feature map is decomposed into high-frequency components and low-frequency components through the fast Fourier transform, and the high-frequency components and low-frequency components are fused through an adaptive weighting method; (2) Design the loss function of the monocular depth estimation model: Construct an adaptive depth bias compensation mechanism based on the depth distribution weight and the edge gradient weight, dynamically allocate the training loss weights of each depth region according to the adaptive depth bias compensation mechanism, and calculate the total loss function to balance the estimation accuracy of the far and near depth regions; (3) Training the monocular depth estimation model: Using the StableDiffusion model as the generation network, and combining it with the monocular depth estimation model for depth feature extraction and denoising training to obtain the trained monocular depth estimation model and noise (4)Perform monocular depth estimation: Input the monocular RGB image to be estimated and the initial noise into the trained monocular depth estimation model, and generate a high-precision depth map through multiple steps of denoising; Step (1) specifically refers to: introducing a frequency-aware fusion module, which uses the fast Fourier transform to decompose the depth feature map to enhance the ability to capture depth edge details; first, perform the fast Fourier transform on the depth feature map z of the input latent space d ∈R H×W×C to convert the spatial domain features to the frequency domain representation: Among them, Z freq represents the depth feature in the frequency domain, u and v are the frequency components in the horizontal and vertical directions respectively, W represents the width of the depth feature map in the latent space, H represents the height of the depth feature map in the latent space, and C represents the number of channels of the depth feature map in the latent space; dx represents the position of the depth feature map in the latent space in the horizontal direction; dy represents the position of the depth feature map in the latent space in the vertical direction; To further separate high-frequency information and low-frequency information, decomposition is performed based on the normalized frequency radius ρ(k), where ρ(k) is defined as: where k x is the frequency component along the horizontal direction in the frequency domain, and k y is the frequency component along the vertical direction in the frequency domain; Define the high-frequency mask M according to the normalized frequency radius ρ(k). high and the low-frequency mask M low : M low (k) = 1 - M high (k) where τ is the set frequency threshold; based on the high-frequency mask M high and the low-frequency mask M low masks, the high-frequency components and the low-frequency components are extracted respectively: Z high = F -1 (M high ·Z freq ) Z low = F -1 (M low ·Z freq ) Among them, F -1 represents the inverse Fourier transform, which is used to convert the high- and low-frequency components from the frequency domain back to the spatial domain; the high-frequency component Z high contains the edge and detail information in the depth feature map, and the low-frequency component Z low retains the global depth structure of the depth feature map; To optimize the utilization of high- and low-frequency components, an adaptive weighting method is adopted for feature fusion, and the fused feature Z FAF is as follows: Z FAF = α1Z high + α2Z low α1+α2=1 where α1 is a learnable parameter used to control the weight of high frequency in the depth feature map; α2 is a learnable parameter used to control the weight of the low-frequency component in the depth feature map; Step (2) specifically refers to: designing an adaptive depth bias compensation mechanism, which is composed of a depth distribution weight map w distance and an edge gradient weight map w edge to solve the problem of the long-tail effect of depth distribution; Depth distribution weight graph w distance Implemented by calculating the distribution density of pixel depth: where d is the original depth map, β is the smoothing parameter, and z mean represents the mean depth of each pixel in the original depth map d, and P is the percentile of the depth distribution; the edge gradient weight map w edge is then calculated by the Sobel operator to obtain the edge gradient magnitude of the original depth map d: where G x and G y are the gradients in the horizontal and vertical directions, respectively; Total loss function L total is as follows: Among them, N represents the number of training samples, and λL var is the variance regularization term, which is used to enhance the diversity of latent features; ∈ θ,i is the predicted noise of the i-th sample, ∈ t,i is the noise added to the i-th sample at time t; w combined = λ base w distance +(1 - λ base )w edge , where λ base is a hyperparameter, which is used to balance the contribution ratio of the depth distribution weight and the edge gradient weight; the dynamic adjustment of w combined can effectively adapt to the changes of different depth regions and edge characteristics, so as to optimize the depth estimation performance.
2. The monocular depth estimation method based on improved depth distribution compensation and frequency domain feature fusion according to claim 1, characterized in that: Step (3) specifically refers to: During the training process, a stable diffusion model is adopted, and depth feature extraction and denoising training are performed in combination with the monocular depth estimation model; First, the input RGB image x and the original depth map d are respectively encoded into the latent space: z x = E(x) z d = E(d) where z x is the RGB feature map of the latent space, and z d is the depth feature map of the latent space. E(x) represents the encoder for encoding the RGB image, and E(d) represents the encoder for encoding the depth map; The forward diffusion process gradually adds Gaussian noise to the deep feature map z of the latent space to generate deep latent representations containing different levels of noise d Among them, \(t\in\{1,\ldots,T\}\) is the current time step, representing the stage of the diffusion process; \(T\) is the maximum number of time steps; is the noise attenuation coefficient at time step \(t\), controlling the balance between the original depth map \(d\) and the noise; denotes Gaussian noise of the standard normal distribution; through forward diffusion, a depth latent representation containing more noise is gradually generated simulates the noise distribution that may appear in the actual input; The reverse denoising process works in cooperation with the monocular depth estimation model to gradually remove noise and restore a high-quality depth latent representation. The mathematical expression of the denoising process is: Among them, represents the noise predicted by the U-Net network, and the input is the depth latent representation at the current time step t and the RGB image representation z of the latent space x ; is the frequency-domain enhancement result of the frequency-aware fusion module for the depth features; is the depth latent representation generated by reverse diffusion; through the step-by-step denoising method, the depth latent representation gradually recovers from the high-noise state to the low-noise state, and finally generates a high-quality denoised depth feature map of the latent space After the reverse diffusion is completed, the decoder D of the variational autoencoder VAE is used to decode the denoised depth feature map of the latent space into the predicted depth map: Among them, is the predicted depth map.
3. The monocular depth estimation method based on improved depth distribution compensation and frequency domain feature fusion according to claim 1, characterized in that: Step (4) specifically includes the following steps in sequence: (4a) Latent space encoding: The input monocular RGB image to be estimated is first encoded by the encoder E of the frozen variational autoencoder VAE to generate a latent space representation; (4b) Initial noise generation: To initiate the reverse diffusion process, sample the initial noise from a standard normal distribution Initial noise Represents the initial state of the deep latent representation, which will be gradually restored to high-quality deep features through multi-step denoising later; (4c) Reverse diffusion denoising process: Remove noise through multiple iterations to generate a gradually optimized depth latent representation; (4d) Dynamic optimization through the adaptive depth bias compensation mechanism: Introduce the adaptive depth bias compensation mechanism to dynamically optimize the depth latent features; After the deep decoding iterative denoising process is completed, a depth feature map of the denoised latent space is finally generated. It is passed to the frozen VAE decoder D and reconstructed into a predicted depth map.
4. An electronic device, comprising: A processor; And A memory in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor executes the monocular depth estimation method based on improved depth distribution compensation and frequency domain feature fusion according to any one of claims 1-3.
5. A computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor executes the monocular depth estimation method based on improved depth distribution compensation and frequency domain feature fusion according to any one of claims 1-3.
Citation Information
Patent Citations
Image depth estimation algorithm based on deep learning and Fourier domain analysis
CN110969653A
Low-illumination image enhancement method based on feature fusion and attention embedding
CN116797488A