Segmentation-Guided Depth Map Refinement for Clearer Boundaries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional depth estimation systems suffer from inaccuracies, inefficiencies, and inflexibility in generating depth maps, leading to blurry boundaries, missing structures, and increased computational burdens in downstream tasks.
Innovation Solution
A depth refinement system utilizing digital segmentation masks to guide a depth refinement machine learning model, employing a self-supervised learning scheme with RGB-D datasets for training, and a layered depth refinement approach to improve accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional machine learning models are used for depth estimation, then depth maps can be generated, but accuracy is poor with blurry boundaries and missing structures
Solution Approach 1:
The patent introduces segmentation masks as an intermediary element that guides the depth refinement process. These masks provide structural information about object boundaries and regions, acting as a mediator between the initial depth map and the refined depth map. The refinement model uses these masks to identify areas needing improvement and apply appropriate refinement strategies, thereby resolving the contradiction between overall accuracy and boundary clarity.
Solution Approach 2:
The patent segments the depth refinement process into multiple layers and stages. The initial depth map is processed through a refinement model that operates in distinct phases: region classification based on segmentation masks, selective refinement of specific areas (foreground, background, boundaries), and composite assembly. This segmentation approach allows different refinement strategies to be applied to different regions, improving both overall accuracy and boundary clarity simultaneously.
2Productivity
If conventional depth estimation systems are used, then depth maps can be generated, but computational burden increases for downstream tasks
Solution Approach 1:
The patent performs depth refinement as a preliminary action before downstream tasks are executed. By pre-processing the initial depth map through the refinement model and producing a high-quality refined depth map with clear boundaries and accurate structures, the system eliminates the need for downstream tasks to perform additional corrective computations. This preliminary refinement reduces the overall computational burden across the entire processing pipeline.
Solution Approach 2:
The patent replaces computationally intensive post-processing operations in downstream tasks with a dedicated refinement model that produces optimized depth maps upfront. Instead of having multiple downstream tasks each attempt to correct depth map deficiencies through their own computational methods, the system substitutes this with a single specialized refinement process that pre-resolves boundary clarity and structural accuracy issues, reducing total computational energy consumption.
3Adaptability or versatility
If conventional depth estimation models are used, then depth maps can be generated, but the system lacks flexibility across different model architectures
Solution Approach 1:
The patent designs the refinement model with universal input and output interfaces that can accept depth maps from various conventional estimation models regardless of their internal architectures. The model uses standardized segmentation mask inputs and produces refined depth maps in a consistent format, enabling it to serve multiple different upstream models. This universality allows the system to be flexibly configured with different model architectures without requiring complex customizations for each case.
Data Source
AI summary
The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing machine learning models to generate refined depth maps of digital images utilizing digital segmentation masks. In particular, in one or more embodiments, the disclosed systems generate a depth map for a digital image utilizing a depth estimation machine learning model, determine a digital segmentation mask for the digital image, and generate a refined depth map from the depth map and the digital segmentation mask utilizing a depth refinement machine learning model. In some embodiments, the disclosed systems generate first and second intermediate depth maps using the digital segmentation mask and an inverse digital segmentation mask and merger the first and second intermediate depth maps to generate the refined depth map.


