Stereo Depth Image Generation via Adaptive Boundary Penalty Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional passive methods for generating dense depth images using stereo-based techniques often produce inaccurate results due to uniform smoothness constraints, leading to discontinuities in depth values, especially for objects with complex geometries like human fingers and hair strands, and are not scalable for large input images.
Innovation Solution
The technique optimizes the cost function by using customized penalty values for boundary pixels based on pixel-wise saliency information, allowing for relaxed depth continuity across object boundaries and preserving geometric details, and employs machine learning models to predict boundary pixels without requiring texture and classification information, enabling efficient computation and scalability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If uniform smoothness constraints are applied for all pixels in conventional passive methods, then the optimization process is simplified, but depth discontinuities occur within the same objects and accuracy deteriorates
Solution Approach 1:
The patent applies different smoothness constraint strengths to different regions of the image. Specifically, it identifies boundary pixels using machine learning models and applies relaxed smoothness constraints to these boundary regions while maintaining stronger constraints in non-boundary regions. This local differentiation allows depth discontinuities to be preserved at object boundaries while maintaining smoothness within homogeneous regions, thereby resolving the contradiction between optimization simplicity and depth map accuracy.
2Manufacturing precision
If conventional passive methods use pixel-wise color information for smoothness constraints, then some geometric details are preserved, but substantial discontinuities still occur within the same objects
Solution Approach 1:
The patent performs preliminary identification of boundary pixels using machine learning models before applying smoothness constraints. By pre-segmenting boundary regions and assigning them different constraint weights, the system prepares the optimization landscape in advance. This preliminary action allows the optimization process to naturally preserve both geometric details and depth continuity, as the boundary pixels are already marked for special treatment with relaxed constraints that prevent artificial discontinuities.
3Measurement precision
If active methods use high-precision hardware for measuring round trip time periods, then measurement accuracy is improved, but cost increases and scalability is reduced
Solution Approach 1:
The patent replaces active hardware-based depth measurement systems (like TOF or LIDAR) with a passive computational approach using stereo image pairs and optimization algorithms. Instead of using complex hardware to measure round trip time periods, the system uses machine learning models for boundary detection and iterative optimization with adaptive smoothness constraints to compute depth maps. This substitution maintains measurement accuracy while significantly reducing hardware complexity and cost.
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating a depth image, comprising obtaining data representing a first image generated by a first sensor and a second image generated by a second sensor, wherein each of the first and second images includes a plurality of pixels; determining, for each pixel of the plurality of pixels included in the first image, whether the pixel is a boundary pixel associated with a boundary of an object that is represented in the first image; determining, from a plurality of candidate penalty values and for each pixel in the first image, an optimized penalty value for the pixel; generating an optimized cost function for the first image based on the optimized penalty values for the plurality of pixels; and generating a depth image for the first image based on the optimized cost function.


