Depth Map Co-Processing for Free Viewpoint Television
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current depth map estimation techniques in free viewpoint television (FTV) suffer from low quality due to errors in occluded regions and sensor noise, leading to misplacement of pixels, especially around object boundaries, which affects the accuracy and quality of synthesized views.
Innovation Solution
A co-processing method that aligns video frames and depth maps to identify and correct incorrect pixel values by using edge detection and alignment to replace erroneous values with accurate ones from neighboring pixels, improving the quality and accuracy of depth maps and video frames.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If depth maps are estimated based on neighboring camera views, then the cost of specialized depth cameras is reduced, but the quality and accuracy of depth maps deteriorates due to sensor noise and estimation errors
Solution Approach 1:
The patent combines video frames and depth maps into a unified co-processing framework, merging color and depth information to mutually reinforce each other's accuracy. Video edges guide depth refinement while depth information enhances video segmentation, creating a synergistic effect that improves overall quality beyond what either modality could achieve alone.
Solution Approach 2:
The patent implements iterative refinement where video edge detection results feed into depth map improvement, and refined depth maps feed back into enhanced video processing. This closed-loop feedback mechanism continuously improves depth map accuracy by using video information as guidance and vice versa, progressively reducing estimation errors.
2Device complexity
If standard depth estimation algorithms are used, then processing complexity is reduced, but the quality of depth maps deteriorates especially around object boundaries and in occluded regions
Solution Approach 1:
The patent applies different processing strategies to different regions of the image based on their specific characteristics. Video edge detection is applied specifically at object boundaries to guide depth refinement where it is most needed, while occluded regions receive specialized handling through propagation from visible areas. This localized approach targets quality improvements where they matter most without uniformly increasing complexity across the entire image.
Solution Approach 2:
The patent segments the image into different regions (visible vs. occluded, edge regions vs. interior regions) and applies specialized processing to each segment. By dividing the depth map into regions requiring different refinement strategies, the system achieves high quality results in critical areas while maintaining computational efficiency through region-specific rather than全局 processing.
3Measurement precision
If brightness constraints are strictly imposed on video frames during depth map estimation, then estimation accuracy is improved, but the flexibility in processing and real-time performance deteriorates
Solution Approach 1:
The patent applies brightness constraints selectively rather than universally across all pixels and processing stages. By applying constraints only where and when they provide the most benefit (e.g., in visible regions, at certain processing iterations), the system achieves improved accuracy without the full computational burden of strict constraint enforcement throughout, thereby maintaining processing flexibility and real-time performance.
Data Source
AI summary
Co-processing of a video frame (32) and its associated depth map (34) suitable for free viewpoint television involves detecting respective edges (70, 71, 80, 81) in the video frame (32) and the depth map (34). The edges (70, 71, 80, 81) are aligned and used to identify any pixels (90-92) in the depth map (34) or the video frame (32) having incorrect depth values or color values based on the positions of the pixels in the depth map (34)or the video frame (32) relative an edge (80) in 5 the depth map (34) and a corresponding, aligned edge (70) in the video frame (32). The depth values or color values of the identified pixels (90-92) can then be corrected in order to improve the accuracy of the depth map (32) or video frame (34).


