3D Video Signal Encoding with Depth Map Subsampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing 3-D video technologies face challenges in providing satisfactory rendering across a variety of display devices, leading to inconsistent depth effects and artifacts when the rendering context differs from the intended context, such as from a cinema to a home environment.
Innovation Solution
The method involves encoding image frames and extra frames using spatial and temporal subsampling of depth components, combining depth and occlusion information, and using a depth map to generate shifted viewpoint pictures, allowing for flexible and efficient rendering on different devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If 3-D video is generated for a particular rendering context (e.g., cinema), then the depth effects are optimized for that context, but the rendering quality deteriorates when displayed on different devices (e.g., home video set)
Solution Approach 1:
The 3-D video signal is segmented into multiple components: a base stereo picture pair and multiple depth maps corresponding to different rendering contexts. Each depth map is specifically optimized for a particular rendering context (e.g., cinema, home video set), allowing the system to select and apply the appropriate depth map based on the target display device, thereby maintaining depth effects accuracy across different contexts without requiring complete signal regeneration.
2Adaptability or versatility
If interpolation or extrapolation is used to adjust depth effects for different rendering contexts, then the adaptability improves, but the computational complexity and cost increase
Solution Approach 1:
Depth maps for different rendering contexts are pre-computed and stored during the 3-D video generation phase. When the 3-D video is played back on different devices, the system simply selects and applies the pre-prepared depth map corresponding to the target rendering context, eliminating the need for real-time interpolation or extrapolation calculations. This preliminary preparation significantly reduces processing complexity while maintaining adaptability across various display devices.
3Adaptability or versatility
If real-time depth adjustment is performed through interpolation or extrapolation, then the adaptability to different rendering contexts improves, but perceptible artifacts are introduced
Solution Approach 1:
Depth maps for different rendering contexts are pre-computed and stored during the 3-D video generation phase. When the 3-D video is played back on different devices, the system simply selects and applies the pre-prepared depth map corresponding to the target rendering context, eliminating the need for real-time interpolation or extrapolation calculations. This preliminary preparation significantly reduces processing complexity while maintaining adaptability across various display devices.
Data Source
Figure 1~2
Figure 3~5
Figure 6~8
AI summary
A 3-D picture signal is provided as follows. An image and depth components comprising a depth map for the image are provided, the depth map comprising depth indication values, a depth indication value relating to a particular portion of the image and indicating a distance between an object at least partially represented by that portion of the image and the viewer. The signal conveys the 3-D picture according to a 3D format having image frames encoding the image. Extra frames (D, D') are encoded that provide the depth components and further data for use in rendering based on the image and the depth components. The extra frames are encoded using spatial and/or temporal subsampling of the depth components and the further data, while the extra frames are interleaved with the image frames in the signal in a Group of Pictures coding structure (GOP).