3D Video Signal Encoding with Depth Map Subsampling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing 3-D video technologies face challenges in providing satisfactory rendering across a variety of display devices, leading to inconsistent depth effects and artifacts when the rendering context differs from the intended context, such as from a cinema to a home environment.

Innovation Solution

The method involves encoding image frames and extra frames using spatial and temporal subsampling of depth components, combining depth and occlusion information, and using a depth map to generate shifted viewpoint pictures, allowing for flexible and efficient rendering on different devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If 3-D video is generated for a particular rendering context (e.g., cinema), then the depth effects are optimized for that context, but the rendering quality deteriorates when displayed on different devices (e.g., home video set)

Engineering Contradiction:
Improvedepth effects accuracyVSAvoidrendering context adaptability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The 3-D video signal is segmented into multiple components: a base stereo picture pair and multiple depth maps corresponding to different rendering contexts. Each depth map is specifically optimized for a particular rendering context (e.g., cinema, home video set), allowing the system to select and apply the appropriate depth map based on the target display device, thereby maintaining depth effects accuracy across different contexts without requiring complete signal regeneration.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If interpolation or extrapolation is used to adjust depth effects for different rendering contexts, then the adaptability improves, but the computational complexity and cost increase

Engineering Contradiction:
Improverendering context adaptabilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Depth maps for different rendering contexts are pre-computed and stored during the 3-D video generation phase. When the 3-D video is played back on different devices, the system simply selects and applies the pre-prepared depth map corresponding to the target rendering context, eliminating the need for real-time interpolation or extrapolation calculations. This preliminary preparation significantly reduces processing complexity while maintaining adaptability across various display devices.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If real-time depth adjustment is performed through interpolation or extrapolation, then the adaptability to different rendering contexts improves, but perceptible artifacts are introduced

Engineering Contradiction:
Improverendering context adaptabilityVSAvoidimage quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

Depth maps for different rendering contexts are pre-computed and stored during the 3-D video generation phase. When the 3-D video is played back on different devices, the system simply selects and applies the pre-prepared depth map corresponding to the target rendering context, eliminating the need for real-time interpolation or extrapolation calculations. This preliminary preparation significantly reduces processing complexity while maintaining adaptability across various display devices.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3101894B1Versatile 3-d picture format
Publication Date: 2021.06.23 KONINKLIJKE PHILIPS NV
  • EP3101894B1 patent drawingFigure 1~2
  • EP3101894B1 patent drawingFigure 3~5
  • EP3101894B1 patent drawingFigure 6~8

AI summary

A 3-D picture signal is provided as follows. An image and depth components comprising a depth map for the image are provided, the depth map comprising depth indication values, a depth indication value relating to a particular portion of the image and indicating a distance between an object at least partially represented by that portion of the image and the viewer. The signal conveys the 3-D picture according to a 3D format having image frames encoding the image. Extra frames (D, D') are encoded that provide the depth components and further data for use in rendering based on the image and the depth components. The extra frames are encoded using spatial and/or temporal subsampling of the depth components and the further data, while the extra frames are interleaved with the image frames in the signal in a Group of Pictures coding structure (GOP).