3D Video Coding Depth Map Filtering for View Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional 3D video coding methods fail for non-coplanar camera arrangements due to inconsistent depth maps, leading to imperfect view synthesis prediction and reduced coding performance, as they rely on block-based approaches that do not account for varying pixel shifts between views.

Innovation Solution

An apparatus and method for encoding and decoding 3D video data using depth maps, which includes a depth map filter to detect edges and generate an auxiliary depth map by solving a boundary value problem, improving the consistency of depth information for view synthesis prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If block-based view synthesis prediction is used for coplanar camera arrangements, then coding performance is maintained, but the method fails for non-coplanar camera arrangements where pixel shifts vary between views

Engineering Contradiction:
Improveapplicability to different camera arrangementsVSAvoidprediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the depth map into multiple depth regions based on depth thresholds, allowing different prediction strategies to be applied to different regions. This enables the system to handle both coplanar and non-coplanar camera arrangements by adapting the prediction method to the local depth characteristics of each region.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically selects between block-based and pixel-based prediction methods based on the camera arrangement type and depth map characteristics. The system adapts its prediction approach in real-time depending on whether the scene contains coplanar or non-coplanar camera arrangements, optimizing performance for each case.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If pixel-based depth maps are used directly for view synthesis prediction in non-coplanar arrangements, then the method can handle varying pixel shifts, but the prediction accuracy deteriorates due to inconsistencies in estimated depth maps

Engineering Contradiction:
Improvehandling of non-coplanar camera arrangementsVSAvoiddepth map consistency
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary depth map enhancement through inter-layer filtering before using the depth map for view synthesis prediction. This preprocessing step improves the consistency and quality of the depth map, ensuring that pixel-based prediction methods produce accurate results for non-coplanar camera arrangements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an inter-layer filtering process as an intermediary step between depth map estimation and view synthesis prediction. This filtering process acts as a mediator that corrects inconsistencies in the estimated depth map, improving the overall prediction accuracy for non-coplanar arrangements.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If depth maps are enhanced through inter-layer filtering, then depth map consistency is improved for pixel-based prediction, but the processing complexity increases

Engineering Contradiction:
Improvedepth map consistencyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies inter-layer filtering selectively to specific depth regions rather than uniformly across the entire depth map. By identifying regions that benefit most from filtering and applying the enhancement only to those areas, the system improves depth map consistency while minimizing the increase in processing complexity.

Inventive Principle:
Principle #3Local quality

4Productivity

If 3D warping is used to map different views to one another, then view synthesis can be achieved, but occlusions occur in the warped view leading to imperfect mapping

Engineering Contradiction:
Improveview synthesis capabilityVSAvoidmapping accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the warped view into occluded and non-occluded regions, allowing different processing strategies to be applied to each. This segmentation enables the system to handle occlusions by identifying affected areas and applying appropriate corrections or alternative prediction methods to those regions.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3459251B1Devices and methods for 3D video coding
Publication Date: 2021.12.22 HUAWEI TECH CO LTD
  • EP3459251B1 patent drawingFigure 1a
  • EP3459251B1 patent drawingFigure 1b
  • EP3459251B1 patent drawingFigure 2a

AI summary

The invention relates to an apparatus (200) for decoding 3D video data, the 3D video data comprising a plurality of texture frames and a plurality of associated depth maps, the apparatus (200) comprising: a first texture decoder (201a) configured to decode a video coding block of a first texture frame associated with a first view; a first depth map decoder (201a) configured to decode a video coding block of a first depth map associated with the first texture frame; a depth map filter (219b) configured to generate an auxiliary depth map on the basis of the first depth map; a first view synthesis prediction unit (221b) configured to generate a predicted video coding block of a view synthesis predicted second texture frame associated with a second view on the basis of the video coding block of the first texture frame and the auxiliary depth map; and a second view synthesis prediction unit (217b) configured to generate a predicted video coding block of a view synthesis predicted second depth map on the basis of the first depth map, wherein the view synthesis predicted second depth map is associated with the view synthesis predicted second texture frame.