Volumetric Video Viewport Synthesis Using Depth Confidence Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for encoding and decoding multi-view plus depth (MVD) frames lack a method to determine the confidence in the information carried by different views, leading to inefficiencies in viewport synthesis, particularly in immersive 6DoF video experiences where accurate depth fidelity is crucial for navigation and immersion.

Innovation Solution

A method is introduced that determines and encodes a parameter representative of the fidelity of depth information for each view, using intrinsic and extrinsic camera parameters, and associates this metadata with the multi-views frame for improved viewport synthesis, allowing for trustworthiness indicators to prioritize reliable camera contributions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple views are combined for viewport synthesis, then the immersion and navigation capability are improved, but the rendering quality deteriorates due to inconsistent depth information from different views

Engineering Contradiction:
Improvenavigation capabilityVSAvoidrendering quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies local quality by assigning different confidence levels to different views based on their depth information quality. Each view is evaluated individually using depth map metrics (depth variance, depth gradient, occlusion ratio) and assigned a confidence weight, allowing the system to selectively trust certain views more than others in the viewport synthesis process, thereby resolving the contradiction between using multiple views and maintaining rendering quality

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameter of view contribution weighting by introducing confidence levels derived from depth information quality metrics. Instead of treating all views equally, the system dynamically adjusts the weight of each view based on parameters such as depth variance, depth gradient, and occlusion ratio, enabling adaptive viewport synthesis that maintains high rendering quality while utilizing multiple views for navigation

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If all views are treated equally in viewport synthesis, then the processing complexity is reduced, but the depth fidelity deteriorates leading to visual artifacts

Engineering Contradiction:
Improveprocessing complexityVSAvoiddepth fidelity
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by pre-evaluating the quality of depth information for each view before the viewport synthesis process. Confidence levels are calculated in advance based on depth map metrics (depth variance, depth gradient, occlusion ratio), allowing the synthesis algorithm to efficiently weight views without adding significant processing complexity during the actual rendering phase, thus maintaining depth fidelity without substantially increasing complexity

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If confidence levels are calculated for each view, then the viewport synthesis quality is improved, but the encoding complexity increases due to additional metadata

Engineering Contradiction:
Improveviewport synthesis qualityVSAvoidencoding complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses confidence levels as an intermediary parameter that mediates between the multiple views and the final viewport synthesis. These confidence levels are derived from simple depth map metrics (variance, gradient, occlusion ratio) and serve as weighting factors in the synthesis process, improving viewport quality without requiring complex encoding schemes since the confidence calculation is based on readily available depth information

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20220345681A1Method and apparatus for encoding, transmitting and decoding volumetric video
Publication Date: 2022.10.27 INTERDIGITAL CE PATENT HOLDINGS SAS
  • US20220345681A1 patent drawing
  • US20220345681A1 patent drawing
  • US20220345681A1 patent drawing

AI summary

Methods, devices and stream for encoding, decoding and transmitting a multi-views frame are disclosed. In a multi-views frame some of the views are more trustable than others. The multi-views frame is encoded in a data stream in association with metadata that comprise, for at least one of the views, a parameter indicating a degree of confidence in the information carried by this view. This information is used at the decoding side to determine the contribution of the view when synthetizing pixels of a viewport frame for a given point of view in the 3D space.