Volumetric Video Encoding via View-Dependent Patch Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in efficiently encoding and decoding volumetric video content due to its large data requirements, which hinder storage and transmission efficiency, especially for immersive 6DoF experiences that demand high bit-rates and storage capacities.

Innovation Solution

The method involves defining a reference viewing box and an intermediate viewing box within a 3D scene, encoding a central view and peripheral patches, and using metadata to describe these boxes, allowing for differential encoding and decoding that reduces data complexity and storage needs without increasing encoding or decoding complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If volumetric video content is encoded with high quality for 6DoF experiences, then viewing quality and immersion are improved, but data size and storage requirements increase significantly

Engineering Contradiction:
Improveviewing qualityVSAvoiddata size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The 3D scene is divided into multiple view-dependent image sets, each corresponding to a specific viewing bounding box. Within each set, images are segmented into patches that can be selectively encoded and transmitted. This segmentation allows the system to provide high-quality viewing experience for specific viewpoints while avoiding the need to encode and store all possible views at full quality, thus reducing overall data size.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different encoding qualities to different regions of the 3D scene based on their importance and view-dependency. Central patches that are visible from multiple viewpoints are encoded with higher quality, while peripheral patches are encoded with lower quality or omitted entirely. This local quality differentiation maintains viewing quality for critical regions while reducing data size for less important areas.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If complete 6DoF volumetric video is provided, then user freedom and immersion are improved, but encoding complexity and computational resources increase

Engineering Contradiction:
Improveuser freedomVSAvoidencoding complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a dynamic encoding scheme where the level of detail and number of view-dependent image sets are adjusted based on the user's current viewpoint and movement within the 3D scene. As the user moves, the system dynamically selects which patches to encode and transmit, providing 6DoF freedom when needed while reducing encoding complexity when the user is stationary or viewing from limited angles.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Instead of encoding complete volumetric data for all possible viewpoints, the patent encodes only the necessary portions (patches) that are visible from the current or predicted viewpoints. This partial action approach provides sufficient 6DoF experience for practical use cases while avoiding the excessive computational complexity of encoding all possible views.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If multiple view-dependent image sets are encoded for different bounding boxes, then parallax and 3DoF+ experience are improved, but transmission bandwidth and processing requirements increase

Engineering Contradiction:
Improveparallax experienceVSAvoidtransmission efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent merges multiple view-dependent image sets into a unified data structure organized by viewing bounding boxes. Patches from different viewpoint sets are combined and encoded together, allowing the system to transmit consolidated data that provides parallax effects while reducing redundant information. This merging improves transmission efficiency by eliminating duplicate data across multiple image sets.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from encoding complete 2D images to encoding 3D patches with spatial coordinates and view-dependency metadata. This dimensional change allows the system to organize data in a multi-dimensional structure that efficiently represents multiple viewpoints, improving transmission efficiency by encoding only the essential 3D geometric information rather than redundant 2D image data for each viewpoint.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11968349B2Method and apparatus for encoding and decoding of multiple-viewpoint 3DoF+ content
Publication Date: 2024.04.23 INTERDIGITAL VC HOLDINGS INC
  • US11968349B2 patent drawing
  • US11968349B2 patent drawing
  • US11968349B2 patent drawing

AI summary

A method for encoding a volumetric video content representative of a 3D scene is disclosed. The method comprises obtaining a reference viewing box and an intermediate viewing box defined within the 3D scene. For the reference viewing bounding box, the volumetric video reference subcontent is encoded as a central image and peripheral patches for parallax. For the intermediate viewing bounding box, the volumetric video intermediate sub-content is encoded as intermediate central patches which are differences between the intermediate central image and the reference central image.