Hierarchical Track Structure for Multi-View VR Content Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding technologies face inefficiencies in processing and delivering 3D spherical content, particularly in virtual reality applications, due to the need to process and transmit the entire spherical content, which is computationally intensive and bandwidth-heavy, even though users typically view only a portion of the content.

Innovation Solution

The implementation of a hierarchical track structure that supports multi-view data operations, such as frame packing, by associating metadata with the track hierarchy and performing derivation operations to generate media data for specific views, allowing for efficient encoding and decoding of stereoscopic content by leveraging a single track hierarchy instead of separate hierarchies for each view.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the entire spherical content is processed and transmitted, then the user can view content at any viewport, but the processing becomes computationally intensive and consumes significant bandwidth

Engineering Contradiction:
Improveviewport flexibilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the spherical content into multiple view-specific tracks, where each track contains content for a specific viewport direction. This allows the system to transmit only the necessary segments based on user viewport, reducing overall processing complexity and bandwidth requirements while maintaining viewport flexibility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic track selection based on user viewport orientation. The system dynamically determines which view-specific tracks to transmit and decode based on the current viewport, allowing adaptability while optimizing processing and bandwidth usage for each user session.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If separate track hierarchies are used for each view, then view-specific processing is simplified, but the overall system complexity and metadata redundancy increase

Engineering Contradiction:
Improveview processing simplicityVSAvoidtrack hierarchy complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent merges multiple view-specific tracks into a single unified track hierarchy structure. This consolidation reduces metadata redundancy by sharing common hierarchical elements across views while maintaining the organizational benefits of view-specific processing, thus simplifying the overall system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal track hierarchy that serves multiple views through a single structure. The hierarchical track system is designed to be multi-functional, supporting different view-specific operations while maintaining a unified metadata framework that reduces redundancy and simplifies processing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of energy

If viewport-dependent processing is implemented, then bandwidth efficiency is improved by delivering only viewed content, but the encoding and decoding complexity increases

Engineering Contradiction:
Improvebandwidth efficiencyVSAvoidencoding complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent performs preliminary organization of content into view-specific tracks during encoding. This preliminary structuring enables efficient viewport-dependent delivery without requiring complex real-time encoding decisions, as the content is pre-organized for rapid selection and transmission based on user viewport.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10869016B2Methods and apparatus for encoding and decoding virtual reality content
Publication Date: 2020.12.15 MEDIATEK SINGAPORE PTE LTD
  • US10869016B2 patent drawing
  • US10869016B2 patent drawing
  • US10869016B2 patent drawing

AI summary

The techniques described herein relate to methods, apparatus, and computer readable media configured to process multi-view multimedia data. The multi-view multimedia data includes a hierarchical track structure with at least a first track, wherein the first track is at a first level in the hierarchical track structure, and includes data for both a first view and a second view of the multi-view multimedia data. Metadata contained within the first track can be determined and used to perform an extraction operation on the first track to generate first media data of a second track, wherein the first media data is for the first view, and second media data of a third track, wherein the second media data is for the second view, wherein the second track and the third track are at a second level in the hierarchical track structure above the first level of the first track.