Viewport Signaling in Point Cloud Video via Region of Interest Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding techniques for omnidirectional content, such as 360-degree video, are inefficient as they process and deliver the entire spherical content, leading to high computational and bandwidth requirements, especially when only a portion of the content is viewed by the user.

Innovation Solution

The method involves encoding and decoding point cloud video data by specifying regions of interest (ROIs) using metadata, allowing for the processing and delivery of only the relevant portions of the content based on user interaction, such as viewpoint and viewport, using 6D spherical and Cartesian coordinates to define these regions within ISOBMFF files.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the entire spherical content is processed and delivered, then the user can view content at any viewport, but the computational load and bandwidth usage increase significantly

Engineering Contradiction:
Improveviewport flexibilityVSAvoidcomputational load
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent divides the spherical content into multiple rectangular regions of interest (ROIs) that can be independently encoded and delivered. Instead of processing the entire sphere, the system segments the content into manageable regions that correspond to potential viewports, allowing selective delivery based on user needs while reducing overall computational burden.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial action by delivering only the necessary portions of spherical content rather than the entire sphere. The system encodes multiple rectangular ROIs that cover the sphere, and only the relevant ROIs are delivered to the client based on predicted or actual viewport positions, significantly reducing bandwidth and computational requirements.

Inventive Principle:
Principle #16Partial or excessive action

2Reliability

If the entire spherical content is processed and delivered, then complete coverage is achieved, but bandwidth consumption increases

Engineering Contradiction:
Improvecontent coverageVSAvoidbandwidth usage
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The spherical content is segmented into multiple rectangular regions of interest that can be independently encoded and delivered. This segmentation allows the system to provide complete coverage by delivering only the specific regions needed for the user's viewport, rather than transmitting the entire spherical content, thus reducing bandwidth consumption while maintaining content coverage reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial action by delivering only the necessary portions of spherical content corresponding to predicted or actual viewports. Multiple rectangular ROIs are pre-encoded to ensure complete coverage potential, but only the relevant subsets are transmitted, optimizing bandwidth usage while preserving the ability to provide complete coverage when needed.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If viewport-dependent processing is implemented, then bandwidth and compute resources are optimized, but the system complexity increases

Engineering Contradiction:
Improveresource efficiencyVSAvoidencoding complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing spherical content into multiple rectangular regions of interest with well-defined boundaries and coordinate systems. This segmentation simplifies the encoding process by allowing independent processing of each region using standard rectangular encoding techniques, reducing overall system complexity while enabling efficient viewport-dependent delivery.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements partial action by encoding and delivering only the necessary rectangular regions corresponding to predicted or actual viewports. This approach optimizes resource efficiency by avoiding processing of unnecessary content, while the use of standardized rectangular encoding methods keeps the implementation complexity manageable compared to processing entire spherical content.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11200700B2Methods and apparatus for signaling viewports and regions of interest for point cloud multimedia data
Publication Date: 2021.12.14 MEDIATEK SINGAPORE PTE LTD
  • US11200700B2 patent drawing
  • US11200700B2 patent drawing
  • US11200700B2 patent drawing

AI summary

The techniques described herein relate to methods, apparatus, and computer readable media configured to encode and/or decode video data. Point cloud video data is received that includes metadata specifying one or more regions of interest of the point cloud video data. A first region of interest is determined from the one or more regions of interest. A portion of the point cloud video data associated with the first region of interest is determined. Point cloud media is generated for viewing by a user based on the determined portion of the point cloud video data associated with the first region of interest.