Sub-picture Base Track Signaling for Video Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding standards, such as ISOBMFF, lack mechanisms to indicate the relationship among sub-picture tracks and the spatial resolution of full video content, making it difficult for file parsers to determine which sub-picture tracks represent the complete video content and requiring all tracks to be fetched for spatial resolution determination.

Innovation Solution

The introduction of a sub-picture base track and track grouping type in the ISO base media file format, which allows for the signaling of spatial resolution and relationship among sub-picture tracks, enabling efficient encoding and decoding of sub-picture bitstreams, particularly in virtual reality applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If all sub-picture tracks are fetched to determine spatial resolution, then accurate spatial resolution determination is achieved, but transmission bandwidth and decoding complexity increase

Engineering Contradiction:
Improvespatial resolution determination accuracyVSAvoidtransmission bandwidth and decoding complexity
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent extracts the spatial resolution information from all sub-picture tracks and places it in a dedicated metadata structure (stsd box) within the base track. This allows the resolution information to be taken out from the actual track data, enabling parsers to obtain spatial resolution without fetching all sub-picture tracks, thus reducing bandwidth and decoding complexity while maintaining measurement accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary action by pre-calculating and storing the spatial resolution of the full picture in the metadata during the encoding phase. This preliminary computation of resolution information allows the decoding device to obtain spatial resolution immediately from the metadata without having to process or fetch all sub-picture tracks, thereby reducing both transmission bandwidth requirements and decoding complexity.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If relationship among sub-picture tracks is not signaled, then file format simplicity is maintained, but ability to determine complete video content representation is lost

Engineering Contradiction:
Improvefile format structureVSAvoidtrack relationship information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary metadata structure (stsd box with track group type) that mediates between the multiple sub-picture tracks and the parser. This intermediary contains the relationship information and spatial resolution data, allowing the parser to understand track relationships without increasing the fundamental file format complexity. The metadata acts as a bridge that organizes track relationships in a standardized, easily accessible manner.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If sub-picture tracks are independently coded, then encoding flexibility and adaptability are improved, but difficulty in determining spatial resolution increases

Engineering Contradiction:
Improveencoding flexibilityVSAvoidspatial resolution detection
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent merges the spatial resolution information from all independently coded sub-picture tracks into a single unified metadata structure associated with the base track. This combining of resolution information allows the parser to detect and measure the spatial resolution of the complete video content by accessing one consolidated metadata source rather than attempting to analyze multiple independent tracks, thereby reducing detection difficulty while preserving encoding flexibility.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11062738B2Signalling of video content including sub-picture bitstreams for video coding
Publication Date: 2021.07.13 QUALCOMM INC
  • US11062738B2 patent drawing
  • US11062738B2 patent drawing
  • US11062738B2 patent drawing

AI summary

In various implementations, modifications and/or additions to the ISOBMFF are provided to process video data. A plurality of sub-picture bitstreams are obtained from memory, each sub-picture bitstream including a spatial portion of the video data and each sub-picture bitstream being independently coded. In at least one file, the plurality of sub-picture bitstreams are respectively stored as a plurality of sub-picture tracks. Metadata describing the plurality of sub-picture tracks is stored in a track box within a media file in accordance with a file format. A sub-picture base track is provided that includes the metadata describing the plurality of sub-picture tracks.