Viewport Dependent Video Coding Bitstream Merging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Viewport-dependent video coding for VR experiences faces challenges in maintaining image quality and file size reduction due to inter-prediction errors caused by resolution mismatches between sub-picture video streams, leading to decoding errors and artifacts.
Innovation Solution
The method involves encoding spherical video sequences into self-referenced sub-picture bitstreams that can be merged using a lightweight bitstream rewriting process, ensuring each bitstream is self-contained and temporally synchronized, with signaling mechanisms like SEI messages to indicate compatibility with multi-bitstream merge functions, thereby preventing decoding errors and allowing for efficient transmission and display.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If viewport-dependent video coding is used to reduce file size, then transmission efficiency is improved, but inter-prediction errors occur due to resolution mismatches between sub-picture streams
Solution Approach 1:
The video stream is divided into multiple sub-picture bitstreams, each representing a different viewport region. Each sub-picture stream is independently encoded with self-contained reference frames, allowing selective transmission based on user viewport while maintaining decoding reliability through independent reference picture sets.
Solution Approach 2:
A mergeable indication syntax element is introduced as an intermediary signal in the bitstream to inform the decoder that sub-picture streams are compatible for merging. This intermediary signaling mechanism enables the decoder to correctly combine multiple sub-picture streams without introducing artifacts, resolving the reliability issue while maintaining transmission efficiency.
2Adaptability or versatility
If multiple sub-picture bitstreams are merged to support viewport coding, then adaptability is improved, but device complexity increases due to merging processing requirements
Solution Approach 1:
The encoder performs preliminary actions by ensuring each sub-picture bitstream is self-contained with complete reference picture sets and by embedding mergeable indication syntax elements during encoding. This preliminary preparation eliminates the need for complex runtime merging operations, reducing decoder complexity while maintaining high adaptability for various viewport configurations.
Solution Approach 2:
The patent changes the parameter representation by introducing syntax elements that indicate mergeability and self-contained properties of sub-picture streams. These parameter changes enable the decoder to identify and process mergeable streams through simple syntax checking rather than complex analysis, reducing computational complexity while preserving viewport adaptability.
3Loss of substance
If sub-picture streams are encoded at different resolutions to optimize bandwidth, then loss of substance is reduced, but inter-prediction artifacts increase due to resolution mismatches
Solution Approach 1:
The video content is segmented into multiple sub-picture streams with different resolutions optimized for their respective viewport regions. Each segment is independently encoded with its own reference picture set, preventing inter-prediction artifacts that would occur if higher-resolution reference pictures from other streams were used, while still achieving overall bandwidth optimization.
Solution Approach 2:
Different sub-picture streams are encoded at different resolutions according to their local importance and viewport usage patterns. This local quality adaptation allows bandwidth optimization for less critical regions while maintaining high quality for central viewport areas, without introducing artifacts because each region uses only its own resolution-appropriate reference pictures.
Data Source
AI summary
A video coding mechanism for viewpoint dependent video coding is disclosed. The mechanism includes mapping a spherical video sequence into a plurality of sub-picture video sequences. The mechanism further includes encoding the plurality of sub-picture video sequences as sub-picture bitstreams to support merging of the plurality of sub-picture bitstreams, the encoding ensuring that each sub-picture bitstream is self-referenced and two or more of the sub-picture bitstreams can be merged to generate a single video bitstream using a lightweight bitstream rewriting process that does not involve changing of any block-level coding results. A mergeable indication is encoded to indicate that the sub-picture bitstream containing the indication is compatible with a multi-bitstream merge function for reconstruction of the spherical video sequence. A set of the sub-picture bitstreams and the mergeable indication are transmitted toward the decoder to support decoding and displaying a virtual reality video viewport.


