Video Tile Motion Prediction for Viewport-Adaptive HEVC Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for viewport-dependent omnidirectional video delivery, such as HEVC Scalable Extension (SHVC) and Constrained Inter-Layer Prediction (CILP), face limitations in hardware compatibility and streaming rate-distortion performance, particularly when finer tile grids are used.
Innovation Solution
An enhanced encoding method that includes encoding an input picture into a coded constituent picture, determining spatial region offsets, and applying motion vectors relative to prediction-unit anchors to derive prediction blocks, improving compatibility and compression efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If SHVC ROI encoding is used to deliver viewport-dependent omnidirectional video, then streaming rate-distortion performance is improved, but hardware compatibility deteriorates due to limited support for inter-layer prediction
Solution Approach 1:
The patent creates a base layer that is a simplified copy of the enhancement layer, encoded without inter-layer prediction dependencies. This base layer can be decoded by any standard HEVC decoder, while the enhancement layer provides the viewport-dependent quality improvements. The base layer acts as a compatible copy that works across all hardware platforms.
Solution Approach 2:
The video stream is segmented into two independent layers: a base layer that provides universal compatibility and an enhancement layer that provides viewport-adaptive quality. By separating the encoding into independent segments rather than relying on inter-layer prediction, the patent achieves both hardware compatibility and improved rate-distortion performance.
2Adaptability or versatility
If CILP encoding is used to deliver viewport-dependent omnidirectional video, then hardware compatibility is improved, but streaming rate-distortion performance deteriorates when finer tile grids are used
Solution Approach 1:
The patent adds a layer dimension to the video coding structure, creating a multi-layer hierarchy where the base layer and enhancement layer operate at different quality dimensions. This dimensional addition allows the system to achieve both compatibility (through the base layer) and superior rate-distortion performance (through the enhancement layer with finer tile grids).
3Ease of manufacture
If conventional HEVC encoding is used without motion-constrained tile sets, then encoding simplicity is maintained, but viewport-adaptive streaming capability is lost
Solution Approach 1:
The patent introduces dynamic viewport adaptation by encoding multiple quality representations of the same spatial region at different resolutions. The system dynamically selects which quality level to transmit based on the current viewport, allowing simple HEVC decoding while achieving adaptive streaming capability through the layered structure.
Data Source
AI summary
A method comprising: encoding an input picture into a coded picture; reconstructing a decoded picture corresponding to the coded picture; encoding a spatial region into a coded tile, the encoding comprising: determining a horizontal offset and a vertical offset indicative of a region-wise anchor position of the spatial region within the decoded picture; encoding the horizontal offset and the vertical offset; determining that a prediction unit at position of a first horizontal coordinate and a first vertical coordinate is predicted relative to the region-wise anchor position; indicating that the prediction unit is predicted relative to a prediction-unit anchor position; deriving a prediction-unit anchor position equal to sum of the first horizontal coordinate and the horizontal offset, and the first vertical coordinate and the vertical offset, respectively; and determining a motion vector for the prediction unit; and applying the motion vector relative to the prediction-unit anchor position to obtain a prediction block.


