Motion-Constrained Tiles for 360-Degree Video Partial Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding techniques face challenges in efficiently compressing 360-degree video data, leading to high bandwidth requirements and increased processing burdens, especially when dealing with viewport-dependent and most interested regions, which are not adequately addressed by existing methods.
Innovation Solution
The proposed solution involves generating and processing media files for viewport-dependent 360-degree video content by dividing video streams into motion-constrained tiles, with specific techniques for encoding and decoding only the tiles required for the current viewport and defining most interested regions for adaptive streaming, transcoding, and cache management, using formats like DASH and ISOBMFF.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If video coding techniques are used to compress video data, then video quality can be maintained with lower bitrate, but the complexity of encoding and decoding processes increases
Solution Approach 1:
The patent divides 360-degree video content into multiple motion-constrained tiles, allowing selective encoding and decoding of only those tiles that fall within the user's current viewport. This segmentation enables partial decoding, where the decoder processes only a subset of the total video data, significantly reducing processing burden and bandwidth requirements while maintaining full viewport quality.
Solution Approach 2:
The patent extracts and identifies the most interested regions (MIR) within the viewport based on user statistics or user-defined parameters. By extracting only these critical regions for full-quality encoding while using lower quality for other viewport areas, the system reduces overall bandwidth requirements while maintaining perceived video quality where users are most likely to focus attention.
2Reliability
If all video data is transmitted for high quality playback, then video quality is maximized, but bandwidth consumption and processing loads increase significantly
Solution Approach 1:
The patent applies different quality levels to different regions of the video content based on their importance. Motion-constrained tiles within the viewport are encoded at high quality, while tiles outside the viewport are encoded at lower quality or not transmitted at all. This local quality differentiation maintains video quality where needed while reducing overall data transmission volume.
Solution Approach 2:
The patent performs preliminary encoding of the entire 360-degree video content into motion-constrained tiles before transmission. This preliminary action organizes the data structure so that during playback, only the necessary tiles for the current viewport need to be decoded and transmitted, rather than decoding and transmitting all video data in real-time.
3Productivity
If viewport-dependent encoding is implemented for 360-degree video, then bandwidth efficiency improves, but the complexity of media file generation and processing increases
Solution Approach 1:
The patent implements dynamic adaptation sets in the media file structure that allow the decoder to dynamically select which motion-constrained tiles to decode based on the user's current viewport orientation. This dynamic approach enables bandwidth efficiency by transmitting only necessary data while managing processing complexity through pre-organized tile structures and adaptation set metadata.
Solution Approach 2:
The patent introduces motion-constrained tiles as an intermediary structure between the encoded video data and the decoder. These tiles serve as intermediate units that can be independently processed, allowing the system to manage complexity by working with discrete tile units rather than entire video frames, while improving bandwidth efficiency through selective transmission.
4Device complexity
If motion-constrained tiles are used for partial decoding, then processing loads are reduced, but the precision of region identification and tile mapping increases
Solution Approach 1:
The patent segments video content into motion-constrained tiles with well-defined boundaries and spatial relationships. This segmentation provides a structured framework that simplifies region identification, as each tile can be independently identified and mapped to viewport coordinates, reducing the precision requirements for overall region identification while maintaining processing efficiency.
Data Source
AI summary
Techniques and systems are provided for processing video data. For example, 360-degree video data can be obtained for processing by an encoding device or a decoding device. The 360-degree video data includes pictures divided into motion-constrained tiles. The 360-degree video data can be used to generate a media file including several tracks. Each of the tracks contain a set of at least one of the motion-constrained tiles. The set of at least one of the motion-constrained tiles corresponds to at least one of several viewports of the 360-degree video data. A first tile representation can be generated for the media file. The first tile representation encapsulates a first track among the several tracks, and the first track includes a first set of at least one of the motion-constrained tiles at a first tile location in the pictures of the 360-degree video data. The first set of at least one of the motion-constrained tiles corresponds to a viewport of the 360-degree video data.


