Video Stream Tile Segmentation for Selective Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding and decoding methods require multiple encoders and decoders for independent video streams, leading to increased complexity, cost, and resource wastage, especially in mobile terminals, where they cannot efficiently decode or display subsets of video streams based on user selection or resource availability.
Innovation Solution
A method for coding a single container video stream that allows independent decoding of multiple component video streams using a single decoder, by partitioning the video into tiles and adding signaling metadata to indicate the content and relationships between tiles, enabling selective decoding and display of specific views or regions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple independent video streams are distributed using traditional coding methods, then each view can be independently decoded, but the device complexity and cost increase due to requiring multiple decoders
Solution Approach 1:
The video stream is segmented into independent tiles, where each tile corresponds to a specific view or region. The encoding structure divides the picture into multiple picture sections (tiles), allowing selective decoding of individual tiles without requiring multiple complete decoders. This segmentation enables independent view decoding while using a single decoder architecture.
Solution Approach 2:
The patent introduces a new dimensional organization by arranging multiple views in a tile-based spatial structure within a single video stream. Instead of treating each view as a separate stream, views are organized in a 2D tile grid structure that can be selectively accessed and decoded, transforming the problem from multiple independent streams to a structured single stream with selective accessibility.
2Quantity of substance
If a single video stream contains multiple component video streams, then bandwidth is reduced, but selective decoding of subsets is currently impossible
Solution Approach 1:
The single video stream is segmented into multiple independent tiles, each representing a specific view or region. This segmentation is achieved through picture section division and tile structure organization, allowing the decoder to selectively decode only the required tiles while ignoring others, thus enabling adaptive bandwidth utilization based on user needs.
Solution Approach 2:
The patent implements dynamic tile decoding capability where the decoder can adaptively select which tiles to decode based on user preferences, device capabilities, or network conditions. The tile structure allows flexible configuration of decodable regions, making the system dynamic rather than static in terms of content delivery.
3Adaptability or versatility
If multiple decoders are used to decode multiple views, then independent decoding is achieved, but energy consumption and computational resources are wasted
Solution Approach 1:
Instead of decoding the entire video stream containing all views, the system enables partial decoding where only the specific tiles corresponding to selected views are decoded. This partial action principle allows the decoder to process only the necessary portions of the data, significantly reducing energy consumption and computational resources while maintaining independent view decoding capability.
Solution Approach 2:
The patent enables extraction of specific tiles from the encoded video stream for selective decoding. The tile-based structure allows the decoder to identify and extract only the required view tiles, discarding or ignoring other tiles, thus taking out only the necessary information for processing and display.
4Quantity of substance
If traditional compression techniques are used on composite videos, then compression is achieved, but no specification-compliant decoder can independently decode component streams
Solution Approach 1:
The patent applies segmentation at the encoding stage by organizing the video data into a tile structure where each tile represents an independent decodable unit. This segmentation is integrated into the compression framework, allowing the compressed stream to maintain both overall compression efficiency and tile-level independence for selective decoding.
Solution Approach 2:
The tile-based structure serves multiple functions simultaneously: it enables compression of multiple views in a single stream, maintains specification-compliant decoding capability, and provides selective tile decoding functionality. This multi-functional design resolves the contradiction by making the same structure serve compression and independent decoding purposes.
Data Source
AI summary
A method is described for generating a video stream by starting from a plurality of sequences of 2D and/or 3D video frames, wherein a video stream generator composes into a container video frame video frames coming from N different sources (S1, S2, S3, SN) and generates a single output video stream of container video frames which is coded by an encoder, wherein said encoder enters into the output video stream a signalling adapted to indicate the structure of the container video frames. A corresponding method for regenerating the video stream is also described.


