Spatially Segmented Composite Video Stream Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video streaming technologies face challenges in decoding multiple video streams simultaneously, especially for lower-end devices with limited hardware decoders, and lack flexibility in dynamically changing the composition of composite video streams.
Innovation Solution
A combiner system that combines multiple video streams into a spatially segmented composite video stream, generating composition metadata and identification metadata to enable dynamic processing and identification of individual video streams, allowing for flexible rendering and processing by the receiver device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If multiple video streams are transmitted separately to a receiver device, then each video stream can be independently decoded, but the receiver device requires multiple hardware decoders which increases device complexity and cost
Solution Approach 1:
Multiple separate video streams are merged into a single composite video stream where each original video stream is represented as an independently decodable spatial segment or tile. This allows the receiver device to use a single hardware decoder to process all video streams simultaneously, eliminating the need for multiple decoder instances while maintaining independent decodability of each video source.
2Device complexity
If video streams are combined into a composite video stream with fixed composition, then decoding is simplified, but the system lacks flexibility to dynamically change the composition of video streams
Solution Approach 1:
The composition of video streams in the composite video stream is made dynamic rather than fixed. The system allows for runtime reconfiguration where video streams can be added, removed, or repositioned within the composite stream. Metadata associated with each spatial segment enables the receiver device to dynamically adapt to composition changes without requiring system reconfiguration, maintaining both decoding simplicity and compositional flexibility.
Solution Approach 2:
The composite video stream is segmented into independently decodable spatial segments (tiles), where each segment corresponds to a specific video stream. This segmentation allows individual segments to be independently processed, added, or removed from the composite stream while maintaining the integrity of other segments, enabling dynamic composition changes without affecting the entire stream.
3Adaptability or versatility
If multiple video streams are received separately, then each stream can be processed independently, but lower-end receiver devices with limited hardware decoders cannot efficiently decode multiple streams simultaneously
Solution Approach 1:
Multiple video streams are combined into a single composite video stream that can be decoded by a single hardware decoder instance. This merging approach enables lower-end receiver devices with limited hardware decoding capability to efficiently process multiple video streams simultaneously, as they only need one decoder rather than multiple separate decoders, thereby improving both device compatibility and decoding efficiency.
Data Source
AI summary
A combiner system may be provided for combining, in a compressed domain, video streams of different media sources in a composite video stream by including a respective video stream as independently decodable spatial segment(s) in the composite video stream. The combiner system may generate composition metadata describing a composition of the spatial segments in the composite video stream and identification metadata comprising identifiers of the respective video streams. A receiver system may obtain decoded video data of a respective video stream based on the composition metadata and a decoding of the composite video stream, and based on the identification metadata, identify a process for handling the decoded video data. Thereby, the composition of spatial segments may dynamically change, while still allowing the receiver device to correctly handle the spatial segments.


