Spatial Scalability via Redundant Pictures and Slice Groups
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conferencing systems face challenges in achieving interoperability between H.264 SVC Scalable Baseline and Baseline/High profile devices due to non-interoperability issues, and transcoding solutions introduce high computational requirements and latency, while also reducing video quality unnecessarily.
Innovation Solution
The method involves encoding video sources using redundant pictures and slice groups to create H.264 Baseline profile conformant bitstreams, allowing for the production of spatially scalable video streams that can be routed to endpoints with varying capabilities without transcoding, ensuring compatibility with millions of deployed Baseline profile devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If transcoding is used to transmit bitstreams of varying quality, then interoperability between different profile devices is improved, but computational requirements and latency increase significantly
Solution Approach 1:
The video stream is segmented into multiple spatial layers with different resolutions. The encoder produces a base layer and enhancement layers, allowing receivers to select appropriate layers based on their capabilities without requiring full transcoding. This segmentation enables interoperability while avoiding the computational overhead of transcoding operations.
Solution Approach 2:
The video is pre-encoded into a scalable format with multiple spatial layers before transmission. This preliminary action allows receivers to directly use the encoded layers without performing transcoding, thereby reducing latency while maintaining adaptability to different device capabilities.
2Adaptability or versatility
If transcoding is used to adapt video to endpoint capabilities, then compatibility with Baseline profile devices is improved, but video quality is unnecessarily reduced
Solution Approach 1:
The system provides dynamic adaptation by allowing receivers to selectively decode only the necessary spatial layers based on their real-time capabilities. High-capability devices can access full-resolution enhancement layers while maintaining Baseline profile compatibility, whereas lower-capability devices can use only the base layer. This dynamic approach preserves video quality for capable devices while ensuring compatibility for all.
3Adaptability or versatility
If H.264 SVC Scalable Baseline profile is used for interoperability, then compatibility with Baseline devices is improved, but device complexity and cost increase due to transcoding MCUs
Solution Approach 1:
Instead of using complex transcoding MCUs that copy and re-encode video streams, the system uses a simpler approach where the encoder directly produces multiple spatial layers in H.264 Baseline conformant format. Receivers can then select and decode the appropriate layer without requiring complex transcoding infrastructure, thereby reducing device complexity and cost while maintaining interoperability.
Data Source
AI summary
Systems and methods for using redundant pictures and slice groups to encode spatially scalable H.264 Baseline profile conformant video and to route that video to endpoints of varying capabilities without using the Scalable Video extension of H.264 or transcoding. Reduced resolution versions of primary coded pictures are encoded as slice groups in a full-resolution composite pictures, which are added to the video bitstream as redundant pictures. A router then processes the spatially scaled video bitstream into separate streams having different resolutions and routes these to endpoints of varying capabilities.


