Spatial Scalability via Redundant Pictures and Slice Groups

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conferencing systems face challenges in achieving interoperability between H.264 SVC Scalable Baseline and Baseline/High profile devices due to non-interoperability issues, and transcoding solutions introduce high computational requirements and latency, while also reducing video quality unnecessarily.

Innovation Solution

The method involves encoding video sources using redundant pictures and slice groups to create H.264 Baseline profile conformant bitstreams, allowing for the production of spatially scalable video streams that can be routed to endpoints with varying capabilities without transcoding, ensuring compatibility with millions of deployed Baseline profile devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If transcoding is used to transmit bitstreams of varying quality, then interoperability between different profile devices is improved, but computational requirements and latency increase significantly

Engineering Contradiction:
ImproveinteroperabilityVSAvoidlatency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The video stream is segmented into multiple spatial layers with different resolutions. The encoder produces a base layer and enhancement layers, allowing receivers to select appropriate layers based on their capabilities without requiring full transcoding. This segmentation enables interoperability while avoiding the computational overhead of transcoding operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The video is pre-encoded into a scalable format with multiple spatial layers before transmission. This preliminary action allows receivers to directly use the encoded layers without performing transcoding, thereby reducing latency while maintaining adaptability to different device capabilities.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If transcoding is used to adapt video to endpoint capabilities, then compatibility with Baseline profile devices is improved, but video quality is unnecessarily reduced

Engineering Contradiction:
ImprovecompatibilityVSAvoidvideo quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system provides dynamic adaptation by allowing receivers to selectively decode only the necessary spatial layers based on their real-time capabilities. High-capability devices can access full-resolution enhancement layers while maintaining Baseline profile compatibility, whereas lower-capability devices can use only the base layer. This dynamic approach preserves video quality for capable devices while ensuring compatibility for all.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If H.264 SVC Scalable Baseline profile is used for interoperability, then compatibility with Baseline devices is improved, but device complexity and cost increase due to transcoding MCUs

Engineering Contradiction:
ImproveinteroperabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Instead of using complex transcoding MCUs that copy and re-encode video streams, the system uses a simpler approach where the encoder directly produces multiple spatial layers in H.264 Baseline conformant format. Receivers can then select and decode the appropriate layer without requiring complex transcoding infrastructure, thereby reducing device complexity and cost while maintaining interoperability.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS8934530B2Spatial scalability using redundant pictures and slice groups
Publication Date: 2015.01.13 VIDYO INC
  • US8934530B2 patent drawing
  • US8934530B2 patent drawing
  • US8934530B2 patent drawing

AI summary

Systems and methods for using redundant pictures and slice groups to encode spatially scalable H.264 Baseline profile conformant video and to route that video to endpoints of varying capabilities without using the Scalable Video extension of H.264 or transcoding. Reduced resolution versions of primary coded pictures are encoded as slice groups in a full-resolution composite pictures, which are added to the video bitstream as redundant pictures. A router then processes the spatially scaled video bitstream into separate streams having different resolutions and routes these to endpoints of varying capabilities.