Temporal Scalability in Hybrid Video Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video compression technologies, such as MPEG-4 Simple Profile and H.264 Baseline Profile, lack effective temporal scalability, which is essential for adapting to varying network conditions and user device capabilities, especially in heterogeneous IP networks and mobile/wireless channels.
Innovation Solution
The introduction of unidirectional predicted temporal scaling frames, which can be omitted without affecting other frames, allowing for adaptive data rate and quality adjustments by encoding and decoding systems, conforming to the MPEG-4 Simple Profile and H.264 Baseline Profile without additional syntax changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If B frames are used for temporal scalability, then temporal scalability is achieved, but device complexity increases and compatibility with simple profiles is lost
Solution Approach 1:
The patent segments the video bitstream into base layer and enhancement layer frames. Base layer frames contain essential video information for basic playback, while enhancement layer frames provide additional temporal scalability. This segmentation allows decoders to process only the base layer for simple profiles, maintaining low complexity while enabling temporal scalability when enhancement layers are available.
Solution Approach 2:
The patent introduces an intermediate frame structure that acts as a mediator between reference frames and predicted frames. These intermediate frames provide prediction references that enable temporal scalability without requiring full B-frame functionality, thus reducing decoder complexity while maintaining adaptability.
2Adaptability or versatility
If multiple independent bitstream versions are generated for different bandwidths, then adaptability to varying network conditions is achieved, but server storage efficiency decreases
Solution Approach 1:
The patent merges multiple bitstream versions into a single scalable bitstream by combining base layer and enhancement layer frames in an interleaved structure. This single bitstream contains all necessary information for different network conditions, eliminating the need to store multiple separate bitstream versions while maintaining full network adaptability.
Solution Approach 2:
The patent creates a universal scalable bitstream that serves multiple functions: it provides base quality for low-bandwidth networks, enhanced quality for medium-bandwidth networks, and full quality for high-bandwidth networks. This single multi-functional bitstream replaces multiple specialized bitstreams, reducing storage requirements while maintaining adaptability across all network conditions.
3Adaptability or versatility
If temporal scalability is implemented in simple profiles, then adaptability to device capabilities is improved, but computational complexity increases
Solution Approach 1:
The patent implements dynamic temporal scalability where the decoder can adaptively adjust the level of enhancement layer processing based on available computational power and device capabilities. When computational resources are abundant, the decoder processes enhancement layers for improved quality; when resources are limited, it processes only base layer frames, maintaining adaptability while controlling power consumption.
Solution Approach 2:
The patent enables device adaptability through parameter changes in the decoding process, such as adjusting motion compensation precision, reference frame buffering, and prediction accuracy based on device capabilities. These parameter adjustments allow temporal scalability to be implemented in simple profiles while controlling computational complexity and power consumption according to specific device requirements.
Data Source
AI summary
The invention is directed to a method and apparatus for providing temporal scaling frames for use in digital multimedia. The method involves using a removable unidirectional predicted temporal scaling frame communication along with intra-coded frames and/or inter-coded frames. The method involves the ability to selectively remove the temporal scaling frame(s) from being transmitted or decoded in order to satisfy, for example, power limits, data rate limits, computational limits or channel conditions. Examples presented include encoders, transcoders and decoders where the decision to drop the removable temporal scaling frames could be made.


