Adaptive Prediction Order for Video Bit Depth Scalability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding solutions, such as H.264/AVC, lack support for bit depth scalability, leading to inefficiencies in compression ratio and operational complexity when delivering video to clients with different bit depth requirements, and fail to combine bit depth scalability with other scalability types like spatial scalability effectively.
Innovation Solution
The method involves dynamic adaptive combination of bit depth scalability with other scalability types, specifically spatial scalability, by upsampling base layer information in two logical steps: texture upsampling and bit depth upsampling, with the prediction order being variable and encoder-controlled, allowing for optimized residual usage based on image characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If separate encoding is performed for different bit depths, then compatibility with different clients is achieved, but compression ratio deteriorates and operational complexity increases
Solution Approach 1:
The patent combines multiple scalability types (spatial scalability and bit depth scalability) into a unified scalable video coding framework. Instead of encoding separate bitstreams for different bit depths, it encodes a single scalable bitstream containing base layer (lower bit depth) and enhancement layer (higher bit depth) data, which can be decoded by clients with different bit depth capabilities from the same source
Solution Approach 2:
The scalable video codec is designed to serve multiple functions: it can deliver video to both 8-bit and 12-bit display devices, support different spatial resolutions, and provide adaptive bit rate streaming. The single encoding process generates a universal bitstream that adapts to different client requirements through selective decoding of base and enhancement layers
2Adaptability or versatility
If separate encoding is performed for different bit depths, then compatibility with different clients is achieved, but operational complexity deteriorates
Solution Approach 1:
The patent merges the encoding processes for different bit depths into a single unified scalable encoding process. The encoder produces one scalable bitstream with hierarchical structure (base layer + enhancement layer) rather than multiple independent bitstreams, significantly reducing the complexity of managing multiple encodings and their synchronization
3Ease of operation
If fixed prediction order is used, then decoding simplicity is maintained, but encoding efficiency deteriorates
Solution Approach 1:
The patent introduces dynamic prediction order selection where the encoder can adaptively choose between two prediction modes (spatial prediction first or bit depth prediction first) based on the characteristics of the video content. This dynamic adaptation improves encoding efficiency while the decoder maintains simplicity by following the same adaptive logic using signaling information from the bitstream
4Adaptability or versatility
If bit depth scalability is added to SVC, then support for high color depth is achieved, but compatibility with existing SVC standards deteriorates
Solution Approach 1:
The patent segments the scalable video structure into distinct base layer and enhancement layer components with clear hierarchical relationships. The base layer contains lower bit depth data compatible with existing SVC standards, while the enhancement layer adds high bit depth information. This segmentation allows incremental integration with existing standards while maintaining backward compatibility
Data Source
AI summary
A scalable video bitstream may have an H.264/AVC compatible base layer (BL) and a scalable enhancement layer (EL), where scalability refers to color bit depth. The H.264/AVC scalability extension SVC provides also other types of scalability, e.g. spatial scalability where the number of pixels in BL and EL are different. According to the invention, BL information is upsampled (TUp,BDUp) in two logical steps in adaptive order, one being texture upsampling and the other being bit depth upsampling. Texture upsampling is a process that increases the number of pixels, and bit depth upsampling is a process that increases the number of values that each pixel can have, corresponding to the pixels color intensity. The upsampled BL data are used to predict the collocated EL. A prediction order indication is transferred so that the decoder can upsample BL information in the same manner as the encoder, wherein the upsampling refers to spatial and bit depth characteristics.


