Scalable Video Encoding with Full-Resolution Residuals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Scalable video coding (SVC) often compromises the rate-distortion performance of the enhancement layer compared to regular H.264/AVC encoding, and existing methods that aim to improve SVC performance typically trade off base layer performance for enhanced enhancement layer performance.
Innovation Solution
The method involves encoding an input video into a scalable format with a spatially downsampled base layer and an enhancement layer at a higher resolution, using full-resolution residual values to optimize motion estimation and mode decision processes in the base layer encoding, and incorporating these values into rate-distortion optimization expressions to produce a bitstream that combines both layers effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If scalable video coding is used to provide multiple layers, then the capability to encode at different resolutions is improved, but the rate-distortion performance of the enhancement layer deteriorates
Solution Approach 1:
The patent performs preliminary full-resolution encoding to generate residual data before creating the scalable bitstream. This preliminary action allows the enhancement layer to inherit high-quality residual information from the base layer encoding, improving rate-distortion performance without sacrificing the base layer's integrity.
Solution Approach 2:
The patent introduces residual data as an intermediary element that bridges the base layer and enhancement layer. By inserting residual information from full-resolution encoding into the scalable bitstream structure, it enables the enhancement layer to achieve better performance without compromising the base layer encoding quality.
2Productivity
If motion estimation is performed using downsampled residuals, then base layer encoding efficiency is improved, but the accuracy of motion vectors deteriorates
Solution Approach 1:
The patent applies different quality levels to different parts of the encoding process. Full-resolution residual data is used for motion estimation in regions where accuracy is critical, while downsampled data is used for general encoding operations. This local quality differentiation maintains motion vector accuracy where needed while preserving overall encoding efficiency.
Data Source
AI summary
An encoder and method for encoding data in a scalable data compression format are described. In particular, process for encoding spatially scalable video are described in which the base layer uses downscaled residuals from a full-resolution encoding of the video in its motion estimation process. The downscaled residuals may also be used in the coding mode selection process at the base layer.


