Multiresolution Video Coding with Spatially Scalable Motion Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video coding technologies lack flexibility in achieving universal scalability, as they sacrifice coding performance due to inflexibility in scalability and cannot support arbitrary combinations of spatial, temporal, and signal-to-noise ratio (SNR) scalability, which is essential for video streaming over variable-bandwidth networks and diverse devices.
Innovation Solution
The introduction of a family of video decomposition processes using subband Motion Compensated Temporal Filtering (MCTF) that intertwines single-level temporal filtering with spatial filtering, enabling the generation of multiresolution video representations with spatially scalable motion vectors, allowing for flexible support of spatial and temporal scalability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional hybrid motion compensation and DCT coding architecture is used, then coding performance is maintained, but scalability flexibility is insufficient and cannot support arbitrary combinations of spatial, temporal, and SNR scalability
Solution Approach 1:
The video signal is segmented into multiple resolution layers through wavelet transform decomposition, creating a multiresolution representation with coarse and fine detail subbands. This segmentation enables independent encoding and transmission of different spatial and temporal resolution layers, providing flexible scalability while maintaining coding efficiency through the inherent energy compaction property of wavelet transform
Solution Approach 2:
The patent introduces temporal dimension to motion vectors by providing motion vectors at multiple temporal resolutions (full temporal resolution and reduced temporal resolution). This dimensional extension allows the system to support both spatial scalability and temporal scalability simultaneously, enabling arbitrary combinations of scalability types without sacrificing coding performance
2Adaptability or versatility
If layered video coding is used to achieve scalability, then different types of scalability are supported, but scalability flexibility is reduced and coding performance is sacrificed
Solution Approach 1:
The wavelet-based multiresolution representation provides a universal framework that simultaneously supports spatial scalability, temporal scalability, and SNR scalability through a single decomposition structure. The same wavelet coefficients can be used to reconstruct video at different spatial resolutions, temporal resolutions, and quality levels, eliminating the need for separate layered coding structures and thereby maintaining coding efficiency
Solution Approach 2:
The patent changes the fundamental parameters of video representation by using wavelet transform coefficients instead of conventional DCT coefficients. This parameter change enables flexible scalability because wavelet coefficients naturally represent the video signal at multiple resolutions and frequencies, allowing arbitrary combinations of spatial, temporal, and quality scalability without the performance loss associated with conventional layered approaches
3Adaptability or versatility
If motion vectors are not provided at multiple spatial resolutions, then coding complexity is reduced, but spatial scalability cannot be supported
Solution Approach 1:
Motion vectors at reduced spatial resolutions are derived in advance from the wavelet decomposition of motion-compensated prediction residuals. By performing the wavelet transform and deriving reduced-resolution motion vectors during the encoding process, the system prepares all necessary motion information for different spatial resolutions beforehand, enabling spatial scalability without requiring complex real-time processing at the decoder
4Adaptability or versatility
If motion vectors are not provided at multiple temporal resolutions, then coding complexity is reduced, but temporal scalability cannot be supported
Solution Approach 1:
Motion vectors at reduced temporal resolutions are derived in advance by applying wavelet transform to motion-compensated prediction residuals across multiple frames. The encoding process preliminarily computes and stores motion vectors at both full and reduced temporal resolutions, enabling the decoder to reconstruct video sequences at arbitrary temporal rates without complex real-time computation
Data Source
AI summary
Wavelet based multiresolution video representations generated by multi-scale motion compensated temporal filtering (MCTF) and spatial wavelet transform are disclosed. Since temporal filtering and spatial filtering are separated in generating such representations, there are many different ways to intertwine single-level MCTF and single-level spatial filtering, resulting in many different video representation schemes with spatially scalable motion vectors for the support of different combination of spatial scalability and temporal scalability. The problem of design of such a video representation scheme to full the spatial/temporal scalability requirements is studied. Signaling of the scheme to the decoder is also investigated. Since MCTF is performed subband by subband, motion vectors are available for reconstructing video sequences of any possible reduced spatial resolution, restricted by the dyadic decomposition pattern and the maximal spatial decomposition level. It is thus clear that the family of decomposition schemes provides efficient and versatile multiresolution video representations for fully scalable video coding.


