Multiresolution Video Coding with Spatially Scalable Motion Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video coding technologies lack flexibility in achieving universal scalability, as they sacrifice coding performance due to inflexibility in scalability and cannot support arbitrary combinations of spatial, temporal, and signal-to-noise ratio (SNR) scalability, which is essential for video streaming over variable-bandwidth networks and diverse devices.

Innovation Solution

The introduction of a family of video decomposition processes using subband Motion Compensated Temporal Filtering (MCTF) that intertwines single-level temporal filtering with spatial filtering, enabling the generation of multiresolution video representations with spatially scalable motion vectors, allowing for flexible support of spatial and temporal scalability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional hybrid motion compensation and DCT coding architecture is used, then coding performance is maintained, but scalability flexibility is insufficient and cannot support arbitrary combinations of spatial, temporal, and SNR scalability

Engineering Contradiction:
Improvescalability flexibilityVSAvoidcoding performance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The video signal is segmented into multiple resolution layers through wavelet transform decomposition, creating a multiresolution representation with coarse and fine detail subbands. This segmentation enables independent encoding and transmission of different spatial and temporal resolution layers, providing flexible scalability while maintaining coding efficiency through the inherent energy compaction property of wavelet transform

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces temporal dimension to motion vectors by providing motion vectors at multiple temporal resolutions (full temporal resolution and reduced temporal resolution). This dimensional extension allows the system to support both spatial scalability and temporal scalability simultaneously, enabling arbitrary combinations of scalability types without sacrificing coding performance

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If layered video coding is used to achieve scalability, then different types of scalability are supported, but scalability flexibility is reduced and coding performance is sacrificed

Engineering Contradiction:
Improvescalability flexibilityVSAvoidcoding efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The wavelet-based multiresolution representation provides a universal framework that simultaneously supports spatial scalability, temporal scalability, and SNR scalability through a single decomposition structure. The same wavelet coefficients can be used to reconstruct video at different spatial resolutions, temporal resolutions, and quality levels, eliminating the need for separate layered coding structures and thereby maintaining coding efficiency

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the fundamental parameters of video representation by using wavelet transform coefficients instead of conventional DCT coefficients. This parameter change enables flexible scalability because wavelet coefficients naturally represent the video signal at multiple resolutions and frequencies, allowing arbitrary combinations of spatial, temporal, and quality scalability without the performance loss associated with conventional layered approaches

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If motion vectors are not provided at multiple spatial resolutions, then coding complexity is reduced, but spatial scalability cannot be supported

Engineering Contradiction:
Improvespatial scalabilityVSAvoidcoding complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Motion vectors at reduced spatial resolutions are derived in advance from the wavelet decomposition of motion-compensated prediction residuals. By performing the wavelet transform and deriving reduced-resolution motion vectors during the encoding process, the system prepares all necessary motion information for different spatial resolutions beforehand, enabling spatial scalability without requiring complex real-time processing at the decoder

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If motion vectors are not provided at multiple temporal resolutions, then coding complexity is reduced, but temporal scalability cannot be supported

Engineering Contradiction:
Improvetemporal scalabilityVSAvoidcoding complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Motion vectors at reduced temporal resolutions are derived in advance by applying wavelet transform to motion-compensated prediction residuals across multiple frames. The encoding process preliminarily computes and stores motion vectors at both full and reduced temporal resolutions, enabling the decoder to reconstruct video sequences at arbitrary temporal rates without complex real-time computation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8477849B2Wavelet based multiresolution video representation with spatially scalable motion vectors
Publication Date: 2013.07.02 NTT DOCOMO INC
  • US8477849B2 patent drawing
  • US8477849B2 patent drawing
  • US8477849B2 patent drawing

AI summary

Wavelet based multiresolution video representations generated by multi-scale motion compensated temporal filtering (MCTF) and spatial wavelet transform are disclosed. Since temporal filtering and spatial filtering are separated in generating such representations, there are many different ways to intertwine single-level MCTF and single-level spatial filtering, resulting in many different video representation schemes with spatially scalable motion vectors for the support of different combination of spatial scalability and temporal scalability. The problem of design of such a video representation scheme to full the spatial/temporal scalability requirements is studied. Signaling of the scheme to the decoder is also investigated. Since MCTF is performed subband by subband, motion vectors are available for reconstructing video sequences of any possible reduced spatial resolution, restricted by the dyadic decomposition pattern and the maximal spatial decomposition level. It is thus clear that the family of decomposition schemes provides efficient and versatile multiresolution video representations for fully scalable video coding.