Multi-view Video Coding Scalable Layer Reference Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-view video coding solutions lack scalability in terms of bitrate, picture rate, and picture quality, offering only view scalability and temporal scalability, which is insufficient for adjusting the requirements of different displays and computational capacities.

Innovation Solution

The method involves coding each view of a multi-view video stream with a scalable video coding scheme, using a reference picture list constructed from marked reference pictures and inter-view only reference pictures, allowing for the selective use of base or enhancement representations for inter-view prediction, and implementing leaky prediction to control potential drift in quality-scalable views.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If multi-view video coding is implemented with high spatial resolution and picture quality, then video fidelity is improved, but bitrate and computational complexity increase significantly

Engineering Contradiction:
Improvepicture qualityVSAvoidbitrate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The video signal is divided into multiple scalable layers (base layer and enhancement layers), where each layer represents a different quality level. The base layer provides essential video content at lower quality, while enhancement layers add incremental improvements. This segmentation allows the system to transmit only the necessary layers based on available bitrate, thereby maintaining picture quality when bandwidth permits while reducing bitrate when bandwidth is constrained.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The scalable video coding structure enables dynamic adaptation of the transmitted bitstream by selectively including or excluding enhancement layers based on network conditions and receiver capabilities. The system can dynamically adjust the effective picture quality and bitrate by pruning the scalable layer representation, allowing flexible response to changing transmission conditions without reencoding the entire video stream.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If multi-view video coding is implemented with high spatial resolution, then video fidelity is improved, but computational complexity increases significantly

Engineering Contradiction:
Improvespatial resolutionVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The video signal is divided into multiple scalable layers (base layer and enhancement layers), where each layer represents a different quality level. The base layer provides essential video content at lower quality, while enhancement layers add incremental improvements. This segmentation allows the system to transmit only the necessary layers based on available bitrate, thereby maintaining picture quality when bandwidth permits while reducing bitrate when bandwidth is constrained.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The scalable video coding structure enables dynamic adaptation of the transmitted bitstream by selectively including or excluding enhancement layers based on network conditions and receiver capabilities. The system can dynamically adjust the effective picture quality and bitrate by pruning the scalable layer representation, allowing flexible response to changing transmission conditions without reencoding the entire video stream.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If scalable video coding with multiple layers is implemented, then flexibility in adjusting bitrate and quality is improved, but reference picture management and inter-view prediction complexity increase

Engineering Contradiction:
Improvebitrate adjustment flexibilityVSAvoidreference picture management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies different quality levels to different parts of the reference picture list by marking reference pictures with layer identifiers. Each reference picture can be associated with a specific scalable layer, allowing the decoder to selectively use base layer or enhancement layer references based on the current picture's requirements. This local differentiation enables precise control over prediction quality while managing reference picture memory efficiently.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces a reference picture list structure that acts as an intermediary between the scalable layer representations and the inter-view prediction process. This reference picture list organizes and manages references from multiple views and layers, providing a structured interface that simplifies the complex task of selecting appropriate references while maintaining flexibility in bitrate and quality adjustment.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If inter-view prediction uses enhancement layer references, then prediction accuracy is improved, but potential quality drift increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidquality consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent modifies the reference picture selection process by introducing layer-specific marking and selection parameters. Instead of uniformly using enhancement layer references, the system dynamically selects references based on the current picture's layer and the desired quality consistency. This parameter-based control allows the system to switch between using enhancement layer references (for higher accuracy) and base layer references (for consistency) depending on the specific coding context.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements a feedback mechanism where the decoder monitors the quality and layer characteristics of reference pictures and adjusts the selection strategy accordingly. When enhancement layer references are used, the system tracks potential drift and can switch to base layer references to maintain quality consistency, creating a closed-loop control system that balances prediction accuracy with quality reliability.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8855199B2Method and device for video coding and decoding
Publication Date: 2014.10.07 NOKIA TECHNOLOGIES OY
  • US8855199B2 patent drawing
  • US8855199B2 patent drawing
  • US8855199B2 patent drawing

AI summary

Embodiments of the present invention relate to video coding for multi-view video content. It provides a coding system enabling scalability for the multi-view video content. In one embodiment, a method is provided for encoding at least two views representative of a video scene, each of the at least two views being encoded in at least two scalable layers, wherein one of the at least two scalable layers representative of one view of the at least two views is encoded with respect to a scalable layer representative of the other view of the at least two views.