Multiview Video Signaling via View Identifier Ranges

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding standards face challenges in efficiently signaling and streaming multiview video data, particularly in adapting to varying display capabilities and network conditions, leading to suboptimal decoding and rendering of multimedia content.

Innovation Solution

The technique involves assigning view identifiers to multiview video data based on camera perspectives, allowing for dynamic selection of representations with varying depth and view numbers, and using HTTP streaming with DASH to adapt to client capabilities and network conditions, ensuring efficient data transmission and utilization of display resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiview video data is transmitted with all available views and depth information, then complete viewing experience is provided, but network bandwidth consumption increases and client device resources are overwhelmed

Engineering Contradiction:
Improveadaptability to display capabilitiesVSAvoiddata transmission volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by signaling view identifier ranges (minimum and maximum view IDs) that specify subsets of views to be transmitted or decoded. This allows different portions of the multiview video data to be selectively included or excluded based on client capabilities and network conditions, rather than transmitting all views uniformly. The view identifier range acts as a local filter that tailors the data quantity to specific client needs.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameter of view identifier ranges to adapt to different display capabilities. By signaling minimum and maximum view IDs in the bitstream, the system dynamically adjusts which views are included in the transmitted data. Clients can request different view identifier ranges based on their display hardware capabilities, thereby changing the quantity of transmitted data while maintaining adaptability to various display configurations.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If view identifier ranges are signaled for all possible client configurations, then complete adaptability is achieved, but signaling overhead and manifest complexity increase

Engineering Contradiction:
Improveadaptability to network conditionsVSAvoidmanifest structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements universality by creating a manifest structure that serves multiple functions: it signals view identifier ranges for different representations, provides depth information, and enables clients to select appropriate configurations based on their capabilities. This single manifest structure handles multiple adaptation scenarios (different display capabilities, network conditions, client types) without requiring separate manifests for each configuration, thereby reducing overall complexity while maintaining broad adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If depth information is included for all views in multiview video representations, then complete 3D viewing experience is provided, but processing requirements and memory usage increase

Engineering Contradiction:
Improveview rendering precisionVSAvoidclient device processing energy
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies the extraction principle by selectively including or excluding depth information for specific views based on client capabilities and representation requirements. Rather than mandating that all views include depth information, the system extracts and transmits depth data only for views that will be rendered in 3D on the client device. This reduces processing energy and memory requirements while maintaining the ability to provide complete 3D viewing experiences when needed.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP2601790B1Signaling attributes for network-streamed video data
Publication Date: 2021.12.22 QUALCOMM INC
  • EP2601790B1 patent drawingFigure 1
  • EP2601790B1 patent drawingFigure 2
  • EP2601790B1 patent drawingFigure 3

AI summary

In one example, an apparatus includes a processor configured to receive video data for two or more views of a scene, determine horizontal locations of camera perspectives for each of the two or more views, assign view identifiers to the two or more views such that the view identifiers correspond to the relative horizontal locations of the camera perspectives, form a representation comprising a subset of the two or more views, and, in response to a request from a client device, send information indicative of a maximum view identifier and a minimum view identifier for the representation to the client device.