Volumetric Video Saliency Streams for Lower Bandwidth Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high bandwidth and computing power requirements for seamless high-quality volumetric video streaming and rendering, along with significant time lags due to large data volumes, make it impractical for immersive applications, especially when viewers make head or body movements.

Innovation Solution

Represent volumetric video using a set of saliency video streams and a base stream, with saliency video streams tracking specific spatial regions and dynamically adapting to viewer movements, accompanied by metadata and disocclusion data to enhance rendering efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If volumetric video is streamed with high quality to support seamless viewer experience, then video quality is improved, but bandwidth requirements increase significantly

Engineering Contradiction:
Improvevideo qualityVSAvoidbandwidth
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent segments the volumetric video representation into multiple saliency video streams, each encoding a specific spatial region or object. Instead of transmitting complete high-resolution views from all camera positions, only the salient regions are encoded at high quality in separate streams, while non-salient regions use lower quality or are omitted. This segmentation allows selective transmission of important visual information, reducing overall bandwidth while maintaining perceived quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by encoding different spatial regions with different quality levels based on their importance. Saliency regions (containing important objects or actions) are encoded at high quality in dedicated saliency video streams, while non-salient background regions use lower quality encoding or are excluded from transmission. This differential quality approach maintains high perceived quality for important content while reducing total bandwidth consumption.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If complete volumetric video data is processed and rendered in real time, then rendering quality is improved, but processing time increases causing significant time lags

Engineering Contradiction:
Improverendering qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent extracts and separates the most important visual information into dedicated saliency video streams, isolating critical content from the complete volumetric dataset. By extracting only the essential salient regions that contribute most to viewer experience, the system reduces the volume of data requiring real-time processing and rendering, thereby decreasing processing time and eliminating perceptible time lags while preserving rendering quality for the extracted important content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by processing and rendering only the necessary salient regions rather than complete volumetric data from all camera positions. The system identifies and processes approximately 20% of the total volumetric data (the salient portions) to achieve 80% of the perceived quality, avoiding the excessive processing time required for complete data while maintaining acceptable rendering quality.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If high resolution video data is transmitted from multiple camera positions, then viewer experience is improved, but device complexity increases

Engineering Contradiction:
Improveviewer experienceVSAvoidcomputing power
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces dynamics by making the video streaming system adaptive to viewer behavior and content characteristics. The encoder dynamically identifies salient regions based on object importance, motion activity, and scene composition, adjusting which regions receive high-quality encoding in saliency streams. This dynamic adaptation allows the system to optimize computing resources based on actual viewer needs and content characteristics, reducing device complexity while maintaining versatile viewer experience across different viewing scenarios.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4165877B1Representing volumetric video in saliency video streams
Publication Date: 2025.07.30 DOLBY LABORATORIES LICENSING CORP
  • EP4165877B1 patent drawingFigure 1A~1B
  • EP4165877B1 patent drawingFigure 1C~1D
  • EP4165877B1 patent drawingFigure 2A

AI summary

Saliency regions are identified in a global scene depicted by volumetric video. Saliency video streams that track the saliency regions are generated. Each saliency video stream tracks a respective saliency region. A saliency stream based representation of the volumetric video is generated to include the saliency video streams. The saliency stream based representation of the volumetric video is transmitted to a video streaming client.