Volumetric Video Saliency Streams for Low-Latency Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high bandwidth and computing power requirements for seamless high-quality volumetric video streaming and rendering, coupled with significant time lags due to large data volumes, hinder viewer experience, especially with viewer movements.
Innovation Solution
Representing volumetric video using a set of saliency video streams and a base stream, with adaptive saliency ranks, to independently stream image data for saliency regions, and utilize disocclusion data to enhance rendering, reducing data volume and resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If volumetric video is streamed with high quality to support all viewpoints, then viewing quality is improved, but bandwidth requirements increase enormously
Solution Approach 1:
The patent segments the volumetric video representation into multiple saliency video streams, each encoding a specific saliency region or viewpoint. Instead of transmitting all possible viewpoints at full resolution, only selected saliency regions are encoded at high quality in separate streams, while other regions use a base stream. This segmentation allows selective transmission of high-quality data only where needed, dramatically reducing total bandwidth requirements.
Solution Approach 2:
The patent applies local quality by assigning different encoding qualities to different spatial regions. Saliency regions that are likely to be viewed receive high-quality encoding in dedicated saliency video streams, while non-saliency regions use lower-quality base stream encoding. This localized quality differentiation maintains high viewing quality for important regions while reducing overall data transmission requirements.
2Productivity
If volumetric video data is compressed to reduce bandwidth, then data transmission efficiency is improved, but computing power requirements for compression and decompression increase
Solution Approach 1:
The patent performs preliminary action by pre-identifying and encoding saliency regions and viewpoints before transmission. The encoder analyzes the volumetric video content in advance to determine which regions are most important, then creates optimized saliency video streams for these regions. This pre-processing allows the decoder to simply decode and render the pre-prepared streams without performing complex real-time analysis, shifting computational burden to the encoding phase and reducing client-side computing requirements.
3Measurement precision
If high quality image content is streamed in real time, then viewing quality is improved, but time lags increase significantly
Solution Approach 1:
The patent extracts and transmits only the essential high-quality saliency region data in separate saliency video streams, rather than transmitting all volumetric video data. The base stream provides a complete but lower-quality representation, while extracted saliency streams provide enhanced quality for specific regions. This extraction approach reduces the total data volume that must be processed in real-time, minimizing time lags while maintaining high quality for viewed regions.
4Measurement precision
If all viewpoints are encoded at full resolution, then viewing quality for any viewpoint is improved, but data volume becomes impractical to support
Solution Approach 1:
The patent applies dynamics by making the video stream composition adaptive and flexible. The system dynamically selects which saliency video streams to transmit based on current viewing conditions, viewer position, and network bandwidth availability. This dynamic approach allows the system to maintain high viewing quality for the currently viewed region while adapting the total data volume to practical transmission limits, rather than statically encoding all viewpoints at full resolution.
Data Source
AI summary
Saliency regions are identified in a global scene depicted by volumetric video. Saliency video streams that track the saliency regions are generated. Each saliency video stream tracks a respective saliency region. A saliency stream based representation of the volumetric video is generated to include the saliency video streams. The saliency stream based representation of the volumetric video is transmitted to a video streaming client.


