Spatially Selective Video Coding for VR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Resource-limited devices, such as battery-operated tablets or smartphones, struggle to efficiently decode and render high data rate panoramic and VR video content due to limited energy, bandwidth, and computational capacity.
Innovation Solution
The implementation of spatially selective encoding methods that dynamically adjust the bitrate and quality of VR or panoramic video content based on the available resources of the client device, by encoding individual image portions with corresponding quality distributions and selectively encoding peripheral portions of the stitched image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If full resolution and full frame high resolution image content is provided, then image quality is improved, but energy consumption and computational load increase beyond device capabilities
Solution Approach 1:
The patent applies local quality by encoding different regions of the panoramic image with different quality levels. The foveal region (central area corresponding to user's direct gaze) is encoded at high resolution, while peripheral regions are encoded at progressively lower resolutions. This resolves the contradiction by maintaining high image quality where the user actually looks while reducing overall energy consumption and computational load through selective quality reduction in less critical areas.
Solution Approach 2:
The patent segments the panoramic image into multiple regions (foveal, intermediate, and peripheral zones) and applies different encoding strategies to each segment. This segmentation allows the system to manage computational resources efficiently by processing only critical regions at full quality while using compressed representations for other areas, thus resolving the energy-quality contradiction.
2Manufacturing precision
If high data rate video content is transmitted, then image quality is improved, but bandwidth requirements exceed available transmission capacity
Solution Approach 1:
By transmitting only the foveal region at full resolution and peripheral regions at reduced resolution, the patent dramatically reduces the total data transmission volume while maintaining perceptual image quality. This resolves the contradiction between image quality and bandwidth requirements by optimizing the spatial distribution of data quality.
Solution Approach 2:
The patent dynamically adjusts the encoded regions and quality levels based on the user's gaze direction and head orientation. As the user moves their gaze, the system dynamically repositions the high-resolution foveal region to match the new viewing direction, ensuring optimal use of limited bandwidth while maintaining quality where needed.
3Manufacturing precision
If complete panoramic image processing is performed, then image quality is improved, but computational complexity exceeds device processing capacity
Solution Approach 1:
The patent extracts and processes only the essential foveal region at full quality, while using pre-encoded or simplified representations for peripheral regions. This extraction approach reduces computational complexity by focusing processing resources on the most visually important areas while minimizing or eliminating processing for less critical regions.
Solution Approach 2:
The patent performs preliminary encoding of peripheral regions at reduced quality before final composition, and pre-calculates projection parameters for different gaze directions. This preliminary action reduces the computational burden during real-time rendering by having less complex processing already completed in advance.
4Manufacturing precision
If high resolution content is displayed, then image quality is improved, but device resources become inadequate for decoding and rendering
Solution Approach 1:
The patent delivers high-resolution content only to the foveal region while providing lower-resolution content for peripheral areas, making the overall resource requirements adaptable to mobile device capabilities. This allows high image quality perception without overwhelming device decoding and rendering resources.
Solution Approach 2:
The patent changes the resolution parameter spatially across different image regions and dynamically adjusts quality levels based on gaze direction and device capabilities. This parameter adaptation enables the system to optimize between image quality and device resource adequacy by flexibly modifying content resolution where perceptual differences are least noticeable.
Data Source
AI summary
A panoramic video frame is partitioned into a plurality of tiles. A viewport corresponding to a field of view within the panoramic video frame is identified. First tiles of the plurality of tiles corresponding to the viewport are encoded at a first bitrate to obtain first encoded tiles. Second tiles of the plurality of tiles outside the viewport are encoded at a second bitrate lower than the first bitrate to obtain second encoded tiles. The first encoded tiles and the second encoded tiles are transmitted to a user device for rendering.


