Spatially Selective Video Coding for VR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Resource-limited devices, such as battery-operated tablets or smartphones, struggle to efficiently decode and render high data rate panoramic and VR video content due to limited energy, bandwidth, and computational capacity.

Innovation Solution

The implementation of spatially selective encoding methods that dynamically adjust the bitrate and quality of VR or panoramic video content based on the available resources of the client device, by encoding individual image portions with corresponding quality distributions and selectively encoding peripheral portions of the stitched image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If full resolution and full frame high resolution image content is provided, then image quality is improved, but energy consumption and computational load increase beyond device capabilities

Engineering Contradiction:
Improveimage qualityVSAvoidenergy consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality by encoding different regions of the panoramic image with different quality levels. The foveal region (central area corresponding to user's direct gaze) is encoded at high resolution, while peripheral regions are encoded at progressively lower resolutions. This resolves the contradiction by maintaining high image quality where the user actually looks while reducing overall energy consumption and computational load through selective quality reduction in less critical areas.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the panoramic image into multiple regions (foveal, intermediate, and peripheral zones) and applies different encoding strategies to each segment. This segmentation allows the system to manage computational resources efficiently by processing only critical regions at full quality while using compressed representations for other areas, thus resolving the energy-quality contradiction.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If high data rate video content is transmitted, then image quality is improved, but bandwidth requirements exceed available transmission capacity

Engineering Contradiction:
Improveimage qualityVSAvoiddata transmission volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

By transmitting only the foveal region at full resolution and peripheral regions at reduced resolution, the patent dramatically reduces the total data transmission volume while maintaining perceptual image quality. This resolves the contradiction between image quality and bandwidth requirements by optimizing the spatial distribution of data quality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts the encoded regions and quality levels based on the user's gaze direction and head orientation. As the user moves their gaze, the system dynamically repositions the high-resolution foveal region to match the new viewing direction, ensuring optimal use of limited bandwidth while maintaining quality where needed.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If complete panoramic image processing is performed, then image quality is improved, but computational complexity exceeds device processing capacity

Engineering Contradiction:
Improveimage qualityVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts and processes only the essential foveal region at full quality, while using pre-encoded or simplified representations for peripheral regions. This extraction approach reduces computational complexity by focusing processing resources on the most visually important areas while minimizing or eliminating processing for less critical regions.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary encoding of peripheral regions at reduced quality before final composition, and pre-calculates projection parameters for different gaze directions. This preliminary action reduces the computational burden during real-time rendering by having less complex processing already completed in advance.

Inventive Principle:
Principle #10Preliminary action

4Manufacturing precision

If high resolution content is displayed, then image quality is improved, but device resources become inadequate for decoding and rendering

Engineering Contradiction:
Improveimage qualityVSAvoiddevice resource adequacy
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent delivers high-resolution content only to the foveal region while providing lower-resolution content for peripheral areas, making the overall resource requirements adaptable to mobile device capabilities. This allows high image quality perception without overwhelming device decoding and rendering resources.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the resolution parameter spatially across different image regions and dynamically adjusts quality levels based on gaze direction and device capabilities. This parameter adaptation enables the system to optimize between image quality and device resource adequacy by flexibly modifying content resolution where perceptual differences are least noticeable.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250039397A1Systems and methods for spatially selective video coding
Publication Date: 2025.01.30 GOPRO INC
  • US20250039397A1 patent drawing
  • US20250039397A1 patent drawing
  • US20250039397A1 patent drawing

AI summary

A panoramic video frame is partitioned into a plurality of tiles. A viewport corresponding to a field of view within the panoramic video frame is identified. First tiles of the plurality of tiles corresponding to the viewport are encoded at a first bitrate to obtain first encoded tiles. Second tiles of the plurality of tiles outside the viewport are encoded at a second bitrate lower than the first bitrate to obtain second encoded tiles. The first encoded tiles and the second encoded tiles are transmitted to a user device for rendering.