Gaze-Based Neural Network Video Quality Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing demand for electronic content consumption, such as streaming video and online gaming, leads to large data volumes that often require compression, resulting in degraded video quality, which can negatively impact viewer experiences.

Innovation Solution

The use of neural networks trained with gaze data to infer attention regions in video frames, allowing for dynamic adjustment of video quality, where high resolution is maintained in areas likely to be viewed and lower resolution or compression is applied to less relevant areas, optimizing video quality without significantly affecting the viewer experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If video data is compressed to reduce file size or bandwidth requirements, then bandwidth requirements are reduced, but video quality is degraded

Engineering Contradiction:
Improvebandwidth requirementsVSAvoidvideo quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The patent applies different quality levels to different regions of the video frame based on gaze data. High-resolution rendering is applied to attention regions where viewers are looking, while lower-resolution or highly compressed regions are applied to non-attention areas. This spatial variation in quality allows bandwidth reduction while preserving perceived video quality in critical viewing areas.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If video data is compressed to reduce file size, then file size is reduced, but video quality is degraded

Engineering Contradiction:
Improvefile sizeVSAvoidvideo quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The system compresses video data at different levels across spatial regions. Attention regions identified through gaze analysis maintain higher quality with less compression, while non-attention regions undergo more aggressive compression. This results in smaller overall file sizes while preserving quality where viewers are actually looking.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If high resolution is maintained across the entire video frame, then video quality is improved, but bandwidth requirements increase

Engineering Contradiction:
Improvevideo qualityVSAvoidbandwidth requirements
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

Instead of uniformly applying high resolution across the entire video frame, the system selectively applies high resolution only to attention regions where viewers are looking, based on gaze data analysis. Non-attention regions use lower resolution, thereby reducing overall bandwidth requirements while maintaining perceived quality in critical areas.

Inventive Principle:
Principle #3Local quality

4Loss of energy

If compression is applied to reduce overall video data size, then bandwidth requirements are reduced, but video quality in critical areas may be degraded

Engineering Contradiction:
Improvebandwidth requirementsVSAvoidvideo quality in critical areas
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The system identifies critical areas through gaze data and applies minimal or no compression to these attention regions, while applying higher compression to non-critical areas. This ensures that video quality is preserved in areas where viewers are actually looking, while still achieving overall bandwidth reduction through compression of less important regions.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20210132688A1Gaze determination using one or more neural networks
Publication Date: 2021.05.06 NVIDIA CORP
  • US20210132688A1 patent drawing
  • US20210132688A1 patent drawing
  • US20210132688A1 patent drawing

AI summary

Apparatuses, systems, and techniques are presented to modify media content using inferred attention. In at least one embodiment, a network is trained to predict a gaze of one or more users on one or more image features based, at least in part, on one or more prior gazes of the one or more users, wherein the prediction is to be used to modify at least one of the one or more image features.