Gaze-Adaptive Image Compression for Low-Latency VR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image compression methods are inadequate for low latency applications like virtual or augmented reality, as they require significant processing time and do not achieve sufficient compression, leading to restricted image resolution and quality, especially in mobile solutions.
Innovation Solution
A method and apparatus that selectively encodes frequency coefficients using a bit encoding scheme based on transmission bandwidth, quality of service, and pixel position, discarding higher frequency components and using varying bit numbers to achieve optimal compression and decompression, allowing for efficient image transmission in wearable digital reality headsets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional image compression methods (JPEG, DCT-based) are used, then bandwidth requirements are reduced, but processing time increases significantly and compression ratio is limited
Solution Approach 1:
The patent extracts and encodes only the most significant frequency coefficients (luminance and selected chrominance coefficients) while discarding less important ones. This selective extraction achieves high compression ratios by transmitting only essential visual information, reducing data volume without requiring full processing of all image data.
Solution Approach 2:
The patent applies different encoding precision to different frequency coefficients based on their visual importance. Luminance coefficients are encoded with higher precision than chrominance coefficients, and coefficients in different frequency ranges use different bit allocations. This local differentiation optimizes compression while preserving perceived image quality.
2Manufacturing precision
If high resolution images are generated to avoid motion sickness, then image quality improves, but processing hardware requirements and bandwidth increase
Solution Approach 1:
The patent transforms image data from spatial domain to frequency domain using DCT, then selectively encodes frequency coefficients with varying bit precision. This parameter transformation allows high-resolution imagery to be represented with fewer bits by concentrating visual information in fewer significant coefficients, reducing bandwidth requirements while maintaining perceived image quality.
3Manufacturing precision
If more frequency coefficients are encoded, then image quality is preserved, but compression ratio decreases
Solution Approach 1:
The patent applies different encoding strategies to different regions of the frequency spectrum. Luminance coefficients (which contribute most to perceived image quality) are encoded with higher precision and retained, while chrominance coefficients are encoded with lower precision or discarded. This local quality differentiation maintains essential visual quality while achieving high compression ratios.
Solution Approach 2:
The patent selectively discards less important frequency coefficients (particularly higher-frequency chrominance coefficients) that contribute minimally to perceived image quality. By discarding these coefficients and retaining only the most visually significant ones, the patent achieves substantial compression while preserving essential image quality.
4Quantity of substance
If compression is increased to reduce bandwidth, then data volume decreases, but image resolution and quality are restricted
Solution Approach 1:
The patent uses variable bit-depth encoding for different frequency coefficients, allocating more bits to luminance coefficients and fewer bits to chrominance coefficients. This parameter differentiation allows aggressive compression of less important data while preserving resolution and quality of visually critical information, achieving high compression without restricting essential image quality.
Data Source
Figure 1
Figure 2A~2B
Figure 3
AI summary
A method of compressing image data from one or more images forming part of digital reality content, the method including obtaining pixel data from the image data, the pixel data representing an array of pixels within the one or more images; determining a position of the array of pixels within the one or more images relative to a defined position, the defined position being at least partially indicative of a point of gaze of the user; and compressing the pixel data at least partially in accordance the determined position so that a degree of compression depends on the determined position of the array of pixels.