Audio-Based Motion Prediction for HMD Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing head-mountable display (HMD) systems face challenges in providing high-quality, immersive experiences due to limitations in image content generation, such as resolution, texture quality, and latency, which are often addressed by techniques like foveal rendering that require additional equipment like gaze tracking cameras.

Innovation Solution

A system and method for localizing and optionally identifying sounds within the environment of an HMD user, which involves capturing audio, determining sound source locations, and predicting user head motion based on sound characteristics and user preferences, thereby optimizing content generation and display.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If high-resolution image content is provided to preserve user immersion, then image quality is improved, but rendering latency increases

Engineering Contradiction:
Improveimage qualityVSAvoidrendering latency
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system predicts future head motion based on current audio and visual attention, then pre-loads and pre-processes image content for the predicted viewpoint before the user actually moves. This preliminary action allows high-resolution content to be ready in advance, maintaining both image quality and low latency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts the field of view and image resolution based on predicted user attention and motion. Instead of rendering the entire high-resolution scene, it focuses computational resources on the predicted area of interest, dynamically optimizing the balance between image quality and rendering speed.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If foveal rendering techniques are used to improve image quality, then texture quality is enhanced, but device complexity increases due to required gaze tracking equipment

Engineering Contradiction:
Improvetexture qualityVSAvoidequipment requirements
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system uses audio attention as an intermediary to infer visual attention and head motion predictions. Instead of requiring complex gaze tracking cameras, it uses audio processing (which is already present in HMDs for spatial audio) to determine where the user is likely to look, thereby inferring the foveal region without additional hardware.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces the mechanical/optical gaze tracking system with an acoustic-based attention detection system. By analyzing audio spatial attention and sound source localization, it substitutes the need for physical eye-tracking cameras with audio processing algorithms, reducing device complexity while maintaining foveal rendering capabilities.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Speed

If the field of view is reduced to improve rendering speed, then latency is reduced, but image immersion quality deteriorates

Engineering Contradiction:
Improverendering speedVSAvoidimage immersion quality
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The system pre-loads high-resolution image content for the predicted field of view based on audio and visual attention analysis. By preparing content in advance for the area the user is likely to look at, it can maintain a wider effective field of view with high immersion quality while keeping actual rendering latency low through selective pre-processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12321508B2Display system and method
Publication Date: 2025.06.03 SONY INTERACTIVE ENTERTAINMENT LLC
  • US12321508B2 patent drawing
  • US12321508B2 patent drawing
  • US12321508B2 patent drawing

AI summary

A system for obtaining content for display to a user of a head-mountable display device, HMD, the system comprising one or more audio detection units operable to capture audio in the environment of the user, a motion prediction unit operable to predict motion of the HMD in dependence upon the captured audio, and a content obtaining unit operable to obtain content for display in dependence upon the predicted motion of the HMD.