Gaze-Adaptive Audio-Visual Output for Sound Image Positioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies for controlling viewing processing of free viewpoint video and stereoscopic images do not adequately utilize user gaze information to enhance content reproduction based on gazing points.
Innovation Solution
An information processing device and method that estimate sound image coordinates and control video and audio output based on user gaze, performing rendering and sound image localization to enhance content reproduction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing technologies for controlling viewing processing of free viewpoint video and stereoscopic images are used, then display control can be performed based on detected eye position and line of sight, but user gaze information is not adequately utilized to enhance content reproduction
Solution Approach 1:
The system implements feedback by detecting user gaze information (eye position and line of sight) and using it to dynamically adjust video output control. The discrimination unit detects gaze points, and this information feeds back to the video output control unit to perform rendering operations including framing and zooming, creating a closed-loop system that continuously adapts to user attention.
Solution Approach 2:
The system enables self-service by automatically performing video rendering operations based on detected gaze information without requiring manual user input. The discrimination unit and video output control unit work together to automatically frame and zoom video content according to where the user is looking, making the system self-adjusting based on user behavior.
2Adaptability or versatility
If video output control performs rendering including framing or zooming processing based on gazing degree, then content reproduction is enhanced, but system complexity increases
Solution Approach 1:
The video output control unit is designed with multi-functionality, performing multiple rendering operations (framing, zooming, and other processing) based on a single gaze detection input. This universal unit handles various video processing tasks unified by the gaze-based control mechanism, reducing the need for separate specialized components for each function.
3Measurement precision
If sound image coordinates are estimated and audio output is controlled to generate sound images at specific coordinates, then realistic viewing experience is provided, but processing requirements increase
Solution Approach 1:
The system applies local quality by estimating sound image coordinates specifically at the gazed location rather than uniformly across the entire field of view. The estimation unit focuses computational resources on calculating sound coordinates only for objects or regions where the user is looking, providing high localization precision where needed while conserving processing power in non-gazed areas.
Data Source
AI summary
Provided is an information processing device that performs processing on a content. An information processing device is provided with an estimation unit that estimates sounding coordinates at which a sound image is generated on the basis of a video stream and an audio stream, a video output control unit that controls an output of the video stream, and an audio output control unit that controls an output of the audio stream so as to generate the sound image at the sounding coordinates. A discrimination unit that discriminates a gazing point of a user who views video and audio is further provided, in which the estimation unit estimates the sounding coordinates at which the sound image of the object gazed by the user is generated on the basis of a discrimination result.


