Dynamic Image Adjustment via Gaze Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital imaging technologies, such as HDR, fail to recreate the dynamic exposure and focus of a scene, leading to an unsatisfying viewing experience, as they often appear 'fake' or 'unreal', and do not effectively adapt to the viewer's gaze or interactions.
Innovation Solution
The system uses gaze tracking and other input methods to swap or adjust image pixels based on viewer interactions, recreating the original scene's brightness, focus, and depth by selecting from a set of images taken with varying exposures and focuses, allowing for dynamic changes in response to viewer actions like gaze direction, touch, or voice input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If HDR technology is used to capture multiple exposures, then all portions of the image are exposed to bring out details, but the viewing experience appears fake or unreal
Solution Approach 1:
The patent applies dynamics by making the image display adaptive and interactive rather than static. The system responds to viewer actions (gaze tracking, touch, voice) by dynamically adjusting which exposure or focus version is displayed, creating a living image that changes based on interaction. This resolves the contradiction by making the HDR image feel more real and engaging through dynamic responsiveness.
Solution Approach 2:
The patent implements feedback loops where the system monitors viewer actions (gaze direction, touch input, voice commands) and uses this feedback to adjust the displayed image in real-time. This feedback mechanism transforms the passive HDR image into an interactive experience, resolving the fake/unreal feeling by creating a responsive dialogue between viewer and image.
2Stability of the object's composition
If static images are displayed, then the image content is fixed and stable, but the viewing experience lacks engagement and adaptability
Solution Approach 1:
The patent transforms static images into dynamic content by implementing multiple versions of the same scene (different exposures, focuses, angles) and using viewer interactions to dynamically select which version to display. This maintains compositional stability through consistent scene content while adding adaptability through interactive selection among multiple variants.
Solution Approach 2:
The patent changes image parameters (exposure level, focus depth, brightness, contrast) based on viewer actions. The system maintains the same scene composition but adapts display parameters in real-time according to gaze tracking, touch input, or voice commands, resolving the contradiction between stability and adaptability.
3Manufacturing precision
If multiple images are captured with varying exposures and focuses, then detailed recreation is possible, but the system complexity increases
Solution Approach 1:
The patent segments the image processing into distinct components: capturing multiple exposures and focuses, storing them as separate image sets, and selectively displaying them based on viewer actions. This segmentation manages complexity by organizing the multi-image system into manageable, independently processed segments rather than attempting to blend them all simultaneously.
Solution Approach 2:
The patent performs preliminary actions by capturing and pre-processing multiple image versions (different exposures, focuses, angles) before the viewing session. These pre-captured variants are stored and ready for rapid selection based on viewer interactions, avoiding the need for complex real-time processing during viewing and thus managing system complexity.
4Device complexity
If traditional imaging systems are used, then the technical implementation is simple, but they fail to recreate the dynamic exposure and focus of a scene
Solution Approach 1:
The patent introduces dynamics into traditional static imaging by implementing multiple captured versions of the scene (different exposures, focuses, angles) and using viewer interactions to dynamically select which version to display. This maintains relative implementation simplicity while achieving dynamic scene recreation that traditional single-image systems cannot provide.
Data Source
AI summary
An embodiment combines the concepts of image enhancement and voice-sound command and control to enhance the experience of viewing images by tracking where the viewer is indicating with his/her voice. The result is to make the viewing experience more like viewing the original scene, or to enhance the viewing experience in new ways beyond the original experience, either automatically, or by interacting with a photographer's previously specified intentions for what should happen when the viewer identifies, with his/her voice sounds, including, but not limited to, words, voice tone, voice inflection, voice pitch, or voice loudness, a particular portion of an image or images taken by that photographer.


