Gaze-Driven Video Recording with Temporal Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video recording technologies in virtual and mixed reality environments fail to accurately capture the user's focus, leading to missed important content outside a fixed recording area and result in either static or jittery recordings that do not accurately represent the user's experience.
Innovation Solution
Implementing gaze-driven recording systems that utilize gaze-tracking sensors to dynamically identify the region of interest within the user's field of view, applying temporal filters to smooth gaze data, and processing video frames accordingly to create a stable and accurate representation of the user's experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a fixed recording area is used, then the recording system is simple and stable, but important content outside the fixed area is missed and the recording does not represent user experience accurately
Solution Approach 1:
The recording system dynamically adjusts the recording area based on real-time gaze data. The region of interest is continuously updated to follow the user's gaze, transforming the static fixed recording area into a dynamic adaptive area that accurately captures user focus while maintaining system simplicity through automated gaze-driven control
2Measurement precision
If gaze data is used to determine region of interest, then the recording accurately represents user experience, but the gaze data may be jittery and unstable
Solution Approach 1:
The system applies a low-pass filter to the gaze data stream in advance, smoothing out high-frequency jitter and instability before the gaze data is used to determine the region of interest. This pre-processing cushioning ensures stable and reliable region selection while maintaining accuracy in representing user focus
3Productivity
If the entire video frame is processed, then all content is captured, but computing resources are wasted on regions not of interest
Solution Approach 1:
The system extracts only the region of interest based on gaze data and processes only this subset of the video frame. By separating the useful information (gazed-at regions) from the irrelevant information (non-gazed regions), the system achieves high computing efficiency while preserving all important content through gaze-driven selection
Data Source
AI summary
Systems and methods for gaze-driven recording of video are described. Some implementations may include accessing gaze data captured using one or more gaze-tracking sensors; applying a temporal filter to the gaze data to obtain a smoothed gaze estimate; determining a region of interest based on the smoothed gaze estimate, wherein the region of interest identifies a subset of a field of view; accessing a frame of video; recording a portion of the frame associated with the region of interest as an enhanced frame of video, wherein the portion of the frame corresponds to a smaller field of view than the frame; and storing, transmitting, or displaying the enhanced frame of video.


