Digital Assistant Gaze Tracking for Video Response Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional electronic digital assistants are unable to intelligently filter video information to provide relevant responses to user inquiries, lacking the capability to computationally process video relative to a user's perception and tailor responses accordingly.
Innovation Solution
An electronic processing device receives a video stream from a video capture device tracking a user's gaze, identifies objects within the stream that remain in view for a threshold period, processes this information using a video processing algorithm, and stores it for later reference to provide tailored responses to user inquiries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional electronic digital assistants source video streams for providing responses to user inquiries, then they can access visual information, but they fail to intelligently filter and process the video information to form relevant responses
Solution Approach 1:
The system performs preliminary actions by detecting user gaze direction and identifying objects of interest before the user makes an inquiry. Video information is pre-processed and stored in association with the user's visual attention, so that when a query is made, the system can quickly retrieve relevant pre-filtered information without having to process the entire video stream in real-time.
Solution Approach 2:
The patent introduces an intermediary video processing system that acts as a mediator between the raw video stream and the digital assistant's response generation. This intermediary component analyzes video content, detects objects, determines user gaze, and filters information based on user attention, thereby enabling the digital assistant to access only relevant video information without direct complex processing.
2Measurement precision
If the electronic digital assistant processes all video information, then comprehensive data is available, but the response accuracy and relevance to user intent deteriorates due to information overload
Solution Approach 1:
The system applies local quality by focusing processing resources on specific regions of interest in the video stream corresponding to the user's gaze direction. Instead of uniformly processing the entire video field, the system identifies and processes only the local area where the user is looking, thereby improving response accuracy while reducing the volume of data that requires processing.
Solution Approach 2:
The patent extracts only the relevant portions of video information that correspond to the user's visual attention and inquiry intent. By taking out and isolating the specific objects and regions the user is examining, the system eliminates unnecessary video data from processing, thus improving response precision without requiring analysis of the complete video stream.
3Adaptability or versatility
If the system tracks user gaze and identifies objects in real-time, then user-specific information is captured, but processing time and computational resources increase
Solution Approach 1:
The system performs gaze tracking and object identification as preliminary actions continuously in the background before explicit user inquiries are made. By pre-detecting and pre-processing video information associated with user gaze, the system captures user-specific context without adding delay to the actual query-response interaction, as the filtering work is already completed when the user asks a question.
Solution Approach 2:
The patent implements continuous gaze tracking and object detection as an ongoing process rather than initiating processing only when a query occurs. This continuous useful action maintains an up-to-date understanding of user attention and context, enabling rapid response generation while the processing occurs continuously in the background, thus minimizing perceived processing delay.
Data Source
AI summary
A process at an electronic computing device that tailors an electronic digital assistant generated inquiry response as a function of previously detected user ingestion of related information includes receiving, from a video capture device configured to track a gaze direction of a first user, a video stream including a first field-of-view of the first user. An object is then identified in the video stream first field-of-view remaining in the first field-of-view for a determined threshold period of time, and the object processed via a video processing algorithm to produce object information, which is then stored. Subsequently, an inquiry is received from the first user for information, and it is determined that the inquiry is related to the object information. The electronic digital assistant then provides a response to the inquiry as a function of the object information.


