Visual Attention Detection for Accurate Digital Assistant Speech Input
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital assistants struggle to accurately determine whether a user's visual attention is directed towards an electronic device while speaking, leading to inefficient and inaccurate responses, which can increase power consumption and require unnecessary user inputs.
Innovation Solution
Concurrently receive audio and video streams to determine user gaze direction, identifying intended speech inputs only when the user is visually attending to the device, thereby initiating tasks through a digital assistant without additional triggers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the device processes all audio inputs without visual attention detection, then the response speed is fast, but the accuracy of identifying intended speech is low
Solution Approach 1:
The patent combines audio stream processing and video stream processing into a unified system that jointly determines user intent. The processor analyzes both audio and video streams concurrently, merging the information to accurately identify when a user is speaking to the device, thereby improving speech identification accuracy without requiring separate independent systems.
Solution Approach 2:
The patent introduces visual attention detection as an intermediary mechanism between the user's speech and the device's response. By detecting whether the user is looking at the device through video stream analysis, the system uses this intermediary information to filter and select relevant speech inputs, preventing incorrect responses while maintaining system manageability.
2Productivity
If the device continuously monitors for user speech, then the responsiveness is high, but the power consumption increases
Solution Approach 1:
The patent implements periodic monitoring of audio and video streams rather than continuous full-processing. The system periodically samples the streams to detect user speech and visual attention, processing data only when needed (when speech and visual attention are detected together), thereby maintaining responsiveness while reducing overall power consumption compared to continuous operation.
Solution Approach 2:
The system uses the user's own visual attention (through gaze detection) as a signal to activate speech processing. When the user looks at the device, it triggers audio processing; when the user looks away, processing is suspended. This self-service mechanism allows the device to respond appropriately without requiring continuous active monitoring, conserving battery life.
3Reliability
If the device responds to all audio inputs, then the productivity is high, but the reliability of correct responses decreases
Solution Approach 1:
The patent employs feedback from video stream analysis to modulate audio processing. The visual attention detection provides feedback that confirms whether the user intends to speak to the device. This feedback loop ensures that only speech accompanied by visual attention (user looking at device) triggers responses, improving reliability by filtering out unintended speech while maintaining productivity through selective processing.
Data Source
AI summary
An example process includes: concurrently receiving an audio stream and a video stream; determining, based on a first portion of the audio stream received within a predetermined duration before a current time and a first portion of the video stream received within the predetermined duration before the current time, whether a visual attention of a user is directed to an electronic device while the user is speaking; and in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is speaking: identifying a second portion of the audio stream to include user speech intended for the electronic device; initiating, by a digital assistant operating on the electronic device, a task based the second portion of the audio stream; and providing an output indicative of the initiated task.


