Virtual Assistant Eye-Gaze Engagement Without Wake Words
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual assistants require attention words for activation, which can be unnatural and cumbersome, especially in multi-person conversations, and do not allow for seamless transition between devices.
Innovation Solution
Implementing eye-gaze technology and non-verbal signals, combined with machine-learning algorithms, to initiate and maintain interactions with virtual assistants, eliminating the need for attention words and enabling interaction across multiple devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If attention words are used to activate virtual assistant, then activation can be reliably detected, but interaction becomes unnatural and cumbersome
Solution Approach 1:
The patent extracts and removes the attention word requirement from the activation mechanism. Instead of requiring users to speak specific wake words, the system uses eye-gaze detection to automatically trigger virtual assistant activation, eliminating the cumbersome verbal command step while maintaining reliable detection through optical sensing.
Solution Approach 2:
The patent replaces the acoustic-based attention word detection system with an optical eye-gaze tracking system. This substitution uses cameras and image processing to detect user intent through gaze direction, replacing the mechanical/verbal activation method with a more natural, non-verbal optical detection approach.
2Ease of operation
If eye-gaze technology is used to initiate interaction, then natural interaction is enabled, but detection precision and reliability become challenging
Solution Approach 1:
The patent implements feedback mechanisms where the virtual assistant provides visual or verbal confirmation when it detects user gaze intent. This feedback loop allows the system to verify detection accuracy and adjust sensitivity parameters, ensuring reliable differentiation between intentional engagement and casual glances while maintaining natural interaction initiation.
Solution Approach 2:
The system performs preliminary calibration and environmental scanning before requiring precise gaze detection. By pre-mapping the user's typical viewing angles and establishing baseline gaze patterns during initial setup, the system improves subsequent detection precision without requiring complex real-time analysis, thus maintaining both accuracy and naturalness.
3Adaptability or versatility
If multiple devices are used for interaction, then versatility is improved, but maintaining consistent engagement across devices becomes complex
Solution Approach 1:
The patent implements a universal eye-gaze detection framework that operates consistently across multiple device types (smartphones, tablets, computers, smart displays). The core gaze-tracking algorithm and engagement detection logic remain the same across platforms, allowing users to interact with any device using the same natural eye-gaze method without requiring device-specific configurations or complex synchronization protocols.
Data Source
AI summary
Systems and methods related to engaging with a virtual assistant via ancillary input are provided. Ancillary input may refer to non-verbal, non-tactile input based on eye-gaze data and/or eye-gaze attributes, including but not limited to, facial recognition data, motion or gesture detection, eye-contact data, head-pose or head-position data, and the like. Thus, to initiate and/or maintain interaction with a virtual assistant, a user need not articulate an attention word or words. Rather the user may initiate and/or maintain interaction with a virtual assistant more naturally and may even include the virtual assistant in a human conversation with multiple speakers. The virtual assistant engagement system may utilize at least one machine-learning algorithm to more accurately determine whether a user desires to engage with and/or maintain interaction with a virtual assistant. Various hardware configurations associated with a virtual assistant device may allow for both near-field and/or far-field engagement.


