Virtual Assistant Eye-Gaze Engagement Without Wake Words

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current virtual assistants require attention words for activation, which can be unnatural and cumbersome, especially in multi-person conversations, and do not allow for seamless transition between devices.

Innovation Solution

Implementing eye-gaze technology and non-verbal signals, combined with machine-learning algorithms, to initiate and maintain interactions with virtual assistants, eliminating the need for attention words and enabling interaction across multiple devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If attention words are used to activate virtual assistant, then activation can be reliably detected, but interaction becomes unnatural and cumbersome

Engineering Contradiction:
Improvenaturalness of interactionVSAvoidcomplexity of activation mechanism
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent extracts and removes the attention word requirement from the activation mechanism. Instead of requiring users to speak specific wake words, the system uses eye-gaze detection to automatically trigger virtual assistant activation, eliminating the cumbersome verbal command step while maintaining reliable detection through optical sensing.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the acoustic-based attention word detection system with an optical eye-gaze tracking system. This substitution uses cameras and image processing to detect user intent through gaze direction, replacing the mechanical/verbal activation method with a more natural, non-verbal optical detection approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If eye-gaze technology is used to initiate interaction, then natural interaction is enabled, but detection precision and reliability become challenging

Engineering Contradiction:
Improvenaturalness of interactionVSAvoidaccuracy of gaze detection
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the virtual assistant provides visual or verbal confirmation when it detects user gaze intent. This feedback loop allows the system to verify detection accuracy and adjust sensitivity parameters, ensuring reliable differentiation between intentional engagement and casual glances while maintaining natural interaction initiation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary calibration and environmental scanning before requiring precise gaze detection. By pre-mapping the user's typical viewing angles and establishing baseline gaze patterns during initial setup, the system improves subsequent detection precision without requiring complex real-time analysis, thus maintaining both accuracy and naturalness.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple devices are used for interaction, then versatility is improved, but maintaining consistent engagement across devices becomes complex

Engineering Contradiction:
Improvecross-device interaction capabilityVSAvoidcomplexity of engagement management
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal eye-gaze detection framework that operates consistently across multiple device types (smartphones, tablets, computers, smart displays). The core gaze-tracking algorithm and engagement detection logic remain the same across platforms, allowing users to interact with any device using the same natural eye-gaze method without requiring device-specific configurations or complex synchronization protocols.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11221669B2Non-verbal engagement of a virtual assistant
Publication Date: 2022.01.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11221669B2 patent drawing
  • US11221669B2 patent drawing
  • US11221669B2 patent drawing

AI summary

Systems and methods related to engaging with a virtual assistant via ancillary input are provided. Ancillary input may refer to non-verbal, non-tactile input based on eye-gaze data and/or eye-gaze attributes, including but not limited to, facial recognition data, motion or gesture detection, eye-contact data, head-pose or head-position data, and the like. Thus, to initiate and/or maintain interaction with a virtual assistant, a user need not articulate an attention word or words. Rather the user may initiate and/or maintain interaction with a virtual assistant more naturally and may even include the virtual assistant in a human conversation with multiple speakers. The virtual assistant engagement system may utilize at least one machine-learning algorithm to more accurately determine whether a user desires to engage with and/or maintain interaction with a virtual assistant. Various hardware configurations associated with a virtual assistant device may allow for both near-field and/or far-field engagement.