Gaze-Activated XR Digital Assistant Sessions With State Animations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital assistant interactions in extended reality environments are inefficient and require multiple user inputs, leading to increased cognitive burden and power consumption.
Innovation Solution
The system detects user gaze at a persistent object to initiate a digital assistant session, displaying animations indicating session initiation and active listening, and modifies display states based on user gaze to facilitate efficient interaction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional digital assistant interactions are used in XR environments, then multiple user inputs are required to initiate and control the assistant, but this increases cognitive burden and the number of steps needed for interaction
Solution Approach 1:
The system performs preliminary actions by continuously monitoring for wake words and detecting user gaze at persistent objects before the user actually needs to interact with the digital assistant. This allows the assistant to be pre-positioned and ready, eliminating the need for multiple sequential inputs. The persistent object display maintains readiness state without requiring additional user actions.
Solution Approach 2:
The digital assistant performs self-service by automatically detecting when it should be activated through wake word recognition and gaze detection at persistent objects. The system autonomously transitions between states (idle, attentive, active) without requiring explicit user commands for each transition. The animation system automatically communicates the assistant's state changes to the user.
2Speed
If the digital assistant is always active and listening, then it can respond immediately to user needs, but this increases power consumption
Solution Approach 1:
The digital assistant employs periodic action by cycling through distinct operational states (idle, attentive, active) rather than remaining continuously active. The system periodically checks for wake words and monitors gaze at persistent objects, then transitions to full listening mode only when necessary. This periodic activation pattern enables immediate response capability while significantly reducing average power consumption compared to continuous operation.
Solution Approach 2:
The system applies dynamics by making the assistant's operational state changeable and adaptive. The digital assistant dynamically adjusts its level of activity based on contextual cues (wake word detection, gaze detection, persistent object presence). This dynamic state management allows the system to optimize between responsiveness and power consumption by transitioning between low-power idle states and high-responsive active states as needed.
3Use of energy by moving object
If the digital assistant requires explicit activation commands, then power consumption is reduced, but the number of user inputs and cognitive burden increase
Solution Approach 1:
The system introduces an intermediary mechanism using persistent objects as mediators between the user and the digital assistant. These persistent objects serve as visual anchors that convey the assistant's availability and readiness state. The animation system acts as another intermediary, providing non-verbal communication about the assistant's state. This intermediary layer eliminates the need for complex explicit activation protocols while maintaining low power consumption, as the mediation occurs through automatic wake word detection and gaze tracking rather than requiring multiple user commands.
Data Source
AI summary
An example process includes: while displaying a portion of an extended reality (XR) environment representing a current field of view of a user: detecting a user gaze at a first object displayed in the XR environment, where the first object is persistent in the current field of view of the XR environment; in response to detecting the user gaze at the first object, expanding the first object into a list of objects including a second object representing a digital assistant; detecting a user gaze at the second object; in accordance with detecting the user gaze at the second object, displaying a first animation of the second object indicating that a digital assistant session is initiated; receiving a first audio input from the user; and displaying a second animation of the second object indicating that the digital assistant is actively listening to the user.


