Virtual Assistant Gaze Detection Mode Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual assistant technologies require a visual avatar for interaction, which consumes processing power and bandwidth, and do not allow seamless voice-based interaction when the user is not looking at the screen.
Innovation Solution
A method that switches between a displayed digital avatar and a voice-only virtual assistant based on user gaze, enabling full voice-based interaction without gestures or facial expressions when the user is not looking at the screen, thereby conserving processing power and bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a visual digital avatar is displayed for virtual assistant interaction, then user engagement and interaction quality are improved, but processing power and bandwidth consumption increase
Solution Approach 1:
The system dynamically switches between two operational modes: a visual avatar mode when the user is looking at the screen, and a voice-only mode when the user is not looking. This dynamic adaptation allows the system to optimize resource consumption based on real-time user engagement state, resolving the contradiction between maintaining high user engagement and reducing processing power usage.
Solution Approach 2:
The system changes the operational parameters of the virtual assistant based on detected user gaze state. When the user is not looking at the screen, the system transitions from rendering visual avatar animations to audio-only processing, effectively changing the output modality parameter to reduce computational load while maintaining functional capability.
2Ease of operation
If a visual digital avatar is displayed for virtual assistant interaction, then interaction quality is improved, but bandwidth usage increases
Solution Approach 1:
The system dynamically adjusts its output modality based on user gaze detection. When the user is not looking at the screen, the system switches from visual avatar rendering to audio-only responses, dynamically reducing bandwidth consumption for数据传输 while maintaining interaction functionality through voice output.
3Ease of operation
If the visual avatar is constantly displayed, then user engagement is maintained, but processing power is wasted when user is not looking at the screen
Solution Approach 1:
The system uses its own sensor resources (camera/gaze detection) to monitor user engagement state and automatically adjusts its operational mode accordingly. This self-service mechanism allows the system to identify when visual display is unnecessary and switch to a lower-resource mode, improving processing efficiency without external intervention.
Solution Approach 2:
The system implements a feedback loop where user gaze detection continuously monitors engagement state and feeds this information back to the virtual assistant module, which then adjusts its operational mode. This closed-loop feedback system ensures that processing resources are optimized based on real-time user attention state.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to a computer-implemented method, comprising: determining, using at least one sensor (24), whether or not a user (22) is looking at a computer program's window (26) on a screen (12) of a device (10); in response to determining that the user is looking at the computer program's window, displaying in the computer program's window a digital avatar (28) capable of interacting with the user with voice combined with at least one of gesture and facial expression; and in response to determining that the user is not looking at the computer program's window, enabling a voice-only virtual assistant (30) of the computer program (20) without displaying the digital avatar.