Gaze-Based Voice Intent Detection in AR/VR Wearables
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice assistant systems in AR/VR environments require repetitive keyword utterance or button pressing, limiting natural interaction and usability.
Innovation Solution
An electronic device that utilizes natural user movements such as gaze, gestures, or facial expressions to detect the intention to utter a voice command, enabling voice recognition without explicit keyword input by analyzing gaze, facial expressions, and gesture data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a specific keyword or button press is used to activate the voice assistant, then the voice assistant can be called, but the interaction becomes repetitive and unnatural
Solution Approach 1:
The system performs preliminary detection of user gaze direction and facial expressions before the voice command is actually uttered. By detecting whether the user is looking at the virtual object and analyzing facial muscle movements in advance, the system prepares to activate the voice assistant proactively, eliminating the need for repeated keyword utterance or button pressing.
2Ease of operation
If natural movements like gaze and gestures are used to detect user intent, then interaction becomes more natural, but the complexity of detection and analysis increases
Solution Approach 1:
The external electronic device performs multiple functions using integrated sensors: detecting gaze direction, analyzing facial expressions, and monitoring gesture movements. By combining these detection capabilities into a single multi-functional system, the patent reduces overall system complexity while enabling natural interaction through multiple modalities.
3Measurement precision
If multiple parameters (gaze, face, gesture) are analyzed to determine user intent, then accuracy improves, but the system complexity increases
Solution Approach 1:
The system divides the user intent detection process into separate functional modules: gaze detection module, facial expression analysis module, and gesture recognition module. Each module independently analyzes one parameter and provides results to a decision-making unit that combines them. This segmentation allows high measurement precision through multi-parameter analysis while managing device complexity through modular architecture.
Data Source
AI summary
An electronic device is provided. The electronic device includes a communication circuit, memory storing one or more computer programs, and one or more processors communicatively coupled to the communication circuit and the memory, wherein the one or more computer programs include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic to receive, from an external electronic device, information indicating detection of user's gaze on a specified virtual object displayed on a display of the external electronic device wearable on at least part of a user's body through the communication circuit, receive information from analysis of the user's gaze, receive, from the external electronic device, first information from analysis of the user's face, second information from analysis of a user's gesture, or third information from analysis of whether the user started utterance, corresponding to the point in time at which the user's gaze has been detected, determine a user's intention to utter a voice command based on whether at least one of the first information, the second information, or the third information, and the information from analysis of the user's gaze satisfy a specified condition, execute a voice recognition application stored in the memory upon determining that there is the intention to utter and control the voice recognition application to be in a state of being capable of receiving a voice command of the user.


