Gaze-Based Voice Intent Detection in AR/VR Wearables

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice assistant systems in AR/VR environments require repetitive keyword utterance or button pressing, limiting natural interaction and usability.

Innovation Solution

An electronic device that utilizes natural user movements such as gaze, gestures, or facial expressions to detect the intention to utter a voice command, enabling voice recognition without explicit keyword input by analyzing gaze, facial expressions, and gesture data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a specific keyword or button press is used to activate the voice assistant, then the voice assistant can be called, but the interaction becomes repetitive and unnatural

Engineering Contradiction:
Improvevoice assistant activationVSAvoidrepeated keyword utterance or button pressing
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary detection of user gaze direction and facial expressions before the voice command is actually uttered. By detecting whether the user is looking at the virtual object and analyzing facial muscle movements in advance, the system prepares to activate the voice assistant proactively, eliminating the need for repeated keyword utterance or button pressing.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If natural movements like gaze and gestures are used to detect user intent, then interaction becomes more natural, but the complexity of detection and analysis increases

Engineering Contradiction:
Improvenatural interactionVSAvoidgaze and gesture analysis
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The external electronic device performs multiple functions using integrated sensors: detecting gaze direction, analyzing facial expressions, and monitoring gesture movements. By combining these detection capabilities into a single multi-functional system, the patent reduces overall system complexity while enabling natural interaction through multiple modalities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If multiple parameters (gaze, face, gesture) are analyzed to determine user intent, then accuracy improves, but the system complexity increases

Engineering Contradiction:
Improveuser intent detection accuracyVSAvoidmulti-parameter analysis system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides the user intent detection process into separate functional modules: gaze detection module, facial expression analysis module, and gesture recognition module. Each module independently analyzes one parameter and provides results to a decision-making unit that combines them. This segmentation allows high measurement precision through multi-parameter analysis while managing device complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240331698A1Electronic device and operation method therefor
Publication Date: 2024.10.03 SAMSUNG ELECTRONICS CO LTD
  • US20240331698A1 patent drawing
  • US20240331698A1 patent drawing
  • US20240331698A1 patent drawing

AI summary

An electronic device is provided. The electronic device includes a communication circuit, memory storing one or more computer programs, and one or more processors communicatively coupled to the communication circuit and the memory, wherein the one or more computer programs include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic to receive, from an external electronic device, information indicating detection of user's gaze on a specified virtual object displayed on a display of the external electronic device wearable on at least part of a user's body through the communication circuit, receive information from analysis of the user's gaze, receive, from the external electronic device, first information from analysis of the user's face, second information from analysis of a user's gesture, or third information from analysis of whether the user started utterance, corresponding to the point in time at which the user's gaze has been detected, determine a user's intention to utter a voice command based on whether at least one of the first information, the second information, or the third information, and the information from analysis of the user's gaze satisfy a specified condition, execute a voice recognition application stored in the memory upon determining that there is the intention to utter and control the voice recognition application to be in a state of being capable of receiving a voice command of the user.