Multi-factor AR Intention Detection for Control Object Positioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In augmented reality (AR) environments, existing object placement controllers often fail to accurately understand user intentions regarding control objects, leading to premature removal or improper positioning, which can distract users and degrade their experience.

Innovation Solution

Implementing a multi-factor intention determination method that uses a combination of gestures, gaze direction, and hand positions to summon and position control objects like menus and keyboards, ensuring they remain visible only when intended by the user, with a timer-based persistence mechanism to manage object visibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a control object is displayed in AR environment based on simple gesture detection, then the ease of operation is improved, but the reliability of user intention determination deteriorates

Engineering Contradiction:
Improveease of operationVSAvoidreliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent combines multiple gesture indicators (hand gesture, eye gaze, head orientation) into a unified control mechanism. The control object is activated only when a specific hand gesture is detected in combination with the user's gaze and head orientation directed toward the same region, merging multiple sensing modalities to improve intention determination reliability while maintaining ease of operation.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If control object is kept visible for extended period, then the user experience is improved, but the loss of time for other tasks increases

Engineering Contradiction:
Improveuser experienceVSAvoidloss of time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a timer-based mechanism that periodically evaluates whether to maintain or remove the control object. The control object is displayed for a predetermined time period after activation, and the system periodically checks for continued user intent through gesture and gaze detection. If no further activation gestures are detected within the time period, the control object is automatically removed, balancing user experience with time efficiency.

Inventive Principle:
Principle #19Periodic action

3Ease of operation

If control object is positioned close to user, then the ease of operation is improved, but the object-generated harmful factors increase due to distraction

Engineering Contradiction:
Improveease of operationVSAvoiddistraction
Core Design Contradiction:
Ease of operationVSObject-generated harmful factors

Solution Approach 1:

The patent positions the control object in a localized region of the AR environment based on the user's gaze and head orientation. Rather than displaying the control object in a fixed global position, the system dynamically places it in the local visual field where the user is already attending, determined by eye and head tracking data. This localized positioning provides ease of operation while minimizing distraction by placing the control object only when and where the user intends to interact.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12067159B2Multi-factor intention determination for augmented reality (AR) environment control
Publication Date: 2024.08.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12067159B2 patent drawing
  • US12067159B2 patent drawing
  • US12067159B2 patent drawing

AI summary

Examples of augmented reality (AR) environment control advantageously employ multi-factor intention determination and include: performing a multi-factor intention determination for summoning a control object (e.g., a menu, a keyboard, or an input panel) using a set of indications in an AR environment, the set of indications comprising a plurality of indications (e.g., two or more of a palm-facing gesture, an eye gaze, a head gaze, and a finger position simultaneously); and based on at least the set of indications indicating a summoning request by a user, displaying the control object in a position proximate to the user in the AR environment (e.g., docked to a hand of the user). Some examples continue displaying the control object while at least one indication remains, and continue displaying the control object during a timer period if one of the indications is lost.