Two-Handed Natural User Interface Gesture Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural user interfaces, particularly for augmented reality head-mounted displays, face challenges in accurately determining user intent and spatially perceiving gestures relative to interface elements, as actions intended for control can correspond to non-interface actions, and the apparent location of interface elements within the user's field of view can be difficult to accurately perceive.

Innovation Solution

The implementation of two-handed interactions where one hand performs a context-setting gesture to define the context for dynamic actions performed by the other hand, utilizing image sensors like depth cameras to detect gestures and establish a virtual interaction coordinate system for precise control of user interface elements, with the context-setting hand providing a real-world reference for making dynamic gestures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If single-hand gestures are used for controlling user interface elements, then the operation is simple, but it is difficult to distinguish intended interactions from non-interface actions and spatial perception is inaccurate

Engineering Contradiction:
Improvespatial perception accuracyVSAvoidgesture recognition complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The gesture recognition process is segmented into two distinct phases: context-setting gesture (defining the interaction plane) and dynamic action gesture (performing the actual control). This segmentation allows the system to first establish a reference frame and then interpret subsequent gestures within that frame, improving spatial perception accuracy while maintaining operational simplicity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The context-setting gesture acts as a preliminary action that defines the interaction plane and coordinate system before the dynamic action gesture is performed. By establishing the reference frame in advance, the system can accurately perceive the spatial relationship between gestures and interface elements, resolving the ambiguity between intended and non-interface actions.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If context-setting gesture is required before dynamic action, then spatial perception and intent expression improve, but the interaction sequence becomes more complex

Engineering Contradiction:
Improveuser intent accuracyVSAvoidinteraction time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The interaction system transitions from a static single-gesture model to a dynamic two-phase model where the first gesture (context-setting) establishes the reference frame and the second gesture (dynamic action) performs the control. This dynamic approach improves reliability by clearly distinguishing user intent while the sequential nature minimizes time loss compared to more complex simultaneous multi-gesture systems.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3030953B1Two-hand interaction with natural user interface
Publication Date: 2021.11.10 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3030953B1 patent drawingFigure 1
  • EP3030953B1 patent drawingFigure 2
  • EP3030953B1 patent drawingFigure 3

AI summary

Two-handed interactions with a natural user interface are disclosed. For example, one embodiment provides a method comprising detecting via image data received by the computing device a context-setting input performed by a first hand of a user. and sending to a display a user interface positioned based on a virtual interaction coordinate system, the virtual coordinate system being positioned based upon a position of the first hand of the user. The method further includes detecting via image data received by the computing device an action input performed by a second hand of the user, the action input performed while the first hand of the user is performing the context-setting input, and sending to the display a response based on the context-setting input and an interaction between the action input and the virtual interaction coordinate system.