AR Wearable Sensor Fusion for Precise 3D Object Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional input techniques in AR/VR/MR environments require high specificity and suffer from high error rates and fatigue due to imprecise commands, making it challenging to accurately interact with virtual or real objects in 3D space.

Innovation Solution

A wearable device dynamically fuses multiple sensor inputs, such as head pose, eye gaze, hand gestures, and voice commands, to anticipate and predict user intent, dynamically adding or removing inputs based on convergence and divergence, reducing reliance on high-resolution sensors and enhancing interaction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional input techniques are used in AR/VR/MR environments, then the system requires high specificity and precision in user inputs, but this leads to high error rates and user fatigue

Engineering Contradiction:
Improveinput precisionVSAvoiderror rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines multiple sensor inputs (head pose, eye gaze, hand gestures, voice commands) into a fused input signal. By merging these different modalities, the system achieves more reliable object selection and command execution, reducing errors that occur with single-input methods while maintaining natural interaction.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system continuously monitors the convergence of multiple sensor inputs and dynamically adjusts which sensors are active based on their contribution to the current task. This feedback mechanism allows the system to reduce noise and errors by deactivating sensors that are not providing useful information, thereby improving overall input reliability.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If multiple sensor inputs are continuously used, then the system can capture comprehensive user intent, but this increases device complexity and computational load

Engineering Contradiction:
Improveinput versatilityVSAvoidsensor fusion complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic sensor selection where the system actively adjusts which sensors are enabled based on the current interaction context and convergence of inputs. This dynamic approach maintains input versatility by having multiple sensors available when needed, while reducing device complexity by deactivating unnecessary sensors during different phases of interaction.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system segments the sensor input processing into distinct phases: initial exploration phase where multiple sensors are active, and convergence phase where only relevant sensors continue. This segmentation allows the system to handle complex multi-sensor data when necessary while simplifying processing during stable interaction states.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If high-resolution sensors are used to ensure accurate input detection, then measurement precision improves, but hardware cost and complexity increase

Engineering Contradiction:
Improvesensor detection accuracyVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines data from multiple lower-resolution or simpler sensors to achieve the equivalent precision of a single high-resolution sensor. By fusing head pose, eye gaze, hand gestures, and voice commands, the system achieves accurate object selection without requiring expensive high-resolution sensors for each individual input modality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses an intermediary processing layer that fuses inputs from multiple sensors before final decision-making. This intermediary fusion process allows simpler sensors to contribute to accurate overall detection, reducing the need for any single sensor to be extremely high-resolution while maintaining overall system accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If the system dynamically selects and fuses sensor inputs, then interaction accuracy improves, but processing time and computational resources increase

Engineering Contradiction:
Improveinteraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary convergence assessment of sensor inputs to determine early whether multiple sensors should be fused. By evaluating the convergence of inputs in advance, the system can quickly decide to use only the most relevant sensors, avoiding the computational overhead of processing all sensors in every situation while maintaining high interaction accuracy when convergence is achieved.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12444146B2Identifying convergence of sensor data from first and second sensors within an augmented reality wearable device
Publication Date: 2025.10.14 MAGIC LEAP INC
  • US12444146B2 patent drawing
  • US12444146B2 patent drawing
  • US12444146B2 patent drawing

AI summary

Examples of wearable systems and methods can use multiple inputs (e.g., gesture, head pose, eye gaze, voice, totem, and/or environmental factors (e.g., location)) to determine a command that should be executed and objects in the three-dimensional (3D) environment that should be operated on. The wearable system can detect when different inputs converge together, such as when a user seeks to select a virtual object using multiple inputs such as eye gaze, head pose, hand gesture, and totem input. Upon detecting an input convergence, the wearable system can perform a transmodal filtering scheme that leverages the converged inputs to assist in properly interpreting what command the user is providing or what object the user is targeting.