AR Wearable Sensor Fusion for Precise 3D Object Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional input techniques in AR/VR/MR environments require high specificity and suffer from high error rates and fatigue due to imprecise commands, making it challenging to accurately interact with virtual or real objects in 3D space.
Innovation Solution
A wearable device dynamically fuses multiple sensor inputs, such as head pose, eye gaze, hand gestures, and voice commands, to anticipate and predict user intent, dynamically adding or removing inputs based on convergence and divergence, reducing reliance on high-resolution sensors and enhancing interaction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional input techniques are used in AR/VR/MR environments, then the system requires high specificity and precision in user inputs, but this leads to high error rates and user fatigue
Solution Approach 1:
The patent combines multiple sensor inputs (head pose, eye gaze, hand gestures, voice commands) into a fused input signal. By merging these different modalities, the system achieves more reliable object selection and command execution, reducing errors that occur with single-input methods while maintaining natural interaction.
Solution Approach 2:
The system continuously monitors the convergence of multiple sensor inputs and dynamically adjusts which sensors are active based on their contribution to the current task. This feedback mechanism allows the system to reduce noise and errors by deactivating sensors that are not providing useful information, thereby improving overall input reliability.
2Adaptability or versatility
If multiple sensor inputs are continuously used, then the system can capture comprehensive user intent, but this increases device complexity and computational load
Solution Approach 1:
The patent implements dynamic sensor selection where the system actively adjusts which sensors are enabled based on the current interaction context and convergence of inputs. This dynamic approach maintains input versatility by having multiple sensors available when needed, while reducing device complexity by deactivating unnecessary sensors during different phases of interaction.
Solution Approach 2:
The system segments the sensor input processing into distinct phases: initial exploration phase where multiple sensors are active, and convergence phase where only relevant sensors continue. This segmentation allows the system to handle complex multi-sensor data when necessary while simplifying processing during stable interaction states.
3Measurement precision
If high-resolution sensors are used to ensure accurate input detection, then measurement precision improves, but hardware cost and complexity increase
Solution Approach 1:
The patent combines data from multiple lower-resolution or simpler sensors to achieve the equivalent precision of a single high-resolution sensor. By fusing head pose, eye gaze, hand gestures, and voice commands, the system achieves accurate object selection without requiring expensive high-resolution sensors for each individual input modality.
Solution Approach 2:
The system uses an intermediary processing layer that fuses inputs from multiple sensors before final decision-making. This intermediary fusion process allows simpler sensors to contribute to accurate overall detection, reducing the need for any single sensor to be extremely high-resolution while maintaining overall system accuracy.
4Measurement precision
If the system dynamically selects and fuses sensor inputs, then interaction accuracy improves, but processing time and computational resources increase
Solution Approach 1:
The system performs preliminary convergence assessment of sensor inputs to determine early whether multiple sensors should be fused. By evaluating the convergence of inputs in advance, the system can quickly decide to use only the most relevant sensors, avoiding the computational overhead of processing all sensors in every situation while maintaining high interaction accuracy when convergence is achieved.
Data Source
AI summary
Examples of wearable systems and methods can use multiple inputs (e.g., gesture, head pose, eye gaze, voice, totem, and/or environmental factors (e.g., location)) to determine a command that should be executed and objects in the three-dimensional (3D) environment that should be operated on. The wearable system can detect when different inputs converge together, such as when a user seeks to select a virtual object using multiple inputs such as eye gaze, head pose, hand gesture, and totem input. Upon detecting an input convergence, the wearable system can perform a transmodal filtering scheme that leverages the converged inputs to assist in properly interpreting what command the user is providing or what object the user is targeting.


