Voice-Enabled Virtual Object Disambiguation in Artificial Reality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional XR systems struggle to effectively select a target virtual object for user input when explicit selection mechanics, such as gaze tracking or ray-based interactions, are absent, leading to inflexibility and a degraded user experience.

Innovation Solution

The implementation of a disambiguation and control layer that uses user voice input, derived intentions, and additional context like interaction history or gaze direction to predict and select a target virtual object, allowing for control without explicit selection mechanics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If explicit selection mechanics (gaze tracking, ray-based interactions) are used to select target virtual objects, then object selection accuracy is improved, but device complexity and user interaction burden increase

Engineering Contradiction:
Improveobject selection accuracyVSAvoidselection mechanism complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the selection function from complex explicit mechanics (gaze tracking, ray casting) and implements it through voice commands alone. The voice-activated selection mechanism removes the need for multiple input modalities while maintaining selection accuracy through natural language processing and spatial awareness of virtual object positions.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces voice as an intermediary between the user and virtual objects. Instead of directly using complex selection mechanics, the voice command serves as a mediator that translates user intent into precise object selection through speech recognition and contextual understanding.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If explicit selection mechanics are required for voice control, then control precision is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvecontrol precisionVSAvoiduser interaction simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent merges voice command processing with spatial awareness and object identification into a unified selection mechanism. By combining these functions, the system achieves precise control without requiring separate explicit selection steps, as the voice command directly targets objects based on their spatial positions and contextual information.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs self-service by automatically determining user intent and target objects through voice analysis and spatial context without requiring additional user input. The system serves itself by inferring selection targets from voice commands alone, eliminating the need for explicit selection mechanics.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If multiple input modalities (gaze, controller, voice) are combined for object selection, then adaptability is improved, but device complexity increases

Engineering Contradiction:
Improveinput method flexibilityVSAvoidmulti-modality system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements universality by making voice the universal input modality that can handle all selection and control functions previously requiring multiple specialized inputs. The voice-activated system serves multiple purposes (selection, navigation, control) through a single input method, reducing overall system complexity while maintaining flexibility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250123799A1Voice-Enabled Virtual Object Disambiguation and Controls in Artificial Reality
Publication Date: 2025.04.17 META PLATFORMS INC
  • US20250123799A1 patent drawing
  • US20250123799A1 patent drawing
  • US20250123799A1 patent drawing

AI summary

Aspects of the present disclosure are directed to applying voice controls to a target virtual object in an artificial reality environment. User controls in an artificial reality environment can take many forms. Some user controls, such as ray casting or gaze tracking, can incorporate selection mechanics to select the artificial reality environment element (e.g., virtual object) that the user is targeting for interaction. Other forms of user controls, such as voice controls, may not include such selection mechanics. Implementations disambiguate user voice input to select a target virtual object and control the target virtual object based on the voice input. For example, a disambiguation and control layer can select the virtual object the user intends to target with voice input, format input for the target virtual object using the voice input, and control the target virtual object via execution of one or more applications that manage the virtual object.