Speech-Controlled Extended Reality with 3D Relationship Visualization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional extended reality systems limit user interaction and collaboration due to cumbersome input mechanisms, such as hand-held controllers, which frustrate users seeking to creatively interact and collaborate in extended reality environments.

Innovation Solution

The system allows users to interact naturally by providing speech inputs to control the extended reality environment, with additional cues like gaze direction and gestures to supplement and disambiguate speech, using a neural network to determine semantic meaning and generate 3D representations of relationships between objects and concepts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If hand-held controllers are used for interaction, then users can select and manipulate objects in the extended reality environment, but the interaction becomes cumbersome and frustrating

Engineering Contradiction:
Improveease of interactionVSAvoidinput mechanism complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces mechanical hand-held controllers with voice-based and gesture-based input systems. Users can interact with the extended reality environment through natural speech commands and body gestures, eliminating the need for physical controllers and significantly improving ease of operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces an intermediary processing layer that captures voice inputs and gesture inputs, interprets them through neural networks, and translates them into appropriate actions within the extended reality environment. This intermediary layer bridges the gap between natural human communication and digital system control.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If conventional input mechanisms are used, then users can control the extended reality environment, but collaboration and creative interaction are limited

Engineering Contradiction:
Improvecollaboration capabilityVSAvoiduser frustration
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system implements multiple input modalities (voice, gestures, and potentially eye tracking) that can be used individually or in combination. This multi-functional approach allows users to interact naturally and collaborate effectively, as different users can leverage their preferred input methods while working together in the same extended reality environment.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If detailed menu navigation and keyboard interaction are required, then precise control is achieved, but computing resources are consumed and user efficiency is reduced

Engineering Contradiction:
Improveuser efficiencyVSAvoidcomputing resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system employs neural networks that can autonomously interpret voice inputs and gestures, reducing the need for extensive menu navigation and keyboard interaction. The intelligent processing systems automatically understand user intent and execute appropriate actions, improving productivity while optimizing computing resource usage through efficient pattern recognition.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11430186B2Visually representing relationships in an extended reality environment
Publication Date: 2022.08.30 META PLATFORMS TECHNOLOGIES LLC
  • US11430186B2 patent drawing
  • US11430186B2 patent drawing
  • US11430186B2 patent drawing

AI summary

Techniques are described herein that enable a user to provide speech inputs to control an extended reality environment, where relationships between terms in a speech input are represented in three dimensions (3D) in the extended reality environment. For example, a language processing component determines a semantic meaning of the speech input, and identifies terms in the speech input based on the semantic meaning. A 3D relationship component generates a 3D representation of a relationship between the terms and provides the 3D representation to a computing device for display. A 3D representation may include a modification to an object in an extended reality environment, or a 3D representation of a concepts and sub-concepts in a mind map in an extended reality environment, for example. The 3D relationship component may generate a searchable timeline using the terms provided in the speech input and a recording of an extended reality session.