Mixed Reality Voice Control via Multi-Dimensional Reference Element

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing mixed reality devices face challenges in interacting with media within the context of mixed reality environments, as conventional approaches like hand, head, and eye tracking are difficult to implement, costly, and limited by light conditions.

Innovation Solution

A system and method for interacting with mixed reality content using voice inputs and device inputs, where a mixed reality device receives voice commands to control the user's view and navigation within the environment, utilizing a multi-dimensional reference element as an overlay to provide visual feedback on the user's position and orientation, allowing seamless switching between virtual and augmented environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If hand, head, and eye tracking are used to interact with media in mixed reality environments, then interaction capability is improved, but device complexity and cost increase

Engineering Contradiction:
Improveinteraction capabilityVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent extracts the interaction control functionality from complex physical tracking devices and relocates it to voice command processing. Users speak commands like 'go to the next scene' or 'pause playback' which are processed by the system to control media playback and navigation, eliminating the need for hand, head, and eye tracking hardware.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces mechanical tracking systems (hand tracking, head tracking, eye tracking) with an acoustic field-based voice recognition system. The mechanical movement detection is substituted with acoustic wave detection and natural language processing, significantly reducing device complexity while maintaining interaction capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If hand, head, and eye tracking are used to interact with media in mixed reality environments, then interaction capability is improved, but implementation difficulty increases

Engineering Contradiction:
Improveinteraction capabilityVSAvoidimplementation difficulty
Core Design Contradiction:
Ease of operationVSEase of manufacture

Solution Approach 1:

The patent extracts the interaction control functionality from complex physical tracking devices and relocates it to voice command processing. Users speak commands like 'go to the next scene' or 'pause playback' which are processed by the system to control media playback and navigation, eliminating the need for hand, head, and eye tracking hardware.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces mechanical tracking systems (hand tracking, head tracking, eye tracking) with an acoustic field-based voice recognition system. The mechanical movement detection is substituted with acoustic wave detection and natural language processing, significantly reducing device complexity while maintaining interaction capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If conventional tracking approaches are used, then interaction control is achieved, but cost becomes prohibitive

Engineering Contradiction:
Improveinteraction controlVSAvoidcost
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent employs voice recognition technology which uses standard microphones and processors already present in most mixed reality devices, rather than expensive specialized tracking hardware. This approach uses readily available, cost-effective components to achieve the same interaction control functionality.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The patent replaces mechanical tracking systems (hand tracking, head tracking, eye tracking) with an acoustic field-based voice recognition system. The mechanical movement detection is substituted with acoustic wave detection and natural language processing, significantly reducing device complexity while maintaining interaction capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If hand, head, and eye tracking are used for interaction, then control precision is improved, but performance limitations occur under certain light conditions

Engineering Contradiction:
Improvecontrol precisionVSAvoidlight condition adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent replaces optical-based tracking systems (hand tracking, head tracking, eye tracking) with an acoustic field-based voice recognition system. Since acoustic waves are not affected by light conditions, the system maintains consistent control precision whether in bright or dark environments, eliminating the light condition limitations of optical tracking.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11256474B2Multi-dimensional reference element for mixed reality environments
Publication Date: 2022.02.22 METRIK LLC
  • US11256474B2 patent drawing
  • US11256474B2 patent drawing
  • US11256474B2 patent drawing

AI summary

Approaches provide for controlling, managing, and/or otherwise interacting with mixed (e.g., virtual and/or augmented) reality content in response to input from a user, including voice input, device input, among other such inputs, in a mixed reality environment. For example, a mixed reality device, such as a headset or other such device can perform various operations in response to a voice command or other such input. In one such example, the device can receive a voice command and an application executing on the device or otherwise in communication with the device can analyze audio input data of the voice command to control the view of content in the environment, as may include controlling a user's “position” in the environment. The position can include, for example, a specific location in time, space, etc., as well as directionality and field of view of the user in the environment. A reference element can be displayed as an overlay to the mixed reality content, and can provide a visual reference to the user's position in the environment.