Mixed Reality Voice Control via Multi-Dimensional Reference Element
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing mixed reality devices face challenges in interacting with media within the context of mixed reality environments, as conventional approaches like hand, head, and eye tracking are difficult to implement, costly, and limited by light conditions.
Innovation Solution
A system and method for interacting with mixed reality content using voice inputs and device inputs, where a mixed reality device receives voice commands to control the user's view and navigation within the environment, utilizing a multi-dimensional reference element as an overlay to provide visual feedback on the user's position and orientation, allowing seamless switching between virtual and augmented environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If hand, head, and eye tracking are used to interact with media in mixed reality environments, then interaction capability is improved, but device complexity and cost increase
Solution Approach 1:
The patent extracts the interaction control functionality from complex physical tracking devices and relocates it to voice command processing. Users speak commands like 'go to the next scene' or 'pause playback' which are processed by the system to control media playback and navigation, eliminating the need for hand, head, and eye tracking hardware.
Solution Approach 2:
The patent replaces mechanical tracking systems (hand tracking, head tracking, eye tracking) with an acoustic field-based voice recognition system. The mechanical movement detection is substituted with acoustic wave detection and natural language processing, significantly reducing device complexity while maintaining interaction capability.
2Ease of operation
If hand, head, and eye tracking are used to interact with media in mixed reality environments, then interaction capability is improved, but implementation difficulty increases
Solution Approach 1:
The patent extracts the interaction control functionality from complex physical tracking devices and relocates it to voice command processing. Users speak commands like 'go to the next scene' or 'pause playback' which are processed by the system to control media playback and navigation, eliminating the need for hand, head, and eye tracking hardware.
Solution Approach 2:
The patent replaces mechanical tracking systems (hand tracking, head tracking, eye tracking) with an acoustic field-based voice recognition system. The mechanical movement detection is substituted with acoustic wave detection and natural language processing, significantly reducing device complexity while maintaining interaction capability.
3Ease of operation
If conventional tracking approaches are used, then interaction control is achieved, but cost becomes prohibitive
Solution Approach 1:
The patent employs voice recognition technology which uses standard microphones and processors already present in most mixed reality devices, rather than expensive specialized tracking hardware. This approach uses readily available, cost-effective components to achieve the same interaction control functionality.
Solution Approach 2:
The patent replaces mechanical tracking systems (hand tracking, head tracking, eye tracking) with an acoustic field-based voice recognition system. The mechanical movement detection is substituted with acoustic wave detection and natural language processing, significantly reducing device complexity while maintaining interaction capability.
4Measurement precision
If hand, head, and eye tracking are used for interaction, then control precision is improved, but performance limitations occur under certain light conditions
Solution Approach 1:
The patent replaces optical-based tracking systems (hand tracking, head tracking, eye tracking) with an acoustic field-based voice recognition system. Since acoustic waves are not affected by light conditions, the system maintains consistent control precision whether in bright or dark environments, eliminating the light condition limitations of optical tracking.
Data Source
AI summary
Approaches provide for controlling, managing, and/or otherwise interacting with mixed (e.g., virtual and/or augmented) reality content in response to input from a user, including voice input, device input, among other such inputs, in a mixed reality environment. For example, a mixed reality device, such as a headset or other such device can perform various operations in response to a voice command or other such input. In one such example, the device can receive a voice command and an application executing on the device or otherwise in communication with the device can analyze audio input data of the voice command to control the view of content in the environment, as may include controlling a user's “position” in the environment. The position can include, for example, a specific location in time, space, etc., as well as directionality and field of view of the user in the environment. A reference element can be displayed as an overlay to the mixed reality content, and can provide a visual reference to the user's position in the environment.


