Speech-Enabled AR Interface Voice Command Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Augmented Reality (AR) user interfaces are limited to specific inputs such as touchscreens or pointing devices, lacking the ability to effectively interpret and respond to voice commands, which restricts their functionality and user interaction.
Innovation Solution
A system that extends speech processing capabilities to AR user interfaces, enabling devices to capture images and display information through voice commands by processing audio data using automatic speech recognition (ASR) and natural language understanding (NLU), allowing users to interact with AR interfaces using voice commands to retrieve information about objects or features, compare vehicles, and receive contextual information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional AR user interfaces use touchscreens or pointing devices for input, then the interface can display information visually, but the functionality and user interaction capability are restricted
Solution Approach 1:
The patent applies universality by enabling the AR interface to accept multiple types of input methods simultaneously - voice commands, touch inputs, and pointer device inputs. The speech processing module processes audio data through ASR and NLU to interpret various command formats, making the system adaptable to different user preferences and situations without requiring separate interfaces for each input type.
Solution Approach 2:
The patent replaces mechanical input methods (touchscreens and pointing devices) with voice-based acoustic input. The speech processing system captures audio data from microphones, converts it to text through automatic speech recognition, and interprets it through natural language understanding, substituting physical interaction with acoustic field interaction to enhance user interaction capability.
2Adaptability or versatility
If speech processing is added to AR interfaces, then voice-controlled interaction is enabled, but system complexity increases
Solution Approach 1:
The patent segments the speech processing functionality into distinct modular components: an automatic speech recognition (ASR) module that converts audio to text, and a natural language understanding (NLU) module that interprets the text commands. This segmentation allows each component to be optimized independently and integrated into the existing AR interface architecture, managing complexity through functional decomposition.
Solution Approach 2:
The patent introduces text data as an intermediary between audio input and command execution. The ASR module converts spoken commands into text, which then serves as a bridge for the NLU module to process and interpret. This intermediary layer simplifies the overall processing pipeline by breaking down the complex task of voice command interpretation into manageable stages with clear interfaces between components.
Data Source
AI summary
A speech interface device is configured to display an Augmented Reality (AR) user interface that displays information specific to a vehicle or other object. For example, the device may capture images of the vehicle and may displaying the AR user interface with labels, graphical elements, visual effects, and/or additional information superimposed above corresponding portions of the vehicle represented in the images. Using a remote system to perform speech processing, the device may respond to a voice command, enabling the AR user interface to display specific information about the vehicle and/or features of the vehicle in response to the voice command. The device may also send position data indicating information about what is displayed on the AR user interface, enabling the remote system to provide information about specific features or components based on where the device is pointed.


