Multi-Level Voice Command Interface for Mobile Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
User interfaces that process audible inputs struggle to be adaptable for both beginner and expert users, as powerful voice command languages can be difficult to learn but simple command languages with visual and audible cues can be slow for expert users.
Innovation Solution
A computing device receives spoken utterances in a multi-level command format, identifying applications and actions, and provides visual or audible prompts if additional input is needed, allowing for both continuous and separate command issuance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a powerful voice command language is used, then the functionality and versatility of the voice interface is improved, but the difficulty of learning and ease of operation deteriorates
Solution Approach 1:
The voice command language is segmented into multiple levels: Level 1 for application identification and Level 2 for action specification. Users can provide commands in a hierarchical manner, starting with the application and then specifying the action, allowing the system to process partial information and provide structured guidance throughout the interaction.
Solution Approach 2:
The system performs preliminary processing by identifying the application from the first level of the command before requiring the second level action specification. Visual representations of identified applications are displayed immediately, and the system prepares to receive additional input, reducing waiting time and providing continuous feedback to the user.
Solution Approach 3:
The system provides visual feedback by displaying representations of identified applications and audible feedback through prompts when additional information is needed. This feedback mechanism guides users through the command structure, making the powerful voice interface more accessible by showing what has been recognized and what is expected next.
2Ease of operation
If a simple command language with visual and audible cues is used, then the ease of operation is improved, but the productivity and interaction speed deteriorates
Solution Approach 1:
The interface dynamically adapts its behavior based on the completeness of the received command. If a complete two-level command is received, the system processes it immediately for fast interaction. If only the first level is received, the system enters a waiting state for additional input, providing visual and audible cues as needed. This dynamic response optimizes both speed and usability.
Solution Approach 2:
The system maintains continuous processing by immediately identifying and displaying the application from the first level of the command while waiting for the second level action specification. This continuous useful action reduces idle time and keeps the interaction flow uninterrupted, improving productivity without sacrificing ease of use.
3Measurement precision
If the system waits for additional spoken utterance, then the accuracy of command recognition is improved, but the loss of time increases
Solution Approach 1:
The system performs preliminary identification of the application from the first level of the command immediately, displaying the visual representation without waiting for the second level action specification. This preliminary action reduces perceived waiting time while the system continues to wait for the complete command in the background, balancing accuracy with responsiveness.
Solution Approach 2:
The system provides immediate visual feedback by displaying the identified application representation as soon as the first level is recognized, and audible feedback through prompts when the second level is expected. This feedback reduces user uncertainty and perceived waiting time, making the accuracy-improving wait feel more manageable and less time-consuming.
Data Source
AI summary
A spoken utterance includes at least a first level of a multi-level command format, in which the first level identifies an application. The spoken utterance may also include a second level of the multi-level command format, in which the second level identifies an action. In response to receiving the spoken utterance at a computing device, a representation of the application identified by the first level is displayed on a display of the computing device. If the spoken utterance includes the second level of the multi-level command format, the action identified by the second level is initiated. If the spoken utterance does not include the second level of the multi-level command format, the computing device waits for a predetermined period of time and provides at least one of an audible or visual action prompt if the second level is not received within the predetermined period of time.


