Multi-Level Voice Command Interface for Mobile Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

User interfaces that process audible inputs struggle to be adaptable for both beginner and expert users, as powerful voice command languages can be difficult to learn but simple command languages with visual and audible cues can be slow for expert users.

Innovation Solution

A computing device receives spoken utterances in a multi-level command format, identifying applications and actions, and provides visual or audible prompts if additional input is needed, allowing for both continuous and separate command issuance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a powerful voice command language is used, then the functionality and versatility of the voice interface is improved, but the difficulty of learning and ease of operation deteriorates

Engineering Contradiction:
Improvefunctionality of voice interfaceVSAvoidease of use for beginner users
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The voice command language is segmented into multiple levels: Level 1 for application identification and Level 2 for action specification. Users can provide commands in a hierarchical manner, starting with the application and then specifying the action, allowing the system to process partial information and provide structured guidance throughout the interaction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary processing by identifying the application from the first level of the command before requiring the second level action specification. Visual representations of identified applications are displayed immediately, and the system prepares to receive additional input, reducing waiting time and providing continuous feedback to the user.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 3:

The system provides visual feedback by displaying representations of identified applications and audible feedback through prompts when additional information is needed. This feedback mechanism guides users through the command structure, making the powerful voice interface more accessible by showing what has been recognized and what is expected next.

Inventive Principle:
Principle #23Feedback

2Ease of operation

If a simple command language with visual and audible cues is used, then the ease of operation is improved, but the productivity and interaction speed deteriorates

Engineering Contradiction:
Improveease of use for beginner usersVSAvoidinteraction speed for expert users
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The interface dynamically adapts its behavior based on the completeness of the received command. If a complete two-level command is received, the system processes it immediately for fast interaction. If only the first level is received, the system enters a waiting state for additional input, providing visual and audible cues as needed. This dynamic response optimizes both speed and usability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system maintains continuous processing by immediately identifying and displaying the application from the first level of the command while waiting for the second level action specification. This continuous useful action reduces idle time and keeps the interaction flow uninterrupted, improving productivity without sacrificing ease of use.

Inventive Principle:
Principle #20Continuity of useful action

3Measurement precision

If the system waits for additional spoken utterance, then the accuracy of command recognition is improved, but the loss of time increases

Engineering Contradiction:
Improveaccuracy of command recognitionVSAvoidwaiting time for additional input
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary identification of the application from the first level of the command immediately, displaying the visual representation without waiting for the second level action specification. This preliminary action reduces perceived waiting time while the system continues to wait for the complete command in the background, balancing accuracy with responsiveness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides immediate visual feedback by displaying the identified application representation as soon as the first level is recognized, and audible feedback through prompts when the second level is expected. This feedback reduces user uncertainty and perceived waiting time, making the accuracy-improving wait feel more manageable and less time-consuming.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8452602B1Structuring verbal commands to allow concatenation in a voice interface in a mobile device
Publication Date: 2013.05.28 GOOGLE LLC
  • US8452602B1 patent drawing
  • US8452602B1 patent drawing
  • US8452602B1 patent drawing

AI summary

A spoken utterance includes at least a first level of a multi-level command format, in which the first level identifies an application. The spoken utterance may also include a second level of the multi-level command format, in which the second level identifies an action. In response to receiving the spoken utterance at a computing device, a representation of the application identified by the first level is displayed on a display of the computing device. If the spoken utterance includes the second level of the multi-level command format, the action identified by the second level is initiated. If the spoken utterance does not include the second level of the multi-level command format, the computing device waits for a predetermined period of time and provides at least one of an audible or visual action prompt if the second level is not received within the predetermined period of time.