Visual Interface for Multi-Token Speech Input Prompting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional methods for vocally interacting with multimodal devices are inefficient due to the need for serial message relaying and lack of utilization of visual displays for prompting and cueing multi-token speech input, leading to user frustration and dissatisfaction.

Innovation Solution

The implementation of an enhanced visual interface that uses visual cues to prompt and cue users for multi-token speech input, allowing minimal prompting and enabling users to construct multi-token speech inputs by highlighting active and incomplete input fields within the interface.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional serial prompting methods are used, then the device can receive speech input, but the interaction becomes slow and causes user frustration

Engineering Contradiction:
Improveinteraction speedVSAvoiduser frustration
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent transitions from traditional auditory-only prompting to a visual dimension by displaying input fields and their locations on a screen. This allows users to see where their speech input will be directed, enabling faster and more intuitive interaction without serial prompting delays.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system pre-displays the interface with input fields visible before the user speaks. This preliminary visual setup allows users to understand the input structure in advance, reducing interaction time and preventing frustration by eliminating the need for sequential verbal prompts.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If visual displays are not utilized for prompting, then the interface remains simple, but multi-token speech input becomes inefficient

Engineering Contradiction:
Improvespeech input efficiencyVSAvoidinterface complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

By adding visual display capability to the speech interface, the system enables users to see input field locations and structures. This visual dimension dramatically improves multi-token speech input efficiency without significantly increasing overall device complexity, as it leverages existing display hardware.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10521186B2Systems and methods for prompting multi-token input speech
Publication Date: 2019.12.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10521186B2 patent drawing
  • US10521186B2 patent drawing
  • US10521186B2 patent drawing

AI summary

A method for prompting user input for a multimodal interface including the steps of providing a multimodal interface to a user, where the interface includes a visual interface having a plurality of input regions, each having at least one input field; selecting an input region and processing a multi-token speech input provided by the user, where the processed speech input includes at least one value for at least one input field of the selected input region; and storing at least one value in at least one input field.