Voice Agent Visual Feedback for Seamless Task Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice agents require users to interrupt their current tasks and provide extensive input to perform actions, lacking integration with other application programs and contextual information, which hinders efficient interaction and user experience.

Innovation Solution

A voice agent that can be invoked without interrupting ongoing tasks, utilizing contextual information to interpret user input and provide visual feedback, allowing partial specification of actions and seamless integration with other application programs, enabling simultaneous interaction with both the voice agent and other applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional voice agents are used to perform actions, then the user can interact with the computing device by voice input, but the user must interrupt current tasks and provide extensive input

Engineering Contradiction:
Improvevoice interactionVSAvoidtask interruption
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The voice agent allows users to perform partial actions through voice input without requiring complete task specification. The system interprets partial voice commands and presents visual representations that complete the action, reducing the need for extensive input while maintaining ease of operation without full task interruption

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary interpretation of voice input to identify relevant application programs and actions before full task execution. By pre-processing the voice command and presenting visual representations in advance, the system prepares the interaction state without forcing complete task interruption, allowing smoother transitions

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If conventional voice agents are used, then voice input can be processed, but integration with other application programs and contextual information is lacking

Engineering Contradiction:
Improvevoice agent functionalityVSAvoidcontextual information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The voice agent system merges with the operating system and application programs to enable seamless integration. The system combines voice input processing with contextual information from active applications, allowing the voice agent to access and utilize information from multiple sources simultaneously without losing contextual information

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The voice agent is designed as a universal interface that can interact with multiple application programs and data sources. By implementing multi-functionality across different applications and contexts, the system maintains adaptability while preserving contextual information through centralized access to system-wide resources

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If visual representations are displayed for voice input interpretation, then user confirmation is improved, but the interface complexity increases

Engineering Contradiction:
Improveinput interpretation accuracyVSAvoidinterface complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements feedback by displaying visual representations of interpreted voice input to confirm understanding before execution. This feedback mechanism improves reliability by allowing users to verify interpretation accuracy, while the visual feedback is integrated into the existing interface framework to minimize additional complexity

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Visual representations serve as an intermediary between voice input and task execution. This mediator layer translates voice commands into visual forms that confirm interpretation without requiring complex interface elements, bridging the gap between voice processing and application interaction while maintaining interface simplicity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10276157B2Systems and methods for providing a voice agent user interface
Publication Date: 2019.04.30 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10276157B2 patent drawing
  • US10276157B2 patent drawing
  • US10276157B2 patent drawing

AI summary

Some embodiments provide techniques performed by at least one voice agent. The techniques include receiving voice input; identifying at least one application program as relating to the received voice input; and displaying at least one selectable visual representation that, when selected, causes focus of the computing device to be directed to the at least one application program identified as relating to the received voice input.