Enhanced Clarification Prompts for Ambiguous Voice Commands

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated assistants often struggle with ambiguous user commands, leading to prolonged interactions, errant inputs, and user frustration due to the limitations of natural language-only clarification prompts, which can result in incorrect disambiguation and abandonment of intended goals.

Innovation Solution

Implementing an enhanced clarification prompt that includes additional audio or visual content, such as musical snippets or images, to assist users in differentiating between candidate responsive actions, selectively provided based on conditions like historical interaction data, semantic similarity, and inverse document frequency, to reduce interaction duration and improve disambiguation accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If natural language-only clarification prompts are used, then the system maintains simplicity in prompt structure, but user interaction time increases and disambiguation accuracy decreases

Engineering Contradiction:
Improveprompt structure complexityVSAvoiduser interaction time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent transitions from one-dimensional text-based clarification prompts to multi-dimensional prompts that incorporate audio snippets, images, and text. This dimensional expansion allows users to disambiguate between candidate actions through multiple sensory channels simultaneously, reducing interaction time without increasing structural complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of manufacture

If natural language-only clarification prompts are used, then the system maintains ease of implementation, but disambiguation accuracy deteriorates

Engineering Contradiction:
Improveimplementation easeVSAvoiddisambiguation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent segments the clarification prompt into distinct modular components: text descriptions, audio snippets, and images. Each component serves a specific disambiguation function and can be independently generated and selected based on the conversation context, maintaining implementation ease while improving accuracy

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If enhanced clarification prompts with audio and visual content are provided, then disambiguation accuracy improves, but system complexity increases

Engineering Contradiction:
Improvedisambiguation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements dynamic selection of clarification prompt types based on real-time analysis of conversation context, user preferences, and candidate action characteristics. The system adaptively determines whether to use text-only, audio-enhanced, image-enhanced, or full multi-modal prompts, optimizing disambiguation accuracy while managing system complexity through context-aware adaptation

Inventive Principle:
Principle #15Dynamics

4Productivity

If enhanced clarification prompts are selectively provided based on conditions, then interaction efficiency improves, but processing requirements increase

Engineering Contradiction:
Improveinteraction efficiencyVSAvoidprocessing resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary analysis of the conversation context, user profile, and candidate actions before generating clarification prompts. This advance preparation allows the system to pre-determine the appropriate prompt type and pre-fetch necessary audio or image resources, improving interaction efficiency while managing processing resources through proactive preparation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230419963A1Selectively providing enhanced clarification prompts in automated assistant interactions
Publication Date: 2023.12.28 GOOGLE LLC
  • US20230419963A1 patent drawing
  • US20230419963A1 patent drawing
  • US20230419963A1 patent drawing

AI summary

Implementations described herein receive audio data that captures a spoken utterance, generate, based on processing the audio data, a recognition that corresponds to the spoken utterance, and determine, based on processing the recognition, that the spoken utterance is ambiguous (i.e., is interpretable as requesting performance of a first particular action exclusively and is also interpretable a second particular action exclusively). In response to determining that the spoken utterance is ambiguous, implementations determine to provide an enhanced clarification prompt that renders output that is in addition to natural language. The enhanced clarification prompt solicits further user interface input for disambiguating between the first particular action and the second particular action. Determining to provide the enhanced clarification prompt includes a current or prior determination to provide the enhanced clarification prompt instead of a natural language (NL) only clarification prompt that is restricted to rendering natural language.