Speech Dialog Management Using Partial Orthography for Ambiguity Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Vehicle speech recognition systems often generate clumsy or difficult-to-follow spoken commands when additional information is needed for recognition, leading to recognition failures.

Innovation Solution

A speech dialog management system that uses partial orthography to analyze ambiguities in user utterances and generates targeted speech prompts to resolve recognition issues, including modules for recognizing speech, identifying ambiguities, and generating prompts based on orthographic differences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech recognition systems request additional information when recognition is ambiguous, then recognition accuracy may improve, but the dialog becomes clumsy and difficult to follow

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddialog usability
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces text-based visual displays as an intermediary between speech recognition and user communication. Instead of using only spoken prompts (which can be clumsy), the system displays text representations of ambiguous speech inputs and possible interpretations on a visual interface, allowing users to clarify their intent through text selection or correction. This intermediary text display resolves the contradiction by maintaining recognition accuracy while dramatically improving dialog usability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions the dialog interaction from a single dimension (spoken audio only) to multiple dimensions by adding visual text display. The system presents speech recognition results, ambiguous alternatives, and correction options as text on a visual interface, creating a two-dimensional interaction space (audio + visual) that resolves the limitations of speech-only dialogs while maintaining or improving recognition accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If speech systems use spoken commands to request clarification, then recognition issues can be addressed, but the commands may fail to resolve the recognition issue

Engineering Contradiction:
Improverecognition resolution effectivenessVSAvoiddialog management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses text display as an intermediary that presents multiple possible interpretations of ambiguous speech inputs. Instead of relying solely on spoken clarification (which may fail), the system shows text-based alternatives and allows users to select or correct the intended meaning. This intermediary text interface increases the reliability of recognition resolution while managing dialog complexity through structured text presentation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary action by displaying text representations of possible recognition results before final interpretation is made. The system presents multiple plausible interpretations in text form, allowing users to confirm or correct the intended meaning before the system proceeds. This preliminary text-based clarification prevents recognition failures downstream and reduces the need for complex multi-turn spoken dialog.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9202459B2Methods and systems for managing dialog of speech systems
Publication Date: 2015.12.01 GM GLOBAL TECHNOLOGY OPERATIONS LLC
  • US9202459B2 patent drawing
  • US9202459B2 patent drawing
  • US9202459B2 patent drawing

AI summary

Methods and systems are provided for managing speech dialog of a speech system. In one embodiment, a method includes: receiving a first utterance from a user of the speech system; determining a first list of possible results from the first utterance, wherein the first list includes at least two elements that each represent a possible result; analyzing the at least two elements of the first list to determine an ambiguity of the elements; and generating a speech prompt to the user based on partial orthography and the ambiguity.