Cooperative Conversational Voice Interface for Natural Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Human-to-Machine interfaces lack intuitive interaction, requiring users to use specific commands and restricting dialogue, failing to bridge the gap between human conversational speech and system understanding, thus inhibiting mass-market adoption.

Innovation Solution

A cooperative conversational voice user interface that processes free-form human utterances, using a speech recognition engine and conversational speech engine to generate adaptive responses, accounting for variations in speech and context, and tolerating noise and imperfect speech, enabling users to interact naturally with systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing speech interfaces use fixed command sets and simple languages, then system understanding is improved, but user interaction intuitiveness deteriorates

Engineering Contradiction:
Improvesystem understandingVSAvoiduser interaction intuitiveness
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent implements dynamic language models that adapt to conversational context, allowing the system to understand evolving user intents rather than relying on static command sets. The language model dynamically adjusts based on conversation history, enabling natural follow-up questions and references without requiring users to repeat full commands.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces a conversational language processor as an intermediary between the speech recognizer and the task execution system. This processor bridges the gap by interpreting natural language utterances in context, resolving ambiguities, and translating user intent into system commands without requiring users to learn specific command syntax.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If existing interfaces require specific commands and phrases, then task execution accuracy is improved, but conversational flexibility deteriorates

Engineering Contradiction:
Improvetask execution accuracyVSAvoidconversational flexibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary context analysis and intent classification before task execution. The conversational language processor pre-processes utterances by identifying user intent, extracting relevant parameters, and resolving ambiguities based on conversation history, ensuring accurate task execution while accepting flexible natural language input.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically changes language model parameters based on conversational context. The system adjusts its understanding thresholds, vocabulary focus, and interpretation strategies according to the conversation state, allowing it to maintain high accuracy for specific tasks while remaining flexible to various expression styles.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If speech interfaces use simple instruction sets, then system reliability is improved, but user learning requirements increase

Engineering Contradiction:
Improvesystem reliabilityVSAvoiduser learning time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The conversational language processor provides self-service by automatically adapting to user language patterns and preferences during the conversation. The system learns from user corrections and preferences in real-time, adjusting its interpretation without requiring explicit user training or memorization of command structures.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements feedback loops where the system monitors user corrections, clarifications, and preference expressions during conversation. This feedback is used to dynamically adjust the language model and interpretation strategies, improving reliability while eliminating the need for upfront user learning of specific commands.

Inventive Principle:
Principle #23Feedback

4Speed

If existing systems process each utterance independently, then processing speed is improved, but contextual understanding deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidcontextual understanding
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The patent segments the conversational processing into parallel components: real-time speech recognition, concurrent context analysis, and integrated intent resolution. The conversational language processor maintains separate but synchronized processing streams for current utterance interpretation and conversation history analysis, combining results to achieve both speed and contextual understanding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a temporal dimension to utterance processing by incorporating conversation history and context as additional processing dimensions. Rather than processing only the current utterance, the system simultaneously analyzes the utterance within the temporal context of the ongoing conversation, maintaining speed through parallel processing while gaining contextual understanding.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11222626B2System and method for a cooperative conversational voice user interface
Publication Date: 2022.01.11 VB ASSETS LLC
  • US11222626B2 patent drawing
  • US11222626B2 patent drawing
  • US11222626B2 patent drawing

AI summary

A cooperative conversational voice user interface is provided. The cooperative conversational voice user interface may build upon short-term and long-term shared knowledge to generate one or more explicit and/or implicit hypotheses about an intent of a user utterance. The hypotheses may be ranked based on varying degrees of certainty, and an adaptive response may be generated for the user. Responses may be worded based on the degrees of certainty and to frame an appropriate domain for a subsequent utterance. In one implementation, misrecognitions may be tolerated, and conversational course may be corrected based on subsequent utterances and/or responses.