Multimodal Utterance Resolution for Ambiguous In-Vehicle Commands

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing in-vehicle electronic assistant systems struggle to interpret vocal utterances with ambiguity, despite accurate speech recognition, due to insufficient understanding of the meaning and intent of the utterance.

Innovation Solution

A multimodal input system that employs natural language processing and contextual disambiguation to resolve ambiguities in vocal utterances by utilizing various context factors, including previous interactions, vehicle status, and gaze data, to determine the intent and object of the command.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition is used to convert vocal utterances to words, then the conversion accuracy is improved, but the understanding of meaning and intent deteriorates

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidmeaning and intent understanding
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent transitions from single-modal speech recognition to multi-modal input processing by incorporating gaze data, haptic input, and contextual information as additional dimensions. This allows the system to resolve ambiguities that speech recognition alone cannot address, thereby recovering the lost meaning and intent information while maintaining high speech recognition accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces contextual information as an intermediary layer between speech recognition and intent understanding. By incorporating previous interactions, vehicle status, and environmental context, the system bridges the gap between accurate word conversion and meaningful intent interpretation, allowing it to disambiguate utterances that would otherwise be misunderstood.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If multiple input modes are accepted for driver assistance, then the ease of operation is improved, but the complexity of the system deteriorates

Engineering Contradiction:
Improvedriver assistance usabilityVSAvoidmultimodal input system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a unified processing framework that handles multiple input modes (speech, gaze, haptic, text) through a single natural language processing system. This universal approach allows the system to accept diverse inputs without requiring separate processing pipelines for each mode, thereby maintaining ease of operation while managing system complexity through consolidation rather than proliferation of components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the multimodal processing into distinct hierarchical levels: input acquisition, contextualization, disambiguation, and execution. By organizing the complex multimodal system into modular segments with clear interfaces, the system maintains ease of operation through consistent user experience while managing complexity through structured organization of processing stages.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If contextual information is used to disambiguate utterances, then the accuracy of intent determination is improved, but the processing time deteriorates

Engineering Contradiction:
Improveintent determination accuracyVSAvoidutterance processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent pre-processes and stores contextual information (vehicle status, previous interactions, environmental data) in ready-to-use formats before utterances occur. This preliminary preparation allows the system to quickly retrieve and apply relevant context during utterance disambiguation without performing extensive real-time analysis, thereby maintaining high intent determination accuracy while minimizing additional processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies contextual disambiguation selectively based on the ambiguity level of the utterance. For clear, unambiguous commands, the system processes them quickly with minimal contextual analysis. For ambiguous utterances, it applies full contextual disambiguation. This partial application approach ensures high accuracy for problematic cases while avoiding unnecessary processing time for clear commands, thus balancing accuracy and speed.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12531057B2Contextual utterance resolution in multimodal systems
Publication Date: 2026.01.20 CERENCE OPERATING CO
  • US12531057B2 patent drawing
  • US12531057B2 patent drawing
  • US12531057B2 patent drawing

AI summary

A system and method of responding to a vocal utterance may include capturing and converting the utterance to word(s) using a language processing method, such as natural language processing. The context of the utterance and of the system, which may include multimodal inputs, may be used to determine the meaning and intent of the words.