Spoken Dialog System Prominence Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current spoken dialog systems are inadequate in understanding human speech due to their insensitivity to prosodic cues, leading to misunderstandings and a lack of intuition in human-machine interaction, especially in noisy environments or when the speaking style differs from expectations.

Innovation Solution

A method and system that analyze both acoustic and visual signals to determine the prominence of parts of an utterance, using prosodic cues to improve speech recognition accuracy and dialog management by identifying and correcting misunderstandings through emphasis detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If prosodic cues are integrated into the spoken dialog system, then speech recognition accuracy and dialog management improve, but system complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments prosodic information into distinct components (pitch, energy, duration) that can be processed independently. Each prosodic feature is extracted and analyzed separately before being integrated into the dialog management system, making the complex task of prosodic analysis more manageable and computationally efficient.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces prosodic cues as an intermediary layer between the raw speech signal and the dialog management system. This intermediary processing layer analyzes prosodic features and provides enhanced information to the dialog manager, improving recognition accuracy without requiring complete reengineering of the entire system.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the system processes both acoustic and visual signals to determine prominence, then understanding accuracy improves, but processing time and computational load increase

Engineering Contradiction:
Improveunderstanding accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary extraction of prosodic features from the acoustic signal before full dialog processing. By pre-processing and identifying prominent segments based on prosodic cues early in the pipeline, the system reduces the computational burden on subsequent processing stages and minimizes overall processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the most relevant prosodic features (pitch, energy, duration) that are necessary for determining prominence, rather than processing all possible acoustic and visual parameters. This selective extraction reduces computational load while maintaining understanding accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If the system uses prosodic cues to identify misunderstood parts, then dialog efficiency improves, but the difficulty of detecting and measuring prominence increases

Engineering Contradiction:
Improvedialog efficiencyVSAvoidprominence detection difficulty
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The system transforms complex prosodic patterns into simplified prominence scores by analyzing changes in pitch, energy, and duration parameters. By monitoring parameter changes and their combinations, the system can identify misunderstood parts efficiently without requiring complex interpretation of raw prosodic data.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback mechanisms where the system continuously monitors prosodic cues during dialog and adjusts its understanding in real-time. When prominence patterns indicate potential misunderstandings, the system can request clarification or rephrase, improving dialog efficiency through adaptive feedback loops.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP2645364B1Spoken dialog system using prominence
Publication Date: 2019.05.08 HONDA RES INST EUROPE
  • EP2645364B1 patent drawingFigure 1~2
  • EP2645364B1 patent drawingFigure 3~4

AI summary

The invention presents a method for analyzing speech in a spoken dialog system, comprising the steps of: accepting an utterance by at least one means for accepting acoustical signals, in particular a microphone, analyzing the utterance and obtaining prosodic cues from the utterance using at least one processing engine, wherein the utterance is evaluated based on the prosodic cues to determine a prominence of parts of the utterance, and wherein the utterance is analyzed to detect either at least one marker feature, e.g. a negative statement, a segment with a very high prominence or both, indicative of the utterance containing at least one part to replace at least one part in a previous utterance, the part to be replaced in the previous utterance being determined based on the prominence determined for the parts of the previous utterance and the replacement parts being determined based on the prominence of the parts in the utterance, and wherein the previous utterance is evaluated with the replacement part(s).