Conversation Evaluation Apparatus Using Overlap and Silence Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing conversation evaluation technologies cannot objectively assess the impression of the first speaker independently from the second speaker, leading to unfair evaluations, especially when the second speaker cuts in while the first speaker is speaking.

Innovation Solution

A conversation evaluation apparatus and method that estimates the starting and ending times of each main speaker's utterance, identifies the timing of switches between main speakers, and evaluates the conversation state based on dialogue information including overlapping and silent segments, using a detection model trained on dialogue data to provide objective evaluations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If the evaluation method uses only the overlapping segment length between two speakers, then the evaluation process is simple, but the evaluation result becomes unfair and inaccurate when one speaker cuts in while the other is speaking

Engineering Contradiction:
Improveevaluation process complexityVSAvoidconversation evaluation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the conversation evaluation into multiple independent components: overlapping segment analysis, silent segment analysis, and turn-switch timing analysis. Each segment is evaluated separately and then integrated to produce a comprehensive assessment, allowing fair evaluation of each speaker's contribution without being biased by interruptions

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds new evaluation dimensions beyond just overlapping segment length. It introduces silent segment length analysis and turn-switch timing analysis as additional dimensions, transforming the evaluation from a single-dimensional metric to a multi-dimensional assessment that captures the full dynamics of speaker interactions

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If the evaluation considers only the second speaker's position, then the evaluation method is straightforward, but it cannot objectively evaluate the first speaker's impression independently

Engineering Contradiction:
Improveevaluation method simplicityVSAvoidfirst speaker's impression information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent inverts the traditional evaluation perspective by analyzing the conversation from both speakers' viewpoints rather than just the second speaker's position. It evaluates how the first speaker's utterance is affected by the second speaker's interruption, providing a balanced assessment that captures both perspectives equally

Inventive Principle:
Principle #13The other way round (Inversion)

3Measurement precision

If the evaluation uses multiple dialogue information parameters, then the evaluation accuracy improves, but the analysis complexity and processing time increase

Engineering Contradiction:
Improveconversation state evaluation accuracyVSAvoidanalysis processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and focuses on the three most critical dialogue information parameters: overlapping segment length, silent segment length, and turn-switch timing. By selecting only these essential parameters rather than analyzing all possible conversation features, it achieves high evaluation accuracy while keeping processing complexity manageable

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240404550A1Non-transitory computer readable medium, and conversation evaluation apparatus and method
Publication Date: 2024.12.05 KK TOSHIBA
  • US20240404550A1 patent drawing
  • US20240404550A1 patent drawing
  • US20240404550A1 patent drawing

AI summary

According to one embodiment, a non-transitory computer readable medium includes computer executable instructions. The instructions, when executed by a processor, cause the processor to perform a method. The method estimates a starting time and an ending time of an utterance of each main speaker relating to a conversation. The method identifies a timing of a switch between the main speakers. The method evaluates a state of the conversation based on dialogue information before and after the identified timing of the switch. The dialogue information includes at least one of a length of an overlapping segment or a length of a silent segment, and includes a length of an utterance segment.