Conversation Structure Analysis via Acoustic Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Call centers face the challenge of analyzing large volumes of customer-agent conversations to improve service quality, as human analysis is unfeasible and existing speech recognition systems are unreliable, especially with accents and poor acoustic conditions.

Innovation Solution

A system and method for analyzing conversations by determining speaker turns, calculating per-turn statistics such as floor-transfer times, within-floor speech ratios, and pause durations, and applying a statistical model like hidden Markov modeling to identify conversation phases without relying on speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition system is used to analyze conversations, then words can be recognized from spoken content, but reliability decreases with accents and poor acoustic conditions

Engineering Contradiction:
Improveword recognition accuracyVSAvoidreliability of generated words
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent extracts and analyzes acoustic features (energy, zero-crossing rate, spectral characteristics) directly from the audio signal without converting speech to text. This bypasses the speech recognition step entirely, eliminating the reliability issues associated with accents and poor acoustic conditions while still enabling conversation quality analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the speech recognition system (mechanical/algorithmic text generation process) with a direct acoustic feature analysis approach. Instead of using speech-to-text conversion, the system analyzes raw acoustic properties of the speech signal to determine conversation phases and quality metrics.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If human analysis of every conversation is performed, then detailed quality information can be obtained, but productivity becomes unachievable due to volume

Engineering Contradiction:
Improvequality analysis detailVSAvoidnumber of conversations analyzed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automated self-analysis of conversations by computing acoustic features and applying statistical models (such as hidden Markov models) to automatically identify conversation phases and quality metrics without human intervention. This allows high-volume processing while maintaining detailed analysis capability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the analysis approach by changing from text-based analysis to acoustic parameter-based analysis. By computing and analyzing acoustic features (energy, zero-crossing rate, spectral characteristics) and applying statistical models, the system achieves both detailed quality assessment and high processing throughput.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If unconstrained speech recognition is used, then speech content can be converted to text, but significant computational resources are required

Engineering Contradiction:
Improvespeech content preservationVSAvoidcomputational resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential acoustic features (energy, zero-crossing rate, spectral characteristics) needed for conversation analysis rather than performing full speech-to-text conversion. This selective extraction significantly reduces computational resource requirements while preserving the information needed to determine conversation phases and quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10354654B2Conversation structure analysis
Publication Date: 2019.07.16 AVAYA INC
  • US10354654B2 patent drawing
  • US10354654B2 patent drawing
  • US10354654B2 patent drawing

AI summary

Embodiments disclosed herein provide systems, methods, and computer readable media for analyzing a conversation between a plurality of participants. In a particular embodiment, a method provides determining a first speaker from the plurality of participants and determining a second speaker from the plurality of participants. The method further provides determining a first plurality of turns comprising portions of the conversation when the first speaker is speaking and determining a second plurality of turns comprising portions of the conversation when the second speaker is speaking The method also provides determining per-turn statistics for turns of the first and second pluralities of turns and identifying phases of the conversation based on the per-turn statistics.