Conversation Structure Analysis via Acoustic Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Call centers face the challenge of analyzing large volumes of customer-agent conversations to improve service quality, as human analysis is unfeasible and existing speech recognition systems are unreliable, especially with accents and poor acoustic conditions.
Innovation Solution
A system and method for analyzing conversations by determining speaker turns, calculating per-turn statistics such as floor-transfer times, within-floor speech ratios, and pause durations, and applying a statistical model like hidden Markov modeling to identify conversation phases without relying on speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition system is used to analyze conversations, then words can be recognized from spoken content, but reliability decreases with accents and poor acoustic conditions
Solution Approach 1:
The patent extracts and analyzes acoustic features (energy, zero-crossing rate, spectral characteristics) directly from the audio signal without converting speech to text. This bypasses the speech recognition step entirely, eliminating the reliability issues associated with accents and poor acoustic conditions while still enabling conversation quality analysis.
Solution Approach 2:
The patent replaces the speech recognition system (mechanical/algorithmic text generation process) with a direct acoustic feature analysis approach. Instead of using speech-to-text conversion, the system analyzes raw acoustic properties of the speech signal to determine conversation phases and quality metrics.
2Measurement precision
If human analysis of every conversation is performed, then detailed quality information can be obtained, but productivity becomes unachievable due to volume
Solution Approach 1:
The system enables automated self-analysis of conversations by computing acoustic features and applying statistical models (such as hidden Markov models) to automatically identify conversation phases and quality metrics without human intervention. This allows high-volume processing while maintaining detailed analysis capability.
Solution Approach 2:
The patent transforms the analysis approach by changing from text-based analysis to acoustic parameter-based analysis. By computing and analyzing acoustic features (energy, zero-crossing rate, spectral characteristics) and applying statistical models, the system achieves both detailed quality assessment and high processing throughput.
3Loss of information
If unconstrained speech recognition is used, then speech content can be converted to text, but significant computational resources are required
Solution Approach 1:
The patent extracts only the essential acoustic features (energy, zero-crossing rate, spectral characteristics) needed for conversation analysis rather than performing full speech-to-text conversion. This selective extraction significantly reduces computational resource requirements while preserving the information needed to determine conversation phases and quality.
Data Source
AI summary
Embodiments disclosed herein provide systems, methods, and computer readable media for analyzing a conversation between a plurality of participants. In a particular embodiment, a method provides determining a first speaker from the plurality of participants and determining a second speaker from the plurality of participants. The method further provides determining a first plurality of turns comprising portions of the conversation when the first speaker is speaking and determining a second plurality of turns comprising portions of the conversation when the second speaker is speaking The method also provides determining per-turn statistics for turns of the first and second pluralities of turns and identifying phases of the conversation based on the per-turn statistics.


