Call Classification Without Transcription
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing and classifying voice calls in businesses, such as call recording and transcription, face challenges like high costs, privacy concerns, and inaccuracies due to poor audio quality, necessitating a more efficient and automated approach.
Innovation Solution
A system that analyzes audio from telephone calls without transcription by identifying speech patterns, measuring network latency, and generating a clustered-frame representation to classify calls, allowing for automated assessment of call outcomes and advertising effectiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If call recording and transcription are used to analyze voice transactions, then businesses can gain insight from customer conversations, but the costs increase and privacy concerns arise
Solution Approach 1:
The patent extracts only the essential features needed for call analysis (audio signal characteristics, speech patterns, key phrases) without requiring complete call recording. This selective extraction approach reduces storage requirements and privacy concerns while maintaining analytical capability
Solution Approach 2:
The analysis system segments calls into manageable units (individual calls, call patterns, aggregated metrics) and processes them through different analytical layers, allowing insight generation without maintaining complete recorded transcripts of all calls
2Productivity
If every call is recorded for analysis, then call outcomes can be evaluated, but the costs and storage requirements increase significantly
Solution Approach 1:
The system applies partial action by analyzing only the necessary portions of calls (key speech segments, critical metrics) rather than processing every call in its entirety, reducing data storage needs while maintaining evaluation capability
Solution Approach 2:
Instead of storing original call recordings, the system creates and analyzes simplified representations (transcripts, metadata, aggregated statistics) that capture essential information with minimal storage requirements
3Loss of information
If transcription is used to identify key words and phrases, then caller intent and needs can be determined, but transcription accuracy decreases with poor audio quality
Solution Approach 1:
The system changes the parameters of analysis by focusing on acoustic features, speech patterns, and audio characteristics rather than relying solely on text transcription accuracy. This allows intent detection through phonetic and prosodic analysis even when transcription is imperfect
Solution Approach 2:
The patent introduces intermediary analysis layers (audio feature extraction, speech pattern recognition) that bridge the gap between raw audio and text interpretation, enabling accurate intent detection without requiring perfect transcription
4Loss of time
If live monitoring is implemented to analyze call outcomes, then real-time insights can be obtained, but additional costs and operational complexity increase
Solution Approach 1:
The system implements self-service analysis where calls automatically undergo processing and analysis without requiring manual monitoring intervention. Automated algorithms extract insights and generate reports, reducing operational complexity while maintaining real-time capability
Data Source
AI summary
A facility and method for analyzing and classifying calls without transcription. The facility analyzes individual frames of an audio to identify speech and measure the amount of time spent in speech for each channel (e.g., caller channel, agent channel). Additional telephony metrics such as R-factor or MOS score and other metadata may be factored in as audio analysis inputs. The facility then analyzes the frames together as a whole and formulates a clustered-frame representation of a conversation to further identify dialog patterns and characterize call classification. Based on the data in the clustered-frame representation, the facility is able to make estimations of call classification. The correlation of dialog patterns to call classification may be utilized to develop targeted solutions for call classification issues, target certain advertising channels over others, evaluate advertising placements at scale, score callers, and to identify spammers.


