Composite Word Context Modeling for Call Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for processing and classifying human-to-human telephone conversations are inefficient and inaccurate due to reliance on full transcription, phonetic transcription, and word spotting approaches, which are computationally expensive and fail to capture the importance and relevance of audio files in classification and search results.
Innovation Solution
A system that integrates recognition and interpretation stages by translating user input into semantically equivalent events, using word spotting and context modeling to classify communications, and assigning confidence levels for accurate and efficient processing of multimedia files, including human-to-human conversations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If full transcription approach is used to convert speech to text, then complete text representation is achieved, but computational cost and processing time increase significantly
Solution Approach 1:
The patent extracts only the essential information needed for classification (specific words and phrases) from the audio stream, rather than transcribing the entire conversation. This selective extraction approach maintains classification accuracy while dramatically reducing processing time and computational resources required.
Solution Approach 2:
Instead of performing complete transcription (excessive action), the system performs partial transcription by identifying and processing only the key words and phrases necessary for call classification. This partial action approach achieves the classification goal with significantly reduced computational overhead.
2Productivity
If phonetic transcription is used to convert audio to phone sequences, then processing speed improves, but transcription accuracy and contextual understanding deteriorate
Solution Approach 1:
The system applies different processing qualities to different parts of the audio signal. Key words and phrases receive detailed analysis with contextual modeling, while other portions use simpler phonetic matching. This local quality differentiation maintains accuracy for important terms while preserving overall processing efficiency.
3Productivity
If word spotting is used to recognize specific words, then processing efficiency improves, but ability to handle arbitrary search queries deteriorates
Solution Approach 1:
The classification system is designed with universal applicability to handle both predefined classification categories and arbitrary user search queries. The context modeling and confidence scoring mechanisms work consistently across different query types, making the system versatile while maintaining efficient processing through selective word spotting techniques.
4Loss of information
If traditional speech recognition is used to transcribe all words, then complete text output is generated, but computational expense and complexity increase
Solution Approach 1:
The system extracts only the classification-relevant information from the audio stream, identifying key words and phrases that indicate call category rather than transcribing the entire conversation. This extraction approach reduces system complexity and computational expense while maintaining the ability to accurately classify calls.
5Measurement precision
If manual review is used to analyze conversations, then classification accuracy is maintained, but cost and efficiency deteriorate
Solution Approach 1:
The system performs automatic classification without requiring manual review, using context modeling and confidence scoring to independently determine call categories. This self-service capability maintains high accuracy while dramatically improving processing efficiency and reducing costs associated with manual analysis.
Data Source
AI summary
A system and method of automatically classifying a communication involving at least one human, e.g., a human-to-human telephone conversation, into predefined categories of interest, e.g., “angry customer, etc. The system automatically, or semi-automatically with user interaction, expanding user input into semantically equivalent events. The system recognizes these events, each of which is given a confidence level, and classifies the telephone conversation based on an overall confidence level. In some embodiments, a different base unit is utilized. Instead of recognizing individual words, the system recognizes composite words, each of which is pre-programmed as an atomic unit, in a given context. The recognition includes semantically relevant composite words and contexts automatically generated by the system. The composite words based contextual recognition technique enables the system to efficiently and logically classifying and indexing large volumes of communications and audio collections such as call center calls, Webinars, live news feeds, etc.


