Automated Call Analysis System Using Multi-Server Emotion and Sentiment Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing calls in call centers are inefficient due to the manual evaluation of only a small percentage of recorded calls, which incurs high costs in time and human resources, and existing transcription software struggles to accurately separate participant voices and preserve non-speech information like emotions and sentiments.
Innovation Solution
A system that sends call recording data to multiple servers to analyze voice and text data, determining emotion and sentiment scores, and generating quality scores and classification data for visualization, enabling automated and comprehensive call analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual evaluation of recorded calls is used, then analysis accuracy can be maintained, but time consumption and human resource costs increase significantly
Solution Approach 1:
The system segments the call analysis process into multiple independent components: voice transcription, emotion recognition, sentiment analysis, and quality scoring. Each component is handled by specialized servers that process specific aspects of the call, enabling parallel processing and significantly reducing overall analysis time while maintaining comprehensive evaluation accuracy.
Solution Approach 2:
The patent introduces an intermediary automated analysis system that acts as a bridge between raw call recordings and final quality assessments. This intermediary system uses AI-driven transcription and emotion recognition servers to pre-process and evaluate calls, reducing the burden on manual reviewers and enabling them to focus on complex cases that require human judgment.
2Loss of energy
If only a small percentage of calls are randomly selected for analysis, then resource costs are reduced, but the representativeness and reliability of analysis results deteriorate
Solution Approach 1:
The system implements self-service automated analysis that can independently evaluate all recorded calls without requiring manual selection. The AI-driven transcription and emotion recognition servers automatically process calls, generate sentiment scores, and identify quality issues, enabling comprehensive analysis of 100% of calls rather than random samples, thereby improving result reliability while controlling resource consumption through efficient automated processing.
3Ease of manufacture
If existing transcription software is used, then call conversion to text is achieved, but the ability to separate participant voices and preserve non-speech information is limited
Solution Approach 1:
The patent replaces traditional mechanical transcription methods with AI-driven voice recognition and emotion recognition technology. The system uses advanced algorithms to not only transcribe speech to text but also to identify and separate different participant voices, detect emotional tones, and preserve non-speech information such as laughter, sighs, and other vocal cues that carry sentiment information.
Solution Approach 2:
The transcription system is designed with multi-functionality, simultaneously performing speech-to-text conversion, speaker diarization (separating participant voices), emotion recognition, and sentiment analysis. This universal system handles multiple analysis tasks in one integrated process, preventing information loss by capturing both speech content and non-speech emotional cues.
Data Source
AI summary
Methods and systems include sending recording data of a call to a first server and a second server, wherein the recording data includes a first voice of a first participant of the call and a second voice of a second participant of the call; receiving, from the first server, a first emotion score representing a degree of a first emotion associated with the first voice, and a second emotion score representing a degree of a second emotion associated with the first voice; receiving, from the second server, a first sentiment score, a second sentiment score, and a third sentiment score; determining a quality score and classification data for the recording data based on the first emotion score, the second emotion score, the first sentiment score, the second sentiment score, and the third sentiment score; and outputting the quality score and the classification data for visualization of the recording data.


