Emotion Prediction System Using Multimodal Keystroke and Text Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current IVR and chat-based systems lack the ability to intelligently route customer interactions based on emotional state, leading to decreased customer satisfaction due to inefficient call or chat handling, as they do not effectively utilize emotion prediction systems that consider multiple inputs such as speech, text, keystrokes, and facial expressions.
Innovation Solution
The development of systems and methods that generate an EmotionPrint from multimodal inputs, including speech, text, and keystrokes, using machine learning models to predict customer emotions in real-time, allowing for personalized interaction routing and performance evaluation of agents and bots.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If rules-based IVR or chat routing is used, then system complexity is reduced and ease of operation is improved, but customer satisfaction deteriorates due to lack of personalization and longer wait times
Solution Approach 1:
The system changes the routing parameters from simple menu selections to include emotion detection results. By incorporating emotion parameters (detected from speech, text, keystrokes, and facial expressions) into the routing decision process, the system maintains operational simplicity while significantly improving customer satisfaction through personalized routing to agents best suited for the customer's emotional state.
Solution Approach 2:
The IVR/chat system is enhanced with multi-functionality by integrating emotion detection capabilities across multiple input modalities (speech, text, keystrokes, facial expressions). This allows the same system to perform both traditional routing functions and emotion-based personalized routing, improving customer satisfaction without requiring completely separate systems.
2Reliability
If emotion prediction based on multimodal input is implemented, then customer satisfaction is improved through personalized routing, but device complexity increases
Solution Approach 1:
The emotion detection system is segmented into separate modules that process different input modalities independently (speech processing module, text processing module, keystroke analysis module, facial expression analysis module). Each module extracts emotion-relevant features from its specific input type, and the results are integrated to form a comprehensive emotion profile. This segmentation reduces overall system complexity by allowing independent development and optimization of each processing component.
Solution Approach 2:
An intermediary emotion analysis layer is introduced between the multi-modal input collection and the routing decision. This intermediary layer aggregates and synthesizes emotion information from various sources (speech, text, keystrokes, facial expressions) and transforms it into actionable routing parameters. The intermediary simplifies the complexity by providing a unified interface between diverse input modalities and the routing system.
3Measurement precision
If real-time emotion analysis is performed on multiple input modalities, then routing accuracy is improved, but processing time and loss of time increase
Solution Approach 1:
The system performs preliminary action by pre-processing and extracting emotion-relevant features from input modalities as they arrive, rather than waiting for complete data collection. Emotion indicators are detected and processed in real-time streams, allowing routing decisions to be made based on emerging emotion patterns without requiring analysis of entire conversation histories, thus reducing processing time while maintaining accuracy.
Solution Approach 2:
The emotion detection and analysis operates continuously across all input modalities simultaneously rather than sequentially. Multiple streams of emotion information (from speech, text, keystrokes, and facial expressions) are processed in parallel and continuously updated, enabling real-time routing decisions without time loss from sequential processing. This continuous parallel action maintains high routing accuracy while minimizing processing delays.
Data Source
AI summary
Systems, apparatuses, methods, and computer program products are disclosed for predicting an emotion in real-time based on a multimodal input including at least (i) an amount of keystrokes over a period of time and (ii) text. An example method may include receiving, by a communications circuitry, a multimodal input from a user including at least (i) an amount of keystrokes over a period of time and (ii) text. The example method may further include generating, by a trained machine learning model of an emotion prediction circuitry and using the multimodal input, an EmotionPrint for the user. The example method may finally include determining, by the emotion prediction circuitry and using the EmotionPrint, a next action.


