Real-time Emotion Tracking via Audio Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing emotion recognition systems in communication systems require a significant amount of audio content, such as a whole phone call, to analyze and classify emotional states, limiting their ability to provide real-time feedback during conversations between customers and company representatives.

Innovation Solution

A device and method for real-time emotion tracking in audio signals, which segments the audio into sequential segments, analyzes each segment for emotional states and confidence scores using lexical and acoustic analyses, and provides user-detectable notifications and actionable conduct based on detected emotional state changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the system analyzes a whole phone call to classify emotional states, then the emotion recognition accuracy is improved, but the real-time feedback capability deteriorates

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidreal-time feedback capability
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the audio signal into sequential segments and processes each segment independently to determine emotional states. This segmentation allows the system to provide real-time feedback on emotional states during ongoing interactions while maintaining reasonable recognition accuracy through continuous analysis of multiple segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis on each audio segment as it arrives, rather than waiting for the complete phone call. This preliminary action enables early detection of emotional states and provides timely feedback to company representatives during the interaction, improving real-time capability while maintaining accuracy through subsequent segment analysis.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If the system processes audio segments sequentially to detect emotional state changes, then the real-time feedback is improved, but the analysis completeness deteriorates

Engineering Contradiction:
Improvereal-time feedbackVSAvoidanalysis completeness
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The system continuously analyzes each incoming audio segment and updates the current emotional state determination based on the most recent segment analysis. This continuous action ensures that the system maintains real-time feedback capability while progressively building a complete understanding of the emotional trajectory throughout the phone call.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system provides feedback notifications to company representatives when emotional state changes are detected in the audio segments. This feedback mechanism ensures that important emotional transitions are communicated in real-time while the system continues to analyze subsequent segments to maintain complete analysis of the overall interaction.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9355650B2Real-time emotion tracking system
Publication Date: 2016.05.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9355650B2 patent drawing
  • US9355650B2 patent drawing
  • US9355650B2 patent drawing

AI summary

Devices, systems, methods, media, and programs for detecting an emotional state change in an audio signal are provided. A plurality of segments of the audio signal is received, with the plurality of segments being sequential. Each segment of the plurality of segments is analyzed, and, for each segment, an emotional state and a confidence score of the emotional state are determined. The emotional state and the confidence score of each segment are sequentially analyzed, and a current emotional state of the audio signal is tracked throughout each of the plurality of segments. For each segment, it is determined whether the current emotional state of the audio signal changes to another emotional state based on the emotional state and the confidence score of the segment.