Real-time Emotion Recognition via Rapid Audio Fingerprinting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current smart devices can only understand semantic speech and fail to comprehend the richer meaning conveyed by non-semantic characteristics of speech, such as emotions, which limits their ability to provide accurate and efficient communication assistance.

Innovation Solution

Systems and methods for recognizing emotions in audio signals in real-time by computing a rapid audio fingerprint, extracting features, and comparing them with defined emotions to determine confidence scores, allowing for the association of emotions with audio signals and initiation of appropriate actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If smart devices use traditional semantic speech recognition, then they can understand basic spoken words, but they fail to comprehend non-semantic characteristics such as emotions conveyed by speech

Engineering Contradiction:
Improveemotional informationVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments audio signals into multiple feature dimensions including pitch, intensity, speech rate, and pause characteristics. Each feature is independently extracted and analyzed to capture different aspects of emotional expression, allowing the system to comprehensively understand non-semantic characteristics without overwhelming system complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that transforms raw audio signals into emotional indicators through feature extraction and analysis. This intermediary layer acts as a bridge between traditional semantic recognition and emotional understanding, enabling the system to comprehend both literal meaning and emotional context

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the system performs comprehensive audio analysis to detect emotions, then accuracy of emotional recognition is improved, but processing time increases

Engineering Contradiction:
Improveemotional detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary feature extraction during audio signal processing, pre-computing pitch, intensity, speech rate, and pause characteristics before emotional analysis. This preliminary action prepares the data in advance, enabling faster and more accurate emotional detection without excessive processing delays

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements periodic analysis of audio features at strategically selected intervals rather than continuous analysis. By analyzing speech at key moments and using representative samples, the system maintains high emotional detection accuracy while significantly reducing overall processing time

Inventive Principle:
Principle #19Periodic action

3Productivity

If the system continuously monitors audio signals for real-time emotion detection, then real-time understanding is achieved, but energy consumption increases

Engineering Contradiction:
Improvereal-time processing capabilityVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent implements continuous monitoring of key audio features such as pitch and intensity that are most indicative of emotional state. By maintaining continuous tracking of these critical parameters while periodically analyzing complete feature sets, the system achieves real-time emotional understanding with optimized energy consumption

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent applies partial analysis by focusing on the most emotionally informative features rather than analyzing all audio characteristics equally. By concentrating computational resources on pitch variation, intensity changes, and speech rate - the most significant emotional indicators - the system achieves effective real-time emotion detection with reduced energy usage

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10068588B2Real-time emotion recognition from audio signals
Publication Date: 2018.09.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10068588B2 patent drawing
  • US10068588B2 patent drawing
  • US10068588B2 patent drawing

AI summary

Systems, methods, and computer-readable storage media are provided for recognizing emotion in audio signals in real-time. An audio signal is detected and a rapid audio fingerprint is computed on a user's computing device. One or more features is extracted from the audio fingerprint and compared with features associated with defined emotions to determine relative degrees of similarity. Confidence scores are computed for the defined emotions based on the relative degrees of similarity and it is determined whether a confidence score for one or more particular emotions exceeds a threshold confidence score. If it is determined that a threshold confidence score for one or more particular emotions is exceeded, the particular emotion or emotions are associated with the audio signal. As desired, various action then may be initiated based upon the emotion/emotions associated with the audio signal.