Emotion Detection from Audio Interruption Moments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional speech emotion recognition systems face low accuracy due to extreme vocal feature variability among individuals and inconsistencies in audio recording devices, requiring large datasets for deep neural network training.

Innovation Solution

A system that detects moments of interruption in dual-channel audio files using deep neural networks to extract vocal features like interruption length and speaking rates, and determines emotion types using machine learning models, reducing the need for extensive training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional SER systems use single-channel audio files with conventional machine learning models, then device complexity is reduced, but measurement precision (emotion detection accuracy) deteriorates due to extreme vocal feature variability and channel inconsistencies

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio analysis task by focusing on specific moments of interruption rather than analyzing the entire audio file. It divides the continuous speech stream into discrete interruption events, extracting features only from these critical segments. This segmentation approach improves accuracy by concentrating on emotionally salient moments while reducing the effective data processing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by analyzing specific local characteristics of interruption moments (such as voice energy, speaking rate, and duration) rather than using global statistics from the entire audio file. By focusing on the local properties of these critical moments, the system achieves higher emotion detection accuracy without requiring complex analysis of all audio data.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If deep neural network approaches are used to capture vocal feature variabilities, then measurement precision improves, but quantity of substance (training data requirements) increases enormously

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidtraining data quantity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the most informative features from interruption moments (voice energy, speaking rate, speaking duration, voice activity ratio) rather than using all possible acoustic features. This selective extraction reduces the dimensionality of the feature space, allowing machine learning models to achieve high accuracy with significantly reduced training data requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by analyzing only a subset of audio data (interruption moments) rather than the complete audio file. By focusing on these partial, emotionally salient moments, the system achieves sufficient training data for accurate emotion recognition without requiring enormous datasets that would be needed for complete audio analysis.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If dual-channel audio analysis is performed to detect moments of interruption, then measurement precision improves, but processing time increases

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the audio analysis into discrete interruption detection tasks rather than continuous processing. By identifying and isolating specific interruption moments, the system can process only these critical segments in detail, reducing overall processing time while maintaining high accuracy through focused analysis of emotionally salient moments.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240371399A1Systems and methods for detecting emotion from audio files
Publication Date: 2024.11.07 CAPITAL ONE SERVICES LLC
  • US20240371399A1 patent drawing
  • US20240371399A1 patent drawing
  • US20240371399A1 patent drawing

AI summary

Disclosed embodiments may include a system that may receive an audio file comprising an interaction between a first user and a second user. The system may detect, using a deep neural network (DNN), moment(s) of interruption between the first and second users from the audio file. The system may extract, using the DNN, vocal feature(s) from the moment(s) of interruption. The system may determine, using a machine learning model (MLM) and based on the vocal feature(s), whether a threshold number of moments of the moment(s) of interruption corresponds to a first emotion type. When the threshold number of moments corresponds to the first emotion type, the system may transmit a first message comprising a first binary indication. When the threshold number of moments do not correspond to the first emotion type, the system may transmit a second message comprising a second binary indication.