Emotion Detection from Audio Interruption Moments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional speech emotion recognition systems face low accuracy due to extreme vocal feature variability among individuals and inconsistencies in audio recording devices, requiring large datasets for deep neural network training.
Innovation Solution
A system that detects moments of interruption in dual-channel audio files using deep neural networks to extract vocal features like interruption length and speaking rates, and determines emotion types using machine learning models, reducing the need for extensive training data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional SER systems use single-channel audio files with conventional machine learning models, then device complexity is reduced, but measurement precision (emotion detection accuracy) deteriorates due to extreme vocal feature variability and channel inconsistencies
Solution Approach 1:
The patent segments the audio analysis task by focusing on specific moments of interruption rather than analyzing the entire audio file. It divides the continuous speech stream into discrete interruption events, extracting features only from these critical segments. This segmentation approach improves accuracy by concentrating on emotionally salient moments while reducing the effective data processing complexity.
Solution Approach 2:
The patent applies local quality by analyzing specific local characteristics of interruption moments (such as voice energy, speaking rate, and duration) rather than using global statistics from the entire audio file. By focusing on the local properties of these critical moments, the system achieves higher emotion detection accuracy without requiring complex analysis of all audio data.
2Measurement precision
If deep neural network approaches are used to capture vocal feature variabilities, then measurement precision improves, but quantity of substance (training data requirements) increases enormously
Solution Approach 1:
The patent extracts only the most informative features from interruption moments (voice energy, speaking rate, speaking duration, voice activity ratio) rather than using all possible acoustic features. This selective extraction reduces the dimensionality of the feature space, allowing machine learning models to achieve high accuracy with significantly reduced training data requirements.
Solution Approach 2:
The patent applies partial action by analyzing only a subset of audio data (interruption moments) rather than the complete audio file. By focusing on these partial, emotionally salient moments, the system achieves sufficient training data for accurate emotion recognition without requiring enormous datasets that would be needed for complete audio analysis.
3Measurement precision
If dual-channel audio analysis is performed to detect moments of interruption, then measurement precision improves, but processing time increases
Solution Approach 1:
The patent segments the audio analysis into discrete interruption detection tasks rather than continuous processing. By identifying and isolating specific interruption moments, the system can process only these critical segments in detail, reducing overall processing time while maintaining high accuracy through focused analysis of emotionally salient moments.
Data Source
AI summary
Disclosed embodiments may include a system that may receive an audio file comprising an interaction between a first user and a second user. The system may detect, using a deep neural network (DNN), moment(s) of interruption between the first and second users from the audio file. The system may extract, using the DNN, vocal feature(s) from the moment(s) of interruption. The system may determine, using a machine learning model (MLM) and based on the vocal feature(s), whether a threshold number of moments of the moment(s) of interruption corresponds to a first emotion type. When the threshold number of moments corresponds to the first emotion type, the system may transmit a first message comprising a first binary indication. When the threshold number of moments do not correspond to the first emotion type, the system may transmit a second message comprising a second binary indication.


