Emotion Classification via Pre-verbal Acoustic Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying emotions in media, such as music, are prone to errors due to reliance on subjective metadata and pre-classified music, leading to misclassification and inaccurate mood identification.
Innovation Solution
The use of pre-verbal expressions as a training set to create a classification model that autonomously identifies emotions in media, focusing on universal emotional features across cultures, with a feature extractor processing audio samples to extract relevant features for a classification engine to predict emotions and moods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If subjective metadata and pre-classified music are used for emotion identification, then the system is simple to implement, but the accuracy of emotion identification deteriorates
Solution Approach 1:
The patent replaces the mechanical/manual system of subjective metadata classification with an acoustic field-based automatic classification system. The system uses acoustic feature extraction and machine learning algorithms to objectively identify emotions in media, eliminating reliance on human-subjective metadata while maintaining implementation feasibility through automated processing pipelines.
Solution Approach 2:
The patent introduces acoustic feature extraction and classification algorithms as intermediary components between the raw audio signal and emotion identification. These intermediaries process the acoustic signal to extract objective features (spectral, temporal, statistical) that serve as reliable indicators of emotional content, bridging the gap between simple implementation and accurate identification.
2Ease of operation
If subjective measures are used to generate metadata, then the system is easier to operate, but the reliability of emotion identification deteriorates
Solution Approach 1:
The patent enables the system to self-service by automatically generating emotion classifications without requiring human operators to manually create or curate metadata. The automated classification engine processes media content independently, extracting acoustic features and determining emotions through algorithmic analysis, thereby eliminating operational complexity while improving reliability through consistent objective measurement.
Solution Approach 2:
The patent implements feedback mechanisms where the classification model is trained on labeled data and continuously improves its accuracy. The system uses feedback from training examples to refine its acoustic feature extraction and classification algorithms, ensuring reliable emotion identification while maintaining ease of operation through automated iterative improvement.
3Device complexity
If pre-classified music is used for training, then the device complexity is reduced, but the measurement precision of emotion identification worsens
Solution Approach 1:
The patent applies preliminary action by pre-training the classification model on a diverse dataset of labeled media with known emotional characteristics. This preliminary training establishes a robust foundation of acoustic feature-emotion relationships that enables accurate identification in subsequent applications without requiring complex real-time adjustments or human intervention.
Solution Approach 2:
The patent utilizes parameter changes by adjusting the acoustic feature extraction parameters and classification model parameters during training and deployment. By optimizing parameters such as spectral resolution, temporal windowing, and feature weighting, the system achieves high measurement precision while maintaining manageable device complexity through parameter-driven adaptability rather than structural complexity.
Data Source
AI summary
Methods and apparatus to identify an emotion evoked by media are disclosed. An example apparatus includes a synthesizer to generate a first synthesized sample based on a pre-verbal utterance associated with a first emotion. A feature extractor is to identify a first value of a first feature of the first synthesized sample. The feature extractor to identify a second value of the first feature of first media evoking an unknown emotion. A classification engine is to create a model based on the first feature. The model is to establish a relationship between the first value of the first feature and the first emotion. The classification engine is to identify the first media as evoking the first emotion when the model indicates that the second value corresponds to the first value.


