Emotion Recognition Using Auditory Attention Cues
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional emotion recognition methods using acoustic signals process the entire speech signal equally, leading to limited performance and computational inefficiencies, as they fail to selectively focus on emotionally salient parts, mimicking human auditory attention mechanisms.
Innovation Solution
The proposed method employs a biologically inspired auditory attention model that extracts multi-scale features and uses a bottom-up saliency-driven approach to detect salient events in speech, combining these with traditional prosodic features for emotion recognition, thereby selectively processing only emotionally relevant segments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional methods process the entire speech signal equally, then all speech information is considered, but computation time increases and recognition accuracy is limited
Solution Approach 1:
The patent segments the speech signal into multiple scales (different time and frequency resolutions) and processes each scale separately to identify salient events. This segmentation allows the system to focus computational resources on emotionally relevant segments rather than processing the entire signal uniformly, thereby improving recognition accuracy while reducing overall computation time.
Solution Approach 2:
The patent extracts salient events from the speech signal based on auditory attention features, separating emotionally significant portions from non-salient portions. By taking out only the relevant segments for detailed emotion analysis, the system achieves higher accuracy without the computational cost of processing the complete signal.
2Reliability
If the entire speech signal is processed, then comprehensive emotion information is captured, but computational efficiency decreases
Solution Approach 1:
The patent performs preliminary processing at multiple scales to identify salient events before conducting detailed emotion recognition. This preliminary action filters the speech signal to highlight emotionally relevant segments, ensuring reliable emotion detection while improving computational efficiency by avoiding detailed analysis of non-salient portions.
Solution Approach 2:
The patent introduces a multi-scale dimension to the speech signal processing, analyzing the signal at different time and frequency resolutions. This dimensional approach allows the system to capture comprehensive emotion information reliably while maintaining computational efficiency by leveraging the complementary information across scales.
3Measurement precision
If traditional prosodic features are used alone, then processing is simple, but emotion recognition performance is limited
Solution Approach 1:
The patent combines traditional prosodic features with auditory attention features extracted at multiple scales to create a composite feature set for emotion recognition. This composite approach leverages the strengths of both feature types, achieving superior recognition accuracy while the multi-scale framework helps manage the complexity through systematic organization.
Solution Approach 2:
The patent creates a multi-functional feature extraction system that simultaneously extracts traditional prosodic features and auditory attention features at multiple scales. This universal approach allows a single system to handle diverse emotion recognition tasks with high accuracy, making the increased complexity worthwhile through its broad applicability and performance.
Data Source
AI summary
Emotion recognition may be implemented on an input window of sound. One or more auditory attention features may be extracted from an auditory spectrum for the window using one or more two-dimensional spectro-temporal receptive filters. One or more feature maps corresponding to the one or more auditory attention features may be generated. Auditory gist features may be extracted from feature maps, and the auditory gist features may be analyzed to determine one or more emotion classes corresponding to the input window of sound. In addition, a bottom-up auditory attention model may be used to select emotionally salient parts of speech and execute emotion recognition only on the salient parts of speech while ignoring the rest of the speech signal.


