Speech Emotion Recognition System for Text Enrichment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems fail to adequately represent human emotional content in speech-to-text conversions, leading to a lack of emotional context in text communications, which can result in confusion and misinterpretation.
Innovation Solution
A method and system for speech emotion recognition that uses a machine learning model to extract acoustic features from speech samples, predict emotion content, and enrich text communications with visual emotion indicators such as color changes, emoticon symbols, and word stretching to convey emotional intensity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional speech recognition systems are used, then speech-to-text conversion is achieved, but emotional content is not represented
Solution Approach 1:
The system segments the speech processing task into multiple components: acoustic feature extraction, emotion prediction using ML models, and visual enrichment. This segmentation allows the system to add emotional analysis capabilities without completely redesigning the speech recognition pipeline, thereby reducing the perceived complexity while recovering emotional information.
Solution Approach 2:
The patent introduces an intermediary machine learning model that acts as a bridge between the conventional speech recognition system and the emotional content representation. This intermediary component processes acoustic features and generates emotion predictions, enabling emotional content recovery without directly modifying the core speech-to-text conversion process.
2Measurement precision
If speaker-dependent speech engines are used, then recognition accuracy is improved, but training time and user setup are increased
Solution Approach 1:
The emotion recognition system operates in a self-service manner by automatically extracting acoustic features and generating emotion predictions without requiring user training or configuration. The machine learning models are pre-trained and can directly process speech input to infer emotional state, eliminating the time-consuming training phase while maintaining accuracy.
Solution Approach 2:
The system performs preliminary actions by pre-training the machine learning models with extensive acoustic feature data before deployment. This preliminary training enables the models to accurately predict emotions from acoustic features during actual use without requiring additional training time for each user or application scenario.
3Loss of information
If emotion recognition processing is added, then emotional context is recovered, but computational load is increased
Solution Approach 1:
The system extracts only the essential acoustic features from speech signals that are most relevant for emotion recognition, rather than processing the entire speech signal in detail. This selective extraction reduces computational energy requirements while still recovering the emotional context effectively by focusing on discriminative features.
Solution Approach 2:
The patent applies partial action by implementing emotion recognition processing only for speech segments where emotional context is most valuable, rather than uniformly processing all speech data at full detail. The system strategically applies computational resources to recover emotional information where it provides maximum benefit, reducing overall energy consumption while maintaining effective emotional context recovery.
Data Source
AI summary
Systems and methods enrich speech to text communications between users in speech chat sessions using a speech emotion recognition model to convert observed emotions in speech samples to enrich text with visual emotion content. The method may include generating a data set of speech samples with labels of a plurality of emotion classes, selecting a set of acoustic features from each of the emotion classes, generating a machine learning (ML) model based on the acoustic features and data set, applying the set of rules based on the selected set of acoustic features and data set, computing a number of rules that have been satisfied, and presenting the enriched text in speech-to-text communications between users in the chat session for visual notice of an observed emotion in the speech sample.


