Emotion-Aware Caption Rendering for Spoken Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic speech recognition systems fail to incorporate sentiment or emotion detection in audio input, resulting in captions that lack visual indicators of tone, inflection, or emotion, providing an incomplete representation of the audio signal.
Innovation Solution
A speech emotion model is used to generate emotion tag data for audio signals, which is then utilized to adjust visual characteristics of captions, such as font, style, and timing, to convey emotional context alongside the text.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional automatic speech recognition systems are used to generate captions, then the captioning process is simple and fast, but the captions lack visual indicators of tone, inflection, or emotion resulting in incomplete representation of the audio signal
Solution Approach 1:
The patent combines multiple systems into a unified captioning system: (1) ASR system for text transcription, (2) speech emotion model for emotion tag generation, and (3) visual characteristic adjustment mechanism. These components work together to produce captions that include both text and emotional context indicators, resolving the contradiction between information completeness and system complexity.
Solution Approach 2:
The captioning system is designed to perform multiple functions simultaneously: it transcribes speech text, detects emotional content, and adjusts visual characteristics to convey emotion. This multi-functional approach allows the system to capture both linguistic and emotional information without requiring separate independent systems, thereby managing complexity while reducing information loss.
2Loss of information
If emotion tag data is processed and visual characteristics are adjusted to convey emotional context, then the captioning provides complete emotional representation, but the processing complexity and computational requirements increase
Solution Approach 1:
The speech emotion model processes audio signals in advance to generate emotion tag data before the captioning stage. This preliminary emotion detection allows the main captioning system to receive pre-processed emotional context, reducing the computational burden during real-time caption generation and display adjustment.
Solution Approach 2:
The system adjusts visual characteristics of captions (such as color, font style, or size) based on emotion tag data parameters. By transforming emotional information into visual parameter modifications rather than complex processing operations, the system efficiently conveys emotional context while managing processing complexity.
3Loss of information
If visual characteristics of captions are adjusted based on emotion tag data, then the captions convey emotional context effectively, but the caption generation time and processing duration increase
Solution Approach 1:
Emotion tag data is generated in advance through the speech emotion model before caption text is finalized. This preliminary emotion analysis enables parallel processing of text transcription and emotion detection, reducing the overall time required to generate emotionally-enriched captions while maintaining information completeness.
Solution Approach 2:
The system automatically adjusts visual characteristics of captions based on emotion tag data without requiring manual intervention or complex real-time analysis. The automated visual adjustment process efficiently conveys emotional context while minimizing additional processing time through rule-based or model-driven parameter modifications.
Data Source
AI summary
Example embodiments of the present disclosure provide for an example method including obtaining, input audio signals including vocal events. The method includes processing, by a speech emotion model, a portion of the input audio signal including one or more vocal events to generate emotion tag data for the vocal event. The method includes obtaining a caption for the vocal event. The method includes adjusting a visual characteristic of the caption based on the emotion tag data. The method includes providing the adjusted caption for display via the graphical user interface of the computing device.


