Emotion-Aware Caption Rendering for Spoken Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic speech recognition systems fail to incorporate sentiment or emotion detection in audio input, resulting in captions that lack visual indicators of tone, inflection, or emotion, providing an incomplete representation of the audio signal.

Innovation Solution

A speech emotion model is used to generate emotion tag data for audio signals, which is then utilized to adjust visual characteristics of captions, such as font, style, and timing, to convey emotional context alongside the text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional automatic speech recognition systems are used to generate captions, then the captioning process is simple and fast, but the captions lack visual indicators of tone, inflection, or emotion resulting in incomplete representation of the audio signal

Engineering Contradiction:
Improveemotional context informationVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent combines multiple systems into a unified captioning system: (1) ASR system for text transcription, (2) speech emotion model for emotion tag generation, and (3) visual characteristic adjustment mechanism. These components work together to produce captions that include both text and emotional context indicators, resolving the contradiction between information completeness and system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The captioning system is designed to perform multiple functions simultaneously: it transcribes speech text, detects emotional content, and adjusts visual characteristics to convey emotion. This multi-functional approach allows the system to capture both linguistic and emotional information without requiring separate independent systems, thereby managing complexity while reducing information loss.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If emotion tag data is processed and visual characteristics are adjusted to convey emotional context, then the captioning provides complete emotional representation, but the processing complexity and computational requirements increase

Engineering Contradiction:
Improveemotional context representationVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The speech emotion model processes audio signals in advance to generate emotion tag data before the captioning stage. This preliminary emotion detection allows the main captioning system to receive pre-processed emotional context, reducing the computational burden during real-time caption generation and display adjustment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system adjusts visual characteristics of captions (such as color, font style, or size) based on emotion tag data parameters. By transforming emotional information into visual parameter modifications rather than complex processing operations, the system efficiently conveys emotional context while managing processing complexity.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If visual characteristics of captions are adjusted based on emotion tag data, then the captions convey emotional context effectively, but the caption generation time and processing duration increase

Engineering Contradiction:
Improveemotional context in captionsVSAvoidcaption generation time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

Emotion tag data is generated in advance through the speech emotion model before caption text is finalized. This preliminary emotion analysis enables parallel processing of text transcription and emotion detection, reducing the overall time required to generate emotionally-enriched captions while maintaining information completeness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically adjusts visual characteristics of captions based on emotion tag data without requiring manual intervention or complex real-time analysis. The automated visual adjustment process efficiently conveys emotional context while minimizing additional processing time through rule-based or model-driven parameter modifications.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250349298A1Expressive Captions for Audio Content
Publication Date: 2025.11.13 GOOGLE LLC
  • US20250349298A1 patent drawing
  • US20250349298A1 patent drawing
  • US20250349298A1 patent drawing

AI summary

Example embodiments of the present disclosure provide for an example method including obtaining, input audio signals including vocal events. The method includes processing, by a speech emotion model, a portion of the input audio signal including one or more vocal events to generate emotion tag data for the vocal event. The method includes obtaining a caption for the vocal event. The method includes adjusting a visual characteristic of the caption based on the emotion tag data. The method includes providing the adjusted caption for display via the graphical user interface of the computing device.