Supplemental Audio Generation for Video Content Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Certain portions of audio-visual content are not suitable for playback in audio-only mode, leading to user inconvenience, increased processing power consumption, and bandwidth usage when users switch back to video mode to understand missing context.

Innovation Solution

A media application generates supplemental audio for content items that are not suitable for audio-only mode, using sensors to detect device orientation and user activity, and accessing metadata from multiple sources to provide text-to-speech and additional context, while also skipping or summarizing content to conserve bandwidth and processing resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If video mode is used to understand content context, then user understanding is improved, but bandwidth consumption and processing power increase

Engineering Contradiction:
Improvecontent context understandingVSAvoidbandwidth and processing power
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The patent extracts only the essential visual information (text and key visual elements) from video content and converts it to audio format through text-to-speech synthesis. This allows users to obtain content context without consuming video bandwidth or processing resources, as only audio data is transmitted and processed.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces text-to-speech synthesis as an intermediary that converts visual text content into audio format. This mediator bridges the gap between video content and audio-only mode, enabling users to understand content context through audio without needing to switch to resource-intensive video mode.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of energy

If audio-only mode is used, then bandwidth and processing power are saved, but user understanding of content context deteriorates

Engineering Contradiction:
Improvebandwidth and processing powerVSAvoidcontent context
Core Design Contradiction:
Loss of energyVSLoss of information

Solution Approach 1:

The patent performs preliminary text-to-speech conversion of essential visual content (such as text overlays, subtitles, and key information) before audio-only playback. This preliminary action ensures that all necessary contextual information is already in audio format, allowing users to consume content fully in audio-only mode without losing understanding.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If video mode is switched on during audio-only playback, then content understanding is improved, but device power consumption increases

Engineering Contradiction:
Improvecontent contextVSAvoiddevice power consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent replaces the mechanical action of switching display screens on with audio-based information delivery. Instead of activating the visual display system (which consumes significant power), the system uses text-to-speech audio output to convey the same information, thereby substituting a high-power mechanical system with a lower-power audio system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If text-to-speech generation is performed, then accessibility and content understanding are improved, but processing time and computational resources increase

Engineering Contradiction:
Improveaudio-only mode compatibilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies partial text-to-speech generation by selecting only the most essential visual elements (such as text overlays, subtitles, and key information) for conversion to audio, rather than converting all visual content. This selective approach maintains audio-only mode compatibility while significantly reducing processing time and computational resource requirements.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240380941A1Supplemental audio generation system in an audio-only mode
Publication Date: 2024.11.14 ADEIA GUIDES INC
  • US20240380941A1 patent drawing
  • US20240380941A1 patent drawing
  • US20240380941A1 patent drawing

AI summary

Systems and methods for generating supplemental audio for an audio-only mode are disclosed. For example, a system generates for output a content item that includes video and audio. In response to determining that an audio-only mode is activated, the system determines that a portion of the content item is not suitable to play in the audio-only mode. In response to determining that the portion of the content item is not suitable to play in the audio-only mode, the system generates for output supplemental audio associated with the content item during the portion of the content item.