AI-Generated Media Playback Enrichment for Live Accessibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for providing additional information in media playback, such as subtitles and audio descriptions, are inadequate for real-time media offerings and often result in an inadequate user experience, particularly for live events, and lack sufficient detail in conventional film subtitles.
Innovation Solution
A method and system that dynamically generates quasi-real-time additional information by separating audio and video tracks, using artificial speech recognition and generative AI to convert speech into text and objects into descriptions, and synchronizes this information with the media playback, allowing users to customize the level of detail and type of information based on their needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If subtitles and audio description are generated in advance and added to video data, then accessibility is provided for users with disabilities, but real-time media offerings such as live sporting events cannot be made accessible
Solution Approach 1:
The system dynamically generates subtitles and audio descriptions in real-time during media playback rather than using pre-generated tracks. The additional information generation algorithm processes the original data stream as it plays, separating audio and video tracks, generating supplementary information dynamically, and synchronizing it with the playback timing, enabling accessibility for live events and dynamic content
Solution Approach 2:
An additional information generation algorithm acts as an intermediary between the original media data stream and the user. This algorithm receives the original audio and video tracks, generates supplementary information (subtitles, audio descriptions, contextual information), and outputs an enriched data stream that provides accessibility without requiring pre-processing of the original content
2Reliability
If 1:1 reproduction of spoken word is provided as subtitles, then accessibility is provided, but user experience is inadequate due to lack of contextual information
Solution Approach 1:
The system merges the original audio track with generated audio descriptions and the original video track with generated visual descriptions and contextual information. The additional information generation algorithm combines speech-to-text conversion with contextual enrichment, merging factual transcription with interpretive descriptions of visual elements, background context, and relevant information that enhances user experience while maintaining accessibility
3Loss of information
If detailed additional information is generated during media playback, then user experience is improved, but data transmission requirements increase
Solution Approach 1:
The system segments the enriched data stream into essential and supplementary information components. The additional information generation algorithm identifies and prioritizes critical accessibility information (such as speech transcription and key visual elements) while optionally providing additional contextual details, allowing flexible data transmission based on user needs and bandwidth availability
4Reliability
If additional information is inserted into original data stream, then accessibility is enhanced, but synchronization issues may occur
Solution Approach 1:
The system implements synchronization mechanisms where the additional information generation algorithm continuously monitors the playback timing and adjusts the generation and insertion of subtitles and audio descriptions accordingly. This feedback loop ensures that the enriched data stream remains synchronized with the original media content, preventing lag or desynchronization issues
Data Source
Figure 1

AI summary
The present invention relates to techniques for dynamically providing additional information during media playback, comprising the following steps: • Receiving an initial data stream of media playback, wherein the media playback comprises an initial video track and/or an initial audio track; • Passing the data stream to a processor unit, wherein an additional information generation algorithm implemented on the processor unit performs the following steps in combination or individually: • With regard to the audio track: Converting the speech of the audio signal of the audio track into a basic text output and generating audio track additional information based on the basic text output and/or the audio signal;and/or ∘ Regarding the video track: Performing pattern recognition of images, whereby objects and/or object movements are recognized from the images and their word descriptions are extracted, wherein the objects and/or object movements are passed as input to a generative AI with the task of generating video track supplementary information; ∘ Creating an enriched media playback data stream, wherein the enriched data stream includes the audio track supplementary information and/or the video track supplementary information, wherein the enriched data stream is displayable on a user's terminal device.;