AI Video Scene Describer for Visual-to-Audio Accessibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video content is often not accessible to individuals with physical impairments, particularly visual impairments, as they cannot detect or fully appreciate on-screen text and unvoiced events.
Innovation Solution
A video scene describer system that utilizes artificial intelligence and algorithms to analyze video content, generating audio description data by combining video insights and embedding data through a visual-language model and a large language model to provide immersive textual or audio descriptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If video content is provided in standard visual format, then the content can be easily displayed and consumed by able-bodied viewers, but individuals with visual impairments cannot access or appreciate the content
Solution Approach 1:
The patent introduces an intermediary system (video scene describer with AI models) that converts visual video content into audio descriptions. This mediator translates visual elements like on-screen text, scenes, and actions into spoken language, enabling visually impaired individuals to access content that would otherwise be inaccessible through standard visual formats alone.
Solution Approach 2:
The system changes the parameter of information delivery from visual to auditory. By transforming the modality of content delivery from sight-based to sound-based, the system enables access for visually impaired users while preserving the core information content of the original video.
2Loss of information
If AI models and algorithms are used to generate audio descriptions, then comprehensive video content analysis and description can be provided, but the processing time and computational resources increase
Solution Approach 1:
The patent segments the video processing task into multiple specialized AI components: a video indexer for initial analysis, a visual-language model for embedding generation, and a large language model for description generation. This segmentation allows each component to focus on specific aspects of video understanding, improving overall efficiency and completeness of the description generation process.
Solution Approach 2:
The video indexer performs preliminary analysis and extraction of key visual elements before the main description generation process. By pre-processing the video content to identify important scenes, objects, and actions, the system reduces the computational burden on subsequent models and accelerates the overall processing time while maintaining description completeness.
Data Source
AI summary
Examples of the present disclosure describe a video scene describer. The video scene describer receives video content data as input and provides audio description (AD) data as output. The video scene describer utilizes one or more components using artificial intelligence (AI) and/or algorithms to analyze and describe the video content data. For example, the video scene describer may include a video indexer component to identify and describe particular aspects of the video content data and generates video insights data based on the analysis. The video indexer provides video insights data to a large language model (LLM) component. The video scene describer may additionally include a visual-language model system, which includes a visual encoder, a relation aggregator, a transformer encoder, and/or transformer. The LLM component synthesizes the video insights data and video embedding data, along with any prompt (e.g., a request or question) or dialogue context, to provide the AD data.


