AI Video Scene Describer for Visual-to-Audio Accessibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video content is often not accessible to individuals with physical impairments, particularly visual impairments, as they cannot detect or fully appreciate on-screen text and unvoiced events.

Innovation Solution

A video scene describer system that utilizes artificial intelligence and algorithms to analyze video content, generating audio description data by combining video insights and embedding data through a visual-language model and a large language model to provide immersive textual or audio descriptions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If video content is provided in standard visual format, then the content can be easily displayed and consumed by able-bodied viewers, but individuals with visual impairments cannot access or appreciate the content

Engineering Contradiction:
ImproveAccessibility to visually impaired individualsVSAvoidLoss of visual information for impaired users
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary system (video scene describer with AI models) that converts visual video content into audio descriptions. This mediator translates visual elements like on-screen text, scenes, and actions into spoken language, enabling visually impaired individuals to access content that would otherwise be inaccessible through standard visual formats alone.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter of information delivery from visual to auditory. By transforming the modality of content delivery from sight-based to sound-based, the system enables access for visually impaired users while preserving the core information content of the original video.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If AI models and algorithms are used to generate audio descriptions, then comprehensive video content analysis and description can be provided, but the processing time and computational resources increase

Engineering Contradiction:
ImproveCompleteness of video content descriptionVSAvoidProcessing time for generating descriptions
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent segments the video processing task into multiple specialized AI components: a video indexer for initial analysis, a visual-language model for embedding generation, and a large language model for description generation. This segmentation allows each component to focus on specific aspects of video understanding, improving overall efficiency and completeness of the description generation process.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The video indexer performs preliminary analysis and extraction of key visual elements before the main description generation process. By pre-processing the video content to identify important scenes, objects, and actions, the system reduces the computational burden on subsequent models and accelerates the overall processing time while maintaining description completeness.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250246176A1Video scene describer
Publication Date: 2025.07.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250246176A1 patent drawing
  • US20250246176A1 patent drawing
  • US20250246176A1 patent drawing

AI summary

Examples of the present disclosure describe a video scene describer. The video scene describer receives video content data as input and provides audio description (AD) data as output. The video scene describer utilizes one or more components using artificial intelligence (AI) and/or algorithms to analyze and describe the video content data. For example, the video scene describer may include a video indexer component to identify and describe particular aspects of the video content data and generates video insights data based on the analysis. The video indexer provides video insights data to a large language model (LLM) component. The video scene describer may additionally include a visual-language model system, which includes a visual encoder, a relation aggregator, a transformer encoder, and/or transformer. The LLM component synthesizes the video insights data and video embedding data, along with any prompt (e.g., a request or question) or dialogue context, to provide the AD data.