Descriptive Audio Indexing for Video Scene Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for retrieving video content using automatic speech recognition are inadequate, as they rely on conversational audio and may not effectively index scenes without specific keywords, limiting the ability to retrieve specific events or actions.
Innovation Solution
A system that converts descriptive audio streams from digital videos into text, aligns the text with the video, and creates an index for searching, utilizing the descriptive audio to enable retrieval of specific scenes and actions, such as 'a man standing' without requiring frame annotation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If automatic speech recognition is used on conversational audio to retrieve video content, then the system can identify mentioned subjects, but it fails to retrieve scenes without specific keywords in dialogue
Solution Approach 1:
The patent segments the video content processing into two distinct indexing approaches: traditional dialogue-based ASR indexing and the new descriptive audio indexing. By separating these functions, the system can use descriptive audio specifically for action and scene description without interfering with dialogue processing, thereby improving both retrieval accuracy for actions and overall search coverage.
Solution Approach 2:
The patent introduces descriptive audio as an intermediary layer between the video content and the search index. This descriptive audio track serves as a mediator that translates visual actions and scenes into textual descriptions, enabling the ASR system to index and retrieve content based on visual events rather than relying solely on dialogue mentions.
2Measurement precision
If frame-by-frame annotation is performed to enable action-based retrieval, then retrieval precision for actions improves, but system complexity and processing time increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-generating descriptive audio tracks during video production or post-production, before the indexing process. These descriptive audio tracks contain pre-processed information about actions and scenes in human language, eliminating the need for complex frame-by-frame analysis during indexing. The ASR system then simply transcribes this pre-existing descriptive audio, dramatically reducing processing complexity while maintaining high retrieval precision.
Solution Approach 2:
The patent replaces the mechanical frame-by-frame visual analysis system with an acoustic-based ASR system that processes descriptive audio. Instead of computationally intensive computer vision algorithms analyzing each video frame, the system uses audio processing and speech recognition to extract action information, significantly reducing device complexity and processing requirements.
3Measurement precision
If descriptive audio streams are converted to text and aligned with video, then indexing accuracy for actions and scenes improves, but processing time and computational resources increase
Solution Approach 1:
The patent applies partial action by selectively processing only the descriptive audio track for indexing purposes, rather than transcribing all audio content including dialogue, music, and sound effects. This selective approach focuses computational resources on the specific audio stream that contains action and scene descriptions, improving indexing accuracy for visual content while reducing overall processing time compared to comprehensive audio transcription.
Data Source
AI summary
Disclosed are systems, methods, and computer readable media for retrieving digital images. The method embodiment includes converting a descriptive audio stream of a digital video that is provided for the visually impaired to text and then aligning that text to the appropriate segment of the digital video. The system then indexes the converted text from the descriptive audio stream with the text's relationship to the digital video. The system enables queries using action words describing a desired scene from a digital video.


