Descriptive Audio Indexing for Video Scene Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for retrieving video content using automatic speech recognition are inadequate, as they rely on conversational audio and may not effectively index scenes without specific keywords, limiting the ability to retrieve specific events or actions.

Innovation Solution

A system that converts descriptive audio streams from digital videos into text, aligns the text with the video, and creates an index for searching, utilizing the descriptive audio to enable retrieval of specific scenes and actions, such as 'a man standing' without requiring frame annotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If automatic speech recognition is used on conversational audio to retrieve video content, then the system can identify mentioned subjects, but it fails to retrieve scenes without specific keywords in dialogue

Engineering Contradiction:
Improveretrieval accuracyVSAvoidsearch coverage
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the video content processing into two distinct indexing approaches: traditional dialogue-based ASR indexing and the new descriptive audio indexing. By separating these functions, the system can use descriptive audio specifically for action and scene description without interfering with dialogue processing, thereby improving both retrieval accuracy for actions and overall search coverage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces descriptive audio as an intermediary layer between the video content and the search index. This descriptive audio track serves as a mediator that translates visual actions and scenes into textual descriptions, enabling the ASR system to index and retrieve content based on visual events rather than relying solely on dialogue mentions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If frame-by-frame annotation is performed to enable action-based retrieval, then retrieval precision for actions improves, but system complexity and processing time increase significantly

Engineering Contradiction:
Improveaction retrieval precisionVSAvoidindexing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-generating descriptive audio tracks during video production or post-production, before the indexing process. These descriptive audio tracks contain pre-processed information about actions and scenes in human language, eliminating the need for complex frame-by-frame analysis during indexing. The ASR system then simply transcribes this pre-existing descriptive audio, dramatically reducing processing complexity while maintaining high retrieval precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical frame-by-frame visual analysis system with an acoustic-based ASR system that processes descriptive audio. Instead of computationally intensive computer vision algorithms analyzing each video frame, the system uses audio processing and speech recognition to extract action information, significantly reducing device complexity and processing requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If descriptive audio streams are converted to text and aligned with video, then indexing accuracy for actions and scenes improves, but processing time and computational resources increase

Engineering Contradiction:
Improveindexing accuracyVSAvoidindexing processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by selectively processing only the descriptive audio track for indexing purposes, rather than transcribing all audio content including dialogue, music, and sound effects. This selective approach focuses computational resources on the specific audio stream that contains action and scene descriptions, improving indexing accuracy for visual content while reducing overall processing time compared to comprehensive audio transcription.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9465870B2System and method for digital video retrieval involving speech recognition
Publication Date: 2016.10.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9465870B2 patent drawing
  • US9465870B2 patent drawing
  • US9465870B2 patent drawing

AI summary

Disclosed are systems, methods, and computer readable media for retrieving digital images. The method embodiment includes converting a descriptive audio stream of a digital video that is provided for the visually impaired to text and then aligning that text to the appropriate segment of the digital video. The system then indexes the converted text from the descriptive audio stream with the text's relationship to the digital video. The system enables queries using action words describing a desired scene from a digital video.