Rich Media Annotation via Audio Metadata Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The proliferation of user-generated video content lacks sufficient metadata for effective searchability, making it difficult and costly for news agencies and forensic agencies to find specific footage, as users often find tagging videos tedious and incomplete.

Innovation Solution

A system for rich media annotation that receives and analyzes audio annotations to extract metadata, using speech recognition and specially trained grammars to identify and associate relevant information with media content, allowing for easier and more convenient metadata addition, storage, and retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If users manually add tags and categories to videos, then some metadata is added, but the process is tedious and tags are incomplete and less than optimal

Engineering Contradiction:
Improvemetadata completenessVSAvoidtagging convenience
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system enables self-service by automatically extracting metadata from audio annotations without requiring manual user intervention. The speech recognition system and trained grammars autonomously process audio content to generate metadata tags, eliminating the tedious manual tagging process while maintaining high completeness of metadata.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual tagging process with an automated speech recognition system. Instead of users manually entering tags, the system uses speech-to-text conversion combined with specially trained grammars to automatically extract and assign metadata, substituting human labor with automated processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If traditional search methods are used on user-generated videos, then search capability is limited, but the volume of unsearchable footage is overwhelming

Engineering Contradiction:
Improvesearch efficiencyVSAvoidsearchability
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system performs preliminary action by extracting and associating metadata with video content at the time of upload or processing, rather than requiring search-time analysis. Audio annotations are processed in advance to generate metadata that is stored and associated with the video, enabling efficient subsequent searches without reprocessing the entire video content.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces metadata as an intermediary between the video content and search queries. Instead of searching through raw video data directly, the system uses extracted metadata from audio annotations as a searchable index, mediating between the overwhelming volume of video content and user search requests.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If multiple cameras record the same event, then a wealth of information is captured, but the clips are not uniformly tagged making them difficult to search

Engineering Contradiction:
Improveinformation availabilityVSAvoidtagging uniformity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system applies universality by using a single standardized metadata extraction process that works across multiple video sources and formats. The same speech recognition system and trained grammars process audio annotations from all cameras uniformly, ensuring consistent metadata generation regardless of the source device or recording conditions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the parameter of tagging consistency by implementing automated metadata extraction that produces uniform output across all video clips. Instead of relying on inconsistent manual tagging, the system transforms the tagging process into a standardized automated operation that maintains consistent metadata structure and quality across all sources.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If forensic agencies search through video clips for evidence, then comprehensive coverage is achieved, but the process requires great difficulty and expense

Engineering Contradiction:
Improveevidence search completenessVSAvoidsearch cost and time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-extracting and indexing metadata from all video content before forensic search is initiated. Audio annotations are processed in advance to create a searchable metadata index, so when forensic agencies need to search for evidence, they can query the pre-processed metadata rather than analyzing raw video content, dramatically reducing search time and cost.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses metadata as an intermediary that enables efficient forensic search. Instead of directly searching through overwhelming volumes of raw video footage, the system uses extracted metadata from audio annotations as a searchable index, mediating between the comprehensive video archive and forensic query requirements while maintaining reliability through the structured metadata approach.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10127231B2System and method for rich media annotation
Publication Date: 2018.11.13 AT&T INTELLECTUAL PROPERTY I L P
  • US10127231B2 patent drawing

AI summary

Disclosed herein are systems, methods, and computer readable-media for rich media annotation, the method comprising receiving a first recorded media content, receiving at least one audio annotation about the first recorded media, extracting metadata from the at least one of audio annotation, and associating all or part of the metadata with the first recorded media content. Additional data elements may also be associated with the first recorded media content. Where the audio annotation is a telephone conversation, the recorded media content may be captured via the telephone. The recorded media content, audio annotations, and/or metadata may be stored in a central repository which may be modifiable. Speech characteristics such as prosody may be analyzed to extract additional metadata. In one aspect, a specially trained grammar identifies and recognizes metadata.