Audio Video Content Segmentation via Metadata Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio/video (A/V) content analysis and interaction technologies are limited, making it difficult for users to discover, navigate, and consume large amounts of A/V content due to the complexity of speech recognition and language analysis, and the linear, non-digestible nature of A/V content, which lacks visual snapshots for quick relevance evaluation.

Innovation Solution

The system employs speech and language analysis components to generate metadata, which is used in creating user interface interaction components, allowing users to view and interact with A/V content in various segments, enabling efficient browsing and navigation through automatically generated metadata.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If A/V content is provided as regular length content (30-60 minute shows), then content completeness is improved, but user navigation and discovery becomes difficult due to linear consumption format

Engineering Contradiction:
Improvecontent completenessVSAvoiduser navigation
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments A/V content into multiple chapters or sections with distinct metadata tags, allowing users to navigate to specific segments of interest rather than consuming content linearly. The system divides a complete show into navigable units while preserving the full content for users who want to watch the entire program.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of interaction by adding metadata-driven navigation layers over the traditional linear A/V playback. Users can jump between chapters, search by topic, and access content non-linearly through metadata indexes, transforming the one-dimensional linear consumption into multi-dimensional navigation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Difficulty of detecting and measuring

If speech recognition and language analysis are implemented for A/V content, then content analysis capability is improved, but system complexity increases due to non-trivial processing requirements

Engineering Contradiction:
Improvecontent analysis capabilityVSAvoidsystem complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The patent performs speech recognition and language analysis in advance during content ingestion, generating metadata indexes and chapter markers before user interaction. This preliminary processing creates reusable metadata structures that enable fast searching and navigation without requiring complex real-time analysis during playback.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces metadata as an intermediary layer between the raw A/V content and the user interface. Instead of directly processing complex speech recognition queries during playback, the system uses pre-generated metadata tags, chapter markers, and topic indexes as intermediaries to enable efficient content retrieval and navigation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If A/V content is provided as short clips with rich metadata, then discoverability is improved, but content completeness deteriorates

Engineering Contradiction:
ImprovediscoverabilityVSAvoidcontent completeness
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent merges the advantages of short clips (rich metadata, easy discovery) with the advantages of long-form content (completeness, context) by organizing A/V content into chaptered structures. Each chapter functions as a discoverable unit with its own metadata, while collectively forming a complete program that preserves full content context and narrative flow.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS7640272B2Using automated content analysis for audio/video content consumption
Publication Date: 2009.12.29 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7640272B2 patent drawing
  • US7640272B2 patent drawing
  • US7640272B2 patent drawing

AI summary

Audio/video (A/V) content is analyzed using speech and language analysis components. Metadata is automatically generated based upon the analysis. The metadata is used in generating user interface interaction components which allow a user to view subject matter in various segments of the A/V content and to interact with the A/V content based on the automatically generated metadata.