Multimedia Word Cloud Navigation via Timestamp Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in efficiently navigating multimedia content, such as audio or video, as they typically rely on a cumbersome hit-and-trial method using seek bars to access specific parts of interest, which can be time-consuming and inefficient.

Innovation Solution

A method and system that extract words from multimedia content with associated timestamps, create a word cloud based on emphasis and occurrence, and allow users to select words to view corresponding multimedia snippets, enabling direct navigation to relevant parts of the content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users navigate multimedia content using a seek bar with hit and trial method, then users can access specific parts of the content, but the navigation process becomes time-consuming and cumbersome

Engineering Contradiction:
Improvenavigation operationVSAvoidtime to access point of interest
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary extraction of words and timestamps from the multimedia content before the user needs to navigate. A word cloud is pre-generated displaying all extractable words with their temporal positions, allowing users to immediately see and select their point of interest without sequential searching or trial-and-error navigation through the seek bar.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary interface element - the word cloud - that mediates between the user and the multimedia content. Instead of directly manipulating the seek bar or scrubbing through content, users interact with the word cloud by selecting words, which then automatically positions the playback at the corresponding timestamp. This intermediary layer abstracts the complex navigation task into a simple selection action.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If a seek bar is displayed to enable navigation, then users have the option to start playback from a point of interest, but the interface complexity increases and user experience deteriorates

Engineering Contradiction:
Improvenavigation flexibilityVSAvoidinterface complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts the essential navigation functionality from the traditional seek bar interface. Instead of requiring users to interpret visual position on a timeline and manually scrub, the system extracts and displays only the meaningful semantic elements (words) from the content. Users select from these extracted words directly, which eliminates the need for a traditional seek bar while preserving and enhancing navigation flexibility.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of navigation from spatial (position on seek bar) to semantic (word selection). Instead of navigating by moving a slider along a time axis, users navigate by selecting semantic keywords that represent their point of interest. This parameter transformation simplifies the interface while maintaining adaptability, as users can jump to any meaningful point in the content through word selection rather than manual scrubbing.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9484032B2Methods and systems for navigating through multimedia content
Publication Date: 2016.11.01 VIDEOKEN INC
  • US9484032B2 patent drawing
  • US9484032B2 patent drawing
  • US9484032B2 patent drawing

AI summary

The disclosed embodiments illustrate methods and systems for processing multimedia content. The method includes extracting one or more words from an audio stream associated with multimedia content. Each word has associated one or more timestamps indicative of temporal occurrences of said word in said multimedia content. The method further includes creating a word cloud of said one or more words in said multimedia content based on a measure of emphasis laid on each word in said multimedia content and said one or more timestamps associated with said one or more words. The method further includes presenting one or more multimedia snippets, of said multimedia content, associated with a word selected by a user from said word cloud. Each of said one or more multimedia snippets corresponds to said one or more timestamps associated with occurrences of said word in said multimedia content.