Caption-Based Video Navigation for Context-Rich Segment Jumping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video navigation interfaces are difficult to use, particularly for longer videos, as they lack intuitive methods for jumping to specific points and often require manual fast forwarding, consuming bandwidth and screen space, especially on small devices.

Innovation Solution

A user interface that enables navigation through video based on displayed text, allowing multiple lines of captions to be overlaid simultaneously, with directional inputs to adjust play position, and the ability to select and search for terms to preload relevant video segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a progress bar with slider is used for video navigation, then the user can jump to specific points, but the interface becomes difficult to use for longer videos and lacks intuitive navigation methods

Engineering Contradiction:
Improvevideo navigation easeVSAvoidnavigation interface complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the navigation interface by introducing chapter markers that divide the video into logical sections. Users can select chapters rather than navigating through continuous time, making the interface simpler and more intuitive for long videos. Each chapter represents a segment of content, allowing users to jump between meaningful sections rather than dealing with a continuous progress bar.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a hierarchical dimension to navigation by introducing chapters as a higher-level organization structure. Instead of only having temporal navigation (progress bar), users now have both chapter-based navigation and time-based navigation, creating a two-dimensional navigation space that improves ease of operation for long videos.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If captions are displayed in a separate window, then captions are readable, but valuable screen space is consumed especially on small devices

Engineering Contradiction:
Improvecaption readabilityVSAvoidscreen space
Core Design Contradiction:
Ease of operationVSArea of stationary object

Solution Approach 1:

The patent merges the caption display with the video playback interface by overlaying captions directly on the video player. This integration eliminates the need for separate caption windows, preserving screen space while maintaining caption readability. The captions are displayed in an overlay layer that combines with the video visual elements.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent moves captions from a separate spatial dimension (separate window) to the same spatial dimension as the video (overlay layer). By using a layered display architecture, captions are positioned in front of the video content without requiring additional window space, thus optimizing screen real estate on small devices.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of operation

If manual fast forwarding is used to find specific video segments, then the user can locate desired content, but a lot of bandwidth is consumed loading irrelevant parts

Engineering Contradiction:
Improvevideo segment locationVSAvoidbandwidth consumption
Core Design Contradiction:
Ease of operationVSLoss of energy

Solution Approach 1:

The patent implements preliminary indexing of video content by extracting and storing keyframes, chapters, and caption information before actual playback. This pre-processing creates a navigation database that allows users to jump directly to desired segments without manually fast-forwarding through the entire video, significantly reducing bandwidth consumption during navigation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary navigation system consisting of chapters and keyframes that mediates between the user's navigation requests and the actual video data. Instead of loading continuous video data during manual fast-forwarding, the system uses these intermediary markers to enable direct jumping to specific segments, reducing unnecessary bandwidth consumption.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of information

If multiple lines of captions are displayed simultaneously, then context is improved, but the interface complexity increases

Engineering Contradiction:
Improvecaption contextVSAvoidinterface complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the caption display into multiple lines that correspond to different time segments of the video. By organizing captions into discrete, time-based lines rather than a continuous stream, the system provides better contextual information while maintaining manageable interface complexity. Each line represents a specific moment or segment, making it easier to navigate and understand.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms where selecting or highlighting a specific caption line provides contextual information about that segment, and the video playback responds by jumping to or pausing at the corresponding time. This interactive feedback loop enhances context while managing complexity through user-initiated navigation rather than requiring complex automatic processing.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12574608B2User interface method and apparatus for video navigation using captions
Publication Date: 2026.03.10 ADEIA GUIDES INC
  • US12574608B2 patent drawing
  • US12574608B2 patent drawing
  • US12574608B2 patent drawing

AI summary

Systems and methods for navigating a video via interaction with text overlaid on the video are described. In one example, a method includes generating for display a video, and generating for display at least one line of text overlaid over the video. Then, in response to receiving a directional user interface input for at least a portion of the at least one line of text, the method includes modifying a play position of the video based on a direction of the directional user interface input for the at least a portion of the at least one line of text.