Spatial Video Navigation via Boundary Detection and Overview Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video navigation techniques are inadequate for educational videos, as they lack spatial navigation solutions, making it difficult for users to navigate based on the content being presented rather than temporal positions.

Innovation Solution

A computer-implemented method that detects boundary events in videos, segments them, generates an overview image, and maps it to video segments, allowing users to navigate through a graphical user interface by selecting portions of the overview image to play associated video segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional timeline-based navigation is used for educational videos, then users can rapidly jump to different time points and maintain awareness of temporal position, but users cannot effectively navigate based on the spatial content being presented in the video

Engineering Contradiction:
Improvespatial navigation capabilityVSAvoidcontent location awareness
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent transforms the one-dimensional temporal navigation (timeline) into two-dimensional spatial navigation by creating an overview image that displays the spatial layout of content within the video. This allows users to navigate based on visual content position rather than just temporal markers, adding a spatial dimension to the navigation interface.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system creates a simplified copy of the video content in the form of an overview image that represents the spatial arrangement of content. This copy allows users to visually locate content before watching, providing content location awareness without requiring users to watch the entire video.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If educational videos are made comprehensive and lengthy to cover all necessary material, then more content is available for learning, but the videos become more difficult to navigate and locate specific content

Engineering Contradiction:
Improveeducational content volumeVSAvoidcontent location difficulty
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments the lengthy video content into distinct regions within the overview image, allowing users to visually identify and navigate to specific content sections. This segmentation maintains the comprehensive nature of the video while making it easier to locate specific topics without having to scan through the entire video duration.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis of the video content to generate the overview image before the user watches the video. This preliminary action extracts and organizes content information in advance, enabling users to quickly locate specific topics without having to navigate through the entire lengthy video sequentially.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10482777B2Systems and methods for content analysis to support navigation and annotation in expository videos
Publication Date: 2019.11.19 FUJIFILM BUSINESS INNOVATION CORP
  • US10482777B2 patent drawing
  • US10482777B2 patent drawing
  • US10482777B2 patent drawing

AI summary

Online educational videos are often difficult to navigate. Furthermore, most video interfaces do not lend themselves to note-taking. Described system detects and reuses boundaries that tend to occur in these types of videos. In particular, many educational videos are organized around distinct breaks that correspond to slide changes, scroll events, or a combination of both. Described algorithms can detect these structural changes in the video content. From these events the system can generate navigable overviews to help users searching for specific content. Furthermore, these boundary events can help the system automatically associate rich media annotations to manually-defined bookmarks. Finally, when manual or automatically recovered spoken transcripts are available, the spoken text can be combined with the temporal segmentation implied by detected events for video indexing and retrieval. This text can also be used to seed a set of text annotations for user selection or be combined with user text input.