Video Segmenting System for Object-Based Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video navigation and interaction methods are limited, making it difficult for users to search for specific content within videos due to the lack of contextual metadata, resulting in inefficient searching and navigation for particular video segments associated with objects.

Innovation Solution

A video segmenting system that generates video segments based on object identifiers, using audio/visual data and metadata to define start and stop frames within a video, allowing users to easily search and navigate to relevant content by identifying objects and keywords, with tags and indicators displayed on user devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users navigate videos using coarse controls (fast forward, rewind, scrub), then basic video playback functionality is maintained, but the ability to search for and locate specific content segments associated with objects is lost

Engineering Contradiction:
Improvevideo navigationVSAvoidcontextual metadata
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The video is segmented into multiple clips based on object appearances and audio transcripts. Each clip is associated with specific objects and timecodes, allowing users to navigate to specific content segments by searching for object names or keywords rather than manually scrubbing through the entire video.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary layer between the user and the video content - a search interface that accepts object names or keywords as input and returns relevant video clips with timecodes. This intermediary enables precise content location without requiring users to manually navigate through the video.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If users manually search through video content to find specific objects or segments, then comprehensive video exploration is possible, but the searching process becomes laborious and inefficient

Engineering Contradiction:
Improvecontent location accuracyVSAvoidsearch time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by automatically generating object annotations, extracting audio transcripts, and creating an indexed database of video clips with associated objects and timecodes before user search. This pre-processing enables instant retrieval of relevant segments when users search for specific objects or keywords.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a simplified copy or representation of the video content in the form of text-based object annotations and audio transcripts. Users can search this textual representation instead of manually viewing the entire video, dramatically reducing search time while maintaining content location accuracy.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If video metadata is enhanced with object identifiers and audio transcripts, then searchability and navigation capability are improved, but the complexity of the video processing system increases

Engineering Contradiction:
Improvevideo search capabilityVSAvoidvideo processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs multi-functional components that perform multiple tasks. For example, the object detection model not only identifies objects but also generates timecodes and annotations. The audio processing system both transcribes speech and extracts keywords for search. This multi-functionality reduces overall system complexity despite enhanced capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses self-service approaches by leveraging pre-trained models for object detection and audio transcription. These models automatically process video content without requiring manual annotation or complex custom processing pipelines, enabling enhanced metadata generation while keeping the system architecture relatively simple.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11120490B1Generating video segments based on video metadata
Publication Date: 2021.09.14 AMAZON TECH INC
  • US11120490B1 patent drawing
  • US11120490B1 patent drawing
  • US11120490B1 patent drawing

AI summary

A video segmenting system identifies a product for sale in a video and determines one or more attributes of audio and video content within the video. The video segmenting system determines a video segment within the video that is associated with the product for sale, based on the attributes. The video segmenting system generates a tag that associates the product for sale with the video segment and sends an indication of the tag to a user device. Once the video is played on a user device, the user device detects a search query about the product for sale. Using the tag, the user device can display a marker on the user device corresponding to the location of the video segment within the video.