Video Content Section Identification via Indexed Text Transcripts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in efficiently finding specific scenes or quotes within video content, as existing search methods are time-consuming and frustrating, especially when the title of the video is unknown or misremembered.

Innovation Solution

The development of indexed sequences of video content, generated using speech-to-text or crowd-sourced transcriptions with timestamp annotations, allows for quick identification of particular portions through text representation and machine learning classifiers, enabling efficient search and generation of custom content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users manually search through video content by selecting different chapters or using seek bars, then they can find specific scenes or quotes, but the search process becomes time-consuming and frustrating

Engineering Contradiction:
Improvescene identification accuracyVSAvoidsearch time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by generating text representations of video content in advance and storing them in an indexed sequence. When a user searches, the system queries the pre-generated text representation instead of analyzing video content in real-time, dramatically reducing search time while maintaining accurate scene identification

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical searching (clicking seek bars, selecting chapters) with automated text-based search. Users input text queries and the system automatically matches them against the indexed text representation, substituting manual navigation with intelligent text processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If users search for video content by title or context, then they can locate specific videos, but the process becomes more time-consuming when the title is unknown or misremembered

Engineering Contradiction:
Improvesearch convenienceVSAvoidsearch time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system introduces text representation as an intermediary between video content and user queries. Instead of directly searching video files or metadata, users search through text descriptions that mediate between their intent and the actual video content, making searches more convenient and faster

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the search parameter from requiring exact titles or contextual knowledge to allowing flexible text queries. The system accepts various text inputs (quotes, descriptions, keywords) and matches them against the indexed text representation, reducing the time users spend remembering exact titles

Inventive Principle:
Principle #35Parameter changes

3Productivity

If the system generates text representations of video content, then search efficiency improves, but the device complexity increases

Engineering Contradiction:
Improvesearch speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system extracts only the essential text representation from video content, separating this searchable information from the full video file. This extraction approach improves search productivity while limiting the added complexity to only the necessary text processing components, rather than requiring complex video analysis systems

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10115433B2Section identification in video content
Publication Date: 2018.10.30 AMAZON TECH INC
  • US10115433B2 patent drawing
  • US10115433B2 patent drawing
  • US10115433B2 patent drawing

AI summary

Video content can be analyzed to identify particular sections of the video content. Speech to text or similar techniques can be used to obtain a transcription of the video content. The transcription can be indexed (e.g., timestamped) to the video content. Information describing how users are interacting with or consuming the video content (e.g., social media information, viewing history data, etc.) can be collected and used to identify the particular sections. Once the particular sections have been identified, other services can be provided. For example, custom trailers and summaries of the video content can be generated based on the identified sections. Additionally, the video content can be augmented to include additional information relevant to the particular sections, such as production information, actor information, or other information. The additional information can be added so as not to interfere with the important sections.