Audio Video Content Search Using Transcripts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search methods for video and audio content are limited, as they primarily rely on program data and metadata, which may not adequately represent the actual content, failing to provide users with specific information they seek.

Innovation Solution

A system that analyzes video and audio content to identify spoken words, generates transcripts, and performs text searches to find matching keywords, providing time indicators for the location of matched words within the content, allowing users to play and highlight relevant sections.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional search methods rely on program data and metadata, then the search system is simple to operate, but the search accuracy and representativeness of actual content deteriorates

Engineering Contradiction:
Improvesearch accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces transcripts as an intermediary layer between the original audio/video content and the search system. These transcripts convert spoken content into searchable text, enabling accurate keyword-based searches without requiring users to analyze raw media files directly. This mediator approach maintains search simplicity while dramatically improving accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual content analysis with automated speech-to-text conversion and computerized text searching. Instead of users manually reviewing video/audio content to find information, the system automatically generates transcripts and enables text-based search, substituting mechanical human analysis with automated computational processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If the system analyzes spoken words to generate transcripts, then the representativeness of actual content improves, but the processing time and computational resources increase

Engineering Contradiction:
Improvecontent representativenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs speech-to-text conversion and transcript generation in advance, before the actual search operation. By pre-processing the content into searchable text format, the system eliminates the need for real-time analysis during user searches, thus reducing perceived processing time while maintaining complete content representativeness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates text copies (transcripts) of the original audio/video content. These transcripts are simplified representations that preserve the semantic information of the spoken content while being much faster to process and search. The system searches the text copy rather than the original media, significantly reducing processing time while maintaining information accuracy.

Inventive Principle:
Principle #26Copying

3Ease of operation

If the system provides time indicators for matched words, then the user can quickly locate specific content segments, but the complexity of result presentation increases

Engineering Contradiction:
Improvecontent location easeVSAvoidresult presentation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces time indicators as an intermediary element that bridges the search results and the original content. These indicators act as navigation markers that guide users to specific locations in the media without requiring complex analysis or interpretation. The time indicator simplifies the user's task of locating content while adding minimal structural complexity to the system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10277953B2Search for content data in content
Publication Date: 2019.04.30 DIRECTV LLC
  • US10277953B2 patent drawing
  • US10277953B2 patent drawing
  • US10277953B2 patent drawing

AI summary

Video and audio content is searchable using a text search. A search component can analyze respective items of content to identify words spoken in the items of content, and generate respective transcripts of the respective words of the items of content based on the analysis. The search component receives a text search comprising a keyword and analyzes the respective transcripts to determine whether a transcript(s) contains a word that matches or substantially matches the keyword. The search component generates a search result(s) associated with the transcript(s) that at least is a substantial match to the keyword. The search component can present a time indicator indicating a time position in proximity to where the word is located in the content of the search result(s), and presentation of the content can start from that time position. The search component can be executed in a set-top box associated with a presentation device.