Structured Video Transcripts for Semantic Query Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video player interfaces lack the ability to semantically understand user queries and provide relevant information due to limitations in keyword-based searches and unstructured transcripts/captions, making it difficult to locate specific content efficiently.

Innovation Solution

A system that generates a semantically-rich, structured document from audio-visual content, incorporating creator-provided text and speaker diarization, using a large language model to process user queries and provide coherent responses, including audio and image data alignment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keyword searches within transcripts/captions are used to locate content in videos, then the ability to search for content is improved, but the ability to semantically understand user queries and provide relevant information is lost

Engineering Contradiction:
Improvesearch accuracyVSAvoidsemantic understanding capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary processing layer between the user query and the video content. This layer includes a speech-to-text conversion module that converts spoken queries into text, and a semantic analysis module that processes the text to understand the user's intent. This intermediary enables the system to go beyond simple keyword matching and provide semantically relevant information from the video content.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If timeline-based video player interfaces are used to scrub through videos, then users can locate particular content, but the process is inefficient and time-consuming

Engineering Contradiction:
Improvecontent location capabilityVSAvoidtime to locate content
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces the mechanical scrubbing action with an automated speech recognition and semantic search system. Instead of manually moving through the video timeline, users can speak their query which is converted to text, processed for semantic meaning, and the system automatically locates and plays the relevant segment. This substitution eliminates the time-consuming manual scrubbing process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If transcripts/captions are generated for dialog in videos, then the ability to search for content is improved, but the user interface lacks the ability to fulfill queries with semantically relevant information

Engineering Contradiction:
Improvecontent accessibilityVSAvoidquery fulfillment capability
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The patent adds a new dimension of semantic processing to the existing transcript-based search. Beyond the textual dimension of transcripts, the system introduces semantic analysis that understands context, intent, and meaning. This additional dimension allows the system to fulfill queries with semantically relevant information rather than just matching keywords in the transcript.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12613915B2Structured video documents
Publication Date: 2026.04.28 GOOGLE LLC
  • US12613915B2 patent drawing
  • US12613915B2 patent drawing
  • US12613915B2 patent drawing

AI summary

A method includes receiving a content feed that includes audio data corresponding to speech utterances and processing the content feed to generate a semantically-rich, structured document. The structured document includes a transcription of the speech utterances and includes a plurality of words each aligned with a corresponding audio segment of the audio data that indicates a time when the word was recognized in the audio data. During playback of the content feed, the method also includes receiving a query from a user requesting information contained in the content feed and processing, by a large language model, the query and the structured document to generate a response to the query. The response conveys the requested information contained in the content feed. The method also includes providing, for output from a user device associated with the user, the response to the query.