Video Caption Editing with Timestamp Preservation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video display technologies lack efficient methods for synchronizing and editing text captions in videos to maintain alignment, synchronization, and timing with the original speech, leading to suboptimal presentation of captions.

Innovation Solution

A computing system that transcribes audio data into text, associates timestamps with each word, and allows users to modify captions while preserving the original timing and pacing, ensuring accurate alignment and synchronization of captions with the video sequence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated transcription is used to generate captions quickly, then productivity is improved, but manufacturing precision deteriorates due to transcription errors

Engineering Contradiction:
Improvecaption generation speedVSAvoidtranscription accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system allows users to self-correct transcription errors directly in the caption editor, enabling them to serve their own needs for accurate captions without requiring manual re-transcription. The preservation of timestamps during editing supports this self-service approach.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system provides feedback by displaying transcribed captions with timestamps that users can review and correct. The ability to see which words have timestamps and how they align with video timing gives users feedback on transcription quality and areas needing correction.

Inventive Principle:
Principle #23Feedback

2Manufacturing precision

If captions are edited to fix errors, then manufacturing precision is improved, but loss of time increases due to manual editing

Engineering Contradiction:
Improvecaption accuracyVSAvoidediting time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary transcription automatically before the user editing phase, so that when users arrive at the editing stage, the captions are already generated and ready for review. This preliminary action reduces the total time required compared to starting from scratch.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a copy of the transcribed text that can be edited without affecting the original transcription process. Users work with this copy, preserving the timestamp structure, which allows efficient editing without re-doing the entire transcription.

Inventive Principle:
Principle #26Copying

3Reliability

If timestamps are preserved during caption editing, then reliability is improved, but device complexity increases

Engineering Contradiction:
Improvecaption synchronizationVSAvoidtimestamp management
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system uses timestamps as an intermediary element that mediates between the original transcription and the edited captions. By preserving timestamps as a separate data structure, the system maintains synchronization information without requiring complex re-synchronization logic during editing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11272137B1Editing text in video captions
Publication Date: 2022.03.08 META PLATFORMS TECHNOLOGIES LLC
  • US11272137B1 patent drawing
  • US11272137B1 patent drawing
  • US11272137B1 patent drawing

AI summary

This disclosure describes techniques that include modifying text associated with a sequence of images or a video sequence to thereby generate new text and overlaying the new text as captions in the video sequence. In one example, this disclosure describes a method that includes receiving a sequence of images associated with a scene occurring over a time period; receiving audio data of speech uttered during the time period; transcribing into text the audio data of the speech, wherein the text includes a sequence of original words; associating a timestamp with each of the original words during the time period; generating, responsive to input, a sequence of new words; and generating a new sequence of images by overlaying each of the new words on one or more of the images.