Automated Caption Synchronization via OCR and Audio Diffing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current captioning systems face challenges in accurately matching spoken words with displayed captions, synchronicity issues, and labor-intensive quality assessment methods prone to human bias and errors, which affect accessibility for individuals who are Hard of Hearing, Deaf, or Deaf-Blind.
Innovation Solution
A computing environment processes video content to generate a descriptive textual file detailing displayed captions and their timing relative to the audio, using optical character recognition and diffing algorithms to align captions with spoken audio, reducing human error and enabling efficient quality assessment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual visual review of audio and video feeds is used to measure caption quality, then quality assessment can be performed, but the process becomes labor intensive and prone to human bias and errors
Solution Approach 1:
The patent replaces the manual mechanical process of visual review with an automated computer-based system that uses optical character recognition (OCR) to extract caption text from video frames and algorithms to compare it with audio transcript data, thereby eliminating human labor and subjectivity while maintaining measurement accuracy
Solution Approach 2:
The system creates a digital copy of the caption text from the video feed through OCR, then compares this copied text against the audio transcript using automated algorithms, enabling precise and efficient quality assessment without manual intervention
2Ease of operation
If captions are displayed simultaneously in blocks of words, then viewing is enabled for hearing-impaired individuals, but accurate matching with spoken words and proper synchronicity becomes difficult to achieve
Solution Approach 1:
The system uses automated comparison between extracted caption text and audio transcript data to provide feedback on timing and accuracy, enabling precise adjustment of caption synchronization to match spoken words while maintaining block display format for accessibility
Solution Approach 2:
The system performs preliminary extraction and comparison of caption text with audio data before final display, allowing synchronization adjustments to be made in advance to ensure accurate matching with spoken words while maintaining accessible block display format
Data Source
AI summary
Disclosed are a method, a system, and a non-transitory computer readable medium for identifying captions in captioned video. A method includes receiving audio and video content from a caption device where the video content includes captioned text, extracting frames of video from the received video content where the frames of video include captioned text, recognizing text from the captioned text in the extracted frames of video, and generating a descriptive textual file including timing information for the recognized text and timing information for the captioned text.


