Subtitle Synchronization Error Detection Using Speech-to-Text

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current media content often experiences synchronization errors between audio and subtitles, leading to immersion issues and inefficiencies in user experience, as conventional methods rely on manual identification and correction.

Innovation Solution

A service provider computer implements a synchronization identification feature that automatically identifies and corrects synchronization errors by parsing subtitle files, using speech-to-text algorithms, and edit distance algorithms to adjust metadata, thereby synchronizing audio and subtitles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual methods are used to identify and correct synchronization errors, then user input is required and errors can be corrected, but the process is inefficient and time-consuming

Engineering Contradiction:
Improveerror correction efficiencyVSAvoidtime to identify and fix errors
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system enables self-service by automatically detecting and correcting synchronization errors without requiring manual user intervention. The service provider computer autonomously analyzes subtitle files, compares audio-subtitle synchronization, identifies timing offsets, and applies corrections to the metadata, making the system self-correcting and eliminating the need for manual error fixing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical operations with automated computational processes. Instead of manual viewing and timing analysis, the system uses speech-to-text algorithms to generate audio captions with timestamps, edit distance algorithms to compare text matching, and automated metadata modification to correct synchronization errors, significantly improving efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If automated speech-to-text and edit distance algorithms are used, then user input is reduced and processing is faster, but computational resources are required

Engineering Contradiction:
Improveautomation of error identificationVSAvoidcomputational resource consumption
Core Design Contradiction:
Extent of automationVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by processing only the necessary portions of audio and subtitle files to identify synchronization errors. Rather than analyzing entire media content, the service provider computer extracts relevant audio segments corresponding to subtitle cues and processes only those portions needed for synchronization detection and correction, reducing overall computational resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If metadata is automatically modified to correct errors, then synchronization is improved and user immersion is maintained, but the subtitle file structure must be precisely manipulated

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidmetadata manipulation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses an intermediary approach by introducing a service provider computer as a mediator between the subtitle file and the correction process. This intermediary system safely parses the subtitle file structure, identifies the specific metadata fields requiring adjustment, calculates the precise timing corrections needed, and applies modifications through controlled metadata manipulation, reducing the risk of errors and simplifying the overall process.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10423660B1System for detecting non-synchronization between audio and subtitle
Publication Date: 2019.09.24 AMAZON TECH INC
  • US10423660B1 patent drawing
  • US10423660B1 patent drawing
  • US10423660B1 patent drawing

AI summary

Techniques for identifying and correcting synchronization errors between audio and subtitles for media content are described herein. For example, a portion of a subtitle file associated with media content may be extracted based on subtitle cues included in the portion of the subtitle file. In embodiments, an audio to text file may be generated from the extracted portion using a speech to text algorithm. A detected subtitle text file may be generated using the subtitle file, the audio to text file, and an edit distance algorithm. In embodiments, one or more synchronization errors between the audio and subtitles for the media content may be identified based on time stamp information associated with the audio to text file and a subtitle cue for the extracted portion of the subtitle file.