Automated Closed Caption Error Detection in Video Assets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multimedia content distribution networks (MCDNs) face challenges in quality control and performance monitoring, particularly with video assets, as user feedback and costly support visits are often relied upon to identify issues, and there is a need for efficient methods to validate closed captioning accuracy.

Innovation Solution

A method and system for monitoring baseband video signals in MCDNs, involving the extraction of text strings from video images and audio tracks, comparison of these strings, and generation of caption errors when matching falls below a predetermined threshold, with automated logging and notification of errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If user feedback and support visits are used to identify quality issues, then quality control can be achieved, but the process becomes costly and time-consuming

Engineering Contradiction:
Improvequality controlVSAvoidtime-consuming
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary automated testing of closed captioning accuracy before content is distributed to users. By extracting text from video frames, converting audio to text, and comparing the two in advance, the system identifies and corrects caption errors before they reach end users, eliminating the need for time-consuming user feedback loops and support visits.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The quality control system performs self-testing and self-validation of closed captioning content. The automated extraction, conversion, and comparison processes enable the system to monitor and validate its own content quality without external user intervention, thereby reducing dependency on user feedback and support visits while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated text extraction and comparison is implemented, then closed captioning accuracy can be validated efficiently, but system complexity increases

Engineering Contradiction:
Improvevalidation efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system employs multi-functional components that perform multiple operations. For example, the text extraction module not only extracts closed caption text from video frames but also handles synchronization with audio timelines. The audio conversion module both converts speech to text and performs initial quality filtering. This multi-functionality reduces the number of separate components needed, thereby managing system complexity while maintaining high validation efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If frequent monitoring of video assets is performed, then quality issues can be detected early, but computational resources are consumed

Engineering Contradiction:
Improveerror detectionVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system implements periodic monitoring at strategically chosen intervals rather than continuous monitoring. Closed captioning validation is performed at key points in the content delivery pipeline and at regular intervals during content playback. This periodic approach ensures timely error detection while allowing computational resources to be released between monitoring cycles, optimizing the balance between reliability and resource consumption.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS8826354B2Method and system for testing closed caption content of video assets
Publication Date: 2014.09.02 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8826354B2 patent drawing
  • US8826354B2 patent drawing
  • US8826354B2 patent drawing

AI summary

A method and system for monitoring video assets provided by a multimedia content distribution network includes testing closed captions provided in output video signals. A video and audio portion of a video signal are acquired during a time period that a closed caption occurs. A first text string is extracted from a text portion of a video image, while a second text string is extracted from speech content in the audio portion. A degree of matching between the strings is evaluated based on a threshold to determine when a caption error occurs. Various operations may be performed when the caption error occurs, including logging caption error data and sending notifications of the caption error.