OCR-Based Video Conformance Using Burn-In Metadata Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In the post-production process of film and television, identifying and matching offline video content with the original camera source frames is a time-consuming and error-prone manual process, especially in marketing campaigns where metadata is stripped during transfers among different partners, making it difficult to recreate high-quality promotional clips from lower resolution 'offline' copies.

Innovation Solution

A method using optical character recognition (OCR) to detect shot boundaries and parse visible character burn-ins from video frames, generating a metadata-rich edit decision list that maps offline video content to the original camera source frames, thereby eliminating the need for manual listing and reducing errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual listing and breakdown processes are used to match offline video content with original camera source frames, then the process can be completed with human judgment and correction, but the process is time-consuming and error-prone

Engineering Contradiction:
Improveaccuracy in matching offline video content with source framesVSAvoidtime required for conformance process
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of listing and breakdown with an automated optical character recognition (OCR) system. The OCR software automatically detects and reads the burned-in metadata from video frames, eliminating the need for manual transcription and matching while significantly reducing errors and time requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables the video content itself to provide the necessary metadata through the burned-in information that is already present in the offline content. The OCR process extracts this self-contained metadata directly from the video frames without requiring external databases or manual intervention, allowing the content to serve its own conformance needs.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If metadata is stripped during transfers among different partners in marketing campaigns, then the offline content can be freely exchanged and edited, but it becomes difficult to map the final cut back to the original high-quality source frames

Engineering Contradiction:
Improveability to exchange and edit offline content among partnersVSAvoidloss of source metadata needed for conformance
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent creates a digital copy of the metadata by extracting the burned-in information from the offline video frames through OCR. This copied metadata is then stored and used to map the edited offline content back to the original source frames, preserving the connection even when the original metadata is stripped during partner exchanges.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary extraction and storage of metadata from the offline content before it is potentially stripped during subsequent transfers. By capturing the burned-in information early in the process, the system ensures that the mapping data is preserved even when later exchanges remove metadata from the video files.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated OCR processes are used to detect shot boundaries and parse burn-in information, then the conformance process becomes faster and more accurate, but the device complexity and processing requirements increase

Engineering Contradiction:
Improvespeed of conformance processVSAvoidcomplexity of OCR and shot boundary detection system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts only the essential metadata information from the video frames using OCR, focusing specifically on the burned-in timecode and scene information. This selective extraction approach simplifies the processing requirements compared to analyzing entire video frames, reducing the computational complexity while maintaining high productivity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments the video analysis process into distinct stages: shot boundary detection, burn-in area identification, character recognition, and metadata extraction. This segmentation allows each component to be optimized independently and processed efficiently, reducing overall system complexity while enabling parallel processing to maintain high productivity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11024341B2Conformance of media content to original camera source using optical character recognition
Publication Date: 2021.06.01 CO 3 METHOD INC
  • US11024341B2 patent drawing
  • US11024341B2 patent drawing
  • US11024341B2 patent drawing

AI summary

A clip of shots is uploaded to a conformance platform. The conformance platform evaluates the clip type and initiates shot boundary evaluation and detection. The identified shot boundaries are then seeded for OCR evaluation and the burned in metadata is extracted into categories using a custom OCR module based on the location of the burn-ins within the frame. The extracted metadata is then error corrected based on OCR evaluation of the neighboring frame and arbitrary frames at pre-computed timecode offsets from the frame boundary. The error corrected metadata and categories are then packaged into a metadata package and returned back to a conform editor. The application then presents the metadata package as an edit decision list with associated pictures and confidence level to the user. The user can further validate and override the edit decision list if necessary and then use it to directly to conform to the online content.