OCR-Based Video Conformance Using Burn-In Metadata Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In the post-production process of film and television, identifying and matching offline video content with the original camera source frames is a time-consuming and error-prone manual process, especially in marketing campaigns where metadata is stripped during transfers among different partners, making it difficult to recreate high-quality promotional clips from lower resolution 'offline' copies.
Innovation Solution
A method using optical character recognition (OCR) to detect shot boundaries and parse visible character burn-ins from video frames, generating a metadata-rich edit decision list that maps offline video content to the original camera source frames, thereby eliminating the need for manual listing and reducing errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual listing and breakdown processes are used to match offline video content with original camera source frames, then the process can be completed with human judgment and correction, but the process is time-consuming and error-prone
Solution Approach 1:
The patent replaces the manual mechanical process of listing and breakdown with an automated optical character recognition (OCR) system. The OCR software automatically detects and reads the burned-in metadata from video frames, eliminating the need for manual transcription and matching while significantly reducing errors and time requirements.
Solution Approach 2:
The system enables the video content itself to provide the necessary metadata through the burned-in information that is already present in the offline content. The OCR process extracts this self-contained metadata directly from the video frames without requiring external databases or manual intervention, allowing the content to serve its own conformance needs.
2Adaptability or versatility
If metadata is stripped during transfers among different partners in marketing campaigns, then the offline content can be freely exchanged and edited, but it becomes difficult to map the final cut back to the original high-quality source frames
Solution Approach 1:
The patent creates a digital copy of the metadata by extracting the burned-in information from the offline video frames through OCR. This copied metadata is then stored and used to map the edited offline content back to the original source frames, preserving the connection even when the original metadata is stripped during partner exchanges.
Solution Approach 2:
The system performs preliminary extraction and storage of metadata from the offline content before it is potentially stripped during subsequent transfers. By capturing the burned-in information early in the process, the system ensures that the mapping data is preserved even when later exchanges remove metadata from the video files.
3Productivity
If automated OCR processes are used to detect shot boundaries and parse burn-in information, then the conformance process becomes faster and more accurate, but the device complexity and processing requirements increase
Solution Approach 1:
The patent extracts only the essential metadata information from the video frames using OCR, focusing specifically on the burned-in timecode and scene information. This selective extraction approach simplifies the processing requirements compared to analyzing entire video frames, reducing the computational complexity while maintaining high productivity.
Solution Approach 2:
The system segments the video analysis process into distinct stages: shot boundary detection, burn-in area identification, character recognition, and metadata extraction. This segmentation allows each component to be optimized independently and processed efficiently, reducing overall system complexity while enabling parallel processing to maintain high productivity.
Data Source
AI summary
A clip of shots is uploaded to a conformance platform. The conformance platform evaluates the clip type and initiates shot boundary evaluation and detection. The identified shot boundaries are then seeded for OCR evaluation and the burned in metadata is extracted into categories using a custom OCR module based on the location of the burn-ins within the frame. The extracted metadata is then error corrected based on OCR evaluation of the neighboring frame and arbitrary frames at pre-computed timecode offsets from the frame boundary. The error corrected metadata and categories are then packaged into a metadata package and returned back to a conform editor. The application then presents the metadata package as an edit decision list with associated pictures and confidence level to the user. The user can further validate and override the edit decision list if necessary and then use it to directly to conform to the online content.


