Content Information Auditing Service Using OCR Error Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Maintaining accurate and comprehensive information about media works, such as cast and crew, is challenging due to the vast amount of data and the difficulty in extracting relevant information from media content, leading to errors and inefficiencies in auditing and identification processes.

Innovation Solution

A content information auditing service that utilizes optical character recognition (OCR) to identify words in media content, filters and corrects errors using predefined rules, and updates information by comparing identified words to a database, providing a user interface for users to interact and correct missing information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual methods are used to maintain media work information, then data accuracy can be maintained through human review, but the process is time-consuming and error-prone due to the vast amount of data

Engineering Contradiction:
Improvedata accuracyVSAvoidtime-consuming process
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual human review processes with automated optical character recognition (OCR) technology to extract information from media content. This substitution of mechanical/automated systems for human manual work enables rapid processing of vast amounts of data while maintaining accuracy through systematic error detection and correction mechanisms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-correction by automatically detecting and correcting errors in extracted information without requiring constant human intervention. The automated process identifies inconsistencies and rectifies them, enabling the system to maintain data accuracy independently while reducing time consumption.

Inventive Principle:
Principle #25Self-service

2Loss of information

If information is extracted from content segments, then comprehensive media work information can be obtained, but the extraction process becomes complex and error-prone

Engineering Contradiction:
Improvecomprehensive information coverageVSAvoidextraction process complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent divides the media content into discrete segments or frames, applying OCR technology to each segment individually. This segmentation approach enables comprehensive information extraction from different parts of the content while simplifying the overall extraction process by breaking it down into manageable, standardized units that can be processed systematically.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If comprehensive cast and crew information is maintained for media works, then complete media work profiles can be created, but data maintenance becomes difficult and error-prone

Engineering Contradiction:
Improvecompleteness of media work informationVSAvoiddata maintenance accuracy
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system incorporates feedback mechanisms where extracted information is continuously verified, cross-checked, and validated against existing data. This feedback loop enables the system to maintain comprehensive cast and crew information while ensuring high reliability by automatically detecting and correcting errors, inconsistencies, or incomplete data through iterative verification processes.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Ensures accurate and complete media work information by filtering out false positives, correcting errors, and completing missing data, enhancing the accuracy and reliability of media content databases.

Implementation Method 1

utilizes optical character recognition (OCR) to identify words in media content

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS10140518B1Content information auditing service
Publication Date: 2018.11.27 IMDBCOM INC
  • US10140518B1 patent drawing
  • US10140518B1 patent drawing
  • US10140518B1 patent drawing

AI summary

Systems and methods are provided for auditing content information for a media work. In embodiments, content information for a plurality of media works that identifies entities associated with each media work may be maintained. In an embodiment, a request to identify a particular media work may be received. One or more words included in a segment of the particular media work may be identified where the segment is configured to be presented. In accordance with at least one embodiment, the one or more identified words may be filtered based on a set of rules to correct errors. An identity of the particular media work may be determined based at least in part on the filtered one or more words and the content information for the plurality of media works.