OCR Text Verification via Layer Cross-Reference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional optical character recognition (OCR) systems are prone to errors when processing low-quality images, leading to inaccuracies in detecting text from scanned documents, which requires manual verification by humans, resulting in tiresome and error-prone processes.

Innovation Solution

A method and system for automatically verifying OCR-detected text in native digital documents by determining the location of the text in both the image and text layers, comparing the detected text with the actual text layer content, and rendering only the accurate text as output when discrepancies are found, thereby improving the accuracy and efficiency of data extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If OCR is performed on low-quality scanned documents, then text detection can be achieved, but accuracy deteriorates due to errors in character recognition

Engineering Contradiction:
Improvetext detection efficiencyVSAvoidOCR accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary verification system that uses multiple OCR engines and cross-referencing mechanisms to validate detected text. The system compares results from different OCR engines and uses confidence scoring to identify and correct errors, thereby maintaining high accuracy while preserving processing efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback loops where detected text is verified against multiple sources and confidence thresholds. When errors are detected through cross-validation, the system automatically adjusts processing parameters or flags results for manual review, creating a self-correcting mechanism that maintains accuracy without sacrificing productivity.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If manual verification of OCR text is performed, then accuracy is improved, but time consumption and human error increase

Engineering Contradiction:
Improvetext verification accuracyVSAvoidverification time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial verification by focusing manual or enhanced computational review only on text regions with low confidence scores or high error probability. The system automatically verifies high-confidence text without additional time cost, while directing verification resources only where needed, thus reducing overall verification time while maintaining accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent implements automated self-verification mechanisms where the system uses multiple OCR engines and cross-referencing to automatically detect and correct errors without human intervention. The system serves itself by identifying inconsistencies and applying correction algorithms, eliminating the need for time-consuming manual verification while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

3Loss of information

If OCR processing is performed on documents with complex formatting and spacing, then comprehensive text extraction is achieved, but error rate increases

Engineering Contradiction:
Improvetext extraction completenessVSAvoidextraction accuracy
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent segments the document processing into distinct stages: initial OCR detection, formatting analysis, confidence scoring, and verification. By dividing the complex processing task into manageable segments, the system can apply specialized algorithms to each stage, improving overall reliability while maintaining complete text extraction from complex documents.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality control by adjusting verification thresholds and processing intensity based on the specific characteristics of different document regions. High-confidence regions are processed quickly with minimal verification, while low-confidence or complex formatting regions receive enhanced verification, optimizing both accuracy and efficiency across the entire document.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11232300B2System and method for automatic detection and verification of optical character recognition data
Publication Date: 2022.01.25 SUREPREP LLC
  • US11232300B2 patent drawing
  • US11232300B2 patent drawing
  • US11232300B2 patent drawing

AI summary

Systems and methods for automatically verifying optical character recognition (OCR) detected text of a native electronic document having an image layer comprising a matrix of pixels and a text layer comprising a sequence of characters. The method includes determining a location of OCR-detected text in the text layer of the native electronic document based on a pixel-based coordinate location of the OCR-detected text in the image layer of the native electronic document. The method also includes applying the location of the OCR-detected text to the text layer of the native electronic document to detect text in the text layer corresponding to the OCR-detected text. The method also includes rendering only the detected text in the text layer as an output when the OCR-detected text does not match the detected text in the text layer, to improve accuracy of the output text.