Document Entity Extraction with OCR Error Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing OCR technologies face difficulties in accurately extracting entity data from scanned documents due to errors caused by font characteristics, formatting issues, and blurring of text during scanning, leading to incorrect recognition of characters.

Innovation Solution

The system employs data extractors guided by an extraction model to identify entity data, which is then processed by experts applying business rules to organize and validate the data, utilizing OCR correction and statistical analyses to improve accuracy and adaptability across different document layouts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If OCR technologies are used to extract entity data from scanned documents, then the extraction process can be automated, but recognition accuracy deteriorates due to font characteristics, formatting issues, and blurring during scanning

Engineering Contradiction:
Improveautomation of entity data extractionVSAvoidcharacter recognition accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent implements a feedback mechanism where extracted entity data is validated against business rules and domain knowledge. The system identifies and corrects OCR errors by comparing extracted data with expected patterns and using statistical analyses to detect and fix recognition mistakes, thereby improving accuracy while maintaining automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary validation layer between OCR extraction and final data output. This intermediary layer includes business rule engines and statistical analysis modules that mediate between raw OCR output and final extracted entities, correcting errors and ensuring data quality before delivery to end users.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If standard OCR extraction is applied to documents with various layouts, then processing speed is maintained, but extraction reliability deteriorates due to layout variations and formatting errors

Engineering Contradiction:
Improveprocessing speedVSAvoidextraction reliability across different layouts
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent dynamically adjusts extraction parameters based on document layout detection and business rule configurations. The system adapts its extraction strategy by changing parameters such as entity type expectations, validation rules, and extraction patterns to match the specific layout and formatting of each document, thereby maintaining high reliability across diverse document types without significantly reducing processing speed.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manual correction of OCR errors is performed, then recognition accuracy is improved, but processing time and labor requirements increase

Engineering Contradiction:
Improvecharacter recognition accuracyVSAvoidprocessing time for error correction
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service error correction where the system automatically identifies and corrects OCR errors using business rules and statistical analyses. The extraction process includes built-in validation that detects and corrects common OCR mistakes without requiring manual intervention, allowing the system to correct its own errors and maintain high accuracy while minimizing time loss.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10755093B2Hierarchical information extraction using document segmentation and optical character recognition correction
Publication Date: 2020.08.25 OPEN TEXT CORPORATION
  • US10755093B2 patent drawing
  • US10755093B2 patent drawing
  • US10755093B2 patent drawing

AI summary

Systems, methods, and media for extracting and processing entity data included in an electronic document are provided herein. Methods may include executing one or more extractors to extract entity data within an electronic document based upon an extraction model for the document, selecting extracted entity data via one or more experts, each of the experts applying at least one business rule to organize at least a portion of the selected entity data into a desired format, and providing the organized entity data for use by an end user.