Document Data Extraction via OCR Anchor Blocks and Confidence Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in efficiently extracting and processing data from voluminous semi-structured and unstructured documents, which are often large and varied, making it difficult to identify relevant information within these documents.

Innovation Solution

A computer system that uses optical character recognition (OCR) and post-OCR processing to identify words and contender values, grouping them into anchor blocks based on spatial positioning, and determining confidence levels through comparisons with pre-defined values, formats, and keywords to extract positively associated values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If optical character recognition (OCR) and post-OCR processing are used to identify words and contender values from documents, then data extraction accuracy is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The document processing is divided into multiple stages: OCR processing to identify contender values, spatial relationship analysis to determine anchor blocks, and confidence level calculation to filter results. This segmentation allows each stage to focus on specific tasks, improving overall accuracy while managing processing time through staged computation rather than monolithic processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary OCR processing and spatial relationship analysis before final confidence level determination. By pre-identifying contender values and their spatial relationships to anchor blocks, the system can efficiently calculate confidence levels without re-processing all document data, thus improving accuracy while reducing total processing time.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If documents are grouped into anchor blocks based on spatial positioning, then data organization and retrieval efficiency are improved, but system complexity increases

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system calculates spatial relationships between contender values and anchor blocks, assigning different weights to different spatial positions. Values closer to anchor blocks receive higher weights, creating local quality variations in the data organization. This approach improves retrieval efficiency by prioritizing relevant data while managing complexity through localized processing rather than global reorganization.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses configurable parameters including spatial relationship thresholds, confidence level thresholds, and weight factors to control the anchoring process. By adjusting these parameters, the system can optimize the balance between data organization quality and system complexity, allowing flexible adaptation to different document types and retrieval requirements without hardcoding complex logic.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If confidence levels are calculated based on spatial relationships and keyword comparisons, then data extraction reliability is improved, but processing complexity increases

Engineering Contradiction:
Improvedata extraction reliabilityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges multiple reliability indicators into a single confidence level calculation: spatial relationship strength, keyword matches, and positional information are combined through weighted averaging. This merging approach improves overall reliability by considering multiple factors simultaneously while managing complexity through a unified calculation framework rather than separate validation systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The confidence level calculation serves multiple functions: it filters extracted data, ranks contender values, and determines inclusion in final results. By making the confidence calculation multi-functional, the system improves reliability through comprehensive validation while reducing overall processing complexity by eliminating the need for separate validation and ranking systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables efficient extraction and identification of relevant data from large document collections, improving data location accuracy and reducing the need for manual input, allowing analysts to quickly find specific information within haystacks of data.

Implementation Method 1

The one or more hardware computer processors can be configured to execute the one or more software modules in order to cause the computer system to, for each page of the compilation, identify words and contender values on the subject page using optical character recognition (OCR) and post-OCR processing

Methodology Applied
Scientific EffectOptical character recognition (OCR):

Data Source

PatentUS9384264B1Analytic systems, methods, and computer-readable media for structured, semi-structured, and unstructured documents
Publication Date: 2016.07.05 TUNGSTEN AUTOMATION CORPORATION
  • US9384264B1 patent drawing
  • US9384264B1 patent drawing
  • US9384264B1 patent drawing

AI summary

A computer system extracts contender values as positively associated with a pre-defined value from a compilation of one or more electronically stored semi-structured document(s) and/or one or more electronically stored unstructured document(s). The computer system performs a multi-dimensional analysis to narrow the universe of contender values from all words on a page of the compilation to the contender value(s) with the highest likelihood of being associated with the pre-defined value. The system's platform allows every user of the system to customize the system according to the user's needs. Various aspects can enable users to mine document stores for information that can be charted, graphed, studied, and compared to help make better decisions.