Document OCR Coordinate Region Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing optical character recognition (OCR) processes for extracting data fields from document images are resource-intensive and prone to errors due to the complexity of handling varying document formats and outlier data field positions, which increases processing time and requires manual correction.

Innovation Solution

A system dynamically tunes OCR processes by allowing users to input and confirm expected image coordinate areas for data fields, storing these corrections for future documents of the same type, enabling data field-specific OCR to automatically extract values with reduced manual intervention and processing time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data field-specific OCR processes are applied only to specific regions of document images, then processing power and time are significantly reduced, but the complexity of OCR process systems increases significantly as the volume and variety of documents increases

Engineering Contradiction:
Improveprocessing speedVSAvoidOCR system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the document image into multiple coordinate regions, each associated with specific data fields. The OCR system processes only the relevant coordinate regions for each document type rather than the entire image, reducing processing time and resource usage while maintaining accuracy through targeted region-specific OCR operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system pre-defines coordinate regions and data field locations for different document types before OCR processing occurs. By establishing these coordinate mappings in advance, the system eliminates the need for complex real-time analysis of document layouts, thereby reducing both processing time and system complexity during actual OCR operations.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If manual entry is used for outlier data fields in positions that are outliers to default or expected coordinate areas, then the processing system can capture all data fields, but user error increases and processing time is consumed

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidmanual entry time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system dynamically adjusts the coordinate regions for outlier data fields based on actual document content rather than relying on fixed default positions. When data fields appear in unexpected locations, the system adapts by identifying and creating new coordinate regions on-the-fly, eliminating the need for manual intervention while maintaining high accuracy for both standard and outlier data fields.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If general OCR processes are applied to all document images, then all data fields can be identified, but the processing is resource-intensive and time-consuming

Engineering Contradiction:
Improvedocument type coverageVSAvoidprocessing resource usage
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent applies different OCR processing strategies to different coordinate regions based on their specific characteristics and data field types. Instead of using a uniform general OCR process across the entire document, the system tailors OCR parameters and methods to each coordinate region, improving processing efficiency while maintaining the ability to handle diverse document types through region-specific optimization.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10346702B2Image data capture and conversion
Publication Date: 2019.07.09 BANK OF AMERICA CORP
  • US10346702B2 patent drawing
  • US10346702B2 patent drawing
  • US10346702B2 patent drawing

AI summary

Embodiments of the present invention provide a system for image capture and conversion. The system receives or captures an image of a resource document comprising image coordinates. The system can then cause a user interface of a computing device to display the image of the resource document and request that a specialist provide an input of an image coordinate area associated with a data field of the resource document. The specialist then provides a selection of boundaries for the coordinate area that encloses a value of the data field within the image of the resource document. The system then applies a data field-specific OCR process to the provided image coordinate area to extract a value of the data field. The extracted value can be presented on the display along with an enlarged view of the image coordinate area to allow the specialist to verify the accuracy of the extracted value.