Document Data Extraction via Bounding Polygon Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data extraction methods, such as generic extractors, often fail to accurately extract data from documents with complex or irregularly shaped fields, leading to incomplete or incorrect digitization.

Innovation Solution

A specialized data extraction method using machine learning models to identify bounding polygons within complex fields, perform optical character recognition, and correct extracted data, which is then combined with data from standard fields to generate accurate input for data processing applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a generic extractor is used to digitize documents, then most fields can be extracted, but complex or irregularly shaped fields cannot be correctly extracted

Engineering Contradiction:
Improveability to handle different field typesVSAvoiddata extraction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the document processing into two distinct pathways: a generic extractor for standard fields and a specialized extractor for complex fields. The system divides the extraction task based on field characteristics, routing simple fields through the generic path while directing complex, irregular, or abnormal fields through the specialized path that uses bounding polygon identification and machine learning models.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different extraction qualities to different parts of the document. Standard fields receive generic extraction treatment, while complex fields receive specialized extraction treatment with higher precision requirements. The system identifies fields with abnormal shapes or complexities and applies targeted processing with bounding polygons and revised extraction logic specifically to those regions.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If a specialized extraction method is used for complex fields, then extraction accuracy improves, but processing complexity increases

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidextraction system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The extraction system is segmented into a generic extractor component and a specialized extractor component. The generic extractor handles the majority of standard fields with simple processing, while the specialized extractor handles only the minority of complex fields requiring bounding polygon identification and machine learning models, thus managing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies the complex specialized extraction process only partially - specifically to fields that are identified as having abnormal shapes, irregular boundaries, or extraction difficulties. The majority of fields undergo simpler generic extraction, avoiding the overhead of complex processing where it is not needed.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If manual verification is performed to ensure extraction accuracy, then data quality improves, but processing time increases

Engineering Contradiction:
Improvedata qualityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs self-verification through automated mechanisms rather than requiring manual review. The specialized extractor uses machine learning models to automatically identify and correct extraction errors in complex fields, generating revised extracted data that is validated through the model's prediction capabilities, thus maintaining high reliability without manual intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback loops where the machine learning model analyzes extracted data, identifies potential errors or inconsistencies, and generates revised extracted data accordingly. This automated feedback mechanism continuously improves extraction accuracy for complex fields without requiring manual verification, maintaining both high reliability and processing speed.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11651606B1Method and system for document data extraction
Publication Date: 2023.05.16 INTUIT INC
  • US11651606B1 patent drawing
  • US11651606B1 patent drawing
  • US11651606B1 patent drawing

AI summary

Certain aspects of the present disclosure provide techniques for extracting data from a document. An example method generally includes identifying a bounding polygon of the region from an electronic image of the document and extracting data from within the bounding polygon of the region. The method further includes generating revised extracted data based on the extracted data, and combining the revised extracted data with other data extracted from the electronic image of the document to generate input data for a data processing application.