Unstructured Document Extraction with ML Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current information extraction methods from unstructured documents are inefficient and require extensive human labor, and existing systems fail to accurately match information across documents and databases due to discrepancies caused by human error or inconsistent updates.

Innovation Solution

A customizable information extraction software that uses machine learning models trained by user-provided documents and metadata, employing a spatial scoring procedure and rich feedback interface to identify attribute-values without assuming document structure, and integrates with databases to highlight discrepancies for user correction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If machine learning models are used to automatically extract information from unstructured documents, then productivity and efficiency are improved, but the system requires extensive customization to each business's specific documents and domain knowledge

Engineering Contradiction:
Improveinformation extraction efficiencyVSAvoidsystem customization complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system allows domain experts to train custom machine learning models by providing example documents and desired output formats. The platform automatically handles model training, optimization, and deployment, enabling businesses to create customized extraction systems without requiring expertise in machine learning engineering or system architecture.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces a specialized platform that acts as an intermediary between domain experts and machine learning systems. This platform translates domain knowledge into trained models through automated workflows, bridging the gap between business requirements and technical implementation while abstracting away complex technical details.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If information is extracted from unstructured documents, then data availability is improved, but discrepancies and errors may exist between extracted information and existing database information

Engineering Contradiction:
Improveinformation accessibilityVSAvoiddata consistency
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system implements a feedback mechanism where extracted information is compared against existing database records, and discrepancies are presented to users for review and correction. User corrections feed back into the system to improve future extraction accuracy and resolve consistency issues between documents and databases.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies preliminary verification steps where the system proactively identifies and flags potential discrepancies between extracted information and existing database records before finalizing the data integration. This prevents erroneous data from being silently accepted and maintains data reliability.

Inventive Principle:
Principle #9Preliminary anti-action

3Measurement precision

If manual review and extraction of information from documents is performed, then accuracy can be verified, but extensive human labor is required which lowers efficiency and productivity

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidhuman labor efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system applies machine learning models to perform the bulk of information extraction automatically, requiring human intervention only for review and correction of uncertain extractions. This partial automation approach maintains high accuracy while dramatically improving productivity by eliminating the need for complete manual review of all documents.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10521464B2Method and system for extracting, verifying and cataloging technical information from unstructured documents
Publication Date: 2019.12.31 AGILE DATA DECISIONS LLC
  • US10521464B2 patent drawing
  • US10521464B2 patent drawing
  • US10521464B2 patent drawing

AI summary

Information extraction methods for use in extracting values from unstructured documents for predetermined or user-specified attributes into structured databases are provided herein. Methods include (a) automatically training machine learning models for extracting values from unstructured documents such that the values of the attributes are known for those training documents but the locations of the values in the documents are not known, (b) making a sustained connection between structured databases and unstructured documents so that the data across those two types of data stores can be cross-referred by the users any time, (c) a graphical interface specialized for rich user feedback to rapidly adapt and improve the machine learning models. The methods allow businesses and other entities or institutions to apply their domain knowledge to train software for extracting information from their documents so that the software becomes customized to those documents both from initial training as well as continuing user feedback.