Machine Learning Pre-fill Engine for Document Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automated techniques for extracting data from documents are not designed to analyze and extract specific fields related to business processes, such as insurance, and require continuous human effort and maintenance, leading to inefficiencies and potential biases.

Innovation Solution

A system utilizing a machine learning algorithm with a pre-fill engine that receives electronic documents, seed dataset documents, and pre-fill questions to determine output data and questions relevant to a particular field of analysis, using terminology, categories, and ontology, and presenting answers through a graphical user interface.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised machine learning techniques are used to extract data from documents, then extraction accuracy can be improved, but continuous human effort and maintenance are required, increasing cost and time consumption

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidtime consumption for maintenance
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically creating seed datasets and training machine learning models in advance. The pre-fill engine generates initial data extracts and questions before human review, so that when documents are processed, the system already has structured frameworks and examples to guide accurate extraction without requiring continuous manual retraining

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service by allowing the machine learning model to autonomously create seed datasets, generate pre-fill questions, and perform data extraction without continuous human intervention. The model uses its own outputs to refine future extractions, creating a self-improving system that reduces dependency on ongoing manual maintenance while maintaining high accuracy

Inventive Principle:
Principle #25Self-service

2Productivity

If existing automated techniques are used for data extraction, then processing speed can be improved, but the techniques are not designed to analyze specific fields related to business processes, leading to incorrect extraction

Engineering Contradiction:
Improveprocessing speedVSAvoidfield-specific extraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system applies local quality by customizing the machine learning model to understand and extract specific fields relevant to particular business processes. Instead of using generic extraction techniques, the system trains domain-specific models that learn the nuances of insurance terminology, document structures, and field relationships, enabling both high speed and high accuracy for field-specific extraction

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system introduces an intermediary layer in the form of a pre-fill engine that acts as a mediator between the raw documents and the final extracted data. This intermediary component translates domain-specific business process requirements into machine-understandable queries and validation rules, enabling the system to maintain both speed and precision by filtering and guiding the extraction process through structured intermediaries

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If human effort is used to create seed datasets for machine learning, then extraction accuracy can be improved, but the process is expensive and error-prone

Engineering Contradiction:
Improveextraction accuracyVSAvoidease of dataset creation
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system implements self-service by enabling the machine learning model to autonomously create seed datasets without human intervention. The pre-fill engine automatically generates training data by processing documents and creating structured examples that the model can learn from, eliminating the expensive and error-prone manual annotation process while maintaining high extraction accuracy through automated data synthesis

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses copying by having the pre-fill engine create synthetic seed datasets that replicate the structure and content patterns of real business documents. These copied datasets serve as training material for the machine learning model, providing accurate extraction examples without requiring humans to manually create and verify each training sample, thus reducing costs and errors

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11580459B2Systems and methods for extracting specific data from documents using machine learning
Publication Date: 2023.02.14 CONVR INC
  • US11580459B2 patent drawing
  • US11580459B2 patent drawing
  • US11580459B2 patent drawing

AI summary

Computer implemented systems and methods are disclosed for extracting specific data using machine learning algorithms. In accordance with some embodiments, a memory device that stores at least a set of computer executable instructions for a machine learning algorithm and a pre-fill engine; and at least one processor that executes the instructions that cause the pre-fill engine to perform functions that include: receiving electronic documents, seed dataset documents, and pre-fill questions; determining output data that enable navigation through the electronic documents using the machine learning algorithm; determining output questions that enable navigation through the electronic documents using the machine learning algorithm; determining output documents to enable navigation through the electronic documents using the machine learning algorithm; and presenting one or more answers for one or more of the output questions using a graphical user interface.