Template-Free Data Extraction Using Context Rules

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Variations in document layouts and designs hinder efficient automatic extraction and transfer of data from documents, requiring manual input into applications, especially when dealing with diverse business and healthcare documents.

Innovation Solution

A system processes text from documents by applying rules to determine context, extracting data without templates, and enabling its use across applications without manual input, allowing for user modifications and rule updates to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If template-based data extraction is used, then data extraction accuracy is improved for standardized documents, but adaptability to varied document layouts deteriorates

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidadaptability to varied document layouts
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs self-learning by automatically analyzing document layouts and extracting data patterns without requiring manual template creation. The machine learning model trains itself on input documents, enabling adaptive data extraction across varied layouts while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts extraction parameters and model configurations based on the specific document layout characteristics detected during analysis, allowing it to adapt to different document types while maintaining high extraction accuracy through parameter optimization.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If manual data entry is required, then data extraction reliability is improved through user verification, but productivity deteriorates due to time-consuming manual input

Engineering Contradiction:
Improvedata extraction reliabilityVSAvoiddata extraction speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements feedback mechanisms where extracted data is verified against multiple sources and cross-checked using machine learning confidence scores. User corrections are fed back into the system to continuously improve extraction accuracy, maintaining reliability while reducing the need for manual verification.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces manual mechanical data entry with automated machine learning-based extraction systems, substituting human operators with intelligent algorithms that can process documents at high speed while maintaining reliability through multiple verification layers.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If custom templates are created for each document type, then data extraction precision is improved, but device complexity increases due to template management

Engineering Contradiction:
Improvedata extraction precisionVSAvoidtemplate management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system employs a universal machine learning model that can handle multiple document types and layouts through a single unified framework. This multi-functional approach eliminates the need for separate custom templates for each document type, reducing management complexity while maintaining extraction precision through adaptive learning.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The extraction system dynamically adapts its parameters and model structure based on the input document characteristics, transitioning from static template-based approaches to dynamic machine learning models that automatically adjust to different document types without requiring separate template configurations.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10366123B1Template-free extraction of data from documents
Publication Date: 2019.07.30 INTUIT INC
  • US10366123B1 patent drawing
  • US10366123B1 patent drawing
  • US10366123B1 patent drawing

AI summary

The disclosed embodiments provide a system that processes data. One example embodiment is a computer-implemented method for processing data. The computer-implemented method includes obtaining text from a document associated with a user, wherein the document was generated based on a template and, with the obtained text intact, applying a set of rules to each term in the obtained text to determine a broad category of a plurality of terms associated with the term. The computer-implemented method further includes applying an additional set of rules to refine the broad category associated with the term to a refined category of fewer terms based on a location in the document of at least one term in the broad category of the plurality of terms, extracting a term from the obtained text using template-independent code developed to process documents generated based on a plurality of templates and enabling use of the term with an application.