Continuous Machine Learning for Document Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for automated document data extraction require laborious manual review and rely on highly structured forms, limiting their ability to efficiently process large datasets of mixed structured, semi-structured, and unstructured text.

Innovation Solution

The implementation of a distributed continuous machine learning approach that utilizes a central repository of term-based machine learning models, incorporating optical character recognition and template-based extraction, to create scalable document processing pipelines capable of handling tens to hundreds of thousands of documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual review and highly structured forms are used for data extraction, then extraction accuracy is improved, but processing efficiency and scalability deteriorate

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables self-service through automated machine learning models that perform data extraction without manual review. The models are trained on annotated documents and automatically extract information from new documents, eliminating the need for continuous human intervention while maintaining extraction accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical review processes are replaced with automated machine learning systems. The patent substitutes human annotators and reviewers with trained ML models that process documents automatically, significantly improving processing efficiency while maintaining or enhancing extraction accuracy through continuous learning.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If traditional machine learning models are trained on annotated documents, then extraction accuracy is improved, but training time and resource requirements increase

Engineering Contradiction:
Improveextraction accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-annotating documents using existing ML models before full training begins. This pre-processing step creates initial training data that accelerates subsequent model training, reducing overall training time while maintaining extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous learning where models are continuously trained on new annotated documents without complete retraining. This continuous useful action allows the system to improve extraction accuracy over time while minimizing training time through incremental updates rather than periodic full retraining.

Inventive Principle:
Principle #20Continuity of useful action

3Productivity

If distributed continuous machine learning approach is implemented, then processing scalability is improved, but system complexity increases

Engineering Contradiction:
Improveprocessing scalabilityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the document processing task into multiple independent machine learning models, each trained on specific document types or extraction tasks. These segmented models can be distributed across multiple computing resources, improving scalability while managing complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal platform that handles multiple document types and extraction tasks through a common distributed ML architecture. This multi-functional system reduces overall complexity by providing a standardized approach that can be applied across different document types rather than requiring separate specialized systems for each.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Adaptability or versatility

If multiple machine learning models are used for different document types, then extraction versatility is improved, but model selection and management complexity increases

Engineering Contradiction:
Improveextraction versatilityVSAvoidmodel management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements feedback mechanisms that automatically evaluate which ML model is most appropriate for a given document based on document type, content analysis, and historical performance data. This feedback-driven model selection automates the complexity of managing multiple models, making the system versatile while simplifying model selection and management.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11645462B2Continuous machine learning method and system for information extraction
Publication Date: 2023.05.09 PWC PRODUCT SALES LLC
  • US11645462B2 patent drawing
  • US11645462B2 patent drawing
  • US11645462B2 patent drawing

AI summary

Methods and systems for artificial intelligence (AI)-assisted document annotation and training of machine learning-based models for document data extraction are described. The methods and systems described herein take advantage of a continuous machine learning approach to create document processing pipelines that provide accurate and efficient data extraction from documents that include structured text, semi-structured text, unstructured text, or any combination thereof.