Automated Document Feature Modeling and Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current electronic document processing systems face challenges in accurately extracting and analyzing information from scanned documents, with typical accuracy around 50%, leading to difficulties in fulfilling contractual obligations and generating business insights, due to manual and time-intensive processes involving human analysts and specialized programming.

Innovation Solution

A user-customizable system for automated document feature modeling and extraction allows users to create and edit data models and extraction rules through a graphical interface, streamlining the process and eliminating the need for human analysts and programming skills, enabling faster template adoption and improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual processes involving human analysts and specialized programming are used, then document feature extraction can be performed, but the process is time-intensive and slow

Engineering Contradiction:
Improvedocument processing speedVSAvoidtime required for manual analysis
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system enables automated self-service document feature extraction using machine learning models that automatically identify and extract relevant information from documents without requiring manual human analysts or specialized programming, thereby significantly reducing processing time and increasing productivity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes (human analysts reviewing documents) with automated electronic systems using machine learning and natural language processing algorithms, transforming the extraction process from labor-intensive to computation-intensive, which dramatically improves speed and efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated feature extraction is implemented, then processing speed improves, but accuracy remains limited to around 50%

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts extraction parameters and model configurations based on document characteristics and feedback from verification processes, optimizing the balance between processing speed and extraction accuracy for different document types and contexts

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent incorporates feedback loops where extraction results are verified and validated, with performance metrics fed back into the system to continuously improve model accuracy and reduce errors while maintaining high processing speeds through iterative optimization

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If custom data models are created to improve extraction accuracy, then the system becomes more adaptable, but the complexity of model creation and maintenance increases

Engineering Contradiction:
Improvecustomization capabilityVSAvoiddata model complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs universal data model templates and pre-configured extraction patterns that can be applied across multiple document types and use cases, reducing the need for complex custom models while maintaining adaptability through parameter customization rather than structural complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11048762B2User-defined automated document feature modeling, extraction and optimization
Publication Date: 2021.06.29 OPEN TEXT CORPORATION
  • US11048762B2 patent drawing
  • US11048762B2 patent drawing
  • US11048762B2 patent drawing

AI summary

Provided herein are systems and methods for user-defined automated document feature modeling, extraction and optimization. In the present disclosure, an end user of an automated document review system can customize and create new data models applicable to a set of focus documents. In addition, an end user of the automated document review system can customize and create new extraction rules applicable to text extraction from the set of focus documents. The user-defined edits to the data model and extraction rules can be further tested in a staging environment, and tested against a ground truth set of documents, before being widely applied to other relevant documents.