Document Extraction System Using Self-Adapting N-Gram Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current document processing systems face challenges in automatically extracting and classifying information from weakly structured documents, particularly in identifying specific fields and correcting OCR errors, due to variability in document formats and languages.

Innovation Solution

The extraction system employs a self-adapting and learning approach using N-gram features, contextual analysis, and statistical scoring methods to identify and extract target fields from documents, with modules for image processing, OCR correction, and validation, enabling iterative refinement and handling of multiple languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional document processing systems are used to extract information from weakly structured documents, then the system structure is simple, but the extraction accuracy is low due to variability in document formats and languages

Engineering Contradiction:
Improveextraction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system dynamically adapts to different document formats and languages by learning from training documents. The extraction module adjusts its parameters and models based on the specific characteristics of each document type, enabling accurate extraction from weakly structured documents with varying formats and languages without requiring a completely different system for each case.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs self-training and self-adaptation by automatically learning from provided training documents. The extraction module improves its own accuracy through iterative learning processes, adjusting its internal models and parameters without external intervention, thereby handling the complexity of varied document formats and languages autonomously.

Inventive Principle:
Principle #25Self-service

2Productivity

If manual methods are used for field identification and OCR error correction, then the accuracy is high, but the processing time and labor cost are excessive

Engineering Contradiction:
Improveprocessing speedVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where extraction results are validated and used to refine the extraction models. The system learns from correct and incorrect extractions, adjusting its parameters to improve accuracy over time while maintaining high processing speeds. This feedback loop enables the automated system to approach manual accuracy levels without the associated time and labor costs.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary training and model adjustment using training documents before actual extraction tasks. By pre-learning the characteristics of different document formats and languages, the system is prepared to accurately extract information from new documents, reducing the need for post-processing corrections and improving overall processing efficiency.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the system processes multiple languages and formats, then the versatility is improved, but the complexity of handling variability increases

Engineering Contradiction:
Improvemulti-language supportVSAvoidhandling complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The extraction module is designed as a universal system that can handle multiple languages and document formats through a single integrated architecture. By implementing language-agnostic feature extraction and using training data from various languages and formats, the system achieves multi-functionality without requiring separate processing pipelines for each language or format type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system manages variability by dynamically adjusting its extraction parameters and models based on the detected document language and format. Through parameter adaptation rather than structural complexity, the system handles diverse document types efficiently, changing its processing characteristics to match the input document properties without increasing overall system complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8321357B2Method and system for extraction
Publication Date: 2012.11.27 HYLAND SWITZERLAND SARL
  • US8321357B2 patent drawing
  • US8321357B2 patent drawing
  • US8321357B2 patent drawing

AI summary

A system and method for extracting information from at least one document in at least one set of documents, the method comprising: generating, using at least one ranking and/or matching processor, at least one ranked possible match list comprising at least one possible match for at least one target entry on the at least one document, the at least one ranked possible match list based on at least one attribute score and at least one localization score.