Document Classification Using Spatial Relations and OCR Templates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document processing systems face challenges in efficiently grouping and classifying scanned paper documents due to optical character recognition (OCR) errors, text differences, graphics, noise, rotations, skewing, and handwriting, which hinder the transformation of information into intelligent content for enterprise applications.

Innovation Solution

A client-server system utilizing a training module, classification module, and location comparison engine that compares documents using textual content and spatial relations of words to create document templates and classify documents, employing a textual distance function to determine document similarity and group them effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional OCR-based document processing is used, then document digitization is achieved, but grouping and classification accuracy deteriorates due to OCR errors, text differences, and layout variations

Engineering Contradiction:
Improvedocument classification accuracyVSAvoidgrouping consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent creates template documents representing ideal document structures and uses these templates as reference copies for comparison. Instead of comparing flawed OCR data directly, the system compares scanned documents against clean template copies, thereby maintaining classification accuracy despite OCR errors and variations in the source documents

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces template documents as an intermediary layer between the scanned documents and the classification process. These templates serve as mediators that bridge the gap between variable scanned documents and the classification system, enabling accurate grouping by comparing documents to templates rather than directly to each other

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If strict text matching is used for document classification, then classification precision is improved, but the system becomes sensitive to OCR errors and text variations

Engineering Contradiction:
Improvetext matching accuracyVSAvoidtolerance to document variations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies different comparison strictness levels to different parts of the document. Critical fields (like document type identifiers) are matched strictly against template text, while variable fields (like specific data values) are allowed to vary. This local differentiation of quality requirements enables both precision and tolerance to variations

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the matching parameters dynamically based on the document being processed. Instead of using a fixed threshold, the system adjusts matching criteria to account for expected OCR error rates and document-specific characteristics, thereby maintaining adaptability while preserving classification precision

Inventive Principle:
Principle #35Parameter changes

3Reliability

If manual document review is performed to ensure accurate classification, then classification reliability is improved, but processing productivity deteriorates

Engineering Contradiction:
Improveclassification accuracyVSAvoiddocument processing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent enables the system to automatically correct its own classification decisions by comparing documents against templates and using confidence scoring. The system self-validates classifications without requiring manual review, thereby maintaining high reliability while achieving automated processing speeds

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements a feedback mechanism where classification results are continuously validated against template expectations and confidence thresholds. Documents that fail validation are automatically reprocessed or flagged for review, creating a self-correcting system that maintains accuracy without requiring universal manual review

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If comprehensive document analysis is performed to handle all document variations, then classification robustness is improved, but system complexity increases

Engineering Contradiction:
Improvehandling of document variationsVSAvoidprocessing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the document analysis process into distinct stages: template creation, feature extraction, comparison, and classification. By dividing the comprehensive analysis into modular segments, the system handles document variations robustly while keeping each processing stage manageable and the overall system complexity controlled

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8595235B1Method and system for using OCR data for grouping and classifying documents
Publication Date: 2013.11.26 OPEN TEXT CORP
  • US8595235B1 patent drawing
  • US8595235B1 patent drawing
  • US8595235B1 patent drawing

AI summary

Document classes for classifying documents are created by comparing the spatial relations of words between a first and second document. If the spatial relations are the same, a document class may be created to classify documents similar to the first and second document. If the spatial relations are different, a first document class may be created to classify documents similar to the first document, and a second document class may be created to classify documents similar to the second document.