Automated Document State Assignment via OCR Text Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual determination of document state in electronic document exchange leads to inconsistent classifications due to varying interpretations and limitations in transmission speed, especially in bulk facsimile messaging environments, resulting in inaccurate urgency or routing determinations.
Innovation Solution
A method for automated fuzzy document state assignment using optical character recognition (OCR) to process and segment text from raster images, generating an index and computing probabilities for classification based on word combinations, with annotations displayed and transmitted for second-level review, ensuring consistent state determination across incremental document receipt.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual determination of document state is used, then individual knowledge and experience can be applied, but inconsistent classifications and inaccurate urgency or routing determinations result
Solution Approach 1:
The patent replaces the manual mechanical review process with an automated optical character recognition (OCR) system that converts document images to text, segments the text, and applies rule-based classification. This substitution eliminates human variability and provides consistent, accurate document state determination across all documents.
Solution Approach 2:
The system enables documents to self-classify by automatically processing their own content through OCR and text segmentation, then applying classification rules to determine their own state. This self-service approach ensures consistent classification without requiring manual review while maintaining high accuracy through automated rule application.
2Measurement precision
If complete document receipt is required for state determination, then accurate classification can be achieved, but processing delays occur in bulk facsimile messaging environments
Solution Approach 1:
The patent performs preliminary OCR processing and text segmentation on documents as they are being received in bulk facsimile messaging environments. By preparing the text data in advance and applying classification rules incrementally, the system can determine document state before complete receipt, eliminating processing delays while maintaining accuracy through progressive analysis.
Solution Approach 2:
The system segments documents into individual pages or sections and processes each segment independently through OCR and text analysis. This segmentation allows parallel processing of multiple document portions simultaneously, enabling state determination to begin before the entire document is received, thus reducing overall processing time while maintaining classification accuracy.
3Productivity
If incremental document receipt is processed, then processing speed is improved, but inconsistent state determinations may occur
Solution Approach 1:
The patent implements feedback mechanisms where the classification system continuously monitors incoming document segments, applies classification rules, and adjusts processing based on accumulated information. This feedback loop ensures that incremental processing maintains consistency by refining state determinations as more document content becomes available, preventing inconsistent classifications.
Solution Approach 2:
The system dynamically adjusts its processing approach based on the completeness and content of received document segments. As documents are incrementally received, the classification rules are applied flexibly, allowing the system to make preliminary determinations when sufficient information is available while maintaining the ability to revise classifications as complete document content is processed, ensuring both speed and consistency.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach ensures uniform and accurate document state classification before the entire document is received, overcoming inconsistencies and speed limitations, enabling reliable urgency and routing decisions.
Implementation Method 1
performing optical character recognition (OCR) upon a page of a document in order to produce parseable text
Data Source
AI summary
Fuzzy document state assignment includes loading into memory a raster image of a document and performing OCR upon a page of a document in order to produce parseable text. The parseable text is then segmented and normalized and an index is generated from the segmented and normalized text. Thereafter, a probability of a particular classification is computed based upon the detection in the index of a combination of words associated with a corresponding classification. Finally, the document is annotated with the particular classification.


