Automated Document State Assignment via OCR Text Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual determination of document state in electronic document exchange leads to inconsistent classifications due to varying interpretations and limitations in transmission speed, especially in bulk facsimile messaging environments, resulting in inaccurate urgency or routing determinations.

Innovation Solution

A method for automated fuzzy document state assignment using optical character recognition (OCR) to process and segment text from raster images, generating an index and computing probabilities for classification based on word combinations, with annotations displayed and transmitted for second-level review, ensuring consistent state determination across incremental document receipt.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual determination of document state is used, then individual knowledge and experience can be applied, but inconsistent classifications and inaccurate urgency or routing determinations result

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent replaces the manual mechanical review process with an automated optical character recognition (OCR) system that converts document images to text, segments the text, and applies rule-based classification. This substitution eliminates human variability and provides consistent, accurate document state determination across all documents.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables documents to self-classify by automatically processing their own content through OCR and text segmentation, then applying classification rules to determine their own state. This self-service approach ensures consistent classification without requiring manual review while maintaining high accuracy through automated rule application.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If complete document receipt is required for state determination, then accurate classification can be achieved, but processing delays occur in bulk facsimile messaging environments

Engineering Contradiction:
Improvestate determination accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary OCR processing and text segmentation on documents as they are being received in bulk facsimile messaging environments. By preparing the text data in advance and applying classification rules incrementally, the system can determine document state before complete receipt, eliminating processing delays while maintaining accuracy through progressive analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system segments documents into individual pages or sections and processes each segment independently through OCR and text analysis. This segmentation allows parallel processing of multiple document portions simultaneously, enabling state determination to begin before the entire document is received, thus reducing overall processing time while maintaining classification accuracy.

Inventive Principle:
Principle #1Segmentation

3Productivity

If incremental document receipt is processed, then processing speed is improved, but inconsistent state determinations may occur

Engineering Contradiction:
Improveprocessing speedVSAvoidstate determination consistency
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the classification system continuously monitors incoming document segments, applies classification rules, and adjusts processing based on accumulated information. This feedback loop ensures that incremental processing maintains consistency by refining state determinations as more document content becomes available, preventing inconsistent classifications.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically adjusts its processing approach based on the completeness and content of received document segments. As documents are incrementally received, the classification rules are applied flexibly, allowing the system to make preliminary determinations when sufficient information is available while maintaining the ability to revise classifications as complete document content is processed, ensuring both speed and consistency.

Inventive Principle:
Principle #15Dynamics

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach ensures uniform and accurate document state classification before the entire document is received, overcoming inconsistencies and speed limitations, enabling reliable urgency and routing decisions.

Implementation Method 1

performing optical character recognition (OCR) upon a page of a document in order to produce parseable text

Methodology Applied
Scientific EffectOptical character recognition: Photoelectric Effect

Data Source

PatentUS20230419018A1Automatic state assignment to documents based on phrase occurrence in text
Publication Date: 2023.12.28 CONCORD III LLC
  • US20230419018A1 patent drawing
  • US20230419018A1 patent drawing
  • US20230419018A1 patent drawing

AI summary

Fuzzy document state assignment includes loading into memory a raster image of a document and performing OCR upon a page of a document in order to produce parseable text. The parseable text is then segmented and normalized and an index is generated from the segmented and normalized text. Thereafter, a probability of a particular classification is computed based upon the detection in the index of a combination of words associated with a corresponding classification. Finally, the document is annotated with the particular classification.