Automated Document Categorization Using OCR Error Feature Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual categorization of large volumes of electronic documents is impractical and inefficient, especially when dealing with low-quality images containing noise and defects, which conventional OCR algorithms struggle to process accurately.

Innovation Solution

An apparatus equipped with an optical character recognition (OCR) algorithm and natural language processing (NLP) algorithm that converts images into text, identifies errors, generates feature vectors incorporating error types, and uses machine learning to assign images to document categories, thereby automating the categorization and assembly of electronic documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual categorization is used for electronic documents, then categorization accuracy can be maintained through human judgment, but processing efficiency and productivity deteriorate due to the impracticality of manually reviewing large volumes of documents

Engineering Contradiction:
Improvecategorization accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent introduces an intermediary system comprising OCR algorithms, error detection mechanisms, and machine learning classifiers that act as a mediator between the physical document and the categorization database. This intermediary automatically processes document images, extracts text, identifies errors, and performs categorization, thereby resolving the contradiction between maintaining accuracy and improving processing efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If conventional OCR algorithms are used on low-quality images, then processing speed can be maintained, but measurement precision and manufacturing precision deteriorate due to noise and defects in the images

Engineering Contradiction:
Improveprocessing speedVSAvoidtext extraction accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent applies preliminary error detection and correction actions before the main OCR processing step. The system first identifies potential errors in the extracted text by analyzing noise patterns and defects in the original image, then corrects these errors before proceeding with categorization. This preliminary action improves text extraction accuracy without significantly compromising processing speed.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated OCR and categorization systems are implemented, then productivity and processing capacity improve, but device complexity and measurement precision requirements increase due to the need for accurate error identification in noisy images

Engineering Contradiction:
Improvedocument processing capacityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the automated document processing system into distinct functional modules: image preprocessing, OCR text extraction, error detection, error correction, and categorization. Each module handles a specific aspect of the processing pipeline, making the overall complex system more manageable and maintainable while preserving high productivity benefits.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If error detection and correction features are added to improve text accuracy, then measurement precision improves, but device complexity and processing time increase

Engineering Contradiction:
Improvetext accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent changes key parameters in the error detection and correction process by adjusting sensitivity thresholds, error type weights, and correction confidence levels. These parameter adjustments allow the system to achieve high text accuracy while controlling processing time, as the system can adaptively skip or simplify certain error correction steps based on image quality and error severity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12033367B2Automated categorization and assembly of low-quality images into electronic documents
Publication Date: 2024.07.09 BANK OF AMERICA CORP
  • US12033367B2 patent drawing
  • US12033367B2 patent drawing
  • US12033367B2 patent drawing

AI summary

An apparatus includes a memory and processor. The memory stores document categories, text generated from an image a physical document page, and a machine learning algorithm. The machine learning algorithm is configured to extract features associated with natural language processing and features associated with the text. The machine learning algorithm is also configured to generate a feature vector that includes the first and second pluralities of features, and to generate, based on the feature vector, a set of probabilities, each of which is associated with a document category and indicates a probability that the physical document from which the text was generated belongs to that document category. The processor applies the machine learning algorithm to the text, to generate the set of probabilities, identifies a largest probability, and assigns the image to the associated document category.