Document Analysis System Filters Binarization Artifacts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for organizing electronic documents are time-consuming, expensive, and require human labor, education, domain expertise, and software knowledge, and do not minimize errors or protect data privacy effectively.
Innovation Solution
A system and method that automatically organizes scanned documents into categorized, searchable, and editable electronic documents using a document analysis system that filters binarized background artifacts, distinguishes text and image features, and employs machine learning for classification, with secure data transfer and storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual organization of document pages is performed, then the pages can be properly ordered, but the process becomes time-consuming and expensive
Solution Approach 1:
The document analysis system automatically performs organization, classification, and indexing of document pages without human intervention. The system extracts text and image features, classifies documents into categories, and generates table of contents and bookmarks autonomously, eliminating the need for manual organization while maintaining high accuracy
Solution Approach 2:
The patent replaces manual mechanical organization processes with automated computer-based image analysis and machine learning systems. The system uses optical character recognition, image feature extraction, and classification algorithms to automatically organize documents, substituting human labor with computational processes
2Quantity of substance
If binarization is applied to reduce file size, then storage efficiency improves, but background artifacts are introduced that interfere with text and image recognition
Solution Approach 1:
The system extracts and removes binarization background artifacts from the processed document images. By identifying and eliminating these artifacts through image analysis and filtering techniques, the system restores recognition accuracy while preserving the file size benefits of binarization
Solution Approach 2:
The patent converts the harmful effect of binarization artifacts into a solvable problem through automated detection and removal. The system uses the artifacts' distinctive characteristics to identify and eliminate them, turning the binarization process from a source of error into an acceptable preprocessing step that maintains both file size efficiency and recognition accuracy
Data Source
AI summary
A method of enhancing electronic documents received from a plurality of users by a document analysis system for improving automatic recognition and classification of the received electronic documents, is provided. For each page of a received electronic document, the method filters the page to infer binarized-background artifacts resulting from the binarization of the original grayscale or color image source document and which reside in the vicinity of binarized text and binarized image features in the page, so that the binarized text and binarized images may be distinguished from the binarized-background artifacts and extracted from the document. The method then uses the extracted features from the filtered document to automatically recognized and classify a document into a document category.


