Document Analysis System Filters Binarization Artifacts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for organizing electronic documents are time-consuming, expensive, and require human labor, education, domain expertise, and software knowledge, and do not minimize errors or protect data privacy effectively.

Innovation Solution

A system and method that automatically organizes scanned documents into categorized, searchable, and editable electronic documents using a document analysis system that filters binarized background artifacts, distinguishes text and image features, and employs machine learning for classification, with secure data transfer and storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual organization of document pages is performed, then the pages can be properly ordered, but the process becomes time-consuming and expensive

Engineering Contradiction:
Improvedocument organization accuracyVSAvoidorganization time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The document analysis system automatically performs organization, classification, and indexing of document pages without human intervention. The system extracts text and image features, classifies documents into categories, and generates table of contents and bookmarks autonomously, eliminating the need for manual organization while maintaining high accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical organization processes with automated computer-based image analysis and machine learning systems. The system uses optical character recognition, image feature extraction, and classification algorithms to automatically organize documents, substituting human labor with computational processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If binarization is applied to reduce file size, then storage efficiency improves, but background artifacts are introduced that interfere with text and image recognition

Engineering Contradiction:
Improvefile sizeVSAvoidrecognition accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system extracts and removes binarization background artifacts from the processed document images. By identifying and eliminating these artifacts through image analysis and filtering techniques, the system restores recognition accuracy while preserving the file size benefits of binarization

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent converts the harmful effect of binarization artifacts into a solvable problem through automated detection and removal. The system uses the artifacts' distinctive characteristics to identify and eliminate them, turning the binarization process from a source of error into an acceptable preprocessing step that maintains both file size efficiency and recognition accuracy

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS8538184B2Systems and methods for handling and distinguishing binarized, background artifacts in the vicinity of document text and image features indicative of a document category
Publication Date: 2013.09.17 GRUNTWORX LLC
  • US8538184B2 patent drawing
  • US8538184B2 patent drawing
  • US8538184B2 patent drawing

AI summary

A method of enhancing electronic documents received from a plurality of users by a document analysis system for improving automatic recognition and classification of the received electronic documents, is provided. For each page of a received electronic document, the method filters the page to infer binarized-background artifacts resulting from the binarization of the original grayscale or color image source document and which reside in the vicinity of binarized text and binarized image features in the page, so that the binarized text and binarized images may be distinguished from the binarized-background artifacts and extracted from the document. The method then uses the extracted features from the filtered document to automatically recognized and classify a document into a document category.