Document Classification System with Failure Image Reporting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional document classification systems face challenges in accurately classifying documents and efficiently handling classification failures, requiring manual intervention and increased operational time due to the need for manual verification of incorrectly classified documents.

Innovation Solution

A document classification system that utilizes an image file as a model for classification through machine learning, incorporates a classification failure image reporter to identify and report failed classifications, and accepts template files containing metadata and region information to facilitate automatic metadata acquisition from images using optical character recognition, thereby automating the classification process and reducing manual effort.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual verification is performed for incorrectly classified documents, then classification accuracy can be improved, but operational time increases significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidoperational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically performs metadata acquisition and classification verification without requiring manual intervention. The classification failure image reporter autonomously identifies failed classifications, and the template acceptor automatically processes template files to extract metadata, eliminating the need for human operators to manually verify each incorrectly classified document while maintaining high classification accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical verification processes with automated optical character recognition (OCR) and machine learning-based classification systems. The document classifier uses machine learning models to automatically verify classifications, and the template acceptor uses OCR technology to automatically extract metadata from images, substituting human manual inspection with automated computational processes that are both faster and equally accurate

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual intervention is required for classification failures, then classification reliability can be maintained, but device complexity increases

Engineering Contradiction:
Improveclassification reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The classification failure image reporter acts as an intermediary component that automatically detects and reports classification failures without requiring manual intervention. It serves as a bridge between the document classifier and the template acceptor, automatically transmitting failure information and coordinating the automated response process, thereby maintaining reliability while avoiding the complexity of manual intervention systems

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements an automated feedback loop where the classification failure image reporter continuously monitors classification results, automatically identifies failures, and feeds this information back to the template acceptor for automated metadata acquisition. This closed-loop feedback system maintains classification reliability through automated correction while keeping the system architecture relatively simple by using standardized feedback mechanisms rather than complex manual intervention protocols

Inventive Principle:
Principle #23Feedback

3Productivity

If automatic metadata acquisition is implemented, then productivity increases, but measurement precision may decrease

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidmetadata accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The template acceptor performs preliminary actions by pre-processing template files and pre-extracting metadata patterns before actual classification occurs. It prepares reference metadata structures in advance, so when automatic metadata acquisition is needed during classification, the system can quickly compare against pre-prepared templates, maintaining both high productivity through automation and high precision through pre-established reference standards

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The automatic metadata acquisition process is segmented into multiple independent stages: image preprocessing, OCR text extraction, metadata pattern matching, and validation. Each segment can be independently optimized and verified, allowing the system to maintain high overall productivity while ensuring precision at each individual stage through specialized processing algorithms and validation rules

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system effectively automates document classification and reduces operational time by enabling the automatic acquisition of metadata from classification failure images, facilitating the creation of template files and improving the efficiency of document processing without the need for manual verification of incorrectly classified documents.

Implementation Method 1

a document classifier that uses an image file, which is a file of an image serving as a model for classifying a document, to classify the document by machine learning

Methodology Applied
Scientific EffectMachine learning:

Implementation Method 2

uses the data file included in the template file to acquire the metadata from an image of the document by optical character recognition

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS11587348B2Document classification system and non-transitory computer readable recording medium storing document classification program
Publication Date: 2023.02.21 KYOCERA DOCUMENT SOLUTIONS INC
  • US11587348B2 patent drawing
  • US11587348B2 patent drawing
  • US11587348B2 patent drawing

AI summary

A document classification system uses an image file as a file of an image serving as a model for classifying a document to classify, by machine learning, an image read from a form as a document by a scanner of an image forming apparatus, and reports a classification failure image as an image of the document when the document is unsuccessfully classified.