Document Image Classification via Segment Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI/ML approaches for document image classification are time-consuming and prone to human error during labeling, leading to miss-classification and waste of resources due to the need for extensive training data sets and manual labeling.

Innovation Solution

A method that divides document images into segments, extracts features, generates numerical coefficients, and compares these to create classification codes without relying on machine learning or artificial intelligence, allowing for quick unsupervised classification and similarity rate calculation between documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If AI/ML approach is used for document image classification, then classification accuracy can be improved, but training time and human labor increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-classification by automatically extracting features and generating classification codes without requiring human labeling of training data. The image processor independently analyzes document images, extracts graphical parameters, and generates classification codes through automated comparison of image segments, eliminating the need for manual data preparation and model training.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual human labeling and iterative model training with an automated image processing system. The image processor uses algorithmic feature extraction and numerical coefficient generation to substitute for human judgment, enabling rapid classification without the time-consuming manual intervention required in traditional AI/ML approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual labeling of training data is performed, then classification model can be trained, but human error and lack of experience lead to miss-classification

Engineering Contradiction:
Improveclassification accuracyVSAvoidlabeling reliability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system eliminates human involvement in the classification process by performing self-service analysis. The image processor automatically extracts graphical parameters from document images, generates numerical coefficients, and creates classification codes through automated comparison of image segments, removing the source of human error entirely.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces human labeling operations with automated image processing mechanisms. The system uses algorithmic feature extraction and objective numerical comparison to substitute for subjective human judgment, ensuring consistent and error-free classification that does not depend on operator experience or expertise.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If extensive training data set is used, then classification model performance improves, but time and human labor for data preparation increase

Engineering Contradiction:
Improveclassification accuracyVSAvoiddata preparation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system generates its own classification codes through automated feature extraction and comparison, eliminating the need for extensive pre-labeled training data. The image processor independently analyzes each document image, extracts graphical parameters, and generates classification codes without requiring a large pre-prepared dataset, thereby improving productivity while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary feature extraction and numerical coefficient generation for each image segment before final classification. By pre-processing the image into segments and extracting graphical parameters in advance, the system efficiently generates classification codes without needing extensive training data, as the preliminary actions prepare the data structure needed for rapid automated classification.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11763586B2Method and system for classifying document images
Publication Date: 2023.09.19 KYOCERA DOCUMENT SOLUTIONS INC
  • US11763586B2 patent drawing
  • US11763586B2 patent drawing
  • US11763586B2 patent drawing

AI summary

A method and system are used for managing and classifying electronic document images. Each of the electronic document images is divided into an array of image segments. The method extracts image features from each of the image segments to obtain numerical coefficients for each of the image segments. The numerical coefficients are compared with each other to generate sub-codes. A classification code is determined as a combination of the sub-codes. The classification codes of a plurality of electronic document images can be stored in a database for further analysis. Based on the classification codes, similarity rates between at two document images can be determined.