Document Classification via Graphical Code Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification methods using OCR and NLP are less efficient and accurate due to their text-based approaches, while image-based methods may not fully leverage topological features for classification.

Innovation Solution

A system that extracts a content sub-region from a document, transforms it into a graphical code using localized OCR and image generation, bypassing NLP, and employs a neural network for image classification, incorporating topological features for improved accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If text-based classification using OCR and NLP is used, then document classification can be performed, but accuracy and efficiency are reduced

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces the text-based NLP classification system with an image-based classification system. Instead of processing extracted text through natural language processing, the system converts document regions to graphical codes (barcodes/QR codes) and uses image classification neural networks, substituting mechanical text processing with optical/image-based recognition for improved accuracy and efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter of classification from text-based features to image-based features. By transforming document content into graphical code representations and using image classification models, the system operates in a different parameter space that yields superior classification performance compared to traditional NLP approaches

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If full-document OCR and NLP processing is performed, then complete text analysis is achieved, but processing time and computational resources increase

Engineering Contradiction:
Improvetext analysis completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts only the relevant content sub-region from the document rather than processing the entire document. By identifying and processing only the necessary portion containing the classification information, the system maintains analysis completeness while significantly reducing processing time and computational resources

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by performing OCR and graphical code generation only on the content sub-region rather than the full document. This selective processing approach achieves the necessary classification information without the excessive time and resource costs of complete document processing

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11669704B2Document classification neural network and OCR-to-barcode conversion
Publication Date: 2023.06.06 KYOCERA DOCUMENT SOLUTIONS INC
  • US11669704B2 patent drawing
  • US11669704B2 patent drawing
  • US11669704B2 patent drawing

AI summary

Document classification techniques are disclosed that convert text content extracted from documents into graphical images and apply image classification techniques to the images. A graphical image of the text (such as a bar-code) may be generated and applied to improve the performance of document classification, bypassing NLP and utilizing more efficient localized OCR than in conventional approaches.