Image Processing Apparatus Topic Word Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing and forming apparatuses struggle to effectively classify images of multiple pages of an original document based on specific topic words, leading to inefficient sorting and organization.

Innovation Solution

An image processing apparatus with an input device, operation device, and control device that recognizes character strings in images using OCR and classifies pages based on input topic words, grouping pages with the same topic words together, and handling pages without matching words into existing groups.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing image processing apparatuses are used to classify document pages, then basic document handling is possible, but accurate classification based on topic words cannot be achieved

Engineering Contradiction:
Improveclassification accuracyVSAvoidtopic-based sorting capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the classification task by processing each page individually and independently. The controller extracts character strings from each page separately, determines topic word presence for each page individually, and classifies pages into different groups based on their specific topic word matches. This page-by-page segmentation enables precise topic-based classification that existing apparatuses cannot achieve.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the classification parameter from basic document handling to topic word-based classification. By introducing topic words as the classification criterion and using OCR to extract and match character strings against these topic words, the system transforms the classification approach to achieve both high accuracy and adaptability to different sorting requirements.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If manual sorting of document pages is performed, then accurate organization is possible, but processing time and labor are excessive

Engineering Contradiction:
Improvedocument sorting efficiencyVSAvoidclassification processing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs self-service classification by automatically extracting character strings from each page using OCR, automatically determining topic word presence, and automatically sorting pages into appropriate groups. This automation eliminates manual sorting labor while maintaining accurate topic-based organization, significantly improving productivity without requiring excessive processing time.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical sorting with an automated optical and computational system. OCR technology converts optical character data into machine-processable text, and the controller automatically compares this text against topic words to determine classification. This substitution of mechanical manual sorting with optical-recognition-based automated sorting dramatically improves efficiency while reducing time loss.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If pages are classified without OCR recognition, then processing is faster, but topic word-based classification cannot be performed

Engineering Contradiction:
Improvetopic word matching accuracyVSAvoidOCR and classification system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary OCR recognition on all pages before classification begins. By pre-extracting character strings from each page and storing them for subsequent topic word comparison, the system ensures that accurate topic word matching can be performed on all pages. This preliminary action maintains reliability while organizing the complex processing into manageable sequential steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces character string extraction via OCR as an intermediary step between image input and topic word classification. The OCR process converts image data into text data that can be compared against topic words, serving as a necessary mediator that enables accurate topic-based classification while managing system complexity through modular processing stages.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of operation

If all pages are processed together, then batch processing is efficient, but individual page classification based on specific topic words is lost

Engineering Contradiction:
Improvetopic-specific grouping capabilityVSAvoidnumber of classification groups
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The patent segments the classification output into multiple topic-specific groups rather than processing all pages into a single batch. Each page is evaluated against topic words and assigned to appropriate groups based on its specific topic word matches. This segmentation enables users to easily access and retrieve pages by specific topics while the system efficiently manages the increased number of classification groups through automated processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11825041B2Image processing apparatus and image forming apparatus capable of classifying respective images of plurality of pages of original document based on plurality of topic words
Publication Date: 2023.11.21 KYOCERA DOCUMENT SOLUTIONS INC
  • US11825041B2 patent drawing
  • US11825041B2 patent drawing
  • US11825041B2 patent drawing

AI summary

An image processing apparatus includes an input device, an operation device, and a control device. The control device functions as a controller. The controller accepts an input of a plurality of topic words through the operation device, recognizes, upon acceptance of input of images of a plurality of pages of the original document through the operation device, a character string in the image of each of the plurality of pages of the original document, makes a determination of whether or not the character string contains at least one of the plurality of input topic words, and classifies the respective images of the plurality of pages of the original document on basis of the individual topic word based on results of the determination for the individual pages of the original document by collecting the images of the pages of the original document containing the same topic word into one group.