Image Processing System for Automatic Document File Naming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document management methods require users to manually set file names for multiple scanned images of the same document form, leading to repetitive and time-consuming work, as existing techniques do not allow for automatic proposal of file names for similar documents.

Innovation Solution

An image processing system that includes an MFP and cloud services for analyzing scanned images, detecting text blocks, and determining similarity between images, allowing for automatic file name proposal and learning results reflection to streamline the file naming process for batches of similar documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual file name setting is performed for each scanned image, then file names can be set accurately, but user time and effort increase significantly when processing multiple documents

Engineering Contradiction:
Improvefile_name_accuracyVSAvoidtime_for_file_name_setting
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary OCR processing and file name extraction on the first scanned image before the user completes manual setting. When a user manually sets a file name for one document in a batch, the system automatically extracts the file name from that image using OCR, stores it as reference data, and then automatically applies this extracted file name to subsequent similar documents in the batch, eliminating the need for repeated manual setting while maintaining accuracy through user-confirmed reference data.

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If file name setting is performed individually for each document, then each file name can be optimized, but the process becomes repetitive and inefficient for batches of similar documents

Engineering Contradiction:
Improveease of file_name_settingVSAvoiddocument_processing_speed
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The system creates a copy of the file name setting from one document and applies it to multiple similar documents. When a user sets a file_name for one document in a batch, the system extracts this file_name using OCR, stores it as reference data, and automatically copies this extracted file_name to subsequent documents with similar characteristics, thereby eliminating repetitive manual work while maintaining the quality of individual file_name optimization.

Inventive Principle:
Principle #26Copying

3Productivity

If automatic file name extraction is implemented, then processing speed increases, but user control over file name accuracy decreases

Engineering Contradiction:
Improvefile_name_processing_speedVSAvoiduser_control_over_file_name
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system implements a feedback mechanism where automatically extracted file names are presented to the user for confirmation or correction. The OCR-extracted file_name from the first document is used as reference data, and when the user manually sets or corrects a file_name, this corrected information feeds back into the system as updated reference data, which then automatically updates subsequent documents. This ensures both high processing speed through automation and maintained user control through confirmation and correction capabilities.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3855717B1Image processing system for computerizing document, control method thereof, and storage medium background
Publication Date: 2023.08.16 CANON KK
  • EP3855717B1 patent drawingFigure 1
  • EP3855717B1 patent drawingFigure 2
  • EP3855717B1 patent drawingFigure 3

AI summary

In an image processing system in which when a paper document is computerized, a file name or the like is set by using a recognized character string obtained by performing OCR processing, so that time and effort of a user when a plurality of documents is computerized en bloc is reduced. Learning data is generated by registering positional information relating to a recognized character string used for setting of a property relating to a scanned image in association with a document form of the scanned image. Then, in a case where the learning data is generated in response to setting of the property being performed for a first scanned image that is selected from a plurality of scanned images included in a list, a scanned image having a document form similar to a document form of the first scanned image is determined among other scanned images included in the list.