Document Splitting via Character String Difference Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing information processing systems require documents with the same character, symbol, or other objects to be prepared in advance for each document set to identify and split documents effectively, which does not significantly reduce the burden associated with document splitting when dealing with batches of paper media.

Innovation Solution

An information processing apparatus that acquires a read image of multiple documents, extracts character strings associated with user-specified items, sets split positions based on differences in character strings between pages, and outputs the read image split into document sets, allowing for automatic identification and reduction of document splitting burden without prior preparation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If documents with the same character, symbol, or other objects are prepared in advance for each document set to identify and split documents, then document splitting can be performed, but the burden associated with splitting documents is not significantly reduced

Engineering Contradiction:
Improveautomatic document splittingVSAvoidpreparation of documents with identical markers
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The invention extracts the identification task from manual document preparation to automatic computer-based processing. The computer reads the batch of paper media, extracts character strings from each document, and automatically identifies document sets based on these extracted features, eliminating the need for manual preparation of documents with identical markers.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system enables documents to identify themselves through their inherent content. By extracting character strings directly from the documents and using these extracted features for identification and grouping, the documents effectively perform self-identification without requiring external markers or manual intervention.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If documents with the same character, symbol, or other objects are prepared in advance for each document set, then documents can be identified and split into sets, but the burden of document preparation and splitting remains high

Engineering Contradiction:
Improvedocument identification accuracyVSAvoidtime for document preparation and splitting
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The computer performs preliminary reading and extraction of character strings from all documents in the batch before the splitting operation. This preliminary action of extracting identification features from the image data enables subsequent automatic grouping and splitting without requiring time-consuming manual preparation of documents with identical markers.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention replaces the mechanical process of manually preparing documents with identical markers and physically sorting them with an automated computer-based system. The computer reads the batch of paper media, extracts character strings, and automatically groups documents into sets based on these extracted features, substituting manual mechanical operations with automated information processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11336779B2Information processing apparatus, and non-transitory computer readable medium
Publication Date: 2022.05.17 FUJIFILM BUSINESS INNOVATION CORP
  • US11336779B2 patent drawing
  • US11336779B2 patent drawing
  • US11336779B2 patent drawing

AI summary

An information processing apparatus includes a processor configured to: acquire a read image and item information, the read image being an image obtained by reading a paper medium including plural documents, the item information being information of an item specified by a user from among plural items contained in the documents; extract a character string from the read image, the character string being associated with the item information; if a character string contained in a page of the read image and extracted from the page differs from a character string extracted from the previous page immediately preceding the page, set a split position, the split position being a position at which to split out a portion of the read image as a set of documents, the portion being a portion of the read image from a page where extraction has begun to the previous page; and output the read image split in accordance with the split position.