Scanned Document Bundle Segmentation Using Title and Identifier Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques fail to accurately divide scanned documents into appropriate standard document bundles, often mistakenly separating documents that should remain together due to mechanical title-based division.
Innovation Solution
An information processing apparatus and method that utilizes title extraction followed by key-value extraction to identify common identifiers within and between document bundles, allowing for the reintegration of documents that should be grouped together.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If documents are divided into bundles based on title extraction only, then the division process is simple and fast, but documents that should be grouped together are mistakenly separated
Solution Approach 1:
The patent divides the document bundling process into two distinct stages: first segmenting documents by title to create initial groups, then further segmenting these groups using identifier matching to achieve accurate standard document bundles. This two-stage segmentation resolves the contradiction by maintaining both speed (first stage) and accuracy (second stage).
Solution Approach 2:
The patent introduces an intermediary step of identifier extraction and matching between the title-based division and the final bundling. This intermediary mechanism (identifier matching) acts as a mediator that corrects the rough grouping from title extraction, ensuring documents with the same identifier are properly grouped together while maintaining processing efficiency.
2Device complexity
If title extraction is used for document division, then the process is straightforward, but it cannot always achieve appropriate division into standard document bundles
Solution Approach 1:
The patent segments the division process into two phases: a simple title extraction phase for initial grouping, and an identifier matching phase for precise bundling. This segmentation allows the system to maintain low complexity in the first phase while achieving high precision in the second phase, resolving the contradiction between simplicity and accuracy.
Solution Approach 2:
The patent performs preliminary title extraction to create initial document groups before applying the more precise identifier matching. This preliminary action reduces the search space for the subsequent precision step, allowing accurate bundling without excessive overall complexity.
3Loss of time
If only title-based division is performed, then processing is quick, but related documents are mistakenly separated into different bundles
Solution Approach 1:
The patent segments the processing into a quick title-based grouping stage followed by a targeted identifier verification stage. This segmentation ensures that most documents are quickly assigned to bundles while only requiring detailed verification for documents within the same title group, minimizing time loss while improving reliability.
Solution Approach 2:
The patent applies partial action by performing identifier matching only within title-based groups rather than across all documents. This partial verification approach maintains quick processing while ensuring accurate association of related documents that share the same identifier within each group.
Data Source
AI summary
An information processing apparatus includes a processor configured to acquire plural scanned images obtained by scanning a bundle of paper media including a plural standard document bundles each of which is a set of a standard document and a related document related to the standard document, extract a title from the plural scanned images, extract, from the plural scanned images, an identifier that is assigned to be common within one standard document bundle and to be different between different standard document bundles, and divide the plural scanned images into bundles such that a scanned image from which the title is extracted is set as a bundle head, scanned images having the same identifier are included in the same bundle, and scanned images having different identifiers are included in different bundles.


