Document Boundary Detection via Anisotropic Diffusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current document processing systems face inefficiencies in automatically identifying document boundaries in bulk collections of digital documents, leading to increased costs and low recognition accuracy due to the need for manual intervention and the use of costly physical separators or ineffective handcrafted rules.
Innovation Solution
A computer-implemented method and system that computes page category scores by considering pairs of consecutive pages, using anisotropic diffusion to refine scores based on boundary probabilities, thereby improving document boundary detection and categorization accuracy without the need for manual intervention or additional physical separators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If physical segmentation with separator sheets is used to mark document boundaries, then document boundary detection accuracy is improved, but cost and manual intensity increase significantly
Solution Approach 1:
The patent creates a digital copy of the document stream and processes this copy to identify boundaries, eliminating the need for physical separator sheets. The system analyzes visual features, text content, and metadata from the digital pages to reconstruct document boundaries without physical intervention
Solution Approach 2:
The patent replaces the mechanical/physical segmentation system with a computational system that uses algorithms to detect document boundaries. Instead of physical separators, the system uses software-based analysis of page characteristics, text patterns, and visual features to automatically identify document boundaries
2Ease of operation
If pages are categorized in isolation using standard techniques, then processing simplicity is maintained, but categorization accuracy deteriorates due to inability to leverage sequential page relationships
Solution Approach 1:
The patent merges the categorization of individual pages with the analysis of page sequences. Instead of processing pages independently, the system combines page-level features with sequential relationships, using models that consider both the content of each page and its position relative to other pages in the document stream
Solution Approach 2:
The patent adds a temporal/sequential dimension to the categorization problem. By considering pages not just in isolation but also in their sequence order and relationship to neighboring pages, the system transforms a one-dimensional (single page) categorization task into a multi-dimensional approach that incorporates sequential context
3Productivity
If handcrafted rules are used to establish page sequence information, then some document reconstruction is achieved, but recognition accuracy remains low with many false positives
Solution Approach 1:
The patent enables the document stream to self-identify its own boundaries through automated analysis of intrinsic characteristics. The system uses machine learning models that automatically detect patterns in visual features, text content, and metadata to reconstruct document boundaries without requiring external rules or manual intervention
Solution Approach 2:
The patent implements iterative refinement where the system continuously improves its boundary detection through feedback loops. The model analyzes page characteristics, compares them against learned patterns, and adjusts its classifications based on the consistency and context of surrounding pages, progressively improving accuracy
Data Source
AI summary
A computer implemented system and method are provided for refining category scores for pages of a sequence of document pages that potentially includes document boundaries. The method uses initial category scores provided by a categorizer that considers one page at a time or concatenated pairs of pages (called bipages). The category scores represent the probability that a page belongs to a particular category. The method uses anisotropic diffusion to refine the initial page category scores using the scores of neighboring pages as a function of the probability that there is a boundary between the pages. The method may be performed iteratively.


