Document Classification Using Page-Level Boundary Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document capture systems struggle to efficiently classify and index unstructured electronic documents, which lack consistency and structure, making it difficult to recognize logical boundaries and classify document types accurately.
Innovation Solution
A classification system that employs a page-level recognition model to identify logical boundaries between documents and separate unstructured pages, followed by a document-level recognition model to classify the separated documents into specific types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If keyword-based document recognition is used, then accuracy for structured documents is high, but effectiveness for unstructured documents deteriorates
Solution Approach 1:
The patent segments the document classification task into multiple stages: first separating structured from unstructured documents, then applying different recognition models to each type. This segmentation allows the system to maintain high accuracy for structured documents while improving effectiveness for unstructured documents through specialized processing.
Solution Approach 2:
The system dynamically adapts its processing approach based on document structure. It uses structured document templates for structured documents and switches to unstructured document processing with logical boundary detection for unstructured documents, allowing the recognition system to adjust its behavior based on the input document type.
2Productivity
If automatic classification is applied to all documents, then processing efficiency is improved, but classification accuracy for unstructured documents deteriorates
Solution Approach 1:
The patent segments the document processing workflow into distinct phases: initial automatic classification using templates, separation of structured and unstructured documents, and selective manual review. This segmentation enables high processing efficiency for structured documents while maintaining accuracy for unstructured documents through targeted manual intervention.
Solution Approach 2:
The system introduces an intermediary manual review stage that acts as a mediator between automatic classification and final document indexing. This intermediary layer allows operators to review and correct classifications for unstructured documents, ensuring accuracy while maintaining overall processing efficiency.
3Measurement precision
If manual classification is used for uncertain documents, then classification accuracy is improved, but processing time increases
Solution Approach 1:
The patent segments documents into structured and unstructured categories, applying automated template-based classification to structured documents for rapid processing, while reserving manual review for unstructured documents. This segmentation minimizes the time spent on manual classification while maintaining accuracy where needed.
Solution Approach 2:
The system enables self-service automatic classification for structured documents using pre-defined templates and patterns, allowing documents to be processed without human intervention. This self-service approach reduces processing time for the majority of structured documents while maintaining high accuracy.
4Ease of manufacture
If document templates with fixed layouts are used, then structured document classification is simplified, but adaptability to unstructured documents deteriorates
Solution Approach 1:
The patent segments the classification system into two distinct processing paths: one for structured documents using fixed templates and patterns, and another for unstructured documents using logical boundary detection. This segmentation allows the system to implement simple template-based classification for structured documents while maintaining adaptability to unstructured documents through a separate processing mechanism.
Data Source
AI summary
A classification system is provided that separates unclassified pages into unclassified, separated documents and classifies the separated documents. The classification system applies a page-level recognition model to the unclassified pages to recognize the logical boundaries between documents and, based on the logical boundaries, separates the pages into unclassified, separated documents. The classification system further applies a document-level recognition model to classify the separated documents.


