Automatic Document Separation Using Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for document or subdocument boundary detection in digital scanning are inefficient, requiring manual insertion of separator pages and rule-based systems that become costly and time-consuming as document volumes increase, and fail to identify document types, limiting automation and processing efficiency.
Innovation Solution
A system using supervised machine learning and probabilistic networks to automatically construct rules for document separation and classification, combining textual and graphical information to generate high-quality separations and identify document types, with configurable constraints and customizable filter rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If manual separator pages are used to sort documents, then document separation is achieved, but processing time and cost increase linearly with document volume
Solution Approach 1:
The patent replaces the mechanical system of physical separator pages with an optical/image processing system that analyzes scanned document images to automatically detect and separate documents based on visual features such as borders, spacing, and content patterns
Solution Approach 2:
The system enables documents to self-identify their boundaries and groupings through automated image analysis, eliminating the need for external manual intervention or physical separator pages to indicate document boundaries
2Reliability
If human constructed rule based systems are used for document categorization, then certain tasks are performed well, but the system becomes cumbersome and expensive as the number of document types and business rules increases
Solution Approach 1:
The patent transforms the approach from using discrete business rules to using continuous image parameters and visual features (such as border detection, spacing measurements, and content density metrics) that can be automatically measured and compared to identify document boundaries and types
Solution Approach 2:
The system replaces the mechanical rule-based classification system with an optical recognition system that uses image processing algorithms to automatically detect document boundaries and characteristics, eliminating the need for manual rule configuration and maintenance
3Loss of information
If separator pages are inserted to identify document types, then processing information is provided, but manual insertion effort and cost increase
Solution Approach 1:
The patent extracts the document type identification function from the physical separator page and embeds it directly into the document image analysis process, allowing the system to identify document types by analyzing visual features within the document images themselves rather than relying on external separator pages
Solution Approach 2:
The system uses automatically generated electronic separator pages or metadata as an intermediary between the scanned document images and the processing system, providing document type identification information without requiring manual physical insertion of separator pages
Data Source
AI summary
A method and system for delineating document boundaries and identifying document types by analyzing digital images of one or more documents, automatically categorizing one or more pages or subdocuments within the one or more documents and automatically generating delineation identifiers, such as computer-generated images of separation pages inserted between digital images belonging to different categories, a description of the categorization sequence of the digital images, or a computer-generated electronic label affixed or associated with said digital images.


