Automatic Document Separation Using Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for document or subdocument boundary detection in digital scanning are inefficient, requiring manual insertion of separator pages and rule-based systems that become costly and time-consuming as document volumes increase, and fail to identify document types, limiting automation and processing efficiency.

Innovation Solution

A system using supervised machine learning and probabilistic networks to automatically construct rules for document separation and classification, combining textual and graphical information to generate high-quality separations and identify document types, with configurable constraints and customizable filter rules.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If manual separator pages are used to sort documents, then document separation is achieved, but processing time and cost increase linearly with document volume

Engineering Contradiction:
Improveautomatic document separationVSAvoidmanual effort per document
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The patent replaces the mechanical system of physical separator pages with an optical/image processing system that analyzes scanned document images to automatically detect and separate documents based on visual features such as borders, spacing, and content patterns

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables documents to self-identify their boundaries and groupings through automated image analysis, eliminating the need for external manual intervention or physical separator pages to indicate document boundaries

Inventive Principle:
Principle #25Self-service

2Reliability

If human constructed rule based systems are used for document categorization, then certain tasks are performed well, but the system becomes cumbersome and expensive as the number of document types and business rules increases

Engineering Contradiction:
Improvedocument categorization accuracyVSAvoidnumber of rules and constraints
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the approach from using discrete business rules to using continuous image parameters and visual features (such as border detection, spacing measurements, and content density metrics) that can be automatically measured and compared to identify document boundaries and types

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system replaces the mechanical rule-based classification system with an optical recognition system that uses image processing algorithms to automatically detect document boundaries and characteristics, eliminating the need for manual rule configuration and maintenance

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If separator pages are inserted to identify document types, then processing information is provided, but manual insertion effort and cost increase

Engineering Contradiction:
Improvedocument type identificationVSAvoidseparator page insertion process
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The patent extracts the document type identification function from the physical separator page and embeds it directly into the document image analysis process, allowing the system to identify document types by analyzing visual features within the document images themselves rather than relying on external separator pages

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses automatically generated electronic separator pages or metadata as an intermediary between the scanned document images and the processing system, providing document type identification information without requiring manual physical insertion of separator pages

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8693043B2Automatic document separation
Publication Date: 2014.04.08 TUNGSTEN AUTOMATION CORPORATION
  • US8693043B2 patent drawing
  • US8693043B2 patent drawing
  • US8693043B2 patent drawing

AI summary

A method and system for delineating document boundaries and identifying document types by analyzing digital images of one or more documents, automatically categorizing one or more pages or subdocuments within the one or more documents and automatically generating delineation identifiers, such as computer-generated images of separation pages inserted between digital images belonging to different categories, a description of the categorization sequence of the digital images, or a computer-generated electronic label affixed or associated with said digital images.