Document Classification Using Page-Level Boundary Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document capture systems struggle to efficiently classify and index unstructured electronic documents, which lack consistency and structure, making it difficult to recognize logical boundaries and classify document types accurately.

Innovation Solution

A classification system that employs a page-level recognition model to identify logical boundaries between documents and separate unstructured pages, followed by a document-level recognition model to classify the separated documents into specific types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keyword-based document recognition is used, then accuracy for structured documents is high, but effectiveness for unstructured documents deteriorates

Engineering Contradiction:
Improvedocument recognition accuracyVSAvoideffectiveness for unstructured documents
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the document classification task into multiple stages: first separating structured from unstructured documents, then applying different recognition models to each type. This segmentation allows the system to maintain high accuracy for structured documents while improving effectiveness for unstructured documents through specialized processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adapts its processing approach based on document structure. It uses structured document templates for structured documents and switches to unstructured document processing with logical boundary detection for unstructured documents, allowing the recognition system to adjust its behavior based on the input document type.

Inventive Principle:
Principle #15Dynamics

2Productivity

If automatic classification is applied to all documents, then processing efficiency is improved, but classification accuracy for unstructured documents deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the document processing workflow into distinct phases: initial automatic classification using templates, separation of structured and unstructured documents, and selective manual review. This segmentation enables high processing efficiency for structured documents while maintaining accuracy for unstructured documents through targeted manual intervention.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary manual review stage that acts as a mediator between automatic classification and final document indexing. This intermediary layer allows operators to review and correct classifications for unstructured documents, ensuring accuracy while maintaining overall processing efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual classification is used for uncertain documents, then classification accuracy is improved, but processing time increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments documents into structured and unstructured categories, applying automated template-based classification to structured documents for rapid processing, while reserving manual review for unstructured documents. This segmentation minimizes the time spent on manual classification while maintaining accuracy where needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system enables self-service automatic classification for structured documents using pre-defined templates and patterns, allowing documents to be processed without human intervention. This self-service approach reduces processing time for the majority of structured documents while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

4Ease of manufacture

If document templates with fixed layouts are used, then structured document classification is simplified, but adaptability to unstructured documents deteriorates

Engineering Contradiction:
Improveclassification system implementationVSAvoidhandling of unstructured documents
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent segments the classification system into two distinct processing paths: one for structured documents using fixed templates and patterns, and another for unstructured documents using logical boundary detection. This segmentation allows the system to implement simple template-based classification for structured documents while maintaining adaptability to unstructured documents through a separate processing mechanism.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250046108A1System and method for separation and classification of unstructured documents
Publication Date: 2025.02.06 OPEN TEXT SA ULC
  • US20250046108A1 patent drawing
  • US20250046108A1 patent drawing
  • US20250046108A1 patent drawing

AI summary

A classification system is provided that separates unclassified pages into unclassified, separated documents and classifies the separated documents. The classification system applies a page-level recognition model to the unclassified pages to recognize the logical boundaries between documents and, based on the logical boundaries, separates the pages into unclassified, separated documents. The classification system further applies a document-level recognition model to classify the separated documents.