Multi-modal Ensemble Deep Learning for Document Image Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current approaches for document splitting rely on human-engineered features and require expert knowledge, making them inefficient for automatically classifying and splitting unstructured document image packages.

Innovation Solution

The use of multi-modal ensemble deep learning techniques to classify each page in a dataset, incorporating multiple independent trained models to predict whether a page is a start page or not, and combining these predictions to generate a final output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If human-engineered features and expert knowledge are used for document splitting, then the system can achieve document separation, but the process requires manual intervention and is inefficient

Engineering Contradiction:
Improvedocument splitting efficiencyVSAvoidautomatic classification capability
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system enables documents to automatically classify and split themselves using deep learning models that autonomously analyze document images and determine page boundaries without human intervention. The multi-modal ensemble models process document features and generate predictions automatically, eliminating the need for manual feature engineering and expert knowledge input.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes (human experts manually analyzing and splitting documents) with automated deep learning systems. Multiple independent trained models process document images through neural networks, substituting human cognitive work with computational algorithms that automatically identify document boundaries and classify pages.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If multiple independent trained models are used for prediction, then the classification accuracy is improved, but the system complexity increases

Engineering Contradiction:
Improvestart page classification accuracyVSAvoidmulti-modal ensemble model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the document classification task into multiple independent models, each specializing in different aspects: image features, font features, margin features, and textual features. Each model processes specific modalities separately, then their predictions are combined through an additional trained model. This segmentation allows each component to focus on its strength while maintaining manageable complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system combines multiple independent model predictions into a composite classification system. Like composite materials that combine different substances to achieve superior properties, the ensemble of multiple models produces more accurate and reliable document classification than any single model could achieve alone, with the additional trained model synthesizing the combined predictions.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS12217525B1Multi-modal ensemble deep learning for start page classification of document image file including multiple different documents
Publication Date: 2025.02.04 FIRST AMERICAN FINANCIAL CORP
  • US12217525B1 patent drawing
  • US12217525B1 patent drawing
  • US12217525B1 patent drawing

AI summary

Multimodal techniques are described for classifying start pages and/or document types of an unstructured document image package. To that end, some implementations of the disclosure relate to a method, including: obtaining a document image file including multiple pages and multiple document types; generating, for each page of the multiple pages, using multiple independent trained models, multiple independent predictions, each of the multiple independent predictions indicating: whether or not the page is a first page of a document, or a document type of the multiple document types that the page corresponds to; and generating, for each page of the multiple pages, based on the multiple independent predictions, using a neural network, a final prediction output indicating whether or not the page is the first page of a document, or one of the multiple document types that the page corresponds to.