Document Classification Using Memory-Enabled Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional document classification models are inefficient in determining suitable candidates for automation and require manual feature engineering, leading to time-consuming and error-prone data extraction from scanned documents.
Innovation Solution
A document classification system using memory-enabled modeling that receives an input image, performs content extraction, selects a matching template, determines a confidence score, and validates the extracted content through image quality, refined content extraction, and layout validation to classify documents as STP or non-STP documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional document classification models are used, then document classification can be performed, but the process requires manual feature engineering which is time-consuming and error-prone
Solution Approach 1:
The patent replaces manual feature engineering (mechanical human operation) with automated machine learning models including a document extractor ML model and a prediction ML model. These models automatically extract features from document images and perform classification without human intervention, thereby eliminating time loss associated with manual feature engineering while maintaining high classification efficiency.
Solution Approach 2:
The system enables self-service classification by training ML models to autonomously perform feature extraction and document classification. The models learn from training datasets and automatically apply extracted features to classify new documents without requiring manual feature engineering for each document, thus improving productivity while reducing time investment.
2Ease of operation
If manual data extraction from scanned documents is performed, then data can be accessed electronically, but the process is time-consuming and error-prone
Solution Approach 1:
The patent replaces manual data extraction operations with automated ML-based extraction. The document extractor ML model automatically extracts relevant data from scanned document images and converts them into electronic format, eliminating the need for manual keying while significantly improving extraction speed and reducing human errors.
Solution Approach 2:
The system introduces ML models as intermediaries between scanned document images and electronic data extraction. These models act as intelligent mediators that automatically interpret document content, extract relevant information, and structure it electronically, thereby maintaining ease of operation while dramatically improving productivity.
3Adaptability or versatility
If conventional document classification models are used, then documents can be categorized, but the models are inefficient in determining suitable candidates for automation
Solution Approach 1:
The patent replaces conventional classification algorithms with advanced ML models that can efficiently identify automation-suitable documents. The prediction ML model analyzes extracted features and determines which documents are appropriate candidates for automated processing, thereby maintaining versatile categorization capability while dramatically improving identification efficiency.
Solution Approach 2:
The system changes the parameters of document analysis by using ML models that learn optimal features and patterns from training data. This enables the system to adaptively identify automation candidates based on learned characteristics rather than fixed conventional rules, improving both versatility and efficiency in determining suitable documents for automation.
Data Source
AI summary
Disclosed is a method of classifying a document for a straight-through processing (STP) using memory enabled modelling. The method includes receiving, performing a content extraction from the input image, and selecting a template with the highest matching probability from a database. The method further includes postprocessing the extracted content based on the predicted template and validating the extracted content. Thereafter, the method includes classifying the document for the STP if the extracted content is successfully validated corresponding to each of the image quality validation, the refined content extraction validation, and the layout validation.


