Hierarchical Document Classification with Fuzzy Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated systems for classifying and identifying metadata in electronic documents face challenges in achieving accurate recognition due to the presence of noise from unrelated document pages, requiring excessive computational resources and often resulting in inaccurate models, especially when metadata is embedded in both textual and graphical/layout information.
Innovation Solution
A multi-stage hierarchical approach involving text-based document classification, image-based metadata recognition, and supplemental fuzzy text matching, utilizing machine learning algorithms to reduce computational requirements and improve accuracy by filtering irrelevant documents and leveraging unique graphical/layout features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing automated systems process all document pages including unrelated content, then comprehensive metadata recognition is attempted, but computational resources are excessively consumed and accuracy decreases due to noise
Solution Approach 1:
The patent segments the document processing task into distinct stages: first identifying the document type (e.g., invoice, receipt), then selectively extracting metadata from relevant sections only. This segmentation allows the system to avoid processing unrelated document pages, reducing computational resource consumption while maintaining comprehensive metadata recognition accuracy.
Solution Approach 2:
The system performs preliminary document type classification before attempting metadata extraction. By identifying the document type first, the system can pre-determine which metadata fields are relevant and which document sections should be processed, thereby avoiding wasteful computation on unrelated content while ensuring all necessary metadata is captured.
2Measurement precision
If existing automated systems process all document pages including unrelated content, then comprehensive metadata recognition is attempted, but noise from unrelated pages reduces recognition accuracy
Solution Approach 1:
The patent extracts and isolates only the relevant portions of the document for metadata recognition based on the identified document type. By taking out and processing only the pertinent sections while excluding unrelated pages, the system eliminates noise from irrelevant content and maintains high metadata recognition accuracy.
Solution Approach 2:
The system applies different processing qualities to different parts of the document based on their relevance. Relevant sections receive full processing attention for accurate metadata extraction, while unrelated sections are either skipped or given minimal processing, thereby reducing noise impact while maintaining overall accuracy.
3Measurement precision
If text-based classification alone is used, then processing is simple, but accuracy is insufficient when metadata is embedded in graphical/layout information
Solution Approach 1:
The patent merges text-based classification with image-based analysis into a unified system. The text-based component identifies document type and structure, while the image-based component extracts metadata from graphical elements and layouts. By combining these approaches, the system achieves high metadata identification accuracy while managing complexity through integrated architecture.
Solution Approach 2:
The system implements a multi-functional classification framework that handles both text-based and image-based metadata extraction through a unified document type classification mechanism. This universal approach allows the same system to process diverse metadata formats (textual and graphical) without requiring separate specialized systems, thereby improving accuracy while controlling complexity.
Data Source
AI summary
A hierarchical document classification system is disclosed. The system includes a text-based document classifier model for classifying an input electronic document into one of a set of predefined document categories. The system further includes an image-based metadata identification model for classifying electronic documents of a particular document category into a set of metadata categories. The system further includes a fuzzy text matcher for supplementing classification accuracy of the image-based metadata identification model to obtain a metadata category for the input electronic document.


