Image Document Supplier Prediction Beyond Text Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models struggle to accurately predict the supplier identity in image-type documents when the supplier information is not explicitly included in the text, often due to noise introduced by scanning and the presence of logos and graphics that hinder text recognition.
Innovation Solution
A system generates non-textual and textual features from image-type documents to predict supplier identity using machine learning models, classifying documents as scanned or not-scanned based on file size variations, and employing additional features like ratios of slices, color channels, and region-of-interest characteristics to enhance prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If text recognition applications are used to extract supplier information from image-type documents, then extraction automation is improved, but recognition accuracy deteriorates when documents contain logos, graphics, or scanning noise
Solution Approach 1:
The system segments the document analysis into multiple independent components: text recognition for explicit supplier fields, non-textual feature extraction for visual patterns, and machine learning prediction for supplier identification. This segmentation allows each component to specialize in specific document elements, improving overall accuracy when text recognition fails due to logos or graphics.
Solution Approach 2:
The machine learning model acts as an intermediary between the extracted features (textual and non-textual) and the final supplier identification output. It processes and synthesizes information from multiple sources, including visual features from logos and graphics, to predict supplier identity even when traditional text recognition fails.
2Measurement precision
If manual scanning of invoices is performed to identify supplier information, then identification accuracy is improved, but processing time and labor costs increase
Solution Approach 1:
The system enables self-service automated processing where the machine learning model independently analyzes document features and predicts supplier identity without human intervention. It processes both textual content and non-textual visual features automatically, eliminating the need for manual scanning while maintaining high accuracy.
Solution Approach 2:
The patent replaces manual mechanical scanning and visual inspection with automated digital processing. The machine learning model uses computational methods to analyze image-type documents, extracting features from logos, graphics, and text simultaneously, thereby substituting human labor with automated intelligent processing.
3Productivity
If existing machine learning models are used to predict supplier identity, then processing speed is improved, but prediction accuracy deteriorates when supplier fields are not explicitly present in the document
Solution Approach 1:
The system transitions from two-dimensional text-based analysis to three-dimensional multi-modal analysis by incorporating non-textual visual features such as logo patterns, graphic elements, and color information. This dimensional expansion enables the model to extract supplier identity information from visual patterns rather than relying solely on explicit text fields, maintaining speed while improving accuracy.
Solution Approach 2:
The machine learning model dynamically adjusts its feature extraction parameters based on document characteristics. When explicit supplier fields are absent, the model shifts to weighting visual and non-textual features more heavily, adapting its analysis parameters to maximize prediction accuracy under different document conditions without sacrificing processing speed.
Data Source
AI summary
Techniques for predicting a missing value in an image-type document are disclosed. A system predicts the identity of a supplier associated with an image-type document in which the supplier's identity may not be extracted by text recognition. When a system determines that the supplier identity cannot be identified using a text recognition application, the system generates a set of machine learning model input features from features extracted from the image-type document to predict the supplier's identity. One input feature is a data file bounds feature indicating whether the image-type document is a scanned document or a non-scanned document. The system predicts a value for the supplier's identity based on the data file bounds value and additional feature values, including color channel characteristics and spatial characteristics of regions-of-interest. The system generates a mapping of values to defined attributes based in part on the predicted value for the supplier's identity.


