Non-Semantic Entity Extraction from Document Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional OCR text extraction methods face challenges in accurately extracting non-semantic text from documents, which is not defined in dictionaries, requiring significant manual effort and lacking efficient automated processes.
Innovation Solution
A method and system for extracting non-semantic entities from document images using a processor-based approach that splits entities into alphabetic and numeric components, determines feature values, and applies prediction techniques to label entities, leveraging trained data for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional OCR text extraction methods are used, then text extraction can be performed, but non-semantic text extraction accuracy is poor and manual effort is required
Solution Approach 1:
The system enables automated extraction of non-semantic text entities without requiring manual intervention. The processor automatically detects, extracts, and classifies entities such as dates, times, percentages, and other non-semantic text elements using integrated prediction models, replacing the need for manual annotation and template creation.
Solution Approach 2:
The patent replaces manual mechanical processes of text extraction with an automated computational system. Instead of manual review and template-based extraction, the system uses machine learning models and algorithms to automatically identify and extract non-semantic entities from document images.
2Ease of manufacture
If templates are formed from text data for future extraction, then extraction can be standardized, but significant manual effort is required
Solution Approach 1:
The system performs preliminary actions by automatically training prediction models on extracted text data before actual extraction tasks. The models learn patterns and characteristics of non-semantic entities during training, enabling standardized automated extraction without requiring manual template creation for each specific extraction task.
Solution Approach 2:
The patent transforms the extraction process by changing from static template-based parameters to dynamic model-based parameters. The system adapts to different document types and extraction requirements through trained prediction models that can adjust their extraction criteria based on learned patterns rather than fixed templates.
3Productivity
If automated extraction models are implemented, then extraction efficiency improves, but system complexity increases
Solution Approach 1:
The system achieves multi-functionality by using a single integrated prediction model framework that can extract various types of non-semantic entities (dates, times, percentages, monetary values, etc.) from different document types. This universal approach improves extraction efficiency across multiple tasks while avoiding the need for separate specialized models for each entity type.
Solution Approach 2:
The patent introduces prediction models as intermediary components between the input document images and the extracted entity data. These models act as mediators that process visual information and transform it into structured entity extracts, simplifying the overall system architecture compared to direct template-matching approaches.
Data Source
AI summary
A method and system of extracting one or more non-semantic entities in a document image including data entities is disclosed. The methodology includes extraction, by a processor, of row entities and corresponding row location based on a text extraction technique from the document image. The row entities are split into split-row entities based on a splitting rule. Semantic entities are determined from alphabetic entities using semantic recognition technique. The non-semantic entities are determined as split-row entities other than semantic entities. Feature values of each feature type for each of the non-semantic entities is determined. The processor further determines a first probability output for non-semantic entities and a second probability output for semantic entities surrounding the non-semantic entities. The system further labels each of the non-semantic entities based on determination of a highest probability value from a sum of the first probability output and the second probability output.


