Dynamic Entity Recognition Engine for Document Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional document processing methods, such as pre-defined extractors, fail to accurately and efficiently process documents due to variations in document structure, layout, and field positions, leading to time-consuming and error-prone manual evaluation.
Innovation Solution
A machine learning-based entity recognition engine that uses a custom and dynamic named entity recognition framework, implemented through a combination of hardware and software, to identify and extract entities from documents by training a model with digitized document object models and tagged files, allowing for dynamic customization and feedback-driven improvements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional pre-defined extractors are used to process documents, then device complexity is reduced, but manufacturing precision (extraction accuracy) deteriorates due to document variations
Solution Approach 1:
The patent implements dynamic extractors that adapt to different document types and structures. The system trains machine learning models on specific document types and dynamically selects or configures extractors based on the input document characteristics, allowing the extraction mechanism to change its behavior to match the document structure rather than using a static pre-defined extractor for all cases.
Solution Approach 2:
The system changes parameters of the extractor based on document type and structure. By training models on specific document types and adjusting extractor configurations dynamically, the system adapts its extraction parameters to match the characteristics of each document, thereby maintaining high accuracy across varied document formats without requiring overly complex fixed extractors.
2Manufacturing precision
If manual evaluation is used to determine logical meaning and extract fields, then manufacturing precision (extraction accuracy) is improved, but productivity deteriorates due to time consumption
Solution Approach 1:
The system implements self-service through automated machine learning models that perform document evaluation and field extraction without human intervention. The trained models automatically determine logical meaning and extract fields from documents, replacing manual evaluation while maintaining high accuracy. This automation enables the system to process large volumes of documents efficiently without the time constraints of manual processing.
Solution Approach 2:
The patent incorporates feedback mechanisms where the system learns from trained data and continuously improves its extraction accuracy. By using training datasets and model refinement, the system captures the expertise of manual evaluators and embeds it in automated processes, allowing high-accuracy extraction to scale to high productivity levels through iterative learning and adaptation.
3Ease of operation
If conventional pre-defined extractors are used, then ease of operation is improved, but reliability deteriorates due to inability to handle document variations
Solution Approach 1:
The patent creates a universal document processing system that handles multiple document types and structures through a single platform. The system trains models on various document types and uses a unified extraction framework that adapts to different formats, maintaining ease of operation while achieving reliable consistent results across diverse documents through multi-functional capability.
Solution Approach 2:
The system dynamically adapts its extraction behavior based on document characteristics while maintaining a consistent user interface. The underlying processing logic changes dynamically to match document types, ensuring reliable results, while the ease of operation is preserved through a standardized interaction model that doesn't require users to understand the complex adaptive mechanisms.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Disclosed herein is a system. The system includes a memory and a processor. The memory stores processor executable instructions for a recognition engine. The processor is coupled to the memory. The processor executes the processor executable to cause the system to define a plurality of baseline entities to be identified from documents in a workflow and digitize the one or documents to generate corresponding document object models. The recognition engine further causes the system to train a model by using as inputs the corresponding document object models and tagged files and determine, using the model, plurality of target entities from target documents.