Dynamic Entity Recognition Engine for Document Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional document processing methods, such as pre-defined extractors, fail to accurately and efficiently process documents due to variations in document structure, layout, and field positions, leading to time-consuming and error-prone manual evaluation.

Innovation Solution

A machine learning-based entity recognition engine that uses a custom and dynamic named entity recognition framework, implemented through a combination of hardware and software, to identify and extract entities from documents by training a model with digitized document object models and tagged files, allowing for dynamic customization and feedback-driven improvements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If conventional pre-defined extractors are used to process documents, then device complexity is reduced, but manufacturing precision (extraction accuracy) deteriorates due to document variations

Engineering Contradiction:
Improveextractor complexityVSAvoidextraction accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent implements dynamic extractors that adapt to different document types and structures. The system trains machine learning models on specific document types and dynamically selects or configures extractors based on the input document characteristics, allowing the extraction mechanism to change its behavior to match the document structure rather than using a static pre-defined extractor for all cases.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes parameters of the extractor based on document type and structure. By training models on specific document types and adjusting extractor configurations dynamically, the system adapts its extraction parameters to match the characteristics of each document, thereby maintaining high accuracy across varied document formats without requiring overly complex fixed extractors.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If manual evaluation is used to determine logical meaning and extract fields, then manufacturing precision (extraction accuracy) is improved, but productivity deteriorates due to time consumption

Engineering Contradiction:
Improveextraction accuracyVSAvoiddocument processing throughput
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system implements self-service through automated machine learning models that perform document evaluation and field extraction without human intervention. The trained models automatically determine logical meaning and extract fields from documents, replacing manual evaluation while maintaining high accuracy. This automation enables the system to process large volumes of documents efficiently without the time constraints of manual processing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where the system learns from trained data and continuously improves its extraction accuracy. By using training datasets and model refinement, the system captures the expertise of manual evaluators and embeds it in automated processes, allowing high-accuracy extraction to scale to high productivity levels through iterative learning and adaptation.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If conventional pre-defined extractors are used, then ease of operation is improved, but reliability deteriorates due to inability to handle document variations

Engineering Contradiction:
Improveextractor usabilityVSAvoidprocessing consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent creates a universal document processing system that handles multiple document types and structures through a single platform. The system trains models on various document types and uses a unified extraction framework that adapts to different formats, maintaining ease of operation while achieving reliable consistent results across diverse documents through multi-functional capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adapts its extraction behavior based on document characteristics while maintaining a consistent user interface. The underlying processing logic changes dynamically to match document types, ensuring reliable results, while the ease of operation is preserved through a standardized interaction model that doesn't require users to understand the complex adaptive mechanisms.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4187452A1Machine learning based entity recognition
Publication Date: 2023.05.31 UIPATH INC
  • EP4187452A1 patent drawingFigure 1
  • EP4187452A1 patent drawingFigure 2
  • EP4187452A1 patent drawingFigure 3

AI summary

Disclosed herein is a system. The system includes a memory and a processor. The memory stores processor executable instructions for a recognition engine. The processor is coupled to the memory. The processor executes the processor executable to cause the system to define a plurality of baseline entities to be identified from documents in a workflow and digitize the one or documents to generate corresponding document object models. The recognition engine further causes the system to train a model by using as inputs the corresponding document object models and tagged files and determine, using the model, plurality of target entities from target documents.