Dynamic Data Extraction Model Updating via User Validation Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional document data extraction systems rely on manual processes, which are time-consuming and prone to errors, and existing prediction models struggle with accuracy when handling documents with varying types of data and layouts.

Innovation Solution

A system that generates data extraction predictions using a prediction model, allows user validation to correct predictions, and uses this validation to generate updated prediction models, improving efficiency and accuracy by creating training documents from user corrections.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual processes are used for document data extraction, then the system can handle various document types, but the process becomes time-consuming and error-prone

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical data extraction processes with an automated prediction model system. The system uses machine learning models to automatically extract data from documents, substituting human manual work with computational processes that are both faster and more accurate when properly maintained.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements a feedback mechanism where user corrections to extraction predictions are captured and used to retrain and update the prediction models. This continuous feedback loop improves model accuracy over time while maintaining automated processing speed, resolving the contradiction between accuracy and time efficiency.

Inventive Principle:
Principle #23Feedback

2Productivity

If existing prediction models are used for data extraction, then automation is achieved, but accuracy deteriorates when handling documents with varying types and layouts

Engineering Contradiction:
Improveextraction efficiencyVSAvoidprediction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent makes the prediction model dynamic by implementing continuous updating mechanisms. Instead of using static models, the system dynamically adapts to varying document types and layouts by retraining models with new data, allowing it to maintain high accuracy across diverse document formats while preserving automation benefits.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the prediction models through continuous retraining with updated data from user corrections. By modifying model parameters based on real-world feedback, the system adapts to different document types and layouts, maintaining both productivity and precision.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If prediction models are continuously updated with user feedback, then accuracy improves, but system complexity increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a self-service mechanism where the system automatically uses user feedback to retrain and update its own prediction models without requiring manual intervention. The system serves itself by automatically incorporating corrections into model improvements, reducing the operational complexity despite the sophisticated update mechanisms.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240028914A1Method and system for maintaining a data extraction model
Publication Date: 2024.01.25 DELL PROD LP
  • US20240028914A1 patent drawing
  • US20240028914A1 patent drawing
  • US20240028914A1 patent drawing

AI summary

Techniques described herein relate to a method for performing data extraction for documents. The method includes obtaining a data extraction request associated with a document; in response to obtaining the request: generating a data extraction prediction using a prediction model and the document; providing the data extraction prediction to a user; obtaining a user validation associated with the data extraction prediction; making a determination that the user validation indicates that the data extraction prediction is not correct; in response to the determination: generating an updated data extraction prediction based on the user validation; and initiating performance of additional document processing using the document based on the updated data extraction prediction.