Confidence Score Calculation for Data Extraction Automation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data extraction models lack accuracy and the ability to re-train based on ground truth, diminishing their predictive impact and requiring human intervention for verification, which is time-consuming and reduces productivity.

Innovation Solution

A system that calculates a reconfigured confidence score by receiving inputs from multiple models, assigning weightage, and generating output confidence scores based on text and labels, allowing for the selection of the most accurate information and providing a final confidence score through an ensemble model that includes a database for additional information storage and retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If current extraction models are used, then data extraction can be automated, but the accuracy is insufficient and requires human intervention

Engineering Contradiction:
Improveautomation of data extractionVSAvoidextraction accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent combines multiple extraction models into an ensemble system where each model processes the input data independently and their results are aggregated. This merging of multiple models' capabilities allows the system to maintain high automation while improving extraction accuracy through collective decision-making, resolving the contradiction between automation extent and measurement precision.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements a feedback mechanism where confidence scores from individual models are calculated and compared, and the final output is selected based on the highest confidence score. This feedback loop allows the system to automatically verify and validate extraction results, maintaining high automation while ensuring improved accuracy through self-verification.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If multiple models are used to improve accuracy, then extraction precision increases, but system complexity increases

Engineering Contradiction:
Improveextraction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the extraction system into independent modular models, each handling specific extraction tasks independently. Each model generates its own confidence score and can be independently trained and optimized. This segmentation allows the system to achieve high extraction accuracy through multiple specialized models while managing complexity through modular, independent components that can be developed and maintained separately.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If manual verification is performed to ensure accuracy, then extraction precision improves, but productivity decreases

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-verification by automatically calculating confidence scores for each model's extraction results and selecting the output with the highest confidence score. This self-service mechanism eliminates the need for manual verification while maintaining high extraction accuracy, thereby preserving productivity and avoiding the productivity loss that would result from human intervention.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12008024B2System to calculate a reconfigured confidence score
Publication Date: 2024.06.11 INFRRD INC
  • US12008024B2 patent drawing
  • US12008024B2 patent drawing
  • US12008024B2 patent drawing

AI summary

A system to calculate a reconfigured confidence score is configured to receive a text, a plurality of labels, and a plurality of confidence scores from a plurality of models and assign a weightage to the inputs received from the plurality of models. The system is configured to select a first text with a first label and retrieve a second text, a third text, and a second label. The system is further configured to generate a first, second and third output confidence score for the first text, second text and third text, and corresponding labels. The system compares the plurality of output confidence scores and generates an output which comprises of the first text, the first label, and a final confidence score, wherein the final confidence score is one among the first, second and third output confidence scores.