Confidence-Based Text Digitization for Illegible Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digitalization methods for physical documents, such as OCR, face challenges in accurately converting illegible, corrected, or imperfectly scanned text from historical documents, requiring significant human intervention and resources due to low confidence levels in text recognition.

Innovation Solution

A computer system analyzes images of physical documents, identifying text above a threshold confidence level for automatic digitization and presenting candidate strings for user selection on lower confidence texts, with optional manual input, and utilizes a database of related records to enhance recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If OCR is used to digitize text from physical documents, then the digitization process can be automated, but the accuracy of text recognition deteriorates due to illegible handwriting, corrections, and scanning errors

Engineering Contradiction:
Improveautomation of digitization processVSAvoidtext recognition accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary human reviewer who acts as a mediator between the automated OCR system and the final digitized output. The OCR system processes documents and generates candidate transcriptions, which are then reviewed and corrected by human experts before being finalized in the database. This intermediary step resolves the contradiction by maintaining automation while improving accuracy through human intervention on uncertain cases.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback mechanisms where human reviewers verify and correct OCR transcriptions, and these corrections are fed back into the system to improve future recognition. The confidence score threshold mechanism also provides feedback by routing low-confidence transcriptions for human review, creating a continuous improvement loop that maintains both automation and accuracy.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If manual review is performed on all transcribed text to ensure accuracy, then text recognition accuracy improves, but the time and resources required increase significantly

Engineering Contradiction:
Improvetext recognition accuracyVSAvoidtime for manual review
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies local quality by differentiating the level of review based on confidence scores. High-confidence transcriptions (above threshold) are accepted automatically without manual review, while low-confidence transcriptions (below threshold) receive focused human attention. This selective approach maintains accuracy where needed while minimizing time loss on reliable transcriptions.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the parameter of review intensity based on the confidence score parameter. By adjusting the threshold parameter, the system can dynamically control what portion of transcriptions require manual review. This parameter-based approach optimizes the balance between accuracy and time investment, ensuring high accuracy overall while reducing total review time compared to reviewing all text manually.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If a low confidence threshold is set for automatic acceptance, then the productivity of digitization increases, but the reliability of the transcribed data decreases

Engineering Contradiction:
Improvedigitization throughputVSAvoidtranscription reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The human reviewer serves as an intermediary quality control mechanism that allows the system to maintain high productivity with a low confidence threshold while ensuring reliability. Transcriptions below the threshold are automatically routed to human reviewers who verify and correct them before finalization, thus maintaining reliability without sacrificing the throughput benefits of the low threshold.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The confidence threshold mechanism creates a feedback loop where transcriptions failing to meet the threshold are automatically identified and routed for human review. This feedback ensures that reliability is maintained by catching errors in low-confidence transcriptions, while the majority of high-confidence transcriptions proceed automatically, preserving productivity.

Inventive Principle:
Principle #23Feedback

4Reliability

If a high confidence threshold is set for automatic acceptance, then the reliability of transcribed data improves, but the productivity of digitization decreases

Engineering Contradiction:
Improvetranscription reliabilityVSAvoiddigitization throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system optimizes the confidence threshold parameter to balance reliability and productivity. By setting the threshold at an appropriate level (e.g., 80-90%), the system maximizes automatic acceptance of reliable transcriptions while minimizing the portion requiring manual review. This parameter optimization ensures high overall reliability while maintaining productivity by avoiding overly conservative thresholds that would require excessive manual intervention.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8908971B2Devices, systems and methods for transcription suggestions and completions
Publication Date: 2014.12.09 ANCESTRY COM OPERATIONS INC
  • US8908971B2 patent drawing
  • US8908971B2 patent drawing
  • US8908971B2 patent drawing

AI summary

Methods, devices and systems are described for transcribing text from artifacts to electronic files. A computer system is provided, wherein the computer system comprises a computer-readable storage device. An image of the artifact is received wherein text is present on the artifact. A first portion of the text is analyzed. Characters representing the first portion of the text are identified at a first confidence level equal to or greater than a threshold confidence level. The characters representing the first portion of the text are stored. A second portion of the text appearing on the artifact is analyzed. A plurality of candidates to represent the second portion of the text are identified at a second confidence level below the threshold confidence level. Finally, the plurality of candidates to a user for selection are presented.