Document Digitization System Using Segmented OCR Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional digitization methods using crowdsourcing platforms are costly due to repetitive verification processes and inconsistent quality, leading to high costs for crowdworkers and platforms for document digitization tasks.
Innovation Solution
A method and system that classify document portions based on character recognition status, creating different types of tasks (data validation, editing, and entry) with varying costs, and assigning these tasks to crowdworkers based on their complexity, optimizing the workflow and reducing unnecessary efforts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If repetitive verification processes are used to assure quality, then digitization quality is improved, but cost increases
Solution Approach 1:
The patent segments the verification process into two distinct levels: automated OCR-based initial recognition and selective crowdworker verification. Only portions that fail automated recognition or fall below quality thresholds are forwarded to crowdworkers, rather than requiring all portions to undergo manual verification. This segmentation resolves the contradiction by maintaining quality assurance through automated means while minimizing costly manual intervention to only when necessary.
Solution Approach 2:
The system enables self-service through automated OCR recognition and self-verification mechanisms. The automated OCR system performs initial digitization and quality checking without human intervention, and the system automatically identifies and routes only problematic portions to crowdworkers. This self-service approach reduces reliance on expensive repetitive manual verification while maintaining quality standards, directly addressing the cost-quality contradiction.
2Reliability
If all document portions are processed through crowdworkers, then quality is assured, but cost increases significantly
Solution Approach 1:
The patent applies local quality by differentiating the processing approach based on the specific characteristics of each document portion. High-confidence portions identified by automated OCR receive minimal or no crowdworker intervention, while low-confidence or problematic portions receive focused manual verification. This localized application of quality assurance resources maintains overall digitization quality while dramatically reducing costs compared to universal crowdworker processing.
Solution Approach 2:
The system employs partial action by applying crowdworker verification only to the extent necessary - specifically to portions that fail automated recognition or fall below quality thresholds. Rather than applying excessive verification to all portions, the system applies verification partially and selectively where needed, optimizing the balance between quality assurance and cost-effectiveness.
3Measurement precision
If character recognition techniques are applied to all portions, then identification accuracy improves, but processing time and cost increase
Solution Approach 1:
The patent applies preliminary action through automated OCR character recognition on all document portions before crowdworker involvement. This preliminary automated recognition establishes a baseline accuracy for all portions, and only portions that fail this preliminary step or fall below quality thresholds are then subjected to additional manual verification. This preliminary automated pass reduces overall processing time while maintaining identification accuracy for the majority of portions.
Data Source
AI summary
A method and a system for digitization of a document are disclosed. The document is scanned to generate an electronic document. One or more characters in a first set of portions of the electronic document are identified, based on a character recognition technique. Each portion in the first set of portions is classified in one or more groups based on at least a status of identification of the corresponding one or more characters. Further, one or more tasks are created for each of the one or more groups. The one or more tasks are transmitted to one or more crowdworkers, based at least on the respective type of the one or more tasks. Further, a response for each of the one or more tasks is received. Based on the received response, a digitized document is generated.


