OCR Character Correction Randomization for Data Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing optical character recognition (OCR) systems face challenges in accurately capturing and correcting form data, particularly when the data includes confidential information, as manual correction methods can compromise privacy and conventional systems lack effective mechanisms to maintain data confidentiality during the correction process.
Innovation Solution
A system is implemented that randomizes the assignment of characters to reviewers for correction, using an autonomous process to ensure privacy by randomly distributing characters that do not satisfy a confidence condition, allowing for secure and efficient correction of OCR errors without revealing sensitive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual correction methods are used to correct OCR errors in form data, then the accuracy of character recognition is improved, but data privacy and confidentiality are compromised
Solution Approach 1:
The patent segments the correction process by separating the OCR output from the form data structure. Characters are extracted as independent units and corrected individually, allowing the system to maintain privacy by not exposing complete form data while still enabling accurate correction of recognition errors.
Solution Approach 2:
The patent introduces an intermediary correction mechanism that operates on extracted characters without requiring access to the original form data. The system uses a correction model that can identify and fix errors in character sequences without exposing the underlying confidential information, thus mediating between accuracy improvement and privacy protection.
2Productivity
If conventional OCR systems are used to capture form data, then data capture is achieved, but processing errors occur and data confidentiality cannot be maintained during correction
Solution Approach 1:
The patent extracts characters from the form data as independent units for correction processing. By taking out only the necessary character information rather than processing complete form data, the system maintains confidentiality while enabling efficient correction of OCR errors through automated character-level processing.
3Measurement precision
If proprietary form-identification models and form-specific field parsers are used, then form data capture accuracy is improved, but system complexity increases and privacy protection becomes more difficult
Solution Approach 1:
The patent creates a universal correction mechanism that can handle various form types without requiring form-specific correction models. The character extraction and correction process is designed to be form-agnostic, reducing system complexity while maintaining high accuracy across different form structures and data types.
Data Source
AI summary
A method for maintaining data privacy includes associating each image of a group of images with a digital character representation to generate a plurality of groups of digital character representations, each one of the plurality of groups of digital character representations associated with a respective data field of a group of data field from one or more documents. The method also includes randomly assigning, to each reviewer of a set of reviewers, a respective subset of digital character representations from the plurality of groups of digital character representations, each digital character representation of the subset failing to satisfy a confidence condition. The method further includes receiving, from each reviewer of the set of reviewers, a message indicating, for each digital character representation of the respective subset of digital character representations, whether the digital character representation is correct or includes a correction.


