Secure Data Entry via Document Segmentation and Randomization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data entry methods, including conventional, outsourcing, and partial automation, are complex, time-consuming, error-prone, and fail to protect sensitive information from exposure and misuse, particularly in offshore operations where security risks and errors are high.
Innovation Solution
The method breaks down scanned document images into small pieces called 'confetti' which are randomly distributed and processed by unskilled workers, ensuring that sensitive information cannot be reassembled, combining automation with human quality checking for secure and efficient data entry.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional or outsourced manual data entry is used, then data can be entered with human judgment and flexibility, but the process becomes complex, time-consuming, error-prone, and exposes sensitive information to security risks
Solution Approach 1:
The patent segments scanned document images into multiple small non-overlapping tiles that are randomly distributed to different workers. Each worker processes only a subset of tiles, making it computationally infeasible to reconstruct the original document from any single worker's data, thus solving both efficiency and security/accuracy contradictions
Solution Approach 2:
The patent introduces an intermediary processing system that receives tiled image fragments, performs automated processing on each tile, and reassembles the results. This intermediary layer enables automated high-speed processing while maintaining data integrity through controlled reconstruction, resolving the contradiction between manual flexibility and automated efficiency
2Manufacturing precision
If skilled workers are used for data entry, then data quality and accuracy improve, but costs increase and security risks increase due to exposure of sensitive information
Solution Approach 1:
By dividing documents into small tiles and randomly distributing them, the system eliminates the need for skilled workers to understand entire documents or contexts. Each worker simply processes isolated tile fragments, removing the complexity of training while maintaining accuracy through automated processing and validation of individual tiles
Solution Approach 2:
The patent treats each tile as a disposable, independent processing unit that is processed once and then discarded or securely deleted. This approach replaces expensive, highly-trained workers with simple, automated tile processing that requires minimal expertise, reducing both training complexity and security management overhead
3Productivity
If document images are distributed to data entry workers, then data entry can be performed, but sensitive information is exposed and can be misused or stolen
Solution Approach 1:
The patent segments complete document images into many small non-overlapping tiles and randomly distributes them to different workers. The mathematical improbability of any single worker receiving enough contiguous tiles to reconstruct a meaningful document provides computational security, enabling high throughput while minimizing information exposure and security risks
Solution Approach 2:
The patent converts the potential harm of data exposure into a benefit by using random tile distribution as a security mechanism. The very act of distributing data fragments to multiple workers, which could enable theft or misuse, instead becomes a protective measure that makes reconstruction infeasible, thus converting the security vulnerability into a security feature
Data Source
AI summary
The present invention includes a method of secure data entry that enables complex data entry work to be performed by unskilled workers that results in data entry with higher productivity, higher quality and higher security than data entry performed by highly skilled workers. The invention identifies data fields on an electronic image of an identified input page, sequences identified data field images, and individually displays data field images for manual data entry. The invention also provides for extracting data from a data field image and displaying extracted data along with the corresponding data field image for approval or correction. Sequenced data field images are optionally reordered or randomized for display and manual entry.


