Document Information Extraction Using Orthogonal Image Capture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The ability to extract useful information from images of documents using optical character recognition is often restricted by image quality, leading to errors and requiring users to perform time-consuming post-acquisition operations such as editing and re-capturing images.
Innovation Solution
An electronic device captures multiple images of a document at different orientations using an integrated imaging device, with associated timestamps and spatial-position information, and performs error correction based on the images to improve information extraction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single image is captured for information extraction, then the image capture process is quick, but the extraction accuracy is low leading to errors
Solution Approach 1:
The system automatically captures multiple images at different orientations during the initial image capture phase, before the user performs any post-acquisition operations. This preliminary multi-orientation capture ensures that high-quality images for accurate information extraction are obtained without requiring subsequent manual intervention or re-capture operations.
Solution Approach 2:
The system dynamically adjusts the image capture process by automatically rotating the device to capture images at multiple orientations (e.g., 0°, 45°, 90°, 135°). This dynamic multi-orientation approach ensures that at least one image will have optimal alignment with the document text, thereby improving extraction accuracy without requiring manual user intervention.
2Measurement precision
If multiple images are captured at different orientations, then the information extraction accuracy is improved, but the image capture process becomes more complex
Solution Approach 1:
The system performs automatic image capture at multiple orientations without requiring manual user intervention. The device autonomously rotates, captures images, and processes them to extract information. This self-service approach simplifies the user experience while maintaining the benefits of multi-orientation capture for improved accuracy.
Solution Approach 2:
The system replaces manual mechanical operations (user manually rotating and re-capturing images) with automated computational control. The device uses software-controlled rotation and automated image selection algorithms to determine the best-oriented image for extraction, substituting complex manual mechanical processes with streamlined automated systems.
3Measurement precision
If post-acquisition operations are performed to correct errors, then the extracted information accuracy is improved, but the overall process time increases
Solution Approach 1:
The system performs error prevention through preliminary multi-orientation image capture, ensuring that high-quality images are obtained before any extraction attempts. By capturing multiple orientations in advance, the system eliminates the need for subsequent error correction operations, thereby maintaining high accuracy while preserving processing throughput.
Solution Approach 2:
The system uses feedback from image quality assessment to automatically select the best-oriented image for information extraction. By analyzing the captured images and identifying the one with optimal text alignment and quality metrics, the system feeds back this selection to the extraction module, ensuring high accuracy without requiring manual user intervention or iterative correction operations.
Data Source
AI summary
During an information-extraction technique, a user of an electronic device may be instructed by an application executed by the electronic device (such as a software application) to acquire images, with different orientations (which are known to the user), of a target location on a document using an imaging sensor, which is integrated into the electronic device. After the user has taken a first image and before the user takes a second image in a different orientation of the electronic device (and, thus, the imaging sensor), the electronic device captures multiple images of the document. Then, the electronic device stores the images with associated timestamps. Moreover, after the user has taken the second image, the electronic device analyzes one or more of the first image, the second image and at least a subset of the images to extract information proximate to the target location on the document.


