Mobile Document Image Processing via Edge Detection and Perspective Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mobile devices face challenges in processing digital images of documents due to limited processing power and image resolution, leading to inefficient image capture and processing, especially with nonlinear distortions and projective effects introduced by camera optics, which are not handled effectively by conventional scanner-based algorithms.
Innovation Solution
The development of algorithms and systems that capture and process digital images on mobile devices by defining candidate edge points, determining field locations and types, and generating metadata labels to convert images into electronic form, allowing for efficient document detection and processing, even with limited resources, and enabling in-line conversion into usable electronic documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional scanner-based image processing algorithms are used on mobile devices, then image processing capability is improved, but processing time and computational cost become prohibitively expensive
Solution Approach 1:
The patent segments the image processing pipeline into distinct modules: document detection, perspective correction, and OCR. Each module processes specific aspects of the image independently, allowing for optimized computation at each stage rather than applying heavy conventional algorithms throughout the entire processing chain.
Solution Approach 2:
The patent applies partial action by implementing lightweight versions of processing algorithms that handle only the essential document capture requirements on mobile devices. Full conventional processing is applied only when necessary, such as when initial lightweight processing fails to achieve satisfactory results.
2Measurement precision
If conventional scanner-based image processing algorithms are used on mobile devices, then image processing capability is improved, but device resource consumption becomes prohibitively expensive
Solution Approach 1:
The processing pipeline is divided into energy-efficient stages: initial document detection using simple edge detection, followed by perspective correction only when needed, and finally OCR processing. This segmentation allows the system to consume minimal energy for common cases while maintaining full capability when required.
Solution Approach 2:
The system performs self-service by using the mobile device's existing camera and basic processing capabilities for the majority of the work, reserving heavy computational resources only for specific correction tasks when the document is captured at an angle or with distortion.
3Ease of operation
If mobile device cameras are used for document capture, then portability and ease of use are improved, but image quality and consistency deteriorate due to nonlinear distortions and projective effects
Solution Approach 1:
The patent introduces an intermediary perspective correction step between image capture and OCR processing. This correction module acts as a mediator that transforms the distorted camera image into a standardized view, eliminating projective effects and nonlinear distortions before the OCR engine processes the text.
Solution Approach 2:
The system dynamically changes image parameters through perspective correction, adjusting the geometric transformation parameters based on detected document boundaries. This allows the same camera to produce consistent, scanner-quality images regardless of capture angle or distance by modifying the spatial parameters of the captured image.
Data Source
AI summary
In various embodiments, methods, systems, and computer program products for capturing and processing digital images captured by a mobile device are disclosed. In one embodiment, a method includes capturing image data using a mobile device, the image data depicting a digital representation of a document; defining, based on the image data, a plurality of candidate edge points corresponding to the document; defining four sides of a tetragon based on at least some of the plurality of candidate edge points; determining a plurality of fields within the tetragon; for each field, determining at least a field location and a field data type; associating each determined field location with each field data type to generate a plurality of metadata labels; and associating the plurality of metadata labels with an image of an electronic form.


