Document Edge Detection and Rectification for Natural Scene Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mobile devices often capture blurred, low-resolution images of documents with perspective effects, causing text to be deformed, rotated, and skewed, making it difficult to extract and align text fields accurately without human intervention or specialized equipment.
Innovation Solution
A processor-based method that detects document edges, computes mapping coordinates, and rectifies selected regions to remove background noise, aligning text fields upright and predictably, enabling efficient extraction and alignment of documents from natural scene images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a mobile device captures a document image, then the image can be obtained for processing, but the image is blurred and low resolution with perspective effects causing text deformation
Solution Approach 1:
The patent replaces manual alignment and specialized equipment with an automated image processing system that uses edge detection, coordinate mapping, and geometric transformation algorithms to correct perspective distortion and extract text from low-quality images
Solution Approach 2:
The system transforms image parameters through coordinate mapping and geometric transformations, converting distorted perspective views into corrected, aligned document images with proper text orientation and positioning
2Measurement precision
If users carefully align text with guidelines or use specialized equipment, then text extraction accuracy improves, but the process complexity and cost increase
Solution Approach 1:
The system performs self-alignment by automatically detecting document edges and computing transformation coordinates without requiring user intervention for manual alignment or specialized equipment, making the process autonomous and accessible with standard mobile devices
Solution Approach 2:
The patent extracts only the essential document regions from the full image by detecting edges and computing mapping coordinates for selected regions, eliminating the need for processing entire images or using complex alignment procedures
3Loss of information
If the entire image is processed, then all content is captured, but processing time and computational resources increase
Solution Approach 1:
The patent segments the image processing task by detecting edges and identifying selected regions that contain document content, then applying transformation operations only to these regions rather than processing the entire image, thereby maintaining information completeness while improving processing efficiency
Data Source
AI summary
The present technology proposes techniques for extracting forms and other types of documents from images taken with a mobile client device. By calculating and making adjustments along a document's detected borders, an input image can be transformed such that the document within the image may be properly aligned and background clutter completely removed. The resulting text fields of the extracted document are thus upright, aligned and locatable at predictable points.


