Document Edge Detection and Rectification for Natural Scene Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Mobile devices often capture blurred, low-resolution images of documents with perspective effects, causing text to be deformed, rotated, and skewed, making it difficult to extract and align text fields accurately without human intervention or specialized equipment.

Innovation Solution

A processor-based method that detects document edges, computes mapping coordinates, and rectifies selected regions to remove background noise, aligning text fields upright and predictably, enabling efficient extraction and alignment of documents from natural scene images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a mobile device captures a document image, then the image can be obtained for processing, but the image is blurred and low resolution with perspective effects causing text deformation

Engineering Contradiction:
Improvetext extraction accuracyVSAvoidimage quality
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent replaces manual alignment and specialized equipment with an automated image processing system that uses edge detection, coordinate mapping, and geometric transformation algorithms to correct perspective distortion and extract text from low-quality images

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system transforms image parameters through coordinate mapping and geometric transformations, converting distorted perspective views into corrected, aligned document images with proper text orientation and positioning

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If users carefully align text with guidelines or use specialized equipment, then text extraction accuracy improves, but the process complexity and cost increase

Engineering Contradiction:
Improvetext alignment accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-alignment by automatically detecting document edges and computing transformation coordinates without requiring user intervention for manual alignment or specialized equipment, making the process autonomous and accessible with standard mobile devices

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent extracts only the essential document regions from the full image by detecting edges and computing mapping coordinates for selected regions, eliminating the need for processing entire images or using complex alignment procedures

Inventive Principle:
Principle #2Taking out (Extraction)

3Loss of information

If the entire image is processed, then all content is captured, but processing time and computational resources increase

Engineering Contradiction:
Improvedocument content completenessVSAvoidprocessing speed
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent segments the image processing task by detecting edges and identifying selected regions that contain document content, then applying transformation operations only to these regions rather than processing the entire image, thereby maintaining information completeness while improving processing efficiency

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8897565B1Extracting documents from a natural scene image
Publication Date: 2014.11.25 GOOGLE LLC
  • US8897565B1 patent drawing
  • US8897565B1 patent drawing
  • US8897565B1 patent drawing

AI summary

The present technology proposes techniques for extracting forms and other types of documents from images taken with a mobile client device. By calculating and making adjustments along a document's detected borders, an input image can be transformed such that the document within the image may be properly aligned and background clutter completely removed. The resulting text fields of the extracted document are thus upright, aligned and locatable at predictable points.