Text Detection in Images Using Multi-Version Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current OCR technologies fail to effectively detect and recognize text in images captured by mobile devices due to poor quality caused by blur, noise, and lighting variations, especially when fonts are unusual or backgrounds are non-uniform.

Innovation Solution

A method and system that applies a series of image-processing techniques to detect text regions, reduce blurring and lighting effects, and create multiple versions of the same text region, which are then sent to an optical character-recognition system for combination into a single text result.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional OCR technology is used on images captured by mobile devices, then the system is simple and easy to implement, but the text detection and recognition accuracy deteriorates due to blur, noise, and lighting variations

Engineering Contradiction:
Improvetext detection accuracyVSAvoidimage processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the text detection and recognition process into multiple distinct stages: blur detection and correction, noise reduction, lighting variation compensation, and finally OCR. This segmentation allows each stage to address specific image quality issues independently, improving overall text detection accuracy while maintaining manageable system complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary image processing actions before OCR recognition, including blur correction, noise reduction, and lighting normalization. These preliminary actions prepare the image data in advance to ensure optimal conditions for the subsequent OCR process, thereby improving text detection accuracy without requiring complex real-time processing during recognition.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple image processing techniques are applied to correct blur and lighting variations, then text recognition accuracy improves, but processing time and computational resources increase

Engineering Contradiction:
Improvetext recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial image processing techniques by selectively correcting only the most significant degradations in each image based on their specific characteristics. Rather than applying all possible processing techniques uniformly, the system identifies and addresses the predominant issues (blur, noise, or lighting variations) to achieve sufficient recognition accuracy with reduced processing time.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs preliminary analysis of image quality characteristics to determine which processing techniques are most needed before applying them. This preliminary assessment allows the system to prioritize and apply only the necessary corrections, reducing unnecessary computational overhead and processing time while maintaining high text recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

3Illumination intensity

If the flash on cell phone cameras is used to illuminate text, then lighting conditions improve, but the flash creates additional illumination variations and is not strong enough

Engineering Contradiction:
Improvetext illuminationVSAvoidillumination variations
Core Design Contradiction:
Illumination intensityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces image processing algorithms as an intermediary to compensate for insufficient and uneven flash illumination. These algorithms analyze the lighting conditions and apply computational corrections to normalize illumination across the image, effectively mediating between the inadequate physical flash output and the requirements for uniform text illumination.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces reliance on the physical flash mechanism with computational image processing methods to achieve proper illumination. Instead of depending on the flash's mechanical light output, the system uses software-based lighting correction techniques to achieve uniform illumination, thereby avoiding the flash's inherent limitations of insufficient strength and uneven distribution.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8977072B1Method and system for detecting and recognizing text in images
Publication Date: 2015.03.10 AMAZON TECH INC
  • US8977072B1 patent drawing
  • US8977072B1 patent drawing
  • US8977072B1 patent drawing

AI summary

Various embodiments of the present invention relate to a method, system and computer program product for detecting and recognizing text in the images captured by cameras and scanners. First, a series of image-processing techniques is applied to detect text regions in the image. Subsequently, the detected text regions pass through different processing stages that reduce blurring and the negative effects of variable lighting. This results in the creation of multiple images that are versions of the same text region. Some of these multiple versions are sent to a character-recognition system. The resulting texts from each of the versions of the image sent to the character-recognition system are then combined to a single result, wherein the single result is detected text.