Text Detection in Images Using Multi-Version Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current OCR technologies fail to effectively detect and recognize text in images captured by mobile devices due to poor quality caused by blur, noise, and lighting variations, especially when fonts are unusual or backgrounds are non-uniform.
Innovation Solution
A method and system that applies a series of image-processing techniques to detect text regions, reduce blurring and lighting effects, and create multiple versions of the same text region, which are then sent to an optical character-recognition system for combination into a single text result.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional OCR technology is used on images captured by mobile devices, then the system is simple and easy to implement, but the text detection and recognition accuracy deteriorates due to blur, noise, and lighting variations
Solution Approach 1:
The patent segments the text detection and recognition process into multiple distinct stages: blur detection and correction, noise reduction, lighting variation compensation, and finally OCR. This segmentation allows each stage to address specific image quality issues independently, improving overall text detection accuracy while maintaining manageable system complexity through modular processing.
Solution Approach 2:
The patent applies preliminary image processing actions before OCR recognition, including blur correction, noise reduction, and lighting normalization. These preliminary actions prepare the image data in advance to ensure optimal conditions for the subsequent OCR process, thereby improving text detection accuracy without requiring complex real-time processing during recognition.
2Measurement precision
If multiple image processing techniques are applied to correct blur and lighting variations, then text recognition accuracy improves, but processing time and computational resources increase
Solution Approach 1:
The patent applies partial image processing techniques by selectively correcting only the most significant degradations in each image based on their specific characteristics. Rather than applying all possible processing techniques uniformly, the system identifies and addresses the predominant issues (blur, noise, or lighting variations) to achieve sufficient recognition accuracy with reduced processing time.
Solution Approach 2:
The patent performs preliminary analysis of image quality characteristics to determine which processing techniques are most needed before applying them. This preliminary assessment allows the system to prioritize and apply only the necessary corrections, reducing unnecessary computational overhead and processing time while maintaining high text recognition accuracy.
3Illumination intensity
If the flash on cell phone cameras is used to illuminate text, then lighting conditions improve, but the flash creates additional illumination variations and is not strong enough
Solution Approach 1:
The patent introduces image processing algorithms as an intermediary to compensate for insufficient and uneven flash illumination. These algorithms analyze the lighting conditions and apply computational corrections to normalize illumination across the image, effectively mediating between the inadequate physical flash output and the requirements for uniform text illumination.
Solution Approach 2:
The patent replaces reliance on the physical flash mechanism with computational image processing methods to achieve proper illumination. Instead of depending on the flash's mechanical light output, the system uses software-based lighting correction techniques to achieve uniform illumination, thereby avoiding the flash's inherent limitations of insufficient strength and uneven distribution.
Data Source
AI summary
Various embodiments of the present invention relate to a method, system and computer program product for detecting and recognizing text in the images captured by cameras and scanners. First, a series of image-processing techniques is applied to detect text regions in the image. Subsequently, the detected text regions pass through different processing stages that reduce blurring and the negative effects of variable lighting. This results in the creation of multiple images that are versions of the same text region. Some of these multiple versions are sent to a character-recognition system. The resulting texts from each of the versions of the image sent to the character-recognition system are then combined to a single result, wherein the single result is detected text.


