Text Recognition in Images Using Ranging Data Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Digital images containing text, especially from urban scenes, are challenging to automatically identify and recognize due to issues with image quality and environmental factors such as low resolution, distortions, and environmental conditions like shadows and contrast effects.
Innovation Solution
A system and method for text recognition in images involving preprocessing, candidate text region detection, enhancement, and optical character recognition, including superresolution processing to improve text identification accuracy and indexing for location-based searching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple low-resolution images are processed individually for text recognition, then the processing complexity is reduced, but the text recognition accuracy deteriorates due to low image quality and small text size
Solution Approach 1:
The patent combines multiple low-resolution images into a single high-resolution composite image by detecting corresponding text regions in each image, aligning them based on spatial coordinates, and superimposing them. This merging process allows the system to achieve high text recognition accuracy while avoiding the need to process each low-resolution image separately, thus resolving the contradiction between accuracy and processing complexity.
Solution Approach 2:
The patent transitions from processing multiple 2D low-resolution images individually to creating a single 2D high-resolution composite image by adding the dimension of image fusion. This dimensional transformation allows text regions that are too small to recognize in individual images to become sufficiently large and clear in the composite image, improving recognition accuracy without proportionally increasing processing complexity.
2Measurement precision
If the image resolution is increased to improve text recognition accuracy, then the text can be identified more accurately, but the image file size and processing requirements increase
Solution Approach 1:
Instead of capturing and storing a single high-resolution image that would require large data volume, the patent merges multiple existing low-resolution images to create a high-resolution composite. This approach achieves the benefit of high resolution for text recognition while avoiding the penalty of large image file sizes, as the composite is generated on-demand from smaller source images.
Solution Approach 2:
The patent creates a high-resolution view from multiple low-resolution views by adding the dimension of image fusion. This allows the system to provide high-resolution text recognition capability without permanently storing large high-resolution images, thus reducing the overall image data volume requirement while maintaining accuracy when needed.
3Measurement precision
If text regions are enhanced by combining multiple images, then the text recognition accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
The patent segments the image processing task by first detecting candidate text regions in each input image separately, then combining only those specific regions into the composite image. This segmentation approach avoids the need to process and combine entire images, significantly reducing processing time and computational resources while still achieving high text recognition accuracy for the regions of interest.
Solution Approach 2:
The patent extracts only the relevant text-containing regions from multiple images rather than processing the entire images. By detecting candidate text regions and extracting only those portions for combination, the system achieves enhanced text recognition accuracy while minimizing processing time and computational resource requirements by working with smaller extracted regions instead of full images.
Data Source
AI summary
Methods, systems, and apparatus including computer program products for recognizing text in images are provided. In one implementation, a computer-implemented method for recognizing text in an image is provided. The method includes receiving a plurality of images. The method also includes processing the images to detect a corresponding set of regions of the images, each image having a region corresponding to each other image region, as potentially containing text. The method further includes combining the regions to generate an enhanced region image and performing optical character recognition on the enhanced region image.


