Small-Text Detection via Local Property Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods fail to effectively detect and process small-font sized text in natural scenes due to difficulties in distinguishing small text from noise and blur, especially when the text is 3-5 pixels in height, as they rely on large visual features that are not evident in such cases.
Innovation Solution
A method that generates input maps reflecting local properties like contrast, gradient, and activity levels to detect small-font sized text by integrating information from active pixel maps, dominant direction maps, and high contrast masks, and processes these regions differently using interpolation and optical zoom for enhanced clarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional methods based on large visual features are used to detect text, then detection is effective for large text, but detection accuracy deteriorates for small text (3-5 pixels in height)
Solution Approach 1:
The patent applies local quality by generating multiple input maps (gradient magnitude map, gradient direction map, Laplacian map, dominant direction map) that capture different local properties of text regions. Each map highlights specific local characteristics such as edge strength, orientation, and curvature, enabling the system to detect small text by integrating these localized features rather than relying on global text patterns.
Solution Approach 2:
The patent transforms the detection problem from relying solely on spatial dimensions to incorporating directional and frequency dimensions. By computing gradient directions, dominant orientations, and applying Laplacian filters, the system adds dimensional information about text structure and orientation, enabling detection of small text that lacks sufficient spatial extent for conventional methods.
2Reliability
If conventional image processing techniques (denoising, sharpening, super-resolution) are applied to text regions, then natural scene areas are enhanced, but text information suffers from undesired artifacts
Solution Approach 1:
The patent implements local quality by creating a text detection mask based on integrated input map analysis, which identifies specific regions containing small text. This mask enables selective application of processing techniques, allowing the system to preserve original text regions while applying enhancement only to non-text areas, thereby preventing artifacts in text regions while still improving overall image quality.
Solution Approach 2:
The patent segments the image into text regions and non-text regions using the detection mask generated from input maps. This segmentation allows differential processing where text regions are protected from conventional enhancement operations that would introduce artifacts, while non-text regions receive full processing benefits.
3Measurement precision
If optical zoom is applied to increase text clarity, then small text becomes more readable, but processing time and complexity increase
Solution Approach 1:
The patent applies local quality by generating an optical zoom mask that identifies only those regions containing small text. This enables selective optical zoom processing where only the detected small text regions are enhanced through optical zoom, while the rest of the image remains unchanged, significantly reducing processing time and computational complexity compared to applying optical zoom to the entire image.
Solution Approach 2:
The patent segments the image into regions requiring optical zoom enhancement and regions that do not, based on the text detection results from input maps. This segmentation allows the system to apply computationally intensive optical zoom processing only to necessary regions, optimizing the balance between text readability improvement and processing efficiency.
Data Source
AI summary
A method for recognizing small-font sized text including receiving digital media of a natural scene, the digital media having at least one frame that includes the small-font sized text; generating input maps having values that reflect local properties of corresponding regions in the at least one frame; and detecting regions of the at least one frame that contain the small-font sized text by integrating information from the input maps. The integrated information may include information located between border lines having active pixels therebetween and gaps having a high ratio of non-ink pixels located below a bottom border line and above a top border line in relation to a dominant direction of the text. The active pixels may be pixels having dense changes in character stroke directions.


