Text Extraction from Digital Images Using Segmentation Masks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image processing techniques struggle to accurately extract text from digital images containing graphical objects due to non-uniform contrast, low resolution, and the presence of sensitive information, often resulting in inaccurate text extraction and inadequate protection of proprietary data.
Innovation Solution
A method involving a mask generation process to separate textual and graphical objects in digital images, followed by character recognition on the transformed image to extract text, while also implementing a workflow approval process to manage sensitive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional OCR techniques are used on digital images, then text extraction can be performed, but the accuracy deteriorates due to non-uniform contrast and low resolution
Solution Approach 1:
The patent segments the digital image into multiple regions based on graphical object boundaries. By dividing the image into distinct regions (text regions and graphical object regions), the system can apply different processing techniques to each region, thereby improving text extraction accuracy while maintaining reliability even in low-resolution images with non-uniform contrast.
Solution Approach 2:
The patent applies different processing qualities to different parts of the image. Text regions receive enhanced processing with adjusted contrast and resolution parameters, while graphical object regions maintain their original characteristics. This local quality approach allows accurate text extraction from specific areas without degrading the overall image quality or requiring uniform processing across the entire image.
2Measurement precision
If imaging processing techniques are used to extract text from digital images with graphical objects, then text can be localized, but processing resources are significantly consumed and accuracy remains poor
Solution Approach 1:
The patent first segments the image to identify graphical object boundaries, then focuses processing only on regions containing text. This segmentation approach avoids applying computationally intensive processing to the entire image, thereby reducing processing resource consumption while maintaining accurate text localization in the relevant regions.
Solution Approach 2:
The patent applies processing actions selectively to only the portions of the image that contain text, rather than processing the entire image. By performing partial processing on text-containing regions only, the system achieves accurate text localization without the excessive resource consumption that would result from processing the complete image.
3Reliability
If techniques to prevent sensitive information access are implemented, then data security is improved, but adequate protection and classification of sensitive textual information in digital images is insufficient
Solution Approach 1:
The patent segments the image to separately identify text regions from graphical object regions. This segmentation enables the system to specifically detect, classify, and protect sensitive textual information (such as credit card numbers, social security numbers, account numbers) while maintaining the ability to process and share non-sensitive portions of the image, thereby providing adequate protection without complete information loss.
Solution Approach 2:
The patent applies different security measures to different regions of the image. Sensitive text regions receive enhanced protection measures such as redaction, masking, or access control, while non-sensitive regions maintain normal accessibility. This local quality approach to security provides adequate protection for sensitive information while preserving the usability of the overall image.
Data Source
AI summary
A method of extracting text from a digital image is provided. The method of extracting text includes receiving a digital image at an image processor where the digital image includes a textual object and a graphical object. A mask is generated based on the digital image. The mask includes a pattern having a first pattern area associated with the textual object and a second pattern area associated with the graphical object. The mask is applied to the digital image creating a transformed digital image. The transformed digital image includes a portion of the digital image associated with the textual object. Character recognition is performed on the portion of the digital image associated with the textual object of the transformed digital image to create a recognized text output.


