Text Recognition in Images Using Preprocessing and Confidence Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic devices face challenges in accurately identifying and processing text from images or videos, particularly under poor lighting conditions, leading to misidentification and misspelling of text items.
Innovation Solution
The device employs image processing techniques such as thresholding and edge detection to enhance text recognition, uses confidence levels and correction algorithms to improve scanning performance, and sends text items to a natural language processing server to determine item types and corresponding actions, with user interfaces for selecting actions based on identified text items.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If image processing techniques are applied to enhance text recognition, then text recognition accuracy is improved, but device processing complexity increases
Solution Approach 1:
The patent applies image processing techniques such as thresholding and edge detection as preliminary steps before text recognition. By pre-processing the image to enhance contrast and detect text boundaries, the system improves recognition accuracy while managing complexity through structured preprocessing pipelines.
Solution Approach 2:
The patent introduces intermediate processing stages including thresholding to create binary images and edge detection to identify text boundaries. These intermediary steps act as mediators between raw image capture and final text recognition, improving accuracy by transforming the image into more recognizable forms.
2Measurement precision
If confidence levels and correction algorithms are used to improve scanning performance, then text identification accuracy is improved, but processing time increases
Solution Approach 1:
The patent implements confidence level assessment and correction algorithms that provide feedback mechanisms. The system evaluates the confidence of each text identification and applies corrections when confidence is low, improving overall accuracy through iterative refinement rather than single-pass processing.
Solution Approach 2:
The patent changes processing parameters dynamically based on confidence levels. When confidence is high, the system accepts results quickly; when confidence is low, it applies correction algorithms and re-evaluation, adjusting the processing intensity based on the specific case to balance accuracy and time.
3Adaptability or versatility
If multiple actions are offered for identified text items, then user versatility is improved, but interface complexity increases
Solution Approach 1:
The patent provides a universal action menu that can be applied to any identified text item regardless of type. The same interface structure handles different text types (locations, contacts, URLs) by offering context-appropriate actions from a standardized set, achieving versatility through a unified multi-functional interface.
Solution Approach 2:
The patent segments the user interface into distinct components: text identification, confidence assessment, action determination, and user selection. By dividing the complex interaction into separate manageable segments, the system provides versatile functionality while maintaining interface clarity through modular design.
Data Source
AI summary
A method and an electronic device are provided for obtaining an image or a video frame, including applying to the image or the video frame, at least one image processing technique, scanning the image or the video frame, to identify a text item, determining an item type for the identified text item, and determining an action, corresponding to the item type.


