Mobile Device Content Recognition via User-Mediated Image Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mobile devices face challenges in processing digital images of documents due to limited processing power, image resolution, and environmental artifacts, leading to poor performance of conventional image processing algorithms, which are computationally expensive and often require network-mediated processing, restricting their utility.
Innovation Solution
A method that uses user-mediated feedback to analyze and classify connected components within a digital image on a mobile device, allowing for efficient content recognition without relying on external processing resources, by displaying the image, receiving user feedback to designate points or regions of interest, and estimating the identity of detected components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional image processing algorithms are used on mobile devices, then processing accuracy may be maintained, but processing time and power consumption become prohibitively high
Solution Approach 1:
The patent segments the image processing task by first detecting document boundaries and extracting the document region from the captured image. This segmentation isolates the relevant content area, allowing subsequent processing to focus only on the document portion rather than the entire image, thereby reducing processing time while maintaining recognition accuracy
Solution Approach 2:
The patent extracts the document region from the captured image by detecting document boundaries and removing surrounding areas. This extraction eliminates unnecessary image data, reducing the computational burden on mobile devices while preserving the essential content needed for accurate recognition
2Measurement precision
If conventional image processing algorithms are used on mobile devices, then processing accuracy may be maintained, but power consumption becomes prohibitively high
Solution Approach 1:
The patent segments the image processing task by first detecting document boundaries and extracting the document region from the captured image. This segmentation isolates the relevant content area, allowing subsequent processing to focus only on the document portion rather than the entire image, thereby reducing processing time while maintaining recognition accuracy
Solution Approach 2:
The patent extracts the document region from the captured image by detecting document boundaries and removing surrounding areas. This extraction eliminates unnecessary image data, reducing the computational burden on mobile devices while preserving the essential content needed for accurate recognition
3Power
If network-mediated processing is used, then processing power limitations are overcome, but network connectivity becomes a requirement and latency increases
Solution Approach 1:
The patent performs preliminary document region extraction and processing on the mobile device before any potential network transmission. By preparing and preprocessing the image data locally, the system reduces the need for network-mediated processing and enables offline operation, thereby eliminating network connectivity requirements while maintaining processing capability
Data Source
AI summary
A method includes: displaying a digital image on a first portion of a display of a mobile device; receiving user feedback via the display of the mobile device; analyzing the user feedback to determine a meaning of the user feedback; based on the determined meaning of the user feedback, analyzing a portion of the digital image corresponding to either the point of interest or the region of interest to detect one or more connected components depicted within the portion of the digital image; classifying each detected connected component depicted within the portion of the digital image; estimating an identity of each detected connected component based on the classification of the detected connected component; and one or more of: displaying the identity of each detected connected component on a second portion of the display of the mobile device; and providing the identity of each detected connected component to a workflow.


