OCR Accuracy via Database Scoring for Business Card Text
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The low resolution and poor optics of built-in digital cameras in cellular telephones, combined with unsteady image acquisition and variable lighting conditions, lead to image defects that result in high errors and uncertainty when using OCR to extract textual information from business cards, which are not typically found in electronic dictionaries and do not follow conventional grammar rules.
Innovation Solution
A system that includes a portable imager to capture business card images, an image segmenter to extract text segments, an OCR to generate textual content candidates, a scoring processor to assign scores based on database queries, and a content selector to choose the most likely candidates for updating a contacts database, utilizing pre-processing techniques and database queries to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If built-in digital cameras with low resolution and poor optics are used to capture business card images, then the device complexity and manufacturing cost are reduced, but the image quality deteriorates leading to high OCR error rates
Solution Approach 1:
The system performs preliminary actions by capturing multiple images of the business card at different positions and orientations before OCR processing. This allows the system to pre-process and combine images to compensate for the poor quality of individual low-resolution captures, thereby improving the accuracy of text extraction without requiring a more complex camera system.
Solution Approach 2:
The patent introduces an intermediary image processing system that acts as a mediator between the low-quality camera capture and the OCR process. This intermediary layer includes image enhancement, noise reduction, and text region detection algorithms that bridge the gap between poor image quality and accurate text recognition, allowing accurate OCR without improving the camera hardware.
2Measurement precision
If multiple images are captured to compensate for unsteady hand-held shooting, then the text extraction accuracy is improved, but the time required for information acquisition increases
Solution Approach 1:
The system applies partial action by capturing a small number of additional images (typically 2-5) rather than requiring extensive multiple captures. This partial excessive action provides sufficient redundancy to compensate for hand-held instability while minimizing the time penalty, striking a balance between accuracy improvement and time efficiency.
Solution Approach 2:
The patent implements skipping by rapidly capturing multiple images in quick succession and then selectively processing only the most useful images for OCR. The system rushes through the image capture phase with automated timing and selection, skipping unnecessary manual review steps, thereby reducing the overall time required while maintaining accuracy benefits from multiple captures.
3Productivity
If OCR is applied directly to low-quality business card images, then the processing speed is maintained, but the error rate increases due to image defects
Solution Approach 1:
The system performs preliminary image enhancement and text region detection before applying OCR. This preliminary processing prepares the low-quality images by enhancing contrast, reducing noise, and isolating text regions, thereby improving OCR reliability without significantly impacting processing speed. The preliminary actions are automated and optimized to minimize additional processing time.
Data Source
AI summary
In a system for updating a contacts database (42, 46), a portable imager (12) acquires a digital business card image (10). An image segmenter (16) extracts text image segments from the digital business card image. An optical character recognizer (OCR) (26) generates one or more textual content candidates for each text image segment. A scoring processor (36) scores each textual content candidate based on results of database queries respective to the textual content candidates. A content selector (38) selects a textual content candidate for each text image segment based at least on the assigned scores. An interface (50) is configured to update the contacts list based on the selected textual content candidates.


