Check OCR Parsing with Expandable Sliding Window Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Optical character recognition (OCR) processes for images of checks often face inaccuracies due to noise, distortions, and inconsistent text structures, leading to unreliable results, especially when dealing with unconstrained images containing information beyond the check.
Innovation Solution
A method involving multiple OCR processes with different segmentation modes, combined with an expandable and sliding window (ESW) technique for fuzzy text-matching, to identify and verify target strings from user input against extracted strings, ensuring accurate verification of check details.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single segmentation mode is used for OCR processing, then the processing speed is maintained, but the accuracy of extracted strings deteriorates due to inconsistent text structures and noise in check images
Solution Approach 1:
The patent applies segmentation by dividing the OCR processing into multiple segmentation modes (e.g., single column mode, multiple columns mode). Each mode is specialized for different text structures found in check images. The system segments the problem space based on text layout characteristics, allowing accurate extraction regardless of the actual text structure in the input image.
Solution Approach 2:
The patent implements dynamics by making the segmentation mode selection adaptive rather than fixed. The system dynamically chooses the appropriate segmentation mode based on analyzing the actual text structure characteristics of the input check image. This dynamic adaptation allows the OCR system to maintain high accuracy across varying text layouts while managing complexity through intelligent mode selection.
2Measurement precision
If multiple OCR processes with different segmentation modes are applied, then the accuracy of extracted strings is improved, but the processing time increases
Solution Approach 1:
The patent applies preliminary action by performing text structure analysis on the check image before executing the full OCR process. This preliminary step identifies the most appropriate segmentation mode in advance, allowing the system to select and apply only the relevant OCR process. This prevents unnecessary execution of multiple OCR processes, thereby reducing processing time while maintaining accuracy.
Solution Approach 2:
The patent implements feedback by using the results of text structure analysis to guide the selection of segmentation modes for OCR processing. The system analyzes characteristics of the input image, provides feedback on the suitable mode, and then applies the corresponding OCR process. This feedback mechanism ensures accurate string extraction while optimizing processing efficiency by avoiding irrelevant OCR operations.
3Reliability
If fuzzy text-matching with expandable and sliding window is used, then the reliability of verification is improved, but the computational complexity increases
Solution Approach 1:
The patent applies segmentation in the matching algorithm by using an expandable and sliding window approach that divides the verification process into incremental steps. Instead of comparing entire strings at once, the algorithm segments the comparison into smaller window-based units that can be evaluated independently. This segmentation makes the complex fuzzy matching more manageable and efficient while maintaining verification reliability.
Solution Approach 2:
The patent implements partial action by using a sliding window that evaluates portions of the text strings incrementally. The expandable window starts small and grows only as needed to capture the necessary matching information. This partial evaluation approach reduces the overall computational complexity compared to exhaustive string comparison, while still achieving reliable verification through accumulated partial matches.
Data Source
AI summary
A method for image processing is disclosed. The method includes: obtaining an image associated with a check; obtaining target strings associated with a payor of the check and based on a user input; obtaining extracted strings by applying multiple optical character recognition (OCR) processes with different segmentation modes to the image; identifying, using an expandable and sliding window (ESW), matches between the plurality of target strings and the plurality of extracted strings; and selecting a winning match from the plurality of matches.


