Pattern-Based Optical Character Recognition for Bulk Text Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Optical Character Recognition (OCR) technologies lack flexibility and ease of use, requiring users to manually select and copy multiple text items from images one at a time, which is laborious and prone to errors, especially when dealing with large datasets like email addresses.

Innovation Solution

Implementing an intelligent text recognition system with a pattern detection mode that allows users to select a pattern, train a model, and automatically identify and list occurrences of that pattern within an image, enabling bulk copying or selective pasting of relevant text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual selection and copying of text items is used, then text extraction is possible, but user efficiency and productivity deteriorate when dealing with multiple text items

Engineering Contradiction:
Improvetext extraction efficiencyVSAvoidmanual text selection complexity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system performs automatic text extraction by detecting patterns themselves without requiring manual user intervention for each text item. The OCR engine automatically identifies and extracts text matching the detected pattern, making the system serve itself rather than requiring continuous user guidance for each extraction task.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of selecting and copying text with an automated computational system. The pattern detection engine and OCR engine work together to automatically identify, select, and extract relevant text items, substituting the manual mechanical interaction with an intelligent automated process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If traditional OCR is used to recognize all text, then text identification is achieved, but user flexibility and ease of use worsen when needing to select specific patterns

Engineering Contradiction:
Improvepattern-specific text extractionVSAvoidmanual text selection process
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system performs preliminary pattern detection and identification before the actual text extraction process. By first detecting the pattern type (email, phone number, URL, etc.) and then using that information to guide the OCR extraction, the system prepares in advance what needs to be extracted, making the subsequent process more efficient and user-friendly.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the text extraction process into distinct phases: pattern detection, pattern classification, and targeted text extraction. This segmentation allows the system to handle different text patterns (emails, phone numbers, URLs) separately and efficiently, providing adaptability while maintaining ease of operation through automated workflow management.

Inventive Principle:
Principle #1Segmentation

3Productivity

If manual iteration for selecting multiple text portions is performed, then text copying is achieved, but time consumption and productivity worsen

Engineering Contradiction:
Improvebulk text extraction speedVSAvoiditerations for text selection
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system maintains continuous automated operation throughout the text extraction process. Once the pattern is detected and the OCR engine is activated, the extraction continues automatically through all matching text items without interruption or manual intervention, eliminating the need for repeated start-stop manual iterations and significantly reducing time loss.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent replaces the iterative manual process of selecting, copying, and pasting text multiple times with a single automated computational process. The system detects the pattern once, then automatically extracts all matching text items in one continuous operation, substituting multiple manual iterations with a single automated workflow.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250022298A1Intelligent and mode-based optical character recognition
Publication Date: 2025.01.16 OMNISSA LLC
  • US20250022298A1 patent drawing
  • US20250022298A1 patent drawing
  • US20250022298A1 patent drawing

AI summary

Disclosed are various embodiments for intelligent text recognition based upon a selected pattern detection mode. First, text can be identified in an image. A pattern detection mode can be selected by a user or autonomously. In some instances, the pattern detection mode can be selected based at least in part on a user account. Next, the text can be parsed for occurrences of a pattern associated with the selected pattern detection mode. A list of occurrences of the pattern can be generated from the text and presented to a user. In some instances, a user can train a model to learn a new pattern.