OCR Screen Region Selection for Fast Application Page Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for identifying pages in graphical user interfaces are CPU and memory intensive, making them unusable on minimum configuration machines, and require different libraries and algorithms for different technologies, leading to slow context identification.
Innovation Solution
Utilizing Optical Character Recognition (OCR) on selected regions of the screen, combined with pre- and post-processing techniques, to identify page information by matching detected words with a word map, and configuring regions, language, and logical conditions for accurate page identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Automation API Interfaces are used to identify pages, then page identification can be achieved, but CPU and memory usage become intensive and identification speed becomes slow
Solution Approach 1:
The patent replaces the mechanical Automation API Interfaces with Optical Character Recognition (OCR) technology to identify pages. Instead of using automation APIs that require memory-intensive operations and multiple technology-specific libraries, the system captures screenshots and uses OCR to extract and compare text content, significantly reducing CPU and memory usage while maintaining identification accuracy
Solution Approach 2:
The patent creates a simplified copy of page identification by capturing visual screenshots and extracting text through OCR, rather than directly interacting with the application through complex automation APIs. This copying approach allows for faster comparison and identification without the overhead of native automation interfaces
2Adaptability or versatility
If different libraries and algorithms are used for different technologies, then each technology can be identified accurately, but device complexity increases and processing becomes slower
Solution Approach 1:
The patent implements a universal OCR-based page identification system that works across different application technologies (SAP, Java, .NET, web applications) without requiring technology-specific libraries or algorithms. The single OCR approach can identify pages in any application that displays text, eliminating the need for multiple specialized automation interfaces
Solution Approach 2:
The patent introduces OCR as an intermediary layer between the screenshot and page identification. Instead of directly using technology-specific automation APIs, the system captures the visual output and uses OCR to extract text, which then serves as a universal identifier across different application platforms and technologies
Data Source
AI summary
A page identification technique is configurable to select regions of a screen for optical character recognition. The size of the regions may be selected to reduce the CPU resources needed to identify a page. Pre-processing and post-processing may be performed to improve the accuracy of the optical character recognition. A word mapping algorithm may include logical condition to determine a page name.


