Virtual Object Detection in Remote Screens via Image Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional robotic process automation (RPA) systems require human interaction, leading to errors due to user-specific mistakes and idiosyncrasies, and struggle with automation of applications where underlying objects are not exposed, especially in remote environments with potential changes in visual appearance.
Innovation Solution
A method and system that detect and define virtual objects on remote screens by capturing and analyzing images to identify actionable objects, employing advanced image processing and OCR technologies, linking objects to their labels, and removing background noise, allowing for accurate automation even without direct access to underlying objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If conventional RPA systems use software robots to interpret user interface, then automation can be achieved, but errors occur due to user mistakes and idiosyncrasies
Solution Approach 1:
The patent replaces conventional RPA's mechanical interpretation of user interfaces with AI-based computer vision and deep learning systems. Instead of software robots mimicking human actions, the system uses neural networks to detect and understand UI elements directly from screen images, eliminating human-like errors while maintaining automation capability
Solution Approach 2:
The system creates digital copies of UI elements by detecting their visual characteristics from screen images. By copying the structural and visual properties of UI elements into a machine-readable format, the system enables reliable automation without requiring direct access to underlying application objects
2Measurement precision
If RPA systems require access to underlying objects for automation, then precise control is achieved, but remote applications with hidden objects become inaccessible
Solution Approach 1:
The patent introduces computer vision technology as an intermediary between the automation system and remote applications. By capturing screen images and using AI to detect UI elements visually, the system bridges the gap between automation requirements and applications where underlying objects are not exposed, enabling remote automation with maintained precision
Solution Approach 2:
The system transitions from accessing applications through their programmatic interface (one dimension) to accessing them through visual screen capture (another dimension). This dimensional shift allows automation of applications previously inaccessible through traditional methods by operating at the visual presentation layer
3Adaptability or versatility
If screen images are captured for analysis, then remote application automation is enabled, but processing complexity increases
Solution Approach 1:
The patent segments the complex image processing task into distinct stages: screen capture, pre-processing (noise removal, contrast enhancement), blob detection for UI element identification, OCR for text recognition, and hierarchical relationship detection. This segmentation manages complexity by breaking down the overall process into specialized, manageable components
Solution Approach 2:
The system performs preliminary actions on captured screen images including noise removal, contrast enhancement, and binarization before proceeding to object detection. These pre-processing steps simplify the subsequent analysis by improving image quality and reducing complexity for the detection algorithms
Data Source
AI summary
Methods and systems that detect and define virtual objects in remote screens which do not expose objects. This permits simple and reliable automation of existing applications. In certain aspects a method for detecting objects from an application program that are displayed on a computer screen is disclosed. An image displayed on the computer screen is captured. The image is analyzed to identify blobs in the image. The identified blobs are filtered to identify a set of actionable objects within the image. Optical character recognition is performed on the image to detect text fields in the image. Each actionable object is linked to a text field positioned closest to a left or top side of the actionable object. The system automatically detects the virtual objects and links each actionable object such as textboxes, buttons, checkboxes, etc. to the nearest label object.


