UI Element Detection Using Unified Selectors, CV, and OCR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing graphical element detection techniques for robotic process automation (RPA) in user interfaces (UIs) are not optimal for all scenarios, as they are typically applied individually and lack versatility across different UI types.
Innovation Solution
A combined series and delayed parallel execution unified target technique is employed, utilizing multiple graphical element detection methods (selectors, computer vision, and optical character recognition) to enhance detection accuracy and adaptability across various UI types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple graphical element detection techniques (selectors, CV, OCR) are applied individually, then each technique can be simple and fast, but the overall detection reliability and adaptability across different UI types deteriorates
Solution Approach 1:
The patent combines multiple graphical element detection techniques (selectors, computer vision, and optical character recognition) into a unified detection system that executes them in a coordinated sequence. The unified target technique integrates these previously separate methods into a single cohesive approach, allowing the system to leverage the strengths of each technique while maintaining overall reliability across different UI types.
Solution Approach 2:
The unified target technique creates a multi-functional detection system that can adapt to various UI types and scenarios. By incorporating multiple detection techniques within a single universal framework, the system gains the ability to handle diverse graphical elements across different applications and interface types, improving overall adaptability without requiring separate specialized systems for each UI type.
2Measurement precision
If a unified target technique combining multiple detection methods is used, then detection accuracy and adaptability improve, but the execution time and processing complexity increases
Solution Approach 1:
The unified target technique executes detection methods in a predetermined sequence, performing preliminary actions with faster techniques (such as selectors) before proceeding to more comprehensive but time-consuming methods (such as computer vision and OCR). This staged approach allows the system to achieve high detection accuracy while minimizing execution time by only invoking additional techniques when necessary.
Solution Approach 2:
The detection system dynamically adjusts its execution based on the specific UI context and element being detected. Rather than always executing all detection methods, the system adaptively selects and sequences techniques appropriate for each scenario, optimizing the balance between detection accuracy and execution time for different graphical elements and UI types.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Graphical element detection using a combined series and delayed parallel execution unified target technique that potentially uses a plurality of graphical element detection techniques, performs default user interface (UI) element detection technique configuration at the application and/or UI type level, or both, is disclosed. The unified target merges multiple techniques of identifying and automating UI elements into a single cohesive approach. A unified target descriptor chains together multiple types of UI descriptors in series, uses them in parallel, or uses at least one technique first for a period of time and then runs at least one other technique in parallel or alternatively if the first technique does not find a match within the time period.