UI Element Recognition Using Agentic AI Without System Hooks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current UI automation technologies face challenges in recognizing applications, screens, and UI elements without system-level or application-level interactions, requiring extensive driver and application-level functionality.
Innovation Solution
Utilizing agentic AI agents that recognize applications, screens, and UI elements through generative AI models, supplemented by optical character recognition (OCR) and listener applications, to infer user and AI agent interactions without relying on system-level or application-level inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If driver level and application level interactions are used to perform UI automation, then automation functionality can be implemented, but system complexity increases and system-level information becomes unavailable
Solution Approach 1:
The patent introduces an intermediary AI agent that acts as a mediator between the user interface and the automation system. This agent captures screenshots and processes them through computer vision and OCR technologies to extract UI element information, thereby eliminating the need for complex driver and application-level interactions while maintaining automation functionality.
Solution Approach 2:
The patent replaces the mechanical system of driver and application-level hooks with an optical-based system using screenshots, computer vision, and OCR. Instead of relying on programmatic access to UI elements through drivers, the system uses image processing and AI to detect and interact with UI elements visually, simplifying the automation architecture.
2Productivity
If system-level or application-level interactions are required for UI automation, then automation can be performed, but the system requires extensive driver and application-level functionality
Solution Approach 1:
The patent introduces an intermediary AI agent that acts as a mediator between the user interface and the automation system. This agent captures screenshots and processes them through computer vision and OCR technologies to extract UI element information, thereby eliminating the need for complex driver and application-level interactions while maintaining automation functionality.
Solution Approach 2:
The patent creates a visual copy of the user interface through screenshots and uses this copy for analysis and interaction. By working with screenshot images rather than direct system-level access, the automation system can identify and interact with UI elements without requiring extensive driver and application-level functionality.
3Measurement precision
If key presses and mouse clicks are captured at kernel hook level, then user interactions can be recognized, but this information is not available in all system configurations
Solution Approach 1:
The patent replaces the mechanical system of kernel hook-based input capture with an optical-based system using screenshots and computer vision. By analyzing visual changes in screenshots and using OCR to detect text and UI elements, the system can recognize user interactions without relying on kernel-level hooks, thereby improving adaptability across different system configurations.
Solution Approach 2:
The patent introduces an intermediary AI agent that acts as a mediator between the user interface and the automation system. This agent captures screenshots and processes them through computer vision and OCR technologies to extract UI element information, thereby eliminating the need for complex driver and application-level interactions while maintaining automation functionality.
Data Source
AI summary
Techniques for recognizing applications, screens, and UI elements and for recognizing user interactions and/or AI agent interactions with the applications, screens, and UI elements using agentic artificial intelligence (AI) are disclosed. An AI model may facilitate such recognition without other system inputs, such as system-level information (e.g., key presses, mouse clicks, locations, operating system operations, etc.) or application-level information (e.g., information from an application programming interface (API) from a software application executing on a computing system).


