Automating GUI Operations via Screenshot Analysis and File Naming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automation software requires users to have extensive knowledge of programming languages to automate operations in graphical user interfaces, making it difficult for users without programming knowledge to automate tasks, and debugging automated processes is time-consuming due to the need for manual verification of each step.
Innovation Solution
The system allows users to automate operations by taking screenshots of graphical user interface elements, naming them with a simple file naming convention, and analyzing the file names and contents to identify operations and locations, enabling automated performance without requiring programming knowledge and reducing debugging time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If conventional automation software uses specialized programming languages to instruct operations, then automation capability is achieved, but user accessibility deteriorates due to requiring extensive programming knowledge
Solution Approach 1:
The patent uses screenshots (visual copies) of graphical user interface elements instead of programming code to represent and automate operations. Users capture visual images of UI elements, and the system processes these images to automate the corresponding operations, eliminating the need for programming knowledge while maintaining automation capability
Solution Approach 2:
The patent replaces the mechanical system of programming language syntax and code writing with an image-based system. Instead of manually writing and debugging code, users interact with the system through screenshots and visual matching, substituting the programming mechanism with a more accessible image-processing mechanism
2Manufacturing precision
If conventional automation methods use element names to control graphical user interface elements, then automation precision is achieved, but ease of operation deteriorates due to unknown or unavailable element names
Solution Approach 1:
The patent uses visual copies (screenshots) of graphical user interface elements to identify and control them, replacing the need for element names. The system matches captured screenshots against current UI screenshots to locate elements, making automation setup intuitive and accessible without requiring knowledge of element identifiers
Solution Approach 2:
The patent changes the identification parameter from text-based element names to image-based visual matching. By transforming the element identification approach from symbolic (names) to visual (screenshots), the system maintains precision while dramatically improving ease of operation
3Ease of operation
If conventional automation methods use element locations to control graphical user interface elements, then ease of operation is improved, but reliability deteriorates due to location changes in the interface
Solution Approach 1:
The patent uses full screenshots of graphical user interface elements rather than just location coordinates. By capturing and matching the visual content of elements, the system can reliably identify elements even when their positions change in the interface, maintaining both ease of operation and automation reliability
4Measurement precision
If conventional automation methods require manual verification of each step for debugging, then debugging accuracy is achieved, but productivity deteriorates due to time-consuming verification
Solution Approach 1:
The patent uses screenshots to visually represent and verify automated operations. Instead of manually checking each code step, the system captures visual states before and after operations, allowing users to verify correctness through image comparison, thereby maintaining debugging accuracy while significantly improving efficiency
Data Source
AI summary
Methods for automating user operations include analyzing a name of an image file to identify an operation to be automated. One or more embodiments compare a content of the image file to a graphical user interface of a computing device to identify a location of a graphical user interface element in the graphical user interface. One or more embodiments perform the operation identified from the name at the location of the graphical user interface element identified from the content of the image file.


