Graph Recognition UI Automation for Cross-Platform Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current software testing frameworks for graphical user interfaces (GUIs) are limited in their ability to automate user interface testing across multiple platforms without requiring specific programming skills for each platform, and they often rely on traditional programming APIs, which can be cumbersome and restrictive.

Innovation Solution

A GUI testing device utilizing a pre-trained convolutional neural network (CNN) and graph recognition technology to identify and interact with GUI elements, allowing for platform-agnostic UI automation by generating and navigating through different states of the GUI based on user input operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional programming APIs are used for UI automation, then platform-specific control and precision are improved, but ease of operation and adaptability across platforms deteriorate

Engineering Contradiction:
Improvecontrol precisionVSAvoidplatform adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by creating a platform-agnostic UI automation system that works across multiple operating systems (Windows, macOS, Linux, Android, iOS) using a single graph recognition-based approach. The system captures GUI elements as graph data structures that are platform-independent, allowing the same automation scripts to operate universally without requiring platform-specific programming APIs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent substitutes the mechanical system of traditional API-based automation with a computer vision-based graph recognition system. Instead of using platform-specific programming interfaces to interact with GUI elements, the system uses image processing and graph neural networks to detect, recognize, and interact with UI elements visually, replacing the need for mechanical API calls with a more adaptable visual recognition approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If platform-specific programming skills are required for UI automation, then control precision over GUI elements is improved, but ease of operation and accessibility deteriorate

Engineering Contradiction:
ImproveGUI element controlVSAvoidoperation simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent applies self-service by enabling the UI automation system to automatically capture, process, and interpret GUI elements without requiring user intervention for programming or configuration. The graph recognition model autonomously identifies UI elements, determines their hierarchical relationships, and generates appropriate interaction sequences, making the system self-sufficient and eliminating the need for users to possess programming skills.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary layer between the user and the GUI elements in the form of a graph recognition model. This intermediary automatically translates visual GUI information into structured graph data representations, which then guide the automation actions. Users interact with this simplified intermediary interface rather than directly programming complex GUI interactions, greatly easing operation while maintaining precise control.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If traditional UI testing frameworks are used, then reliability of testing is improved, but device complexity and difficulty of implementation increase

Engineering Contradiction:
Improvetesting reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by breaking down the complex UI automation system into distinct modular components: a graph capture module that extracts GUI elements and their relationships, a graph recognition module that processes the captured data using neural networks, and an interaction execution module that performs automated actions. This segmentation allows each component to be independently optimized and maintained, reducing overall system complexity while preserving testing reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-training the graph recognition neural network model with大量 GUI data from various platforms before deployment. This pre-training establishes a robust foundation for recognizing diverse UI elements and their relationships, allowing the system to reliably handle new applications and interfaces without requiring complex reconfiguration or additional programming, thereby simplifying implementation while maintaining high reliability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11599449B2Framework for UI automation based on graph recognition technology and related methods
Publication Date: 2023.03.07 CITRIX SYSTEMS INC
  • US11599449B2 patent drawing
  • US11599449B2 patent drawing
  • US11599449B2 patent drawing

AI summary

A GUI testing device may be configured to execute a testing state machine for interacting with a software application to generate an initial screen of a GUI. The GUI testing device may be configured to determine a current state in the testing state machine based upon a matching trigger target in the initial screen to a given state. The current state may include an operation, and the operation may associate with a trigger target to operate on. The trigger may include a source state, a destination state, and a trigger target. The operation may include a user input operation, and an operation trigger target. The GUI testing device may be configured to perform the operation on the matching trigger target in the initial screen to generate a next screen of the GUI, and advance from the current state to a next state based upon the trigger.