A GUI test script amplification method based on multi-agent decision

CN122817072APending Publication Date: 2026-09-25NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610883821.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明要解决的问题是:当前基于测试脚本的录制回放方法一定程度上仍需人工编写新的脚本,缺乏对现有脚本的利用机制

Benefits of technology

[0006]本发明要解决的问题是:当前基于测试脚本的录制回放方法一定程度上仍需人工编写新的脚本,缺乏对现有脚本的利用机制。我们的发明能够理解现有测试脚本及应用界面所蕴含的测试意图与语义信息,并将其组织为状态图模型;同时利用多智能体协同决策自动化地生成扩增测试路径,并转化为扩增测试脚本,从而提高整个测试脚本集合的覆盖度和有效性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817072A_ABST
    Figure CN122817072A_ABST
Patent Text Reader

Abstract

A GUI test script augmentation method based on multi-agent decision-making includes a script playback module, an application exploration module, a semantic understanding module, and an augmented test script generation module. The script playback module gradually replays the original test script on a real mobile device and collects information, converting the execution trajectory into an initial GUI state graph. The application exploration module continuously explores the application under test based on the initial graph, forming a more complete GUI state graph. The semantic understanding module uses a large language model to extract the test intent from the original test path; at the same time, it semantically annotates the representative states in the GUI state graph. The augmented test script generation module uses the semantic information of the original test path and the predefined path augmentation strategy as the intent constraint, combines the semantic annotations in the GUI state graph that has been constructed, and uses a large language model-driven multi-agent collaborative decision to generate test steps, which are then executed simultaneously on a real mobile device to generate an augmented test script. The augmented test script generation module also includes a long-term memory mechanism, which can learn and provide feedback in multiple rounds of generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of mobile application GUI testing in software testing, and specifically relates to recording and playback technology. Background Technology

[0002] Modern mobile applications are characterized by increasing functional complexity, diverse interfaces, and high-frequency iterations, posing unprecedented challenges to the field of mobile application testing. The Graphical User Interface (GUI), as the core entry point and medium for user interaction in mobile applications, requires thorough testing to ensure its functional correctness and maintain application usability and user experience. Traditional GUI testing typically relies on manual operation and verification of the application's GUI. To improve the efficiency of GUI testing, automated GUI testing methods have been increasingly researched and developed. Among these, the mobile application GUI test script recording and playback method provides an automated solution. It records user actions on the device, generates corresponding test scripts, and can automatically replay and verify these actions on the device using the scripts and a testing framework, bringing convenience to GUI testing.

[0003] While test script-based recording and playback methods improve the efficiency of GUI testing to some extent, they still largely rely on manually writing or recording test scripts. Considering that existing mobile applications typically have many features and frequent iterations, manually developing test scripts is obviously time-consuming and labor-intensive, and may introduce potential errors. Furthermore, a single test script can only cover a limited number of page states and business processes, lacking a flexible extension mechanism. When verifying similar but not identical business processes, testers often need to write new test scripts from scratch, resulting in existing test scripts not being fully utilized.

[0004] In recent years, significant progress has been made in the field of Large Language Models (LLMs), providing new opportunities to alleviate the aforementioned problems. Mainstream LLMs, such as DeepSeek and ChatGPT, have demonstrated a certain level of natural language understanding capabilities. In the field of GUI testing, LLMs can understand information such as the GUI semantics and test logic of applications. Therefore, many researchers have begun to incorporate LLMs into GUI testing work, further improving the automation and intelligence of GUI testing. However, no research has yet focused on using LLMs to extend existing GUI test scripts.

[0005] Therefore, based on the aforementioned work and challenges, this invention proposes a GUI test script augmentation method based on multi-agent decision-making to automatically generate augmented scripts from the existing GUI test script set of the application under test, thereby improving the coverage and defect detection capability of the entire test script set for the application under test. First, this invention replays the existing test scripts to collect their execution information. Then, it automatically explores the application under test, organizing the above information into a GUI state diagram. Next, this invention uses a large language model to perform semantic understanding on each state node in the state diagram and extracts knowledge such as the test intent contained in the existing test scripts. Then, relying on the constructed state diagram and the knowledge extracted from the existing test scripts, combined with a predefined path augmentation strategy, this invention uses multi-agent collaborative decision-making driven by a large language model to generate test steps, which are then executed synchronously on a real mobile device, forming candidate augmented test paths. Finally, this invention deduplicates and filters the candidate test paths to generate the final augmented test script, realizing an automated augmentation process from existing test scripts to augmented test scripts. Summary of the Invention

[0006] The problem this invention aims to solve is that current test script-based recording and playback methods still require manual writing of new scripts to some extent, lacking a mechanism for utilizing existing scripts. Our invention can understand the testing intent and semantic information contained in existing test scripts and application interfaces, and organize them into a state diagram model; at the same time, it uses multi-agent collaborative decision-making to automatically generate augmented test paths and transform them into augmented test scripts, thereby improving the coverage and effectiveness of the entire test script set.

[0007] The technical solution of this invention is: a GUI test script augmentation method based on multi-agent decision-making, characterized by its ability to automatically generate augmented test scripts based on existing test scripts and the application under test, utilizing multi-agent collaborative decision-making, thereby improving the coverage and effectiveness of the entire test set. This generation method comprises the following four modules:

[0008] 1) Script Replay Module: This module replays the original script on the device and collects execution information to form an initial GUI state diagram. It parses the operation steps and test assertions in the original test script and replays them step-by-step on a real mobile device. During each execution step, it parses the script steps using Locator parsing and bounding box IoU matching strategies, records the transition relationships between states, and transforms the execution trajectory into an initial GUI state diagram.

[0009] 2) Application Exploration Module: This module is responsible for exploring the application, forming a more complete GUI state graph, providing more sufficient state space and transition information for subsequent path expansion. Based on the initial state graph formed by the original script replay, this module integrates global state graph information, prioritizing access to events that lead to more unexplored areas, continuously executing unaccessed events on the device, and writing new states (including page screenshots and component trees) and transition relationships into the state graph, thereby exploring the application. Exploration stops when the coverage reaches the configured threshold or after several consecutive rounds of no coverage gain.

[0010] 3) Semantic Understanding Module: This module is responsible for semantic understanding of the raw script information and GUI state diagram extracted earlier, providing semantic anchors and context for subsequent path expansion generation. This module uses a large language model to perform semantic understanding on each raw test path, extracting information such as test intent summary, step description, preconditions, and key steps to form the raw test intent. It also uses the large language model to perform multimodal semantic understanding on screenshots of representative states and component trees in the GUI state diagram, generating descriptions including function type and main business intent, forming a semantically annotated GUI state diagram. Representative states are clustered using similarity based on screenshots and component trees of each state in the GUI state diagram, with one state selected from each cluster as the representative state set.

[0011] 4) Augmentation Test Script Generation Module: This module is responsible for generating augmentation test paths and instantiating them into test scripts, and is the core module of this invention. This module utilizes the semantic information and augmentation strategies of the original test paths as intent constraints. Within a nested structure of an outer multi-round strategy loop (using four types of augmentation strategies: round-robin scheduling path augmentation, data augmentation, deep augmentation, and branch augmentation) and an inner single-step decision loop, it uses a semantically annotated GUI state graph as its structural and semantic basis. Through collaborative decision-making by four clearly defined intelligent agents—context observers, action decision-makers, oracle generators, and termination decision-makers—it executes synchronously on a real mobile device, gradually generating candidate augmentation paths. The outer loop incorporates a long-term memory mechanism to achieve cross-round learning and feedback capabilities. Finally, it performs deduplication and filtering on the generated candidate augmentation paths, instantiating the filtered paths into directly runnable test scripts.

[0012] The features of this invention are:

[0013] 1. A novel automated expansion method for GUI test scripts based on multi-agent decision-making is proposed for the first time;

[0014] 2. In terms of automated expansion of GUI test scripts, a nested path generation mechanism was designed for the first time, consisting of an outer multi-round policy loop and an inner multi-agent collaborative single-step decision loop. By having agents with clearly defined responsibilities advance sequentially and combining cross-round long-term memory constraints, the expansion path is aligned with the original test intent while achieving differential exploration.

[0015] 3. This paper proposes a representative state extraction method based on screenshot and component tree similarity clustering for the first time. Combined with multimodal semantic understanding of large language models, it constructs global semantic annotations for GUI state graphs, which is beneficial for subsequent testing and expansion using multi-agent decision-making. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the implementation of the present invention.

[0017] Figure 2 Flowchart for script playback and application exploration modules.

[0018] Figure 3 This is a flowchart of the semantic understanding module.

[0019] Figure 4 Generate a module flowchart for the augmentation test script. Detailed Implementation

[0020] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.

[0021] This patent expands the GUI test script through multi-agent decision-making driven by a large language model, mainly using LLM multi-agent technology, visual language model technology, GUI state graph technology and similarity clustering technology.

[0022] 1. LLM Multi-Agent

[0023] In this invention, we utilize LLM (Multi-Agent Modeling) technology to collaboratively drive the core decision-making process of augmented path generation. Tasks such as observation, decision-making, assertion generation, and termination determination are assigned to clearly defined sub-agents for collaborative completion. An LLM agent is an intelligent system based on a large language model, capable of understanding natural language instructions, generating coherent text, and performing complex tasks. It extends the capabilities of traditional language models by integrating external tools and data sources, achieving more accurate and efficient task processing. Building upon this, the multi-agent system further decomposes complex tasks into multiple sub-agents with independent responsibilities. These sub-agents exchange information and collaborate through structured interfaces, preserving the semantic understanding and reasoning capabilities of individual agents while enhancing the overall task's controllability and scalability through division of labor and collaboration.

[0024] 2. Visual Language Model (VLM)

[0025] In this invention, we use a visual-language model to perform multimodal semantic understanding on GUI screenshots and component trees of mobile applications, generating structured descriptions of the application interface's functional types, key business elements, and main business intentions. A visual-language model (VLM) is a type of large-scale multimodal model capable of processing both image and text information. By jointly training a visual encoder and a language model, the model gains the ability to understand image content and describe or reason using natural language. VLM can identify objects, layouts, text, and their semantic relationships in images, and generate coherent answers or structured outputs by combining them with textual context. Its applications cover areas such as image question answering, visual reasoning, document understanding, and interface analysis, and it is one of the core directions of current multimodal artificial intelligence research.

[0026] 3. GUI State Diagram

[0027] In this invention, we use GUI state diagrams to formally model the runtime interface behavior of mobile applications. We organize the interface states, interaction events, and state transition relationships collected during script playback and application exploration into a directed graph structure, providing a structured foundation for subsequent semantic understanding and path expansion. GUI state diagrams are an extension of the finite state machine concept in the field of graphical user interfaces. State nodes represent snapshots of the application's interface at a given moment (including control hierarchy and screenshots), and directed edges represent interface state transitions triggered by user interaction events, thus abstracting the application's dynamic interactive behavior into a computable graph model. State diagrams are widely used in mobile application testing, GUI automated exploration, and application behavior analysis, systematically capturing the application's reachable state space and providing a formal basis for test path planning and coverage analysis.

[0028] 4. Similarity Clustering

[0029] In this invention, we employ a multimodal similarity clustering algorithm that integrates component tree structural features and screenshot visual features to group GUI states in the state graph and extract representative states from each cluster, thereby achieving semantic understanding of the application interface at a low cost. This clustering method utilizes two complementary types of information: component tree structural features characterize the control hierarchy and layout skeleton of the interface, measuring structural similarity between states by recursively comparing the structural matching degree of subtrees; screenshot visual features capture the visual appearance of the interface at the pixel and layout levels. By fusing these two features, we can effectively distinguish interface variants with similar structures but different content. Multimodal clustering methods improve the accuracy and generalization ability of clustering by integrating heterogeneous information sources, and have broad application prospects in image retrieval, document classification, and interface understanding.

[0030] The GUI test script augmentation method based on a large language model of the present invention mainly includes:

[0031] Script replay module: Based on the GUI state graph model, it records the state nodes and event transition relationships during the script execution process, and transforms the execution trajectory of the original test script into a structured state graph skeleton.

[0032] Application Exploration Module: Based on the GUI state graph, continuously explore the unvisited states of the application, constantly expand the state nodes and transition relationships, and form a more complete state space.

[0033] Semantic understanding module: It uses multimodal similarity clustering to group the GUI states in the state graph and extract representative states. It calls the visual language model to perform semantic understanding on the screenshots and component trees of the representative states, and uses the large language model to perform semantic understanding on the original test path.

[0034] Augmentation Test Generation Module: Utilizes LLM multi-agent to generate candidate augmentation paths in a nested structure of outer policy loop and inner single-step decision loop, and combines a visual language model to understand the current interface screenshot at each decision step, and finally instantiates the filtered paths into an executable test script.

Claims

1. A method for expanding GUI test scripts based on multi-agent decision-making, characterized in that: Based on the existing test scripts and the mobile application to be tested provided by the user, the method automatically generates augmented test scripts to improve the coverage and defect detection capabilities of the test script set for the application to be tested. The method includes the following four modules in sequence: a script replay module, which replays the original test scripts on a real mobile device to form an initial GUI state diagram; and an application exploration module, which continuously explores the unaccessed states of the application based on the initial state diagram to form a more complete GUI state diagram. The semantic understanding module uses a large language model to perform semantic understanding on the state graph and the original test path respectively; the augmentation test script generation module uses a large language model-driven multi-agent collaborative decision-making to generate candidate augmentation paths, and after deduplication and filtering, instantiates them into GUI test scripts that can be run directly.

2. The GUI test script augmentation method based on a large language model according to claim 1, characterized in that: In the script replay module, the operation steps of the original script are parsed and replayed step by step on the device, recording the state of each interface and the transition relationship. The execution trajectory is transformed into an initial GUI state diagram, and each state node contains a screenshot of the corresponding interface and component tree information.

3. The GUI test script augmentation method based on a large language model according to claim 1, characterized in that: In the application exploration module, events that guide more unexplored areas are prioritized based on the information in the global state graph, and the state nodes and transition relationships are continuously expanded. Exploration stops when the coverage reaches the configured threshold or when there is no coverage gain for several consecutive rounds.

4. The GUI test script augmentation method based on a large language model according to claim 1, characterized in that: In the semantic understanding module, multimodal similarity clustering is used to extract representative states from the state graph, a visual language model is called to generate semantic descriptions of the states, and a large language model is used to extract information such as test intent summaries for each original test path.

5. The GUI test script augmentation method based on a large language model according to claim 1, characterized in that: The amplification test script generation module presents a nested structure with two layers of loops. The inner layer generates candidate amplification paths step by step through multi-agent collaborative decision-making driven by a large language model. The outer loop combines long-term memory mechanism to achieve cross-round learning. Finally, the generated candidate amplification paths are instantiated into executable test scripts after deduplication and screening.