Script granularity GUI test migration method based on large language model scene understanding

By using a large language model to understand the scene information in the test scripts and screenshots, we automatically generate the test scripts for the target application, solving the test migration problems of cross-platform and cross-vertical applications, and achieving efficient test migration and automated playback.

CN120086126APending Publication Date: 2025-06-03NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411888636.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Current recording and playback technology is difficult to make breakthroughs in complex cross-platform playback, especially in the test migration between the same vertical applications.

Method used

Use the Large Language Model (LLM) to understand the scene information in the test scripts and screenshots, and automatically generate the test scripts for the target application, so as to realize cross-platform and cross-vertical applications. Specifically, it includes the LLM-Agent scene extraction module, the visual enhancement module, the script preprocessing module and the automation exploration module.

Benefits of technology

It realizes test migration across platforms and across vertical applications, reduces the cost of script reuse and improves test efficiency and automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086126A_ABST
    Figure CN120086126A_ABST
Patent Text Reader

Abstract

A script granularity GUI test migration method based on large language model scene understanding comprises an LLM-Agent scene extraction module, a visual enhancement module, a script preprocessing module and an automatic exploration module. The visual enhancement module is used for extracting picture character information and recognizing component coordinates so as to reduce the understanding difficulty of a large model on pictures. The script preprocessing module fully analyzes an abstract syntax structure of a script and extracts basic information required by scene understanding, such as an operation target and an operation category. And the LLM-Agent scene extraction module extracts a functional scene according to the preprocessed script and the screenshot, and automatically explores the to-be-tested application to generate a test script after forming overall understanding of the scene. And the automatic exploration module is responsible for operating specific equipment and supporting automatic operation of the whole framework, and comprises error correction, termination judgment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software testing, and particularly relates to recording and playback technology. After extracting scenarios from test scripts and application screenshots, playback scripts are automatically generated based on the overall understanding of the scenarios. Background Art

[0002] According to the statistics of the China Internet Network Information Center, in March 2024, the active user scale of the mobile Internet in China reached 1.232 billion, a year-on-year increase of more than 20 million, and the total usage duration of the whole network of mobile Internet users reached 203.96 billion hours, a year-on-year increase of 11.6 billion hours. While the number of Internet users is increasing and their stickiness is enhancing, mobile applications have become increasingly complex and diverse. The traditional GUI testing method for mobile applications is to manually write test scripts, which relies on specific test interfaces. When the scale and quantity of test tasks surge, it will consume a considerable amount of time and labor costs.

[0003] The GUI test script recording and playback technology for mobile applications provides an automated solution for this. It can record the operations of users on the device to generate corresponding test scripts and automatically playback these operations on the same device. This greatly reduces the time overhead of test migration. Nowadays, many test frameworks such as Appium have provided the recording and playback function, and testers can perform automated tests after recording a script once, which brings convenience to GUI testing.

[0004] However, the fragmentation problem of mobile applications is very serious. Developers often release versions for multiple platforms such as Android and IOS, and adapt to multiple models under this platform, adapting to different sizes and resolutions. Therefore, we expect the recording and playback tool to have good cross-platform effects. But this is obviously not supported by traditional recording and playback technology.

[0005] In recent years, large language models represented by ChatGPT have emerged continuously. They have powerful semantic understanding and reasoning capabilities, and some large models can even effectively process multi-modal information. Therefore, many researchers have begun to introduce LLM into the development of recording and playback tools in order to make breakthroughs in cross-platform and cross-application aspects.

[0006] Based on the above work, to solve the test migration of similar functional scenarios between applications in the same vertical category and reduce the significant costs brought by script reuse, the present invention, based on the overall granularity of test cases (scripts), utilizes the capabilities of large language models to identify and extract the functional scenarios to be tested presented in the test cases, and migrates the tests of the extracted test functional scenarios to similar functional scenarios in the same vertical category. To enhance the ability of large language models to understand the functional scenarios to be tested, the present invention fully exploits the rich information in the interface images and combines it with the test scripts as the learning corpus for the large model. After forming the scenarios of this type of function, the large language model automatically generates test commands based on the screenshots of the target application, generates playback test scripts, and finally realizes recording and playback across platforms and across applications in the same vertical category. Summary of the Invention

[0007] The problem to be solved by the present invention is that the current recording and playback technology mainly performs mapping based on steps and controls, and it is difficult to make breakthroughs in complex cross-platform playback. Our invention can understand the rich scenario information contained in its scripts and screenshots and automatically play them back to form migration scripts, ultimately solving the problem of recording and playback across platforms and across applications in the same vertical category.

[0008] The technical solution of the present invention is: a script granularity GUI test migration method based on large language model scenario understanding, which is characterized in that it can automatically generate test scripts for the target application according to the provided scripts and screenshots. This generation method includes the following four modules:

[0009] 1) LLM-Agent Scenario Extraction Module: Responsible for extracting scenarios from the processed scripts and application screenshots, perceiving the target application state after forming a mature understanding of the overall scenario, and gradually generating test commands. This module converts the application screenshot into a url that the LLM can directly understand, and organizes it with the processed script into a prompt to provide to the LLM for scenario extraction. After the scenario extraction is completed, the LangGraph sub-module maintains the understanding of the scenario and explores the application to be tested based on this understanding to obtain information such as the coordinates of the controls to be tested and the operation types. The Structured Output sub-module organizes this basic information to form test commands in this GUI state.

[0010] 2) Visual Enhancement Module: Identifies and extracts text information segments existing in the screenshot, and filters the identified text using a preset stop word list. The text extraction part uses optical character recognition (OCR) technology, which supports pictures containing Chinese, English, or mixed text. Regarding the control processing in the screenshot, after preprocessing operations such as zooming, grayscale conversion, and binarization, the Canny technology is used to extract the edges of the picture, and the mutual relationships between the controls are fully analyzed to remove tiny contour lines, and the recognition of containers and nested relationships is processed. Among them, we will omit tiny controls, which is completed by calculating the ratio of the control size to the application screenshot size.

[0011] 3) Script preprocessing module: This module aims to remove the specific framework information in the test script and retain the important test scenario information: type (specific operation or assertion), operation target (a specific control, which is represented as XPath and element_id in the appium framework), target content (mostly appears in assertions), operation type (click event, input event), others (input value, value judged in assertions). This module parses the test script through AST, extracts function calls and method parameters, and determines the "type" based on function calls such as click and assertEqual to split the test script. For each split part, extract the "operation target" and "target content" from the method parameters, and identify the "operation type" through a simple classification model SVM.

[0012] 4) Automated exploration module: To better support automated playback, this mainly responsible for detecting the quality of scenario extraction, automated playback on specific devices, error correction during playback, and automated judgment of playback termination. The playback on devices is relatively simple, while the quality detection, error correction, and termination judgment are more complex. We introduce the idea of multi-agent, use independent evaluation Agents to complete the above scenario judgment and fine-tuning, and deliver the analysis results to the scenario extraction Agent to further adjust its understanding and instructions.

[0013] The features of the present invention are as follows:

[0014] 1. In the field of automated recording and playback, for the first time, a large language model is used to understand the GUI screenshots of mobile applications.

[0015] 2. For the first time, recording and playback are performed based on the granularity of the entire test script, rather than a simple mapping of steps.

[0016] 3. For the first time, test migration of applications in the same vertical category and similar functions is realized, achieving a breakthrough in cross-device and cross-platform test migration. Description of the Drawings

[0017] Figure 1 is the implementation flow chart of the present invention.

[0018] Figure 2 is the detailed flow chart of the LLM-Agent scenario extraction module.

[0019] Figure 3 is the detailed flow chart of the visual enhancement module.

[0020] Figure 4 is the detailed flow chart of the script preprocessing module.

[0021] Figure 5 is the detailed flow chart of the automated exploration module. Detailed implementation manners

[0022] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.

[0023] This patent realizes GUI test migration by understanding the overall scenario of the script through a large language model, mainly using LLM-Agent technology and image understanding technology. The specific key technologies involved are LLM-Agent technology, OCR technology, Canny edge detection, machine learning SVM classification model, AST technology, etc.

[0024] 1. LLM-Agent

[0025] In the present invention, we use LLM-Agent to extract the functional scenarios contained in the script and screenshots, and apply them to the control of the automated process. LLM-Agent is an intelligent system based on a large language model (LLM), which can understand natural language instructions, generate coherent text, and execute complex tasks. It extends the capabilities of traditional language models by integrating external tools and data sources, such as search engines, databases, or computing modules, to achieve more accurate and efficient task processing. Its applications cover fields such as content generation, question answering, code writing, and data analysis, and it has strong autonomous reasoning and context adaptation capabilities. The core advantage of LLM-Agent lies in combining powerful language understanding capabilities with multi-modal interaction, thus supporting cross-domain applications and dynamic task collaboration.

[0026] 2. Recognize text information

[0027] In the present invention, we use OCR technology to recognize the rich text information existing on the controls in the application screenshots. OCR refers to the process in which electronic devices (such as scanners or digital cameras) check printed characters, determine their shapes by detecting dark and bright patterns, and then translate the shapes into computer text using character recognition methods.

[0028] 3. Spatial isolation analysis

[0029] In the present invention, we use Canny technology to segment the application screenshots and extract each individual control in the image for analysis. The Canny edge detection algorithm first applies Gaussian filtering to smooth the image to remove noise; secondly, it finds the intensity gradient of the image; then it applies non-maximum suppression technology to eliminate false edge detections; then it applies a double-threshold method to determine possible boundaries; finally, it uses hysteresis technology to track the boundaries.

[0030] 4. SVM

[0031] In the present invention, we use an SVM model to classify script operations to assist the preprocessing process of the script. SVM (Support Vector Machine) is a supervised learning algorithm applicable to classification and regression tasks. The core idea is to divide the data by constructing an optimal hyperplane to maximize the classification margin. SVM uses kernel functions to process non-linear data, maps the data into a high-dimensional space through transformation, and thus finds a solution for linear separation. Its advantages are that it performs well on high-dimensional data, has strong generalization ability, and can effectively handle small-sample problems. It is widely used in fields such as image recognition and text classification.

[0032] 5. AST

[0033] In the present invention, we use AST to extract the syntax structure of the test script and extract important information required for scene understanding from the abstracted script structure. AST (Abstract Syntax Tree) is a tree-like data structure representing the source code structure, used to describe the syntax relationships in a program. Each node represents a structural element in the program, such as an operator, variable, or statement. By removing details (such as parentheses and whitespace) from the code, AST highlights the syntax structure of the program, facilitating optimization, analysis, and transformation by compilers or analysis tools. It is a key intermediate representation in the compilation process and is widely used in compilation, static analysis, code optimization, and the development of automation tools.

[0034] The script granularity GUI test migration method based on large language model scene understanding of the present invention mainly includes:

[0035] Scene extraction module: Utilize the LLM to understand the overall scene implied by the script and the screenshot, and generate test scripts for the application under test based on the understanding of the scene.

[0036] Visual enhancement module: Use OCR to extract text information from the screenshot, and use Canny to identify the controls in the screenshot and provide the specific coordinates of the controls to the LLM.

[0037] Script preprocessing module: Analyze the abstract structure of the script based on AST, extract the basic information required for scene extraction, and classify the script operations to facilitate LLM processing.

[0038] Automated exploration module: Operate on a specific device based on commands such as ADB, and use the LLM-Agent to judge the degree of scene understanding and adjust the understanding and instructions.

[0039] In this example, for the scene extraction part, we adopt the gpt-4o model of OPENAI and configure related modules such as memory and structured_out based on this model. For the text extraction part of visual enhancement, we adopt publicly available OCR services to identify the Chinese-English mixed text in the screenshot, obtain the specific text, extraction confidence, and the coordinates of the text.

Claims

1. A script-granular GUI test migration method based on large language model scene understanding, characterized in that The application screenshots and scripts provided by the user are automatically replayed to generate a test script. The steps of using the framework are as follows: 1) Users upload test scripts and related GUI screenshots; 2) This method extracts scenes from scripts and screenshots uploaded by users, mainly including: 2.1) Visual enhancement module, which is responsible for extracting text information from images, identifying components and their coordinates, to reduce the difficulty of understanding images for large models. 2.2) The script preprocessing module fully analyzes the abstract syntax structure of the script, discards redundant framework information and extracts the basic information required for scene understanding, such as operation objectives, operation categories, etc. 2.3) The LLM-Agent scenario extraction module extracts functional scenarios based on preprocessed scripts and screenshots. After an overall understanding of the scenario, it automatically explores the application to be tested and finally generates a test script. 2.4) The automated exploration module is responsible for the control operations of specific devices and supports the automated operation of the entire framework, including scene understanding quality detection, error correction, and termination judgment. 3) Present the generated test script to the user.

2. The script-granularity GUI test migration method based on large language model scenario understanding according to claim 1 is characterized in that: In the visual enhancement module, text extraction uses optical character recognition (OCR) technology, which supports images containing Chinese, English or mixed texts; component recognition uses the Canny edge detection algorithm, and performs line simplification, nested recognition and other processing tasks, ultimately accurately identifying component contours and providing LLM with component coordinate information.

3. The script-granularity GUI test migration method based on large language model scenario understanding according to claim 1 is characterized in that: In the script preprocessing module, AST parses the abstract syntax structure of the test script, obtains information such as the operation target from the function call and method parameters, and determines the category of the operation through the SVM model.

4. The script-granularity GUI test migration method based on large language model scenario understanding according to claim 1 is characterized in that: In the LLM-Agent scenario extraction module, LLM-Agent organizes the processed screenshots and scripts into prompts and extracts the functional scenarios. After fully understanding the functional scenarios, LLM-Agent generates reasonable test commands based on the GUI state of the application to be tested, realizes automatic playback, and finally organizes the commands into test scripts.

5. The script-granularity GUI test migration method based on large language model scenario understanding according to claim 1 is characterized in that: In the automated exploration module, the control of the playback device is mainly done by using platform debugging commands such as ADB to complete operations such as screenshots and clicks. This module is equipped with an independent evaluation agent to judge the quality of scene extraction and play a role in error correction during playback and playback termination judgment.