Multi-agent-based scene granularity interface test case generation method

Through the multi-agent collaboration framework and multi-modal large language model, scenario-based GUI test cases are generated, which solves the problem that traditional technology is difficult to understand actual usage scenarios, and achieves comprehensive and reasonable test generation.

CN120086125APending Publication Date: 2025-06-03NANJING UNIV
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202411888527.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional automated GUI test generation technology lacks understanding of actual usage scenarios, and it is difficult to generate complex interactive action sequences covering real business logic or reasonable input text that meets the requirements of usage scenarios, resulting in the inability to complete test generation.

Method used

The multi-agent-based scenario granular interface test case generation method is adopted, and the multi-modal large language model is used to understand the application state and test scenarios, and the multi-agent collaboration framework is used to intelligently generate scenario-based GUI test cases. The method includes five main stages: observation, decision-making, execution, result checking and feedback, and recording, to ensure the business logic support and scenario coverage generated by the test.

Benefits of technology

It realizes GUI test generation with business logic understanding from the user's perspective, which can accurately identify user behavior patterns and generate corresponding test sequences, ensuring the comprehensiveness of mobile application GUI tests and the rationality of test cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086125A_ABST
    Figure CN120086125A_ABST
Patent Text Reader

Abstract

The invention discloses a scene granularity interface test case generation method based on multiple agents, and the method comprises a plurality of large language model agents and related components, namely context memory, an observer, a decision maker, an executor, a supervisor, a recorder and a to-be-tested application. An observer identifies the interface control, detection and attribute identification of the control are effectively focused, and it is ensured that accurate data input is provided for subsequent testing steps. Based on the output of an observer, a decision maker generates a detailed plan of the next operation according to the current page state, and an operation strategy can be independently formulated, so that test generation is supported by business scene logic. And an executor analyzes and executes the generated plan operation, carries out test operation strictly according to the plan, and ensures the traceability and repeatability of the test process. Furthermore, the supervisor verifies whether the operation is executed according to the plan and meets the requirements of the target scene, and makes corresponding feedback according to the check result, thereby ensuring that the wrong operation is corrected in time, avoiding the continuous accumulation of errors, and ensuring the effectiveness of each test step. The recorder is responsible for storing the execution instruction and the context information of the test process to the context memory, and the operation history and result in the test process can be traced through the test context information, so that a better decision can be made in the subsequent iteration, and the result export after the test is completed can be effectively supported. And after the recording is finished, the test generation enters the next iteration until the coverage of the test scene is finished.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software testing, and particularly relates to automated GUI testing technology. Aiming at the scenario granularity, through the coordination and cooperation of multiple intelligent agents based on large language models, test cases that meet the requirements of the target scenario are generated on the basis of understanding the business logic and interface semantics contained in the application under test. Background Art

[0002] With the continuous progress of mobile device and mobile Internet technologies, mobile applications have also developed rapidly. The rapid iteration cycle has continuously promoted the evolution of mobile applications. This phenomenon has not only brought rich and diverse application types and a more abundant user experience, but also posed new challenges and higher requirements for the quality assurance of mobile applications. Among many quality assurance methods, GUI testing plays an indispensable role. Manual testing is completed through manual interaction with the application program, which requires a large amount of human and time resources. Although manual testing has an intuitive advantage in identifying GUI interaction and user experience problems, its low efficiency and high cost often make it difficult to meet the needs of rapid iteration in modern software development.

[0003] In existing research, automated test generation explores the application under test through specific strategies and generates test cases, achieving a higher level of automation, greatly improving efficiency, and reducing the possibility of human errors. Common exploration strategies include random-based, model-based, and learning-based methods. Current automated GUI testing methods are mainly optimized for improving code coverage. However, existing methods have limitations in simulating the complexity and coherence of real user operations, and it is also difficult for them to generate specific inputs to effectively trigger relevant responses of the target application, resulting in the inability of the test to comprehensively simulate various possible user interaction scenarios, thus affecting the coverage. In addition, using code coverage as the main evaluation metric is not sufficient to fully reflect the integrity of the test, because this metric only focuses on code-level execution and ignores functional diversity and user experience.

[0004] Existing methods lack an in-depth understanding of the application's functional scenarios and cannot effectively identify and simulate the actual usage scenarios of the application, thus unable to comprehensively measure the performance and functions of the application under complex and diverse interaction modes. The lack of scenario understanding not only affects the improvement of the coverage metric, but may also lead to test blind spots for certain functions. Based on the above challenges encountered in the existing work and automated GUI testing generation technology, the present invention incorporates a scenario-aware strategy into automated GUI testing to accurately identify user behavior patterns and generate corresponding test sequences, and better understand the semantics contained in GUI controls and layouts from the perspective of human testers, ensuring the comprehensiveness of mobile application GUI testing. Summary of the Invention

[0005] The problem to be solved by the present invention is that traditional automated GUI test generation techniques are difficult to generate complex interaction action sequences covering real business logics or reasonable input texts meeting the requirements of usage scenarios due to the lack of understanding of actual usage scenarios, thus unable to perform complete test generation for the application under test.

[0006] The technical solution of the present invention is: a method for generating scenario-grained interface test cases based on multi-agent, characterized by taking a large language model as the core driver, using a multi-modal large language model to understand the application state and test scenarios, and adopting a multi-agent collaboration framework to intelligently generate scenario-based GUI test cases. This test generation method adopts an iterative generation strategy, and each iteration includes five main stages:

[0007] a) Observation: Identify diverse and complex GUI controls in the interface. In this stage, the system effectively focuses on the detection and property identification of the controls, providing accurate data input for subsequent test steps.

[0008] b) Decision-making: Generate a detailed plan for the next operation according to the current page state. In this stage, the system independently formulates operation strategies, enabling the test generation to be supported by business scenario logics, and also enabling the test to adapt to different application scenarios and control combinations, improving the flexibility and pertinence of the test.

[0009] c) Execution: Parse and execute the generated planned operations. In this stage, the test operations are strictly carried out according to the plan, and the traceability and repeatability of the test process are ensured. In addition, separating the execution process from the decision-making process can reduce the risk of failures caused by insufficient plans during execution.

[0010] d) Result checking and feedback: Verify whether the operations are executed as planned and meet the requirements of the target scenario, and make corresponding feedback according to the checking results, such as continuing to wait for the application response, reverting back to the decision-maker or entering the recorder. In this stage, the system timely corrects incorrect operations, avoids the continuous accumulation of errors, and ensures the effectiveness of each test step.

[0011] e) Recording: Save the execution instructions and the context information of the test process to the context memory. Through the test context information, the operation history and results during the test process can be traced, so as to make better decisions in subsequent iterations, and effectively support the export of results after the test generation is completed. After this stage ends, the test generation enters the next iteration until the coverage of the test scenario is completed.

[0012] The characteristics of the present invention are:

[0013] 1. For the first time, a multi-agent system is used for scenario-grained automated GUI test generation.

[0014] 2. Combine traditional computer vision technology and large language model technology to understand GUI interface information and assist in GUI test generation.

[0015] 3. Aiming at scenario coverage, GUI test generation from the user perspective with business logic understanding is realized. Brief Description of the Drawings

[0016] Figure 1 It is a flowchart of implementation.

[0017] Figure 2 It is a schematic diagram of logical decision-making.

[0018] Figure 3 It is a schematic diagram of control positioning.

[0019] Figure 4 It is a schematic diagram of control prediction.

[0020] Figure 5 It is a schematic diagram of control positioning correction.

[0021] Figure 6 It is a schematic diagram of self-correction.

[0022] Figure 7 It is a schematic diagram of real-time loading verification.

[0023] Figure 8 It is a schematic diagram of state transition verification.

[0024] Figure 9 It is a schematic diagram of test completion verification. Detailed Implementation Manner

[0025] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.

[0026] This patent realizes the generation of automated GUI test cases based on scenario granularity in the following way: a multi-agent collaboration system is constructed through multiple large language model agents and related components, and the application under test is automatically explored and corresponding test cases are generated with scenarios as the goal. Its core modules (roles) specifically include context memory, observer, decision maker, executor, supervisor, and recorder.

[0027] 1. Context Memory: Due to the stateless nature of large language models, to support the intelligent agent in continuously generating scenario-based tests, the context memory module manages the processes that support test generation and the context it generates. Context memory forms the core of the large language model agent's work. The memory content is updated during the processes of control recognition, action execution, and recording, while action decision-making and application state verification rely on the memory content to generate prompt information and interact with the large language model. Drawing on the theory of human memory, context memory is divided into three parts, each storing different types of information generated during the test process:

[0028] a) Long-term Memory Area: Contains all device, application, and scenario information, which are persistent data that can be loaded from the file system or database.

[0029] b) Working Memory Area: Involves the current test generation target scenario, the application and device under test, the test operations that have been executed, and the interactive dialogue with the large language model. This part of the information exists throughout the test task.

[0030] c) Short-term Memory Area: Records the current application state, the state before the most recent operation, and the page control recognition results. This information is updated as the test task progresses, and the application state can be represented by GUI images.

[0031] 2. Observer: The core task of the observer is control recognition. Control recognition mainly targets the GUI images of application pages, and its process is divided into three main stages. First, traditional computer vision algorithms are used to extract graphical controls from the GUI images. Second, the OCR algorithm is used to recognize text controls in the screenshots. Finally, the recognition results of graphical and text controls are integrated and filtered to obtain the final control recognition result. In the stage of integrating and filtering graphical and text controls, the present invention takes into account the merging of text fragments, the exclusion of unreasonable graphical controls, and the integration of related graphics and text. The recognition of text controls uses the OCR recognition algorithm, and the text recognition granularity is mainly based on words and phrases. Therefore, the recognition results often contain many text fragments. The present invention merges these text fragments into complete sentences or paragraphs by calculating distances, thereby improving the accuracy and practicality of the OCR recognition results. For control images, the control recognition results generated by traditional computer vision algorithms are often relatively rough and contain redundant and unreasonable results, which need to be removed. In addition, the recognition results of graphical controls from the status bar at the top of the page (such as battery icons, network signals, etc.), those that are too large (possibly containing multiple actual controls) or too small (noisy data in recognition) also need to be removed. On the GUI interface, text and images are interrelated rather than isolated. A common example is that there is a corresponding text label next to or below a certain graphical control, and the two share the same click response but are recognized as different controls. In response to this situation, the present invention merges the relevant graphical and text controls into a single control entity according to a preset distance threshold. After obtaining the final control recognition result, the present invention will visually display these controls by drawing bounding boxes on the original screenshot according to the position coordinates of each control. This processed screenshot not only shows the boundaries of the controls but also provides the IDs of the controls, forming an intuitive visual representation of the control recognition result, which is recorded in the short-term memory area together with the original GUI image and the set of recognized control objects. This visualization is crucial for the subsequent action decision-making process and is also convenient for multimodal large language model interaction recognition.

[0032] 3. Decision Maker: The core tasks of the decision maker are divided into two parts: logical decision-making and control positioning. In addition, it also involves the implementation of self-correction mechanisms for these two parts.

[0033] a) Logical Decision: The logical decision-making process relies on the excellent capabilities demonstrated by multi-modal large language models in multiple visual tasks and the rich information contained in the page GUI images, which are sufficient to accurately describe the current page state. Therefore, the present invention uses the original GUI image of the page as the basis for analyzing and judging the application state. By combining the information such as the set of operation types, target scenarios, and executed operations to construct prompts, directly utilize the large language model for visual decision-making to determine the most suitable next operation. During this process, the results generated by logical decision-making will include in detail: the name of the operation type, the description of the operation intention, the description of the target control, the text to be input (only for input operations), and the scrolling direction (only for scrolling operations). It should be noted that, in order to optimize the decision-making and execution efficiency, the present invention has considered the click action required to activate the input area when processing input operations. Therefore, it is necessary to avoid generating the corresponding click action separately when using the large language model for decision-making through prompts.

[0034] b) Control Location: In the control location stage, in addition to the direct target control location, it also involves remedies for failed direct location or corrections to incomplete direct location results.

[0035] i. Control Location: After completing the preliminary logical decision-making, the information required to perform the next operation can be obtained, including the operation type, operation intention, control description, and special parameters related to the operation type. However, the pixel-level position of the target control on the page still cannot be accurately located, which limits the ability to directly implement the logical decision. Therefore, further control location is required to supplement the missing control detail information based on the logical decision result of the current page and the recognized control set to ensure the executability of the decision result. During the control location process, the primary task is to establish a direct mapping to match the abstract target control identified in the logical decision-making stage with the specific control recognition result on the page. Given that the direct use of multi-modal large language models for visual parsing has been adopted in the logical decision-making stage, this strategy will continue to be used to ensure the coherence of the entire process. Specifically, we use the visualization result of control recognition as a prompt to guide the large language model to analyze the GUI image with control annotations and bounding boxes, compare it with the original GUI image, and identify and output the accurate ID of the target control to achieve control matching. Based on inheriting the dialogue context with the large language model in the logical decision-making stage, this process aims to construct a logically coherent and natural reasoning chain.

[0036] ii. Control Prediction: During the control recognition process, some controls, such as open text input boxes, are difficult to ensure that the recognition results cover all existing controls due to the lack of clear visual boundaries. This may lead to the failure to find the ID of the target control during the control positioning stage. We can use the relative position information of the recognized controls to construct "virtual controls" to predict the possible positions of the target controls. This strategy needs to be implemented in combination with the control positioning process and the dialogue context of the large language model.

[0037] iii. Control Positioning Correction: Before performing an input operation, the input area must first be activated by clicking. However, the effective click position may deviate from the position of the matched target control. Therefore, it is necessary to correct the position of the target control. Before performing the input action, the click position needs to be corrected according to the relative position to the recognized target control (if no correction is required, the correction result is the target control itself).

[0038] c) Self-Correction: If the most recent decision operation fails to pass the state transition verification, that is, it does not advance the test generation of the test scenario, the self-correction mechanism is activated to ensure the smooth progress of the test generation process. The self-correction process is divided into two main steps: cause analysis and decision adjustment. In the cause analysis stage, potential problems are identified based on the difference between the operation result after the most recent decision and the expected state. To facilitate the analysis of the large language model's reasoning, we have summarized three common problem causes in GUI testing for it to choose from: (1) The logical decision is correct, but there are positioning / prediction / position correction errors in the control positioning, resulting in a mismatch between the actual execution and the expectation. (2) Although the logical decision is reasonable in itself, necessary pre-operations are omitted, resulting in the application being unable to respond normally. (3) The logical decision is itself incorrect at a higher level, not adapting to the current application state or the requirements of the target scenario, resulting in a deviation between the application feedback and the expectation. After completing the cause analysis, decision modification needs to be carried out according to different problem causes. For cause 1, follow the dialogue context in the corresponding control positioning and the large language model, point out the problem to the large language model through prompts, and ask it to provide the modified answer. For causes 2 and 3, follow the dialogue context in the corresponding logical decision and the large language model, point out the problem through prompts, and ask it to adjust and provide the answer again. Through the self-correction mechanism, we ensure that each action decision can accurately push the test generation process in the correct direction.

[0039] 4. Executors and Recorders: The present invention supports four common operations: click, input, scroll, and return. These operations cover the vast majority of requirements during the testing process. The execution process of these operations starts from the parsing of the action decision result, then implements the specific operation by calling the system interface (such as ADB in the Android system), and finally records the execution result in the working memory area of the context memory. Among all the supported operation types, except for the return operation, additional information needs to be extracted during the parsing phase for other operations. The click operation requires specifying the target control; the input operation not only requires the target control and the text to be input, but also needs to locate the input position relative to the target control; the scroll operation requires specifying the target control (if not specified, it defaults to the entire screen) and the scroll direction. It should be specifically noted that the input operation needs to be carried out in two steps: first, locate the input box by clicking to obtain the focus; then perform text input.

[0040] 5. Supervisors: The main responsibility of the supervisor is to verify the application state after each step of the test operation is executed to ensure that the generation of test cases is reasonable and meets the requirements of the target scenario, which is specifically divided into three steps: real-time loading verification, state transition verification, and test completion verification:

[0041] a) Real-time loading verification: The execution of the most recent action decision result may cause the application to enter the loading state, for example, when triggering a network service. In this case, the application is in an unstable state and may change at any time. Therefore, it is necessary to wait for the application to be fully loaded and enter the stable state before making the next decision. Making a decision prematurely may misjudge the effect of the decision execution and wrongly activate the self-correction mechanism. Using a simple method of fixed waiting time can solve this problem, but it will face the problems of fixed duration selection and reduced test generation efficiency. We utilize the visual understanding ability of the multimodal large language model for real-time loading verification, by directly analyzing the GUI image of the current page to determine whether the application is in the loading state. According to the verification result, the present invention dynamically adjusts the response waiting time, thereby optimizing the processing flow.

[0042] b) State transition verification: The goal of state transition verification is to evaluate whether the operations resulting from the most recent action decision have achieved the expected effects, that is, to check whether the application page has undergone the expected valid changes. We introduce a verification method based on a multimodal large language model, which makes full use of the model's visual understanding and analysis capabilities to determine whether the transition from the previous page to the current page triggered by the most recent operation performed meets the expectations in a given target scenario. If the verification fails, the system will trigger a self-correction mechanism. Before that, we also need to additionally confirm whether the most recent operation has caused actual changes to the page, that is, has truly affected the state of the application page. It should be noted that "actual changes" do not include page changes that may occur automatically, such as battery level display updates, carousel transformations, advertisement pop-ups, etc. If actual changes to the page are detected, we perform a corresponding inverse operation to eliminate the impact of the most recent operation, providing a basis for subsequent re-decision-making and execution.

[0043] c) Test completion verification: The goal of test completion verification is to evaluate whether the test process has reached its expected end point. Since this end point cannot be accurately specified before test generation, real-time verification is required during the exploration process of test generation. This step ensures that once the target scenario has been fully tested, further test generation can be terminated in a timely manner. We adopt a method based on a multimodal large language model. By integrating the target scenario recorded in the context memory, the list of executed operations, the screenshots of the current page, and the screenshots of the previous page, we use the model's visual understanding and analysis capabilities to determine whether the test has been completed.

[0044] The complete workflow of the present invention is as follows: The user provides the application to be tested and inputs the description of the target scenario and the relevant private data and requirements coupled with the scenario. The multi-agent automated GUI test system for scenario granularity is started and begins to iteratively generate tests. After the test generation is completed, the generated test cases are integrated into a test script for output. During each iteration, the observer extracts and annotates control information from the screen screenshots; the decision maker makes real-time test decisions based on the information provided by the observer and the target scenario requirements in the context memory, and maps the decision results to an executable test operation; the executor executes the corresponding test operation on the device under test; the supervisor checks the application state after the test operation is executed. When the state does not meet the scenario requirements, it starts the self-correction mechanism, rolls back the workflow to the decision maker, who makes a new decision. When the state meets the scenario requirements, first the recorder records the relevant information, and then it enters the next iteration.

Claims

1. A method for generating scenario-granular interface test cases based on multi-agent, characterized in that: Based on the target test scenario description and the specified application to be tested input by the user, a multi-agent system and its related components are used to automatically generate scenario-level interface test cases for the application to be tested. The following agents are introduced into the system: a) Observer: extracts and annotates control information from screenshots as one of the decision bases; Executor performs the corresponding test operations on the device under test; Supervisor checks the application status after the test operation is executed, and starts the self-correction mechanism when the status does not meet expectations, and rolls back the workflow to the decision maker, who makes a new decision. When the status meets expectations, the recorder records the relevant information before entering the next iteration; b) Decision maker: Makes real-time test decisions based on the information provided by the observer and the target scenario requirements in the context memory, and maps the decision results into an executable test operation; c) Executor: Analyzes the decision results of the decision maker and executes the corresponding test operations on the device under test; d) Supervisor: Checks the application status after the test operation is executed. When the status does not meet the scenario requirements, it starts the self-correction mechanism and rolls back the workflow to the decision maker, who makes a new decision. When the status fully meets the scenario requirements, it actively terminates the test and outputs the test case. e) Recorder: Records relevant information during the test generation process for subsequent test operations and generation of final test cases.

2. A method for generating scenario-granular interface test cases based on multi-agents according to claim 1, characterized in that: The multi-agent automated test generation system constructs five agents: observer, decision maker, executor, supervisor and recorder. The agent memory is managed by a unified context memory module. It uses a multimodal large language model as the agent basis, combines traditional computer vision technology, OCR technology and prompt word engineering technology, and simulates the human thinking process to build an application state understanding and test case generation process based on visual understanding.

3. The method for generating scenario-granular interface test cases based on multi-agent according to claim 1, characterized in that: Conduct interface testing based on scenario granularity, introduce users' language description of the target scenario in the test, define test objectives at scenario granularity, restrict test objectives, generate automated test cases from the user's perspective and business logic, and embed this strategy into the multi-agent system through prompt word engineering, extract semantic understanding of the target scenario from the large language model that supports the operation of the agent, and generate test operations that meet the requirements of the target scenario.

4. The method for generating scenario-granular interface test cases based on multi-agent according to claim 1, characterized in that: The decision-making process of test case generation is centered on observers and decision makers. The observer captures the screenshot of the application to be tested, and automatically identifies and annotates the controls in the screenshot by combining traditional computer vision technology and OCR technology. The decision maker then analyzes the status of the application to be tested based on the recognition and annotation results, and generates the content of the next operation based on the requirements and completion status of the target scenario and maps it into an executable test operation.

5. The method for generating scenario-granular interface test cases based on multi-agent according to claim 1, characterized in that: The inspection and correction process of test case generation is centered on supervisors and decision makers. During the test case generation process, the supervisor verifies the execution results of each test operation, rolls back the operations that deviate from the target scenario, and the decision maker makes a new decision to ensure that the final generated complete test case meets the requirements of the target scenario.

Citation Information

Cited By

  • Self-adaptive error correction method and device for operation error, equipment and storage medium

    CN120725772A

  • Automatic case self-healing method and system for operating system

    CN120849300A

  • Operating System Automation Use Case Self-Healing Methods and Systems

    CN120849300B

  • LLM-based natural language driven GUI test rehearsal and test script generation method

    CN120973696A

  • Graphical user interface test case generation method based on multi-agent framework

    CN121187955A