LLM-driven GUI test method and system with experience accumulation and adaptive learning capabilities
By constructing a three-layer architecture of contextual memory, reflective memory, and strategic memory, the problems of experience accumulation and cross-application migration in existing LLM-driven GUI testing are solved, enabling efficient and in-depth GUI testing and improving testing efficiency and bug detection capabilities.
Patent Information
- Application Number
- CN202510995854.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-11
AI Technical Summary
Existing LLM-driven GUI automation testing methods lack systematic testing experience accumulation and cross-task reuse mechanisms, resulting in frequent invalid interaction loops during testing, difficulty in deeply exploring error-prone functional areas, and inability to transfer testing experience from similar applications to new scenarios, leading to low bug detection efficiency and insufficient coverage.
We construct an LLM-driven GUI testing method with experience accumulation and adaptive learning capabilities. By building contextual memory, reflective memory, and policy memory, we achieve multimodal structured encoding of GUI interaction trajectories and cross-application policy transfer. We also adopt an on-demand memory recall mechanism to guide test decisions in real time.
It significantly improves testing efficiency and coverage, enables in-depth exploration of complex functional areas, reduces repetitive operations, enhances cross-application testing versatility and bug detection capabilities, and reduces testing time and costs.
Smart Images

Figure CN120929368A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated graphical user interface (GUI) testing in software development testing, and specifically to an LLM-driven GUI testing method and system with experience accumulation and adaptive learning capabilities. Background Technology
[0002] As mobile applications become increasingly larger and more complex, automated GUI testing is becoming a crucial step in ensuring software quality. GUI testing techniques based on Large Language Models (LLMs), which interpret GUI screenshots and generate actions, are widely used in industry and academia. However, existing methods typically test each application in isolation, lacking experience accumulation and reuse mechanisms. This makes it difficult to deeply explore potential problem areas during testing and to effectively trigger error-prone functions.
[0003] Traditional automated GUI testing methods rely on predefined exploration scripts or hard-coded strategies, resulting in limited adaptability. While existing LLM-based methods possess multimodal understanding capabilities, they are limited to single-session testing and cannot leverage historical testing experience to optimize strategies. In reality, many applications share common UI patterns and functional modules. Testers need to perform tests across multiple versions, updates, or usage scenarios, and the accumulated experience is crucial for determining which interactions might improve coverage and which functions are more prone to errors. Therefore, there is an urgent need for a dynamic, experience-driven memory mechanism that allows LLM-based GUI testing systems to reflect on past behaviors over time to summarize problem patterns and reuse learned strategies to more effectively guide future testing.
[0004] The shortcomings of existing technologies are as follows:
[0005] 1) Lack of test experience accumulation and dynamic reuse mechanisms: Existing LLM-based GUI testing methods generally adopt a single-session isolated testing mode, lacking a systematic storage of historical experience and a cross-task reuse architecture. During testing, the agent cannot record and reuse efficient interaction paths discovered in previous tests, often resulting in repetitive and ineffective exploration. For example, in multiple tests of financial apps, abnormal input test trajectories for similar transfer functions (such as crashes caused by inputting non-numeric characters) were not structured and stored. Each test required retracing the basic process, resulting in more than 30% of the testing time being wasted on repetitive operations. This "memoryless" testing mode prevents the system from iteratively optimizing its strategies. Especially when facing deep functions of complex applications (such as settings modules with multi-level nested menus), the lack of guidance from successful historical exploration experience usually results in insufficient test depth, covering less than 50% of the surface page, leading to a large number of hidden bugs (such as logical conflicts in deep settings items) not being triggered.
[0006] 2) Weak ability to abstract common knowledge and transfer strategies across applications: Existing solutions lack semantic-level abstraction mechanisms for UI patterns (such as standardized layouts of top navigation bars and bottom operation bars) and functional modules (such as common processes like user registration and shopping cart checkout) across different applications. When testing similar applications, it is impossible to transfer the "message sending function boundary value testing strategy" (such as the interface freezing due to sending excessively long texts) summarized in a social app to a communication application with a similar architecture, resulting in the duplication of testing work for similar issues across different applications. For example, in e-commerce application testing, a parameter conflict bug found in the "promotional coupon stacking" function of an app, due to the lack of a cross-application strategy transfer mechanism, still needs to be explored from scratch when testing another e-commerce app of the same type, resulting in an average increase of more than 40% in the detection time of such common risk points, and a difference of 30%-50% in bug detection coverage across applications.
[0007] 3) Insufficient Multimodal Memory Representation and Dynamic Policy Adjustment Mechanisms: Existing methods lack a unified structured encoding system for multimodal data such as screenshots, view hierarchies, and action logs in GUI testing, failing to support efficient memory retrieval based on semantic similarity. When the test agent encounters changes in interface layout (such as button position adjustments after version updates) or functional logic changes, it often gets stuck in an invalid operation loop because it cannot quickly match the interaction strategies of similar interfaces in the past. For example, after a news app update, the "favorite" button was moved from the bottom of the article to the upper right corner. Traditional methods, which rely on fixed coordinate clicks, continue to click the old position, resulting in more than 15% of the test time being wasted on invalid interactions. In addition, the existing memory mechanism and LLM inference process lack sufficient coordination depth. The memory retrieval results cannot dynamically adjust the weights according to the current inference needs of LLM (such as functional understanding and anomaly prediction). For example, when generating test sequences of complex processes, LLM has difficulty calling the complete interaction trajectory of similar functions in the past, resulting in logical gaps in policy generation and a high omission rate of 25%-30% for key test steps (such as boundary condition verification for abnormal scenarios). Summary of the Invention
[0008] This invention aims to address the core technical problems existing in current LLM-driven GUI automated testing: Current methods lack a systematic accumulation of testing experience and a cross-task reuse mechanism, leading to frequent invalid interaction loops during testing, difficulty in deeply exploring error-prone functional areas, and an inability to transfer testing experience from similar applications to new scenarios, resulting in low bug detection efficiency and insufficient coverage. Specifically, existing technologies cannot effectively capture and structure the storage of functional-level test trajectories, cannot summarize problem patterns through historical behavior analysis, and cannot synthesize universal cross-application testing strategies. This prevents test agents from dynamically calling relevant memories to optimize exploration paths when facing changes in interface layout or functional logic adjustments, leading to a high rate of missed bugs. Furthermore, existing solutions lack unified representation and efficient retrieval capabilities for multimodal test data, and the collaboration between memory mechanisms and the LLM inference process is insufficient, making it difficult to dynamically inject the most relevant historical experience based on real-time GUI states. There is an urgent need to construct a memory mechanism that supports the dynamic accumulation, abstraction, and cross-application migration of testing experience to improve the efficiency, coverage, and bug detection capabilities of automated GUI testing.
[0009] The technical solution adopted in this invention is as follows:
[0010] An LLM-driven GUI testing method with experience accumulation and adaptive learning capabilities includes the following steps:
[0011] Contextual memory is constructed based on GUI interaction patterns;
[0012] By analyzing episodic memories and summarizing behavioral characteristics, reflective memories are constructed.
[0013] By summarizing and reflecting on the testing strategies in memory, we can construct strategy memory;
[0014] Retrieve relevant memories from episodic memory, reflective memory, and strategic memory and inject them into LLM prompts to guide GUI testing decisions in real time.
[0015] Furthermore, the step of constructing interactive-level contextual memory based on GUI interaction trajectories includes:
[0016] The continuous GUI interaction trajectory is segmented into discrete functional units. The segmentation process analyzes the visual-text pairs of the screenshot through a multimodal large language model to identify the index range of the corresponding complete functional unit.
[0017] By constructing a directed interaction graph, the operation flow within each functional unit is captured. Each node in the directed interaction graph corresponds to a GUI state, and the directed edges represent the interaction that transforms the application from one GUI state to another. Each edge is marked with an operation and the corresponding component identifier.
[0018] Each functional unit is multimodal structured encoded to form contextual memory.
[0019] Furthermore, the multimodal structured encoding of each functional unit includes:
[0020] A dual encoder architecture based on the CLIP model is adopted, in which the labeled screenshots and text sequences are encoded by the CLIP visual encoder and text encoder respectively to generate visual embeddings and text embeddings.
[0021] A two-layer graph convolutional network (GCN) is used to process directed interaction graphs and obtain behavioral structure embeddings.
[0022] Visual embeddings, text embeddings, and behavioral structure embeddings are connected and projected onto a unified latent space through a fully connected layer, forming vectors that are stored in contextual memory.
[0023] Furthermore, the construction of reflective memory by analyzing episodic memories and summarizing behavioral characteristics includes:
[0024] Merge semantically similar test tracks based on activity name sequences, action-component pair sequences, and cross-step GUI page structures;
[0025] Problematic operations are identified and recorded in each functional unit and embedded into the structured metadata of reflective memory. Problematic operations include unresponsive operations, redundant operations, and operations that trigger defects.
[0026] Natural language reflections are generated for each functional unit, and these reflections summarize test behaviors and problem operations.
[0027] Furthermore, the construction of strategy memory by summarizing and reflecting on test strategies includes:
[0028] Input the application's name, category, and description, along with a set of natural language reflections, integrate them into a structured prompt, and request MLLM to generate a concise and reusable test strategy. The output consists of three parts: a global test plan, a feature-level strategy, and actions that may trigger defects, thus generating an application-level strategy summary.
[0029] The natural language strategy summaries for each application are grouped, and the summaries within each category are concatenated. Then, MLLM synthesis is used to generate a summary that integrates domain-level testing strategies, indicating the identification of recurring patterns, common defect sources, and best practices applicable to the group of applications.
[0030] Furthermore, the retrieval of relevant memories from episodic memory, reflective memory, and strategic memory employs an on-demand memory selection mechanism, retrieving relevant memories only under specific runtime conditions, and defines the following three types of triggers:
[0031] Cold start: At the beginning of testing a new application, retrieve application-level policies from policy memory;
[0032] Test Stall: During the test, monitor the test trajectory and detect stagnation. When stagnation is detected, use the latest functional unit code to query the reflex memory.
[0033] Function transition: After completing a functional unit, it is expected to transition to a new function, retrieve the contextual memory entries related to the current GUI page, and encode them using the same visual-language encoder as the contextual memory encoding.
[0034] Furthermore, when a strategy memory is retrieved during a cold start, the application-level strategy memory is inserted into the prompt as a high-level exploration guide; during test pauses, reflective memories are injected along with the prompt, providing behavioral references and describing the execution trajectory of similar functional units; during function transitions, functional units matching the starting page and the current page are retrieved from contextual memories, and the corresponding functional unit test trajectories are inserted into the prompt.
[0035] An LLM-driven GUI testing system with experience accumulation and adaptive learning capabilities, comprising:
[0036] The contextual memory building module is used to construct contextual memories based on GUI interaction patterns.
[0037] The reflective memory building module is used to construct reflective memories by analyzing contextual memories and summarizing behavioral characteristics;
[0038] The strategy memory building module is used to construct strategy memory by summarizing and reflecting on the test strategies in memory;
[0039] The runtime memory retrieval module is used to retrieve relevant memories from episodic memory, reflective memory, and strategic memory and inject them into the LLM prompts to guide GUI testing decisions in real time.
[0040] The beneficial effects of this invention are as follows:
[0041] 1) Experience Accumulation and Efficient Reuse. This invention, through its ingenious three-layer memory architecture, completely solves the core pain points of existing LLM-based GUI testing methods in terms of experience accumulation and reuse. Specifically, contextual memory can capture fine-grained functional-level test trajectories, recording in detail every interaction between the agent and the GUI; reflective memory, based on this, summarizes the problem patterns, repetitive or redundant behaviors that occur during the testing process, extracting lessons learned for future reference; and strategy memory, with a high-level perspective, can synthesize and extract universal exploration strategies across applications. These layered and synergistic memories are intelligently retrieved and dynamically injected into the LLM's prompts at runtime, enabling the test agent to efficiently reuse previously successful operation sequences, avoiding the repeated execution of invalid or inefficient actions, thereby significantly improving the efficiency and intelligence level of testing.
[0042] 2) Significantly Enhances Testing Depth and Breadth. The memory mechanism of this invention empowers the GUI test agent to conduct more in-depth and comprehensive explorations, particularly adept at triggering defects that are more likely to cause errors or are hidden in deep functionalities. Through continuous learning and experience accumulation from test history, this invention can intelligently identify and prioritize test paths that historical data indicates are more prone to bugs. This process allows the test agent to go beyond surface-level interactions and strategically delve into the complex functions of the application, thereby significantly improving the ability to discover potential problems, avoiding missing deep defects that could lead to serious consequences, and significantly enhancing the depth of test coverage.
[0043] 3) Superior Cross-Application Testing Capabilities. A key innovation of this invention is the introduction of a strategy memory layer, which enables test agents to synthesize and apply universal exploration strategies across different applications. This means the invention is no longer limited to isolated, zero-based testing of a single application. Instead, it can cleverly transfer and apply efficient exploration patterns, successful interaction paradigms, and insights into potential problems learned from one application to other new applications. This cross-application knowledge transfer capability greatly improves the versatility and efficiency of testing, allowing testing teams to achieve comprehensive and efficient coverage of multiple applications with less effort, thus demonstrating significant practical value in real-world application scenarios.
[0044] 4) High integration and wide applicability. This invention can be designed and implemented as a lightweight, modular plug-in, a feature that allows for extremely convenient integration into various existing LLM-based GUI testing methods. Its flexible design ensures seamless compatibility with different baseline testing methods, greatly reducing the barriers to deployment and application. Attached Figure Description
[0045] Figure 1 This is a flowchart of the steps of the method of the present invention.
[0046] Figure 2 It represents the baseline and the activity coverage under different memory pool sizes using the method of this invention. Detailed Implementation
[0047] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0048] 1. The main contents and innovative points of this invention are as follows:
[0049] 1) Structured capture and encoding of interactive contextual memory. Breaking through the limitations of traditional test logs, this invention innovatively divides continuous GUI interaction trajectories into discrete functional units and performs multimodal structured encoding on each unit, including visual, textual, and structural information, laying the foundation for subsequent memory retrieval and analysis.
[0050] 2) Abstraction and summarization of functional-level reflective memory. Addressing the challenge of effectively utilizing repetitive behaviors and problem patterns during testing, this invention introduces a reflective memory layer. By merging semantically similar test trajectories, detecting problematic behaviors, and generating natural language summaries, it achieves high-level abstraction and behavioral perception insights into testing experience.
[0051] 3) Cross-application generalization and guidance of application-level strategy memory. Addressing the difficulty of generalizing testing knowledge to different applications or versions using existing methods, this invention constructs a strategy memory layer. By aggregating reflections and summaries from single applications and even across applications, it comprehensively generates high-level testing strategies, providing reusable guidance for future testing tasks.
[0052] 4) On-demand retrieval and injection of dynamic memory recall mechanism. This invention designs a dynamic, on-demand memory recall mechanism that can intelligently retrieve relevant memories and inject them into the LLM prompts based on specific conditions such as GUI status, stall signals, or function transitions during the testing process, thereby guiding test decisions in real time and improving test efficiency and relevance.
[0053] The LLM-driven GUI testing method of this invention, characterized by experience accumulation and adaptive learning capabilities, is called MemoDroid. Its implementation revolves around the construction, management, and dynamic invocation of its three-layer memory mechanism, aiming to endow LLM-based GUI testing agents with the ability to learn and accumulate experience over a long period. Its core lies in transforming raw interaction data during the testing process into reusable and reflective structured knowledge, and guiding test decisions at runtime through an intelligent retrieval mechanism. The architecture diagram of this invention is shown below. Figure 1 As shown.
[0054] Step 1: Building Interactive Contextual Memory
[0055] The contextual memory of this invention aims to capture the test trajectory of the GUI agent at the application functional granularity, thereby avoiding the storage of complete session logs or unstructured operation history. The construction process includes the following three key components:
[0056] 1) Process 1: Functional Unit Segmentation
[0057] During automated GUI testing, the test agent generates continuous GUI states and operation sequences. To organize these trajectories into contextual memory, this invention segments them into discrete functional units, each representing an independent interaction sequence. The segmentation process uses a multimodal large language model (MLLM) to analyze the visual-text pairs of screenshots, identifying the index range of the corresponding complete functional unit. Specifically, firstly, the screenshots captured for each interaction step are labeled with components, identifying target components and marking them with red rectangles; simultaneously, a text sequence is constructed to describe each interaction, including the page index, activity name, operation type, and the most informative component identifier. These visual-text pairs are input into the MLLM, which identifies the functional units in the continuous steps.
[0058] 2) Process 2: Interaction Graph Construction
[0059] To structurally capture the operational flow within each functional unit, this invention constructs a directed interaction graph G = (V, E), where each node corresponds to a GUI state (i.e., a screenshot and its associated view hierarchy), and directed edges represent interactions that transition the application from one GUI state to another. Each edge is labeled with an action (e.g., click, scroll, input) and a corresponding component identifier (component text, label, or ID) to preserve transition semantics and component-level behavior.
[0060] 3) Process 3: GUI Information Encoding
[0061] To enable memory retrieval and similarity comparison between functional units, each segmented unit is encoded as a compact multimodal representation. This invention employs a CLIP-based dual-encoder architecture, encoding labeled screenshots and text sequences using a CLIP visual encoder and a text encoder, respectively, to generate visual and text embeddings. Furthermore, a two-layer graph convolutional network (GCN) is used to process the interaction graph, obtaining behavioral structure embeddings. Finally, the visual, text, and structure embeddings are concatenated and projected into a unified latent space through a fully connected layer, forming the final vector stored in context memory for future retrieval, clustering, and similarity matching.
[0062] Step Two: Generation of Functional-Level Reflective Memory
[0063] To achieve experiential abstraction and promote the reuse of higher-level memories, this invention constructs functional-level reflective memories by analyzing episodic memories and summarizing their behavioral characteristics. This generation process includes the following three aspects:
[0064] 1) Process 1: Functional Trajectory Merging
[0065] This invention reduces redundancy by merging functionally similar units (i.e., the functional units obtained through segmentation in step one) based on activity name sequences, operation-component pair sequences, and cross-step GUI page structures. A similarity score is obtained by weighted combination of Levenshtein similarity of activity sequences, Jaccard similarity of operation-component sets, and perceptual hash matching scores between GUI pages. Units with scores exceeding a predefined threshold (e.g., 0.8) are considered equivalent and merged. A JSON object is then constructed to summarize the behavior in the merging trajectory.
[0066] 2) Process 2: Problem Operation Detection
[0067] This invention identifies and records problematic actions within each functional unit. These problematic actions are stored as part of a reflective memory, providing actionable insights for strategy generation and defect identification. Problematic actions primarily include: unresponsive actions (the GUI remains unchanged after interaction), redundant actions (the same action is repeatedly executed on the same component in multiple steps), and actions that trigger defects (causing application crashes, etc.). All detected problematic actions are embedded in the structured metadata of the reflective memory to generate natural language reflections and inform strategy-level planning.
[0068] 3) Process 3: Natural Language Reflective Generation
[0069] This invention generates high-level natural language reflections for each functional unit, summarizing test behaviors and problematic actions. By using the functional unit's image and test behavior summary (including aggregated test behavior data and problematic action detections during the merging process) as input, a cue containing visual and structured information is constructed. This cue is then fed into an MLLM (Multi-Level Modeling) to instruct it to generate concise reflections describing observed problems, behavioral summaries, causes, and recommendations. These reflections provide both human-readable insights and LLM-compatible summaries that can be invoked in future tests.
[0070] Step 3: Summarizing Application-Level Memory Strategies
[0071] To support application-level planning and facilitate experience transfer across applications, this invention constructs a strategy memory layer to store high-level testing strategies derived from reflective memory. These strategies summarize functional test priorities, repetitive interaction risks, and defect-triggering actions common in specific application categories or use cases. This summarization process includes the following two aspects.
[0072] 1) Process 1: Application-level strategy summary generation
[0073] The input includes the application's name, category, and description, as well as a natural language reflection set. This invention integrates these inputs into a structured prompt and requests the MLLM (Multi-Level Modeling) to generate a concise and reusable test strategy. The output consists of three parts: a global test plan (indicating test priorities and a suggested execution order), a function-level strategy (providing specific test guidance for critical operations), and operations that may trigger defects. The generated strategy is stored in text format as part of the application strategy memory, serving as reusable guidance when retesting the same application or testing applications with similar structure or functionality.
[0074] 2) Process 2: Cross-application strategy abstraction
[0075] To further enhance the generality of testing knowledge, this invention abstracts higher-level strategies by aggregating application-level summaries from multiple applications belonging to the same category (such as e-commerce, social media, or tools). This invention groups the natural language strategy summaries for each application and concatenates the summaries within each category. Then, MLLM synthesis is used to generate a summary that integrates domain-level testing strategies, indicating the identification of recurring patterns, common defect sources, and best practices applicable to that group of applications.
[0076] Step 4: Runtime memory call
[0077] Runtime memory recall refers to the process by which MemoDroid dynamically retrieves relevant memory entries during testing. This mechanism enables the system to leverage accumulated experience to guide decisions at contextually appropriate times. MemoDroid triggers memory recall when a new application test begins, when testing stalls, or when a feature transitions. It analyzes the current GUI state, recent interaction history, and overall application metadata to retrieve relevant memories from episodic memories, reflective memories, and policy memories. The retrieved content is organized into cues and injected into the LLM's input to enhance its reasoning and action selection without modifying the underlying testing framework.
[0078] 1) Process 1: Select and memorize as needed
[0079] To enable memory utilization during testing, MemoDroid employs an on-demand memory selection mechanism, retrieving memories only under specific runtime conditions. This selective strategy avoids unnecessary overhead while ensuring guidance is provided when most effective. MemoDroid defines three types of triggers: cold start, test stagnation, and feature transition.
[0080] Cold start: In the initial phase of testing a new application, MemoDroid retrieves application-level policies from its policy memory. This involves retrieving application-based textual information (including application name, category, application description, and a list of activity names), which is encoded into sentence embeddings and compared with the application textual information of existing policy memory entries using cosine similarity.
[0081] Mid-term test stagnation: During testing, MemoDroid monitors the test trajectory and detects stagnation based on the following criteria: (1) more than 60% of recent actions return to previously explored screens; (2) fewer than 3 unique action types occur in the past 10 steps; (3) no new functional units are completed in the past 20 steps. When stagnation is detected, MemoDroid queries the reflective memory using the latest functional unit encoding (including the visual, textual, and graph structure information of that functional unit).
[0082] Function Transition: After completing a functional unit, MemoDroid anticipates transitioning to a new function. It retrieves episodic memory entries related to the current GUI page, including screenshots and text descriptions of the current page, and encodes them using the same visual-language encoder used for episodic memory encoding. These embeddings are matched against the first-page embeddings of previously tested functional units in the memory. The retrieved episodic memories contain natural language summaries describing functions previously explored from similar UI pages, including function descriptions and action sequences.
[0083] 2) Process 2: Memory cue generation
[0084] After retrieving relevant memories, MemoDroid organizes them into structured cues and injects them before the LLM cues as input for LLM-based GUI testing to guide test decisions. When policy memories are retrieved during cold start, application-level policy memories are inserted into the cues as high-level exploration guidance. Memory prefixes include system instructions indicating that they originate from previous testing of similar applications and should be considered when planning initial operations. During test pauses, reflective memories are injected along with the cues, providing behavioral references and describing the execution trajectories of similar functional units. During function transitions, MemoDroid retrieves functional units from contextual memories that match the starting page with the current page. The corresponding functional unit test trajectories are inserted into the cues.
[0085] This invention aims to address key issues in existing automated GUI testing methods, such as low efficiency, poor versatility, and insufficient utilization of historical bug reports. Through its unique three-layer memory mechanism, it brings revolutionary progress to graphical user interface (GUI) testing based on large language models (LLMs). It effectively overcomes the limitations of traditional methods, enabling test agents to learn and evolve from experience, thus demonstrating significant benefits in multiple aspects.
[0086] 1) Experience Accumulation and Efficient Reuse. This invention, through its ingenious three-layer memory architecture, completely solves the core pain points of existing LLM-based GUI testing methods in terms of experience accumulation and reuse. Specifically, contextual memory can capture fine-grained functional-level test trajectories, recording in detail every interaction between the agent and the GUI; reflective memory, based on this, summarizes the problem patterns, repetitive or redundant behaviors that occur during the testing process, extracting lessons learned for future reference; and strategy memory, with a high-level perspective, can synthesize and extract universal exploration strategies across applications. These layered and synergistic memories are intelligently retrieved and dynamically injected into the LLM's prompts at runtime, enabling the test agent to efficiently reuse previously successful operation sequences, avoiding the repeated execution of invalid or inefficient actions, thereby significantly improving the efficiency and intelligence level of testing.
[0087] 2) Significantly Enhances Testing Depth and Breadth. The memory mechanism of this invention empowers the GUI test agent to conduct more in-depth and comprehensive explorations, particularly adept at triggering defects that are more likely to cause errors or are hidden in deep functionalities. Through continuous learning and experience accumulation from test history, this invention can intelligently identify and prioritize test paths that historical data indicates are more prone to bugs. This process allows the test agent to go beyond surface-level interactions and strategically delve into the complex functions of the application, thereby significantly improving the ability to discover potential problems, avoiding missing deep defects that could lead to serious consequences, and significantly enhancing the depth of test coverage.
[0088] 3) Superior Cross-Application Testing Capabilities. A key innovation of this invention is the introduction of a strategy memory layer, which enables the test agent to synthesize and apply universal exploration strategies across different applications. This means MemoDroid is no longer limited to isolated, zero-based testing of a single application. Instead, it can cleverly transfer and apply efficient exploration patterns, successful interaction paradigms, and insights into potential problems learned from one application to other new applications. This cross-application knowledge transfer capability greatly improves the versatility and efficiency of testing, allowing testing teams to achieve comprehensive and efficient coverage of multiple applications with less effort, thus demonstrating significant practical value in real-world application scenarios.
[0089] 4) High Integration and Wide Applicability. This invention is designed and implemented as a lightweight, modular plugin, a feature that allows for extremely convenient integration into various existing LLM-based GUI testing methods. Its flexible design ensures seamless compatibility with different baseline testing methods, significantly lowering the barriers to deployment and application. To verify its effectiveness, MemoDroid underwent extensive and rigorous evaluation on real-world applications from 15 different application categories. For example... Figure 2 As shown, the test results of various baselines after using this invention are encouraging, fully demonstrating that MemoDroid can significantly improve the overall performance of GUI testing, especially in key metrics such as activity coverage and code coverage, proving its strong potential and wide applicability in practical applications.
[0090] Another embodiment of the present invention provides an LLM-driven GUI testing system with experience accumulation and adaptive learning capabilities, comprising:
[0091] The contextual memory building module is used to construct contextual memories based on GUI interaction patterns.
[0092] The reflective memory building module is used to construct reflective memories by analyzing contextual memories and summarizing behavioral characteristics;
[0093] The strategy memory building module is used to construct strategy memory by summarizing and reflecting on the test strategies in memory;
[0094] The runtime memory retrieval module is used to retrieve relevant memories from episodic memory, reflective memory, and strategic memory and inject them into the LLM prompts to guide GUI testing decisions in real time.
[0095] The above division of modules is merely illustrative. In practical applications, the functions described above can be assigned to different functional modules as needed to complete all or part of the functions described in the aforementioned method. The specific working process of each module can be found in the corresponding process in the aforementioned method embodiments, and will not be repeated here.
[0096] Another embodiment of the present invention provides a computer device (computer, server, smartphone, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing steps of the method of the present invention.
[0097] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) that stores a computer program, which, when executed by a computer, implements the steps of the method of the present invention.
[0098] Another embodiment of the present invention provides a computer program product, the computer program product including a computer program, which, when executed by a computer, implements the steps of the method of the present invention.
[0099] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and to implement it accordingly. Those skilled in the art will understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification; the scope of protection of the present invention is defined by the claims.
Claims
1. A testing method for LLM-driven GUIs with experience accumulation and adaptive learning capabilities, characterized in that, Includes the following steps: Contextual memory is constructed based on GUI interaction patterns; By analyzing episodic memories and summarizing behavioral characteristics, reflective memories are constructed. By summarizing and reflecting on the testing strategies in memory, we can construct strategy memory; Retrieve relevant memories from episodic memory, reflective memory, and strategic memory and inject them into LLM prompts to guide GUI testing decisions in real time.
2. The method according to claim 1, characterized in that, The step of constructing interactive-level contextual memory based on GUI interaction trajectories includes: The continuous GUI interaction trajectory is segmented into discrete functional units. The segmentation process analyzes the visual-text pairs of the screenshot through a multimodal large language model to identify the index range of the corresponding complete functional unit. By constructing a directed interaction graph, the operation flow within each functional unit is captured. Each node in the directed interaction graph corresponds to a GUI state, and the directed edges represent the interaction that transforms the application from one GUI state to another. Each edge is marked with an operation and the corresponding component identifier. Each functional unit is multimodal structured encoded to form contextual memory.
3. The method according to claim 2, characterized in that, The multimodal structured coding of each functional unit includes: A dual encoder architecture based on the CLIP model is adopted, in which the labeled screenshots and text sequences are encoded by the CLIP visual encoder and text encoder respectively to generate visual embeddings and text embeddings. A two-layer graph convolutional network (GCN) is used to process directed interaction graphs and obtain behavioral structure embeddings. Visual embeddings, text embeddings, and behavioral structure embeddings are connected and projected onto a unified latent space through a fully connected layer, forming vectors that are stored in contextual memory.
4. The method according to claim 1, characterized in that, The method of constructing reflective memory by analyzing episodic memories and summarizing behavioral characteristics includes: Merge semantically similar test tracks based on activity name sequences, action-component pair sequences, and cross-step GUI page structures; Problematic operations are identified and recorded in each functional unit and embedded into the structured metadata of reflective memory. Problematic operations include unresponsive operations, redundant operations, and operations that trigger defects. Natural language reflections are generated for each functional unit, and these reflections summarize test behaviors and problem operations.
5. The method according to claim 1, characterized in that, The process of summarizing and reflecting on test strategies to construct strategy memory includes: Input the application's name, category, and description, along with a set of natural language reflections, integrate them into a structured prompt, and request MLLM to generate a concise and reusable test strategy. The output consists of three parts: a global test plan, a feature-level strategy, and actions that may trigger defects, thus generating an application-level strategy summary. The natural language strategy summaries for each application are grouped, and the summaries within each category are concatenated. Then, MLLM synthesis is used to generate a summary that integrates domain-level testing strategies, indicating the identification of recurring patterns, common defect sources, and best practices applicable to the group of applications.
6. The method according to claim 1, characterized in that, The retrieval of relevant memories from episodic memory, reflective memory, and strategic memory employs an on-demand memory selection mechanism, retrieving relevant memories only under specific runtime conditions, and defines the following three types of triggers: Cold start: At the beginning of testing a new application, retrieve application-level policies from policy memory; Test Stall: During the test, monitor the test trajectory and detect stagnation. When stagnation is detected, use the latest functional unit code to query the reflex memory. Function transition: After completing a functional unit, it is expected to transition to a new function, retrieve the contextual memory entries related to the current GUI page, and encode them using the same visual-language encoder as the contextual memory encoding.
7. The method according to claim 6, characterized in that, When a strategy memory is retrieved during a cold start, the application-level strategy memory is inserted into the prompt as a high-level exploration guide; during test pauses, reflective memories are injected along with the prompt, providing behavioral references and describing the execution trajectory of similar functional units; during function transitions, functional units that match the starting page and the current page are retrieved from contextual memories, and the corresponding functional unit test trajectories are inserted into the prompt.
8. An LLM-driven GUI testing system with experience accumulation and adaptive learning capabilities, characterized in that, include: The contextual memory building module is used to construct contextual memories based on GUI interaction patterns. The reflective memory building module is used to construct reflective memories by analyzing contextual memories and summarizing behavioral characteristics; The strategy memory building module is used to construct strategy memory by summarizing and reflecting on the test strategies in memory; The runtime memory retrieval module is used to retrieve relevant memories from episodic memory, reflective memory, and strategic memory and inject them into the LLM prompts to guide GUI testing decisions in real time.
9. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 7.