APP automatic testing method based on multi-modal perception and Agent system
Through multimodal perception technology, the changes in the APP interface are perceived in real time, the UI tree is built, and action sequences are generated through the neural-symbol collaboration mechanism, which solves the problems of insufficient dynamic interface adaptation capabilities, visual semantic understanding faults, high version iteration maintenance costs, and lack of cross-platform generalization capabilities in the existing technology, and efficient and flexible APP automation testing is achieved.
Patent Information
- Application Number
- CN202510563714.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing automated testing technology shows the problems of insufficient dynamic interface adaptation capabilities, visual semantic understanding faults, high version iteration maintenance costs, and lack of cross-platform generalization capabilities in cross-device, cross-version and dynamic scenarios.
The APP automation testing method based on multimodal perception is adopted, and the APP screenshots and dynamic response detection are carried out through the real-time perception layer. The control bounding box detection, visual semantic coding and text extraction are combined with YOLOv5, CLIP and OCR models to carry out control bounding box detection, visual semantic coding and text extraction are constructed to build a UI tree containing hierarchical topological relationships, and the action sequence is generated through the neural-symbol collaboration mechanism through the dynamic inference layer, and the execution optimization layer converts the action sequence into platform-specific operation code.
It realizes efficient adaptation in cross-device, cross-version and dynamic scenarios, reduces maintenance costs, improves generalization capabilities and execution efficiency, and can dynamically update the knowledge base to quickly adapt to different application scenarios.
Smart Images

Figure CN120086109A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image or video recognition or understanding, and particularly to an APP automated testing method and an Agent system based on multi-modal perception. Background Art
[0002] With the complication of the mobile application ecosystem and the acceleration of the multi-terminal integration trend, the automated testing technology is facing increasingly severe challenges. The current mainstream automated testing frameworks (such as Appium, UI Automator, etc.) mainly operate based on predefined control coordinates or hierarchical structure attributes, and their technical limitations are gradually emerging in cross-device, cross-version, and dynamic scenarios, specifically manifested as the following problems: Insufficient dynamic interface adaptation ability: Traditional tools rely on static coordinates or control attributes to locate elements, and cannot effectively handle scenarios such as dynamic interface rendering, resolution adaptive adjustment, or screen rotation, resulting in frequent failures of test cases when migrating devices or updating layouts; Visual semantic understanding gap: Existing solutions lack the ability to model the association between the visual features of interface elements (such as icon styles, color contrasts) and functional semantics, and it is difficult to recognize the intentions in complex interaction scenarios (such as distinguishing the functional differences between the "shopping cart" icon and the "favorite" icon), which limits the intelligent level of test logic; High cost of version iteration and maintenance: Static scripts and fixed test knowledge bases cannot adapt to the interface logic changes in the rapid iteration of APPs (such as control hierarchy reconstruction, adjustment of business process jump rules), and manual repositioning of elements and correction of test logic are required, significantly increasing the maintenance overhead; Lack of cross-platform generalization ability: The differences in UI description systems and interaction protocols between heterogeneous Agents such as Android and iOS force developers to independently design test scripts for different platforms, resulting in low test logic reuse rate and difficulty in multi-terminal consistency verification.
[0003] Although existing research has tried to improve the robustness of control positioning through OCR text matching or image template recognition, it still has not broken through the limitations of single-modal information processing, and lacks the ability of dynamic reasoning about interface context semantics. In addition, most existing knowledge-driven testing methods use predefined rule bases, and it is difficult to achieve real-time environment perception and adaptive decision-making, which restricts the long-term evolution ability of the automated testing Agent system in complex scenarios. Summary of the Invention
[0004] The purpose of the present invention is to provide an APP automated testing method and an Agent system based on multi-modal perception, aiming to overcome the above problems existing in the prior art.
[0005] To achieve the purpose, the present invention provides the following technical solutions: An APP automated testing method based on multimodal perception, comprising the following steps: Step S1: Start and connect the APP to be tested, and at the same time, the causal chain learner is initialized according to the dynamic knowledge base.
[0006] Step S2: The real-time perception layer takes real-time screenshots of the APP's screen interface.
[0007] Step S3: Use the screenshots for dynamic response detection.
[0008] Step S4: Through dynamic response detection, when an interface change is detected, the screenshots are processed in three ways: performing control bounding box detection; performing visual semantic encoding on interface elements; synchronously extracting text content to generate function labels.
[0009] Step S5: Integrate the output results of the three paths, construct a UI tree containing a hierarchical topological relationship, and pass the UI tree to the dynamic inference layer.
[0010] Step S6: When the dynamic inference layer receives a task instruction, the dynamic inference layer parses the task instruction through a neuro-symbolic collaborative mechanism, combines it with the UI tree to generate an action sequence, and then performs symbolic system verification and permission status check on the action sequence.
[0011] Step S7: Update the action sequence to the dynamic document, save it to the real-time knowledge base, and at the same time pass the action sequence to the execution optimization layer.
[0012] Step S8: The execution optimization layer converts the actions of the action sequence into platform-specific operation codes; executes the operation codes, and stores the operation path and operation results in the multimodal memory pool.
[0013] Step S10: Repeat steps S3 - S8 until the current task instruction is completed.
[0014] Furthermore, in the above step S4, the YOLOv5 model is used for control bounding box detection; the CLIP model is used for visual semantic encoding of interface elements; the OCR module is used to synchronously extract text content to generate function labels.
[0015] Furthermore, in the above step S5, the format of the UI tree is XML or JSON.
[0016] Furthermore, in the above step S6, the dynamic inference layer parses the task instruction through a neuro-symbolic collaborative mechanism, combines it with the UI tree to generate an action sequence, and then performs symbolic system verification and permission status check on the action sequence. Specifically: Semantically parse the task instructions issued by the user using a large language model to obtain the instruction intent; the causal chain learner generates an action sequence based on the UI tree and the instruction intent; the symbol system performs symbol system verification and permission status check on the action sequence, repairs existing errors, and ensures the permission to execute the action.
[0017] Further, in the above step S6, the generation of the action sequence is further optimized by combining the temporal memory stored in the multi-modal memory pool.
[0018] Further, in the above step S6, after receiving a task instruction, close the task instruction entry; Between the above step S5 and step S6, it further includes: the dynamic reasoning layer detects whether the current task instruction is completed; if completed, save the detection result to the real-time knowledge base, open the task instruction entry, and receive the next task instruction; if not completed, continue to process the current task instruction, and the task instruction entry remains closed.
[0019] Further, the above step S8 further includes, when the symbol system verification fails or the operation result is incorrect, automatically mark the abnormal path, trigger path rollback, feedback to the causal chain learner, and at the same time update the avoidance strategy to the real-time knowledge base.
[0020] An APP automated testing Agent system (i.e., an intelligent agent) based on multi-modal perception, used to execute any of the above APP automated testing methods; this Agent system is based on a large language model and includes a real-time perception layer, a dynamic reasoning layer, an execution optimization layer, a real-time knowledge base, and a multi-modal memory pool; The above real-time perception layer is used to convert the real-time screenshot of the APP's screen interface into a UI tree containing a hierarchical topological relationship; The above dynamic reasoning layer is used to parse the task instructions through a neural-symbolic collaborative mechanism, generate an action sequence in combination with the UI tree, and then perform symbol system verification and permission status check on the action sequence; The above execution optimization layer is used to convert the action sequence into platform-specific operation codes and execute the operation codes; The above real-time knowledge base is used to store updated dynamic documents, avoidance strategies, and save test results; The above multi-modal memory pool is used to store the operation paths and operation results of executing the operation codes.
[0021] Further, the above real-time perception layer includes a dynamic response detection module, a YOLOv5 model, a CLIP model, an OCR module, and a UI structure generator. The dynamic response detection module is used to analyze the difference degree of interface changes through the SSIM algorithm; when interface changes are detected, the YOLOv5 model is used for control boundary box detection, the CLIP model is used for visual semantic encoding of interface elements, and the OCR module is used to synchronously extract text content and generate function labels; the above UI structure generator is used to fuse the output results of the three models and construct a UI tree containing hierarchical topological relationships. The above dynamic inference layer includes a multi-functional module, a causal chain learner, a symbol system, and a conflict elimination engine. The above multi-functional module is used for LLM task parsing and task status detection; the above causal chain learner is used to generate an action sequence according to the UI tree and task instructions; the above symbol system is used to verify the action sequence, and the above conflict elimination engine is used to automatically mark the abnormal path, trigger path rollback, and feedback to the causal chain learner when it detects that the symbol system verification fails or the operation result is incorrect. In addition, when the causal chain learner detects an abnormal interface, it will also trigger path rollback and generate a rollback action sequence.
[0022] The present invention has the following beneficial effects compared with the prior art: Compared with the prior art, the present invention realizes a real-time perception-dynamic inference-execution closed-loop through a three-level architecture, and has significant improvements in indicators such as generalization ability, maintenance cost, exception coverage, and execution efficiency. The present invention can dynamically update the knowledge base according to the needs of tasks to ensure rapid adaptation in different application scenarios, and automatically optimize its own decision-making process by continuously updating historical information and operation results. Description of the Drawings
[0023] Figure 1 It is a flowchart of the APP automated testing method in the present invention (i.e., the interaction diagram of the Agent system).
[0024] Figure 2 It is a principle block diagram of realizing visual-semantic alignment through the CLIP model in the present invention.
[0025] Figure 3 It is a flowchart of hierarchical verification acceleration of the memory component in the present invention. Detailed Embodiments
[0026] The following describes the detailed embodiments of the present invention with reference to the drawings. To fully understand the present invention, many details are described below, but for those skilled in the art, the present invention can be implemented without these details.
[0027] As Figures 1 - 3 shown, a multi-modal perception-based APP automated testing method includes the following steps: Step S1: Start and connect the APP under test, and at the same time, the causal chain learner is initialized according to the dynamic knowledge base.
[0028] Specifically, this APP automated testing method starts with the establishment of a connection between the APP automated testing Agent system and the APP. This connection can be a wired connection or a wireless connection.
[0029] Step S2: The real-time perception layer takes real-time screenshots of the APP's screen interface.
[0030] Specifically, after the connection is established, the real-time perception layer continuously captures images of the APP screen interface, and the capture frequency can be 15 frames per second (i.e., 15fps).
[0031] Step S3: Use the screenshots for dynamic response detection.
[0032] Specifically, the real-time perception layer includes a dynamic response detection module. The dynamic response detection module analyzes the difference degree of the captured screenshots through the SSIM algorithm to detect the dynamic response of the screen interface during user operations or system changes, so as to be able to timely discover interface changes, or problems such as interface anomalies or non-responses.
[0033] Step S4: When interface changes are detected through dynamic response detection, the screenshots are processed in three ways: perform control bounding box detection; perform visual semantic encoding on interface elements; synchronously extract text content to generate function labels.
[0034] Such as Figure 1 and Figure 2 shown, specifically, the real-time perception layer includes the YOLOv5 model, the CLIP model, and the OCR module. When interface changes are detected, the screenshots are processed in three ways. Specifically: the YOLOv5 model performs control bounding box detection; Figure 2 The principle block diagram for the CLIP model to achieve vision-semantic alignment. The CLIP model performs visual semantic encoding on interface elements (such as mapping a red circular button to the "close" function); the OCR module synchronously extracts text content to generate function labels.
[0035] Step S5: Integrate the output results of the three paths, construct a UI tree containing hierarchical topological relationships, and pass the UI tree to the dynamic inference layer.
[0036] Specifically, the real-time perception layer includes a UI structure generator. The UI structure generator integrates the output results of the three models, namely the YOLOv5 model, the CLIP model, and the OCR module, to construct a UI tree containing hierarchical topological relationships, thereby converting the screenshots of the screen interface into structured data. Among them, the UI tree can be in XML format or JSON format, or other structured languages.
[0037] Step S6: The dynamic inference layer detects whether the current task instruction is completed; if the task instruction is completed, save the detection result to the real-time knowledge base, open the task instruction entry, and receive the next task instruction in the task instruction sequence; if not completed, continue to process the current task instruction.
[0038] The above task instructions include preset test instructions and can also be operation instructions issued by users. In actual production, first use a large number of preset test instructions as task instructions to perform coverage training on the functions of the APP. When the test accuracy reaches a relatively high level, such as 95%, it means that the Agent system meets the conditions for release or going online. After the Agent system goes online, an unknown APP can be selected as the test object for testing, and the user constructs a test instruction set for the APP under test to test the APP under test.
[0039] Step S7: When the dynamic inference layer receives a task instruction, the dynamic inference layer parses the task instruction through a neural-symbolic collaborative mechanism and generates an action sequence in combination with the UI tree.
[0040] Specifically, the signal entry of the dynamic inference layer has a multifunctional module, which has two states: waiting and running, corresponding to two functions: LLM task parsing and task status detection, specifically: When no task instruction is input, the multifunctional module is in the waiting state and enables the LLM task parsing function, which is used to perform LLM task parsing on the task instruction at this time. After the LLM task parsing is completed, the multifunctional module automatically switches from the waiting state to the running state and enables the task status detection function, which is used to detect whether the current task instruction is completed at this time, that is, to execute Step S6; after the current task instruction is completed, the multifunctional module automatically switches from the running state to the waiting state, waiting for the input of the next task instruction, and so on.
[0041] Of course, the above multifunctional module can also be replaced by a task status detection module and an LLM task parsing module.
[0042] Specifically, Step S7 includes the following sub-steps: Step S701: After the dynamic inference layer receives a task instruction, close the task instruction entry. At the same time, the multifunctional module uses the large language model (i.e., LLM) to perform semantic parsing on the task instruction to obtain the instruction intention.
[0043] Step S702: The causal chain learner generates a series of actions that conform to the context causal relationship according to the UI tree and the instruction intention (such as "Modify the avatar → Click on the photo album → Select a photo", or "Click on the search bar → Enter keywords → Filter results"), that is, the action sequence.
[0044] Among them, the causal chain is that the Agent system executes predefined actions (such as clicking, swiping, etc.), records the interface changes after the operations, and establishes a causal chain of "action → element → result" (for example: click on the "search bar" → pop up the keyboard → enter text → display the result list).
[0045] Step S703: The causal chain learner further optimizes the action sequence by combining the temporal memories stored in the multi-modal memory pool.
[0046] Step S704: The symbol system verifies the action sequence, and when the symbol system verification fails, it triggers the conflict resolution engine to repair possible errors, and then performs a permission status check to ensure that there is permission to execute the action.
[0047] Symbol system verification refers to checking whether the action conforms to the UI rules, such as "when entering picture information in the chat window, the UI rule requires that there should be a '+' button on the right side of the input box". The conflict resolution engine refers to the update of scenario-based rules, that is, the functional changes of the same element in different scenarios (such as the "share button" is to share a link on the product page and save content on the article page), and updates the document through historical interaction data and marks the conditional trigger logic.
[0048] Step S8: Update the action sequence to the dynamic document, save it to the real-time knowledge base, and at the same time pass the action sequence to the execution optimization layer.
[0049] Step S9: The execution optimization layer converts the actions in the action sequence into platform-specific operation codes; executes the operation codes, and stores the operation path and operation results in the multi-modal memory pool.
[0050] Specifically, the execution optimization layer includes a cross-device abstraction module, which converts the actions in the action sequence into platform-specific operation codes by the cross-device abstraction module.
[0051] Preferably, step S9 further includes: when the symbol system verification fails or the operation result is incorrect, automatically mark the abnormal path and update the avoidance strategy to the real-time knowledge base.
[0052] Specifically, when the symbol system verification fails or the operation result is incorrect, the conflict resolution engine will trigger a path rollback, feedback to the causal chain learner, let the causal chain learner perform reinforcement learning, and at the same time automatically mark the abnormal path and update the avoidance strategy to the real-time knowledge base. In addition, when the causal chain learner detects an abnormal interface (such as a crash pop-up window or an advertisement page), it will also trigger a path rollback and generate a rolled-back action sequence.
[0053] Step S10: Repeat steps S3 - S9 until the current task instruction is completed.
[0054] Specifically, after the current task instruction is completed, the causal chain learner is initialized according to the dynamic knowledge base, and the multi-functional module starts the LLM task parsing function, allowing the reception of the next task instruction, so that all task instructions in the task instruction sequence can be completed one by one according to steps S1 - S10.
[0055] Simple example: Example 1. E-commerce APP product search test (training phase) 1. Initialization: Connect to the electronic device with the e-commerce APP installed. If the electronic device is an iOS device, use the official Xcode toolchain and the open-source libimobiledevice library; if the electronic device is an Android device, use the Android Debug Bridge (ADB) to establish a connection with the electronic device. Load the industry semantic rule set of the e-commerce APP (for example: ecommerce_rules_v2). For the sake of brevity, the initialization steps are omitted in the description of other embodiments.
[0056] 2. Perception stage: Parse the e-commerce product search page based on YOLOv5 and Tesseract (OCR model). Its Python code is as follows (including three elements: dynamic perception, element association, and rule verification).
[0057] # E-commerce APP product search page parsing system (YOLOv5 + Tesseract solution) import cv2 import pytesseract import torch import numpy as np from yolov5.models.experimental import attempt_load from yolov5.utils.general import non_max_suppression class SearchPageAnalyzer: def __init__(self): # Initialize the multi-modal parsing module self.device = 'cuda' if torch.cuda.is_available() else 'cpu' self.yolo = self._init_yolo_model('yolov5s_ecommerce.pt') self.ocr_config = r'--oem 3 --psm 6 -l chi_sim+eng' # Dynamic perception parameters self.min_ocr_confidence = 0.7 self.element_classes = ['search_box', 'filter_btn', 'product_card', 'price_tag'] # Rule verification configuration self.required_elements = { 'search_page': ['search_box','search_btn'], 'product_card': ['product_image', 'price_tag'] } def _init_yolo_model(self, weights_path): """Load the customized YOLOv5 model""" model = attempt_load(weights_path).to(self.device) model.eval() return model def _dynamic_preprocess(self, img): """Dynamic perception preprocessing""" # CLAHE contrast enhancement lab = cv2.cvtColor(img, cv2.COLOR_BGR2LAB) l, a, b = cv2.split(lab) clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8)) cl = clahe.apply(l) enhanced = cv2.merge((cl,a,b)) # Edge-preserving filtering return cv2.bilateralFilter(enhanced, 9, 75, 75) def _element_relationship_analysis(self, elements): """Element relationship analysis (UI structure parsing)""" # Establish the hierarchy by sorting according to the Y coordinate sorted_elements = sorted(elements, key=lambda x: x['position'][1]) hierarchy = {} for idx, elem in enumerate(sorted_elements): x1, y1, x2, y2 = elem['position'] center = ((x1+x2) / 2, (y1+y2) / 2) # Find the parent container for ancestor in reversed(sorted_elements[:idx]): a_x1, a_y1, a_x2, a_y2 = ancestor['position'] if a_x1<center[0]<a_x2 and a_y1<center[1]<a_y2: hierarchy[elem['id']] = ancestor['id'] break return hierarchy def _rule_validation(self, elements, page_type): """Symbolic rule validation""" missing = [] for req in self.required_elements.get(page_type, []): if not any(e['type'] == req for e in elements): missing.append(req) if missing: raise ValueError(f"Violation of {page_type} rule: missing required element {missing}") def parse_search_page(self, screenshot): """Complete parsing process""" # Dynamic perception stage processed_img = self._dynamic_preprocess(screenshot) img_tensor = torch.from_numpy(processed_img).permute(2,0,1).float().div(255).unsqueeze(0) # YOLOv5 element detection with torch.no_grad(): pred = self.yolo(img_tensor.to(self.device))[0] detections = non_max_suppression(pred, 0.5, 0.5)[0] elements = [] for det in detections.cpu().numpy(): x1, y1, x2, y2, conf, cls = det elem_type = self.element_classes[int(cls)] # Element ROI extraction roi = screenshot[int(y1):int(y2), int(x1):int(x2)] # Tesseract OCR recognition gray = cv2.cvtColor(roi, cv2.COLOR_BGR2GRAY) text = pytesseract.image_to_string(gray, config=self.ocr_config).strip() # Build dynamic element document elements.append({ "id": f"elem_{len(elements)}", "type": elem_type, "position": (x1, y1, x2, y2), "text": text, "confidence": float(conf) }) # Element association analysis hierarchy = self._element_relationship_analysis(elements) # Rule validation (compliance with symbolic constraints) self._rule_validation(elements,'search_page') return { "version": "ecommerce_1.0", "page_type": "search_page", "elements": elements, "hierarchy": hierarchy, "validation": { "status": "passed", "timestamp": "2024-03-20T14:30:00Z" } }
[0058] 3. Inference and decision-making stage: Generate a set of action sequences: click on the search bar → enter "smartphone" → select price sorting → verify the results; Symbolic system verification: confirm that the search button exists and is not grayed out.
[0059] 4. Execution stage: Convert the instruction to an ADB command: adb shell input tap 520 1800; Update the detection result list. If the price sorting label does not appear, trigger the exception recovery strategy.
[0060] Example 2: Abnormal handling on the payment page (training phase) Abnormal detection: Detect an unknown pop-up window (similarity to the known pop-up window < 0.7) through the SSIM algorithm; Judge payment failure: The causal chain learner matches the scenario of the "payment failure pop-up window" in the historical records and detects an unexpected interface (i.e., an abnormal interface); Automatic recovery (i.e., path rollback): Execute the preset path of clicking the retry button → checking the network status → returning to the product page.
[0061] Example 3: Online detection of the Agent system The Agent system maintains test knowledge through dynamic documents, optimizes the operation path through causal chain learning, and ensures operation reliability through symbolic verification, achieving intelligent test coverage of the entire process from requirements to online. When the test accuracy rate of the Agent system reaches a relatively high level, such as 95%, it means that the Agent system meets the release or online conditions. After the Agent system is deployed and goes online, an unknown APP can be selected as the test object for testing. The user constructs a test instruction set for the APP under test and tests the APP under test.
[0062] E-commerce APP test process code: def test_checkout_flow(): # Initialize the multi-modal perception Agent Agent = TestAgent("amazon.apk") Agent.start_recording() # Execute the test try: Agent.execute("Search for product iPhone15") Agent.execute("Add to cart") if Agent.detect_element("Insufficient stock prompt"): raise TestError("Stock exception") Agent.execute("Enter the settlement page") # Symbolically verify the payment conditions SymbolValidator.check( "total_amount > 0", "shipping_address.exists()" ) finally: report = Agent.generate_report() upload_to_jira(report).
[0063] The present invention also discloses an APP automated testing Agent system based on multi-modal perception for executing the above APP automated testing method.
[0064] The work of the Agent system is divided into two stages: training and deployment online. In the training stage, for some specific APPs, the Agent system first learns and identifies the visual and semantic perception of page elements in a multi-modal perception manner by executing a preset test instruction set, generates a UI tree, and stores the UI tree in a multi-modal memory pool and a real-time knowledge base. Combining the current task requirements and the information of the UI tree, it generates an action sequence, then converts the action sequence into an operation code, and executes the operation code, etc. In the deployment stage, the Agent system selects an unknown APP as the test object for testing, and the user constructs a test instruction set for the APP under test to test the APP under test. The Agent system optimizes its decision-making process by continuously updating historical information and operation results.
[0065] The Agent system is based on a large language model (LLM) and includes a real-time perception layer, a dynamic reasoning layer, an execution optimization layer, a real-time knowledge base, a dynamic knowledge base, and a multi-modal memory pool.
[0066] Among them, the real-time perception layer is used to convert the real-time screenshot of the APP's screen interface into a UI tree containing a hierarchical topological relationship.
[0067] Specifically, the real-time perception layer includes a dynamic response detection module, a YOLOv5 model, a CLIP model, an OCR module, and a UI structure generator. The dynamic response detection module is used to analyze the difference degree of interface changes through the SSIM algorithm; when an interface change is found, the YOLOv5 model is used for control boundary box detection, the CLIP model is used for visual semantic encoding of interface elements, and the OCR module is used to synchronously extract text content and generate function labels; the UI structure generator is used to fuse the output results of the three models to construct a UI tree containing a hierarchical topological relationship.
[0068] The dynamic reasoning layer is used to parse task instructions through a neural-symbolic collaboration mechanism, combine with the UI tree to generate an action sequence, and then perform symbolic system verification and permission status check on the action sequence.
[0069] Specifically, the dynamic reasoning layer includes a multi-functional module, a causal chain learner, a symbol system, and a conflict resolution engine. The multi-functional module is used for LLM task parsing and task status detection; the causal chain learner is used to generate an action sequence based on the UI tree and task instructions; the symbol system is used to verify the action sequence, and the conflict resolution engine is used to automatically mark the abnormal path, trigger path rollback, and feedback to the causal chain learner when it detects that the symbol system verification fails or the operation result is incorrect.
[0070] The execution optimization layer is used to convert the action sequence into platform-specific operation codes and execute the operation codes.
[0071] It can be seen that the Agent system realizes a real-time perception-dynamic reasoning-execution closed-loop through a three-level architecture, and has significant improvements in 1) generalization ability (reuse rate of cross-APP test scripts); 2) maintenance cost (adaptation time of test cases after interface changes); 3) exception coverage (automatically identify and handle 23 common exception scenarios (such as permission pop-ups, network timeouts)); 4) execution efficiency, etc.
[0072] The real-time knowledge base is used to store action sequences and existing knowledge, etc., and belongs to short-term memory. It maintains a dynamic document for each APP, recording the function descriptions of current page elements (such as "return button → go back to the previous page") and associated operations (such as "the form needs to be filled out before submission" or "the payment process includes the state of a verification code pop-up window").
[0073] The dynamic knowledge base belongs to long-term and general memory. It accumulates general rules (such as "all setting entrances include a gear icon") based on test task sets (instructions, processes, results) and causal chain learning. After the task instructions are executed, the real-time knowledge base will be integrated into the dynamic knowledge base.
[0074] The multi-modal memory pool is used to store the operation paths and operation results of executing operation codes, etc., and belongs to short-term hybrid memory. It integrates visual perception, operation history, semantic parsing records, and physical interaction features into temporal memory, establishes a cross-dimensional feature index, and assists the Agent system in predicting the next action (such as associating the button color with the historical click success rate, or in an e-commerce APP, completing the coherent operations of search → browse products → add to cart → settlement).
[0075] It can be seen that the real-time knowledge base, dynamic knowledge base, and multi-modal memory pool are memory components. Through the long-term pattern precipitation of the dynamic knowledge base, the context awareness of the real-time knowledge base, and the cross-dimensional association of the multi-modal memory pool, the Agent system can achieve the following features: 1) Pre-generation of operation paths to reduce the inference requests of the large language model (i.e., LLM). 2) Abnormal quick positioning to shorten the problem diagnosis time through the memory backtracking mechanism. 3) Cross-scenario knowledge reuse, where existing knowledge can be reused when new devices are adapted. This memory component effectively solves the problems of "repeated exploration" and "context break" existing in traditional AI Agents, enabling the Agent system to maintain high-efficiency decision-making capabilities in continuous learning.
[0076] Figure 3 For the hierarchical verification acceleration process of the memory component, its hierarchical design is the core path to reduce the inference process. The key steps are represented by code, as shown in the following example: Pattern pre-matching (Python code) class MemoryRetriever: def match_operation_pattern(self, current_state): # Multi-modal memory pool gives priority to matching visual_hash = self._generate_visual_hash(current_state.screenshot) matched = self.memory_pool.search(visual_hash, threshold=0.9) if matched: # Directly call the historical successful path return matched['action_chain'] else: # Fuzzy matching of the dynamic knowledge base return self.dynamic_kb.fuzzy_match(current_state.elements)。
[0077] More specifically, taking the handling of APP pop-ups as an example, the real-time data structures of each memory component: 1. When a file is uploaded and a "storage space insufficient" pop-up window is detected, retrieve the historical operation records of the multi-modal memory pool: { "visual_hash": "a83d2e", "action_chain": {"action": "click", "target": "Clean immediately"}, {"expect": "progress_bar", "timeout": 30} , "success_rate": 92.7% }。
[0078] 2. Matching and pre-storing rules for the dynamic knowledge base: rule = { "condition": "dialog.text contains 'Storage space'", "recommend_action": "click('Clean immediately')", "exception_cases": ["Backup mode", "Agent is being updated"] } 3. The real-time knowledge base records the current element and confirms the executable action: { "element_type": "dialog", "buttons": ["Clean immediately", "Do not process for now"], "context": "file_upload_page" }。
[0079] The above are only the specific implementation manners of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection scope of the present invention.
Claims
1. An APP automated testing method based on multimodal perception, characterized in that: The following steps are involved: Step S1: Start and connect the APP under test, and the causal chain learner is initialized according to the dynamic knowledge base; Step S2: The real-time perception layer takes a real-time screenshot of the APP screen interface; Step S3: Using the screenshot to perform dynamic response detection; Step S4: When the interface changes through dynamic response detection, the screenshot is processed in three ways: control bounding box detection; visual semantic encoding of interface elements; synchronous extraction of text content and generation of function labels; Step S5: Fuse the three output results, construct a UI tree containing hierarchical topological relationships, and pass the UI tree to the dynamic reasoning layer; Step S6: When the dynamic reasoning layer receives the task instruction, it parses the task instruction through the neural-symbolic collaborative mechanism, combines the UI tree, generates an action sequence, and then performs symbol system verification and permission status check on the action sequence; Step S7: Update the action sequence to the dynamic document and save it to the real-time knowledge base, and pass the action sequence to the execution optimization layer; Step S8: The execution optimization layer converts the actions of the action sequence into platform-specific operation codes; executes the operation codes, and stores the operation path and operation results in the multimodal memory pool; Step S9: Repeat steps S3-S8 until the current task instruction is completed.
2. According to claim 1, an APP automated testing method based on multimodal perception is characterized in that: In step S4, the YOLOv5 model performs control bounding box detection; the CLIP model performs visual semantic encoding on the interface elements; The OCR module synchronously extracts text content and generates function labels.
3. According to claim 1, an APP automated testing method based on multimodal perception is characterized in that: In the step S5, the format of the UI tree is XML or JSON.
4. According to claim 1, an APP automated testing method based on multimodal perception is characterized in that: In step S6, the dynamic reasoning layer parses the task instructions through the neural-symbolic collaborative mechanism, and generates an action sequence in combination with the UI tree, and then performs a symbol system check and a permission status check on the action sequence, specifically: Use a large language model to semantically parse the task instructions issued by the user to obtain the instruction intent; The causal chain learner generates an action sequence based on the UI tree and instruction intent; the symbol system performs symbol system verification and permission status check on the action sequence, fixes existing errors, and ensures that there is permission to perform the action.
5. The APP automated testing method based on multimodal perception according to claim 1 or 4 is characterized in that: In step S6, the generation of the action sequence is further optimized by combining the temporal memory stored in the multimodal memory pool.
6. The APP automated testing method based on multimodal perception according to claim 1 is characterized in that: In the step S6, after receiving a task instruction, the task instruction entry is closed; Between step S5 and step S6, the following further includes: the dynamic reasoning layer detects whether the current task instruction is completed; if completed, saves the detection result to the real-time knowledge base, opens the task instruction entry, and receives the next task instruction; If not completed, continue processing the current task instruction.
7. The APP automated testing method based on multimodal perception according to claim 1 is characterized in that: The step S8 also includes, when the symbol system verification fails or the operation result is wrong, automatically marking the abnormal path, triggering the path rollback, feeding back to the causal chain learner, and updating the avoidance strategy to the real-time knowledge base.
8. An APP automated testing Agent system based on multimodal perception, characterized by: Used to execute the APP automated testing method as described in any one of claims 1 to 7; the Agent system is based on a large language model and includes a real-time perception layer, a dynamic reasoning layer, an execution optimization layer, a real-time knowledge base and a multimodal memory pool; The real-time perception layer is used to convert the real-time screenshot of the APP screen interface into a UI tree containing a hierarchical topological relationship; The dynamic reasoning layer is used to parse task instructions through a neural-symbolic collaborative mechanism, and generate action sequences in combination with the UI tree, and then perform symbol system verification and permission status check on the action sequences; The execution optimization layer is used to convert the action sequence into platform-specific operation codes and execute the operation codes; The real-time knowledge base is used to store updated dynamic documents, avoidance strategies and save test results; The multimodal memory pool is used to store the operation path and operation result of executing the operation code.
9. The APP automated testing Agent system based on multimodal perception according to claim 8 is characterized by: The real-time perception layer includes a dynamic response detection module, a YOLOv5 model, a CLIP model, an OCR module and a UI structure generator. The dynamic response detection module is used to perform difference analysis on interface changes through the SSIM algorithm. When an interface change is found, the YOLOv5 model is used to detect the control bounding box, the CLIP model is used to perform visual semantic encoding on the interface elements, and the OCR module is used to synchronously extract text content and generate function labels. The UI structure generator is used to fuse the output results of the three models and construct a UI tree containing hierarchical topological relationships. The dynamic reasoning layer includes a multifunctional module, a causal chain learner, a symbol system and a conflict elimination engine. The multifunctional module is used to perform LLM task parsing and task status detection; the causal chain learner is used to generate an action sequence according to the UI tree and task instructions; the symbol system is used to verify the action sequence, and the conflict elimination engine is used to automatically mark the abnormal path, trigger the path rollback, and feedback to the causal chain learner when it detects that the symbol system verification fails or the operation result is wrong.
Citation Information
Patent Citations
Product UI function identification and defect detection method based on multi-modal large model
CN118467349A
Automatic testing method and system based on deep learning
CN118860859A
Low-code application automatic test system and method based on machine learning
CN119473834A
Electric power cross-modal knowledge fusion multi-agent cooperative processing method and system
CN119477235A
Multi-modal Action Transform model and intelligent task execution method thereof
CN119494078A
Cited By
Action sequence execution method, electronic equipment, storage medium and program product
CN120653167A
Action sequence execution method, electronic device, storage medium, and program product
CN120653167B
Multi-modal large model training method, device and equipment for terminal adjustment
CN121119195A
Multimodal large model training method, device and equipment for terminal adjustment
CN121119195B
Mobile application intelligent agent method
CN121125830A