Task processing method and device based on GUI agent, computer device and storage medium

CN122840256APending Publication Date: 2026-09-29INDUS INTELLIGENT TECHNOLOGY (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611143273.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]本申请实施例提供了基于GUI代理的任务处理方法、装置、计算机设备及存储介质,可以解决现有技术中存在现有的 GUI 智能体无法满足用户深层次的需求的技术问题

Benefits of technology

[0049]本申请实施例提供了基于GUI代理的任务处理方法、装置、计算机设备及存储介质。其中,方法包括响应于推理指令,获取当前屏幕截图以及任务要求;基于所述当前屏幕截图检索节点,得到与所述当前屏幕截图的所反映的UI状态匹配的目标节点;通过调用大模型基于所述任务要求在预设的离线记忆流程图中查询图细节,筛选确定出由多个关键节点有序组成的节点列表;以所述目标节点为起点节点,搜索中间经过所述节点列表中的全部节点直至终点节点的最短路径,得到有序边序列;通过调用大模型基于所述有序边序列中边的特异属性,填充对应的路径参数;按照所述有序边序列的顺序,依次基于所述路径参数在本地执行边对应的脚本。本申请针对特定任务,离线预构建应用级工作流图记忆,并在运行时通过截图检索快速定位当前 UI 状态节点,减少大模型推理步数和端到端延迟。离线应用的专属工作流图记忆,在图上路径规划后执行边脚本,以此替代常规的逐步大模型在线推理。使得GUI智能体满足用户的深层次需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840256A_ABST
    Figure CN122840256A_ABST
Patent Text Reader

Abstract

This application discloses a task processing method, apparatus, computer device, and storage medium based on a GUI agent. The method includes acquiring a current screenshot and task requirements; retrieving nodes from the current screenshot to obtain a target node; querying graph details in a pre-defined offline memory flowchart based on the task requirements and filtering the node list; searching for the shortest path starting from the target node to obtain an ordered edge sequence; filling path parameters based on the specific properties of the edges in the ordered edge sequence; and executing scripts locally based on the path parameters. This application pre-builds an application-level workflow graph memory offline for specific tasks and quickly locates the current UI state node at runtime through screenshot retrieval, reducing the number of inference steps for large models and end-to-end latency. The offline application's dedicated workflow graph memory executes edge scripts after path planning on the graph, replacing the conventional step-by-step online inference of large models, enabling the GUI agent to meet the user's deeper needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimodal large model technology, and in particular to a task processing method, apparatus, computer device and storage medium based on GUI agent. Background Technology

[0002] Existing GUI agents rely on memoryless or trajectory-level retrieval. The large model needs to participate in every step of the task; for each click on a control or page transition, the system must take a new screenshot for the large model to recognize and infer the next action. This means that for tasks with dozens of steps, the large model must perform dozens of inferences, resulting in extremely long overall processing times. Each task execution requires starting from scratch, leading to numerous large model calls and high end-to-end latency, making it unsuitable for speed-sensitive automated deployment scenarios. In short, existing GUI agents fail to meet the deeper needs of users. Summary of the Invention

[0003] This application provides a task processing method, apparatus, computer device, and storage medium based on a GUI agent, which can solve the technical problem that existing GUI agents cannot meet the deep-seated needs of users.

[0004] In a first aspect, embodiments of this application provide a task processing method based on a GUI agent, which includes:

[0005] In response to inference commands, retrieve the current screenshot and task requirements;

[0006] Based on the current screenshot, a target node matching the UI state reflected in the current screenshot is obtained.

[0007] By calling the large model to query the graph details in the preset offline memory flowchart based on the task requirements, a node list composed of multiple key nodes in an orderly manner is determined.

[0008] Starting from the target node, search for the shortest path that passes through all nodes in the node list until the destination node, and obtain an ordered edge sequence.

[0009] The corresponding path parameters are filled in by calling the large model based on the specific properties of the edges in the ordered edge sequence;

[0010] Execute the scripts corresponding to the edges locally in the order of the ordered edge sequence, based on the path parameters.

[0011] In some embodiments, following the sequential execution of the scripts corresponding to the edges locally based on the path parameters according to the ordered edge sequence, the process includes:

[0012] Determine whether to complete the corresponding task based on the aforementioned task requirements;

[0013] If the corresponding task is not completed, search the graph path between the uncompleted node and the endpoint node;

[0014] If the search result indicates an invalid path, a reasoning instruction is generated, and the process returns to the steps of obtaining the current screenshot and task requirements in response to the reasoning instruction, until the corresponding task is determined to be completed based on the task requirements.

[0015] In some embodiments, if the search result indicating the path is invalid, a reasoning instruction is generated, and the process returns to the steps of obtaining the current screenshot and task requirements in response to the reasoning instruction, until the corresponding task is determined to be completed based on the task requirements, including:

[0016] If the search results indicate an invalid path in the graph, the node adjustment tool is invoked to adjust the nodes, and / or the edge adjustment tool is invoked to adjust the edges. The node adjustment includes at least one of modifying or deleting a node, and the edge adjustment includes at least one of modifying or deleting an edge.

[0017] Trigger the generation of inference instructions;

[0018] The return response is in response to the inference instruction, obtains the current screenshot and the steps required by the task, and searches for a path based on the adjusted nodes and / or adjusted edges until the corresponding task is determined to be completed based on the task requirements.

[0019] In some embodiments, the path parameters include edge parameter declarations, and the scripts corresponding to the edges are executed locally based on the path parameters in the order of the ordered edge sequence, including:

[0020] Following the order of the ordered edge sequence, syntactic analysis and segmentation are performed character by character to obtain multiple semantic units;

[0021] Based on the types of multiple semantic units, identify the sentence types and construct the corresponding node tree;

[0022] Write the default value of each parameter in the parameter declaration of the edge into the runtime variable table;

[0023] Based on the runtime variable table, the corresponding node tree is traversed, and the script for each corresponding node is executed.

[0024] In some embodiments, the corresponding node tree includes locating the target text, traversing the corresponding node tree based on the runtime variable table, and executing the script for each corresponding node, including:

[0025] Based on the runtime variable table, the corresponding node tree is traversed. When the target text node is reached, all matching reference targets are collected based on the target text.

[0026] The reference target is located to obtain its reference coordinates on the screen;

[0027] Calculate the similarity between the reference target and the target text and normalize it to obtain the similarity normalization score; and calculate the distance normalization score based on the geometric distance between the coordinates of the reference target and the target text.

[0028] A comprehensive score is determined based on the similarity normalized score and the distance normalized score;

[0029] The reference target with the highest overall score is used as the final positioning target, and the corresponding script is executed.

[0030] In some embodiments, before querying graph details in a preset offline memory flowchart based on the task requirements by calling a large model, and filtering to determine a node list composed of multiple key nodes in an ordered manner, the process includes:

[0031] In response to the command to build the offline memory flowchart, perform initialization processing and application registration;

[0032] Based on the initial perception of the application's initial interface, create the root node;

[0033] Based on the exploration strategy of the current map planning, simulate human interaction on the current interface;

[0034] Record the state transitions during interactive execution, extract the execution actions based on the state transitions to construct action edges between two nodes, and obtain the original graph links;

[0035] The original graph link is corrected to obtain the offline memory flowchart.

[0036] In some embodiments, by invoking a large model to query graph details in a preset offline memory flowchart based on the task requirements, a node list composed of multiple key nodes in an ordered manner is determined, including:

[0037] By calling a large model to perform semantic parsing on the task requirements, the task constraints are extracted.

[0038] Based on the task constraints, a semantic scan is performed on all nodes in the preset offline memory flowchart to recall multiple key nodes that conform to the task semantics.

[0039] Based on the dependencies of multiple key nodes and the pointers in the offline memory flowchart, a node list is obtained by sorting.

[0040] Secondly, embodiments of this application also provide a task processing device based on a GUI agent, comprising:

[0041] The acquisition unit is used to acquire the current screenshot and task requirements in response to inference instructions;

[0042] The retrieval unit is used to retrieve nodes based on the current screenshot and obtain target nodes that match the UI state reflected in the current screenshot.

[0043] The node filtering unit is used to query the graph details in the preset offline memory flowchart by calling the large model based on the task requirements, and filter to determine a node list composed of multiple key nodes in an orderly manner.

[0044] The path planning unit is used to search for the shortest path from the target node to the end node, passing through all nodes in the node list, to obtain an ordered edge sequence.

[0045] The path filling unit is used to fill the corresponding path parameters by calling the large model based on the special properties of the edges in the ordered edge sequence;

[0046] An execution unit is configured to execute the scripts corresponding to the edges locally, based on the path parameters, in the order of the ordered edge sequence.

[0047] Thirdly, embodiments of this application also provide a task processing computer device based on a GUI agent, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.

[0048] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the above-described method.

[0049] This application provides a task processing method, apparatus, computer device, and storage medium based on a GUI agent. The method includes: responding to an inference command by acquiring a current screenshot and task requirements; retrieving nodes based on the current screenshot to obtain a target node matching the UI state reflected in the screenshot; querying graph details in a pre-set offline memory flowchart based on the task requirements using a large model to filter and determine an ordered list of nodes composed of multiple key nodes; using the target node as a starting node, searching for the shortest path passing through all nodes in the node list to the endpoint node to obtain an ordered edge sequence; filling corresponding path parameters based on the specific attributes of the edges in the ordered edge sequence using a large model; and executing the scripts corresponding to the edges locally according to the order of the ordered edge sequence and the path parameters. This application pre-builds an application-level workflow graph memory offline for specific tasks and quickly locates the current UI state node at runtime through screenshot retrieval, reducing the number of large model inference steps and end-to-end latency. The offline application's dedicated workflow graph memory executes edge scripts after path planning on the graph, thus replacing the conventional step-by-step online inference of the large model. This enables the GUI agent to meet the user's deeper needs. Attached Figure Description

[0050] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart illustrating the task processing method based on a GUI agent provided in an embodiment of this application;

[0052] Figure 2 A schematic block diagram of a GUI agent-based task processing device provided in the embodiments of this application;

[0053] Figure 3 A schematic block diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0056] It should also be understood that the terminology used in this application specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this application specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0057] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0058] Existing GUI agents rely on memoryless or trajectory-level retrieval. The large model needs to participate in every step of the task; for each click on a control or page transition, the system must take a new screenshot for the large model to recognize and infer the next action. This means that for tasks with dozens of steps, the large model needs to perform dozens of inferences, resulting in extremely long overall processing times. Each task execution requires starting from scratch, leading to numerous large model calls and high end-to-end latency, making it unsuitable for speed-sensitive automated deployment scenarios. In short, existing GUI agents fail to meet the deeper needs of users.

[0059] To address the aforementioned issues, this application provides a task processing method, apparatus, computer device, and storage medium based on a GUI agent. This application targets GUI automation agent systems, applicable to automated operation scenarios in desktop, mobile, and web applications. In existing technologies, a large model needs to participate in every step of the task: every time a control is clicked or a page is navigated, the system needs to take a new screenshot for the large model to recognize and reason about the next step. In other words, if a task has dozens of steps, the large model has to perform dozens of reasoning operations, resulting in a very long overall time consumption. This application, for specific applications, pre-builds an application-level workflow graph memory offline and quickly locates the current UI state at runtime through screenshot vector retrieval, thereby reducing the number of large model reasoning steps and end-to-end latency. A directed workflow graph describing the UI state transitions of the application is built offline in one go—that is, an application-specific workflow graph memory—and the edges of the graph serve as directly executable operation script units. During actual task execution, the current UI node is quickly located through screenshot vector similarity retrieval, and the edge scripts are executed sequentially after path planning on the graph, thus replacing the conventional step-by-step online reasoning of the large model.

[0060] In some embodiments, this GUI agent-based task processing method is applied to a GUI agent-based task processing computer device. The device deploys a GUI agent, which is an AI assistant capable of autonomously understanding graphical interfaces and operating like a human. It simulates human perception, decision-making, and execution, translating natural language instructions into specific interface operations to help users complete various tasks. It understands the interface by analyzing screenshots or UI structure trees, recognizing elements such as buttons, input boxes, and menus. This makes it independent of specific control IDs and adaptable even to interface changes. It utilizes the reasoning capabilities of a large model to understand user instructions and autonomously plans the operational steps required to complete the goal. It executes the planned actions by simulating clicks, inputs, and swipes, interacting with the software.

[0061] The GUI agent-based task processing computer device can be a terminal or a server, and the terminal can be a smartphone, tablet, handheld computer, or laptop, etc.

[0062] Figure 1 This is a flowchart illustrating the task processing method based on a GUI agent provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps S110-S160:

[0063] S110, In response to the inference command, obtain the current screenshot and task requirements;

[0064] The types of task requirements are numerous. For example, in e-commerce operations, this could involve automated tasks such as cross-platform product listing and price monitoring. In financial risk control, it could involve automating the collection and preliminary analysis of anti-fraud data. In industrial internet scenarios, it could involve closed-loop control tasks for equipment status monitoring and anomaly warning. In enterprise office scenarios, it could involve automating complex business processes such as schedule changes. The specific task requirements are closely linked to the actual work scenarios in which they are applied. Furthermore, the task requirements differ across different application scenarios.

[0065] Capture a screenshot of the current screen to extract visual information. Taking a screenshot is fundamental to achieving visual perception capabilities. A screenshot is needed to see the interface in order to perform analysis and operations.

[0066] S120. Based on the current screenshot, retrieve the target node that matches the UI state reflected in the current screenshot.

[0067] The GUI agent's operating mechanism is not a simple image comparison, but an intelligent retrieval process involving feature extraction. Specifically, this retrieval and matching process can be broken down into the following four stages:

[0068] Based on the current screenshot, it is not searched as a whole image; instead, semantic deconstruction is performed first. All visible text in the screenshot is extracted using OCR (Optical Character Recognition), while an object detection model is used to identify the types and relative positions of UI elements such as buttons, input boxes, and dropdown menus. Subsequently, this information is compiled in real-time into a set of observation state token sequences. These token sequences are equivalent to a digital fingerprint of the current interface, containing the hash values ​​of key identifying text, the type distribution of visible controls, and the coordinate hierarchy of core operational elements.

[0069] A rapid coarse screening of all state nodes stored in the offline memory graph is performed based on the observed state token sequence. This screening avoids comparing pixels frame by frame, instead comparing feature vectors. Hash index matching is prioritized using key text anchors such as pop-up titles, unique button names, and the number of controls to quickly filter out candidate nodes with the most similar visual structures. This step is extremely fast, significantly reducing the scope of precise calculations required and avoiding the performance overhead of traversing the entire offline image.

[0070] Among the candidate nodes initially selected, a refined multi-factor scoring mechanism is activated to calculate a matching score between each candidate node and the current screenshot. This score includes factors such as text hit rate and control layout topology. Text hit rate is given a high weight.

[0071] The text hit rate is calculated by comparing the core guiding text and button labels on the interface with the text stored in the offline node memory. The control layout topology is determined by checking the relative positions of the input box and the confirmation button. Finally, the candidate node with the highest score is selected as the target node. The visual embedding is extracted from the current screenshot, and the most similar nodes are retrieved from the FAISS index. The node with the highest similarity score is used as the starting point for the current UI state.

[0072] In some embodiments, the weight of dynamically changing elements such as timestamps, user avatars, and scrolling list content is automatically reduced to prevent misjudgments caused by changes in irrelevant numbers.

[0073] In some embodiments, if the score is higher than the high threshold of 90%, it is considered a successful exact match. The node is then directly bound to the current process position and the next action edge pointed to by the node is prepared for execution.

[0074] If the score falls between 50% and 90% of the low-to-medium threshold, it is considered a successful fuzzy match. At this point, based on the fault-tolerant anchor points pre-embedded in the offline graph, the abnormal elements of the current interface are automatically attributed to pop-ups, advertisements, or loading delays, and an attempt is made to trigger bypass actions to eliminate interference, restoring the interface to the standard state before confirming again.

[0075] If the score is below 50% of the lower threshold, the match is considered to have failed, meaning the current interface has no record in the memory. In this case, the retrieval action returns an empty result, directly triggering an invalid graph path signal, which then initiates a replanning or help-seeking mechanism.

[0076] This application does not create a global index for the graph, but rather a fast hash index for the core business entry point of the application. When a task starts, the system only needs to match the entry hash value to load the entire offline workflow memory graph belonging to the application in milliseconds, without the need for network parsing.

[0077] In some embodiments, this application implements a dedicated workflow based on an offline memory flowchart. By using the offline memory flowchart at runtime, the current UI state can be quickly located through screenshot vector retrieval, thereby reducing the number of inference steps for large models and end-to-end latency.

[0078] In some embodiments, an offline memory flowchart needs to be built in advance. Before the task is actually run, the user's operational experience on a specific GUI application is transformed into a storable, reusable, and evolvable static topology diagram through recording, abstraction, and structuring. This diagram does not rely on a real-time network or online large model and is stored entirely locally. Building this offline memory flowchart requires going through the following standardized construction stages.

[0079] Prior to S130, steps A1 through A5 are included:

[0080] A1. Responding to the command to build the offline memory flowchart, perform initialization processing and application registration;

[0081] The model calls the API to obtain the graph structure, retrieves a screenshot of the current interface, and locates the current node position. Responding to the offline memory flowchart construction command, it performs initialization and application registration. It initiates an offline exploration task, assigning a unique graph identifier to the target application. It creates a new configuration entry for the application in the graph registry, initializes an empty JSON workflow graph file and a blank FAISS vector index library, preparing to receive data generated by subsequent explorations.

[0082] A2. Based on the initial perception of the application's initial interface, create the root node;

[0083] Develop an exploration plan, perceive the initial state, and create a root node. Open the target application and build an agent to perceive the application's initial interface, such as the homepage. Automatically capture the current screen and save it locally. Based on the current interface, the agent generates semantic description text, covering page hierarchy, main functions, etc. Use a visual language model to extract image vectors from the screenshot and automatically write them into the FAISS vector index. Generate the root node and persist it to a JSON nodes collection.

[0084] In some embodiments, the next action is output, such as clicking an element to enter a new page. If no existing node is matched, a new node is created, an edge is created, and the next action is output. When the model objective is completed, the exploration plan is automatically rewritten.

[0085] A3. Based on the exploration strategy of the current map planning, simulate human interaction on the current interface;

[0086] The agent autonomously explores and makes decisions by examining the adjacency list of the current graph to identify unexplored functional branches, thus formulating a non-repetitive exploration strategy. It then outputs the next interactive action, such as clicking a navigation bar or expanding a dropdown menu. The agent translates its exploration decisions into concrete interactive actions, invoking keyboard and mouse controllers to simulate human actions on the current interface, such as left-clicking, text input, and page scrolling.

[0087] A4. Record the state transitions during interactive execution, extract the execution actions based on the state transitions to construct action edges between two nodes, and obtain the original graph links;

[0088] After physical interaction, the application's UI changes, navigating to a new page or popping up a new window. The agent records this state transition process. Action boundaries are extracted and normalized to create edges. The user's actions from one state to the next are analyzed and extracted as edges connecting two nodes to obtain the original graph links. Physical operations such as left mouse clicks, keyboard input, and scrolling are transcribed into a unified set of action instructions. Preconditions for executing each action are also recorded. For example, clicking must wait until the loading animation disappears; these conditions serve as hard checks for the edge's validity during later execution. First, the agent re-evaluates the new interface. If it's a completely new interface, the node creation tool is invoked again. Then, the agent uses the edge creation tool to connect the preceding and following action logic. The interface nodes before and after the action are entered. Operation parameters are entered, converting the just-executed physical action into a GAPL language script, serving as the core execution logic of the edge. A description is entered, explaining the script's business meaning in natural language, such as "clicking to enter the settings page." After the edge is built, it will automatically attempt to compile and run on the image of the interface node before the action occurs to ensure that the operation is consistent with expectations. If it is inconsistent, it will be regenerated.

[0089] A5. Perform graph correction on the original graph link to obtain the offline memory flowchart.

[0090] Through continuous iterative exploration, the agent constantly expands the graph network. It also possesses self-correction and management capabilities. If the agent discovers an invalid exploration path, such as clicking a dead link or an inaccurate node description, it can dynamically prune and correct it using node or edge modification / deletion tools. After exploration is complete, the system converts the graph containing complete nodes and edges into JSON format for persistent storage and saves the final vector index file. At this point, the offline memory of the target application is officially built and can be activated and invoked by the path planning agent during the task execution phase.

[0091] The above are the steps for automatically constructing an offline memory flowchart. In some embodiments, a semi-automatic construction method using manual recording can also be used to construct the offline memory flowchart.

[0092] As an example, we record every valid user interaction with the GUI interface in offline recording mode. The starting point is to walk through the process. In offline recording mode, we first record every valid user interaction with the GUI interface, such as buttons, input fields, and dropdown menus. Whenever the user clicks or inputs, and the interface undergoes a significant change, such as a pop-up appearing, a page redirect, or a list refresh, the system captures the semantic features of the GUI interface at that critical moment. We record the action at that time, as well as key state nodes, including control hierarchy, key text labels, and element IDs. Each node is not recorded as rigid screen coordinates but is abstracted as a semantic anchor. For example, for a password input field on a login page or an order confirmation submit button, we ensure that the offline graph is insensitive to pixel offsets in the interface.

[0093] Then, the user's control actions for switching states are analyzed, and the corresponding state nodes are recorded. These actions are extracted into action edges between two nodes, resulting in the original link containing state nodes and action edges. Next, the action boundaries are extracted and normalized to establish edges. The user's actions from one state to the next are analyzed and extracted into edges connecting two nodes. Simultaneously, the preconditions for executing the action are recorded. For example, clicking must wait until the loading animation disappears; these conditions will serve as hard validation criteria for the validity of this edge during later execution.

[0094] Next, the original links are simplified to obtain simplified links. This simplification process compresses the paths and eliminates redundancy, as the original recorded operations often contain numerous meaningless intermediate pages, such as accidental back clicks or rapidly flashing transition pages. The offline build engine performs topology simplification on the original links. It directly connects to core functional nodes, eliminating purely visual transition nodes that do not change business data. Simultaneously, it identifies convergence points where multiple different operation paths lead to the same functional endpoint, merges redundant branches, and retains only the most stable main path and one alternative bypass, preventing the memograph from becoming excessively large.

[0095] Based on user-input anomaly markers, corresponding labels are assigned to the simplified link. These labels include invalid path labels and anomaly anchor point labels. Pre-embedded anomaly anchor points and rollback markers are key differences between offline memory graphs and ordinary flowcharts. During construction, corresponding labels are assigned to two types of pre-embedded special nodes: invalid path markers and rollback snapshot anchor points. These correspond to invalid path labels and anomaly anchor point labels, respectively. An invalid path marker is placed at the end of a path that is known to cause errors or freezes. For example, clicking submit without selecting a protocol will permanently leave the user stuck in an error pop-up.

[0096] A rollback snapshot anchor point refers to forcibly setting a state rollback point before performing critical destructive operations, such as deletion or submitting a payment. It records the UI element attributes at that moment, serving as the physical basis for determining which node to revert to and restart from should a replanning be triggered later.

[0097] Finally, the simplified links and their corresponding labels are compressed to obtain the offline memory flowchart, which is then stored locally. A memory sequence and index are constructed to enable offline storage. The simplified links include the aforementioned state nodes and action edges. The corresponding labels include exception markers and rollback points. The offline memory flowchart is a complete directed graph, which is then transformed into highly compressed offline data packets and stored locally.

[0098] S130. By calling the large model to query the graph details in the preset offline memory flowchart based on the task requirements, a node list composed of multiple key nodes in an orderly manner is determined.

[0099] S130 includes steps S1301-S1303:

[0100] S1301. By calling the large model, the task requirements are semantically parsed to extract the task constraints.

[0101] S1302. Based on the task constraints, perform a semantic scan on all nodes in the preset offline memory flowchart and recall multiple key nodes that conform to the task semantics.

[0102] S1303. Based on the dependencies of multiple key nodes and the pointers in the offline memory flowchart, sort them to obtain a node list.

[0103] In the GUI intelligent agent architecture, the process of querying and filtering an ordered list of nodes in the offline memory flowchart by calling a large model involves the large model drawing the most reasonable route on the static offline flowchart. First, intent deconstruction and query condition generation are performed. Instead of directly feeding the entire offline flowchart to the large model, the user's natural language task is first assigned to the large model for intent deconstruction. The large model extracts key constraints. Simultaneously, these constraints are compiled into query keys for the offline flowchart metadata index. Then, based on semantics, graph nodes are coarsely screened. The large model, carrying the query keys, performs a global semantic scan of all nodes in the offline memory flowchart. Since each node in the offline graph pre-stores a semantic token,

[0104] Ignoring irrelevant branches, all nodes that conform to the task semantics are recalled to form a candidate node pool. Secondly, based on graph topology-based path sorting and chain concatenation, the large model, combined with the directed edge connections recorded in the offline graph, performs path deduction on the candidate node pool.

[0105] The model examines the dependencies between candidate nodes. If multiple paths in the offline graph can complete the task, the model selects the main path with the fewest nodes and the most direct state transition, based on time or number of operation steps in the task constraints, and eliminates detour-oriented candidate branches. Finally, the large model arranges these nodes into a strictly linear sequence list according to the edges in the offline graph.

[0106] In some embodiments, a closed-loop check is performed before outputting the final node list. This check verifies whether the first node in the list is the standard entry point for the application and whether the last node contains the task's final state. It also confirms that there must be a valid action edge connecting any two adjacent nodes in the offline graph. After the check passes, this list, composed of multiple key nodes in an ordered manner, is locked and passed to the executor. The executor then searches for and confirms each node step by step according to the screenshot node retrieval logic. The last node is the endpoint node.

[0107] S140. Taking the target node as the starting node, search for the shortest path that passes through all nodes in the node list until the end node, and obtain an ordered edge sequence.

[0108] Path planning is performed using the target node as the starting node and the destination node as the ending node. The path planning agent reads the node category descriptions and adjacency relationships in the graph structure, and autonomously iterates through the graph details according to the user task requirements, ultimately determining an ordered list of milestone nodes and the destination node. The shortest path is then searched, passing through all nodes in the node list until the destination node, resulting in an ordered edge sequence.

[0109] For example, Dijkstra's algorithm is used to search for the shortest path from the current node to the destination through all milestones in the directed workflow graph, resulting in an ordered sequence of edges. The path planning agent interacts with the workflow graph via ReAct, outputting a list of milestone nodes and the destination node. Then, find_shortest_paths() is called to calculate the actual execution path from the current node to the destination, passing through all milestones, in the directed workflow graph.

[0110] Return the list of paths. If the number of paths found is not equal to 1, that is, there are no paths or there is ambiguity, skip the subsequent parameter filling, log the warning and wait for replanning to be triggered.

[0111] S150. By calling the large model based on the special properties of the edges in the ordered edge sequence, the corresponding path parameters are filled in;

[0112] By calling the large model to read the specific attributes of each side of the path, the necessary text / numerical parameters are filled in according to the task instructions.

[0113] Without altering the execution path, specific data is filled into each action along the path based on the context of the current task. Specific attribute placeholders in the edge sequence are parsed to identify reserved parameter gaps, which are labeled with business semantics.

[0114] The large model employs dual context injection. When filling in parameters, it doesn't mechanically replace text but performs context alignment. The large model re-examines the ultimate goal of the entire task to verify the reasonableness of the values ​​to be filled. It references historical trajectories (Tokens) already completed in the current task. If a previous step has just been completed, the corresponding parameters are extracted from the OCR results of the previous step's screenshot when filling the current edge, ensuring the continuity of parameter input. Different strategies are adopted to generate the final filled values ​​for specific special attributes.

[0115] As an example, for input box class attributes, the specific text to be filled in is generated based on the adjacent text labels of the control recorded in the offline graph, combined with the personal information database provided by the user.

[0116] As an example, regarding option logic reasoning, if the current edge points to a dropdown selection or slider, the offline image only records the attribute requiring the selection of departure time. Based on the list of options recognized by OCR from the current screenshot, combined with the user's requirements, it infers which specific text option in the list should be clicked, and fills its index coordinates into the action parameters.

[0117] S160. Execute the scripts corresponding to the edges locally according to the order of the ordered edge sequence and based on the path parameters.

[0118] The edge scripts are executed one by one in the order of the path. After each edge is compiled by the compilation chain, the interpreter calls the positioning engine to perform each step of the operation.

[0119] The path parameters include the parameter declarations of the edges, and S160 includes S1601-S1604:

[0120] S1601. Following the order of the ordered edge sequence, perform syntactic analysis and segmentation character by character to obtain multiple semantic units;

[0121] The GAPL compiler chain transforms the script strings stored on the graph edges into a sequence of operations that can be executed on the screen.

[0122] First, lexical analysis and segmentation are performed. The lexical analysis layer segments the script string into a sequence of tokens by scanning each character. Each token carries its type, original value, and row and column number, which facilitates subsequent error location.

[0123] As an example, if case-insensitive keyword matching is used, non-standard capitalization in the generated script from a large model will not cause parsing errors. As another example, variable name identifiers retain their original capitalization to facilitate finding the corresponding parameter configuration in the parameter declaration table. As yet another example, two-character operators such as =, <=, ==, and = are merged and identified by looking up one character while scanning the current character, avoiding accidental splitting into two independent tokens.

[0124] S1602. Based on the types of multiple semantic units, identify the sentence types and construct the corresponding node tree;

[0125] Then, syntax parsing is performed. The syntax parsing layer implements a recursive descent algorithm, which does not require pre-construction of the analysis table. It directly identifies the statement type based on the current token type and constructs the corresponding AST node.

[0126] Specifically, the parser loops through the tokens until the end of the script, sequentially adding all top-level statements to the root node. When encountering a move statement, it reads the main target identifier and checks if a constraint type follows. If it's IN, it reads the region constraint name; if it's NEAR / NEAR_X / NEAR_Y, it reads the anchor name, writing both the constraint type and anchor name into the node. Parentheses are optional and do not affect semantics. When encountering a conditional statement, it sequentially reads the conditional expression and each branch block, supporting optional else statements. When encountering a loop statement, it reads the loop variable name and the loop upper bound expression, then recursively parses the loop body. When encountering a block structure, it recursively parses the internal statements, naturally supporting arbitrary nesting depths. Conditional expressions only support single-level comparison operations, do not involve arithmetic operations, and the result is either true or false.

[0127] S1603. Write the default value of each parameter in the parameter declaration of the edge into the runtime variable table;

[0128] S1604. Based on the runtime variable table, traverse the corresponding node tree and execute the script for each corresponding node.

[0129] Finally, the interpreter initializes using the parameter declaration table of the graph edges and the underlying screen operation driver as input. Before execution, the interpreter writes the default value of each parameter in the parameter declaration table into the runtime variable table. Subsequent runtime only needs to update the corresponding values ​​in the variable table to dynamically populate the parameters; the script itself does not need to be recompiled. Subsequently, the interpreter traverses the AST node tree, selecting the execution logic for each node according to its type.

[0130] If the node type is a matching threshold, the matching confidence threshold can be specified individually for each image parameter during declaration; otherwise, the default value will be used. Image edges for different controls and different screenshot qualities can be set differently to avoid mismatches or missed matches.

[0131] If the node type is a loop statement, the current loop count is written to the variable table in each iteration. The move, input, and other statements in the loop body can read this count through the variable name to realize parameterized loop operation.

[0132] If the node type is a conditional statement, the comparison expression is evaluated. If the result is true, the then branch is executed; otherwise, the elif branches are checked in turn, and if none of them are satisfied, the else branch is executed.

[0133] The execution layer calls the underlying screen operation driver through a unified interface, which is completely decoupled from the specific implementation method. Each layer of the compilation chain can be tested independently, and it is also easy to replace the underlying execution environment.

[0134] In some embodiments, the corresponding node tree includes locating the target text, and S1604 includes steps B1-B5:

[0135] Existing localization methods based on template matching or OCR cannot automatically select the nearest neighbor when the same control appears repeatedly in the interface. This requires manual intervention or hard-coded indexing and cannot disambiguate duplicate elements. Moreover, existing methods either rely on task-level memory, such as the EchoTrail-GUI class, which expands its retrieval library with use, or they lack memory altogether, requiring pre-built graph memory for each cold start. There is still no systematic solution for application-level pre-built graph memory.

[0136] Previous positioning technologies relied solely on visual features such as OCR text or pure image templates for matching. Once multiple identical controls appeared on the interface, such as a table with the same edit icon at the end of each row, the technology would be unable to make a choice.

[0137] This application proposes a GUI automation path language as a side-script description language, including a lexer, parser, AST, interpreter compilation chain, and introduces spatial relationship constraint operators to handle the disambiguation localization problem of duplicate controls in the interface. This application introduces mandatory constraints based on the spatial physical dimension. Anchor-based spatial relationship operators such as NEAR / NEAR_X / NEAR_Y are designed in the GAPL scripting language. During localization, the system not only evaluates visual similarity but also calculates the geometric distance based on the anchor point and the specific customer name in the current row. A joint scoring mechanism is used to hard-add spatial distance weights at the algorithm's underlying level, increasing the overall score of candidate controls closest to the anchor point. This logically eliminates interference from identical controls in other rows.

[0138] B1. Traverse the corresponding node tree based on the runtime variable table. When traversing to the target text node, collect all matching reference targets based on the target text.

[0139] The interpreter calls the localization engine. The multi-candidate disambiguation process for spatial relationship constraints is as follows: for example, when the MOVE TO statement contains NEAR / NEAR_X / NEAR_Y constraints, all matching candidates are collected for the main target and sorted in descending order of confidence.

[0140] In this context, the distance in NEAR mode is Euclidean distance; the distance in NEAR_X mode is horizontal distance; and the distance in NEAR_Y mode is vertical distance.

[0141] B2. Locate the reference target to obtain its reference coordinates on the screen;

[0142] Locate the reference target and obtain its screen coordinates.

[0143] B3. Calculate the similarity between the reference target and the target text and normalize it to obtain the similarity normalization score, and calculate the distance normalization score based on the geometric distance between the coordinates of the reference target and the target text.

[0144] A comprehensive score is calculated for each candidate. The similarity between the reference target and the target text is calculated and normalized to obtain a similarity normalized score. Based on the geometric distance between the coordinates of the reference target and the target text, a distance normalized score is calculated.

[0145] B4. Determine the comprehensive score based on the similarity normalized score and the distance normalized score;

[0146] As an example, the overall score is calculated as 0.5 × similarity normalized score + 0.5 × distance normalized score. The candidate with the highest overall score is selected as the final localization target.

[0147] B5. Use the reference target with the highest comprehensive score as the final positioning target and execute the corresponding script.

[0148] The essence of executing the corresponding script is to combine static offline memory, dynamically real-time perceived current screenshot coordinates, and injected context data, and transform them into a precise, controllable, and observable physical interaction action through the operating system's underlying API. The execution result is fed back to the state machine in real time, driving the process to the next node or triggering an exception handling branch.

[0149] The executor, based on the action instructions containing the complete parameters of the final target location, converts the abstract target control into specific pixel coordinates on the screen. The executor then compares the current real-time screenshot with the anchor point features of that node in offline memory. Based on the edge features of the control and nearby text anchor points in the current screenshot, the executor dynamically calculates the latest center click coordinates.

[0150] After locking onto the specific coordinates, the executor needs to translate high-level actions such as clicks, swipes, and inputs into low-level instruction sequences that the operating system can execute. For example, on Operating System 1, this would be translated into SendInput or mouse_event system calls; on Operating System 2, it would be mapped to CGEvent; and on Operating System 3, it would be mapped to ADB commands or XCTest.

[0151] For text input, the script does not simulate typing word by word. Instead, it uses the clipboard or an accessibility service interface to quickly fill in the text, ensuring stable and reliable long text input. During script compilation, timeout circuit breaker and exception handling instructions are automatically injected.

[0152] In some embodiments, to ensure the stability of the host system, the execution script is not run directly in the main process of the agent, but is instead delivered to a separate sandbox executor for execution.

[0153] Following S160, steps S170–S190 are included:

[0154] S170. Determine whether to complete the corresponding task based on the task requirements;

[0155] S180. If the corresponding task is not completed, search the graph path between the unfinished node and the endpoint node.

[0156] If the task is not completed and the path is invalid, a replanning process is triggered. The invalid path is abandoned, and a visible temporary passage is used instead. This ensures the long-term stability of the offline memory graph and gives the system flexible adaptability in the face of UI changes or abnormal pop-ups.

[0157] When the task is determined to be incomplete and the graph path is invalid during edge execution, replanning is triggered, returning to the previous action, that is, responding to the inference instruction, obtaining the current screenshot and the steps required by the task.

[0158] When the executor completes an edge action, the supervision module compares the real-time screenshot token with the offline graph node token. If the similarity falls below the threshold and the global task completion flag is set to False, an emergency stop is initiated. All current mouse and keyboard simulations are immediately frozen, and no further attempts are made to execute the next preset edge to prevent the system from straying further down the wrong path. The supervision module encapsulates the core contradiction leading to the decision into a failure context package. This failure context package, along with the core data, is submitted to the large model inference engine.

[0159] The system retrieves the user's initial ultimate goal at the start of the task, along with a list of planned but unexecuted remaining nodes. Upon receiving the screenshot and remaining task steps, the inference engine no longer relies on the old node order in the offline graph but instead performs vision-based zero-shot path deduction. It parses the operable elements of the current interface, analyzes the screenshot, identifies the text labels of all buttons and input boxes on the current interface, and analyzes which operations logically approach the ultimate goal.

[0160] S190. If the search result indicator path is invalid, a reasoning instruction is generated, and the steps of obtaining the current screenshot and task requirements in response to the reasoning instruction are returned until the corresponding task is determined to be completed based on the task requirements.

[0161] In some embodiments, the large model will discard the entire list of remaining nodes after the original planned graph path becomes invalid. Based on the current interface element, a very short temporary remedial action is generated directly. Return to the inference command to complete the inference.

[0162] S190 includes S1901-S1903:

[0163] S1901. If the search result indicates that the path in the graph is invalid, call the node adjustment tool to adjust the nodes, and / or call the edge adjustment tool to adjust the edges, wherein the node adjustment includes at least one of modifying a node or deleting a node, and the edge adjustment includes at least one of modifying an edge or deleting an edge;

[0164] Through continuous iterative exploration, the graph network is constantly expanded. It possesses self-correction and management capabilities.

[0165] If the agent finds that an exploration path is invalid, such as by clicking a dead link or by an inaccurate node description, it can call the node modification / deletion tool or the edge modification / deletion tool at any time to perform dynamic pruning and correction.

[0166] After the exploration is complete, the graph containing all nodes and edges is converted into JSON format for persistent storage, and the final vector index file is saved. At this point, the offline memory of the target application is officially built and can be activated and invoked by the path planning agent during the task execution phase at any time.

[0167] S1902, Trigger the generation of inference instructions;

[0168] S1903. Return to the steps of responding to the inference instruction, obtaining the current screenshot and task requirements, and searching for a path based on the adjusted nodes and / or adjusted edges until the corresponding task is determined to be completed based on the task requirements.

[0169] This application provides a task processing method, apparatus, computer device, and storage medium based on a GUI agent. The method includes: responding to an inference command by acquiring a current screenshot and task requirements; retrieving nodes based on the current screenshot to obtain a target node matching the UI state reflected in the screenshot; querying graph details in a preset offline memory flowchart based on the task requirements using a large model to filter and determine an ordered list of nodes composed of multiple key nodes; using the target node as a starting node, searching for the shortest path passing through all nodes in the node list to the endpoint node to obtain an ordered edge sequence; filling in corresponding path parameters based on the specific attributes of the edges in the ordered edge sequence using a large model; and executing the scripts corresponding to the edges locally according to the path parameters in the order of the ordered edge sequence. This application pre-builds an application-level workflow graph memory offline for specific tasks and quickly locates the current UI state node at runtime through screenshot retrieval, reducing the number of large model inference steps and end-to-end latency. The offline application's dedicated workflow graph memory executes edge scripts after path planning on the graph, replacing the conventional step-by-step online inference of the large model. This allows the GUI agent to meet the user's deeper needs.

[0170] Figure 2 This is a schematic block diagram of a task processing device based on a GUI agent provided in an embodiment of this application. Figure 2 As shown, corresponding to the above-described GUI agent-based task processing method, this application also provides a GUI agent-based task processing apparatus 600. This GUI agent-based task processing apparatus 600 includes a unit for executing the above-described GUI agent-based task processing method, and can be configured in terminals such as desktop computers, tablet computers, and laptops. Specifically, please refer to... Figure 2 The GUI agent-based task processing device 600 includes an acquisition unit 601, a retrieval unit 602, a filtering unit 603, a path planning unit 604, a path filling unit 605, and an execution unit 606, wherein:

[0171] Acquisition unit 601 is used to respond to inference commands and acquire the current screenshot and task requirements;

[0172] The retrieval unit 602 is used to retrieve nodes based on the current screenshot and obtain target nodes that match the UI state reflected in the current screenshot.

[0173] The filtering unit 603 is used to query the graph details in the preset offline memory flowchart based on the task requirements by calling the large model, and filter to determine a node list composed of multiple key nodes in an orderly manner;

[0174] The path planning unit 604 is used to search for the shortest path from the target node to the end node, passing through all nodes in the node list, to obtain an ordered edge sequence.

[0175] The path filling unit 605 is used to fill the corresponding path parameters by calling the large model based on the special properties of the edges in the ordered edge sequence;

[0176] The execution unit 606 is used to execute the scripts corresponding to the edges locally according to the order of the ordered edge sequence and based on the path parameters.

[0177] In some embodiments, after executing the scripts corresponding to the edges locally based on the path parameters in the order of the ordered edge sequence, the execution unit 606 is specifically used for:

[0178] Determine whether to complete the corresponding task based on the aforementioned task requirements;

[0179] If the corresponding task is not completed, search the graph path between the uncompleted node and the endpoint node;

[0180] If the search result indicates an invalid path, a reasoning instruction is generated, and the process returns to the steps of obtaining the current screenshot and task requirements in response to the reasoning instruction, until the corresponding task is determined to be completed based on the task requirements.

[0181] In some embodiments, if the search result indicating the path is invalid, a reasoning instruction is generated, and the process returns to the steps of obtaining the current screenshot and task requirements in response to the reasoning instruction, until the corresponding task is determined to be completed based on the task requirements. The execution unit 606 is specifically used for:

[0182] If the search results indicate an invalid path in the graph, the node adjustment tool is invoked to adjust the nodes, and / or the edge adjustment tool is invoked to adjust the edges. The node adjustment includes at least one of modifying or deleting a node, and the edge adjustment includes at least one of modifying or deleting an edge.

[0183] Trigger the generation of inference instructions;

[0184] The return response is in response to the inference instruction, obtains the current screenshot and the steps required by the task, and searches for a path based on the adjusted nodes and / or adjusted edges until the corresponding task is determined to be completed based on the task requirements.

[0185] In some embodiments, the path parameters include edge parameter declarations, and the scripts corresponding to the edges are executed locally based on the path parameters in the order of the ordered edge sequence. The execution unit 606 is specifically used for:

[0186] Following the order of the ordered edge sequence, syntactic analysis and segmentation are performed character by character to obtain multiple semantic units;

[0187] Based on the types of multiple semantic units, identify the sentence types and construct the corresponding node tree;

[0188] Write the default value of each parameter in the parameter declaration of the edge into the runtime variable table;

[0189] Based on the runtime variable table, the corresponding node tree is traversed, and the script for each corresponding node is executed.

[0190] In some embodiments, the corresponding node tree includes locating the target text, traversing the corresponding node tree based on the runtime variable table, and executing the script for each corresponding node. The execution unit 606 is specifically used for:

[0191] Based on the runtime variable table, the corresponding node tree is traversed. When the target text node is reached, all matching reference targets are collected based on the target text.

[0192] The reference target is located to obtain its reference coordinates on the screen;

[0193] Calculate the similarity between the reference target and the target text and normalize it to obtain the similarity normalization score; and calculate the distance normalization score based on the geometric distance between the coordinates of the reference target and the target text.

[0194] A comprehensive score is determined based on the similarity normalized score and the distance normalized score;

[0195] The reference target with the highest overall score is used as the final positioning target, and the corresponding script is executed.

[0196] In some embodiments, before querying graph details in a preset offline memory flowchart based on the task requirements by calling a large model to filter and determine a node list composed of multiple key nodes in an ordered manner, the device includes an offline memory flowchart construction unit, which is specifically used for:

[0197] Before querying the graph details in the preset offline memory flowchart based on the task requirements by calling the large model, and filtering to determine the node list composed of multiple key nodes in an ordered manner, the following steps are taken:

[0198] In response to the command to build the offline memory flowchart, perform initialization processing and application registration;

[0199] Based on the initial perception of the application's initial interface, create the root node;

[0200] Based on the exploration strategy of the current map planning, simulate human interaction on the current interface;

[0201] Record the state transitions during interactive execution, extract the execution actions based on the state transitions to construct action edges between two nodes, and obtain the original graph links;

[0202] The original graph link is corrected to obtain the offline memory flowchart.

[0203] In some embodiments, the filtering unit 603, when performing a process of querying graph details in a preset offline memory flowchart based on the task requirements by calling a large model, filters and determines a node list composed of multiple key nodes in an ordered manner, specifically for:

[0204] By calling a large model to perform semantic parsing on the task requirements, the task constraints are extracted.

[0205] Based on the task constraints, a semantic scan is performed on all nodes in the preset offline memory flowchart to recall multiple key nodes that conform to the task semantics.

[0206] Based on the dependencies of multiple key nodes and the pointers in the offline memory flowchart, a node list is obtained by sorting.

[0207] In summary, the GUI agent-based task processing device 600 in this embodiment of the application obtains the current screenshot and task requirements in response to inference instructions; retrieves nodes based on the current screenshot to obtain a target node that matches the UI state reflected in the current screenshot; queries graph details in a preset offline memory flowchart based on the task requirements by calling a large model, and filters and determines a node list composed of multiple key nodes in an ordered manner; searches for the shortest path from the target node to the endpoint node by passing through all nodes in the node list, obtaining an ordered edge sequence; fills the corresponding path parameters based on the specific attributes of the edges in the ordered edge sequence by calling a large model; and executes the scripts corresponding to the edges locally according to the order of the ordered edge sequence based on the path parameters. This application pre-builds an application-level workflow graph memory offline for specific tasks and quickly locates the current UI state node at runtime through screenshot retrieval, reducing the number of large model inference steps and end-to-end latency. The dedicated workflow graph memory for offline applications executes edge scripts after path planning on the graph, thereby replacing the conventional step-by-step online inference of the large model. This allows the GUI agent to meet the deeper needs of users.

[0208] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned GUI agent-based task processing device and its various units can be referred to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity, these details will not be repeated here.

[0209] The aforementioned GUI agent-based task processing device can be implemented as a computer program, which can, for example... Figure 3 It runs on the computer device shown.

[0210] Please see Figure 3 , Figure 3 This is a schematic block diagram of a computer device 700 provided in an embodiment of this application. The computer device 700 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.

[0211] See Figure 3 The computer device 700 includes a processor 702, a memory, and a network interface 705 connected via a system bus 701. The memory may include a non-volatile storage medium 703 and internal memory 704.

[0212] The non-volatile storage medium 703 may store an operating system 7031 and a computer program 7032. The computer program 7032 includes program instructions that, when executed, cause the processor 702 to perform a task processing method based on a GUI agent.

[0213] The processor 702 provides computing and control capabilities to support the operation of the entire computer device 700.

[0214] The internal memory 704 provides an environment for the execution of the computer program 7032 in the non-volatile storage medium 703. When the computer program 7032 is executed by the processor 702, the processor 702 can execute a task processing method based on a GUI agent.

[0215] This network interface 705 is used for network communication with other devices. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 700 to which the present application is applied. The specific computer device 700 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0216] The processor 702 is used to run a computer program 7032 stored in the memory to perform the following steps: in response to inference instructions, obtain a current screenshot and task requirements; retrieve nodes based on the current screenshot to obtain a target node that matches the UI state reflected in the current screenshot; query graph details in a preset offline memory flowchart based on the task requirements by calling a large model to filter and determine a node list composed of multiple key nodes in an ordered manner; search for the shortest path from the target node to the endpoint node by using the target node as the starting node, passing through all nodes in the node list; fill in the corresponding path parameters based on the specific properties of the edges in the ordered edge sequence by calling a large model; and execute the scripts corresponding to the edges locally according to the order of the ordered edge sequence based on the path parameters.

[0217] It should be understood that in the embodiments of this application, the processor 702 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0218] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0219] It should be noted that the specific embodiments are merely illustrative examples intended to aid in understanding the technical solutions of the present invention, and are not intended to limit the scope of protection of the present invention. The scope of protection of this application is determined by the contents of the claims. Specific embodiments, parameter ranges, or technical features described in the specification should not be construed as undue limitation or expansion of the scope of protection of the claims. The technical effects described in the specification are only used to illustrate the innovation of the present invention. Any technical solution that does not simultaneously possess all the technical features of the present invention, even if it claims to solve the same technical problem, does not fall within the scope of protection of the present invention.

[0220] Therefore, this application also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the following steps: in response to inference instructions, acquire a current screenshot and task requirements; retrieve nodes based on the current screenshot to obtain a target node matching the UI state reflected in the current screenshot; query graph details in a preset offline memory flowchart based on the task requirements by calling a large model; filter and determine an ordered list of nodes composed of multiple key nodes; using the target node as the starting node, search for the shortest path passing through all nodes in the node list until the ending node, obtaining an ordered edge sequence; fill in the corresponding path parameters based on the specific properties of the edges in the ordered edge sequence by calling a large model; and execute the scripts corresponding to the edges locally according to the order of the ordered edge sequence and the path parameters.

[0221] The storage medium can be any computer-readable storage medium that can store program code, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0222] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0223] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0224] The steps in the methods of this application embodiment can be adjusted, merged, or deleted according to actual needs. The units in the apparatus of this application embodiment can be merged, divided, or deleted according to actual needs. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0225] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.

[0226] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A task processing method based on a GUI agent, characterized in that, The method includes: In response to inference commands, retrieve the current screenshot and task requirements; Based on the current screenshot, a target node matching the UI state reflected in the current screenshot is obtained. By calling the large model to query the graph details in the preset offline memory flowchart based on the task requirements, a node list composed of multiple key nodes in an orderly manner is determined. Starting from the target node, search for the shortest path that passes through all nodes in the node list until the destination node, and obtain an ordered edge sequence. The corresponding path parameters are filled in by calling the large model based on the specific properties of the edges in the ordered edge sequence; Execute the scripts corresponding to the edges locally in the order of the ordered edge sequence, based on the path parameters.

2. The method according to claim 1, characterized in that, Following the ordered edge sequence, and after executing the scripts corresponding to the edges locally based on the path parameters, the process includes: Determine whether to complete the corresponding task based on the task requirements; If the corresponding task is not completed, search the graph path between the uncompleted node and the endpoint node; If the search result indicates an invalid path, a reasoning instruction is generated, and the process returns to the steps of obtaining the current screenshot and task requirements in response to the reasoning instruction, until the corresponding task is determined to be completed based on the task requirements.

3. The method according to claim 2, characterized in that, If the search result indicates an invalid path, a reasoning instruction is generated, and the process returns to the steps of obtaining the current screenshot and task requirements in response to the reasoning instruction, until the corresponding task is determined to be completed based on the task requirements, including: If the search results indicate an invalid path in the graph, the node adjustment tool is invoked to adjust the nodes, and / or the edge adjustment tool is invoked to adjust the edges. The node adjustment includes at least one of modifying or deleting a node, and the edge adjustment includes at least one of modifying or deleting an edge. Trigger the generation of inference instructions; The return response is in response to the inference instruction, obtains the current screenshot and the steps required by the task, and searches for a path based on the adjusted nodes and / or adjusted edges until the corresponding task is determined to be completed based on the task requirements.

4. The method according to claim 1, characterized in that, The path parameters include edge parameter declarations. Following the ordered edge sequence, the scripts corresponding to the edges are executed locally based on the path parameters, including: Following the order of the ordered edge sequence, syntactic analysis and segmentation are performed character by character to obtain multiple semantic units; Based on the types of multiple semantic units, identify the sentence types and construct the corresponding node tree; Write the default value of each parameter in the parameter declaration of the edge into the runtime variable table; Based on the runtime variable table, the corresponding node tree is traversed, and the script for each corresponding node is executed.

5. The method according to claim 4, characterized in that, The corresponding node tree includes locating the target text, traversing the corresponding node tree based on the runtime variable table, and executing the script for each corresponding node, including: Based on the runtime variable table, the corresponding node tree is traversed. When the target text node is reached, all matching reference targets are collected based on the target text. The reference target is located to obtain its reference coordinates on the screen; Calculate the similarity between the reference target and the target text and normalize it to obtain the similarity normalization score; and calculate the distance normalization score based on the geometric distance between the coordinates of the reference target and the target text. A comprehensive score is determined based on the similarity normalized score and the distance normalized score; The reference target with the highest overall score is used as the final positioning target, and the corresponding script is executed.

6. The method according to claim 1, characterized in that, Before querying the graph details in the preset offline memory flowchart based on the task requirements by calling the large model, and filtering to determine the node list composed of multiple key nodes in an ordered manner, the following steps are taken: In response to the command to build the offline memory flowchart, perform initialization processing and application registration; Based on the initial perception of the application's initial interface, create the root node; Based on the exploration strategy of the current map planning, simulate human interaction on the current interface; Record the state transitions during interactive execution, extract the execution actions based on the state transitions to construct action edges between two nodes, and obtain the original graph links; The original graph link is corrected to obtain the offline memory flowchart.

7. The method according to claim 1, characterized in that, By calling the large model to query graph details in the preset offline memory flowchart based on the task requirements, a node list composed of multiple key nodes in an ordered manner is determined, including: By calling a large model to perform semantic parsing on the task requirements, the task constraints are extracted. Based on the task constraints, a semantic scan is performed on all nodes in the preset offline memory flowchart to recall multiple key nodes that conform to the task semantics. Based on the dependencies between multiple key nodes and the pointers in the offline memory flowchart, a node list is obtained by sorting.

8. A task processing device based on a GUI agent, characterized in that, The device includes: The acquisition unit is used to acquire the current screenshot and task requirements in response to inference instructions; The retrieval unit is used to retrieve nodes based on the current screenshot and obtain target nodes that match the UI state reflected in the current screenshot. The filtering unit is used to query the graph details in the preset offline memory flowchart based on the task requirements by calling the large model, and filter to determine a node list composed of multiple key nodes in an orderly manner; The path planning unit is used to search for the shortest path from the target node to the end node, passing through all nodes in the node list, to obtain an ordered edge sequence. The path filling unit is used to fill the corresponding path parameters by calling the large model based on the special properties of the edges in the ordered edge sequence; An execution unit is configured to execute the scripts corresponding to the edges locally, based on the path parameters, in the order of the ordered edge sequence.

9. A task processing computer device based on a GUI agent, characterized in that, The method includes a memory, a processor, and a GUI agent-based task handler stored in the memory and executable on the processor, wherein the processor executes the GUI agent-based task handler to implement the steps of the GUI agent-based task handling method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a program that implements a GUI agent-based task processing method, which is executed by a processor to implement the steps of the GUI agent-based task processing method as described in any one of claims 1 to 7.