A GUI agent skill self-evolution method and system based on execution feedback
Patent Information
- Application Number
- CN202611174895.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-04
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]针对现有GUI智能体在长程动态界面任务中存在的技能表示不结构化、技能执行时静态不可改、失败反馈难以沉淀、外部反思成本较高以及技能检索匹配不准确等问题,本发明提出一种基于执行反馈的GUI智能体技能自进化方法及系统
[0048]1、提高GUI长程任务成功率。本发明将静态技能转化为可在部署阶段持续修订的程序化知识,使GUI智能体能够从真实执行失败中获得可操作改进。在MobileWorld、AndroidWorld和OSWorld三个覆盖移动端与桌面端的GUI基准上,本发明方法均能在无需训练的情况下提升多个基座模型性能,最高增益分别达到+16.2%、+6.0%和+10.5%。
Smart Images

Figure CN122840102A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of artificial intelligence and software engineering, and in particular relates to a method and system for the self-evolution of GUI intelligent agent skills based on execution feedback. Background Technology
[0002] In recent years, GUI agents based on multimodal large language models have been able to observe screenshots, parse accessible trees or interface text according to natural language task instructions, and output actions such as clicking, inputting, swiping, waiting, and responding, thereby completing digital tasks in mobile, desktop, or web environments. To improve the efficiency of long-term tasks, existing methods are gradually introducing skill or experience memory mechanisms to save operation processes, location cues, and historical experiences as external knowledge for reuse in subsequent similar tasks.
[0003] However, real-world GUI environments are dynamic and non-stationary. Application interfaces may experience pop-ups, permission prompts, loading delays, changes in control positions, scrolling of list content, expired accessibility trees, and differences in device resolution. If the agent relies solely on a static plan generated or retrieved once before execution, the plan is prone to failure during execution, leading to repeated clicks, incorrect input, premature termination, or timeouts.
[0004] Existing skill-based GUI agents suffer from four main shortcomings: First, skills are often stored as single long texts or loosely structured prompts, with a mix of planned steps, location hints, exception recovery rules, and failure cases, making it difficult to make localized revisions for specific errors; second, skills are mostly generated or retrieved before execution, and even if it is found during execution that the plan is not suitable for the current interface, there is a lack of an immediate editing mechanism; third, some methods rely on external reflection models or independent reflection processes, increasing latency and costs, and the reflection results are weakly coupled with skill execution; fourth, actionable experiences from failure trajectories often remain in logs and are not converted into searchable, reusable, and auditable skill assets.
[0005] Therefore, there is an urgent need for a GUI agent skill management and task execution mechanism that enables skills to be represented as structured, locally editable procedural knowledge, and to reflect, revise and reuse them based on real execution feedback during the deployment and inference phases, while avoiding retraining the model or introducing a complex multi-model pipeline. Summary of the Invention
[0006] To address the problems of unstructured skill representation, static immutability during skill execution, difficulty in accumulating failure feedback, high external reflection costs, and inaccurate skill retrieval and matching in existing GUI agents for long-term dynamic interface tasks, this invention proposes a GUI agent skill self-evolution method and system based on execution feedback.
[0007] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution:
[0008] In a first aspect, the present invention provides a method for the self-evolution of GUI agent skills based on execution feedback, comprising the following steps:
[0009] S1. Based on the GUI task instructions input by the user, perform metadata retrieval and reuse judgment on the existing skill library. If the skill package in the existing skill library matches the target GUI task contained in the GUI task instructions, reuse the skill package; otherwise, create a new skill package for the GUI agent to complete the target GUI task.
[0010] S2. In the t-th interaction step of the i-th round of interaction in executing the target GUI task, the GUI agent acts as the executor, reads the GUI task instructions, and generates the t-th GUI action based on the historical interaction trajectory and the skill pack obtained in S1, which is then applied to the target execution environment. If the current GUI interface state is found to be inconsistent with the skill pack, then the skill pack is subject to immediate local revision.
[0011] S3. When the i-th round of execution fails or the upper limit of the number of interaction steps is reached, the failure reflection process of information isolation is initiated. The critic in an independent session diagnoses the failure reason of the failed interaction trajectory. Then, the executor completes the trajectory-level revision of the skill package based on the feedback of the critic and registers or updates the revised skill package to the existing skill library. The critic and the executor share the same base model parameters.
[0012] S4. After multiple rounds of interaction, if the target GUI task is executed successfully, the skill pack is marked as verified and the skill pack's retrieval metadata is updated; if the target GUI task fails but the skill pack has been revised, the skill pack is marked as pending review for user confirmation.
[0013] Based on the above scheme, each step can be implemented in the following preferred manner.
[0014] As a preferred embodiment of the first aspect mentioned above, the specific process for metadata retrieval and reuse determination in S1 is as follows:
[0015] S11. Traverse the skill packs in the existing skill library, obtain the search metadata of the current skill pack, and extract the query term set containing the search metadata. The skill package will be constructed from the retrieval information extracted from the retrieval metadata. Search metadata collection Let the set of overlapping words be denoted as Calculate GUI task instructions Base match score with this skill pack :
[0016]
[0017] in, Size of the set;
[0018] S12. Add a preset application matching bonus to the base matching score. Preset keyword matching rewards and divergent punishment Receive GUI task instructions The final search score for that skill pack :
[0019]
[0020] in, This is the clipping function;
[0021] S13. Compare the final search score with the preset search threshold: When the final search score is greater than the search threshold, mark the corresponding skill pack as verified and reuse the skill pack; when the final search score is less than or equal to the search threshold, create a new skill pack according to the GUI task instructions.
[0022] As a preferred option of the first aspect mentioned above, in S1, the specific process for creating a new skill package is as follows: the reusable procedural knowledge contained in the skill package is split into three types of skill files: retrieval metadata, accessibility tree assistance tools, failure case set, and executable plan file, backup location and identification strategy file, and failure recovery rule file, and a directory structure for the skill package is established.
[0023] As a preferred embodiment of the first aspect mentioned above, the skill set includes meta_info.json, ally_utils, plan.md, backup.md, recover.md, and failure_examples; metadata is written in meta_info.json; executable steps and verification points from the initial interface screenshot to task completion are written in plan.md; alternative control text and interface layout are written in backup.md; pop-ups, permission requests, loading failures, loss of input focus, or page redirection errors are written in recover.md; and failure_examples are diagnostically confirmed failure cases.
[0024] As a preferred embodiment of the first aspect mentioned above, in S2, the process of performing real-time partial revisions to the skill pack is as follows:
[0025] S21. When the executor finds that the current GUI interface state is inconsistent with the skill package content, it determines the type of inconsistency: if the control text or interface layout has changed, it modifies backup.md in the skill package; if an unexpected pop-up, permission request, loading failure, or page jump error occurs, it modifies recover.md in the skill package; if the current execution plan lacks confirmation, checking, scrolling, or rollback steps, it modifies plan.md in the skill package.
[0026] S22. Through the restricted tool interface, perform read, write, append, search, list, or create_failure operations on the skill files to be revised in the skill package to obtain the skill package after the i-th round of immediate local revision; where read is used to read the skill files in the skill package; write is used to overwrite the skill files in the skill package; append is used to add content to the skill files in the skill package; search is used to search for keywords in the skill package; list is used to list the skill package directory; create_failure is used to create a failure case file and write the failure interaction trajectory, structured diagnostic results, and anti-duplicate rules.
[0027] As a preferred embodiment of the first aspect mentioned above, S22 also provides a read-only accessibility tree observation tool for returning the accessibility tree of the current GUI interface, enabling the executor to read and revise skill files within a preset range, while avoiding arbitrary modification of external resources of the skill pack.
[0028] As a preferred embodiment of the first aspect mentioned above, the specific process of S3 is as follows:
[0029] S31. A GUI agent acts as a critic and diagnoses the cause of failure of the failed interaction trajectory based on the GUI task instructions, observation sequence, and action sequence in the i-th round of failed interaction trajectory, obtaining a structured diagnostic result. The failed interaction trajectory also includes the failure location and visible evidence representing the failure of the target GUI, the observation sequence. The action sequence consists of observations of each interaction step in the i-th round. It consists of GUI actions in each interaction step of the i-th round;
[0030] S32. Based on the execution trajectory of round i, the skill pack revised locally in round i, and the structured diagnostic results fed back by the critic, the executor generates the revised skill pack for round i+1, and determines whether the failed interaction trajectory can be used as a failure case: if it can, it considers the wrong action, wrong answer, or wrong location to have a risk of recurrence, and creates a failure case file through create_failure in subsequent interaction steps, while writing the GUI task instruction, structured diagnostic results, and failed interaction trajectory into the failure case set; if it cannot, it does not create a failure case file through create_failure, nor does it write it into the failure case set.
[0031] As a preferred embodiment of the first aspect mentioned above, in S31, the information set of the constraint critic is... To ensure that the critic can only read information from the information set during diagnosis, the following information isolation conditions must be met:
[0032]
[0033]
[0034]
[0035] in, A set of GUI task instructions; For skill packs; This represents the inference chain within the actuator; This is the true answer; It is the set of all information that the executor can access during runtime.
[0036] As a preferred embodiment of the first aspect mentioned above, in S31, the executor, based on the initial skill pack of the i-th round of interaction, completes a single round of interaction in the target execution environment, from start to successful execution, or from start to execution failure, or from start to the maximum number of execution steps, to obtain the execution trajectory of the i-th round.
[0037] Secondly, the present invention provides a GUI agent skill self-evolution system based on execution feedback, comprising:
[0038] The skill package confirmation module is used to perform metadata retrieval and reuse judgment on the existing skill library based on the GUI task instructions input by the user. When the skill package in the existing skill library matches the target GUI task contained in the GUI task instructions, the skill package is reused; otherwise, a new skill package is created for the GUI agent to complete the target GUI task.
[0039] The local revision module is used in the t-th interaction step of the i-th round of interaction in the execution of the target GUI task. The GUI agent acts as the executor, reads the GUI task instructions, and generates the t-th step GUI action based on the historical interaction trajectory and the skill pack obtained by S1, which is then applied to the target execution environment. If the current GUI interface state is found to be inconsistent with the skill pack, the skill pack is immediately revised locally.
[0040] The trajectory-level revision module is used to initiate a failure reflection process with information isolation when the i-th round of execution fails or the upper limit of interaction steps is reached. The critic in an independent session diagnoses the failure reason of the failed interaction trajectory, and then the executor completes the trajectory-level revision of the skill pack based on the feedback of the critic, and registers or updates the revised skill pack to the existing skill library. The critic and the executor share the same base model parameters.
[0041] The management module is used to mark the skill pack as verified and update the skill pack's retrieval metadata if the target GUI task is successfully executed after multiple rounds of interaction; if the target GUI task fails but the skill pack has been revised, the skill pack is marked as pending review for user confirmation.
[0042] Thirdly, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the GUI agent skill self-evolution method based on execution feedback as described in any of the solutions of the first aspect above.
[0043] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the GUI agent skill self-evolution method based on execution feedback as described in any of the solutions of the first aspect above.
[0044] Fifthly, the present invention provides a computer electronic device, which includes a memory and a processor;
[0045] The memory is used to store computer programs;
[0046] The processor is configured to, when executing the computer program, implement the GUI agent skill self-evolution method based on execution feedback as described in any of the embodiments of the first aspect above.
[0047] Compared with the prior art, the present invention has the following advantages:
[0048] 1. Improve the success rate of long-term GUI tasks. This invention transforms static skills into procedural knowledge that can be continuously revised during the deployment phase, enabling GUI agents to gain actionable improvements from real-world execution failures. On three GUI benchmarks covering mobile and desktop platforms—MobileWorld, AndroidWorld, and OSWorld—the method of this invention can improve the performance of multiple pedestal models without training, with maximum gains of +16.2%, +6.0%, and +10.5%, respectively.
[0049] 2. Reduced difficulty of skill maintenance and revision. This invention, by splitting plans, backup locations, recovery rules, accessibility tools, and failure cases into different files, enables the mapping of different types of errors to corresponding files for local revision. In experiments, the structured skill package increased the success rate on MobileWorld from 66.67% to 69.52% compared to single-file skill representations.
[0050] 3. Reduce cascading failures caused by dynamic interfaces. This invention utilizes an instant local revision mechanism during task execution, allowing the GUI agent to immediately correct skill files upon detecting pop-ups, loading delays, control movement, or plan expiration. Ablation experiments show that removing instant local revisions reduced the MobileWorld success rate from 69.52% to 62.86%.
[0051] 4. Improve the quality of failure reflection and reduce reliance on external models. The critic in this invention diagnoses failures based solely on information visible in the execution trajectory, avoiding inauthentic revisions caused by using true answers or hiding context. Ablation experiments show that the success rate decreased from 69.52% to 60.95% after removing the critic.
[0052] 5. Enhance the reusability of similar tasks. Metadata-based skill retrieval outperforms full-text matching of long texts. In 12 MobileWorld reuse cases, the metadata retrieval of this invention achieved an average score of 0.8753, with a hit rate of 12 / 12 at a threshold of 0.6; the full-text retrieval achieved an average score of 0.4050, with a hit rate of only 1 / 12.
[0053] 6. It possesses the advantages of being training-free and cross-model adaptable. This invention does not rely on retraining or fine-tuning model parameters, but achieves self-evolution only through skill file revisions during the inference and deployment phases. Therefore, it can be used as an external framework to connect to different closed-source or open-source GUI models, and can be incrementally deployed in existing GUI automation systems. Attached Figure Description
[0054] Figure 1 This is a flowchart of the steps of the method of the present invention;
[0055] Figure 2 This is a diagram illustrating the skill pack's catalog.
[0056] Figure 3 This is a diagram illustrating the revision path of skill documents. Detailed Implementation
[0057] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.
[0058] In the description of this invention, it should be understood that the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature.
[0059] like Figure 1 As shown, in a preferred embodiment of the present invention, the above-mentioned GUI agent skill self-evolution method based on execution feedback includes the following steps S1 to S4. The specific implementation process of each step will be described in detail below.
[0060] S1. Based on the GUI task instructions input by the user, perform metadata retrieval and reuse judgment on the existing skill library. If the skill package in the existing skill library matches the target GUI task contained in the GUI task instructions, reuse the skill package; otherwise, create a new skill package for the GUI agent to complete the target GUI task.
[0061] It should be noted that in step S1 of this invention, the target GUI task can run in a target execution environment such as a mobile application, desktop software, web application, virtual machine environment, or real device environment. At each step, the GUI agent receives the current screen screenshot and the accessibility tree, and outputs GUI actions such as click, long press, double click, drag, swipe, text input, back, home, wait, answer, and end. It can also call the skill packs defined in the existing skill library of this invention.
[0062] It should be noted that the specific process for metadata retrieval and reuse judgment in step S1 of this invention is as follows:
[0063] S11. Traverse the skill packs in the existing skill library, obtain the search metadata of the current skill pack, and extract the query term set containing the search metadata. The skill package will be constructed from the retrieval information extracted from the retrieval metadata. Search metadata collection Let the set of overlapping words be denoted as Calculate GUI task instructions Base match score with this skill pack :
[0064]
[0065] in, The size of the set.
[0066] In this embodiment S11, the retrieval metadata for each skill pack in the existing skill library includes a skill identifier. Task Intent Target application Operating platform Search keywords Reusable input parameter slots Usage history And the verification status label indicating whether the skill pack has been verified. .
[0067] In this embodiment S11, the first term of the basic matching score is used to characterize the query coverage, and the second term is used to characterize the set tightness between the GUI task instructions and the retrieval metadata of the candidate skill pack.
[0068] S12. Add a preset application matching bonus to the base matching score. Preset keyword matching rewards and divergent punishment Receive GUI task instructions The final search score for that skill pack :
[0069]
[0070] in, This is the clipping function.
[0071] In this embodiment S12, in addition to two additional rewards based on the basic matching score, a divergence penalty is also applied to key concepts unique to GUI task instructions but missing from the skill pack. In the final retrieval score calculation, this embodiment does not directly perform full-text matching of long skill texts, but rather indexes and scores the skill metadata. This avoids incorrect matching of long plan texts from different tasks due to shared common GUI terms, and also allows applications, platforms, task intents, keywords, and parameter slots to have a more direct impact on reuse decisions.
[0072] S13. Compare the final search score with the preset search threshold. Comparison: When the final search score is greater than the search threshold, the corresponding skill pack is marked as verified and reused; when the final search score is less than or equal to the search threshold, a new skill pack is created according to the GUI task instructions.
[0073] In this embodiment S13, the above-mentioned retrieval threshold can be adjusted according to the size of the skill base, task similarity requirements, or the cost of misuse. Preferably, the retrieval threshold... Take 0.6.
[0074] In S13 of this invention, skills are not saved as a single long text prompt, but rather as a structured, locally editable, and auditable multi-file package. This design avoids mixing planned steps, location prompts, exception recovery rules, and failure experiences in the same long text, making it difficult to subsequently make local revisions for specific errors.
[0075] It should be noted that, in step S1 of this invention, the specific process for creating a new skill package is as follows: The reusable procedural knowledge contained in the skill package is broken down into retrieval metadata. Accessibility Tree Assistance Tool A collection of failure cases and an executable plan file Backup location and identification strategy documents Failure recovery rule file These three skill files, and establish the directory structure of the skill pack.
[0076] In this embodiment, as Figure 2 As shown, the executable plan file mainly answers "what to do"; the backup location and identification strategy file mainly answers "how to find"; and the failure recovery rule file mainly answers "how to recover when the status does not meet expectations". When the error is due to a missing high-level process, this embodiment prioritizes revising the executable plan file; when the error is due to inaccurate control identification or location, this embodiment prioritizes revising the backup location and identification strategy file; when the error is due to pop-ups, loading, page redirection, or permission blocking, this embodiment prioritizes revising the failure recovery rule file.
[0077] In this embodiment, the skill package includes at least meta_info.json, ally_utils, plan.md, backup.md, recover.md, and failure_examples. Specifically, this embodiment writes retrieval metadata in meta_info.json; writes executable steps and verification points from the initial interface screenshot to task completion in plan.md; writes alternative control text and interface layout (serial number, OCR keywords, or parent-child node relationships) in backup.md; writes failure recovery rules for pop-ups, permission requests, loading failures, loss of input focus, or page jump errors in recover.md; and writes diagnosed and confirmed failure cases in failure_examples. The alternative control text refers to which buttons or icons can be used as substitutes when the button or icon specified in plan.md does not exist.
[0078] S2. In the t-th interaction step of the i-th round of the target GUI task, the GUI agent acts as the executor, reads the GUI task instructions, and generates the t-th GUI action based on the historical interaction trajectory and the skill pack obtained in S1. It operates on the target execution environment; if the current GUI interface state is found to be inconsistent with the skill pack, then an immediate local revision is performed on the skill pack.
[0079] It should be noted that in step S2 of this invention, the historical interaction trajectory express , They represent the first Step-by-step observation; They represent the first Step action; observation at step t From the current screenshot and barrier-free tree composition.
[0080] It should be noted that in step S2 of the present invention, as Figure 3 As shown, the process of performing immediate local revisions to a skill pack is as follows:
[0081] S21. When the executor finds that the current GUI interface state is inconsistent with the skill package content, it determines the type of inconsistency: if the control text or interface layout has changed, then revise backup.md in the skill package; if an unexpected pop-up, permission request, loading failure, or page jump error occurs, then revise recover.md in the skill package; if the current execution plan lacks confirmation, checking, scrolling, or rollback steps, then revise plan.md in the skill package.
[0082] S22. Perform read, write, append, search, list, or create_failure operations on the skill files to be revised in the skill package through the restricted tool interface to obtain the skill package after the i-th round of real-time partial revision. Among them, read is used to read skill files in the skill pack; write is used to overwrite skill files in the skill pack; append is used to append content to skill files in the skill pack; search is used to search for keywords in the skill pack; list is used to list the skill pack directory; create_failure is used to create a failure case file and write failure interaction traces, structured diagnostic results and rules to prevent duplicates.
[0083] Furthermore, S22 also provides a read-only accessibility tree observation tool to return the accessibility tree of the current GUI interface, enabling the executor to read and revise skill files within a preset range, while avoiding arbitrary modification of external resources of the skill pack.
[0084] It should be noted that in step S22 of this invention, the restricted tool interface is used to ensure that skills are editable and the scope of revision is controllable. This interface only allows operation on skill files or directories within the skill package, and does not allow modification of system prompts, external data, environmental states, model parameters, or resources outside the skill package mode. In this embodiment, the read-only accessibility tree observation tool only enhances interface observation and does not modify the skill package; the restricted tool interface is used for skill revision. Thus, "interface understanding" and "skill editing" are separated at the tool level, making the skill self-evolution process controllable and auditable.
[0085] S3. When the i-th round of execution fails or the upper limit of the number of interaction steps is reached, the failure reflection process of information isolation is initiated. The critic in an independent session diagnoses the failure reason of the failed interaction trajectory. Then, the executor completes the trajectory-level revision of the skill package based on the feedback of the critic and registers or updates the revised skill package to the existing skill library. The critic and the executor share the same base model parameters.
[0086] It should be noted that in step S3 of this invention, the critic does not use a stronger external supervision model, nor does it accept true answers, hidden annotations, internal thought chains of the executor, or the full text of the current skill pack, in order to avoid using information not available at deployment time for inaccurate diagnosis. The critic determines the cause of failure solely based on GUI task instructions and failure interaction trajectories.
[0087] In step S3 of this embodiment, each GUI task is allowed a maximum of 50 interaction steps. After a failed attempt, a maximum of two rounds of skill revision are allowed. The above values are preferred experimental settings and can be adjusted according to task length, system resources, or reliability requirements.
[0088] It should be noted that the specific process of step S3 in this invention is as follows:
[0089] S31. A GUI agent acts as a critic and diagnoses the cause of failure of the failed interaction trajectory based on the GUI task instructions, observation sequence, and action sequence in the i-th round of failed interaction trajectory, obtaining a structured diagnostic result. The failed interaction trajectory also includes the failure location and visible evidence representing the failure of the target GUI, the observation sequence. The action sequence consists of observations of each interaction step in the i-th round. It consists of GUI actions in each interaction step of the i-th round.
[0090] In this embodiment S31, the structured diagnostic result includes at least whether the target GUI task was successful, the failure type, the initial error step, visible evidence, the direct cause of the failure, the root cause of the failure, and actionable revision suggestions. The failure type may include planning gaps, location errors, verification omissions, outdated accessibility trees, environmental errors, unhandled abnormal states, or other errors. In this way, failure experiences that were originally confined to logs are transformed into searchable, reusable, and auditable skill assets.
[0091] Furthermore, in this embodiment S31, the information set of the constraint critic... To ensure that the critic can only read information from the information set during diagnosis, the following information isolation conditions must be met:
[0092]
[0093]
[0094]
[0095] in, A set of GUI task instructions; For skill packs; This represents the inference chain within the actuator; This is the true answer; It is the set of all information that the executor can access during runtime.
[0096] The aforementioned information isolation condition means that the critic only uses information visible in the executed trajectory and does not read the full text of the skill pack, the internal inference chain of the executor, or the true answer.
[0097] S32, The actuator follows the execution trajectory of the i-th round. The skill pack after the i-th round of real-time partial revisions and the structured diagnostic results from the critic feedback Generate the revised skill pack for round i+1. The system determines whether a failed interaction trajectory can be considered a failure case. If it can, the system considers the incorrect action, incorrect answer, or incorrect location to have a risk of recurrence and creates a failure case file using create_failure in subsequent interaction steps. At the same time, the GUI task instructions, structured diagnostic results, and failed interaction trajectories are written into the failure case set. If it cannot, the system does not create a failure case file using create_failure, nor does it write the failure case into the failure case set.
[0098] In this embodiment S32, the executor, based on the initial skill pack of the i-th round of interaction, completes a single round of interaction in the target execution environment, from start to successful execution, or from start to execution failure, or from start to the maximum number of execution steps, to obtain the execution trajectory of the i-th round.
[0099] It should be noted that in steps S3-S4 of this invention, skill revision includes immediate local revision during the i-th round of interaction execution and trajectory-level revision after the i-th round of interaction execution fails. The specific process is as described above and will not be repeated here. Among them, immediate local revision is used to handle local state deviations such as pop-ups, loading delays, changes in control text, changes in target position, expired accessibility tree nodes, or continuous failures in a certain step; trajectory-level revision is used to handle planning gaps, verification omissions, or recurring erroneous actions.
[0100] S4. After multiple rounds of interaction, if the target GUI task is executed successfully, the skill pack is marked as verified, and the search metadata of the skill pack is updated (specifically, the usage history and verification status labels in the search metadata are updated); if the target GUI task fails but the skill pack has been revised, the skill pack is marked as pending review for user confirmation.
[0101] It should be noted that, in this embodiment, the aforementioned base model can be a closed-source multimodal large language model or an open-source GUI-specific model. It is worth noting that the method of this invention is not limited to a specific multimodal large language model; any model capable of reading GUI screenshots, interface text, task instructions, and skill context and outputting structured GUI actions can be used as an executor or critic.
[0102] To better demonstrate the specific implementation and technical effects of the present invention, the GUI agent skill self-evolution method based on execution feedback shown in steps S1 to S4 of the above preferred implementation is applied to a specific example.
[0103] Example
[0104] The specific implementation process of the GUI agent skill self-evolution method based on execution feedback used in this embodiment is as described above and will not be repeated here.
[0105] First, this embodiment verifies GUI tasks on the MobileWorld mobile platform. The experiment selects a subset of GUI tasks from MobileWorld and compares the task success rates of multiple base models without using this invention and with EvoSkill-GUI. The results are shown in Table 1. Claude-Sonnet-4.6 improved from 57.1% to 67.6%; Qwen3.6-Plus improved from 53.3% to 69.5%; Qwen3.6-35B-A3B improved from 32.4% to 44.8%; and MAI-UI-8B improved from 29.5% to 37.1%.
[0106] Table 1. Comparison of results for different methods in MobileWorld
[0107]
[0108] Secondly, this embodiment verified skill reuse on AndroidWorld related tasks. This embodiment constructed 116+116 AndroidWorld task flows. The first stage used seed 30 tasks, and the second stage used related seed 42 variant tasks. The results are shown in Tables 2 and 3. In the first stage, the skill reuse rate was 37.9%, and the success rate increased from 68.1% to 70.7%; in the second stage, the reuse rate reached 100.0%, and the success rate increased from 55.2% to 61.2%. Of the 89 initially failed tasks, this invention recovered 27 overall, of which 19 were recovered through skill reuse.
[0109] Table 2. Failure Recovery Analysis by Skill Source on AndroidWorld
[0110]
[0111] Table 3. Results of the AndroidWorld Two-Phase Skills Building and Reuse Assessment
[0112]
[0113] Finally, this embodiment was validated on the OSWorld desktop client. The experiment used the OSWorld-Verified benchmark and evaluated applications such as Chrome, GIMP, LibreOffice, Thunderbird, VLC, and VS Code. The results are shown in Table 4. After incorporating this invention into GUI-Owl-1.5-8B, the overall success rate increased from 46.7% to 54.8%, an improvement of +8.1; after incorporating this invention into Qwen3-VL-8B-Instruct, the overall success rate increased from 23.8% to 34.3%, an improvement of +10.5.
[0114] Table 4. Comparison of results from different methods in OSWorld
[0115]
[0116] It should also be noted that the GUI agent skill self-evolution method based on execution feedback in the above embodiments can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a GUI agent skill self-evolution system based on execution feedback, corresponding to the GUI agent skill self-evolution method based on execution feedback provided in the above embodiments, comprising:
[0117] The skill package confirmation module is used to perform metadata retrieval and reuse judgment on the existing skill library based on the GUI task instructions input by the user. When the skill package in the existing skill library matches the target GUI task contained in the GUI task instructions, the skill package is reused; otherwise, a new skill package is created for the GUI agent to complete the target GUI task.
[0118] The local revision module is used in the t-th interaction step of the i-th round of interaction in the execution of the target GUI task. The GUI agent acts as the executor, reads the GUI task instructions, and generates the t-th step GUI action based on the historical interaction trajectory and the skill pack obtained by S1, which is then applied to the target execution environment. If the current GUI interface state is found to be inconsistent with the skill pack, the skill pack is immediately revised locally.
[0119] The trajectory-level revision module is used to initiate a failure reflection process with information isolation when the i-th round of execution fails or the upper limit of interaction steps is reached. The critic in an independent session diagnoses the failure reason of the failed interaction trajectory, and then the executor completes the trajectory-level revision of the skill pack based on the feedback of the critic, and registers or updates the revised skill pack to the existing skill library. The critic and the executor share the same base model parameters.
[0120] The management module is used to mark the skill pack as verified and update the skill pack's retrieval metadata if the target GUI task is successfully executed after multiple rounds of interaction; if the target GUI task fails but the skill pack has been revised, the skill pack is marked as pending review for user confirmation.
[0121] It is understood that the GUI agent skill self-evolution method based on execution feedback described in S1-S4 above can essentially be implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the GUI agent skill self-evolution method based on execution feedback provided in the above embodiments, which includes a computer program / instruction. When the computer program / instruction is executed by a processor, it can implement the GUI agent skill self-evolution method based on execution feedback as described in the above embodiments.
[0122] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the GUI intelligent agent skill self-evolution method based on execution feedback provided in the above embodiments, which includes a memory and a processor;
[0123] The memory is used to store computer programs;
[0124] The processor is configured to implement the GUI agent skill self-evolution method based on execution feedback in the above embodiments when executing the computer program.
[0125] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0126] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the GUI agent skill self-evolution method based on execution feedback provided in the above embodiments. The storage medium stores a computer program, which, when executed by a processor, can implement the GUI agent skill self-evolution method based on execution feedback in the above embodiments.
[0127] Specifically, in the computer-readable storage medium of the above three embodiments, the stored computer program is executed by a processor, which can perform the aforementioned steps S1 to S4.
[0128] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.
[0129] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0130] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.
[0131] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.
Claims
1. A method for self-evolution of GUI agent skills based on execution feedback, characterized in that, Includes the following steps: S1. Based on the GUI task instructions input by the user, perform metadata retrieval and reuse judgment on the existing skill library. If the skill package in the existing skill library matches the target GUI task contained in the GUI task instructions, reuse the skill package; otherwise, create a new skill package for the GUI agent to complete the target GUI task. S2. In the t-th interaction step of the i-th round of interaction in executing the target GUI task, the GUI agent acts as the executor, reads the GUI task instructions, and generates the t-th GUI action based on the historical interaction trajectory and the skill pack obtained in S1, which is then applied to the target execution environment. If the current GUI interface state is found to be inconsistent with the skill pack, then the skill pack is subject to immediate local revision. S3. When the i-th round of execution fails or the upper limit of the number of interaction steps is reached, the failure reflection process of information isolation is initiated. The critic in an independent session diagnoses the failure reason of the failed interaction trajectory. Then, the executor completes the trajectory-level revision of the skill package based on the feedback of the critic and registers or updates the revised skill package to the existing skill library. The critic and the executor share the same base model parameters. S4. After multiple rounds of interaction, if the target GUI task is executed successfully, the skill pack is marked as verified and the search metadata of the skill pack is updated. If the target GUI task fails but the skill pack has been revised, the skill pack will be marked as pending review for user confirmation.
2. The GUI agent skill self-evolution method based on execution feedback as described in claim 1, characterized in that, In S1, the specific process for metadata retrieval and reuse determination is as follows: S11. Traverse the skill packs in the existing skill library, obtain the search metadata of the current skill pack, and extract the query term set containing the search metadata. The skill package will be constructed from the retrieval information extracted from the retrieval metadata. Search metadata collection Let the set of overlapping words be denoted as Calculate GUI task instructions Base match score with this skill pack : ; in, Size of the set; S12. Add a preset application matching bonus to the base matching score. Preset keyword matching rewards and divergent punishment Receive GUI task instructions The final search score for that skill pack : ; in, This is the clipping function; S13. Compare the final search score with the preset search threshold: When the final search score is greater than the search threshold, mark the corresponding skill pack as verified and reuse the skill pack; when the final search score is less than or equal to the search threshold, create a new skill pack according to the GUI task instructions.
3. The GUI agent skill self-evolution method based on execution feedback as described in claim 1, characterized in that, In S1, the specific process for creating a new skill pack is as follows: the reusable procedural knowledge contained in the skill pack is broken down into three types of skill files: retrieval metadata, accessibility tree assistance tools, failure case collection, and executable plan file, backup location and identification strategy file, and failure recovery rule file, and the directory structure of the skill pack is established.
4. The GUI agent skill self-evolution method based on execution feedback as described in claim 3, characterized in that, The skill set includes meta_info.json, ally_utils, plan.md, backup.md, recover.md, and failure_examples. Meta_info.json contains the retrieved metadata; plan.md contains the executable steps and verification points from the initial screenshot to task completion; backup.md contains the alternative control text and interface layout; recover.md contains pop-ups, permission requests, loading failures, lost input focus, or page redirection errors; and failure_examples contains diagnosed and confirmed failure cases.
5. The GUI agent skill self-evolution method based on execution feedback as described in claim 4, characterized in that, In S2, the process of performing real-time partial revisions to skill packs is as follows: S21. When the executor finds that the current GUI interface state is inconsistent with the skill pack content, it determines the type of inconsistency: if the control text or interface layout has changed, it modifies backup.md in the skill pack. If unexpected pop-ups, permission requests, loading failures, or page redirection errors occur, revise the recover.md file in the skill pack; If the current execution plan lacks confirmation, inspection, rolling, or rollback steps, revise the plan.md file in the skill pack; S22. Through the restricted tool interface, perform read, write, append, search, list, or create_failure operations on the skill files to be revised in the skill package to obtain the skill package after the i-th round of immediate local revision; where read is used to read the skill files in the skill package; write is used to overwrite the skill files in the skill package; append is used to add content to the skill files in the skill package; search is used to search for keywords in the skill package; list is used to list the skill package directory; create_failure is used to create a failure case file and write the failure interaction trajectory, structured diagnostic results, and anti-duplicate rules.
6. The GUI agent skill self-evolution method based on execution feedback as described in claim 5, characterized in that, S22 also provides a read-only accessibility tree observation tool, which returns the accessibility tree of the current GUI interface, enabling the executor to read and revise skill files within a preset range, while avoiding arbitrary modification of external resources of the skill pack.
7. The GUI agent skill self-evolution method based on execution feedback as described in claim 5, characterized in that, The specific process of S3 is as follows: S31. A GUI agent acts as a critic and diagnoses the cause of failure of the failed interaction trajectory based on the GUI task instructions, observation sequence, and action sequence in the i-th round of failed interaction trajectory, obtaining a structured diagnostic result. The failed interaction trajectory also includes the failure location and visible evidence representing the failure of the target GUI, the observation sequence. The action sequence consists of observations of each interaction step in the i-th round. It consists of GUI actions in each interaction step of the i-th round; S32. Based on the execution trajectory of round i, the skill pack revised locally in round i, and the structured diagnostic results fed back by the critic, the executor generates the revised skill pack for round i+1, and determines whether the failed interaction trajectory can be used as a failure case: if it can, it considers the wrong action, wrong answer, or wrong location to have a risk of recurrence, and creates a failure case file through create_failure in subsequent interaction steps, while writing the GUI task instruction, structured diagnostic results, and failed interaction trajectory into the failure case set; if it cannot, it does not create a failure case file through create_failure, nor does it write it into the failure case set.
8. The GUI agent skill self-evolution method based on execution feedback as described in claim 7, characterized in that, In S31, the information set of the constraint critic To ensure that the critic can only read information from the information set during diagnosis, the following information isolation conditions must be met: ; ; ; in, A set of GUI task instructions; For skill packs; This represents the inference chain within the actuator; This is the true answer; It is the set of all information that the executor can access during runtime.
9. The GUI agent skill self-evolution method based on execution feedback as described in claim 7, characterized in that, In S31, the executor, based on the initial skill pack of the i-th round of interaction, completes a single round of interaction in the target execution environment, from start to successful execution, or from start to execution failure, or from start to the maximum number of execution steps, and obtains the execution trajectory of the i-th round.
10. A GUI intelligent agent skill self-evolution system based on execution feedback, characterized in that, include: The skill package confirmation module is used to perform metadata retrieval and reuse judgment on the existing skill library based on the GUI task instructions input by the user. When the skill package in the existing skill library matches the target GUI task contained in the GUI task instructions, the skill package is reused; otherwise, a new skill package is created for the GUI agent to complete the target GUI task. The local revision module is used in the t-th interaction step of the i-th round of interaction in the execution of the target GUI task. The GUI agent acts as the executor, reads the GUI task instructions, and generates the t-th step GUI action based on the historical interaction trajectory and the skill pack obtained by S1, which is then applied to the target execution environment. If the current GUI interface state is found to be inconsistent with the skill pack, then an immediate partial revision is performed on the skill pack; The trajectory-level revision module is used to initiate a failure reflection process with information isolation when the i-th round of execution fails or the upper limit of interaction steps is reached. The critic in an independent session diagnoses the cause of failure of the failed interaction trajectory, and the executor completes the trajectory-level revision of the skill pack based on the feedback of the critic. The revised skill pack is then registered or updated to the existing skill library. The critic and the executor share the same base model parameters. The management module is used to mark the skill pack as verified and update the skill pack's retrieval metadata if the target GUI task is successfully executed after multiple rounds of interaction. If the target GUI task fails but the skill pack has been revised, the skill pack will be marked as pending review for user confirmation.