Graphical interface agent training method and device and storage medium
By introducing dynamic task progress prediction and fine-grained reward allocation in the training of graphical interface agents, the reward sparsity problem is solved, and the training efficiency and task completion rate are improved.
Patent Information
- Application Number
- CN202510146642.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-23
AI Technical Summary
Graphical interface agents face reward sparsity problems in reinforcement learning, resulting in inefficiency in training and learning difficulties.
Dynamic task progress prediction and fine-grained reward allocation mechanism are introduced. By calculating the task prediction progress of each operation step in real time, sparse rewards are broken down into continuous distributed progress reward values, providing immediate feedback and dense reward signals.
The frequency of policy gradient update is improved, and a training paradigm for co-evolution of task understanding and operation execution is formed, which significantly improves the training efficiency and task completion rate of graphical interface agents.
Smart Images

Figure CN120031099A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method, device, storage medium and program product for a graphical interface intelligent agent. Background Art
[0002] Graphical User Interface (GUI) agents have become an important component of automating the interaction with complex systems, especially in task-oriented environments such as web platforms. Typically, these agents are trained using reinforcement learning (RL) techniques, where the agent learns to perform tasks by interacting with the environment and receiving feedback in the form of rewards. However, there are still several challenges in the training of GUI agents, especially in the scarcity of paired instruction and evaluation data and the sparsity of rewards in reinforcement learning.
[0003] The industry has not yet proposed a better solution to the above problems. Summary of the invention
[0004] The present application provides a graphical interface intelligent agent training method, device, storage medium and program product, which are used to at least solve the problem that the reward sparsity of reinforcement learning in the current related technology cannot meet the training needs of GUI intelligent agents.
[0005] In a first aspect, an embodiment of the present application provides a method for training a graphical interface intelligent agent, comprising: obtaining a graphical interface interaction task sample, wherein the graphical interface interaction task sample comprises a graphical interface interaction task trajectory and a corresponding interaction task completion result; for each graphical interface operation step in the graphical interface interaction task trajectory, determining a predicted task progress corresponding to the graphical interface operation step, and calculating a corresponding progress reward value based on the predicted task progress and the interaction task completion result; and training the graphical interface intelligent agent based on the progress reward value corresponding to each of the graphical interface operation steps.
[0006] In a second aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the training method of a graphical interface intelligent agent of any embodiment of the present application.
[0007] In a third aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the training method of a graphical interface intelligent agent of any embodiment of the present application are implemented.
[0008] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the method for training a graphical interface agent of any embodiment of the present application.
[0009] The beneficial effects of the embodiments of the present application are: By introducing dynamic task progress prediction and fine-grained reward allocation mechanism, the task prediction progress of each operation step is calculated in real time, and the sparse rewards that originally only appear at the end of the task are disassembled into continuously distributed progress reward values, so that the GUI agent can obtain immediate feedback when completing key intermediate steps, forming a dense reward signal gradient. Therefore, compared with the delayed reward mechanism of traditional RL, the high-level task goals are decomposed into quantifiable progress indicators through trajectory sequence modeling, and then a fine-grained reward function is constructed to drive policy optimization, which can effectively improve the frequency of policy gradient updates and form a training paradigm for the co-evolution of task understanding and operation execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0011] Figure 1 A flowchart showing an example of a method for training a graphical interface agent according to an embodiment of the present application is shown; Figure 2 A schematic diagram showing a comparison of the operating principles of an example of ProgRM and ORM provided according to an embodiment of the present application is shown; Figure 3 A schematic diagram of an example workflow of ProgRM provided according to an embodiment of the present application is shown; Figure 4 It is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION
[0012] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0013] It should be noted that a major challenge in training GUI agents is the scarcity of reward data. Due to the complexity of the task environment controlled by UI, existing research usually relies on manual evaluation to determine whether the task is successfully completed and assign rewards accordingly. This often leads to a limited number of task trajectories with available reward annotations, which seriously restricts the learning process of the agent. Although task-oriented datasets such as WikiHow provide instruction-action pairs, they usually lack the fine-grained evaluation criteria required for effective reward modeling.
[0014] Another key challenge is the sparsity of rewards in traditional reinforcement learning settings. In many tasks, rewards are only provided when the task is successfully completed, leaving the agent with little feedback during its performance. The sparsity of rewards significantly hinders the learning process because the agent has little guidance on how to improve its actions during task performance.
[0015] In current related technologies, reward models provide task feedback in RL training to help agents adjust their behavior based on feedback from the environment. However, in traditional RL models, rewards are given only when the task is completed. This sparse reward setting limits the ability of agents to gradually adjust their behavior during task execution.
[0016] The ORM (Outcome Reward Models) method usually provides a discrete reward scalar after the task is completed, but ORM lacks fine-grained evaluation of each step in the task process, resulting in low learning efficiency. Although ORM simplifies reward evaluation, it only provides rewards when the task is completed, so there is a lack of feedback for each step. It is difficult for the agent to learn efficiently from the intermediate steps, resulting in low training efficiency.
[0017] The sparse rewards of the task limit the effective learning of the agent for each task step, especially in complex long-sequence tasks, where fewer reward signals make it difficult for the agent to accurately predict and optimize each step of behavior. In addition, existing reward models usually cannot provide continuous feedback, and the sparse reward signals make it difficult to learn and optimize the impact of each step in the task, thereby reducing the learning efficiency of the agent, so that the agent can only adjust its behavior through fewer discrete reward signals, which performs poorly in complex environments.
[0018] It should be noted that in current related technologies, improved reward distribution methods are usually adopted, such as using more efficient gradient optimization algorithms to accelerate learning. In addition, by introducing more efficient reward signals, such as through exploration strategies in reinforcement learning algorithms (such as imitation learning, inverse reinforcement learning, etc.), the impact of sparse rewards can be reduced. In addition, through hardware and computing optimization, increasing computing resources can improve training efficiency. However, most of the methods in the above-mentioned related technologies only rely on rewards after the task is completed, and lack fine control over the intermediate progress.
[0019] It should be understood that the purpose of the above description of the current related art is only to facilitate the public to better understand the inventive spirit and motivation of the present application, and is not to be regarded as a limitation of the present application. In addition, the technical solutions described in the above-mentioned current related art are not prior art, and may also be undisclosed technical solutions, such as solutions under research or in the laboratory stage.
[0020] In view of the deficiencies in the above-mentioned current related technologies, a progress reward model (ProgRM) is proposed in the embodiments of the present application, which provides a continuous reward signal through the progress prediction of each step, solves the sparse reward problem, can effectively improve the reward density, and makes the model training more efficient.
[0021] Figure 1 An operational flow chart showing an example of a method for training a graphical interface agent according to an embodiment of the present application is shown.
[0022] like Figure 1 As shown, in step S110, a graphical interface interactive task sample is obtained, and the graphical interface interactive task sample includes a graphical interface interactive task trajectory and a corresponding interactive task completion result.
[0023] In some embodiments, the operation logs of actual users in software or mobile applications are used to collect the various operation steps of users when completing specific tasks and the final task completion status (success, failure, exception, etc.), so as to obtain corresponding graphical interface interaction task samples. Furthermore, a large number of graphical interface interaction task samples can be automatically generated using simulators or test platforms to cover common operations, edge scenarios, and erroneous operations. By collecting diverse samples from real users and simulated environments, it is ensured that the training data covers a wide range of task scenarios, user behaviors, and abnormal situations.
[0024] It should be understood that in addition to the operation track, it is also possible to collect information in combination with various interactive behaviors such as mouse click position, scrolling, dragging, keyboard input, and the status information of interface elements (such as whether a button is available, whether a window pops up). In addition, the task types of graphical interface interactive tasks can be diverse, such as logging into an application, implementing specific functional instructions, etc., which should not be limited here.
[0025] In step S120, for each graphical interface operation step in the graphical interface interactive task trajectory, the task prediction progress corresponding to the graphical interface operation step is determined, and the corresponding progress reward value is calculated according to the task prediction progress and the interactive task completion result.
[0026] In some embodiments, for each graphical interface operation step, key features are extracted from the current interface state, such as the state of interface elements, the distribution of operable areas, and the state changes before and after the user operation. Then, the progress reward model (ProgRM) is used to model the interface state and operation sequence to predict the progress value of the step in the overall task. The progress reward model can adopt various non-restrictive deep learning network architectures, such as convolutional neural networks, recurrent neural networks, or Transformer models.
[0027] Regarding the definition of progress value, it can be to combine the contextual information of the previous and next operations, make a comprehensive prediction of the progress of the current step, and map the task progress to a standardized value (for example, between 0 and 1), where 0 means that the task has not started or there is no progress, and 1 means that the task has been completed. In addition, the interactive task completion result can be the successful completion of the task or the failure of the task. If the task is successfully completed, it can be more inclined to positive rewards, while if the task fails to be completed, it can be more inclined to negative rewards or penalties.
[0028] For example, if the current step brings the task state significantly closer to the expected final state (for example, the progress value increases significantly), a higher positive reward is given. If the operation fails to promote the progress of the task or causes the task state to regress, a negative reward or deduction is given.
[0029] In step S130, the graphical interface agent is trained according to the progress reward value corresponding to each graphical interface operation step.
[0030] Here, various reinforcement learning methods can be used to train the graphical interface agent, such as deep Q network (DQN), policy gradient, Actor-Critic, etc. Exemplarily, each sample trajectory is converted into a data sequence structure containing multiple states-actions-rewards, and the graphical interface agent is trained using the collected task samples.
[0031] Through the embodiments of the present application, by utilizing the fine-grained feedback of the progress reward value, the intelligent agent can gradually learn to choose appropriate operations in a complex interface environment. The task prediction progress and progress reward mechanism provide the intelligent agent with more detailed operation guidance, which helps to quickly focus on the effective operation path and improve the completion rate of interactive tasks.
[0032] Regarding the specific details of determining the predicted progress of the task, in some embodiments, the environmental state observation value corresponding to the graphical interface operation step is obtained, and the corresponding task prediction progress is calculated based on the environmental state observation value and the graphical interface operation step. Specifically, before and after each graphical interface operation step, the system automatically captures the current interface state. For example, for an interface based on a web page or mobile application, the attributes of each element (such as element type, position, size, state, etc.) can be obtained by parsing the DOM tree or UI tree. For unstructured graphical interfaces, a visual description of the environmental state is obtained by taking screenshots and extracting image features using computer vision technology (such as convolutional neural networks). Therefore, by obtaining the interface state in a fine-grained manner, it is beneficial to make an accurate assessment of the progress contribution of each operation step and improve the accuracy of the prediction results of the task implementation progress.
[0033] Regarding the specific calculation details of the progress reward value, when the interactive task completion result indicates that the graphical interface interactive task trajectory is a successful task completion trajectory, the progress reward value corresponding to the current graphical interface operation step is calculated based on the progress difference between the task predicted progress of the current graphical interface operation step and the task predicted progress of the previous graphical interface operation step.
[0034] For example, the progress value may be standardized within a range of 0 to 1, and then the order information of each operation step in the interaction trajectory is attached to ensure that the previous and next operations can be distinguished during the progress calculation process.
[0035] In some implementations, for two consecutive operation steps in the same task trajectory, the predicted progress of the previous graphical interface operation step is recorded as , the predicted progress of the current graphical interface operation step is Then, the progress difference is calculated as ,pass It can reflect the contribution of the current operation to the overall progress of the task, calculate the progress difference between consecutive operation steps, accurately quantify the contribution of each operation to the completion of the task, and avoid the problem of sparse rewards caused by the overall reward. For the trajectory of successful task completion, it can form a phased intensive reward signal to guide the agent to identify the impact of different operation steps on the progress of the task.
[0036] In some examples of the embodiments of the present application, each graphical interface operation step has a corresponding progress reward label, and the progress reward label is determined according to the track step sequence number of the graphical interface operation step in the corresponding graphical interface interaction task track.
[0037] In some embodiments, for each task execution, a complete interactive track can be constructed in chronological order, each track contains multiple discrete graphical interface operation steps, and records the corresponding timestamp, operation type, interface status and other information. For each operation step in the same interactive task track, a continuous serial number (for example: step 1, step 2, ..., step N) is automatically assigned according to the order in which it occurs. During the preprocessing process, the serial number of each operation step is used as a basic identifier and stored together with other state information, operation type and other data to form a structured data record. In some examples, the serial number of the operation step in the task track can be used to characterize the progress of the task. For example, a low serial number indicates that the task has just started and the reward value is low; in addition, as the serial number increases, the task stage gradually advances, and the reward value can also increase accordingly. Therefore, a progress reward label is automatically generated according to the serial number of the operation step, and phased feedback is provided for each operation step, which can guide the intelligent agent to better learn and grasp the progress of the task.
[0038] Regarding the details of the agent training, in one example of the embodiment of the present application, it can be to use a general type of loss function and optimize the training with the goal of minimizing the gap between the predicted reward and the actual reward value. However, the above method has better performance in the case of predicting discrete reward categories, but has little effect in the case of continuous rewards.
[0039] In another example of the embodiment of the present application, the rewards of the graphical interface operation steps corresponding to the same trajectory step number in the first graphical interface interaction task trajectory and the second graphical interface interaction task trajectory are compared in pairs to model the training loss value of the corresponding graphical interface operation step. Then, the graphical interface agent is trained according to the training loss value.
[0040] Specifically, for the operation steps corresponding to the same trajectory step number from two different graphical interface interaction task trajectories, their true reward values often reflect the relative contribution of the operation to the progress of task completion. By comparing these reward values, a training loss reflecting the ranking relationship is constructed. In this embodiment, it is not directly required to predict the absolute value, but the model is required to correctly reflect the relative high and low rewards between the two operation steps. In this way, for continuous reward values, the model only needs to capture the relative relationship of "which step is better" to obtain a more stable training signal.
[0041] It should be noted that training data from multiple task trajectories together can make full use of operation data from different scenarios and users, and improve the generalization and robustness of the model. In some examples, by combining trajectory pairs of positive and negative trajectories, it can ensure that the model can distinguish between successful and failed task completion results. In particular, by comparing the sequence numbers of the same trajectory steps in pairs, the agent obtains more fine-grained feedback signals, which helps to quickly identify more effective operation strategies at key steps.
[0042] In some examples of the embodiments of the present application, the loss function used to train the graphical interface agent is a cross entropy loss function or a mean square error loss function. Through the cross entropy loss function, different rewards or operation states can be effectively distinguished; through the mean square error for continuous reward value regression, the difference in reward values can be accurately captured. Therefore, through the selection of the above loss function, a stable and obvious gradient signal can be provided, ensuring that the model converges quickly in complex graphical interface tasks and effectively reducing error accumulation.
[0043] In the embodiment of the present application, the problem of sparse rewards in traditional RL is solved by providing a continuous, progress-based reward signal through the progress reward model (ProgRM). Unlike ORM, ProgRM provides a progress-based reward value for each step, and the reward value is calculated based on the progress difference between the current state of the task and the goal. In this way, the agent can receive more frequent feedback signals during the task execution, thereby accelerating the learning process and improving the success rate of task completion.
[0044] Figure 2 A schematic diagram comparing the operating principles of an example of ProgRM and ORM provided according to an embodiment of the present application is shown.
[0045] like Figure 2 As shown in the figure, the ProgRM model provides denser and more informative feedback during task execution. ProgRM predicts the agent's progress in completing the task at each step, providing a continuous reward signal to guide the agent throughout the process. Unlike traditional ORM, which usually provides a reward for each trajectory at the end of the task, ProgRM allows each step in the task trajectory to contribute to the reward, thereby improving learning efficiency and increasing the density of the reward. By evaluating the agent's progress relative to the task goal, ProgRM provides fine-grained feedback, effectively alleviating the problem of sparse rewards.
[0046] It should be noted that reward models (RMs) are crucial for training RL agents, especially in complex tasks that require structured feedback. In mathematical reasoning tasks, two key reward models are involved: ORM verifies and strengthens each reasoning step through step-by-step verification; Progress Reward Model (PRM) provides cumulative rewards based on the agent's progress in the problem-solving process. These models not only guide the large language model (LLM) to improve the final answer, but also help improve the reasoning process behind it, ensuring more reliable and effective problem solving.
[0047] The concept of reward models is extended by introducing autonomous evaluation, where the agent uses intrinsic feedback to improve itself, reducing the reliance on external supervision. This framework enables the agent to independently adjust its behavior, which is particularly useful in environments where rewards are sparse. In adaptive reward models, the agent evolves as it learns in a dynamic web environment. Curriculum-based reinforcement learning enables the reward model to adapt to the increase in task complexity, helping the agent to effectively cope with complex environments. Therefore, the increasing importance of reward models in enabling LLMs to handle complex decision-making tasks, especially in environments where rewards are sparse or there is no direct reward.
[0048] A note on progress rewards: Progress reward models are designed to guide agents to learn in environments where rewards are sparse or where there are long-term dependencies. The ELE (Explore Like Experts) method uses expert demonstrations to train a progress model, and the agent learns to optimize its exploration behavior based on expert trajectories. The model predicts the temporal distance between observations, providing meaningful progress signals even in the absence of explicit rewards.
[0049] In ELE, the progress model acts as an auxiliary reward to help the agent explore more efficiently in complex environments, such as the game NetHack, where traditional reward signals are sparse. This approach contrasts with traditional reinforcement learning methods, where progress is often difficult to track without explicit feedback.
[0050] In summary, progress reward models like ELE provide a valuable mechanism for training agents in environments with sparse rewards, helping them make more informed decisions and improving their exploration capabilities.
[0051] I. Methods 1.1 Progress Reward Model The Progress Reward Model (ProgRM) provides continuous feedback to the agent during task execution by predicting the agent's progress in completing the task at each step. and corresponding actions , the model outputs the predicted progress value of this step The reward for each action is obtained by calculating the progress difference between consecutive steps. Specifically, for the trajectory ,in Represents the time step The progress reward is calculated as: , Formula (1) in, It is The predicted progress value of the step, is the predicted progress of the next step. For failed trajectories, all rewards are set to 0, indicating no progress: , Formula (2) Therefore, the progress reward model learns to predict the agent’s progress towards task completion and calculates rewards based on the difference in progress between consecutive steps.
[0052] Figure 3 A schematic diagram of an example workflow of ProgRM provided according to an embodiment of the present application is shown.
[0053] like Figure 3 As shown, for a complex task trajectory, the agent predicts an action a t Afterwards, the environment gives the observations after execution o t , ProgRM gives the reward value based on actions and observations r t As the progress score of the current task, the agent decides the next action based on the feedback observation and reward value a t+1 .
[0054] 1.2 Model Training A modified version of the LLaMA model is used to train the progress reward model. The basic architecture follows the LLaMA structure, but the last layer is modified to output a scalar value representing the reward of the current state.
[0055] make represents the loss function of the training model, where represents the model parameters. The training process involves minimizing the prediction reward With real rewards There are two possible loss functions used for this task: 1.2.1 Loss Function 1: Cross Entropy Loss In this case, we aim to minimize the cross entropy between the predicted reward and the true reward value. The loss function is defined as: , Formula (3) in, is the reward for prediction, is the real reward. This formula is suitable for predicting discrete reward categories, but in this technical scenario, the reward is continuous.
[0056] 1.2.2 Loss Function 2: Pairwise Comparison Loss Another approach is to model the reward as a pairwise comparison between different trajectories. Here, the difference between the predicted rewards for two trajectories can be calculated: , Formula (4) in, It can be cross entropy or mean square error (MSE), and They are the first This formulation is particularly effective for training positive and negative trajectories, ensuring that the model can distinguish between successful and failed task completions.
[0057] 1.3 Data and label creation The training data of the progress reward model consists of task trajectories and corresponding reward pairs. The reward is distributed based on the success or failure of the task.
[0058] For a successful trajectory, the reward of the last step is assigned as 1, and the rewards of the intermediate steps are uniformly distributed between 0 and 1. Specifically, for a trajectory containing T steps, The reward for the step is calculated as: , Formula (5) For failed trajectories, all rewards are set to 0, indicating that the task was not completed: , Formula (6) We generate training labels by collecting a large number of task trajectories from the WikiHow dataset, which are preprocessed to include both successful and failed task executions. In addition, we augment the data by generating pseudo trajectories, especially for negative cases, using techniques such as pairing mismatched instructions and actions, or applying random strategies to simulate failure scenarios.
[0059] The final dataset consists of a combination of real trajectory data and pseudo trajectory data, ensuring the robustness of the reward model training process.
[0060] Through the embodiments of the present application, the training overhead of large language models can be significantly reduced, solving the sparse reward problem in reinforcement learning. Furthermore, the solution can break through the bottleneck of the high overhead of large language model training, provide strong support for more intelligent artificial intelligence systems in the future, and promote the widespread implementation of AI applications.
[0061] II. Experiment 2.1 Experimental Setup 2.1.1 Environment Experiments were conducted in two different environments, each designed to test the performance of the task completion agent in different scenarios.
[0062] WebArena: A benchmark environment for testing agents that perform tasks that involve interacting with web pages, filling out forms, or navigating websites to complete specified goals.
[0063] WikiHow: A set of instructional tasks that require an agent to follow a sequence of steps to complete various tasks, such as “how to make pasta” or “how to fix a bicycle.” This environment is designed to evaluate an agent’s ability to follow and execute multi-step processes.
[0064] 2.1.2 Basic Model For the base model, we used LLaMA 3.1 models of different sizes, namely 8B and 70B. This enabled us to explore the performance of the model under different computing powers and understand how model scaling affects the performance of reward prediction.
[0065] 2.1.3 Baseline Model ProgRM is evaluated against several baseline methods: • Simple Outcome Reward Model (ORM): This baseline model provides sparse rewards only at the end of the task, that is, the reward is set to 1 when the task is successfully completed and 0 otherwise. This is a traditional reward model commonly used in reinforcement learning.
[0066] 2.2 Evaluation as an Evaluator In this experimental setting, the performance of ProgRM is evaluated by comparing the rewards of task trajectories predicted by ProgRM with a standard evaluation (ideal evaluation). The specific goal is to evaluate the consistency between the predicted reward scores and the manually assigned scores, as this directly reflects the quality of the reward model.
[0067] The dataset is divided into a training set and a test set. The model is trained on the training set and then evaluated on the test set using metrics such as mean squared error (MSE) and Pearson correlation coefficient.
[0068] Evaluation Metrics The following evaluation metrics are used to evaluate the performance of the reward model: • Mean Squared Error (MSE): Measures the average squared difference between the predicted reward and the ideal reward. Lower MSE indicates better performance.
[0069] • Pearson Correlation Coefficient (PCC): Measures the linear correlation between the predicted reward and the ideal reward. The higher the PCC, the better the consistency between the model’s predictions and human evaluations.
[0070] Table 1. Evaluation of reward model as evaluator The results in Table 1 show that ProgRM consistently outperforms the baseline models, achieving the lowest MSE and highest PCC, indicating that the predicted rewards are highly consistent with the manually annotated evaluation results.
[0071] 2.3 Evaluation as a Reward Model In this setting, ProgRM is used as a reward model to train a reinforcement learning agent. The model predicts rewards for each step of the task trajectory, and these rewards are then used to guide the agent's learning process. The key question here is whether providing process-based rewards (including intermediate rewards at each step) can improve task performance more effectively than the standard sparse reward setting (i.e., providing rewards only at the end of the task).
[0072] Training settings In this experiment, the agent is trained using two different reward settings: • Sparse rewards: rewards are provided only at the last step of a task. If the task is completed successfully, the reward is set to 1; otherwise, it is set to 0.
[0073] • Process Reward: Provides rewards at each step of the task, and the reward at each step is based on the progress predicted by ProgRM. Specifically, the reward is calculated as the difference in the predicted progress from the previous step to the current step.
[0074] The agents are evaluated on the WikiHow test set, measuring their performance by task completion rate and the number of steps required to complete a task.
[0075] Table 2. Comparison of sparse reward and process reward task performance As shown in Table 2, the agent trained with the process reward model achieved 90% task completion rate, compared to 75% for the sparse reward model. In addition, the average number of steps required to complete the task is lower in the process reward model (9.3 steps vs. 12.5 steps), indicating that the process reward helps the agent learn more efficiently and effectively.
[0076] It should be noted that future work will focus on improving the robustness of progress estimation and exploring how to combine ProgRM with other reward learning techniques. In addition, extending ProgRM to more diverse and complex task environments will help evaluate its scalability and applicability in real-world scenarios, such as interactive web applications and multi-agent systems.
[0077] The progress reward model (ProgRM) provided in the embodiment of the present application provides continuous and step-by-step feedback during task execution, solving the problems of reward sparsity and limited training data for graphical user interface (GUI) agents in reinforcement learning (RL). Experimental results on WebArena and WikiHow show that ProgRM significantly outperforms the traditional output reward model (ORM) in task evaluation and agent training, improves task completion rate and reduces execution time. ProgRM makes learning more efficient, especially for larger models, by enhancing the density of rewards. Overall, ProgRM provides an effective solution to the problem of sparse rewards in RL and paves the way for more efficient GUI agent training.
[0078] In the embodiments of the present application, a solution to the reward sparsity problem in GUI agent training is provided based on a progress reward model (ProgRM); a novel task progress estimation method is proposed, and the reward model is trained using positive and negative task trajectories; the effectiveness of ProgRM as an evaluator is verified by evaluating the consistency of the automatic evaluation results of ProgRM with standard evaluation indicators; and the effectiveness of ProgRM as a reward model is verified by testing its ability to train reinforcement learning agents and improve task completion performance.
[0079] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of actions combined, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0080] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the above-mentioned training methods for a graphical interface intelligent agent in the present application.
[0081] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any of the above-mentioned training methods for a graphical interface intelligent agent.
[0082] In some embodiments, the embodiments of the present application also provide an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a training method for a graphical interface intelligent agent.
[0083] Figure 4 FIG. 1 is a schematic diagram of the hardware structure of an electronic device for executing a training method for a graphical interface agent provided by another embodiment of the present application. Figure 4 As shown, the device includes: One or more processors 410 and memory 420, Figure 4 A processor 410 is taken as an example.
[0084] The device for executing the method for training a graphical interface agent may further include: an input device 430 and an output device 440 .
[0085] The processor 410, the memory 420, the input device 430 and the output device 440 may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.
[0086] The memory 420 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the training method of the graphical interface agent in the embodiment of the present application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 420, that is, the training method of the graphical interface agent in the above method embodiment is implemented.
[0087] The memory 420 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 420 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 420 may optionally include a memory remotely arranged relative to the processor 410, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0088] The input device 430 may receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 440 may include a display device such as a display screen.
[0089] The one or more modules are stored in the memory 420, and when executed by the one or more processors 410, the training method of the graphical interface agent in any of the above method embodiments is executed.
[0090] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present application.
[0091] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication equipment: This type of equipment is characterized by having mobile communication functions and its main purpose is to provide voice and data communications. This type of terminal includes: smart phones, multimedia phones, functional phones, and low-end phones.
[0092] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have mobile Internet access features. These terminals include: PDA, MID and UMPC devices, etc.
[0093] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0094] (4) Other onboard electronic devices with data interaction functions, such as on-board devices installed in vehicles.
[0095] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for training a graphical interface agent, comprising: Acquire a graphical interface interactive task sample, wherein the graphical interface interactive task sample includes a graphical interface interactive task trajectory and a corresponding interactive task completion result; For each graphical interface operation step in the graphical interface interactive task track, determine the task prediction progress corresponding to the graphical interface operation step, and calculate the corresponding progress reward value according to the task prediction progress and the interactive task completion result; The graphical interface agent is trained according to the progress reward value corresponding to each of the graphical interface operation steps.
2. The method according to claim 1, wherein: The determining of the task prediction progress corresponding to the graphical interface operation step includes: Obtaining the environmental state observation value corresponding to the graphical interface operation step; According to the environmental state observation value and the graphical interface operation steps, the corresponding task prediction progress is calculated.
3. The method according to claim 2, wherein: The calculating of the corresponding progress reward value according to the task prediction progress and the interactive task completion result includes: When the interactive task completion result indicates that the graphical interface interactive task trajectory is a successful task completion trajectory, the progress reward value corresponding to the current graphical interface operation step is calculated based on the progress difference between the task predicted progress of the current graphical interface operation step and the task predicted progress of the previous graphical interface operation step.
4. The method according to claim 3, wherein: Each of the graphical interface operation steps has a corresponding progress reward label, and the progress reward label is determined according to the track step sequence number of the graphical interface operation step in the corresponding graphical interface interactive task track.
5. The method according to claim 1, wherein: The calculating of the corresponding progress reward value according to the task prediction progress and the interactive task completion result includes: When the interactive task completion result indicates that the graphical interface interactive task trajectory is a task failure completion trajectory, the progress reward value corresponding to each graphical interface operation step is set according to a preset value.
6. The method according to claim 1, wherein: The step of training the graphical interface agent according to the progress reward value corresponding to each of the graphical interface operation steps includes: Compare the rewards of the graphical interface operation steps corresponding to the same trajectory step number in the first graphical interface interaction task trajectory and the second graphical interface interaction task trajectory in pairs to model the training loss value of the corresponding graphical interface operation steps; The graphical interface agent is trained according to the training loss value.
7. The method according to claim 6, wherein: The loss function used to train the graphical interface agent is a cross entropy loss function or a mean square error loss function.
8. A storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Platform, method and system for collecting training data of graphical user interface proxy model
CN121051471A
Intelligent agent self-evolution training method and system based on codes and fused with graph extension
CN121279348A
Reinforcement learning training method, system construction method, system, equipment and product
CN121387436A
Display method, electronic device, readable storage medium and program product
CN121833099A
Intelligent agent training method and device and graphical user interface operation method and device
CN122389921A