Intelligent agent model training method, human-computer interaction program processing method, device, equipment, storage medium and program product
By employing a phased sample selection and training strategy, the problem of instability in the early stages of agent model training was solved, improving the model's generalization ability and robustness, and promoting the smooth convergence and optimization of model parameters.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-11
- Publication Date
- 2026-04-14
AI Technical Summary
Existing agent models suffer from policy instability in the early stages of training, leading to significant gradient fluctuations during model parameter updates and affecting their generalization ability in complex scenarios.
By employing a phased sample selection and training strategy, the first action sequence samples are generated and selected using the first prompt word to construct the target sample set. Then, when the training progress meets the conditions, the second action sequence samples are generated using the second prompt word to continue training, thereby promoting the smooth convergence and optimization of model parameters.
It improves the generalization ability and robustness of the intelligent agent model, and promotes the accuracy of model decision-making and the adaptability of scenarios.
Smart Images

Figure CN121860091A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method for an intelligent agent model, a processing method for human-computer interaction programs, an apparatus, a device, a storage medium, and a program product. Background Technology
[0002] With the development of artificial intelligence technology, agent models are widely used in automated interaction tasks. Among related technologies, reinforcement learning (RL) is used to train agent models, such as proximal policy optimization (PPO) and group relative policy optimization (GRPO) algorithms. However, because the model's policy is unstable in the early stages of training, the feedback signals generated by interactions often contain uncertainties, leading to significant gradient fluctuations during model parameter updates. This, in turn, affects the generalization ability of the finally trained agent model in complex scenarios. Summary of the Invention
[0003] This application provides a training method for an intelligent agent model, a processing method for a human-computer interaction program, an apparatus, a device, a storage medium, and a program product, which can improve the generalization ability and robustness of the intelligent agent model through a phased sample selection and training strategy.
[0004] The technical solution of this application embodiment is implemented as follows: This application provides a method for training an intelligent agent model, the method comprising: The agent model to be trained generates multiple first action sequence samples based on a preset first prompt word. From the plurality of first action sequence samples, first action sequence samples that meet preset sample selection conditions are selected to form a first target sample set; The agent model is trained based on the first target sample set; If the training progress of the trained agent model meets the training phase switching conditions, then the trained agent model generates multiple second action sequence samples based on a preset second prompt word to form a second target sample set. The agent model is trained based on the second target sample set.
[0005] This application provides a method for processing human-computer interaction programs, the method comprising: An action sequence is generated based on a third prompt word using a pre-trained agent model. The actions in the action sequence are used to operate a human-computer interaction program to complete a preset task. The agent model is trained using the agent model training method provided in the embodiments of this application.
[0006] This application provides a training device for an intelligent agent model, comprising: The first training module is used to generate multiple first action sequence samples based on a preset first prompt word using the agent model to be trained. The first training module is further configured to select first action sequence samples that meet preset sample selection conditions from the plurality of first action sequence samples to form a first target sample set; The first training module is further configured to train the agent model based on the first target sample set; The second training module is used to generate multiple second action sequence samples based on a preset second prompt word, to form a second target sample set, if the training progress of the trained agent model meets the training phase switching conditions. The agent model is trained based on the second target sample set.
[0007] This application provides a human-computer interaction program processing device, including: The data processing module is used to generate an action sequence based on a third prompt word using a pre-trained agent model. The actions in the action sequence are used to operate a human-computer interaction program to complete a preset task. The agent model is trained using the agent model training method provided in the embodiments of this application.
[0008] This application provides an electronic device, including: Memory is used to store executable instructions or computer programs. The processor is configured to execute computer-executable instructions or computer programs stored in the memory to implement the training method for the intelligent agent model provided in the embodiments of this application, or to implement the processing method for the human-computer interaction program provided in the embodiments of this application.
[0009] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implement the training method for the intelligent agent model provided in this application, or the processing method for the human-computer interaction program provided in this application.
[0010] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the training method for the intelligent agent model provided in this application, or the processing method for the human-computer interaction program provided in this application.
[0011] The embodiments of this application have the following beneficial effects: By generating and filtering action sequence samples based on the first cue word, a first target sample set that meets specific conditions is constructed for the agent model's initial training. This allows the selected samples to guide the model in establishing a stable basic strategy during the early stages of training. Furthermore, when the training progress meets the switching conditions, the agent model is trained again using a second target sample set generated based on the second cue word. This phased, progressive training mechanism enables the agent model to adaptively adjust the focus of its learning content according to its own training state, promoting the smooth convergence and optimization of model parameters, thereby improving the decision-making accuracy and scene generalization ability of the agent model. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the architecture of the training system for the intelligent agent model provided in the embodiments of this application; Figure 2A This is a first structural schematic diagram of the electronic device provided in an embodiment of this application; Figure 2B This is a schematic diagram of the second structure of the electronic device provided in the embodiments of this application; Figure 3A This is a schematic diagram of the first process of the training method for the intelligent agent model provided in the embodiments of this application; Figure 3B This is a schematic diagram of the second process of the training method for the intelligent agent model provided in the embodiments of this application; Figure 3C This is a schematic diagram of the third process of the training method for the intelligent agent model provided in the embodiments of this application; Figure 3D This is a schematic diagram of the fourth process of the training method for the intelligent agent model provided in the embodiments of this application; Figure 4A This is a schematic diagram of the first interface of the human-computer interaction program provided in the embodiments of this application; Figure 4B This is a schematic diagram of the second interface of the human-computer interaction program provided in the embodiments of this application; Figure 5 This is a schematic diagram illustrating the principle of the training method for the intelligent agent model provided in the embodiments of this application; Figure 6 This is a schematic diagram illustrating the calculation principle of the sample evaluation parameters provided in the embodiments of this application; Figure 7 This is a schematic diagram illustrating the principle of action prediction of the intelligent agent model provided in the embodiments of this application.
[0013] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0016] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0017] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0018] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0019] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0020] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0021] 1) An intelligent agent model, also known as an intelligent agent, refers to an artificial intelligence system that possesses the capabilities of environmental perception, goal understanding, dynamic decision-making, and interactive execution, and can autonomously advance to achieve specific tasks through multiple closed-loop processes. In the embodiments of this application, the core features of the intelligent agent model are "autonomy" and "interactivity." In human-computer interaction scenarios (such as ordering drinks), the intelligent agent can perceive the interface environment through image encoding (such as recognizing button positions, interface states, etc.), understand user instructions through text encoding (such as "order a medium drink with ice F," etc.), and then dynamically select appropriate actions (such as clicking, swiping, etc.) in combination with a preset set of actions, ultimately driving the service from initiation to completion.
[0022] 2) The interface of a human-computer interaction program refers to the interface used to provide human-computer interaction functions. Examples include Graphical User Interface (GUI), Augmented Reality (AR) interface, Virtual Reality (VR) interface, Voice User Interface (VUI), Interactive Projection Interface (using projection technology to display information on a flat surface), Eye-tracking Interface (an interface controlled by detecting the user's gaze), Holographic Interface (a three-dimensional hologram formed by projecting images using holographic projection technology, allowing viewing of stereoscopic images without special glasses), Multimodal Interface (an interface combining multiple interaction methods, such as tactile, visual, and auditory interaction), and Brain-Machine Interface (BMI) interface, etc.
[0023] 3) Large Language Models (LLMs), also known as large models, are large-scale language models designed to understand and generate human language. They are trained on massive amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. The defining characteristic of large language models is their sheer size, containing billions of parameters that help them learn complex patterns in language data. They are typically based on deep learning architectures. Large language models refer to deep learning models trained on vast amounts of text data, containing billions or even more parameters. They can be used to generate natural language text and understand its meaning. Through training, the models learn the statistical regularities and semantic relationships of language to build a vast language knowledge base, thereby simulating human language understanding and generation capabilities.
[0024] 4) Visual Language Model (VLM) is a multimodal large model that can simultaneously understand image and text information and make decisions.
[0025] 5) Group Relative Policy Optimization (GRPO) is a reinforcement learning method that aims to optimize the action policy of an agent model and improve the decision-making stability and convergence efficiency of the agent model in complex tasks.
[0026] With the development of artificial intelligence technology, agent models are widely used in automated interaction tasks. Among related technologies, reinforcement learning (RL) is used to train agent models, such as proximal policy optimization (PPO) and group relative policy optimization (GRPO) algorithms. However, because the model's policy is unstable in the early stages of training, the feedback signals generated by interactions often contain uncertainties, leading to significant gradient fluctuations during model parameter updates. This, in turn, affects the generalization ability of the finally trained agent model in complex scenarios.
[0027] This application provides a training method for an intelligent agent model, a processing method for a human-computer interaction program, an apparatus, a device, a computer-readable storage medium, and a computer program product. These methods can improve the generalization ability and robustness of the intelligent agent model through a phased sample selection and training strategy. The following describes exemplary applications of the electronic devices provided in this application. These electronic devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers.
[0028] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the training system for the intelligent agent model provided in this application embodiment. Figure 1 The system involves server 100, terminal device 200, and network 300. Terminal device 200 is connected to server 100 through network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.
[0029] In some embodiments, the training method for the intelligent agent model provided in this application can be implemented collaboratively by a server and a terminal device. For example, the terminal device 200 sends a first prompt word and a second prompt word to the server 100. The server 100 receives the first prompt word and the second prompt word, trains the intelligent agent model using the training method provided in this application, and sends the intelligent agent model to the terminal device 200. The terminal device 200 receives the intelligent agent model and implements the human-computer interaction program processing method provided in this application based on the intelligent agent model.
[0030] In other embodiments, the training method for the intelligent agent model provided in this application can be implemented independently by the terminal device. The terminal device 200 calls the first prompt word and the second prompt word in its local database, obtains the trained intelligent agent model through the training method for the intelligent agent model provided in this application, and implements the human-computer interaction program processing method provided in this application based on the trained intelligent agent model.
[0031] In some embodiments, the human-computer interaction program processing method provided in this application can be implemented by a terminal device or a server alone. The terminal device 200 or the server 100 invokes a local intelligent agent model and generates an action sequence based on a third prompt word using the human-computer interaction program processing method provided in this application. The actions in the action sequence are used to operate the human-computer interaction program to complete a preset task. The intelligent agent model can be an intelligent agent model obtained by the terminal device 200 using the intelligent agent model training method provided in this application, or it can be an intelligent agent model trained by the server 100 using the intelligent agent model training method provided in this application and sent to the terminal device 200.
[0032] Here, server 100 can be a single server. In this case, the training method for the intelligent agent model and the processing method for the human-computer interaction program provided in this application embodiment can be implemented by the same server. Server 100 can also be a cluster of servers. In the case where server 100 is a server cluster, the training method for the intelligent agent model and the processing method for the human-computer interaction program provided in this application embodiment can be implemented by different servers. This application embodiment does not impose any limitations on this.
[0033] The intelligent agent model provided in this application can be applied to various complex scenarios such as graphical interface interaction, content generation, and automated operation and maintenance. Examples are given below.
[0034] 1) In the product ordering scenario, for example, a user sends a product ordering instruction to the intelligent agent model via voice (e.g., "Order me a tomato beef brisket rice from XX brand") or text input on a terminal device (such as a mobile phone or computer); the intelligent agent model calls a human-computer interaction program (such as a mini-program of XX takeaway platform, etc.) and generates an action sequence through the human-computer interaction program processing method provided in the embodiments of this application, so that the intelligent agent model operates the human-computer interaction program to complete the ordering instruction and realize automated ordering.
[0035] For example, the agent model first outputs a predicted action based on the initial interface state (such as "click the search bar and enter 'XX brand'"). After this action is executed on the human-computer interaction interface, the application jumps to the search results page. The agent model then takes the state of this new interface (search results page) as input and outputs a predicted action (such as "identify and click the matching store entrance"). This process continues, with each action's generation relying on the real-time interface feedback after the previous action's execution. Through iterative execution, the agent model can dynamically adapt to changes such as page loading and pop-ups, until finally clicking the "Submit Order" button to complete the entire action sequence.
[0036] 2) Data retrieval scenario, for example, a user sends a data retrieval command to the intelligent agent model on a terminal device (such as an office computer or tablet) via voice (such as "help me check my flight boarding pass") or text input; the intelligent agent model calls the corresponding human-computer interaction program (such as the ticket service module of the XX travel application client, the mini-program of the XX travel platform, etc.), and generates an action sequence through the human-computer interaction program processing method provided in the embodiments of this application, so that the intelligent agent model operates the human-computer interaction program to complete the data retrieval command and realize automated flight boarding pass query.
[0037] For example, during execution, the intelligent agent model perceives all interactive nodes on the current interface (such as buttons and menus) in real time and determines which node is closest to the target information, "boarding pass." If the current interface is the homepage, the intelligent agent model executes the action of "clicking 'My Trip'"; if a login verification or advertising pop-up appears on a subsequent interface, the intelligent agent model inserts temporary actions such as "enter verification code" or "close pop-up" based on the current interface state. This non-fixed, real-time observation-based sequence of actions ensures that the model can accurately find the target data path in complex application hierarchies, ultimately obtaining the boarding pass information and presenting it to the user.
[0038] 3) Code-assisted generation scenario, for example, a user sends code development instructions to the intelligent agent model via text input on a terminal device (such as "Please write a script for data deduplication in Python and save it locally"); the intelligent agent model calls the Integrated Development Environment (IDE) or code editor application, and generates an action sequence through the human-computer interaction program processing method provided in the embodiments of this application, so that the intelligent agent model can operate the code editor to complete the code writing.
[0039] For example, the generation of action sequences can employ a global planning strategy. This is because the logic of code writing and file saving is highly deterministic and does not heavily rely on real-time visual feedback from the interface. Upon receiving instructions, the agent model directly plans a list of actions containing complete steps, such as "generating a complete Python deduplication algorithm code block," "calling the editor interface to write the code to the buffer," "executing the file saving command (Ctrl+S or calling the API)," and "naming the file as deduplicate.py." After receiving this action sequence, executors are scheduled in batches to complete the operations sequentially, without waiting for environmental feedback after each line of code is entered. This enables highly automated automation of the entire process from logic construction to file storage at an extremely high speed.
[0040] Taking the use of electronic devices (such as servers and terminal devices) for training intelligent agent models as an example, see [link to relevant documentation]. Figure 2A , Figure 2A This is a first structural schematic diagram of the electronic device provided in an embodiment of this application. Figure 2A The illustrated electronic device 100-1 includes at least one processor 110-1, a memory 130-1, and at least one network interface 120-1. The various components in the electronic device 100-1 are coupled together via a bus system 140-1. It is understood that the bus system 140-1 is used to implement communication between these components. In addition to a data bus, the bus system 140-1 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2A The general designated all buses as Bus System 140-1.
[0041] The processor 110-1 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0042] The memory 130-1 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 130-1 may optionally include one or more storage devices physically located away from the processor 110-1.
[0043] The memory 130-1 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 130-1 described in this application embodiment is intended to include any suitable type of memory.
[0044] In some embodiments, memory 130-1 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0045] Operating system 131-1 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 132-1 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 120-1, such as Bluetooth, WiFi, and Universal Serial Bus (USB). In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2A A training device 133 for an agent model stored in memory 130-1 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a first training module 1331 and a second training module 1332. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0046] Taking the processing of human-computer interaction programs by electronic devices (such as servers and terminal devices) as an example, see [link to relevant documentation]. Figure 2B , Figure 2B This is a schematic diagram of the second structure of the electronic device provided in the embodiments of this application. Figure 2B The illustrated electronic device 100-2 includes at least one processor 110-2, a memory 130-2, and at least one network interface 120-2. The various components in server 100-2 are coupled together via a bus system 140-2. It is understood that the bus system 140-2 is used to implement communication between these components. In addition to a data bus, the bus system 140-2 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2BAll buses are labeled as Bus System 140-2. For detailed explanations of Processor 110-2 and Memory 130-2, please refer to the above text; they will not be repeated here.
[0047] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2B A processing device 134 for a human-computer interaction program stored in memory 130-2 is shown. This program can be software in the form of programs and plug-ins, and includes the following software module: a data processing module 1341. The function of this module will be described below.
[0048] In some embodiments, the terminal device or server can implement the training method for the intelligent agent model and the processing method for the human-computer interaction program provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.
[0049] The following will describe the training method of the intelligent agent model provided in the embodiments of this application, with the server as the execution subject, by combining the exemplary application and implementation of the server provided in the embodiments of this application.
[0050] See Figure 3A , Figure 3A This is a schematic diagram of the first process of the training method for the intelligent agent model provided in the embodiments of this application, which will be combined with Figure 3A The steps shown are explained.
[0051] In step 101, multiple first action sequence samples are generated based on a preset first prompt word using the agent model to be trained.
[0052] In some embodiments, during the first training phase, multiple first action sequence samples are generated based on a preset first prompt word using the agent model to be trained.
[0053] Here, the agent model to be trained refers to a machine learning model built on a deep neural network, possessing multimodal perception and decision-making capabilities. It can receive data from one or more modalities (such as images, natural language text, etc.) as input, and after processing internal network parameters, output decision actions or text responses to change the state of the environment. In this embodiment, the agent model can be a visual language model (VLM) pre-trained with massive amounts of image and text data. It not only possesses the ability to understand the semantics of natural language instructions but also the ability to recognize and locate visual elements (such as icons, buttons, and text areas) in a graphical user interface (GUI).
[0054] Here, the first cue word refers to the initial set of input data given to the agent model to trigger it to perform a specific task. It is the contextual basis for the agent model to generate action sequences, which clarifies the target task that the agent model needs to complete and the environmental state at the start of the task, thereby limiting the search space for the agent model's subsequent action planning.
[0055] Here, environmental state refers to the aggregated description of the operational environment perceived by the intelligent agent model. For example, in human-computer interaction scenarios, environmental state can be represented by the interface presentation information and system context information of electronic devices (such as mobile phones and computers) at the moment of interaction. This can include not only visual screen pixel information (such as interface screenshots, reflecting the layout, color, shape, and other visual features of UI elements), but also structured descriptive information of the interface (such as the Document Object Model (DOM) tree, auxiliary function node information, etc.), used to accurately describe the attributes (such as position coordinates, size, control type, current text content, clickability, etc.) of each interactive control (such as buttons, input boxes, and sliders) in the interface. Environmental state is the objective basis for the intelligent agent model to determine which stage of the business process it is currently in, identify what available operations are available, and evaluate the effect of the previous action.
[0056] For example, in human-computer interaction scenarios (such as automated control of mobile terminal applications, office assistance of desktop software, etc.), the first prompt word can be a multimodal data combination including the interface image (Visual State) of the human-computer interaction program at the current moment and the natural language instruction (Language Instruction) describing the user's needs. For example, the first prompt word could consist of a screenshot of a mobile phone desktop and the text "Open the food delivery app and order a fried chicken".
[0057] Here, a human-computer interaction program refers to an application or functional module that can run on electronic devices (such as mobile phones and computers) and allows users to obtain specific services through interface operations. It has a visual interface and interactive elements (such as buttons, input boxes, drop-down menus, etc.). For example, a human-computer interaction program can be of various types such as mini-programs, apps, H5, and web applications. This application embodiment does not limit the types. Users can interact with the human-computer interaction program through clicks, inputs, and other operations to obtain services such as placing orders, inquiries, and handling business.
[0058] In some embodiments, the first prompt word includes a task instruction; multiple first action sequence samples are generated based on the preset first prompt word, which can be achieved by performing the following processes multiple times: performing semantic understanding on the first prompt word to obtain at least one subtask corresponding to the task indicated by the task instruction in the first prompt word; performing actions to process each subtask, and combining the actions corresponding to each subtask into a first action sequence sample.
[0059] Here, a subtask refers to a smaller, more specific, and independently executable individual operation step or stage goal that is broken down into the task instructed by the task instruction. For example, for the task of "purchasing a black XX brand mobile phone with 256G memory", subtasks may include "opening the homepage of the shopping program", "entering XX brand mobile phone in the search bar to search", "clicking to enter the details page of the target product from the product list", "selecting the color (black) and storage capacity (256G) specifications", "clicking to buy now and confirming the shipping address", and "submitting the order and completing the payment", etc.
[0060] In some embodiments, semantic understanding of the first prompt word to obtain at least one subtask corresponding to the task indicated by the task instruction in the first prompt word can be achieved in the following way: semantic understanding of the task instruction to obtain the task description text of the task; generating a thought chain corresponding to the task based on the task description text, wherein the thought chain includes multiple subtasks, and the multiple subtasks are ordered according to a logical progressive relationship.
[0061] Here, "thought chain" refers to a structured output that simulates the human thinking and problem-solving process. It presents the reasoning steps, decision-making processes, and action planning required from the initial intention to the final goal in a coherent and logical sequence. In the embodiments of this application, the thought chain is a detailed planning blueprint that decomposes a preset task into a series of ordered sub-tasks, guiding each subsequent action of the intelligent agent.
[0062] For example, semantic understanding of instructions to obtain the task description text of a preset task can be achieved in the following way: First, intent recognition is performed, that is, the agent analyzes the entire instruction to determine the user's target intent (e.g., "online shopping", "check the weather"); second, entity extraction, also known as slot filling, is performed, that is, the agent extracts the key information parameters necessary to complete the target intent from the instruction (for example, in the online shopping intent, core entities that clearly indicate shopping needs, such as "product name" (e.g., XX mobile phone), "specification attributes" (e.g., black, 256G memory), and "target platform" (e.g., XX mall), are extracted); finally, structured output is performed, integrating the identified target intent and the extracted entities into a structured data (e.g., a JSON object), which serves as the task description text of the preset task.
[0063] For example, suppose the preset task indicated in the instruction is "Help me buy a black XX brand mobile phone with 256GB of memory from XX Mall", the task description text can be generated in the following way: Intent recognition: By analyzing the core semantics of the instruction, the intelligent agent can determine that the user's target intent is "online shopping" (i.e., the core task of completing the purchase of goods). Entity extraction: Extract the key information necessary to complete "online shopping" from the instruction information. The platform entity is "XX Mall", the product name entity is "XX Brand Mobile Phone", the color attribute entity is "Black", and the specification attribute entity is "256G Memory" (these are the core parameters for achieving the intention of "precise shopping"). Structured output: The identified "online shopping" intent is integrated with the extracted entity information into structured data. The final task description text (in JSON format for example) is: {"intent": "online shopping", "shopping platform": "XX Mall", "product name": "XX brand mobile phone", "color": "black", "storage specification": "256G"}.
[0064] For example, generating a thought chain corresponding to a pre-defined task based on a task description text can be achieved in the following way: First, task decomposition: After receiving the structured task description text, the intelligent agent model, based on its pre-learned knowledge of world operations and processes from massive amounts of data, breaks down the high-level task into multiple logically independent sub-tasks (for example, decomposing "online shopping" into "opening the shopping app," "searching for specific products," "selecting colors and memory specifications," and "clicking to buy now"). Second, step sequencing and dependency analysis: The intelligent agent model not only generates sub-tasks, but more importantly, it sequences them according to the logical and operational dependencies between them, ensuring that the completion of the previous sub-task is a prerequisite for the execution of the next sub-task. Finally, action planning: The model organizes the sequenced sub-tasks into a complete action plan, which details every step from beginning to end. The complete presentation of this plan is what is known as the thought chain.
[0065] Following the example above, a thought chain can be represented as follows: Subtask 1: Open the specified XX Mall APP (as the entry point for task execution, and the basis for all subsequent interactive operations); Subtask 2: Enter "XX brand mobile phone" in the "search bar" at the top of the APP and click the search button (fill in the core entity information to provide the necessary parameters for locating the product); Subtask 3: In the search results list, click on the product image or title with the highest matching degree (to initially filter the target and enter the product details page); Subtask 4: On the product details page, click the "Select Specifications" button, and in the pop-up window, click to select the "Black" color block and the "256G" capacity button (matching the attribute requirements in the task description); Subtask 5: After confirming that the specifications are correct, click the "Buy Now" button (confirm the specific product and proceed to the subsequent checkout process). Subtask 6: Proceed to the order confirmation page and verify the default shipping address and contact information (completing the logistics requirements for shopping and preparing for order submission); Subtask 7: After confirming that the amount is correct, click "Submit Order".
[0066] The above thought process clearly presents the sequence of subtasks from "launching the application" to "submitting the order": Task breakdown: The "purchase mobile phone" process is broken down into 7 logically independent sub-tasks (covering the complete process of "searching for products → browsing details → selecting specifications → confirming receipt → completing payment"). Step ordering: Strictly follow the operational dependency relationship of "first search and locate → then determine specifications → then confirm order → final closed loop" (for example, you must enter the details page before you can select color / memory; you must select the specifications before you can generate a valid order). Action Plan: Each subtask corresponds to a specific interactive operation (such as "clicking the color block" or "entering keywords"), forming an "operation guide" that can be executed directly, perfectly matching the core requirements in the task description text (XX brand, black, 256G).
[0067] By semantically understanding task instructions (including intent recognition, entity extraction, and structured output) to obtain the task description text, and then generating a corresponding thought chain based on the task description text, the task can be decomposed into multiple subtasks ordered in a logically progressive manner. These subtasks strictly follow the operation dependency relationship, and each subtask corresponds to a specific interactive operation, clearly presenting the entire process sequence from task initiation to completion. This not only fully matches the core requirements in the task description text (such as product name, color, and specifications in a shopping task), but also forms a directly executable "operation guide," ensuring that the intelligent agent proceeds step by step in a logical order and completes the task instructed by the task instructions.
[0068] In some embodiments, the intelligent agent model is used to interact with the human-computer interaction program; the actions for processing each subtask can be implemented in the following way: for each subtask, the following processing is iteratively performed until a preset condition is met: for the t-th time step, based on the interface image of the human-computer interaction program and the subtask at the t-th time step, the input data for the t-th time step is constructed; based on the input data at the t-th time step, action prediction is performed to obtain the action at the t-th time step; and the action at the t-th time step is executed to obtain the interface image of the human-computer interaction program at the (t+1)-th time step, where t≥0.
[0069] Here, the interface image of a human-computer interaction program refers to a snapshot of the application's visible area obtained by the agent at the current time step through screen capture technology, application programming interface (API) calls, or document object model (DOM) tree parsing.
[0070] In some embodiments, the interface image includes not only pixel-level RGB image data, but also structured information superimposed to enhance understanding, such as the bounding box coordinates of UI controls or functional description labels of controls. This application does not limit the specific format of the interface image; it can be a bitmap in PNG or JPG format, or preprocessed tensor data.
[0071] For example, based on the interface image and subtask of the human-computer interaction program at time step t, the input data for time step t can be constructed in the following way: combine the interface image (as a representation of the environment state), subtask, and task instructions at time step t as the input data for time step t.
[0072] For example, let's assume the interface image is... The subtask is described as follows: The task instructions are First, Adjust to the preset size (e.g.) And then converted into a sequence of visual feature vectors by a visual encoder. ;Will and Tokenization is performed, and the tokenized text is converted into a sequence of text feature vectors through a text embedding layer. Next, the visual feature vector sequence With text feature vector sequence The sequences are concatenated along the length dimension to form a joint input embedding. , dimension This joint input embedding That is, the input data used at the t-th time step.
[0073] In some embodiments, the preset conditions include any one of the following: the action at the t-th time step is the end action, wherein the end action is used to indicate that the subtask has been completed; the t-th time step is a preset maximum number of time steps; the descriptive text of the interface image of the human-computer interaction program at the (t+1)-th time step indicates that the subtask has been completed.
[0074] For example, the preset conditions include any one or a combination of the following to control the termination of the iteration loop: First, the action at time step t is identified as the end-of-episode action. The end-of-episode action represents the agent's perception that the subtask has been completed. For example, the action set includes a special stopping action. This condition is satisfied when the agent model outputs the stopping action with the highest probability.
[0075] Second, the t-th time step reaches the preset maximum number of time steps (e.g., (to prevent getting stuck in an infinite loop).
[0076] Third, the subtask of describing the interface image of the human-computer interaction program at time step t+1 has been completed.
[0077] For example, determining whether the descriptive text of the interface image of the human-computer interaction program at time step t+1 indicates that the subtask has been completed can be achieved by introducing a visual language understanding mechanism, which will be explained in detail below.
[0078] The determination process can be achieved by combining a visual language model (VLM) with semantic similarity calculation. First, the interface image at time step t+1 is input into a pre-trained visual language model (e.g., BLIP-2), and the visual language model outputs a natural language interface summary text (Description Text) based on the image content.
[0079] Assuming the current subtask is "Complete payment for goods in the shopping app," the interface image at time step t+1 displays a page containing a green checkmark icon and the words "Payment Successful." The VLM-generated interface summary text... For example, it could be expressed as: "The current screen displays the payment success page, including order details and a 'Return to Homepage' button."
[0080] Next, retrieve the "Task Completed" status description text. Regarding the aforementioned shopping task, For example, it could be: "The payment process has ended, and the screen displays a payment success confirmation message."
[0081] For example, retrieve the "task completed" status description text. This can be achieved by leveraging the reasoning capabilities of Large Language Models (LLMs) to predict the interface features or system feedback that should be displayed when a subtask is successfully executed, based on the description of the subtask and its contextual information. For example, a prompt word including a description of the subtask could be constructed, such as: "Given subtask..." "Please describe the text or visual state that would typically be displayed on the user interface (UI) when the task is successfully completed." Input this prompt into the large language model, and the natural language text output by the large language model will serve as the input. .
[0082] This mechanism, capable of generating target state descriptions based on causal reasoning, enables the agent to establish accurate stopping criteria even when facing zero-shot tasks. For the aforementioned shopping task, the large language model generates [the following] based on semantic deduction of "complete payment". For example, it could be: "The payment process has ended, and the screen displays a payment success confirmation message."
[0083] Example, get This can also be achieved by consulting a pre-built task-state knowledge base. The knowledge base stores metadata for a large number of historical successful trajectories (Expert Trajectories). This application's embodiments are not limited. The specific source can be generated by the model, or it can be a predefined rule or a search result.
[0084] Subsequently, the summary text of the computing interface. and Semantic similarity between two texts can be achieved by inputting them into a text encoder to obtain high-dimensional semantic vectors. and (For example, the dimension is 768).
[0085] semantic similarity It can be obtained by calculating the cosine similarity of two vectors, for example, expressed by formula (1): (1) in, express and Perform dot product operation. The L2 norm represents the magnitude of a vector.
[0086] Finally, the calculated semantic similarity With preset threshold (For example ) for comparison. Assuming calculations show that... and The cosine similarity is 0.92. Because... If the logical judgment satisfies the condition of "the text representation subtask has been completed", the agent model will then terminate the iteration loop of the current subtask.
[0087] In some embodiments, action prediction based on the input data at time step t to obtain the action at time step t can be achieved by: performing feature encoding on the input data at time step t to obtain encoded features; performing feature mapping on the encoded features to obtain an action probability distribution, wherein the action probability distribution includes the probability value of each action in a preset action set being an action at time step t; and determining the action at time step t based on the action probability distribution.
[0088] For example, see Figure 7 , Figure 7 This is a schematic diagram illustrating the principle of action prediction of the intelligent agent model provided in the embodiments of this application. Figure 7 This demonstrates a single-step interaction loop of an intelligent agent model within a human-computer interaction program (environment). At time step t, the current interface image is acquired and, combined with a subtask (e.g., "entering XX brand mobile phone in the search bar"), the "input data for time step t" is constructed. This input data is fed into the intelligent agent model, undergoing sequential "feature encoding" and "feature mapping" processes, ultimately outputting an "action probability distribution." Based on this action probability distribution, the intelligent agent model determines the specific action for time step t (e.g., ...). Figure 7 (See the "click search box" shown). After this action is executed, the environment state changes, generating the "interface image at time step t+1", thus entering the loop for the next time step.
[0089] Following the previous example, input the data. It is processed through multiple layers of self-attention mechanism. In each layer, the query matrix is used. Key matrix Sum matrix Calculate the attention weights, for example, as shown in formula (2): (2) Through this mechanism, the intelligent agent model can capture the long-distance dependencies between specific UI elements and text commands in an image, ultimately outputting context-aware encoded features. .
[0090] Secondly, feature mapping is performed on the encoded features to obtain the action probability distribution, where the action probability distribution includes the probability value of each action in the preset action set being the action at the t-th time step.
[0091] For example, this can be achieved using a multilayer perceptron (MLP) as the prediction head, which encodes high-dimensional features. Projected onto the dimension of the action space size.
[0092] For example, assume the size of the action set is... (For example, actions such as clicking coordinate areas, swiping the screen, and inputting text). The feature map outputs a... Logits vector of dimension To obtain the probability distribution, the vector... Apply the Softmax function for normalization and calculate the th... The probability of an action ,in, It is the th in the Logits vector The value of each element.
[0093] Finally, based on the action probability distribution, the action at time step t is determined.
[0094] In some embodiments, a greedy decoding strategy can be used to directly select the action index with the highest probability value as the current action, i.e. .
[0095] In other embodiments, to increase exploratory power, nucleus sampling or temperature sampling strategies can be used to randomly select actions from the probability distribution. This application does not limit the specific sampling strategy; it can be determined based on the specific application scenario's requirements for determinism or diversity.
[0096] By encoding and fusing multimodal features of the interface images, subtasks, and task instructions of the human-computer interaction program, a context-aware action probability distribution can be generated, thereby achieving accurate action prediction. Through an iterative "perception-decision-execution-feedback" closed-loop mechanism, the intelligent agent model can dynamically adapt to real-time changes in the interface state. Furthermore, by introducing a task termination determination mechanism based on visual language models and semantic similarity, the reliability of the task closed loop is improved.
[0097] In step 102, first action sequence samples that meet preset sample selection conditions are selected from multiple first action sequence samples to form a first target sample set.
[0098] In some embodiments, selecting first action sequence samples that meet preset sample selection conditions from a plurality of first action sequence samples can be achieved in the following way: for each first action sequence sample, calculate the sample evaluation parameter of the first action sequence sample, and select the first action sequence samples whose sample evaluation parameter is greater than or equal to the preset evaluation parameter threshold as the first action sequence samples that meet the sample selection conditions.
[0099] Here, the sample evaluation parameter refers to a numerical indicator used to quantify the quality, effectiveness, or degree of matching between the first action sequence sample and the first cue word (task intent). It can be a probability score reflecting the semantic consistency between the execution path and the instruction, a ratio reflecting the correctness of the action sequence structure, or a confidence value reflecting the success of the final task completion state.
[0100] Here, the sample selection criteria are based on one or more threshold values (such as the lower limit of the score) set by the sample evaluation parameters, which are used to filter out low-quality, irrelevant or incorrectly executed action sequences, thereby ensuring that the retained first target sample set has high training reference value.
[0101] In some embodiments, the sample evaluation parameter of the first action sequence sample can be calculated by: calculating the advantage value of the first action sequence sample and using the advantage value as the sample evaluation parameter of the first action sequence sample.
[0102] Here, the advantage value refers to a numerical value calculated based on a pre-set reinforcement learning policy, used to measure the superiority of an action sequence sample relative to a baseline. Advantage values can be, for example, within-group standardized advantage values calculated using the Group Relative Policy Optimization (GRPO) algorithm (i.e., normalized based on the mean and variance of the reward values of samples within the same group), generalized advantage estimation (GAE) values calculated using the Proximal Policy Optimization (PPO) algorithm combined with value function estimation, or the difference relative to the historical moving average baseline calculated using the Reinforce++ algorithm.
[0103] In some embodiments, the advantage value of a first action sequence sample can be calculated by: calculating the reward value of each first action sequence sample; determining a baseline value based on the reward values corresponding to multiple first action sequence samples; and determining the advantage value of each first action sequence sample based on the reward value and the baseline value corresponding to the first action sequence sample.
[0104] For example, the reward value for each first action sequence sample can be calculated using a pre-defined rule model or a trained reward model. For instance, if the first action sequence sample successfully completes the task indicated by the first cue word, the reward value is 1; if it fails to complete the task or triggers an error, the reward value is 0. Intermediate rewards can also be introduced, assigning a continuous value between 0 and 1 based on the length of the action sequence, the safety of the operation, or its conformity to human habits.
[0105] For example, the baseline value can be determined based on the reward values corresponding to multiple first action sequence samples. This can be achieved by calculating the arithmetic mean of the reward values corresponding to multiple first action sequence samples as the baseline value.
[0106] For example, for each first action sequence sample, the advantage value of the first action sequence sample can be determined based on the reward value and the baseline value corresponding to the first action sequence sample. This can be achieved by calculating the difference between the reward value and the baseline value as the advantage value.
[0107] Here, the advantage value is used to characterize the "relative superiority" of a particular first action sequence sample compared to the average level. If the advantage value is positive, it means that the first action sequence sample performs better than the average level and should be retained or reinforced in subsequent training; if the advantage value is negative, it means that the first action sequence sample performs worse than expected and should be filtered or suppressed.
[0108] For example, in order to unify the evaluation scale of different tasks, the calculation of the advantage value can also be achieved through standardization, for example, by using the following formula (3): (3) in, Indicates the first The advantage value of the first action sequence sample. Indicates the first The reward value of the first action sequence sample. This represents the set of reward values for this group of samples (i.e., multiple first action sequence samples) generated based on the same first cue word. Indicates the baseline value (average). This represents the standard deviation of the reward values for that group. To prevent extremely small constants from being divided by zero, the advantage value obtained in this way is used as a sample evaluation parameter, which can more robustly screen out the truly outstanding "high-quality" samples generated in this batch, without being affected by the absolute difficulty of the task itself (which leads to a low overall score).
[0109] By calculating the reward value of the first action sequence sample and normalizing it to the advantage value using the group average baseline, the absolute score difference caused by different task difficulties can be eliminated. This accurately quantifies the relative superiority of each first action sequence sample compared to the average level, thereby effectively screening out first action sequence samples that truly possess high-quality policy characteristics and suppressing inefficient or erroneous first action sequence samples. This makes the model training process more robust, has higher convergence efficiency, and can adaptively optimize decision-making strategies in complex and variable tasks.
[0110] In some embodiments, the agent model is used to interact with a human-computer interaction program; see also Figure 3B The calculation of the sample evaluation parameters of the first action sequence sample can be achieved through the following steps 201 to 204, which are explained in detail below.
[0111] In step 201, the page layout information of the human-computer interaction program is obtained when each action in the first action sequence sample is executed.
[0112] In some embodiments, obtaining page layout information can be achieved by calling the operating system's Accessibility API or by parsing the Document Object Model (DOM) tree. Page layout information can be presented in a tree structure, where each leaf node represents a specific interactive control (such as a button, text box, image, etc.); each node contains the control's attribute information, such as the control's class, text content, content description, resource ID, and its bounding box on the screen.
[0113] For example, assuming the current time step's interface is a product details page in an e-commerce application, the obtained page layout information can be represented as a nested JSON object or XML fragment. For instance, for an "Add to Cart" button, the extracted layout information fragment might be: {node_id: 105, class: "android.widget.Button", text:"Add to Cart", bounds: [450, 1800, 650, 1900], visible: true}. If the first action sequence contains... Each action will collect corresponding data. A snapshot of the page layout structure like a frame.
[0114] In other embodiments, when the underlying structured layout information cannot be directly obtained (e.g., in a remote desktop or game interface), page layout information can also be obtained through computer vision technology. For example, optical character recognition (OCR) is performed on the current interface screenshot to extract text blocks and their coordinates, and an object detection model is used to identify non-text controls (such as icons and input boxes). Finally, all the identified elements and their position information are reconstructed into page layout information.
[0115] In step 202, for each action in the first action sequence sample, the target area corresponding to the action is determined based on the operation coordinates of the action and the page layout information corresponding to the action.
[0116] In some embodiments, determining the target region is a hit-testing or coordinate mapping process. This involves using normalized coordinates or absolute pixel coordinates recorded in the first action sequence sample. This involves geometrically matching the bounding boxes of each control in the page layout information. For example, a depth-first search (DFS) strategy can be used to traverse the leaf nodes of the layout tree, find the bottom-level control (Leaf Node) that contains the coordinate point and has the smallest area, and consider it as the target area (UI Element) of the actual operation.
[0117] For example, suppose the coordinates of a click action are... Traversing the page layout information, it was found that the coordinates fell within both a parent container Linear Layout (range [0, 1000, 1080, 2000]) and a child control Button (range [450, 1800, 650, 1900]). Since the child control Button is in a deeper hierarchy and has a smaller area, the target area of the action was determined to be the child control Button.
[0118] In other embodiments, if the action type is scrolling or dragging, then start and end coordinates are involved. In this case, determining the target area can be based on the control where the start coordinates are located, or on the scrollable container with the largest coverage area of the scroll trajectory. This application does not limit the specific geometric determination logic; it only needs to establish a correspondence between actions and interface elements.
[0119] In step 203, an execution path description is generated based on the description information of the target region corresponding to each action.
[0120] In some embodiments, generating an execution path description is a process of converting a series of discrete interactive events into a natural language narrative. First, key semantic attributes of each control identified as a target area are extracted, such as text labels, function descriptions, or resource IDs. Next, the action type (e.g., click, input, drag, etc.) is templated and concatenated with the control's semantic attributes to form a single-step description. Finally, all single-step descriptions are concatenated in chronological order to generate a complete execution path description text.
[0121] For example, suppose the first action sequence contains three actions. Action 1: The target area is the search box, the action is to enter "running shoes", and the generated step description is "enter 'running shoes' in the search box". Action 2: The target area is the search button (icon), the action is to click it; the generated step description is "click the search icon". Action 3: The target area is the first item in the product list, the action is to click it; the generated step description is "click the first search result".
[0122] Based on the above steps, the generated overall execution path description (Trajectory Caption) is: "Enter 'running shoes' in the search box, then click the search icon and select the first search result."
[0123] In other embodiments, an end-to-end multimodal large model (VLM) can also be used to generate execution path descriptions. Screenshot sequences and action trajectory markers corresponding to the action sequence are input into the VLM, and the model's image understanding capabilities are used to directly output a summary text. This text summary can capture implicit logic arising from interface changes (e.g., the user clicked cancel due to a stockout pop-up), thus providing richer contextual information than template splicing.
[0124] In step 204, the semantic matching degree between the execution path description and the first prompt word is calculated, and the semantic matching degree is used as the sample evaluation parameter.
[0125] In some embodiments, semantic matching is calculated by mapping the text to a high-dimensional semantic vector space. For example, a pre-trained text embedding model (such as BERT, Transformer, etc.) is used to convert the "first prompt word" (specifically a task instruction, such as "buy me a pair of men's running shoes") and the execution path description generated in step 203 into fixed-length feature vectors. Subsequently, the consistency of their semantic intent is quantified by calculating a similarity metric (such as cosine similarity) between the two vectors.
[0126] For example, suppose the embedding vector of the task instruction in the first prompt is... The embedding vector describing the execution path is ,in, The dimension is a vector (e.g., 768). Sample evaluation parameters. The cosine similarity can be used for calculation, and the result is... The range of values is within Between these values, the closer the value is to 1, the more closely the actual execution path matches the user's intent, and the higher the sample quality.
[0127] In other embodiments, semantic matching can be calculated using a Large Language Model (LLM) as a scorer (LLM-as-a-judge). For example, a scoring prompt word can be constructed, including the task instruction and execution path description in the first prompt word. The LLM is then required to output a specific numerical score (e.g., 0 to 10 points) or confidence probability based on whether the execution path effectively completes the task corresponding to the instruction. This method can handle more complex logical judgments, such as identifying situations where the path description and instruction do not perfectly match, but the actual result is the same.
[0128] By acquiring page layout information and accurately mapping action coordinates to specific interface elements, low-dimensional physical action sequences are transformed into execution path descriptions with high-level semantics, thus realizing the semantic expression of action data. Furthermore, by calculating the semantic matching degree between this path description and user instructions, the consistency between action sequences and user intentions can be accurately quantified from the semantic understanding level, thereby providing more granular and interpretable evaluation feedback for intelligent agent models.
[0129] In some embodiments, the agent model is used to interact with a human-computer interaction program; see also Figure 3C The calculation of the sample evaluation parameters of the first action sequence sample can also be achieved through the following steps 301 to 303, which are explained in detail below.
[0130] In step 301, the longest common action sequence between the first action sequence sample and the reference action sequence is taken as the target action sequence, and the length of the target action sequence is taken as the target length. The reference action sequence is a pre-configured action sequence corresponding to the first prompt word.
[0131] In some embodiments, determining the target action sequence involves identifying the maximum overlap between a first action sequence sample and a reference action sequence that maintains a consistent relative order. First, each action in both the first and reference action sequences needs to be converted into a comparable feature identifier. This feature identifier can be a tuple or hash value consisting of action type (e.g., Click, Scroll), operation object identifier (e.g., Resource-ID, XPath), and operation parameters (e.g., Text Content). Next, a comparison matrix is constructed using dynamic programming. The set of action elements common to both the first and reference action sequences, maintaining a consistent relative order, is calculated using state transition equations; this set represents the target action sequence. Here, the longest common subsequence allows for discontinuous intervals within the sequence. This means that as long as the core critical steps are in the correct order, irrelevant redundant operations (e.g., accidental touches or unnecessary swipes) or missing non-critical steps will not completely block the match, but will affect the total length of the common sequence.
[0132] For example, suppose the reference action sequence contains 5 action steps, represented as follows: The first action sequence sample (the predicted sequence generated by the agent model) contains 6 action steps, represented as follows: ,in, and This is an erroneous operation. Construct a two-dimensional dynamic programming matrix. ,in, express forward elements and forward The length of the common sequence of elements. Through calculation, the common action sequence is identified as follows: At this point, the target action sequence is... The number of actions included is 4, therefore the target length is determined. .
[0133] In other embodiments, the comparison of actions can also employ a fuzzy matching strategy. When determining whether two actions are "common actions," the operation coordinates or parameters are not required to be completely identical. Instead, the similarity between single-step actions is calculated. If the similarity exceeds a threshold (e.g., the Euclidean distance between coordinates is less than 20 pixels, or the edit distance of the input text is less than 2), it is considered a common action. This increases the robustness of the evaluation algorithm to different screen resolutions or UI tweaks.
[0134] In step 302, a first length of the first action sequence sample is determined, and a second length of the reference action sequence is determined, and the maximum value of the first length and the second length is taken as the third length.
[0135] In some embodiments, determining the first length and the second length is a process of counting discrete action units included in the sequence. First length It characterizes the total number of steps actually executed by the agent, reflecting the total workload including correct steps, redundant steps, and erroneous steps; the second length This represents the total number of optimal steps in either the expert demonstration or the standard answer. The maximum of the two is taken as the third length. This design aims to construct a stringent normalized denominator. This denominator implicitly includes a penalty mechanism for two types of bias: if the sequence generated by the agent is too long ( This indicates that redundant operations exist. Using it as the denominator will lower the final score; if the sequence generated by the agent is too short ( This indicates that a step was omitted, although Small, but useful As a denominator, it will also limit the final score.
[0136] For example, following the example of step 301 above: First action sequence sample It includes 6 actions, therefore the first length Reference action sequence It includes 5 actions, therefore the second length .Compare and Take the maximum value Therefore, the third length is determined. .
[0137] In step 303, the ratio between the target length and the third length is calculated to obtain the sample evaluation parameters of the first action sequence sample.
[0138] In some embodiments, calculating the ratio is the process of generating the final normalized score, and the range of values for the sample evaluation parameters is [range missing]. A value of 1 indicates that the sample sequence is completely identical to the reference sequence (with neither redundancy nor missing values); a value of 0 indicates that the two are completely unrelated.
[0139] Continuing with the previous example, the target length The third length Sample evaluation parameters = .
[0140] In other embodiments, a non-linear penalty coefficient may be introduced when calculating the ratio. For example, a higher weighting is given to matches corresponding to critical steps (such as the final "submit order" action). Assuming a critical action match is counted as 2 length units and a regular action as 1 length unit, the weighted ratio is recalculated. and This allows samples missing key steps to receive significantly lower evaluation parameters, ensuring that the selected samples are correct in the core business logic.
[0141] By extracting the longest common subsequence between the first action sequence sample and the reference action sequence, the consistency of action logic in relative timing can be accurately captured while tolerating interference from non-critical redundant operations or local noise (such as accidental touches). At the same time, by selecting the maximum value of the two sequence lengths as the normalization denominator, a dual penalty mechanism is constructed for operation redundancy (too many steps) and missing steps (task incomplete), ensuring that the sample evaluation parameters can take into account both the accuracy of task completion and the efficiency of the operation path.
[0142] In some embodiments, the agent model is used to interact with a human-computer interaction program; see also Figure 3D The calculation of the sample evaluation parameters of the first action sequence sample can also be achieved through the following steps 401 to 403, which are explained in detail below.
[0143] In step 401, the task feedback result after executing the first action sequence sample in the human-computer interaction program is obtained.
[0144] In some embodiments, obtaining task feedback results refers to the process of capturing and parsing the final state of the human-computer interaction program after the agent model has executed the last action in the first action sequence sample. This involves recording a screenshot, the page hierarchy tree (DOM Tree or Accessibility Node Tree), and system log information at the moment the first action sequence sample ends. This information collectively constitutes the task feedback result, indicating whether the task was successfully completed or what errors were encountered. To ensure the stability of the feedback result, a preset waiting time window (e.g., 500ms to 2000ms) can be introduced after the last action is executed, waiting for the page to load or the animation to finish rendering before collecting a state snapshot.
[0145] For example, suppose the task instruction is "Book a train ticket to Beijing for tomorrow." After the agent clicks "Confirm Payment," it captures the current page layout information. If the obtained page layout information includes the text "Ticket issued successfully" or a specific green checkmark icon (Icon ID: success_tick) is detected, then this page element information constitutes the task feedback result. If a dialog box containing the text "Insufficient balance" pops up, then this error message constitutes the task feedback result.
[0146] In other embodiments, task feedback results can also be obtained by detecting underlying network communication or application callback interfaces. This is achieved by intercepting HTTP / HTTPS response messages sent by the human-computer interaction program and analyzing the returned status code and response body. For example, if an API-returned data packet containing {status: 200, message: "order_created"} (indicating that an order has been successfully created or a task instruction has been effectively executed) is captured, the structured data can be directly used as the task feedback result.
[0147] In step 402, based on the task feedback results, the matching result between the first action sequence sample and the task instruction is determined.
[0148] In some embodiments, the task feedback result includes the interface image of the human-computer interaction program; based on the task feedback result, the matching result between the first action sequence sample and the task instruction is determined, which can be achieved by: performing feature encoding on the task instruction to obtain task features, and performing feature encoding on the interface image to obtain interface features; calculating the semantic similarity between the task features and the interface features, and using the semantic similarity as the matching result.
[0149] In some embodiments, task instructions are processed by a pre-trained text encoder. The text encoder converts the instructions in natural language form into a dense vector in a high-dimensional space, i.e., task features. Task features contain semantic information such as the task's intent, the object of operation, and the desired state. Simultaneously, the interface image, which serves as the task feedback, is processed by a pre-trained image encoder. The image encoder extracts information such as visual elements, layout structure, and text texture from the image, converting it into another high-dimensional vector in the same feature space, i.e., interface features.
[0150] Next, the similarity measure between the two feature vectors is calculated, for example, using cosine similarity. The higher the semantic similarity value, the more semantically the state presented by the interface image matches the expected result described by the task instruction.
[0151] For example, the task instruction is "Open Settings page". Scenario 1: After the first action sequence is executed, the screen remains on the "Settings" main menu interface. The image encoder recognizes the gear icon and the "Settings" title, and the semantic similarity between the generated interface features and the task features including the "Settings" semantics is 0.92.
[0152] In step 403, the sample evaluation parameters are determined based on the matching results.
[0153] In some embodiments, semantic similarity is normalized to obtain sample evaluation parameters.
[0154] For example, since the original semantic similarity scores may be distributed across different numerical ranges, or to adapt to the requirements of subsequent model training on the input data distribution, normalization or rescaling is necessary. One approach is to use Min-Max Normalization, mapping the matching results to a closed interval [0, 1]. If the original similarity scores are mainly distributed between [-1, 1], and negative values indicate complete irrelevance, all negative values can be truncated to 0, or linearly mapped to scores below 0.5.
[0155] For example, see Figure 6 , Figure 6 This is a schematic diagram illustrating the calculation principle of the sample evaluation parameters provided in the embodiments of this application. For example... Figure 6As shown, the calculation process of sample evaluation parameters adopts a two-stream processing architecture. The left-hand flow processes the "task instruction in the first prompt word" (e.g., the text "payment completed"), performing feature encoding through a text encoder to extract "task features" in the form of a high-dimensional vector. The right-hand flow processes the "task feedback result" (i.e., the interface image including the execution result) after the action is executed, performing feature encoding through an image encoder to extract "interface features" in the form of a high-dimensional vector. Subsequently, the task features and interface features are converged to the "semantic similarity calculation" module (e.g., calculating cosine similarity). ,in, Indicates interface features, The task features are represented, and the original "matching result" (e.g., a similarity value of 0.92) is obtained. Finally, the matching result (similarity) is normalized to map the similarity to a specified interval (e.g., [0,1]), thereby outputting the final sample evaluation parameters used to measure sample quality.
[0156] Another approach is to use a non-linear transformation, such as the sigmoid function. This transforms semantic similarity into a probability value, representing the confidence that the first action sequence sample can successfully complete the task. Through this process, high-confidence matches tend to be close to 1, while low-confidence matches tend to be close to 0, thus widening the score gap between good and bad samples.
[0157] In other embodiments, determining the matching result is a binary classification or multi-level scoring process based on a predefined rule set or discriminator model. "Success condition sets" and "failure condition sets" are predefined for different task types. The task feedback results obtained in step 401 (e.g., text information recognized from interface images using optical character recognition technology, control resource IDs or attribute values in the page layout tree, etc.) are scanned to determine whether the included keywords, key element IDs, or image features match the aforementioned condition sets. If no failure condition is matched but at least one success condition is matched, the matching result is determined to be a "Positive Match" and assigned a higher sample evaluation score; otherwise, it is determined to be a "Negative Match".
[0158] For example, let the set of success keywords corresponding to the task instruction be: The set of text content extracted from the task feedback results in step 401 is as follows: Perform set intersection operation. Due to the intersection empty set And detected The error message includes the phrase "system busy". Therefore, the matching result is determined to be "task failed", and the corresponding sample evaluation parameters (such as success rate indicators) are set to 0 or a small penalty value (such as -1). If the match result is "task successful", the sample evaluation parameter is set to 1.
[0159] By capturing the task feedback results of human-computer interaction programs and using feature encoding technology to calculate the semantic similarity between interface features and task instructions or to match them through key rules, alignment from "execution results" to "user intent" is achieved; unstructured task feedback results are transformed into quantifiable sample evaluation parameters, thereby providing objective sample evaluations of the intelligent agent model based on the task achievement effect.
[0160] In other embodiments, for certain specific scenarios (such as payment result pages and form submission pages), the interface images for "task success" and "task failure" may have extremely high visual similarity overall (e.g., 90% of the pixel area is the same, with differences only in the color of the pop-up text or status bar icon). Directly calculating the semantic similarity of the entire image can easily lead to misjudgment (i.e., the Hard Negative problem). Therefore, determining sample evaluation parameters based on task feedback results can also be achieved in the following ways: First, the target region is obtained. For example, by performing pixel-level difference calculations between the interface image (current frame) after the execution of the first action sequence sample and the interface image (previous frame) of the previous time step or a preset success / failure template image, the changed region is extracted as the target region (or difference region); or, using the attention mechanism (Attention Map) of the Visual Language Model (VLM), the saliency region (Saliency Region) in the interface that is highly related to the semantics of the task instructions is extracted as the target region, such as the area of a newly popped-up dialog box or the text area where the color changes abruptly.
[0161] Secondly, information is extracted from the target area to obtain regional information. For example, feature encoding is performed on the image content within the target area, and the encoded features are used as regional information. Alternatively, optical character recognition (OCR) is performed on the text within the target area, and the recognized text is used as regional information. For instance, when the target area is a "payment result pop-up," the focus is on extracting the text "payment successful" or "insufficient balance" within the target area.
[0162] Finally, the matching degree between the region information and the task instructions is calculated, and this matching degree is used as a sample evaluation parameter. This can be achieved by calculating the similarity between the feature vector of the target region and the feature vector of the task instructions, or by determining whether the text in the target region contains the keywords expected by the task instructions.
[0163] By focusing on the most significant or changing target areas in the interface for evaluation, the interference of background images (such as common APP backgrounds) on similarity calculation can be effectively eliminated, and subtle changes in visual state can be accurately captured, thereby greatly improving the ability to identify highly similar negative samples and ensuring the accuracy of sample evaluation parameters.
[0164] See also Figure 3A In step 103, the agent model is trained based on the first target sample set.
[0165] Here, training an agent model based on the first target sample set refers to using reinforcement learning (RL) algorithms to update the weight parameters of the agent model (policy network) based on the selected high-quality action sequence samples (i.e., the first target sample set), so as to maximize the probability of the agent model generating high-reward actions when facing the same or similar environmental states in the future.
[0166] In the first training phase (or "imitation phase") of training the agent model on the first target sample set, the agent model learns only from the first target sample set that is evaluated as positive (i.e. the sample evaluation parameter is greater than or equal to the preset evaluation parameter threshold). The goal of the first training phase is to enable the agent model to quickly imitate and absorb "successful" experiences, prioritize mastering the core path and correct operation to complete the task, and lay a solid foundation for the subsequent exploration and optimization phases.
[0167] In some embodiments, training an agent model based on a first target sample set can be achieved through reinforcement learning training algorithms such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO).
[0168] For example, when the advantage value of the first action sequence sample is used as the sample evaluation parameter for the first action sequence sample to filter samples, a first target sample set is obtained, and an agent model is trained based on the first target sample set, the objective function for model training can be expressed by formula (4): (4) in, This is the objective function for the first training phase; the goal of training is to find a set of optimal model parameters. To maximize the value of the function; Indicates the current policy The complete action sequence (trajectory) generated by the agent model currently being trained. The expectation is calculated by taking the average value of the first action sequence sample (i.e., the first action sequence sample). It is an indicator function. Representative trajectory The dominance value. Only when the dominance value of the entire trajectory is... When the value is greater than or equal to 0 (i.e., the trajectory performs better than or equal to the average level), the function takes a value of 1, and the subsequent objective function term is included in the calculation; otherwise, it takes a value of 0, which is equivalent to ignoring this "bad" trajectory. This ensures that in the first training phase, the model only learns from "good" samples; It is the clipped surrogate objective. In time step The advantage value represents the action to be performed in this state. Compared to the average quality of movement. These are importance sampling weights, representing the new strategy. and old strategies For the same action The ratio of the probability of adoption. Yes The clipping result, for example, ... Limited to a preset small interval Inner and Multiplication. By taking This operation can prevent the steps of a single parameter update from being too large, thereby avoiding drastic fluctuations in policy performance and ensuring the stability of the training process. It is a regularization term used to constrain the update range of the policy. Indicates the current strategy With a reference policy The KL divergence (KL divergence) measures the degree of difference between the distributions of two policies. Reference policy It can be the initial pre-trained model or the stable model from the previous iteration; It is a hyperparameter that controls the strength of regularization. Its function is to penalize updates that deviate too far from the reference policy, prevent the agent model from "forgetting" the general capabilities it has already mastered when learning new samples, and ensure that the behavior of the agent model does not change drastically and uncontrollably.
[0169] Here, the new strategy Old strategy Reference Strategy The following logical relationship exists between the agent model to be trained and the agent model: First, the agent model to be trained is essentially the one that carries parameters. The strategy network.
[0170] Secondly, the new strategy This refers to the agent model whose parameters are being adjusted and optimized in the current gradient update step. In other words, the training process involves continuous updates. parameters ,make Evolve towards higher rewards based on the old strategy.
[0171] Again, the old strategy This refers to collecting the current batch of first action sequence samples (i.e. This is a historical version of the agent model used in the current round of parameter update calculations. The parameters of the old policy remain fixed (and do not participate in gradient backpropagation) during this round of parameter update calculations. Their role is to provide the original probability distribution of actions at the sampling time, so that the ratio can be used to... New Quantitative Strategies Compared to the sampling strategy The extent of the shift that supports the calculation of importance sampling.
[0172] Finally, refer to the strategy This refers to the initial agent model before fine-tuning in this stage of reinforcement learning (e.g., the model after supervised fine-tuning), or the agent model saved at the end of the previous large training stage. During training, the reference policy remains frozen (without updating parameters) as an anchor or benchmark, forcing the new policy being trained through the KL divergence term. Do not deviate too far from the initial language distribution, so as to ensure that the agent model does not lose basic language fluency and logical coherence while pursuing task success rate.
[0173] In summary, by maximizing the objective function defined by formula (4), the agent model can stably learn from the selected high-quality positive samples in the first training stage, which can not only effectively improve the policy performance of completing the task, but also maintain the generalization ability and stability of the model through regularization constraints.
[0174] In step 104, if the training progress of the trained agent model meets the training phase switching conditions, then the trained agent model generates multiple second action sequence samples based on the preset second prompt word to form a second target sample set, and trains the agent model based on the second target sample set.
[0175] In some embodiments, the training progress of the agent model can be determined by the following method: determining the training progress based on the cumulative training time steps of the trained agent model.
[0176] In some embodiments, training progress is a scalar value used to measure the state of the model's training lifecycle, for example, a value ranging from [0, 1] or [0%, 100%]. Determining progress based on cumulative training timesteps is an "open-loop" or pre-planned progress estimation method. Here, "cumulative training timesteps" refers to the number of times the agent model parameters are updated (Iterations / Steps).
[0177] In some embodiments, the training progress is determined based on the cumulative training time steps of the trained agent model, which can be achieved by: obtaining the current cumulative training time steps of the agent model and the preset total training time steps; calculating the ratio between the cumulative training time steps and the total training time steps; and determining the ratio as the training progress.
[0178] For example, assuming a preset total number of training steps ( The training time is 1000 steps. When the model reaches step 200, the current cumulative step count is obtained. Calculate the ratio: At this point, the training progress is determined to be 20%.
[0179] By calculating the ratio of cumulative training time steps to the preset total training time steps, a progress metric with low computational cost and monotonically increasing characteristics is constructed. This provides a stable time benchmark for model training that is not affected by performance fluctuations, thereby ensuring that the agent can accurately trigger the switching of training phases according to the preset rhythm, effectively improving the stability and controllability of multi-stage training strategy execution.
[0180] In some embodiments, the training progress can also be determined by: calculating the sample evaluation parameters of each first action sequence sample; fusing the sample evaluation parameters corresponding to multiple first action sequence samples to obtain fused evaluation parameters; and using the ratio between the fused evaluation parameters and the preset target evaluation parameters as the training progress.
[0181] Here, the implementation method for calculating the sample evaluation parameters of each first action sequence sample can be found in the description above, and will not be repeated here.
[0182] For example, the fusion evaluation parameters obtained by fusing the sample evaluation parameters corresponding to multiple first action sequence samples can be achieved in any of the following ways: calculate the arithmetic mean, weighted average, etc. of the sample evaluation parameters corresponding to multiple first action sequence samples, and use them as the fusion evaluation parameters.
[0183] For example, a target evaluation parameter is preset, which represents the expected evaluation parameter value when the agent model "meets the training phase switching conditions". The ratio between the fused evaluation parameter and the preset target evaluation parameter is used as the training progress.
[0184] By calculating the ratio of the fusion evaluation parameters to the preset target evaluation parameters, a closed-loop progress measurement mechanism based on the actual performance of the model is constructed. The multi-sample fusion smooths out the random fluctuations of a single sampling, which can accurately reflect the gap between the current convergence degree of the agent model and the expected target. This ensures that the switching of the training phase is driven by the substantial improvement of the model's capabilities, effectively avoiding the problems of insufficient or overtraining that may occur under the fixed-step strategy, and enhancing the adaptability and robustness of the training process to the model state.
[0185] In some embodiments, if the training progress of the trained agent model meets the training phase switching conditions, the second training phase is entered. In the second training phase, the trained agent model generates multiple second action sequence samples based on preset second prompt words to form a second target sample set. The second target sample set includes positive samples (or positive samples) that meet the sample selection conditions and negative samples (or negative samples) that do not meet the sample selection conditions.
[0186] Here, the training phase switching condition refers to a critical point set according to the training progress of the agent model, used to trigger the transition from the first training phase (imitation phase) to the second training phase (discrimination phase). For example, the training phase switching condition can be set as follows: the current number of training steps accounts for a preset percentage of the total number of training steps (training progress) that reaches a preset percentage threshold (e.g., 10%~30%). The purpose of setting this switching point is to introduce more comprehensive feedback signals in a timely manner after the agent model has established a basic and stable behavioral pattern using positive samples.
[0187] Unlike the first training phase, which only utilizes positive samples, the second training phase introduces action sequence samples with negative dominance values (i.e., performance below average) into the training process. This aims to establish a full-spectrum feedback mechanism. In this phase, all second action sequence samples are combined into a second target sample set, and sample selection based on sample evaluation parameters is no longer performed. That is, the second target sample set includes both positive samples whose sample evaluation parameters (e.g., dominance values) meet preset selection conditions (e.g., greater than or equal to 0) and negative samples whose sample evaluation parameters do not meet preset selection conditions (e.g., less than 0).
[0188] The goal of the second training phase is to enable the agent model to learn to "distinguish" between optimal and suboptimal behaviors, building upon the stable behavioral patterns established in the first training phase. By involving all samples in gradient calculation, the agent model can clearly identify which operations are inefficient, erroneous, or logically broken (e.g., clicking on invalid areas or outputting out of format). This reduces the probability of generating these actions during parameter updates (i.e., suppressing suboptimal behaviors), significantly enhancing the model's generalization ability and robustness in complex scenarios.
[0189] Here, the second cue word refers to the initial set of input data given to the agent model to trigger it to perform a specific task. It is the contextual basis for the agent model to generate action sequences, which clarifies the target task that the agent model needs to complete and the environmental state at the start of the task, thereby limiting the search space for the agent model's subsequent action planning.
[0190] It should be noted that the "first prompt word" and "second prompt word" in this application embodiment are not fundamentally different in terms of data attributes. "First" and "second" are only used to distinguish temporal logic, that is, they refer to the prompt word data input to the agent model in the "first training phase" and "second training phase," respectively. This does not mean that the second prompt word is more difficult or complex than the first prompt word, or belongs to a different business domain. The second prompt word can be exactly the same as the first prompt word, or it can be sampled from different batches of the same dataset.
[0191] In some embodiments, the second prompt word includes a task instruction; multiple second action sequence samples are generated based on the preset second prompt word, which can be achieved by performing the following process multiple times: performing semantic understanding on the second prompt word to obtain at least one subtask corresponding to the task indicated by the task instruction in the second prompt word; performing actions to process each subtask; and combining the actions corresponding to each subtask into a second action sequence sample.
[0192] For specific implementation details, please refer to the description of step 101 above, which will not be repeated here.
[0193] For example, when training an agent model based on a second target sample set, the objective function for model training can be expressed by formula (5): (5) in, It is the objective function for the second training phase.
[0194] In other embodiments, before generating multiple second action sequence samples based on a preset second prompt word, the following processing may also be performed: if the training progress is greater than or equal to a preset training progress threshold, it is determined that the training progress of the agent model meets the training phase switching condition; if the training progress is less than the training progress threshold, it is determined that the training progress of the agent model does not meet the training phase switching condition, and the process of generating multiple first action sequence samples based on a preset first prompt word is initiated.
[0195] Here, the training progress threshold refers to a pre-set time point or proportional boundary used to divide the "imitation phase" (first training phase) and the "discrimination phase" (second training phase). The training progress threshold is a hyperparameter that can be set based on experience or the complexity of the specific task.
[0196] For example, if the training progress is less than the training progress threshold, it is determined that the training progress of the agent model does not meet the conditions for switching training phases, and the process switches to generating multiple first action sequence samples based on a preset first prompt word. This means that the system remains in the loop of the first training phase (imitation phase). In this case, the closed loop of "generating samples → filtering samples → updating the model using only the filtered samples" continues to be executed.
[0197] The purpose of this logical design is to establish a "cold start" protection mechanism. In the early stages of training, the agent model's policy is not yet stable, and the quality of the generated action sequences varies greatly. Introducing negative feedback (negative-dominant samples) too early can easily lead to excessive gradient variance, causing model parameter oscillations or even policy collapse. By forcing the model to remain in the first training phase, it ensures that the agent model has sufficient time to follow only the "successful" models, accumulating enough correct behavioral patterns (such as correct UI click logic and basic task understanding), laying a solid foundation for subsequent advanced training.
[0198] Here, if the training progress is greater than or equal to a preset training progress threshold, the training progress of the agent model is determined to meet the training phase switching conditions, and it officially enters the second training phase (discrimination phase). At this time, samples are generated based on preset second prompt words, and sample screening is canceled (introducing all positive and negative samples) for training. This switch marks a qualitative change in the training strategy from "simple imitation" to "comprehensive optimization." The agent model begins to actively learn how to avoid errors based on a certain level of discriminative ability, thereby improving its adaptability to complex scenarios.
[0199] Through steps 101 to 104, action sequence samples are generated and filtered based on the first cue word to construct a first target sample set that meets specific conditions for preliminary training of the agent model. This allows the filtered samples to guide the model in establishing a stable basic strategy during the initial training phase. Furthermore, when the training progress meets the switching conditions, the agent model is trained again using a second target sample set generated based on the second cue word. This phased, progressive training mechanism enables the agent model to adaptively adjust the focus of its learning content according to its own training state, promoting the smooth convergence and optimization of model parameters, thereby improving the decision-making accuracy and scene generalization ability of the agent model.
[0200] The following will describe the processing method of the human-computer interaction program provided in the embodiments of this application, with the server as the execution subject, by combining the exemplary application and implementation of the server provided in the embodiments of this application.
[0201] In some embodiments, an action sequence is generated based on a third prompt word using a pre-trained agent model. The actions in the action sequence are used to operate a human-computer interaction program to complete a preset task. The agent model is trained using the agent model training method provided in the embodiments of this application.
[0202] Here, the third cue word refers to the contextual information input into the agent model during the inference phase, which is intended to describe the user's desired interaction intent.
[0203] It should be noted that an action in an action sequence refers to a decision or execution step output by the intelligent agent model based on the perceived environmental state and user instructions, in order to achieve a specific goal. Actions can be categorized as both direct control of the virtual environment or physical devices (such as clicking, inputting, dragging, etc.) and information generation and processing (such as code writing, data format conversion, etc.).
[0204] Furthermore, an action sequence refers to a set of actions with temporal logical relationships or causal dependencies, generated by an agent model based on its understanding and decomposition of complex tasks. An action sequence is not simply a stack of actions, but rather a complete operational flow planned by the agent model to transform an initial state into a target state. In some embodiments, each subsequent action in the action sequence depends on the execution result of the preceding action.
[0205] In some embodiments, the third prompt is multimodal data, such as: a screenshot of the current interface of the human-computer interaction program (Visual State), the user's natural language instructions, and optional historical interaction history.
[0206] Here, the historical interaction trajectory refers to a series of state-action pairs generated during the interaction between the agent model and the human-computer interaction program before the current moment. The historical interaction trajectory is used to provide the agent model with temporal contextual memory to avoid repeatedly executing the same invalid actions or getting stuck in an infinite loop.
[0207] For example, let's assume the current time step is... Then the historical interaction trajectory It can be represented as: .in, Indicates the first The interface state features of the step (which can be image features or DOM tree features). Indicates the agent model in the first... The action instructions to be taken.
[0208] For example, when inputting into the model, each item in the historical interaction trajectory is mapped to the corresponding token sequence and temporal positional encoding is added so that the agent model of the Transformer architecture can capture the causal logic of the operation.
[0209] In some embodiments, the third prompt word includes a natural language instruction (i.e., a task instruction); generating an action sequence based on the third prompt word can be achieved by: performing semantic understanding on the second prompt word to obtain at least one subtask corresponding to the task indicated by the natural language instruction in the second prompt word; executing actions for processing each subtask; and combining the actions corresponding to each subtask into an action sequence.
[0210] For specific implementation details, please refer to the description of step 101 above, which will not be repeated here.
[0211] In some embodiments, the actions in the action sequence are used to operate the human-computer interaction program to complete a preset task. Specifically, this is manifested in the discrete action codes output by the intelligent agent model being converted into control events that the operating system can recognize. Actions include action type and action parameters.
[0212] Examples of action types include: Click, Long Press, Scroll, Type Text, and Terminate. If the agent model outputs an action of SCROLL(start_x=0.5, start_y=0.8, end_x=0.5, end_y=0.2), it will be displayed at screen coordinates... arrive Simulate finger swipe events along the path to drive the content of the human-computer interaction program to scroll, thus completing the actual operation of the interface.
[0213] The embodiments of this application do not limit the specific network architecture of the intelligent agent model. In other embodiments, an end-to-end architecture based on a multimodal large language model (MLLM) can be adopted to directly map screenshots and instructions into JSON format action instructions without explicitly outputting subtask text. This implicit thought chain processing method is also within the protection scope of this application.
[0214] For example, see Figure 4A , Figure 4A This is a schematic diagram of the first interface of the human-computer interaction program provided in this embodiment of the application. In this example scenario, the intelligent agent model is browsing an encyclopedia webpage about "Name A". The interface content includes a top navigation area (including the "XX Encyclopedia" logo and a "Search" function), a middle "Notification Bar" content, and a bottom text area. The left side of the text area is a table of contents index, including navigation items such as "Personal Experience", "Major Achievements", and "List of Works"; the right side contains the corresponding biographical text and portrait of the person.
[0215] See Figure 4B , Figure 4B This is a schematic diagram of the second interface of the human-computer interaction program provided in the embodiments of this application. Figure 4B This diagram illustrates the results of an agent model's visual perception of an operating environment (a webpage). The agent model identifies all interactive or important elements in the interface and generates a bounding box for each element (as shown by the rectangles around each element in the diagram) to determine its precise position on the screen. Simultaneously, each element is assigned a unique identifier (as shown by the numerical code in the diagram). For example, the identifier '9' represents the 'List of Works' link in the directory. The agent's subsequent action planning can be based on these identifiers, such as generating an instruction like 'CLICK(9)', thereby achieving precise operation on the target.
[0216] For example, in order for the agent to understand and operate the graphical user interface, the original screenshot is first structured and parsed using an object detection module or optical character recognition technology: First, element localization and bounding box generation: The agent model identifies all interactive elements (such as buttons, links, and input boxes) or elements containing important information (such as images and text blocks) in the interface, and generates a visual bounding box for each element (e.g., ...). Figure 4B (The rectangular black borders surrounding each element). These bounding boxes precisely define the effective clickable area (HitBox) or information reading area of the element.
[0217] Second, identifier allocation and mapping: Based on the defined boundaries, the agent model assigns a globally unique numerical identifier (Unique Identifier / Index ID, e.g., ...) to each detected element. Figure 4B (The numerical identifier is attached to the upper right corner of the bounding box). Transforming the continuous pixel coordinate space into a discrete ID symbol space significantly reduces the difficulty of decision-making.
[0218] based on Figure 4B Based on the perception results shown, the agent's action planning process is as follows: Suppose the user's instruction is "View A's list of works". The agent first reads the data through a visual encoder. Figure 4B By combining semantic understanding capabilities, the agent matches the intent keyword "list of works" in the natural language command with text elements on the interface. The agent observes that the text "list of works" is located in the left-hand directory and is marked with identifier '9'.
[0219] At this point, the agent model does not need to predict complex screen coordinates (x, y), but directly generates decision instructions based on unique identifiers. For example, the next instruction in the output action sequence is CLICK(9). Subsequently, this instruction is mapped by the interpreter to a click event on the center position of the bounding box corresponding to the identifier '9' on the screen, thereby accurately completing the page jump operation. Similarly, if a search operation is required, after the agent recognizes that the identifier of the search box is '2' and the identifier of the search button is '3', it can generate a combination of actions TYPE(2, "query content") and CLICK(3).
[0220] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0221] With the rapid development of large language models and artificial intelligence technologies, intelligent agents (called GUI Agents) that can autonomously perceive and operate graphical user interfaces (GUIs) are gradually becoming a key technology for realizing the automation of human-computer interaction.
[0222] Specifically, a GUI Agent is an intelligent agent capable of simulating human behavior. It understands screen interface information (such as screenshots and control layout trees) through visual perception and uses a Large Language Model (LLM) or Multimodal Large Model (VLM) as its decision-making "brain" for reasoning and planning, thereby autonomously executing interactive operations such as clicking, swiping, and text input. GUI Agents possess semantic understanding capabilities of the interface and can dynamically generate and execute action sequences in complex cross-application environments based on natural language instructions, thereby completing complex tasks such as software testing, office automation, and intelligent assistant services.
[0223] However, training methods in reinforcement learning for GUI agents still face significant challenges. On one hand, the vast action space and unstructured nature of the GUI environment make models prone to severe gradient fluctuations and unstable convergence during training. On the other hand, while some existing technologies attempt to introduce static curriculum learning methods—that is, pre-dividing task difficulty batches for phased training—this static scheduling strategy often lacks flexibility and struggles to achieve an effective balance between training stability and model generalization ability based on the model's real-time learning status. These limitations ultimately lead to existing GUI agents exhibiting insufficient robustness and poor performance when facing complex interaction scenarios with intricate logic and changing environments.
[0224] In view of this, this application provides a method for training an intelligent agent model, which is described in detail below.
[0225] For example, see Figure 5 , Figure 5 This is a schematic diagram illustrating the principle of the training method for the intelligent agent model provided in the embodiments of this application, as shown below. Figure 5 As shown, the training architecture of the intelligent agent model training method provided in this application embodiment includes three parts: a first training stage (imitation stage), training progress detection and switching logic, and a second training stage (discrimination stage). The specific data flow is as follows: In the first training phase (imitation phase), the agent model (policy network) to be trained receives a preset first cue word and generates multiple first action sequence samples under the current policy. These first action sequence samples are then filtered to determine if they meet preset sample filtering conditions (e.g., whether the task is completed or the score is achieved). First action sequence samples that do not meet the conditions (e.g., invalid or erroneous first action sequence samples) are directly discarded; those that meet the filtering conditions are included in the first target sample set, which has the characteristic of "containing only positive samples". Next, the agent model updates its parameters based on the samples in the first target sample set, completing a closed-loop iteration aimed at quickly establishing a basic policy.
[0226] In the middle part of the architecture, there is a training progress detection and switching controller. The training progress detection and switching controller obtains the training progress of the agent model in real time (e.g., the cumulative time steps) and determines whether the current training progress meets the preset training stage switching conditions: if it is determined to be "no", the first stage (i.e., the imitation stage) is maintained, and the process of screening and updating parameters based on positive samples continues; if it is determined to be "yes", the second stage (i.e., the identification stage) is triggered, and the subsequent process is instructed to cancel the sample screening mechanism.
[0227] In the second training phase (discrimination phase), the agent model (current policy) trained in the first phase continues to generate multiple second action sequence samples based on a preset second cue word. At this stage, no sample selection is performed; all generated samples (i.e., containing both positive and negative samples) directly constitute the second target sample set. Finally, the agent model updates its parameters based on all samples in the second target sample set. During this process, through reinforcement learning algorithms (such as GRPO), the agent model can reinforce behaviors corresponding to positive samples and suppress behaviors corresponding to negative samples, thereby further improving the model's generalization ability and robustness while ensuring stability.
[0228] In summary, the training method for the intelligent agent model provided in this application uses a two-stage course learning strategy consisting of a first training stage and a second training stage. It utilizes sample advantage signals as intrinsic training guidance, balances training stability and generalization ability, and is compatible with various reinforcement learning optimization algorithms such as GRPO and PPO.
[0229] The following is a detailed description of the implementation details of the training method for the intelligent agent model provided in the embodiments of this application.
[0230] I. Preliminary Preparations Define training objectives: optimize the visual, planning, and interactive capabilities of the GUI Agent to improve training stability and generalization ability.
[0231] Prepare input data: GUI task dataset, which includes visual states, natural language commands and corresponding action sequences.
[0232] Initialization parameters: Strategy Set the maximum input / output token length to 1024.
[0233] Reference Strategy KL penalty coefficient β (value 0.02), clipping factor The value ranges from 0.1 to 0.2.
[0234] The stage switching point s (corresponding to the preset training progress threshold mentioned above) is 10%~30% of the total training steps, with a default of 30%.
[0235] Optimizer parameters: AdamW optimizer, learning rate 5×10⁻⁶ -6 Batch size 128, training rounds 3.
[0236] Choose the target optimization algorithm: options include GRPO, PPO, RLOO, or Reinforce++, and specify the corresponding advantage value calculation method.
[0237] II. Sample Generation and Odds Calculation 1. Sample Generation: Input includes GUI task queries containing visual states (such as the initial GUI interface image) and language commands. Through the current strategy Sample to generate batch action sequences (As shown in the multiple first action sequence samples above).
[0238] Advantage value calculation: Calculate the advantage value for each sample pair based on the selected optimization algorithm. Advantages : PPO: Calculated by summing the discounted time difference residuals using the generalized advantage estimation (GAE). GRPO: The relative advantage formula for grouping is used to normalize the sample advantage values within the same group. RLOO / Reinforce: Uses the original advantage value calculation logic of this algorithm.
[0239] Dominance labeling: Based on the calculation results, the samples are classified and labeled. Then it is marked as a positive dominance sample; if Then it is marked as a negative dominance sample.
[0240] III. Stage Assessment and Training Implementation With preset switching points Based on this core principle, the training mode is dynamically switched by real-time detection of the proportion of current training steps to total training steps (i.e., training progress). Switching point The default setting is 30%, which can be flexibly adjusted within the range of 10% to 30% depending on the complexity of the GUI task, without the need for additional model structure adaptation.
[0241] (a) Imitation phase, current training progress < s This stage is the initial training phase. The core objective is to quickly establish a stable model foundation through positive feedback. This can be regarded as a "cold start" protection mechanism for the model, aiming to avoid drastic gradient oscillations caused by policy instability in the early stages.
[0242] 1. Sample selection mechanism: An indicator function is used to filter the generated samples, retaining only samples with positive dominance. ), directly remove negatively dominant samples ( ), so that it does not participate in gradient calculation.
[0243] 2. Training Objective Function: Based on the basic pruning and replacement objective (Surrogate Objective), a KL divergence regularization term is added to prevent the model from crashing due to excessive policy updates and to ensure that the model distribution always closely follows the reference distribution. The objective function for this stage is in the form of formula (4) mentioned above.
[0244] 3. Gradient Calculation and Update: For the objective function in the first stage... Calculate the gradient to ensure that the model only strengthens positive dominant samples. The gradient formula can be expressed by formula (6): (6) in, For indicator functions, This represents the positive advantage value of the sample. After each update, the current policy model is saved as the old policy for the next round. Then, regenerate samples and iterate until the training progress reaches the switching point. .
[0245] (ii) Identification phase, (current training steps / total training steps) ≥ s Once the model has successfully passed the initial stage and possesses basic GUI interaction capabilities, it enters the discrimination stage. This stage introduces negatively dominant samples to train the model to distinguish between "optimal behavior" and "suboptimal / incorrect behavior," thereby improving its generalization ability in complex scenarios.
[0246] 1. Sample usage rules: The sample screening mechanism is cancelled, and all positive and negative dominant samples are fully included in the training to make full use of the complete feedback signal spectrum.
[0247] 2. Training Objective Function: Retain the pruning substitution objective and KL regularization term used in the imitation stage, but remove the positive dominance sample indicator function. The form of the objective function in this stage is referenced in formula (5) above. At this time, the negative dominance sample ( ) will be allowed to participate in the calculation, which will help to suppress suboptimal behavior.
[0248] 3. Gradient Calculation and Update: Parameters are updated using the gradient ascent method. The gradient formula is shown below: (7) This gradient is an unbiased estimate. Positive dominance terms drive the model to reinforce correct interactions (such as precise clicks and formatted output), while the gradient signal generated by negative dominance terms effectively suppresses erroneous behaviors (such as clicks on invalid areas and logical breaks). Training continues iteratively at this stage until the preset total number of training steps is reached, without adjusting hyperparameters, ultimately outputting the optimized agent model.
[0249] The training method for the intelligent agent model provided in this application has the following beneficial effects: 1. Improve training stability: By filtering out high-noise negative signals in the early stages of training, gradient variance is significantly reduced, effectively avoiding gradient oscillations or policy collapse problems common in reinforcement learning, and achieving a smooth "cold start".
[0250] 2. Significantly Enhanced Generalization Ability: By introducing complete positive and negative dominance signals during the identification phase, the agent model learns to finely distinguish between optimal and suboptimal behaviors, building upon its stable capabilities. This strengthens accurate operations through positive samples and suppresses erroneous logic through negative samples, thereby significantly improving the robustness of the GUI Agent's inference in complex cross-modal scenarios.
[0251] 3. Strong algorithm compatibility: No need to modify the core logic of reinforcement learning algorithms such as GRPO and PPO. It can be seamlessly integrated through a phased course scheduling mechanism, with low technical implementation costs and good versatility.
[0252] The following description continues to illustrate the exemplary structure of the training device 133 for the intelligent agent model provided in this application embodiment as a software module. In some embodiments, such as... Figure 2A As shown, the software modules in the training device 133 for the agent model stored in memory 130-1 may include: The first training module 1331 is used to generate multiple first action sequence samples based on a preset first prompt word using the agent model to be trained.
[0253] In some embodiments, the first training module 1331 is further configured to select first action sequence samples that meet preset sample selection conditions from the plurality of first action sequence samples to form a first target sample set.
[0254] In some embodiments, the first training module 1331 is further configured to train the agent model based on the first target sample set; The second training module 1332 is used to generate multiple second action sequence samples based on a preset second prompt word to form a second target sample set if the training progress of the trained agent model meets the training phase switching condition, and to train the agent model based on the second target sample set.
[0255] In some embodiments, the first training module 1331 is further configured to calculate a sample evaluation parameter for each first action sequence sample, and to use the first action sequence samples whose sample evaluation parameters are greater than or equal to a preset evaluation parameter threshold as first action sequence samples that meet the sample selection conditions.
[0256] In some embodiments, the agent model is used to interact with a human-computer interaction program; the first training module 1331 is further used to obtain the page layout information of the human-computer interaction program when each action in the first action sequence sample is executed; for each action in the first action sequence sample, based on the operation coordinates of the action and the page layout information corresponding to the action, determine the target area corresponding to the action; generate an execution path description based on the description information of the target area corresponding to each action; calculate the semantic matching degree between the execution path description and the first prompt word, and use the semantic matching degree as the sample evaluation parameter.
[0257] In some embodiments, the first training module 1331 is further configured to take the longest common action sequence between the first action sequence sample and the reference action sequence as the target action sequence, and take the length of the target action sequence as the target length, wherein the reference action sequence is a pre-configured action sequence corresponding to the first prompt word; determine a first length of the first action sequence sample, and determine a second length of the reference action sequence, and take the maximum value between the first length and the second length as a third length; calculate the ratio between the target length and the third length to obtain the sample evaluation parameter of the first action sequence sample.
[0258] In some embodiments, the agent model is used to interact with a human-computer interaction program, and the first prompt word includes a task instruction; the first training module 1331 is further used to obtain a task feedback result after executing the first action sequence sample in the human-computer interaction program; based on the task feedback result, determine the matching result between the first action sequence sample and the task instruction; and determine the sample evaluation parameters based on the matching result.
[0259] In some embodiments, the first training module 1331 is further configured to perform feature encoding on the task instructions to obtain task features, and to perform feature encoding on the interface image to obtain interface features; calculate the semantic similarity between the task features and the interface features, and use the semantic similarity as the matching result; and normalize the semantic similarity to obtain the sample evaluation parameters.
[0260] In some embodiments, the first training module 1331 is further configured to obtain the current cumulative training time steps of the agent model and the preset total training time steps; calculate the ratio between the cumulative training time steps and the total training time steps; and determine the ratio as the training progress.
[0261] In some embodiments, the first training module 1331 is further configured to determine that the training progress of the agent model meets the training phase switching condition if the training progress is greater than or equal to a preset training progress threshold; and to determine that the training progress of the agent model does not meet the training phase switching condition if the training progress is less than the training progress threshold, and to proceed to the process of generating multiple first action sequence samples based on a preset first prompt word.
[0262] In some embodiments, the first prompt word includes a task instruction; the first training module 1331 is further configured to perform the following processes multiple times to obtain the plurality of first action sequence samples: performing semantic understanding on the first prompt word to obtain at least one subtask corresponding to the task indicated by the task instruction in the first prompt word; performing the action for processing each subtask, and combining the actions corresponding to each subtask into the first action sequence sample.
[0263] In some embodiments, the first training module 1331 is further configured to perform semantic understanding on the task instruction to obtain the task description text of the task; and generate a thought chain corresponding to the task based on the task description text, wherein the thought chain includes multiple sub-tasks, and the multiple sub-tasks are ordered according to a logical progressive relationship.
[0264] In some embodiments, the intelligent agent model is used to interact with the human-computer interaction program; the first training module 1331 is further used to iteratively perform the following processing for each sub-task until a preset condition is met: for the t-th time step, based on the interface image of the human-computer interaction program at the t-th time step and the sub-task, construct the input data for the t-th time step, perform action prediction based on the input data for the t-th time step to obtain the action for the t-th time step, and execute the action for the t-th time step to obtain the interface image of the human-computer interaction program at the (t+1)-th time step, where t≥0.
[0265] In some embodiments, the first training module 1331 is further configured to perform feature encoding on the input data at the t-th time step to obtain encoded features; perform feature mapping on the encoded features to obtain an action probability distribution, wherein the action probability distribution includes the probability value of each action in a preset action set being an action at the t-th time step; and determine the action at the t-th time step based on the action probability distribution.
[0266] The following description continues to illustrate the exemplary structure of the human-computer interaction program processing device 134 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2B As shown, the software modules in the human-computer interaction program processing device 134 stored in the memory 130-2 may include: The data processing module 1341 is used to generate an action sequence based on a third prompt word using a pre-trained agent model, wherein the actions in the action sequence are used to operate a human-computer interaction program to complete a preset task, and the agent model is trained by the agent model training method provided in the embodiments of this application.
[0267] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the intelligent agent model training method described in this application embodiment, or the human-computer interaction program processing method described in this application embodiment.
[0268] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the training method of the intelligent agent model provided in this application embodiment, or execute the processing method of the human-computer interaction program provided in this application embodiment, for example... Figure 3A The training method for the agent model is shown.
[0269] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0270] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0271] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0272] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0273] In summary, by generating and filtering action sequence samples based on the first prompt word in this embodiment, a first target sample set meeting specific conditions is constructed for preliminary training of the agent model. This allows the selected samples to guide the model in establishing a stable basic strategy during the initial training phase. Furthermore, when the training progress meets the switching conditions, the agent model is trained again using a second target sample set generated based on the second prompt word. This phased, progressive training mechanism enables the agent model to adaptively adjust the focus of its learning content according to its training state, promoting the smooth convergence and optimization of model parameters, thereby improving the decision-making accuracy and scene generalization ability of the agent model.
[0274] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for training an intelligent agent model, characterized in that, The method includes: The agent model to be trained generates multiple first action sequence samples based on a preset first prompt word. From the plurality of first action sequence samples, first action sequence samples that meet preset sample selection conditions are selected to form a first target sample set; The agent model is trained based on the first target sample set; If the training progress of the trained agent model meets the training phase switching conditions, then the trained agent model generates multiple second action sequence samples based on a preset second prompt word to form a second target sample set. The agent model is trained based on the second target sample set.
2. The method according to claim 1, characterized in that, The step of selecting first action sequence samples that meet preset sample selection conditions from the plurality of first action sequence samples includes: For each first action sequence sample, a sample evaluation parameter is calculated. First action sequence samples whose sample evaluation parameter is greater than or equal to a preset evaluation parameter threshold are considered as first action sequence samples that meet the sample selection conditions.
3. The method according to claim 2, characterized in that, The intelligent agent model is used to interact with the human-computer interaction program; The calculation of the sample evaluation parameters for the first action sequence sample includes: Obtain the page layout information of the human-computer interaction program when each action in the first action sequence sample is executed; For each action in the first action sequence sample, the target area corresponding to the action is determined based on the operation coordinates of the action and the page layout information corresponding to the action; Based on the description information of the target region corresponding to each action, an execution path description is generated; Calculate the semantic matching degree between the execution path description and the first prompt word, and use the semantic matching degree as the sample evaluation parameter.
4. The method according to claim 2, characterized in that, The calculation of the sample evaluation parameters for the first action sequence sample includes: The longest common action sequence between the first action sequence sample and the reference action sequence is taken as the target action sequence, and the length of the target action sequence is taken as the target length. The reference action sequence is a pre-configured action sequence corresponding to the first prompt word. A first length is determined for the first action sequence sample, and a second length is determined for the reference action sequence, and the maximum value between the first length and the second length is taken as a third length; The ratio between the target length and the third length is calculated to obtain the sample evaluation parameters of the first action sequence sample.
5. The method according to claim 2, characterized in that, The intelligent agent model is used to interact with the human-computer interaction program, and the first prompt word includes task instructions; The calculation of the sample evaluation parameters for the first action sequence sample includes: Obtain the task feedback result after executing the first action sequence sample in the human-computer interaction program; Based on the task feedback results, the matching result between the first action sequence sample and the task instruction is determined; The sample evaluation parameters are determined based on the matching results.
6. The method according to claim 5, characterized in that, The task feedback results include the interface images of the human-computer interaction program; Determining the matching result between the first action sequence sample and the task instruction based on the task feedback result includes: The task instructions are feature-encoded to obtain task features, and the interface images are feature-encoded to obtain interface features; Calculate the semantic similarity between the task features and the interface features, and use the semantic similarity as the matching result; Determining the sample evaluation parameters based on the matching results includes: The semantic similarity is normalized to obtain the sample evaluation parameters.
7. The method according to claim 1, characterized in that, The training progress is determined in the following way: Perform any of the following processes: The training progress is determined based on the cumulative training time steps of the trained agent model. For each first action sequence sample, calculate the sample evaluation parameters of the first action sequence sample; fuse the sample evaluation parameters corresponding to multiple first action sequence samples respectively to obtain fused evaluation parameters, and use the ratio between the fused evaluation parameters and the preset target evaluation parameters as the training progress.
8. The method according to claim 7, characterized in that, The determination of the training progress based on the cumulative training time steps of the trained agent model includes: Obtain the current cumulative training time steps of the agent model, and the preset total training time steps; Calculate the ratio between the cumulative training time steps and the total training time steps; The ratio is determined as the training progress.
9. The method according to claim 1, characterized in that, Before generating multiple second action sequence samples based on a preset second cue word, the method further includes: If the training progress is greater than or equal to a preset training progress threshold, then the training progress of the agent model is determined to meet the training phase switching condition. If the training progress is less than the training progress threshold, it is determined that the training progress of the agent model does not meet the training stage switching condition, and the process of generating multiple first action sequence samples based on a preset first prompt word is initiated.
10. The method according to any one of claims 1 to 9, characterized in that, The first prompt word includes a task instruction; The generation of multiple first action sequence samples based on a preset first prompt word includes: Perform the following process multiple times to obtain the plurality of first action sequence samples: Perform semantic understanding on the first prompt word to obtain at least one subtask corresponding to the task indicated by the task instruction in the first prompt word; The action for processing each of the subtasks is executed, and the actions corresponding to each of the subtasks are combined into the first action sequence sample.
11. The method according to claim 10, characterized in that, The step of semantically understanding the first prompt word to obtain at least one subtask corresponding to the task indicated by the task instruction in the first prompt word includes: Perform semantic understanding on the task instructions to obtain the task description text; A thought chain corresponding to the task is generated based on the task description text, wherein the thought chain includes multiple sub-tasks, and the multiple sub-tasks are ordered according to a logical progressive relationship.
12. The method according to claim 10, characterized in that, The intelligent agent model is used to interact with the human-computer interaction program; The execution of the action for processing each of the sub-tasks includes: For each of the subtasks, the following processing is performed iteratively until a preset condition is met: For the t-th time step, based on the interface image of the human-computer interaction program at the t-th time step and the sub-task, the input data for the t-th time step is constructed. Based on the input data of the t-th time step, action prediction is performed to obtain the action at the t-th time step, and the action at the t-th time step is executed to obtain the interface image of the human-computer interaction program at the (t+1)-th time step, where t≥0.
13. The method according to claim 12, characterized in that, The preset conditions include any one of the following: The action at the t-th time step is the ending action, wherein the ending action is used to indicate that the subtask has been completed; The t-th time step is the preset maximum number of time steps; The descriptive text of the interface image of the human-computer interaction program at the (t+1)th time step indicates that the subtask has been completed.
14. The method according to claim 12, characterized in that, The action prediction based on the input data at the t-th time step to obtain the action at the t-th time step includes: The input data at the t-th time step is feature-encoded to obtain the encoded features; The encoded features are mapped to obtain an action probability distribution, wherein the action probability distribution includes the probability value of each action in the preset action set being the action at the t-th time step; Based on the action probability distribution, the action at the t-th time step is determined.
15. A method for processing human-computer interaction programs, characterized in that, The method includes: An action sequence is generated based on a third prompt word using a pre-trained agent model, wherein the actions in the action sequence are used to operate a human-computer interaction program to complete a preset task, and the agent model is trained by the method described in any one of claims 1 to 14.
16. A training device for an intelligent agent model, characterized in that, The device includes: The first training module is used to generate multiple first action sequence samples based on a preset first prompt word using the agent model to be trained. The first training module is further configured to select first action sequence samples that meet preset sample selection conditions from the plurality of first action sequence samples to form a first target sample set; The first training module is further configured to train the agent model based on the first target sample set; The second training module is used to generate multiple second action sequence samples based on a preset second prompt word, to form a second target sample set, if the training progress of the trained agent model meets the training phase switching conditions. The agent model is trained based on the second target sample set.
17. A processing device for a human-computer interaction program, characterized in that, The device includes: A data processing module is used to generate an action sequence based on a third prompt word using a pre-trained agent model, wherein the actions in the action sequence are used to operate a human-computer interaction program to complete a preset task, and the agent model is trained by the method described in any one of claims 1 to 14.
18. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the method of any one of claims 1 to 14, or implements the method of claim 15.
19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method according to any one of claims 1 to 14, or the method according to claim 15.
20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method of any one of claims 1 to 14, or the method of claim 15.
Citation Information
Cited By
Weld intelligent defect detection model training method, detection method and electronic equipment
CN122199516A