Intelligent agent optimization method and device, equipment, storage medium and product

By recording task execution trajectories and extracting execution experience data within the agent, target training samples are generated, and the target large model of the agent is optimized. This solves the problem of low efficiency in existing technologies and achieves efficient adaptive optimization of the agent.

CN120996185APending Publication Date: 2025-11-21BEIJING QIHOOD TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511029951.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

现有智能体优化方式耗时耗力,难以适应快速变化的任务需求和复杂应用场景,效率低下。

Method used

在智能体内部的目标大模型执行任务过程中,自动记录任务执行轨迹,提取执行经验数据,生成目标训练样本以优化智能体。

Benefits of technology

It significantly improves the efficiency of training sample acquisition, enhances the ability of the agent to cope with rapidly changing task requirements and complex application scenarios, and realizes the continuous adaptive optimization of the agent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996185A_ABST
    Figure CN120996185A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent agent optimization method and device, equipment, a storage medium and a product, relates to the technical field of artificial intelligence, and discloses a method for recording a task execution track of an intelligent agent in a process of executing a task based on a target large model in the intelligent agent; extracting execution empirical data from the task execution track, wherein the execution empirical data comprises problems existing in the task execution process, generation reasons of the problems and solutions; and generating a target training sample based on the execution experience data, the target training sample being used for training the target large model to optimize the agent. According to the method, the efficiency of obtaining the training sample can be remarkably improved, so that the optimization efficiency of the intelligent agent is improved, and the ability of the intelligent agent to deal with rapidly changing task requirements and complex application scenes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, storage medium and product for optimizing intelligent agents. Background Technology

[0002] With the development of artificial intelligence technology, intelligent agents, as systems capable of autonomously perceiving their environment and performing tasks, have been widely used in various fields. However, in practical applications, the optimization process of intelligent agents still faces the bottleneck of insufficient self-evolution capabilities.

[0003] Current agent optimization methods rely on developers collecting task failure cases, manually creating training samples for these cases, and then using these samples to train the agent's internal large model to optimize it. This approach is not only time-consuming and labor-intensive but also extremely inefficient, making it difficult to adapt to rapidly changing task requirements and complex application scenarios. Therefore, a more efficient agent optimization method is urgently needed.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this application is to provide an agent optimization method, apparatus, device, storage medium, and product that can significantly improve the efficiency of acquiring training samples, thereby improving the optimization efficiency of the agent and enhancing the agent's ability to cope with rapidly changing task requirements and complex application scenarios.

[0006] To achieve the above objectives, this application proposes an agent optimization method, the method comprising:

[0007] During the execution of tasks based on the target large model inside the agent, the task execution trajectory of the agent is recorded;

[0008] Execution experience data is extracted from the task execution trajectory, including problems encountered during task execution, the causes of the problems, and solutions.

[0009] Based on the execution experience data, target training samples are generated, which are used to train the target large model to optimize the agent.

[0010] Optionally, extracting execution experience data from the task execution trajectory includes:

[0011] Obtain reflection prompts, which instruct the analysis of the task execution trajectory to generate the execution experience data, and the reflection prompts indicate the target content and target format of the execution experience data;

[0012] By extracting a large model based on experience, the task execution trajectory is analyzed based on the reflection prompts to generate the execution experience data.

[0013] Optionally, obtaining reflection prompts includes:

[0014] If the task result in the task execution trajectory is a failure, an error correction prompt is obtained. This prompt is used to indicate the steps that led to the task failure and to generate a correct solution; or...

[0015] If the task result in the task execution trajectory is successful, an optimization prompt word is obtained. The optimization prompt word is used to indicate the steps that affect the task execution efficiency and generate an optimized solution.

[0016] Optionally, extracting execution experience data from the task execution trajectory includes:

[0017] Activate the optimization engine after the task has finished executing;

[0018] The optimization engine obtains the reflection prompts, and the experience extraction model is invoked to analyze the task execution trajectory based on the reflection prompts in order to generate the execution experience data.

[0019] Optionally, generating target training samples based on the execution experience data includes:

[0020] Extract at least one task description and the target action trajectory corresponding to each task description from the execution experience data;

[0021] Each task description is used as the sample input, and the target action trajectory corresponding to each task description is used as the sample output to generate at least one target training sample. Each target training sample includes the sample input and the corresponding sample output.

[0022] Optionally, the solution in the execution experience data includes the task objective, at least one step, and the processing results of each step;

[0023] The step of extracting at least one task description and the target action trajectory corresponding to each task description from the execution experience data includes:

[0024] The task objective is taken as the first task description, and the first step in the solution is taken as the target action trajectory corresponding to the first task description.

[0025] The task objective, the first step, and the processing result of the first step are used as the second task description, and the second step in the solution is used as the target action trajectory corresponding to the second task description.

[0026] This process continues until the target number of task descriptions and target action trajectories are obtained, where the target number is the number of steps contained within the solution.

[0027] Optionally, generating target training samples based on the execution experience data includes:

[0028] The execution experience data is subjected to quality inspection, which includes at least one of format inspection, scheme validity inspection, scheme performance inspection, and semantic quality inspection.

[0029] If the quality inspection is passed, target training samples are generated based on the execution experience data.

[0030] Optionally, the quality inspection of the execution experience data includes:

[0031] Execute the task in the sandbox environment according to the solutions in the aforementioned experience data;

[0032] If the task is successfully executed, the execution experience data is determined to have passed the validity test of the scheme.

[0033] Optionally, the number of execution experience data extracted from the task execution trajectory is multiple, and the solutions in different execution experience data are different; the quality detection of the execution experience data includes:

[0034] For any solution in the execution experience data, obtain multiple performance evaluation scores:

[0035] Based on the multiple performance evaluation scores, a comprehensive performance score is determined;

[0036] The execution experience data of the solution with the highest overall performance score is determined through the solution performance test.

[0037] Optionally, the performance evaluation score includes at least one of the following:

[0038] A basic score, wherein the value of the basic score is greater when the solution is valid than when the solution is invalid;

[0039] An efficiency score, which is negatively correlated with the number of steps in the solution;

[0040] The cost score is negatively correlated with the number of times external tools are invoked in the solution.

[0041] Optionally, the quality inspection of the execution experience data includes:

[0042] The semantic quality of the execution experience data is evaluated by assessing the large model, and a semantic quality score for the execution experience data is obtained.

[0043] If the semantic quality score is greater than the semantic score threshold, the execution experience data is determined to have passed the semantic quality detection.

[0044] Optionally, after generating the target training samples based on the execution experience data, the method further includes:

[0045] Obtain the training effect parameters of the target training samples on the target large model;

[0046] The model reward parameters are determined based on the training performance parameters.

[0047] Based on the task execution trajectory, the execution experience data, and the model reward parameters, training samples are generated for training the large experience extraction model.

[0048] Optionally, determining the model reward parameters based on the training performance parameters includes:

[0049] If the training performance parameter indicates an improvement in the performance of the target large model, the model reward parameter is set to a positive value; or,

[0050] If the training effect parameters indicate that the performance of the target large model has not improved, the model reward parameter is set to zero or a negative value.

[0051] Optionally, recording the task execution trajectory of the intelligent agent during the execution of a task based on a target large model within the intelligent agent includes:

[0052] During the execution of a task based on the target model within the intelligent agent, the task execution trajectory is recorded by a task trajectory recorder. The task execution trajectory includes the task objective, thought process, action, observation, and task result.

[0053] Furthermore, to achieve the above objectives, this application also proposes an intelligent agent optimization device, the device comprising:

[0054] The trajectory recording module is used to record the task execution trajectory of the intelligent agent during the execution of a task based on a target large model inside the intelligent agent;

[0055] An experience extraction module is used to extract execution experience data from the task execution trajectory. The execution experience data includes problems existing in the task execution process, the causes of the problems, and solutions.

[0056] The sample generation module is used to generate target training samples based on the execution experience data. The target training samples are used to train the target large model to optimize the agent.

[0057] Optionally, the experience extraction module includes:

[0058] The prompt word acquisition unit is used to acquire reflection prompt words, wherein the reflection prompt words indicate the analysis of the task execution trajectory to generate the execution experience data, and the reflection prompt words indicate the target content and target format of the execution experience data;

[0059] The experience generation unit is used to analyze the task execution trajectory based on the reflection prompts by extracting a large model of experience in order to generate the execution experience data.

[0060] Optionally, the prompt word acquisition unit is used to acquire error correction prompt words when the task result in the task execution trajectory is a failure, the error correction prompt words being used to indicate the identification and analysis of the steps that caused the task failure and to generate a correct solution; or, when the task result in the task execution trajectory is a success, acquire optimization prompt words, the optimization prompt words being used to indicate the identification and analysis of the steps that affect the task execution efficiency and to generate an optimized solution.

[0061] Optionally, the experience extraction module is used to activate the optimization engine when the task execution ends; obtain the reflection prompt words through the optimization engine, and call the experience extraction big model to analyze the task execution trajectory based on the reflection prompt words to generate the execution experience data.

[0062] Optionally, the sample generation module includes:

[0063] An information extraction unit is used to extract at least one task description and the target action trajectory corresponding to each task description from the execution experience data.

[0064] The sample generation unit is used to take each task description as sample input and the target action trajectory corresponding to each task description as sample output to generate at least one target training sample. Each target training sample includes the sample input and the corresponding sample output.

[0065] Optionally, the solution in the execution experience data includes the task objective, at least one step, and the processing results of each step;

[0066] The information extraction unit is used to take the task objective as the first task description and the first step in the solution as the target action trajectory corresponding to the first task description; take the task objective, the first step, and the processing result of the first step as the second task description and the second step in the solution as the target action trajectory corresponding to the second task description; and so on, until a target number of task descriptions and target action trajectories are obtained, where the target number is the number of steps contained in the solution.

[0067] Optionally, the sample generation module includes:

[0068] A quality inspection unit is used to perform quality inspection on the execution experience data, and the quality inspection includes at least one of format inspection, scheme validity inspection, scheme performance inspection and semantic quality inspection;

[0069] The sample generation unit is used to generate target training samples based on the execution experience data if the quality detection is passed.

[0070] Optionally, the quality inspection unit is used to execute a task in a sandbox environment according to the solution in the execution experience data; if the task is successfully executed, it determines that the execution experience data has passed the solution validity test.

[0071] Optionally, the number of execution experience data extracted from the task execution trajectory is multiple, and the solutions in different execution experience data are different;

[0072] The quality detection unit is used to obtain multiple performance evaluation scores for any solution in any execution experience data; determine a comprehensive performance score based on the multiple performance evaluation scores; and determine the execution experience data to which the solution with the highest comprehensive performance score belongs passes the solution performance detection.

[0073] Optionally, the performance evaluation score includes at least one of the following:

[0074] A basic score, wherein the value of the basic score is greater when the solution is valid than when the solution is invalid;

[0075] An efficiency score, which is negatively correlated with the number of steps in the solution;

[0076] The cost score is negatively correlated with the number of times external tools are invoked in the solution.

[0077] Optionally, the quality detection unit is used to perform semantic quality assessment on the execution experience data by evaluating a large model to obtain a semantic quality score for the execution experience data; if the semantic quality score is greater than a semantic score threshold, the execution experience data is determined to have passed the semantic quality detection.

[0078] Optionally, the sample generation module is further configured to obtain the training effect parameters of the target training samples on the target large model; determine the model reward parameters based on the training effect parameters; and generate training samples for training the experience extraction large model based on the task execution trajectory, the execution experience data, and the model reward parameters.

[0079] Optionally, the sample generation module is configured to set the model reward parameter to a positive value when the training effect parameter indicates an improvement in the performance of the target large model; or, when the training effect parameter indicates no improvement in the performance of the target large model, set the model reward parameter to zero or a negative value.

[0080] Optionally, the trajectory recording module is used to record the task execution trajectory through a task trajectory recorder during the execution of a task based on the target large model inside the intelligent agent. The task execution trajectory includes the task objective, thinking, actions, observations, and task results.

[0081] Optionally, the device further includes:

[0082] The model training module is used to fine-tune the target large model based on the generated target training samples when the number of generated target training samples reaches a threshold, so as to obtain an optimized agent.

[0083] In addition, to achieve the above objectives, this application also proposes an intelligent agent optimization device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the intelligent agent optimization method as described above.

[0084] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the intelligent agent optimization method described above.

[0085] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the intelligent agent optimization method described above.

[0086] One or more technical solutions proposed in this application have at least the following technical effects:

[0087] The agent optimization scheme provided in this application automatically records the agent's task execution trajectory during task execution based on the target large model within the agent. Key execution experience data is extracted from this trajectory, including problems encountered during task execution, their causes, and corresponding solutions. Based on this experience data, target training samples are automatically generated to optimize the target large model within the agent. This effectively transforms the "immediate experience" acquired by the agent during task execution into long-term "permanent capabilities," thereby enabling continuous adaptive optimization of the agent in practical applications. Compared to traditional methods that rely on manually collecting failure cases and manually writing training samples, this scheme significantly improves the efficiency of obtaining training samples, thereby enhancing the efficiency of model training and agent optimization, reducing reliance on developer experience, and strengthening the agent's ability to cope with rapidly changing task requirements and complex application scenarios. Attached Figure Description

[0088] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0089] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0090] Figure 1 This is a schematic diagram of the structure of an intelligent agent provided in an embodiment of this application;

[0091] Figure 2 This is a flowchart illustrating the first embodiment of the intelligent agent optimization method of this application;

[0092] Figure 3 This is a flowchart illustrating the second embodiment of the intelligent agent optimization method of this application.

[0093] Figure 4 This is a flowchart illustrating the third embodiment of the intelligent agent optimization method of this application;

[0094] Figure 5 This is a flowchart illustrating the fourth embodiment of the intelligent agent optimization method of this application;

[0095] Figure 6 A schematic diagram illustrating an agent optimization process provided in an embodiment of this application;

[0096] Figure 7 A schematic diagram illustrating a sample generation process provided in an embodiment of this application;

[0097] Figure 8 This is a schematic diagram of the module structure of the intelligent agent optimization device according to an embodiment of this application;

[0098] Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the intelligent agent optimization method in the embodiments of this application.

[0099] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0100] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0101] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0102] Figure 1 This is a schematic diagram of the structure of an intelligent agent provided in an embodiment of this application. See also... Figure 1 The intelligent agent comprises a target big model 101, interaction objects 102, a task trajectory recorder 103, and an optimization engine 104. The target big model 101 is the core of the intelligent agent, used for thinking, decision-making, and generating task execution steps. Interaction objects 102 refer to objects that interact with the intelligent agent during task execution, such as web browsers, code interpreters, API (Application Programming Interface) collections, and robot control systems. The task trajectory recorder 103 records each step of the intelligent agent during task execution, generating the task execution trajectory. The optimization engine 104 triggers the self-updating of its internal target big model 101 after the intelligent agent completes the task.

[0103] The agent in this embodiment possesses a high degree of self-evolution capability. This is achieved as follows: during the execution of a task based on the target large model 101 within the agent, the task trajectory recorder 103 records the agent's task execution trajectory. The optimization engine 104 extracts execution experience data from the task execution trajectory and generates target training samples based on this data. Then, the target large model 101 within the agent is trained using these target training samples to optimize the agent.

[0104] The solution provided in this application is applicable to various scenarios. Taking an office scenario as an example: when an agent performs a document organization task, the task trajectory recorder captures the complete execution trajectory in real time. If the agent makes an important file classification error during execution, the optimization engine will deeply analyze the task execution trajectory, thereby clarifying the "misplaced files" problem in the generated execution experience data, locating the root cause of the "fuzzy classification rules," and constructing a refined classification logic system in the solution. Based on the training samples generated from this execution experience data, the decision logic of the target large model can be optimized in a targeted manner. When the agent performs similar tasks subsequently, the task execution steps output by the model will automatically incorporate the enhanced classification strategy, significantly improving the accuracy and standardization of document organization.

[0105] Taking a smart home scenario as an example: When the intelligent agent performs the task of adjusting the air conditioner temperature, the task trajectory recorder captures the complete execution trajectory in real time. If the intelligent agent experiences discomfort due to failure to recognize changes in the number of people in the room, the optimization engine will deeply analyze the task execution trajectory to identify the "temperature adjustment deviation" problem in the generated execution experience data and pinpoint the root cause of the "lack of occupant perception." The solution's temperature adjustment strategy will then incorporate steps to acquire relevant sensor data to determine the real-time number of people and adjust the temperature threshold accordingly. Training samples generated based on this execution experience data can be used to specifically optimize the decision logic of the target large model. When the intelligent agent performs similar tasks subsequently, the task execution steps output by the model will automatically incorporate the enhanced perception strategy, significantly improving the comfort and intelligence of environmental control.

[0106] Of course, the solution provided in this application embodiment can also be applied to other scenarios, such as checking the weather or buying train tickets, and this application embodiment does not limit this.

[0107] Figure 2 This is a flowchart illustrating the first embodiment of the intelligent agent optimization method of this application. (Refer to...) Figure 2 Taking the intelligent agent optimization device as the executing entity as an example, the intelligent agent optimization method includes the following steps S10 to S30:

[0108] Step S10: During the execution of tasks based on the target large model inside the agent, record the task execution trajectory of the agent.

[0109] An intelligent agent is an entity with autonomous decision-making capabilities, such as software or hardware systems, that can perceive its environment and achieve specific goals by performing actions. Examples include self-driving cars and chatbots.

[0110] The target large model is the core large model upon which the intelligent agent relies. It can be a deep learning model with a large number of parameters and strong generalization and reasoning capabilities. Examples include large language models such as GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers), or large neural network models in multimodal domains. It is the core component for the intelligent agent to perform task processing, and the agent relies on its powerful semantic understanding and generative capabilities to execute complex tasks.

[0111] The task execution trajectory records the entire process data of the agent from the start to the end of the task, which may include input data, intermediate decision-making processes, model output, user feedback, system logs, etc. These trajectories record how the agent completes the task step by step.

[0112] For example, the intelligent agent utilizes the ReAct (Reasoning and Acting) framework to reason, break down tasks, and invoke tools to execute them. The ReAct framework is a chain-of-thought framework that enables the intelligent agent to think and act simultaneously, much like a human. Its structure includes thinking, acting, and observing. Thinking is the agent's internal reasoning process; the target model performs logical reasoning based on the current state and historical information to generate task execution steps. Acting involves invoking tools to perform specific operations based on the reasoning results, such as calling APIs or querying databases. Observing involves the results returned by the tools. This entire process repeats cyclically until the task is completed. Correspondingly, the task execution trajectory can include the task objective, each step of thinking, acting, and observing, as well as the final task result. This makes the entire task execution process transparent, providing accurate and complete data support for subsequent extraction of execution experience data.

[0113] For example, the task execution trajectory includes the following:

[0114] Mission objective: Book a flight from Beijing to Shanghai for 2 PM tomorrow.

[0115] Think about it: To book a flight, I must first search for available flights.

[0116] Action: Search for flights (origin = Beijing, destination = Shanghai, date = tomorrow, time = 14:00).

[0117] Observation: Flight CA1234 was found, with a ticket price of 1500 yuan.

[0118] Thinking: There are available tickets for this flight at a reasonable price, so I should book it.

[0119] Action: Purchase a ticket (flight identifier = CA1234).

[0120] Task result: Flight booking failed.

[0121] Optionally, during the execution of tasks based on the large target model within the agent, a task trajectory recorder is used to record the task execution trajectory. A task trajectory recorder is a module or component used to automatically record the agent's behavioral trajectory during task execution. It supports structured or unstructured data formats, such as JSON (JavaScript Object Notation), text logs, etc. Introducing a task trajectory recorder to record the task execution trajectory during agent operation enables in-depth analysis of the task execution process. When a task fails or the output does not meet expectations, the task execution trajectory can be used to quickly identify the erroneous or optimizable steps, facilitating the subsequent extraction of execution experience data.

[0122] Step S20: Extract execution experience data from the task execution trajectory. The execution experience data includes problems that exist during task execution, the causes of the problems, and solutions.

[0123] Execution experience data is key information extracted from the task execution trajectory, and it contains knowledge related to problem discovery and resolution. For example, the problems include: errors, failures, and abnormal behaviors that occur during task execution; the causes include: an analysis of the reasons for these problems, such as model inference errors, environmental interference, and knowledge gaps; and the solutions include: measures taken to resolve these problems, such as revised task execution steps, i.e., revised action sequences.

[0124] Step S30: Based on the execution experience data, generate target training samples. The target training samples are used to train the target large model to optimize the agent.

[0125] Target training samples are training data constructed from execution experience data to optimize a target large model. They may include input examples, expected outputs, error outputs, error type labels, correction suggestions, contextual information, etc. Using target training samples can enhance the model's understanding and ability to cope with specific problems, thereby improving the overall performance of the agent.

[0126] For example, a large model can be constructed using samples to generate target training samples based on execution experience data. For instance, preset sample construction prompts can be used to instruct the model to generate target training samples based on execution experience data, specifying the target content and format of the target training samples. This allows for the automated generation of training samples through a large model constructed from samples. It should be noted that the large model constructed from samples can be a target large model within the agent, meaning the target large model can be used as the sample construction large model. Alternatively, it can be another large model besides the target large model; this embodiment does not limit this.

[0127] Optionally, after generating target training samples based on execution experience data, the method further includes: in response to the number of generated target training samples reaching a threshold, fine-tuning the target large model based on the generated target training samples to obtain an optimized agent.

[0128] The threshold for the number of target training samples can be set according to the needs of the application scenario, for example, it can be set to 100. This application embodiment does not limit this.

[0129] For example, PEFT (Parameter-Efficient Fine-Tuning) techniques such as LoRA (Low-Rank Adaptation) and QLoRA (Quantized Low-Rank Adaptation) can be used to fine-tune the training of a large target model. LoRA adds a low-rank submatrix next to the weight matrix of the original model. During training, only the parameters of the low-rank submatrix are updated, while the original model weights remain frozen. During inference, the parameters of the low-rank submatrix are linearly combined with the original weights to adapt the model's capabilities. This approach reduces the number of parameters that need to be trained, thereby lowering the fine-tuning cost. QLoRA is an improved version of LoRA, combining quantization techniques with low-rank adaptation to further optimize memory usage and training efficiency.

[0130] This application embodiment automates and iteratively stages the model optimization process by setting a threshold for the number of target training samples and triggering fine-tuning training of the target large model when this threshold is reached. On one hand, this mechanism avoids the problem of insufficient sample size affecting training results, ensuring that each fine-tuning is based on sufficient and representative sample data, thus improving the stability and effectiveness of training. On the other hand, the periodic, batch training method reduces the computational resource consumption caused by frequent training, improving system operating efficiency. Furthermore, this design helps the agent continuously accumulate experience and periodically evolve its capabilities in practical applications, thereby better adapting to dynamic task requirements and complex environmental changes, enhancing the agent's autonomous learning and adaptive optimization capabilities.

[0131] The agent optimization scheme provided in this application automatically records the agent's task execution trajectory during task execution based on the target large model within the agent. Key execution experience data is extracted from this trajectory, including problems encountered during task execution, their causes, and corresponding solutions. Based on this experience data, target training samples are automatically generated to optimize the target large model within the agent. This effectively transforms the "immediate experience" acquired by the agent during task execution into long-term "permanent capabilities," thereby enabling continuous adaptive optimization of the agent in practical applications. Compared to traditional methods that rely on manually collecting failure cases and manually writing training samples, this scheme significantly improves the efficiency of obtaining training samples, thereby enhancing the efficiency of model training and agent optimization, reducing reliance on developer experience, and strengthening the agent's ability to cope with rapidly changing task requirements and complex application scenarios.

[0132] Based on the first embodiment described above, a second embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description and will not be repeated hereafter. (Refer to...) Figure 3 In the second embodiment, step S20 includes steps S201 to S202:

[0133] Step S201: Obtain reflection prompts. The reflection prompts indicate the analysis of the task execution trajectory to generate execution experience data, and the reflection prompts indicate the target content and target format of the execution experience data.

[0134] A reflection prompt is a pre-defined instruction text used to guide a large model in a deep analysis and summary of the task execution trajectory. This prompt explicitly indicates the target content of the analysis and the format requirements of the output results, thereby ensuring the consistency and usability of the generated empirical data. The target content in the reflection prompt can include problem identification, root cause analysis, and solutions. The target format can be a structured text format, such as JSON.

[0135] Optionally, the process of obtaining reflection prompts includes: if the task result in the task execution trajectory is failure, obtaining error correction prompts to indicate the steps that led to the task failure and to generate a correct solution; or, if the task result in the task execution trajectory is success, obtaining optimization prompts to indicate the steps that affect the task execution efficiency and to generate an optimized solution.

[0136] Task outcome refers to the final state reached by the agent after completing a task, categorized into "success" and "failure." Error correction prompts are specific prompts used when the task outcome is failure. They guide the experience extraction model to focus on the key steps leading to failure in the task execution trajectory, analyze the causes, and generate corresponding corrective solutions. These prompts emphasize error identification and correction. Optimization prompts are used when the task outcome is success. They guide the experience extraction model to identify potential efficiency bottlenecks or redundant operations during task execution and propose optimization schemes to improve execution efficiency. These prompts focus on performance improvement and process optimization.

[0137] For example, the error correction prompt is: Analyze the following failed task execution trajectory. Identify the key error steps that led to the failure and generate a correct, better sequence of actions. Output in the format of "Problem-Root Cause Analysis-Solution". Task execution trajectory: {Content of task execution trajectory}. The execution experience data generated by the model includes the following:

[0138] Problem: Flight booking failed due to insufficient funds.

[0139] Cause analysis: The failure to check the balance before booking the flight resulted in unnecessary operation failure.

[0140] The solution includes the following:

[0141] {User}: I'm booking a flight from Beijing to Shanghai for 2 PM tomorrow.

[0142] {Model Output}: Thought: To book a flight, I must first search for available flights. Action: Search for flights (origin = Beijing, destination = Shanghai, date = tomorrow, time = 14:00).

[0143] {Observation}: Flight CA1234 was found, with a ticket price of 1500 yuan.

[0144] {Model Output}: Thinking: To buy this ticket, I must first verify my account balance to ensure I have the ability to pay before I can make the booking. Action: Search my account balance and compare it to the ticket price.

[0145] {Model Output}: Thinking: Sufficient balance, book flight now. Action: Purchase ticket (Flight ID = CA1234).

[0146] For example, the optimization prompt might be: Analyze the following successful task execution trajectory. Consider whether there is a simpler, more efficient path to achieve the same goal. Generate a correct and better sequence of actions. Output in the format of "Problem-Cause Analysis-Solution". Task execution trajectory: {Content of task execution trajectory}. The execution experience data generated by the model includes the following:

[0147] Problem: The solution is redundant and inefficient.

[0148] Cause analysis: The API was called twice: first today's weather, then tomorrow's weather.

[0149] The solution includes the following:

[0150] {User}: Check today's and tomorrow's weather and send reminders.

[0151] {Model Output}: Consideration: Directly calling the weather forecast API to obtain two days' worth of weather data at once reduces one API call and is more efficient. Action: Call the weather forecast API with the parameters (City = Beijing, Number of Days = 2).

[0152] {Observation}: Beijing weather: Today, sunny, 22 degrees Celsius; tomorrow, cloudy, 19 degrees Celsius.

[0153] {Model Output}: Consideration: After obtaining the weather forecast, a reminder needs to be sent to the user. Action: Send a reminder to the user (Target = User, Content = "Beijing Weather: Today Sunny, 22 degrees Celsius; Tomorrow Cloudy, 19 degrees Celsius").

[0154] This application's embodiments achieve fine-grained control over the experience extraction process by dynamically selecting different types of prompts—correction prompts or optimization prompts—based on task results. Compared to uniformly using fixed prompts, this method can more specifically guide the experience extraction model to focus on either "failure attribution and repair" or "efficiency improvement of successful paths," thereby generating more valuable experience data. On the one hand, it improves the accuracy and practicality of experience extraction; on the other hand, it also helps to build higher-quality training samples, promoting the simultaneous evolution of the target model in dealing with failure scenarios and optimizing successful processes, further enhancing the agent's adaptability and practical application value.

[0155] Step S202: Through experience extraction of the large model, the task execution trajectory is analyzed based on reflection prompts to generate execution experience data.

[0156] The experience extraction big model refers to a large model specifically designed to process task execution trajectories and generate execution experience data based on reflection prompts. It can be the target big model within the agent, meaning the target big model is used as the experience extraction big model. Alternatively, it can be other big models besides the target big model; this application's embodiments do not impose any limitations on this.

[0157] Optionally, execution experience data can be extracted from the task execution trajectory, including: activating the optimization engine when the task execution is completed; obtaining reflection prompts through the optimization engine; and calling the experience extraction big model to analyze the task execution trajectory based on the reflection prompts to generate execution experience data.

[0158] Task completion refers to the state where the system determines the task flow has terminated after the agent completes a task. This state can be due to successful task completion or task failure caused by an error or exception. It is a key node that triggers subsequent experience extraction and optimization processes.

[0159] The optimization engine is a control module responsible for coordinating and driving the experience extraction and model optimization processes. It determines whether to initiate the experience extraction process based on task execution and calls relevant components, such as the experience extraction large model, to acquire experience. Simultaneously, the optimization engine can also coordinate and drive subsequent training sample generation and fine-tuning of the target large model, thereby achieving adaptive optimization of the agent.

[0160] In this embodiment, by automatically activating the optimization engine after task execution, and having it coordinate the acquisition of reflection prompts and the invocation of the experience extraction model for experience analysis, the automation and closed-loop management of the experience extraction process are achieved. Compared to manual intervention, this mechanism improves the timeliness of system response and processing efficiency, ensuring that effective experience is automatically accumulated after each task. Furthermore, by modularizing the experience extraction process, the system's scalability and flexibility are enhanced, improving the agent's self-evolution capabilities and application adaptability.

[0161] Optionally, the experience extraction model in this embodiment can also be trained and optimized. Specifically, after generating target training samples based on execution experience data, the training effect parameters of the target training samples on the target large model are obtained; the model reward parameters are determined based on the training effect parameters; and training samples for training the experience extraction model are generated based on the task execution trajectory, execution experience data, and model reward parameters. Then, the experience extraction model is fine-tuned based on these training samples, enabling it to continuously optimize its experience extraction strategy and generate higher-quality execution experience data. For example, the behavior of the experience extraction model in extracting execution experience data is itself considered a strategy. The experience extraction model can employ RL (Reinforcement Learning) algorithms such as PPO (Proximal Policy Optimization) and GRPO (Group Relative Policy Optimization) to update the execution experience data extraction strategy, thereby extracting execution experience data that maximizes the reward.

[0162] Training effectiveness parameters refer to a series of quantitative indicators used to evaluate the training effectiveness after training a large target model using target training samples. These indicators include the magnitude of accuracy improvement, changes in task success rate, and improvements in inference efficiency. These parameters reflect the quality and effectiveness of the target training samples. For example, after training a large target model using target training samples, training effectiveness parameters can be determined based on the model's performance changes on a set of benchmark tasks.

[0163] The model reward parameter is a feedback signal determined based on the training effect parameters, used to measure the contribution of the target training samples to the improvement of model performance. As the "reward" in the reward mechanism of reinforcement learning, it guides large models to learn valuable experiences more efficiently.

[0164] This application's embodiments obtain the training effect parameters of the target large model from the target training samples, determine the model reward parameters based on the training effect parameters, and generate training samples for training the experience extraction large model based on the task execution trajectory, execution experience data, and model reward parameters. A closed-loop feedback mechanism is established, which can dynamically evaluate the effectiveness of the training samples and optimize the experience extraction large model itself accordingly. This adaptive training mechanism not only improves the accuracy and practicality of experience extraction but also achieves co-evolution between the experience extraction large model and the target large model.

[0165] Optionally, the model reward parameter is determined based on the training effect parameters, including: setting the model reward parameter to a positive value when the training effect parameters indicate an improvement in the performance of the target large model; or setting the model reward parameter to zero or a negative value when the training effect parameters indicate no improvement in the performance of the target large model. A positive reward parameter is given when the training effect parameters show an improvement in model performance, indicating that the training samples corresponding to the execution experience data have a positive impact on the model, and the strategy for generating this execution experience data should be encouraged and strengthened. A zero or negative reward parameter is given when the training effect parameters show no improvement or even a decrease in model performance, indicating that the sample quality corresponding to the execution experience data is low or misleads the model's learning, and the strategy for generating this execution experience data should be weakened or excluded.

[0166] This application's embodiments achieve quantitative feedback on the value of experience data by dynamically adjusting the model reward parameters based on training results. This mechanism can effectively guide large experience extraction models to extract high-quality experience more accurately in the future, avoiding inefficient or erroneous experience extraction strategies.

[0167] This application embodiment obtains reflection prompts and utilizes a large-scale experience extraction model to analyze the task execution trajectory based on these prompts, generating execution experience data. This achieves intelligent and standardized analysis and experience extraction of the task execution trajectory. Compared to methods relying on manual summarization, this method not only improves the efficiency and consistency of experience extraction but also ensures that the generated execution experience data meets the requirements for subsequent training sample construction in terms of content completeness and format standardization.

[0168] Based on the first embodiment of this application described above, a third embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description, and will not be repeated hereafter. See also... Figure 4 In the third embodiment, step S30 includes steps S301 to S302.

[0169] Step S301: Perform quality checks on the execution experience data. The quality checks include at least one of the following: format checks, scheme validity checks, scheme performance checks, and semantic quality checks.

[0170] Quality inspection refers to the process of evaluating the content quality and usability of execution experience data before it is transformed into training samples. Quality inspection may include at least one of the following: Format inspection: Verifying whether the execution experience data conforms to pre-defined structural specifications. Solution effectiveness inspection: Determining whether the solutions proposed in the execution experience data can truly solve the problems described in the experience data. Solution performance inspection: Evaluating the efficiency and cost of the solutions in the execution experience data in practical applications, such as whether they improve task completion speed or resource utilization, and whether they reduce resource utilization costs. Semantic quality inspection: Checking whether the language expression of the execution experience data is clear, logically sound, and free from ambiguity and errors. It also checks whether the execution experience data accurately identifies the root problems and causes in the task execution trajectory and whether the proposed solutions are likely to succeed.

[0171] Optionally, the implementation method for validating the solution based on the execution experience data is as follows: Execute the task in a sandbox environment according to the solution in the execution experience data; if the task is successfully executed, determine that the execution experience data has passed the solution validity test. Here, the sandbox environment is an isolated, controlled operating environment used to simulate the tasks of the intelligent agent to verify the effectiveness of the solution.

[0172] For example, the action sequence in the solution is extracted, a sandbox environment is launched, and the agent is allowed to strictly follow the action sequence to perform the task in the sandbox environment. If the task is successfully completed, the solution is considered effective, and the execution experience data is determined to pass the solution effectiveness test.

[0173] This application's embodiments achieve automated verification of the effectiveness of solutions by reproducing the solutions contained in the execution experience data in a sandbox environment and observing whether they can successfully complete the corresponding tasks. Compared with relying solely on text analysis or manual judgment, this method has higher objectivity and accuracy, effectively filtering out false, erroneous, or infeasible execution experience data, thereby improving the overall quality of training samples.

[0174] Optionally, the number of execution experience data extracted from the task execution trajectory can be multiple, and the solutions in different execution experience data are different. For example, for a task execution trajectory, multiple experience extraction operations are performed to obtain multiple execution experience data. Alternatively, the reflection prompts instruct the experience extraction model to extract multiple execution experience data at once for a task execution trajectory. Correspondingly, the implementation method for solution performance testing of execution experience data is as follows: for any solution in any execution experience data, obtain multiple performance evaluation scores; based on the multiple performance evaluation scores, determine the comprehensive performance score; determine the execution experience data to which the solution with the highest comprehensive performance score belongs to pass the solution performance test.

[0175] A performance evaluation score is a quantitative assessment of a solution across a specific dimension; a single solution may correspond to multiple performance evaluation scores. The overall performance score is a weighted or combined score derived from these multiple performance evaluation scores, used to measure the overall quality of a solution.

[0176] Optionally, the performance evaluation score includes at least one of the following:

[0177] The base score is a numerical value that indicates a valid solution is greater than that indicating an invalid solution. For example, a valid solution has a base score of 1, while an invalid solution has a base score of -1. It reflects the "feasibility" and "effectiveness" of the solution.

[0178] The efficiency score, which is negatively correlated with the number of steps in a solution, is used to evaluate the solution's "execution speed" and "logical simplicity." For example, the efficiency score is -(number of steps × α). α is a small penalty coefficient for the number of steps, such as 0.01. If the solution has 10 steps, the efficiency score is -0.1.

[0179] The cost score is negatively correlated with the number of times external tools are called in the solution. It reflects the system resource consumption of the solution during execution. External tools can include APIs with call costs or latency. The more times these external tools are called, the greater the system overhead, and the lower the cost score. For example, the cost score is -(API call count × β), where β is a small API call penalty coefficient, such as 0.02. If the solution requires 3 calls to a paid API, the cost score is -0.06. This score reflects the solution's "resource utilization efficiency" and "economic efficiency."

[0180] For example, the overall performance score R = base score + efficiency score + cost score; for instance, in the execution experience data, solution A: task execution successful, containing 5 steps and 2 API calls. R = 1.0 - (5 × 0.01) - (2 × 0.02) = 0.91; in the execution experience data, solution B: task completion successful, containing 12 steps and 5 API calls. R = 1.0 - (12 × 0.01) - (5 × 0.02) = 0.78; in the execution experience data, solution C: task execution failed, containing 8 steps and 3 API calls. R = -1.0 - (8 × 0.01) - (3 × 0.02) = -1.14.

[0181] This application embodiment obtains multi-dimensional performance evaluation scores for different solutions from multiple execution experience datasets and calculates their comprehensive performance score, thereby achieving a systematic evaluation and optimization of solution performance. Compared with single-index judgment, this approach can more comprehensively and objectively reflect the actual application value of each solution, helping to screen out truly efficient and feasible experience data for subsequent training sample construction and model optimization. Furthermore, by introducing performance evaluation scores of multiple dimensions such as basic score, efficiency score, and cost score, it is not only possible to accurately determine whether a solution is effective, but also to further measure its operational efficiency and resource consumption in actual deployment, thereby screening out high-quality experience data that is both effective and efficient, and low-cost.

[0182] Optionally, the semantic quality detection of the execution experience data can be implemented by: evaluating the semantic quality of the execution experience data through a large evaluation model to obtain a semantic quality score for the execution experience data; if the semantic quality score is greater than the semantic score threshold, the execution experience data is determined to have passed the semantic quality detection.

[0183] The evaluation big model is a large model specifically designed for analyzing the semantic quality of text. It can understand the language structure and logical relationships in the execution experience data and provide a semantic quality score accordingly. The evaluation big model can be the target big model, i.e., the target big model is used as the evaluation big model, or it can be other large language models. This application does not limit this.

[0184] For example, the evaluation of a large model can be guided by evaluation prompts to perform semantic quality scoring on execution experience data. For example, the evaluation prompts could be: "You are an agent analyst. Please evaluate the quality of the following execution experience data content. Does it accurately identify the root problem in the task execution trajectory? Is its proposed solution logically sound and likely to succeed? Please rate it from 1 (poor) to 5 (excellent) and return the score and a brief reason in JSON format. Task execution trajectory: {content of the task execution trajectory}, Execution experience data: {content of the execution experience data}."

[0185] It's important to note that the evaluation prompts can differ between "error correction" and "optimization" type execution experience data. For example, for "error correction" type data, the evaluation prompts primarily indicate scoring based on the accuracy of identified error points and causes, and the effectiveness of the solution in correcting the errors. For "optimization" type data, the evaluation prompts primarily indicate scoring based on the accuracy of identified areas for improvement and root cause analysis, and the rationality of the optimization plan. This approach makes the semantic quality scoring more aligned with the actual application scenarios of the execution experience data, thereby enhancing the relevance and practical value of the selected data.

[0186] Semantic quality scores reflect the quality of semantic expression in execution experience data. Higher scores indicate clearer language and more rigorous logic in the execution experience data. A semantic score threshold is a pre-defined scoring standard used to determine whether execution experience data passes the semantic quality check. Only data with scores exceeding this threshold is considered to have sufficient semantic quality to proceed to the next stage of processing.

[0187] This application's embodiments introduce a large-scale evaluation model to score the semantic quality of execution experience data and set a semantic score threshold as a screening criterion, effectively improving the quality control capability of execution experience data in terms of language expression and logical structure. Compared to methods that rely solely on format checks or manual review, this method achieves automated semantic-level evaluation, improving detection efficiency. Furthermore, high-semantic-quality execution experience data helps improve the understanding and generalization ability of training samples, thereby further enhancing the learning effect and decision-making ability of the target large-scale model. This mechanism provides a solid foundation for building stable, efficient, and self-evolving intelligent agent systems, enhancing the system's adaptability and practicality in complex task scenarios.

[0188] Optionally, the format detection content for execution experience data includes: detecting whether the execution experience data is in a valid JSON format; detecting whether it contains predefined key fields, namely the problem, the cause of the problem, and the solution; and detecting whether the field content is empty or an invalid tool name. The format detection mechanism provided in this application embodiment effectively ensures the usability, integrity, and legality of execution experience data by standardizing data structures and content requirements.

[0189] For example, when quality inspection includes format inspection, solution validity inspection, solution performance inspection, and semantic quality inspection, the inspection order can be format inspection first, then solution validity inspection, followed by solution performance inspection and semantic quality inspection. This inspection order realizes a layer-by-layer screening mechanism for execution experience data from "basic usability" to "content value," ensuring that the final selected execution experience data is not only compliant in form, but also reliable in content, clearly expressed, and has practical optimization value.

[0190] In step S302, if the quality check passes, target training samples are generated based on execution experience data. The target training samples are used to train the target large model to optimize the agent.

[0191] This approach introduces a multi-dimensional quality control mechanism before generating target training samples, ensuring that only compliant execution experience data is used to generate training samples for model training. This effectively avoids the negative impact of low-quality, invalid, or even misleading data on the model. This not only improves the overall quality of training samples but also enhances the stability and reliability of model optimization. The quality control mechanism provides a "filter" for the entire experience-driven self-evolutionary process, contributing to the construction of a more efficient, accurate, and controllable agent optimization system.

[0192] Based on the first embodiment of this application described above, a fourth embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description, and will not be repeated hereafter. See also... Figure 5 In the fourth embodiment, step S30 includes steps S303 to S304.

[0193] Step S303: Extract at least one task description and the target action trajectory corresponding to each task description from the execution experience data.

[0194] A task description is a natural language description of the task that the agent needs to complete. It may include information such as task objectives, input conditions, or environmental states. It is used as the "input" part in the training samples so that the model can learn and understand the task requirements.

[0195] The target action trajectory refers to a series of actions taken by the agent during task execution, representing the specific operational path for solving the task. It serves as the "output" part of the training samples, allowing the model to learn how to correctly complete the task.

[0196] Optionally, the solution in the execution experience data includes a task objective, at least one step, and the processing results of each step. Correspondingly, at least one task description and the target action trajectory corresponding to each task description are extracted from the execution experience data, including: using the task objective as the first task description and the first step in the solution as the target action trajectory corresponding to the first task description; using the task objective, the first step, and the processing result of the first step as the second task description and the second step in the solution as the target action trajectory corresponding to the second task description; and so on, until a target number of task descriptions and target action trajectories are obtained, where the target number is the number of steps contained in the solution.

[0197] This solution proposes a progressive task description and action trajectory extraction method. Specifically, the task description is constructed progressively from the solution, adding completed steps and their results at each step, and using the next step to be executed as the target action trajectory. For example, the first step only provides the task objective and requires the model to output the first step; the second step provides the task objective, the first step, and the result of the first step, requiring the model to output the second step; and so on, until all steps are converted into training samples.

[0198] For example, solutions implemented from experience data include the following:

[0199] {User}: I'm booking a flight from Beijing to Shanghai for 2 PM tomorrow.

[0200] {Model Output}: Thought: To book a flight, I must first search for available flights. Action: Search for flights (origin = Beijing, destination = Shanghai, date = tomorrow, time = 14:00).

[0201] {Observation}: Flight CA1234 was found, with a ticket price of 1500 yuan.

[0202] {Model Output}: Thinking: To buy this ticket, I must first verify my account balance to ensure I have the ability to pay before I can make the booking. Action: Search my account balance and compare it to the ticket price.

[0203] {Observation}: Account balance is greater than 1500 yuan.

[0204] {Model Output}: Thinking: Sufficient balance, book flight now. Action: Purchase ticket (Flight ID = CA1234).

[0205] {Observation} Flight booking successful.

[0206] Based on this execution experience data, the following three training samples can be extracted:

[0207] The input for training sample 1 is:

[0208] {User}: I'm booking a flight from Beijing to Shanghai for 2 PM tomorrow.

[0209] The output of training sample 1 is:

[0210] Thinking: To book a flight, I must first search for available flights. Action: Search for flights (origin = Beijing, destination = Shanghai, date = tomorrow, time = 14:00).

[0211] The input for training sample 2 is:

[0212] {User}: Booking a flight from Beijing to Shanghai tomorrow at 2 PM. {Model Output}: Thinking: To book a flight, I must first search for available flights. Action: Search for flights (origin = Beijing, destination = Shanghai, date = tomorrow, time = 2 PM). {Observation}: Found flight CA1234, ticket price 1500 yuan.

[0213] The output of training sample 2 is:

[0214] Thought: To buy this ticket, I must first verify my account balance to ensure I have the ability to pay before I can make the booking. Action: Search my account balance and compare it to the ticket price.

[0215] The input for training sample 3 is:

[0216] {User}: Booking a flight from Beijing to Shanghai tomorrow at 2 PM. {Model Output}: Thought: To book a flight, I must first search for available flights. Action: Search for flights (origin = Beijing, destination = Shanghai, date = tomorrow, time = 2 PM). {Observation}: Found flight CA1234, ticket price 1500 yuan. {Model Output}: Thought: To purchase this ticket, I must first verify my account balance to ensure I have the ability to pay before I can make the booking. Action: Search my account balance and compare it to the ticket price. {Observation}: Account balance is greater than 1500 yuan.

[0217] The output of training sample 3 is:

[0218] Thinking: Sufficient balance, book flight now. Action: Purchase ticket (flight identifier = CA1234).

[0219] Training sample 1 teaches the model to generate the first step (including the first thought and action) upon receiving the initial task. Training sample 2 teaches the model to continue executing the task based on the result of the first step, learning to generate the second step (including the second thought and action). Training sample 3 teaches the model to continue executing the task based on the result of the second step, learning to generate the third step (including the third thought and action).

[0220] It should be noted that some steps in the solution proposed from the original execution experience data may not have processing results. In this case, the processing results of the steps recorded when running the solution in the sandbox environment can be obtained to supplement the content of the solution, making it more complete, so as to generate target training samples.

[0221] This application's embodiments construct task descriptions and target action trajectories in a progressive manner, enabling training samples to more realistically reflect the agent's dynamic reasoning process and behavioral decision-making logic during task execution. Compared to single-input-output sample construction methods, this phased, gradual introduction of historical information helps the model better understand the task background and contextual dependencies, thereby improving its continuous decision-making ability in real-world tasks. Furthermore, this method enhances the model's ability to model multi-step tasks, giving it stronger planning and adaptability when facing complex processes.

[0222] Step S304: Take each task description as sample input and the target action trajectory corresponding to each task description as sample output to generate at least one target training sample. Each target training sample includes sample input and corresponding sample output. The target training samples are used to train the target large model to optimize the agent.

[0223] The target training sample is a data unit consisting of sample input and sample output, used to train the target large model so that it can automatically generate the correct action path based on the task description.

[0224] This application's embodiments automatically extract task descriptions and their corresponding target action trajectories from execution experience data and construct them into target training samples in an "input-output" format. This transforms the agent's behavioral experience in real-world tasks into supervised learning data that can be used for model training. This approach ensures the authenticity and executableness of the training samples, enabling the target large model to master the key paths for task solving through imitation learning, thereby improving its task understanding and decision-making execution capabilities. Simultaneously, the entire sample generation process requires no manual intervention and is highly automated, providing a solid data foundation and implementation path for the agent to achieve experience-driven continuous self-evolution.

[0225] Figure 6 This is a schematic diagram illustrating an agent optimization process provided in an embodiment of this application. (Reference) Figure 5The agent optimization process mainly includes the agent's task execution process and the agent's adaptive optimization process. The task execution process includes: receiving a new task, generating thoughts and actions, calling tools, and obtaining observation results. The processes of generating thoughts and actions, calling tools, and obtaining observation results are executed cyclically until the task result is obtained. A task trajectory recorder records the entire task execution process, generating a task execution trajectory, and storing it in a task trajectory database. After the task is completed, the optimization engine is activated. The optimization engine retrieves the task execution trajectory from the task trajectory database and extracts execution experience data based on the trajectory. Then, the execution experience data undergoes quality checks. If the quality check passes, target training samples are generated based on the execution experience data. Then, the target large model is fine-tuned using the target training samples to obtain the fine-tuned model weights. Finally, the target large model in the agent is updated based on the model weights, thus completing the agent's adaptive optimization.

[0226] Figure 7 This is a schematic diagram illustrating a sample generation process provided in an embodiment of this application. (Reference) Figure 6 First, execution experience data is generated based on the task execution trajectory. Then, it is determined whether the execution experience data passes format detection; if not, it is discarded. If yes, the effectiveness of the solutions in the execution experience data is verified in a sandbox environment, and it is determined whether the solution effectiveness detection passes. If not, the execution experience data is discarded. If yes, it is determined whether semantic quality detection and solution performance detection pass; if not, the execution experience data is discarded. If yes, the selected high-quality execution experience data is converted into target training samples and stored in the fine-tuning dataset. The target training samples in the fine-tuning dataset are used to train the target large model.

[0227] Current mainstream intelligent agents rely primarily on two levels for their knowledge and behavioral patterns when performing tasks:

[0228] First, static model weights: As long-term memory, these weights are solidified once through large-scale pre-training and supervised fine-tuning, and usually remain unchanged throughout the agent's lifecycle. The core capabilities of the model come from this, but it cannot adapt to new or unseen task scenarios or environmental rules.

[0229] Second, the short-term context: As short-term working memory, an agent records the trajectory of its observations, thoughts, and actions during a single task. This "in-context learning" method is temporary; once the task ends, the context is cleared, and valuable successful experiences or lessons learned from failures are lost and cannot be retained as permanent abilities.

[0230] This fragmented architecture leads to the following core problems:

[0231] Experience cannot be internalized: intelligent agents cannot "truly" learn from past mistakes. Even if they are prompted to reflect in context through complex prompting engineering, this reflection cannot be transformed into instinctive behavior for the next task, leading to repeated failures to learn from past mistakes in similar scenarios.

[0232] Poor adaptability and reliance on manual iteration: The evolution of the agent heavily depends on the developer. Developers need to collect a large number of failure cases, manually write high-quality fine-tuning training samples, and then periodically retrain and fine-tune the model over long periods. This process is inefficient and cannot respond in real time to new problems encountered by the agent in real business scenarios.

[0233] Growth bottleneck: As the agent interacts more with the environment, it accumulates more and more "experience". However, due to the lack of an effective mechanism for internalizing experience, its core capabilities have not grown accordingly, resulting in a huge waste of experience.

[0234] Therefore, there is an urgent need for a new technological paradigm to break the static attributes of intelligent agents and establish an automated, low-cost transformation path from "immediate experience" to "permanent capabilities," enabling intelligent agents to achieve adaptive evolution through continuous interaction with the environment, just like biological organisms.

[0235] The solution provided in this application is to construct a closed-loop self-evolving system for an intelligent agent. This allows the model to learn and evolve through practice. Leveraging the model's powerful language understanding and generation capabilities, after completing a task, the model reviews and reflects on its execution trajectory, proactively generating high-quality execution experience data to guide its own improvement. Subsequently, the system transforms this high-quality execution experience data into training samples, and through lightweight fine-tuning or reinforcement learning techniques, "injects" the knowledge it contains into the model weights, completing the internalization of experience. Through this automated closed-loop optimization mechanism of "execution, recording, reflection, refinement, filtering, and internalization," the weights of the core large model in the intelligent agent are continuously and incrementally updated, thereby transforming the valuable experience gained from a one-time task into permanent, generalizable capabilities.

[0236] This plan will bring at least the following beneficial effects:

[0237] Improving model adaptability: For example, the next time the model encounters the "book a flight" task, it will "instinctively" check its balance first, rather than just remembering it in the context. It has learned to adapt to the implicit rule in the environment that "you can't buy it if you don't have enough money".

[0238] Reduced supervision: Humans no longer need to write rules or provide large amounts of supervisory data for every possible failure scenario. An intelligent agent can create high-quality learning data for itself through a single failure, achieving self-supervision and iteration.

[0239] Continuous growth: As the number of tasks performed increases, the agent continuously engages in a "reflection-fine-tuning" cycle, and its knowledge and skills continue to grow like a snowball, becoming more and more "sophisticated" and efficient.

[0240] Another point to note is that the above examples are only for understanding this application and do not constitute a limitation on the intelligent agent optimization method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0241] This application also provides an intelligent agent optimization device, please refer to... Figure 8 The intelligent agent optimization device includes:

[0242] The trajectory recording module 10 is used to record the task execution trajectory of the intelligent agent during the execution of a task based on a target large model inside the intelligent agent;

[0243] The experience extraction module 20 is used to extract execution experience data from the task execution trajectory. The execution experience data includes problems that exist during task execution, the causes of the problems, and solutions.

[0244] The sample generation module 30 is used to generate target training samples based on execution experience data. The target training samples are used to train the target large model to optimize the agent.

[0245] Optionally, the experience extraction module 20 includes:

[0246] The prompt word acquisition unit is used to acquire reflection prompt words. The reflection prompt words indicate the analysis of the task execution trajectory to generate execution experience data, and the reflection prompt words indicate the target content and target format of the execution experience data.

[0247] The experience generation unit is used to extract large models of experience and analyze the task execution trajectory based on reflection prompts to generate execution experience data.

[0248] Optionally, the prompt word acquisition unit is used to acquire error correction prompt words when the task result in the task execution trajectory is failure. The error correction prompt words are used to indicate the steps that led to the task failure and generate a correct solution. Alternatively, when the task result in the task execution trajectory is success, the unit can acquire optimization prompt words. The optimization prompt words are used to indicate the steps that affect the task execution efficiency and generate an optimized solution.

[0249] Optionally, the experience extraction module 20 is used to activate the optimization engine when the task execution is completed; obtain reflection prompts through the optimization engine, and call the experience extraction big model to analyze the task execution trajectory based on the reflection prompts to generate execution experience data.

[0250] Optionally, the sample generation module 30 includes:

[0251] The information extraction unit is used to extract at least one task description and the target action trajectory corresponding to each task description from the execution experience data.

[0252] The sample generation unit is used to take each task description as sample input and the target action trajectory corresponding to each task description as sample output to generate at least one target training sample. Each target training sample includes sample input and corresponding sample output.

[0253] Optionally, the solution in the execution experience data includes the task objective, at least one step, and the processing results of each step;

[0254] The information extraction unit is used to take the task objective as the first task description and the first step in the solution as the target action trajectory corresponding to the first task description; take the task objective, the first step, and the processing result of the first step as the second task description and the second step in the solution as the target action trajectory corresponding to the second task description; and so on, until the task descriptions and target action trajectories of the target number are obtained, where the target number is the number of steps contained in the solution.

[0255] Optionally, the sample generation module 30 includes:

[0256] The quality inspection unit is used to perform quality inspection on the execution experience data. The quality inspection includes at least one of the following: format inspection, scheme validity inspection, scheme performance inspection, and semantic quality inspection.

[0257] The sample generation unit is used to generate target training samples based on execution experience data, provided that the quality check passes.

[0258] Optionally, the quality inspection unit is used to execute a task in a sandbox environment according to the solution in the execution experience data; if the task is successfully executed, it determines that the execution experience data has passed the solution validity test.

[0259] Optionally, the number of execution experience data extracted from the task execution trajectory is multiple, and the solutions in different execution experience data are different;

[0260] The quality inspection unit is used to obtain multiple performance evaluation scores for a solution in any execution experience data: based on the multiple performance evaluation scores, a comprehensive performance score is determined; the execution experience data to which the solution with the highest comprehensive performance score belongs passes the solution performance inspection.

[0261] Optionally, the performance evaluation score includes at least one of the following:

[0262] The basic score is greater when the solution is valid than when the solution is invalid.

[0263] Efficiency score, which is negatively correlated with the number of steps in the solution;

[0264] Cost score, which is negatively correlated with the number of times external tools are called in the solution.

[0265] Optionally, the quality detection unit is used to perform semantic quality assessment on the execution experience data by evaluating the large model to obtain a semantic quality score for the execution experience data; if the semantic quality score is greater than the semantic score threshold, the execution experience data is determined to have passed the semantic quality detection.

[0266] Optionally, the sample generation module 30 is also used to obtain the training effect parameters of the target training samples on the target large model; determine the model reward parameters based on the training effect parameters; and generate training samples for training experience extraction of the large model based on the task execution trajectory, execution experience data and model reward parameters.

[0267] Optionally, the sample generation module 30 is used to set the model reward parameter to a positive value when the training effect parameter indicates that the performance of the target large model has improved; or, when the training effect parameter indicates that the performance of the target large model has not improved, to set the model reward parameter to zero or a negative value.

[0268] Optionally, the trajectory recording module 10 is used to record the task execution trajectory through a task trajectory recorder during the execution of a task based on a target large model within the intelligent agent. The task execution trajectory includes the task objective, thoughts, actions, observations, and task results.

[0269] Optionally, the device further includes:

[0270] The model training module is used to fine-tune the target large model based on the generated target training samples when the number of generated target training samples reaches a threshold, so as to obtain the optimized agent.

[0271] The agent optimization device provided in this application, employing the agent optimization method in the above embodiments, can solve the technical problems of agent optimization methods in related technologies being not only time-consuming and labor-intensive but also extremely inefficient, making it difficult to adapt to rapidly changing task requirements and complex application scenarios. Compared with the prior art, the beneficial effects of the agent optimization device provided in this application are the same as those of the agent optimization method provided in the above embodiments, and other technical features in the agent optimization device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0272] This application provides an intelligent agent optimization device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the intelligent agent optimization method in the above embodiments.

[0273] The following is for reference. Figure 9 The diagram illustrates a structural schematic suitable for implementing the intelligent agent optimization device in the embodiments of this application. The intelligent agent optimization device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The intelligent agent optimization device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0274] like Figure 9 As shown, the agent optimization device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the agent optimization device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the agent optimization device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows agent optimization devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0275] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0276] The agent optimization device provided in this application, employing the agent optimization method described in the above embodiments, can solve the technical problems of agent optimization methods in related technologies, which are not only time-consuming and labor-intensive but also extremely inefficient, making them difficult to adapt to rapidly changing task requirements and complex application scenarios. Compared with the prior art, the beneficial effects of the agent optimization device provided in this application are the same as those of the agent optimization method provided in the above embodiments, and other technical features of this agent optimization device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0277] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0278] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0279] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the agent optimization method in the above embodiments.

[0280] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0281] The aforementioned computer-readable storage medium may be included in the agent optimization device; or it may exist independently and not assembled into the agent optimization device.

[0282] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by the agent optimization device, the agent optimization device: records the task execution trajectory of the agent during the execution of a task based on the target large model inside the agent; extracts execution experience data from the task execution trajectory, the execution experience data including problems existing in the task execution process, the causes of the problems, and solutions; and generates target training samples based on the execution experience data, the target training samples being used to train the target large model to optimize the agent.

[0283] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0284] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0285] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0286] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described agent optimization method. This solves the technical problem that agent optimization methods in related technologies are not only time-consuming and labor-intensive but also extremely inefficient, making them difficult to adapt to rapidly changing task requirements and complex application scenarios. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the agent optimization method provided in the above embodiments, and will not be repeated here.

[0287] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the intelligent agent optimization method described above.

[0288] The computer program product provided in this application solves the technical problem that agent optimization methods in related technologies are not only time-consuming and labor-intensive, but also extremely inefficient, making it difficult to adapt to rapidly changing task requirements and complex application scenarios. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the agent optimization method provided in the above embodiments, and will not be repeated here.

[0289] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for optimizing an intelligent agent, characterized in that, The method includes: During the execution of tasks based on the target large model inside the agent, the task execution trajectory of the agent is recorded; Execution experience data is extracted from the task execution trajectory, including problems encountered during task execution, the causes of the problems, and solutions. Based on the execution experience data, target training samples are generated, which are used to train the target large model to optimize the agent.

2. The method as described in claim 1, characterized in that, The extraction of execution experience data from the task execution trajectory includes: Obtain reflection prompts, which instruct the analysis of the task execution trajectory to generate the execution experience data, and the reflection prompts indicate the target content and target format of the execution experience data; By extracting a large model based on experience, the task execution trajectory is analyzed based on the reflection prompts to generate the execution experience data.

3. The method as described in claim 2, characterized in that, The process of obtaining reflection prompts includes: If the task result in the task execution trajectory is a failure, an error correction prompt is obtained. This prompt is used to indicate the steps that led to the task failure and to generate a correct solution; or... If the task result in the task execution trajectory is successful, an optimization prompt word is obtained. The optimization prompt word is used to identify and analyze the steps that affect the task execution efficiency and generate an optimized solution.

4. The method as described in claim 2, characterized in that, The extraction of execution experience data from the task execution trajectory includes: Activate the optimization engine after the task has finished executing; The optimization engine obtains the reflection prompts, and the experience extraction model is invoked to analyze the task execution trajectory based on the reflection prompts in order to generate the execution experience data.

5. The method as described in claim 1, characterized in that, The step of generating target training samples based on the execution experience data includes: Extract at least one task description and the target action trajectory corresponding to each task description from the execution experience data; Each task description is used as the sample input, and the target action trajectory corresponding to each task description is used as the sample output to generate at least one target training sample. Each target training sample includes the sample input and the corresponding sample output.

6. The method as described in claim 5, characterized in that, The solutions in the execution experience data include the task objective, at least one step, and the processing results of each step; The step of extracting at least one task description and the target action trajectory corresponding to each task description from the execution experience data includes: The task objective is taken as the first task description, and the first step in the solution is taken as the target action trajectory corresponding to the first task description. The task objective, the first step, and the processing result of the first step are used as the second task description, and the second step in the solution is used as the target action trajectory corresponding to the second task description. This process continues until the target number of task descriptions and target action trajectories are obtained, where the target number is the number of steps contained within the solution.

7. An intelligent agent optimization device, characterized in that, The device includes: The trajectory recording module is used to record the task execution trajectory of the intelligent agent during the execution of a task based on a target large model inside the intelligent agent; An experience extraction module is used to extract execution experience data from the task execution trajectory. The execution experience data includes problems existing in the task execution process, the causes of the problems, and solutions. The sample generation module is used to generate target training samples based on the execution experience data. The target training samples are used to train the target large model to optimize the agent.

8. An intelligent agent optimization device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the agent optimization method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the agent optimization method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the agent optimization method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Automatic reward function design system and method based on large language model

    CN118036675A

  • Large model agent knowledge enhancement method and device based on self-awareness

    CN120218172A

Cited By

  • Method and device for intelligent computing cloud platform to realize agent self-evolution through computing power

    CN121614114A

  • Intelligent agent self-error correction method and system based on track comparison and positioning

    CN122173330A

  • An agent self-correction method and system based on trajectory comparison and positioning

    CN122173330B