Method and device for evaluating intelligent agent, equipment and medium
By sending interactive guidance information to the agent and obtaining hierarchical tracking data, the problem of inaccurate agent evaluation in the prior art is solved, and fine-grained evaluation of the agent's internal execution process is realized, improving the accuracy and transparency of the evaluation.
Patent Information
- Application Number
- CN202511096458.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies are insufficient for a comprehensive and automated assessment of the internal execution processes and performance of intelligent agents, resulting in inaccurate and opaque assessment results and making it difficult to identify and locate specific defects in intelligent agents.
By sending interactive guidance information to the agent, hierarchical structure tracking data is obtained during the execution of the test task. Based on this data, specific parts are selected for evaluation to obtain the evaluation results of the agent's internal execution process.
It enables fine-grained end-to-end evaluation of agents, enhances the interpretability and transparency of evaluation results, and can more accurately discover and locate the defects of agents.
Smart Images

Figure CN120930798A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a method, apparatus, electronic device, computer program product, and computer-readable storage medium for evaluating intelligent agents. Background Technology
[0002] In recent years, breakthroughs in artificial intelligence technology, particularly in Large Language Models (LLMs), have greatly propelled the development and application of agent technology (also known as intelligent agent technology). An intelligent agent, as an AI system capable of perceiving its environment, making autonomous decisions, and invoking tools to perform complex tasks, typically includes core components such as Large Language Models (LLMs), tool calling capabilities, and a domain-specific knowledge base. Thanks to its high degree of autonomy and task processing capabilities, agent technology has demonstrated broad application prospects in numerous fields, including customer service, data analysis, automated processes, and scientific research assistance, and continues to expand into deeper and more complex application scenarios. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the subsequent detailed description section. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] At least one embodiment of this disclosure provides a method for evaluating an intelligent agent, comprising: sending first interactive guidance information to the intelligent agent, wherein the first interactive guidance information is used to guide the intelligent agent to perform a first test task; obtaining tracking data generated by the intelligent agent during the execution of the first test task, wherein the tracking data has a hierarchical structure; selecting a first data portion from the tracking data based at least on the hierarchical structure of the tracking data; and obtaining an evaluation result of the internal execution process of the intelligent agent performing the first test task based on the first data portion.
[0005] At least one embodiment of this disclosure provides a computer-readable storage medium having instructions stored thereon that, when executed by a processor, cause the processor to perform the method described above.
[0006] At least one embodiment of this disclosure provides a computer program product, including a computer program or instructions, wherein the computer program or instructions, when executed by a processor, implement the method described above.
[0007] At least one embodiment of this disclosure provides an electronic device, including: one or more processors; and one or more memories storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform the method as described above.
[0008] At least one embodiment of this disclosure provides an apparatus for evaluating an intelligent agent, comprising: a sending module configured to send first interactive guidance information to the intelligent agent, wherein the first interactive guidance information is used to guide the intelligent agent to perform a first test task; an obtaining module configured to obtain tracking data generated by the intelligent agent during the execution of the first test task, wherein the tracking data has a hierarchical structure; a selection module configured to select a first data portion from the tracking data, at least based on the hierarchical structure of the tracking data; and an evaluation module configured to obtain an evaluation result of the internal execution process of the intelligent agent performing the first test task based on the first data portion. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0010] Figure 1 This illustration schematically depicts an application scenario of the method and apparatus for evaluating intelligent agents provided in at least one embodiment of this disclosure.
[0011] Figure 2 A flowchart illustrating a method for evaluating an intelligent agent provided in at least one embodiment of this disclosure is shown schematically.
[0012] Figure 3 The illustration schematically depicts a scenario of client-agent interaction for evaluating an agent, provided by at least one embodiment of this disclosure.
[0013] Figure 4 This schematic diagram illustrates a structural block diagram of an apparatus for evaluating an intelligent agent provided in at least one embodiment of the present disclosure;
[0014] Figure 5 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0016] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0017] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0019] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0021] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0022] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and their authorization should be obtained. Relevant users may include any type of rights holder, such as individuals, enterprises, or groups.
[0023] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information, thereby enabling the relevant user to choose whether to provide information to the software or hardware such as the electronic device, application, server, or storage medium that performs the operation of the technical solution disclosed herein based on the prompt message.
[0024] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.
[0025] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0026] With the development of computer and electronic technologies, artificial intelligence (AI) technology has rapidly developed and been widely applied. Taking AI models with a large number of parameters (also known as large AI models, or simply "large models") constructed by AI networks as an example, AI models have been widely used in search engines, intelligent agents, related vertical industries, and basic disciplines, promoting the intelligent development of various industries. For example, the accuracy of content generated by AI models can be improved by combining knowledge bases with AI models. It should be noted that in this embodiment, the AI model includes any one or a combination of multiple types of large language models, large visual models, large audio models, and multimodal large models.
[0027] With the development of artificial intelligence models, AI-driven intelligent agents (also known as AI agents or AIAgents) have been widely applied. Especially with the emergence of AI models, various industries have begun deploying AI model-driven intelligent agents (also known as AI model agents) to accelerate output. For example, the content generated by AI model agents can be constrained to ensure that they output desired content. AI model agents are typically able to perceive information in their environment, make decisions, and take actions to achieve specific goals or tasks.
[0028] An intelligent agent, as an artificial intelligence system capable of perceiving the environment, making autonomous decisions, and calling tools to perform complex tasks, typically includes key components such as a large language model (LLM), tool calling capabilities, and a domain-specific knowledge base.
[0029] An LLM (Language Learning Model) can be the core inference engine of an intelligent agent, acquiring general language understanding and generation capabilities through pre-training on massive amounts of text data. For example, an LLM can undertake key functions such as task parsing, logical reasoning, and decision planning within an intelligent agent. Tool invocation can be a key mechanism for expanding the capabilities of an intelligent agent. For example, through predefined application programming interfaces (APIs), an intelligent agent can actively invoke external tools (e.g., calculators, databases, specialized software, or application data interfaces (APIs) to perform operations that an LLM cannot directly execute (e.g., real-time data queries, complex calculations). A knowledge base can be a dedicated information storage and retrieval system for an intelligent agent, storing structured / unstructured domain knowledge (e.g., product manuals, industry rules, private data). For example, through vectorized embedding and similarity matching techniques, an intelligent agent can dynamically retrieve precise task-related information, supplementing the static knowledge deficiencies of an LLM.
[0030] During user interaction with an AI model or agent, it is necessary to define the role, capabilities, rules, and behavioral boundaries of the AI model or agent through interactive guidance information (or sequences of interactive instructions, sequences of behavioral guidance parameters, sequences of model control instructions, etc.), thereby constraining the content generated by the AI model or agent. Interactive guidance information (or sequences of interactive instructions, sequences of behavioral guidance parameters, sequences of model control instructions) may include prompts. In the following text, interactive guidance information, sequences of interactive instructions, sequences of behavioral guidance parameters, and sequences of model control instructions are used interchangeably.
[0031] As the application scenarios of intelligent agents become increasingly diversified, the complexity of tasks increases significantly, and the interaction logic of their internal components (e.g., LLM, tool calls, knowledge bases) becomes increasingly complex, it is necessary to effectively evaluate the performance and quality of intelligent agents and adjust or improve them based on the evaluation results.
[0032] The performance of intelligent agents can be evaluated through manual assessment. For example, professional evaluators can subjectively judge the agent's output based on preset scoring criteria (e.g., task completion, answer accuracy, logical coherence). While manual assessment can capture some subtle differences at the semantic level, its inherent limitations severely restrict its application effectiveness. For instance, manual assessment methods are highly subjective, inefficient, difficult to apply on a large scale, and struggle to evaluate the quality of the agent's internal execution processes.
[0033] Furthermore, end-to-end automated evaluation methods can be used to assess the performance of the agent. For example, the agent's input and final output can be directly evaluated automatically. Performance evaluation results can be obtained by analyzing the agent's input and final output using pre-defined rules or LLM-based evaluation methods. However, such methods have a single evaluation dimension, primarily focusing on the agent's final output and neglecting the performance evaluation of the agent's internal execution process. Moreover, because such methods ignore the agent's internal execution process, it is difficult to discover and locate specific defects in the agent.
[0034] Therefore, the methods of evaluating the performance of intelligent agents using manual evaluation or end-to-end automatic evaluation have significant shortcomings in terms of cost, efficiency, and comprehensiveness of evaluation results. In this regard, this disclosure aims to provide a technical solution that can achieve comprehensive and automated evaluation of intelligent agents and guide them to discover problems and carry out continuous optimization.
[0035] To at least partially solve or alleviate at least one of the aforementioned technical problems, at least one embodiment of this disclosure provides a method for evaluating an intelligent agent, comprising: sending first interactive guidance information to the intelligent agent, wherein the first interactive guidance information is used to guide the intelligent agent to execute a first test task; obtaining trace data generated by the intelligent agent during the execution of the first test task, wherein the trace data has a hierarchical structure; selecting a first data portion from the trace data based at least on the hierarchical structure of the trace data; and obtaining an evaluation result of the internal execution process of the intelligent agent executing the first test task based on the first data portion.
[0036] Based on the method for evaluating an intelligent agent provided in at least one embodiment of this disclosure, at least one embodiment of this disclosure also provides an apparatus, electronic device, computer program product, and computer-readable storage medium for the method of evaluating an intelligent agent.
[0037] According to at least one embodiment of the present disclosure, a method, apparatus, electronic device, computer program product, and computer-readable storage medium for evaluating an intelligent agent can assess the complex execution chain within the intelligent agent by obtaining and filtering tracking data during the agent's task execution. By performing fine-grained end-to-end evaluation of the intelligent agent, its performance can be evaluated more comprehensively. Furthermore, the above method can enhance the interpretability and transparency of the evaluation results, facilitating the discovery and localization of defects in the intelligent agent.
[0038] The solutions in this disclosure can be applied to products used to evaluate intelligent agents or other types of evaluation products, or they can be a functional module of an intelligent agent or artificial intelligence model. It is understood that before using the technical solutions disclosed in the embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and their authorization should be obtained. Relevant users include, but are not limited to: the intelligent agent or artificial intelligence model business entity and its users, as well as other possible related parties.
[0039] The embodiments and some examples of this disclosure will now be described in detail with reference to the accompanying drawings.
[0040] Figure 1 The illustration shows an application scenario of the method and apparatus for evaluating the interaction process of an intelligent agent provided in at least one embodiment of the present disclosure.
[0041] like Figure 1 As shown, the application scenario 100 of this embodiment includes a user 101 and a client device 102. The user 101 may include, for example, a user of the intelligent agent, an evaluator of the intelligent agent, or other relevant parties. The client device 102 may include, for example, any electronic device capable of providing services to the user based on the intelligent agent, evaluating the performance of the intelligent agent, and providing an interactive page for user operation and use. Examples include mobile phones, tablets, portable computers, desktop computers, smart wearable devices, smart home appliances, or smart vehicle terminals, etc. The embodiments disclosed herein do not limit this.
[0042] For example, in response to user 101's interaction with the interactive page, client device 102 can obtain the interaction guidance information provided by the interaction, and based on the interaction guidance information, obtain the response content corresponding to the interaction guidance information, and output the response content to user 101. For example, the interaction guidance information may include information entered on the interactive page, or it may include information obtained by converting speech. For example, the interaction guidance information may instruct a test task to be performed by an intelligent agent, and the intelligent agent may perform the corresponding test task based on the interaction guidance information and return the execution result of the test task to user 101.
[0043] For example, the client device 102 runs an operating system, which may install instant messaging clients, search clients, artificial intelligence assistants, and other clients. These clients may be clients used to execute the methods according to the embodiments of this disclosure. These clients may include at least a target application capable of acquiring response content generated by an intelligent agent in response to interactive guidance information, or a target application integrating a mini-program capable of acquiring response content generated by an intelligent agent in response to interactive guidance information, or a web page integrating the function of acquiring response content generated by an artificial intelligence model in response to interactive guidance information. The embodiments of this disclosure do not limit this.
[0044] For example, in some embodiments, the client device 102 may have a server locally deployed to support the operation of the intelligent agent, so that the client on the terminal device can send instructions / requests to the server, so that the server responds to the instructions / requests by using the intelligent agent to generate response content with interactive guidance information, and sends the response content to the client.
[0045] In at least one embodiment of this disclosure, the client device 102 can obtain tracking data generated by the agent used by the server in the client device 102 during the execution of a corresponding test task. The agent's tracking data can refer to complete, time-series, and detailed log information recording the agent's interactions with the external environment (including user 101, other agents, tools, knowledge bases, etc.) and changes in its internal decision-making state during the execution of the test task. The agent's tracking data can include the specific processing path of the interaction guidance information flow through the agent's internal processing and the intermediate results generated when the agent executes the test task. For example, a software development kit (SDK) can be deployed in the server of the client device 102 to collect or obtain the agent's tracking data.
[0046] In other embodiments, such as Figure 1 As shown, application scenario 100 may also involve a server 103, which can communicate with the client device 102 via wired or wireless communication links. For example, it can be a local area network server, a wide area network server, or a cloud server. For instance, the server 103 can run a background server that supports the client installed on the client device 102. The client on the terminal device can respond to interactive operations by sending instructions / requests to the server on the server 103, or by sending or receiving data, so that the server on the server can respond to the instructions / requests by using an intelligent agent to generate response content corresponding to the interactive guidance information, and send the response content to the client on the client device 102. For example, the interactive guidance information can instruct a test task to be executed by an intelligent agent. The intelligent agent can execute the corresponding test task based on the interactive guidance information and return the execution result of the test task to the user 101.
[0047] In at least one embodiment of this disclosure, the application scenario 100 may further include a server on server 103, providing interactive guidance information to the intelligent agent used by the server, enabling the intelligent agent to generate response content based on the test task indicated in the interactive guidance information. In at least one embodiment of this disclosure, client device 102 can obtain tracking data generated by the intelligent agent used by the server on server 103 during the execution of the corresponding test task. For example, an SDK can be deployed on the server on server 103 to collect or obtain the tracking data of the intelligent agent.
[0048] For example, the method for evaluating an intelligent agent provided in at least one embodiment of this disclosure can be implemented in software, hardware, firmware, or any combination thereof.
[0049] For example, the method for evaluating an intelligent agent provided in at least one embodiment of this disclosure is applicable to a client device 102 or a server 103, which can load and execute the method for evaluating an intelligent agent. The embodiments of this disclosure do not limit this.
[0050] For example, the client device 102 may include a central processing unit (CPU) or graphics processing unit (GPU), digital signal processor (DSP), neural network processing unit (NPU), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, storage units, etc. The client device 102 is also equipped with an operating system, application programming interfaces (APIs) (e.g., OpenGL (Open Graphics Library), Metal, etc.), etc., and implements the method for evaluating intelligent agents provided in the embodiments of this disclosure by running code or instructions.
[0051] For example, the client device 102 may also include an output component, such as a display component, which may be a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a quantum dot light-emitting diode (QLED) display, or a light-emitting diode display (e.g., a micro-LED display), etc. The embodiments disclosed herein are not limited in this regard. For example, the display component may display an interactive page, including a configuration page for configuring the evaluation process for the agent, and a result presentation page displaying the evaluation results of the agent.
[0052] The following will combine Figures 2-3 A method for evaluating an intelligent agent, provided in at least one embodiment of the present disclosure, will be described in detail.
[0053] Figure 2 A flowchart illustrating a method for evaluating an intelligent agent provided in at least one embodiment of the present disclosure is shown.
[0054] like Figure 2 As shown, the method 200 for evaluating an intelligent agent in this embodiment includes steps S210 to S240. For example, the executing entity of this method for evaluating an intelligent agent can be an electronic device with a corresponding client deployed, or any electronic device communicating with the client; the embodiments of this disclosure do not limit this.
[0055] In step S210, a first interactive guidance message is sent to the agent, wherein the first interactive guidance message is used to guide the agent to perform a first test task.
[0056] According to embodiments of this disclosure, the first interactive guidance information may be information obtained in response to a user's interactive operation, such as multimedia information in text, image, and / or audio form. The following description uses a text prompt as an example of the first interactive guidance information being in text form. This first interactive guidance information may be a prompt (also called a Prompt) generated by a guidance model for a specific output; it may be a question, a piece of text, or a formatted instruction.
[0057] According to embodiments of this disclosure, first interactive guidance information can guide an agent to perform a first test task. The first test task can be configured by a user based on testing needs (e.g., evaluation metrics for the agent). For example, the first interactive guidance information can be an immediate request or question entered by the user during the test, used to convey a specific first test task. The first interactive guidance information can be dynamic, different in each round of dialogue, and can be used, for example, to "tell" the agent the current test task. As another example, the first interactive guidance information can be selected by the user from a predetermined test set based on testing needs (e.g., evaluation metrics for the agent). The test set can include one or more test tasks testing different evaluation metrics or performance of the agent, and each of the one or more test tasks can correspond to one or more first interactive guidance messages. By selecting the appropriate test task, the corresponding first interactive guidance information can be sent to the agent.
[0058] In step S220, tracking data generated by the agent during the execution of the first test task is obtained, wherein the tracking data has a hierarchical structure.
[0059] According to embodiments of this disclosure, an SDK can be deployed on the server using the intelligent agent. For example, tracking data of the intelligent agent can be collected or obtained through the corresponding SDK. This disclosure is not limited thereto; other methods of collecting tracking data are also possible. As mentioned above, the tracking data of the intelligent agent can refer to complete, time-series, and detailed log information recording the interaction between the intelligent agent and the external environment and changes in its internal decision-making state during the execution of a test task. The tracking data can reflect not only the input and output of the intelligent agent but also the internal execution chain or internal execution process of the intelligent agent.
[0060] According to embodiments of this disclosure, various tracking data can be collected via an SDK during the execution of a first test task by an agent. This tracking data can have a hierarchical structure. For example, the tracking data may include tracking data indicating the agent's inputs and outputs, tracking data indicating the agent's tool selection, tracking data indicating the agent's tool invocation process, tracking data indicating the agent's tool return results, and tracking data indicating the model's generated results. At least some or all of the tracking data indicating the agent's inputs and outputs, the tracking data indicating the agent's tool selection, the tracking data indicating the agent's tool invocation process, the tracking data indicating the agent's tool return results, and the tracking data indicating the model's generated results can each have different levels.
[0061] According to embodiments of this disclosure, the tracking data indicating the input and output of the intelligent agent can reflect the first interactive guidance information received by the intelligent agent and the response content generated by the first test task indicated by the first interactive guidance information. The final result generated by the intelligent agent in performing the first test task can be obtained through the tracking data indicating the input and output of the intelligent agent.
[0062] According to embodiments of this disclosure, the tracking data instructing the agent to select tools can reflect the process by which the agent identifies and selects the most suitable tool from the available tool set to solve the first test task or a sub-test task derived from the first test task, based on first interactive guidance information or a first test task. For example, the agent can select a suitable application as the tool to be invoked based on the first test task or a sub-test task, but this disclosure is not limited thereto.
[0063] According to embodiments of this disclosure, the tracking data indicating the tool invocation process of the agent can reflect the invocation parameters used by the agent when invoking the corresponding tool. For example, after the agent selects a tool, the agent can generate corresponding invocation parameters to trigger the invoked tool.
[0064] According to embodiments of this disclosure, tracking data of the results returned by the tool instructing the agent can reflect the corresponding results received by the agent from the invoked tool. For example, the tool can perform a corresponding task based on the invocation parameters and generate and send the tool return results to the agent.
[0065] According to embodiments of this disclosure, the tracking data indicating the model's generation results can reflect the results generated by the LLM model configured in the agent. For example, during the agent's execution of a first test task, the LLM may generate intermediate results once or multiple times as needed.
[0066] In step S230, a first data portion is selected from the tracking data, based at least on the hierarchical structure of the tracking data.
[0067] As described above, tracking data can have a hierarchical structure. For example, a first data portion can be selected from the tracking data based on its hierarchical structure, targeting the agent's evaluation metrics. For instance, the first data portion may include at least one of the following: tracking data indicating the agent's inputs and outputs; tracking data indicating the agent's tool selection; tracking data indicating the agent's tool invocation process; tracking data indicating the agent's tool return results; and tracking data indicating the model's generated results.
[0068] In S240, based on the first data section, the evaluation results of the internal execution process of the agent performing the first test task can be obtained.
[0069] As described above, a first data portion can be selected from the tracking data based on a hierarchical structure, targeting the evaluation metrics. The evaluation metrics can be related to the internal execution process of the first test task. Therefore, by obtaining the tracking data and further selecting the first data portion based on the hierarchical structure of the tracking data, a more accurate and granular evaluation of the agent's internal execution process of the first test task can be obtained.
[0070] Figure 3 The illustration schematically depicts a scenario of client-agent interaction for evaluating an agent, provided by at least one embodiment of this disclosure. For example... Figure 3 As shown, the client 3100 used to evaluate the agent and the agent 3200 to be evaluated can interact and invoke the artificial intelligence model 3300 used to evaluate the agent 3200.
[0071] In step S3001, the client 3100 may send first interactive guidance information to the agent 3200, wherein the first interactive guidance information is used to guide the agent to perform the first test task.
[0072] According to embodiments of this disclosure, client 3100 can present an interface for configuring test tasks to a user. In one embodiment, client 3100 can present an input box for receiving user input on a first test task. For example, the user can input a specific first test task in the input box based on evaluation metrics. In another embodiment, client 3100 can present one or more items indicating one or more test tasks included in a predetermined test set. For example, the user can select one or more test tasks by selecting one or more items based on evaluation metrics. Each of the one or more test tasks may correspond to one or more first interactive guidance messages. By configuring the corresponding test task, the user can send the corresponding first interactive guidance message to agent 3200. In one embodiment, at least a portion of the evaluation metrics may be related to the internal execution process of the agent performing the test task.
[0073] In step S3002, in response to the agent 3200 receiving the first interactive guidance information from the client 3100, the agent 3200 can execute the first test task. For example, the agent 3200 can process the first interactive guidance information to understand the first test task. The agent 3200 can plan a path to process the first test task. For example, the agent 3200 can decompose the first test task into one or more sub-test tasks. The agent 3200 can then invoke a large language model, tools, and knowledge base to execute one or more sub-test tasks, thereby completing the first test task. The process by which the agent 3200 executes the first test task is exemplary and not restrictive; the process may include more or fewer steps.
[0074] According to embodiments of this disclosure, the agent 3200 can generate raw tracking data during the internal execution of the first test task. For example, an SDK for the agent 3200 can be pre-configured. Figure 3 (Not shown in the diagram). For example, the SDK can be configured at agent 3200, but this disclosure is not limited thereto; the SDK can operate independently of agent 3200. The SDK can obtain trajectory data generated by agent 3200 when performing the first test task. The raw tracking data may not have a hierarchical structure, or at least a portion of the raw tracking data may not have a hierarchical structure.
[0075] In step S3003, the SDK can send the raw tracking data generated by the agent 3200 or the standardized tracking data to the client 3100. According to one embodiment of this disclosure, the SDK can standardize the raw tracking data to obtain tracking data with a hierarchical structure.
[0076] Although steps S3001-S3003 illustrate the process of client 3100 sending first interactive guidance information to agent 3200 and obtaining tracking data once, this disclosure is not limited thereto. For example, client 3100 may send first interactive guidance information to agent multiple times to obtain multiple tracking data generated by agent during multiple executions of first test task. Client 3100 may select tracking data from the generated multiple tracking data for subsequent evaluation based on predetermined conditions. Such predetermined conditions may be one or more of the following: predetermined conditions related to the generation time of tracking data, predetermined conditions related to the quantity of multiple tracking data, and predetermined conditions related to keywords. For example, predetermined conditions related to the generation time of tracking data may refer to selecting tracking data generated during a predetermined time period from multiple tracking data. For example, predetermined conditions related to the quantity of multiple tracking data may refer to selecting tracking data a predetermined number of times from multiple generated tracking data. For example, predetermined conditions related to keywords may refer to selecting tracking data with predetermined keywords from multiple generated tracking data.
[0077] In step S3004, the client 3100 can receive raw tracking data or standardized tracking data from the agent 3200 from the SDK. According to embodiments of this disclosure, the agent 3200 can standardize the raw tracking data to obtain tracking data with a hierarchical structure.
[0078] According to embodiments of this disclosure, for example in one example, tracking data with a hierarchical structure may include four levels (in other examples, there may be more or fewer), such as a root span level, a prompt span level, a model span level, and a tool span level, but this disclosure is not limited thereto. (See also...) Figure 2 At least a portion of the described tracking data for the input and output of the agent, the tracking data for the tool selection of the agent, the tracking data for the tool invocation process of the agent, the tracking data for the tool return results of the agent, and the tracking data for the model generation results may have or be abstracted into different levels.
[0079] According to embodiments of this disclosure, the root span level tracking data represents the complete lifecycle of a first test task in agent 3200 from triggering to final completion. The root span level tracking data includes the ID of the first test task, start and end times, final execution status (e.g., success, failure, abort, etc.), and some optional other information (e.g., total time elapsed).
[0080] In addition, the client 3100 can obtain the execution trajectory data of the agent 3200 from the agent 3200 or the SDK. The execution trajectory data records the complete execution trajectory of the agent 3200 in executing the first test task, such as the internal execution process or intermediate execution flow of the agent 3200.
[0081] According to embodiments of this disclosure, the tracking data of the prompt word span can record the process by which the agent 3200 processes the first interactive guidance information input by the user, etc. The tracking data of the prompt word span can capture the original first interactive guidance information and the key preprocessing steps that the agent 3200 may perform to execute subsequent steps (e.g., invoking an LLM or tool). These key preprocessing steps may include, but are not limited to, parsing the first interactive guidance information, intent recognition, supplementing and fusing contextual information, and security filtering; this disclosure is not limited thereto.
[0082] According to embodiments of this disclosure, the model span level tracking data can record a single invocation process of the agent 3200 to the LLM. The model span level tracking data may include, but is not limited to, the specific request content sent to the LLM (e.g., the processed first interaction guidance information, context, instructions, etc.) and the original response content returned by the LLM. Furthermore, the model span level tracking data can record the internal "thinking process" that the LLM may generate before generating the final response, such as thought chain steps, self-verification, and candidate option evaluation.
[0083] According to embodiments of this disclosure, tool span-level tracking data records a single tool invocation process of agent 3200. The tool span-level tracking data can annotate the name of the invoked tool, the tool invocation parameters used to invoke the tool, the result returned after tool execution (tool output upon success or error message upon failure), and the execution status of the tool invocation (success, failure, timeout, etc.).
[0084] In step S3005, client 3100 may configure the evaluation task. For example, client 3100 may present a configuration page for configuring the evaluation task to the user. For example, it may present the user with at least one evaluation metric and at least one evaluator for evaluating the at least one evaluation metric. For example, the evaluator may be used to evaluate the various steps in the internal execution process of agent 3200 performing the first test task. The evaluator used may be determined based on the user's selection of the evaluator or evaluation metric.
[0085] According to embodiments of this disclosure, at least one evaluation metric may include at least one of single-step and multi-step metrics. Single-step metrics can be used to evaluate a single step in the execution process of the agent 3200, such as tool selection, parameter correctness, tool execution, etc. Multi-step metrics can be used to evaluate multiple steps or the entire execution process of the agent 3200, such as task completion, trajectory accuracy, etc.
[0086] According to embodiments of this disclosure, the single-step metric may include at least one of a tool selection evaluation metric, a tool invocation parameter evaluation metric, and a tool execution evaluation metric. The tool selection evaluation metric can be used to evaluate whether the tool selected by the agent is appropriate and whether it can effectively promote the solution of the problem (e.g., the first test task and its sub-tasks). The tool invocation parameter evaluation metric can be used to evaluate whether the tool parameters provided by the agent 3200 are correct and meet the tool's requirements. The tool execution evaluation metric can be used to evaluate whether the tool invoked by the agent 3200 executes successfully and whether errors occur.
[0087] According to embodiments of this disclosure, the multi-step metrics may include at least one of a task completion evaluation metric and a trajectory accuracy evaluation metric. The task completion evaluation metric can be used to assess whether the agent 3200 successfully completed the first test task and whether it achieved the expected goal. The trajectory accuracy evaluation metric can be used to assess whether the execution trajectory of the agent 3200 is accurate, for example, whether it is consistent with a reference trajectory. The trajectory accuracy evaluation metric can also be used to assess whether the agent's execution trajectory is efficient and whether there are redundant steps, etc.
[0088] According to embodiments of this disclosure, at least one evaluator can be used to evaluate at least one evaluation metric of the agent 3200. The at least one evaluator may include, but is not limited to, a strict version of the tool selection quality evaluator, a lenient version of the tool selection quality evaluator, a tool parameter correctness evaluator, an agent task completion evaluator, and an agent trajectory quality evaluator. According to embodiments of this disclosure, the strict version of the tool selection quality evaluator and the lenient version of the tool selection quality evaluator can be used to evaluate the tool selection evaluation metric of the agent 3200. According to embodiments of this disclosure, the tool parameter correctness evaluator can be used to evaluate the tool invocation parameter evaluation metric of the agent 3200. According to embodiments of this disclosure, the agent task completion evaluator can be used to evaluate the task completion evaluation metric of the agent 3200. According to embodiments of this disclosure, the agent trajectory quality evaluator can be used to evaluate the trajectory accuracy evaluation metric of the agent 3200.
[0089] According to at least one embodiment of this disclosure, at least one evaluator may include second interactive guidance information for guiding an artificial intelligence model 3300 to output at least one evaluation result for at least one evaluation metric based on a selected first data portion. That is, at least one evaluator may include a selected first data portion and second interactive guidance information to guide the artificial intelligence model 3300 to evaluate the first data portion. In another embodiment, at least one evaluator may further include reference tracking data corresponding to the selected first data portion. For example, the reference tracking data may be used to compare with the selected first data portion so that the artificial intelligence model 3300 can determine at least one evaluation result.
[0090] According to at least one embodiment of this disclosure, the rigorous version of the tool selection quality evaluator may be in the following form, but the embodiments of this disclosure are not limited thereto:
[0091]
[0092]
[0093] The examples of the evaluators described above are for reference and understanding only.
[0094] According to embodiments of this disclosure, a more lenient version of the tool selection quality evaluator relative to the strict version described above can be in the following form, but embodiments of this disclosure are not limited thereto:
[0095]
[0096]
[0097] The above evaluator example is for reference and understanding only.
[0098] According to embodiments of this disclosure, the tool parameter correctness evaluator may be in the following form, but embodiments of this disclosure are not limited thereto:
[0099]
[0100]
[0101] The above examples are for reference and understanding only.
[0102] According to embodiments of this disclosure, the agent task completion evaluator may be in the following form, but embodiments of this disclosure are not limited thereto:
[0103]
[0104]
[0105] The above evaluator example is for reference and understanding only.
[0106] According to at least one embodiment of this disclosure, the agent trajectory quality evaluator may be in the following form, but the embodiments of this disclosure are not limited thereto:
[0107]
[0108]
[0109] The above evaluator example is for reference and understanding only.
[0110] According to embodiments of this disclosure, at least one evaluator can be associated with at least one level of tracking data having a hierarchical structure. For example, different evaluators can be associated with different levels. A first data portion having at least one level can be selected from the tracking data based on at least one level associated with the used or selected evaluator. According to another embodiment of this disclosure, a first data portion can be selected from the tracking data based on at least one of predetermined conditions related to the generation time of the tracking data and predetermined conditions related to keywords. For example, a first data portion can be selected from the tracking data based on the time period during which the user expects to conduct the evaluation. As another example, a first data portion can be selected from the tracking data based on keywords associated with the evaluator.
[0111] In step S3005, the client 3100 can add the selected first data portion to at least one evaluator. For example, the first data portion can be added as a variable to be evaluated to a pre-configured evaluator for subsequent processing.
[0112] In step S3006, the client 3100 may send at least one evaluator to the artificial intelligence model 3300 used to evaluate the agent. Although in Figure 3 In the scenario shown, the artificial intelligence model 3300 is presented as being set up independently of the client 3100, but the disclosure is not limited thereto. For example, the artificial intelligence model 3300 may be integrated into the client 3100 or may run on the server running the intelligent agent.
[0113] According to embodiments of this disclosure, the artificial intelligence model 3300 can be a content generation model based on prompt words, such as a large language model, a large visual model, or a multimodal large model. For example, it can be a model based on a Transformer architecture, a model built on a recurrent neural network, or a model built on an attention mechanism. Embodiments of this disclosure do not limit this.
[0114] In step S3007, the artificial intelligence model 3300 can evaluate the tracking data included in the evaluator based on the received evaluator. For example, the artificial intelligence model 3300 can evaluate the tracking data included in the evaluator based on the second interactive guidance information in the received evaluator to generate at least one evaluation result.
[0115] In step S3008, the artificial intelligence model 3300 may send at least one generated evaluation result to the client 3100. The client 3100 may receive at least one evaluation result from the artificial intelligence model 3300. For example, at least one evaluation result may be generated for a portion of the internal execution process of the agent 3200 performing the first test task.
[0116] In step S3009, the client 3100 may determine the comprehensive evaluation result of the agent 3200 based on at least one evaluation result. For example, the client 3100 may determine the comprehensive evaluation result of the agent 3200 during the internal execution process of performing the first test task based on at least one evaluation result.
[0117] According to embodiments of this disclosure, client 3100 can perform statistical analysis on at least one evaluation result. For example, it can perform statistical analysis on at least one evaluation result for different test tasks, different agent versions, and different time periods. Client 3100 can generate trend graphs expressing the trend of at least one evaluation result changing with test tasks, agent versions, time periods, etc., can generate charts expressing the numerical values of at least one evaluation result (e.g., task completion rate, tool selection accuracy, etc.), and can generate detail pages displaying at least one evaluation result, etc., but this disclosure is not limited thereto.
[0118] According to embodiments of this disclosure, client 3100 can determine, based on at least one evaluation result, defects existing in the internal execution process of agent 3200 in performing the first test task, such as incorrect tool selection, incorrect tool calling parameters, etc., and this disclosure is not limited thereto.
[0119] Figure 4 The schematic diagram illustrates a structural block diagram of an apparatus for evaluating an intelligent agent provided in at least one embodiment of the present disclosure.
[0120] like Figure 4As shown, the apparatus 400 for evaluating an intelligent agent in this embodiment includes a sending module 410, an obtaining module 420, a selection module 430, and an evaluation module 440. For example, these units or modules can be implemented by hardware (e.g., circuit) modules or software modules, firmware modules, etc. The following embodiments are similar and will not be repeated. For example, these units or modules can be implemented by a central processing unit (CPU), a general-purpose graphics processor (GPGPU), a graphics processing unit (GPU), a tensor processor (TPU), a field-programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, along with corresponding computer instructions.
[0121] The sending module 410 can be configured to send first interactive guidance information to the agent, wherein the first interactive guidance information is used to guide the agent to perform a first test task.
[0122] The acquisition module 420 can be configured to acquire tracking data generated by the agent during the execution of the first test task, wherein the tracking data has a hierarchical structure.
[0123] Selection module 430 can be configured to select a first data portion from the tracking data, based at least on a hierarchical structure of the tracking data.
[0124] The evaluation module 440 can be configured to obtain an evaluation result of the internal execution process of the agent performing the first test task based on the first data portion.
[0125] In at least one embodiment of this disclosure, the obtaining module 420 may be specifically configured to obtain raw tracking data generated by the agent during the internal execution of the first test task; and to standardize the raw tracking data to obtain tracking data with a hierarchical structure.
[0126] In at least one embodiment of this disclosure, the raw tracking data is obtained through a pre-configured software development kit (SDK).
[0127] In at least one embodiment of this disclosure, the tracking data having a hierarchical structure includes tracking data at one or more levels, the one or more levels including at least one of the following: root span level; cue word span level; model span level; and tool span level.
[0128] In at least one embodiment of this disclosure, the selection module 430 may be specifically configured to select at least one evaluator, wherein the at least one evaluator is used to evaluate at least one evaluation metric of the agent, and the at least one evaluator is associated with at least one level of tracking data having a hierarchical structure; and to select a first data portion having at least one level from the tracking data.
[0129] In at least one embodiment of this disclosure, the selection module 430 may be specifically configured to select a first data portion from the tracking data based on at least one of predetermined conditions related to the generation time of the tracking data and predetermined conditions related to keywords.
[0130] In at least one embodiment of this disclosure, at least one evaluation metric includes one or more of the following: single-step metrics and multi-step metrics. The single-step metrics include at least one of the following: tool selection evaluation metrics; tool invocation parameter evaluation metrics; and tool execution evaluation metrics. Alternatively, the multi-step metrics include at least one of the following: task completion evaluation metrics; and trajectory accuracy evaluation metrics.
[0131] In at least one embodiment of this disclosure, the apparatus 400 for evaluating an intelligent agent may further include an adding module, which may be configured to add a selected first data portion to at least one evaluator; and send at least one evaluator to an artificial intelligence model for evaluating the intelligent agent, wherein the at least one evaluator includes second interactive guidance information for guiding the artificial intelligence model to output at least one evaluation result for at least one evaluation metric based on the selected first data portion.
[0132] In at least one embodiment of this disclosure, at least one evaluator further includes reference tracking data corresponding to a selected first data portion, wherein the reference tracking data is used to compare with the selected first data portion to determine at least one evaluation result.
[0133] In at least one embodiment of this disclosure, the evaluation module 440 may be specifically configured to receive at least one evaluation result from an artificial intelligence model and, based on the at least one evaluation result, determine the evaluation result of the agent.
[0134] In at least one embodiment of this disclosure, the evaluation module 440 may be specifically configured to determine, based on at least one evaluation result, a defect existing in the internal execution of the agent performing the first test task.
[0135] In at least one embodiment of this disclosure, the sending module 410 may be specifically configured to send first interactive guidance information to the agent multiple times; obtain multiple tracking data generated by the agent during the multiple executions of the first test task; and select tracking data from the generated multiple tracking data for evaluation based on predetermined conditions.
[0136] In at least one embodiment of this disclosure, the predetermined conditions include one or more of the following: predetermined conditions related to the generation time of the tracking data; predetermined conditions related to the quantity of multiple tracking data; and predetermined conditions related to keywords.
[0137] It should be noted that, for clarity and brevity, this disclosure does not show all the constituent units of the device 400 for evaluating intelligent agents. To achieve the necessary functions of the device for evaluating intelligent agents, those skilled in the art can provide or configure other constituent units (not shown) according to specific needs, and this disclosure does not impose any limitations on this.
[0138] Figure 5 A schematic diagram of the structure of an electronic device (e.g., a terminal device or a server) 500 suitable for implementing embodiments of the present disclosure is shown.
[0139] refer to Figure 5 The terminal devices in this disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (e.g., vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0140] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0141] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0142] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0143] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0144] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0145] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0146] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: send first interactive guidance information to the intelligent agent, wherein the first interactive guidance information is used to guide the intelligent agent to perform a first test task; obtain tracking data generated by the intelligent agent during the execution of the first test task, wherein the tracking data has a hierarchical structure; select a first data portion from the tracking data based at least on the hierarchical structure of the tracking data; and obtain an evaluation result of the internal execution process of the intelligent agent performing the first test task based on the first data portion.
[0147] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0149] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily constitute a limitation on the unit or module itself.
[0150] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0151] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0152] One embodiment of this disclosure provides a non-transitory readable storage medium having computer instructions stored thereon that, when executed by a processor, perform one or more steps of the various methods and additional aspects described above.
[0153] For example, the non-temporarily readable storage medium may be any combination of one or more computer-readable storage media, such as a computer-readable storage medium containing program code for performing the various methods described above.
[0154] For example, when the program code is read by a computer, the computer can execute the program code stored in the computer storage medium to perform one or more steps of the various methods and additional aspects described above, such as those according to at least one embodiment of the present disclosure.
[0155] For example, the non-transitory readable storage medium may include a memory card of a smartphone, a storage component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, and other non-transitory readable storage media or any combination thereof.
[0156] Embodiments of this disclosure also provide a computer program product. The computer program product may include a computer program or instructions. When executed by a processor, the computer program or instructions can implement the methods described above, which will not be repeated here for the sake of brevity.
[0157] According to one or more embodiments of this disclosure, Example 1 provides a method for evaluating an intelligent agent, comprising:
[0158] Send first interactive guidance information to the intelligent agent, wherein the first interactive guidance information is used to guide the intelligent agent to execute a first test task;
[0159] Obtain tracking data generated by the agent during the execution of the first test task, wherein the tracking data has a hierarchical structure;
[0160] Based at least on the hierarchical structure of the tracking data, a first data portion is selected from the tracking data;
[0161] Based on the first data portion, an evaluation result is obtained of the internal execution process of the agent performing the first test task.
[0162] According to one or more embodiments of this disclosure, Example 2 provides a method, wherein obtaining the tracking data generated by the agent during the execution of the first test task in Example 1 includes:
[0163] Obtain the raw tracking data generated by the agent during its internal execution of the first test task; and
[0164] The original tracking data is standardized to obtain tracking data with a hierarchical structure.
[0165] According to one or more embodiments of this disclosure, Example 3 provides a method in which the raw tracking data in Example 2 is obtained through a pre-configured software development kit (SDK).
[0166] According to one or more embodiments of this disclosure, Example 4 provides a method wherein the hierarchical tracking data in Example 1 includes tracking data of one or more levels, said one or more levels including at least one of the following:
[0167] Root span level;
[0168] The scope of the prompt words spans multiple levels;
[0169] Model span levels; and
[0170] Tool spanning levels.
[0171] According to one or more embodiments of this disclosure, Example 5 provides a method in which selecting a first data portion from the tracking data in Example 1 includes:
[0172] At least one evaluator is selected, wherein the at least one evaluator is used to evaluate at least one evaluation metric of the agent, and the at least one evaluator is associated with at least one level of the tracking data having a hierarchical structure;
[0173] Select a first data portion having at least one level from the tracking data.
[0174] According to one or more embodiments of this disclosure, Example Six provides a method in which selecting a first data portion from the tracking data in Example Five includes:
[0175] A first data portion is selected from the tracking data based on at least one of predetermined conditions related to the generation time of the tracking data and predetermined conditions related to keywords.
[0176] According to one or more embodiments of this disclosure, Example 7 provides a method wherein the at least one evaluation metric in Example 5 includes one or more of the following: single-step metrics and multi-step metrics; wherein the single-step metrics include at least one of the following: tool selection evaluation metrics, tool invocation parameter evaluation metrics, and tool execution evaluation metrics; and the multi-step metrics include at least one of the following: task completion evaluation metrics and trajectory accuracy evaluation metrics.
[0177] According to one or more embodiments of this disclosure, Example Eight provides a method, wherein the method in Example Five further includes:
[0178] Add the selected first data portion to the at least one evaluator;
[0179] The at least one evaluator is sent to the artificial intelligence model used to evaluate the agent.
[0180] The at least one evaluator includes second interactive guidance information for guiding the artificial intelligence model to output at least one evaluation result for the at least one evaluation metric based on the selected first data portion.
[0181] According to one or more embodiments of this disclosure, Example Nine provides a method wherein the at least one evaluator in Example Eight further includes reference tracking data corresponding to the selected first data portion.
[0182] The reference tracking data is used to compare with the selected first data portion to determine the at least one evaluation result.
[0183] According to one or more embodiments of this disclosure, Example 10 provides a method in which obtaining the evaluation result of the agent in Example 8 includes:
[0184] Receive at least one evaluation result from the artificial intelligence model.
[0185] Based on the at least one evaluation result, the evaluation result of the agent is determined.
[0186] According to one or more embodiments of this disclosure, Example 11 provides a method, wherein the method in Example 10 further includes:
[0187] Based on the at least one evaluation result, defects in the agent's internal execution process of the first test task are determined.
[0188] According to one or more embodiments of this disclosure, Example Twelve provides a method, wherein the method in Example One further includes:
[0189] Send the first interaction guidance information to the intelligent agent multiple times;
[0190] Obtain multiple tracking data generated by the intelligent agent during the multiple executions of the first test task;
[0191] Based on predetermined conditions, tracking data is selected from the generated multiple tracking data sets for evaluation; wherein the predetermined conditions include one or more of the following:
[0192] Pre-defined conditions related to the time of generation of tracking data;
[0193] Predetermined conditions related to the quantity of the plurality of tracking data; and
[0194] Pre-defined conditions related to keywords.
[0195] According to one or more embodiments of the present disclosure, Example Thirteen provides a computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the processors to perform the method as described in any one of Examples One to Twelve.
[0196] According to one or more embodiments of this disclosure, Example Fourteen provides an electronic device, including:
[0197] One or more processors; and
[0198] One or more memories storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform the method as described in any one of Examples 1 to 12.
[0199] According to one or more embodiments of this disclosure, Example Fifteen provides an apparatus for evaluating an intelligent agent, comprising:
[0200] The sending module is configured to send first interactive guidance information to the intelligent agent, wherein the first interactive guidance information is used to guide the intelligent agent to perform a first test task;
[0201] The acquisition module is configured to acquire tracking data generated by the agent during the execution of the first test task, wherein the tracking data has a hierarchical structure;
[0202] The selection module is configured to select a first data portion from the tracking data, based at least on the hierarchical structure of the tracking data.
[0203] The evaluation module is configured to obtain an evaluation result of the internal execution process of the agent performing the first test task based on the first data portion.
[0204] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0205] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0206] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for evaluating an agent, comprising: Send first interactive guidance information to the intelligent agent, wherein the first interactive guidance information is used to guide the intelligent agent to execute a first test task; Obtain tracking data generated by the agent during the execution of the first test task, wherein the tracking data has a hierarchical structure; Based at least on the hierarchical structure of the tracking data, a first data portion is selected from the tracking data; Based on the first data portion, an evaluation result is obtained of the internal execution process of the agent performing the first test task.
2. The method according to claim 1, wherein, The tracking data generated by the agent during the execution of the first test task includes: Obtain the raw tracking data generated by the agent during its internal execution of the first test task; and The original tracking data is standardized to obtain tracking data with a hierarchical structure.
3. The method according to claim 2, wherein, The raw tracking data was obtained through a pre-configured software development kit (SDK).
4. The method according to claim 1, wherein, The hierarchical tracking data includes one or more levels of tracking data, and the one or more levels include at least one of the following: Root span level; The scope of the prompt words spans multiple levels; Model span levels; and Tool spanning levels.
5. The method according to claim 1, wherein, Selecting a first data portion from the tracking data includes: At least one evaluator is selected, wherein the at least one evaluator is used to evaluate at least one evaluation metric of the agent, and the at least one evaluator is associated with at least one level of the tracking data having a hierarchical structure; Select a first data portion having at least one level from the tracking data.
6. The method according to claim 5, wherein, Selecting a first data portion from the tracking data includes: A first data portion is selected from the tracking data based on at least one of predetermined conditions related to the generation time of the tracking data and predetermined conditions related to keywords.
7. The method according to claim 5, wherein, The at least one evaluation metric includes one or more of the following: single-step metrics and multi-step metrics. The single-step metrics include at least one of the following: tool selection evaluation metrics, tool invocation parameter evaluation metrics, and tool execution evaluation metrics; The multi-step metrics include at least one of the following: task completion assessment metrics and trajectory accuracy assessment metrics.
8. The method according to claim 5, further comprising: Add the selected first data portion to the at least one evaluator; The at least one evaluator is sent to the artificial intelligence model used to evaluate the agent. The at least one evaluator includes second interactive guidance information for guiding the artificial intelligence model to output at least one evaluation result for the at least one evaluation metric based on the selected first data portion.
9. The method according to claim 8, wherein, The at least one evaluator also includes reference tracking data corresponding to the selected first data portion. The reference tracking data is used to compare with the selected first data portion to determine the at least one evaluation result.
10. The method according to claim 8, wherein, Obtaining the evaluation results for the agent includes: Receive at least one evaluation result from the artificial intelligence model. Based on the at least one evaluation result, the evaluation result of the agent is determined.
11. The method of claim 10, further comprising: Based on the at least one evaluation result, defects in the agent's internal execution process of the first test task are determined.
12. The method according to claim 1, further comprising: Send the first interaction guidance information to the intelligent agent multiple times; Obtain multiple tracking data generated by the intelligent agent during the multiple executions of the first test task; Based on predetermined conditions, select tracking data from the generated multiple tracking data sets for evaluation; The predetermined conditions include one or more of the following: predetermined conditions related to the generation time of the tracking data, predetermined conditions related to the quantity of the plurality of tracking data, and predetermined conditions related to keywords.
13. A computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method as described in any one of claims 1-12.
14. An electronic device comprising: One or more processors; as well as One or more memories storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform the method as described in any one of claims 1-12.
15. An apparatus for evaluating an intelligent agent, comprising: The sending module is configured to send first interactive guidance information to the intelligent agent, wherein the first interactive guidance information is used to guide the intelligent agent to perform a first test task; The acquisition module is configured to acquire tracking data generated by the agent during the execution of the first test task, wherein the tracking data has a hierarchical structure; The selection module is configured to select a first data portion from the tracking data, based at least on the hierarchical structure of the tracking data. The evaluation module is configured to obtain an evaluation result of the internal execution process of the agent performing the first test task based on the first data portion.
Citation Information
Patent Citations
Data analysis method and device based on distributed link tracking and electronic equipment
CN114185708A
Consensus transaction tracking method and device, equipment and storage medium
CN115129554A
Method and device for evaluation, electronic equipment and computer program product
CN118708921A
Intelligent agent evaluation method and device and storage medium
CN118916662A
Method for debugging intelligent agent through computing power of intelligent computing center
CN120179535A