Intelligent agent quality evaluation method and device, computer equipment and readable storage medium

By combining task execution records and user preference operation records, the method for evaluating the quality of intelligent agents solves the problem of inaccurate evaluation in the intelligent agent market and achieves accurate ranking and value reflection of intelligent agents.

CN121958045APending Publication Date: 2026-05-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2026-01-06
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing methods for evaluating the quality of intelligent agents suffer from traffic monopoly, where early-released, mediocre agents dominate the rankings for a long time, making it difficult for new, high-quality agents to surpass them. Furthermore, usage volume does not directly equate to quality, leading to inaccurate evaluations.

Method used

By identifying multiple agents from the agent market, evaluating the task completion quality factor based on task execution records, and evaluating the user preference factor by combining user preference operation records, a quantitative evaluation of the agents is obtained, and the agents are ranked based on this result.

Benefits of technology

This improves the accuracy of agent quality assessment, ensures the reasonable ranking of high-quality agents in the market, avoids traffic monopoly, and reflects the actual value of agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958045A_ABST
    Figure CN121958045A_ABST
Patent Text Reader

Abstract

The invention relates to an agent quality evaluation method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: determining a plurality of agents from an agent market; for each agent, according to the task execution record of the agent, evaluating the task completion quality of the agent, and determining a task completion quality factor; according to the user preference operation record of the intelligent agent, performing user preference evaluation on the intelligent agent, and determining a user preference factor; according to the task completion quality factor and the user preference factor, performing quantitative evaluation on the intelligent agent to obtain a quantitative evaluation result of the intelligent agent; sorting the plurality of agents based on the quantitative evaluation results of the plurality of agents to obtain an agent quality evaluation result; the intelligent agent quality evaluation result is used for updating the arrangement sequence of the multiple intelligent agents in the intelligent agent market. By adopting the method, the accuracy of intelligent agent quality evaluation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for evaluating the quality of intelligent agents. Background Technology

[0002] With the development of computer technology, an intelligent agent market has emerged. An intelligent agent market is a platform where developers can publish, users can discover, and enable various intelligent agents. Each intelligent agent describes its functionality through structured metadata, and users can enable one or more intelligent agents to assist in completing tasks. An intelligent agent is an autonomous software entity capable of perceiving its environment, planning, invoking tools (or its own capabilities), and performing actions to achieve a goal. An intelligent agent typically encapsulates complex logic and capabilities for solving specific types of problems.

[0003] In related technologies, the quality assessment of intelligent agents in the intelligent agent market is usually done by sorting multiple intelligent agents in descending order according to their total historical usage or total downloads, and then displaying them to the user.

[0004] However, traditional methods are prone to creating traffic monopolies. Early-released agents, even those with mediocre capabilities, can maintain their top rankings for a long time due to cumulative effects, making it difficult for new, high-quality agents to surpass them. Furthermore, usage volume does not directly equate to quality; some highly entertaining but poorly functional agents may also rank highly, leading to inaccurate agent quality assessments. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for assessing intelligent agents that can improve the accuracy of intelligent agent quality assessment, in response to the above-mentioned technical problems.

[0006] Firstly, this application provides a method for evaluating the quality of an intelligent agent, including:

[0007] Identify multiple agents from the agent market;

[0008] For each agent, the quality of the agent's task completion is evaluated based on the agent's task execution record, and a task completion quality factor is determined.

[0009] Based on the user preference operation records of the intelligent agent, the user preference is evaluated to determine the user preference factor;

[0010] The agent is quantitatively evaluated based on the task completion quality factor and the user preference factor to obtain the quantitative evaluation result of the agent.

[0011] Based on the quantitative evaluation results of each of the multiple agents, the multiple agents are ranked to obtain the agent quality evaluation results; the agent quality evaluation results are used to update the ranking order of the multiple agents in the agent market.

[0012] Secondly, this application also provides an intelligent agent quality assessment device, comprising:

[0013] The agent identification module is used to identify multiple agents from the agent market;

[0014] The task quality assessment module is used to assess the quality of the task completed by each agent based on the agent's task execution record, and determine the task completion quality factor.

[0015] The user preference evaluation module is used to evaluate the user preferences of the intelligent agent based on the user preference operation records of the intelligent agent and determine the user preference factor.

[0016] The quantitative evaluation module is used to perform quantitative evaluation on the agent based on the task completion quality factor and the user preference factor, and obtain the quantitative evaluation result of the agent.

[0017] The sorting module is used to sort the multiple agents based on their respective quantitative evaluation results to obtain agent quality evaluation results; the agent quality evaluation results are used to update the order of the multiple agents in the agent market.

[0018] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0019] Identify multiple agents from the agent market;

[0020] For each agent, the quality of the agent's task completion is evaluated based on the agent's task execution record, and a task completion quality factor is determined.

[0021] Based on the user preference operation records of the intelligent agent, the user preference is evaluated to determine the user preference factor;

[0022] The agent is quantitatively evaluated based on the task completion quality factor and the user preference factor to obtain the quantitative evaluation result of the agent.

[0023] Based on the quantitative evaluation results of each of the multiple agents, the multiple agents are ranked to obtain the agent quality evaluation results; the agent quality evaluation results are used to update the ranking order of the multiple agents in the agent market.

[0024] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0025] Identify multiple agents from the agent market;

[0026] For each agent, the quality of the agent's task completion is evaluated based on the agent's task execution record, and a task completion quality factor is determined.

[0027] Based on the user preference operation records of the intelligent agent, the user preference is evaluated to determine the user preference factor;

[0028] The agent is quantitatively evaluated based on the task completion quality factor and the user preference factor to obtain the quantitative evaluation result of the agent.

[0029] Based on the quantitative evaluation results of each of the multiple agents, the multiple agents are ranked to obtain the agent quality evaluation results; the agent quality evaluation results are used to update the ranking order of the multiple agents in the agent market.

[0030] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0031] Identify multiple agents from the agent market;

[0032] For each agent, the quality of the agent's task completion is evaluated based on the agent's task execution record, and a task completion quality factor is determined.

[0033] Based on the user preference operation records of the intelligent agent, the user preference is evaluated to determine the user preference factor;

[0034] The agent is quantitatively evaluated based on the task completion quality factor and the user preference factor to obtain the quantitative evaluation result of the agent.

[0035] Based on the quantitative evaluation results of each of the multiple agents, the multiple agents are ranked to obtain the agent quality evaluation results; the agent quality evaluation results are used to update the ranking order of the multiple agents in the agent market.

[0036] The aforementioned intelligent agent quality assessment method, apparatus, computer equipment, computer-readable storage medium, and computer program product, based on identifying multiple intelligent agents from the intelligent agent market, assess the quality of task completion for each intelligent agent by evaluating its task execution records and determining task completion quality factors, thus achieving quality assessment from the task quality dimension. Furthermore, by evaluating user preferences based on the intelligent agent's user preference operation records and determining user preference factors, quality assessment from the user preference dimension can be achieved. Therefore, by combining task completion quality factors and user preference factors, and considering both dimensions, accurate quantitative assessment of the intelligent agent can be obtained, yielding quantitative assessment results. Subsequently, based on the individual quantitative assessment results of multiple intelligent agents, they can be ranked to obtain the overall intelligent agent quality assessment result, thereby improving the accuracy of intelligent agent quality assessment. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a diagram illustrating the application environment of an agent quality assessment method in one embodiment.

[0039] Figure 2 This is a flowchart illustrating an agent quality assessment method in one embodiment;

[0040] Figure 3 This is a schematic diagram illustrating the selection of a target task in one embodiment;

[0041] Figure 4 This is a schematic diagram illustrating the quantitative evaluation results of the agent obtained in one embodiment;

[0042] Figure 5 This is a schematic diagram of task completion quality assessment prompts in one embodiment;

[0043] Figure 6 This is a schematic diagram illustrating the first completion quality score obtained in one embodiment;

[0044] Figure 7 This is a core architecture diagram of the competitive elimination algorithm engine in one embodiment;

[0045] Figure 8 This is a schematic diagram illustrating the calculation of multiple evaluation factors in one embodiment;

[0046] Figure 9 This is a flowchart illustrating the agent quality assessment method in another embodiment;

[0047] Figure 10 This is a schematic diagram of the market front-end display interface in one embodiment;

[0048] Figure 11 This is a schematic diagram of the market front-end display interface in another embodiment;

[0049] Figure 12 This is a structural block diagram of an agent quality assessment device in one embodiment;

[0050] Figure 13 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0052] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0053] In order to clearly describe the technical solution of this application and facilitate understanding of the technical solution of this application, the key concepts involved in this application will be explained below.

[0054] 1. Smart agent market.

[0055] The Agent Marketplace is a platform for developers to publish, users to discover, and enable various agents. Each agent describes its functionality through structured metadata, and users can enable one or more agents to assist in completing tasks as needed.

[0056] 2. Task execution record.

[0057] Task execution records refer to structured data records generated by an intelligent agent during the planning, decision-making, and execution of tasks. They are used to record the complete task lifecycle and support subsequent analysis and optimization; they can be understood as the agent's "task archive." For example, task execution records may include task identifier, user identifier, agent identifier, task objective, agent output results, timestamp, and task status.

[0058] 3. User preference operation records.

[0059] User preference operation records refer to structured data records generated after a user initiates a preference operation on an intelligent agent. These records are used to document user preferences for the intelligent agent and support subsequent analysis and optimization; they can be understood as the intelligent agent's "user preference profile." For example, user preference operation records may include user identifier, intelligent agent identifier, preference operation, and timestamp.

[0060] The agent quality assessment method provided in this application can be applied to, for example... Figure 1 In the application environment shown, the data storage system stores the data that server 102 needs to process. The data storage system can be set up independently, integrated into server 102, or placed in the cloud or on other network servers. Server 102 identifies multiple agents from the agent market. For each agent, it evaluates the quality of task completion based on the agent's task execution records, determines a task completion quality factor, evaluates user preferences based on the agent's user preference operation records, determines a user preference factor, and performs a quantitative evaluation of the agent based on the task completion quality factor and user preference factor, obtaining a quantitative evaluation result. Based on the quantitative evaluation results of multiple agents, the agents are ranked to obtain an agent quality evaluation result. The agent quality evaluation result is used to update the ranking order of multiple agents in the agent market. Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0061] In one exemplary embodiment, such as Figure 2 As shown, a method for evaluating the quality of an intelligent agent is provided. This embodiment illustrates the application of this method to a server. It is understood that this method can also be applied to a terminal, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes steps 202 to 210. Wherein:

[0062] Step 202: Identify multiple intelligent agents from the intelligent agent market.

[0063] The Agent Marketplace is a platform for developers to publish, users to discover, and enable various agents. Each agent describes its functions through structured metadata, and users can enable one or more agents to assist in completing tasks as needed.

[0064] An intelligent agent is an autonomous software entity capable of perceiving its environment, planning, invoking tools (or its own capabilities), and executing actions to achieve its goals. An intelligent agent typically encapsulates complex logic and capabilities for solving specific types of problems.

[0065] For example, when agent quality assessment is required, the server identifies multiple agents from the agent marketplace. It is understood that agents in the agent marketplace refer to those that developers have published to the marketplace.

[0066] Step 204: For each agent, evaluate the quality of the agent's task completion based on the agent's task execution record, and determine the task completion quality factor.

[0067] Task execution records refer to structured data records generated by an agent during the planning, decision-making, and execution of tasks. They are used to record the complete task lifecycle and support subsequent analysis and optimization; they can be understood as the agent's "task archive." For example, a task execution record may include task identifier, user identifier, agent identifier, task objective, agent output result, timestamp, and task status. Specifically, the format of a task execution record can be: {task_id (task identifier), user_id (user identifier), agent_id (agent identifier), task_description (task objective), agent_output (agent output result), timestamp (timestamp), status (task status)}.

[0068] Here, the task identifier is a unique identifier assigned to a task, used to distinguish it from other tasks. The user identifier is a unique identifier assigned to the user who submitted the task, used to distinguish the user who submitted the task from other users. The agent identifier is a unique identifier assigned to the agent that performs the task, used to distinguish the agent that performs the task from other agents.

[0069] The task objective, expressed in natural language, refers to the specific tasks the agent needs to complete, the expected final state, and possible constraints. It serves as the initial starting point and ultimate evaluation criterion for all the agent's planning, decision-making, and actions. The agent's output refers to the result produced after executing the task; that is, the result of the agent's task processing. The timestamp records the time point in time when the agent processed the task. The task state describes the agent's performance on the task. For example, the task state could be "task executed successfully" or "task executed unsuccessfully."

[0070] Evaluating the quality of an agent's task completion refers to the process of assessing the quality of its output based on pre-defined multi-dimensional standards (such as relevance, completeness, and usability) after the agent completes the task, and then providing quantitative indicators. The task completion quality factor is an indicator used to evaluate the quality of the agent's task completion. For example, the task completion quality factor can specifically be a task completion quality score.

[0071] For example, based on the determination of multiple agents, for each agent, the server will evaluate the quality of the agent's task completion based on the agent's task execution record and determine the task completion quality factor.

[0072] In practical applications, for each agent, the task execution platform outputs the agent's task execution record after the agent completes a task. The server subscribes to and receives event messages from the task execution platform in real time, based on which the server can obtain the agent's task execution record. It can be understood that when an agent completes multiple tasks, the task execution record output by the task execution platform is actually a task execution record stream, including the individual task execution records of each completed task.

[0073] In practical applications, based on the task execution records of the agent, the server can evaluate the task completion quality of multiple completed tasks of the agent according to the task execution records, obtain the task completion quality score of each of the multiple completed tasks, and then determine the task completion quality factor based on the task completion quality scores of each of the multiple completed tasks.

[0074] In a specific application, for each completed task, the task execution record can be determined based on the task execution record. Then, by analyzing the task execution record, the task completion quality can be evaluated, resulting in a task completion quality score. Based on the individual task completion quality scores of multiple completed tasks, the server can determine the agent's first completion quality score for the current quality evaluation period. Finally, based on this first completion quality score, a task completion quality factor is obtained.

[0075] In a specific application, based on the task completion quality scores of multiple completed tasks, the server determines the task quantity window for the current quality assessment period. According to the task quantity window, multiple target tasks are selected from the multiple completed tasks. Then, based on the task completion quality scores of the multiple target tasks, the agent's first completion quality score for the current quality assessment period is determined.

[0076] The task count window for the current quality assessment period can be a value that is less than or equal to the number of completed tasks. The method for determining this value can be configured according to the actual application scenario. For example, the task count window can be a predefined fixed number of tasks, a predefined task count ratio, or the number of times the agent is used in the current quality assessment period. The fixed task count and the task count ratio can be configured according to the actual application scenario.

[0077] For example, let's take 100 completed tasks as an example. Figure 3 As shown, if the task quantity window is a predefined fixed number of tasks, and that fixed number of tasks is 50, then the server can select the 50 most recently completed tasks from multiple completed tasks as multiple target tasks. If the task quantity window is a predefined task quantity ratio, and that task quantity ratio is 40%, then the server can select the 40 most recently completed tasks from multiple completed tasks as multiple target tasks. If the task quantity window is based on the number of times the agent is used, and the task quantity window is 30, then the server can select the 30 most recently completed tasks from multiple completed tasks as multiple target tasks.

[0078] In a specific application, based on the determined first completion quality score, the server can use the first completion quality score as the task completion quality factor, or it can combine the first completion quality score with the agent's second completion quality score in the historical quality evaluation period to determine the task completion quality factor. In this embodiment, no specific limitation is made here.

[0079] Step 206: Based on the user preference operation records of the agent, evaluate the user preference of the agent and determine the user preference factor.

[0080] User preference operation records refer to structured data records generated after a user initiates a preference operation on an agent. These records are used to document user preferences for the agent and support subsequent analysis and optimization; they can be understood as the agent's "user preference profile." For example, a user preference operation record may include a user identifier, agent identifier, preference operation, and a timestamp. Specifically, the format of a user preference operation record could be: {user_id (user identifier), agent_id (agent identifier), action: "enable / disable" (preference operation: enable / disable), timestamp (timestamp)}.

[0081] Here, the user identifier is a unique identifier assigned to the user who initiates a preference operation on an agent, used to distinguish the user who initiates the preference operation from other users. The agent identifier is a unique identifier assigned to the agent on which the user initiates a preference operation, used to distinguish the agent on which the user initiates a preference operation from other agents. A preference operation is an operation initiated by the user that represents a preference for an agent. For example, a preference operation can be "enable" or "disable," where "enable" means enabling the agent and "disable" means disabling the agent. The timestamp is used to record the time point at which the user initiates a preference operation on the agent.

[0082] Among them, the user preference factor refers to the indicator used to evaluate the agent from the perspective of user preferences. For example, the user preference factor can be specifically a user preference score. It can be understood that user preference actions towards the agent refer to the preference signals expressed by the user through proactive behaviors such as enabling or disabling the agent. This is a strong feedback that directly reflects the actual value of the agent to the user.

[0083] For example, the server will evaluate the agent's user preferences based on the agent's user preference operation records to determine the user preference factor.

[0084] In practical applications, for each agent, after a user makes a preference action on the front-end interface, a user preference action record is generated for the agent. The server subscribes to and receives event messages from the front-end interface in real time. Based on this, the server can obtain the agent's user preference action record. It can be understood that when multiple users make preference actions, the user preference action record of the agent generated by the front-end interface is actually a stream of user preference action records, including the individual user preference action records of multiple users.

[0085] In practical applications, based on the user preference operation records of the agent, the server can count the number of times the agent is enabled and disabled, and evaluate the user preference of the agent based on the number of times the agent is enabled and disabled to determine the user preference factor.

[0086] Step 208: Quantitatively evaluate the agent based on the task completion quality factor and user preference factor to obtain the quantitative evaluation result of the agent.

[0087] Quantitative evaluation refers to the objective analysis and comparison of the quality of an intelligent agent through clearly defined indicators and quantitative methods. The quantitative evaluation result refers to the objective reflection of the agent's quality after the quantitative evaluation. For example, the quantitative evaluation result can specifically be a quantitative evaluation score.

[0088] For example, the server will perform a quantitative evaluation of the agent based on the task completion quality factor and the user preference factor, and obtain the quantitative evaluation result of the agent.

[0089] In specific applications, such as Figure 4 As shown, when both the task completion quality factor and the user preference factor are numerical values, the server can quantitatively evaluate the agent by weighting these factors, thus obtaining the agent's quantitative evaluation result. The first factor weight coefficient of the task completion quality factor and the second factor weight coefficient of the user preference factor can be configured according to the actual application scenario.

[0090] In specific applications, such as Figure 4 As shown, the server can also acquire additional factors, and combine them with task completion quality factors, user preference factors, and other factors to quantitatively evaluate the agent, obtaining a quantitative evaluation result for the agent. In a specific application, these other factors can specifically include at least one of usage trend factors and timeliness factors. When the task completion quality factor, user preference factor, and other factors are all numerical values, the server can quantitatively evaluate the agent by weighting these factors to obtain a quantitative evaluation result for the agent. The first factor weight coefficient of the task completion quality factor, the second factor weight coefficient of the user preference factor, and the third factor weight coefficients of the other factors can be configured according to the actual application scenario.

[0091] Step 210: Based on the quantitative evaluation results of each agent, the agents are ranked to obtain the agent quality evaluation results; the agent quality evaluation results are used to update the ranking order of the agents in the agent market.

[0092] Among them, the agent quality evaluation result refers to the result formed after ranking multiple agents, which can objectively reflect the quality of the multiple agents, that is, the agent quality ranking result.

[0093] For example, based on the quantitative evaluation results of multiple agents, the server sorts the multiple agents to obtain agent quality evaluation results, which are used to update the ranking order of multiple agents in the agent market.

[0094] The aforementioned agent quality assessment method, based on identifying multiple agents from the agent market, evaluates the quality of task completion for each agent by analyzing its task execution records, determining a task completion quality factor. This enables quality assessment from the task quality dimension. Furthermore, by analyzing the agent's user preference operation records, it assesses user preferences by determining a user preference factor, achieving quality assessment from the user preference dimension. By combining the task completion quality factor and the user preference factor, and considering both dimensions, an accurate quantitative assessment of the agent can be achieved, yielding a quantitative assessment result. Based on these quantitative assessment results, multiple agents can be ranked to obtain the final agent quality assessment result, thus improving the accuracy of agent quality assessment.

[0095] In an exemplary embodiment, for each agent, the quality of the agent's task completion is evaluated based on the agent's task execution record to determine a task completion quality factor, including:

[0096] For each agent, based on the agent's task execution records, determine the individual task execution records for each of the agent's multiple completed tasks;

[0097] For each completed task, based on the task execution record of the completed task, the task completion quality is evaluated to obtain the task completion quality score of the completed task;

[0098] Based on the task completion quality scores of each of the multiple completed tasks, a task completion quality factor is determined.

[0099] For example, for each agent, the server can determine the task execution records of each of the agent's multiple completed tasks from the agent's task execution records. Then, for each completed task, the server can perform a task completion quality assessment based on the task execution records of the completed task to obtain a task completion quality score. Based on the task completion quality scores of each of the multiple completed tasks, the server can determine the agent's first completion quality score in the current quality assessment period. Finally, based on the first completion quality score, the server can obtain a task completion quality factor.

[0100] In practical applications, for each completed task, the server can asynchronously invoke the large language model evaluation service based on the task execution record to evaluate the task completion quality and obtain a task completion quality score. This evaluation can be performed immediately upon receiving the task execution record, and the resulting task completion quality score can be directly used when evaluating the agent's task completion quality. This asynchronous evaluation method saves computational resources when evaluating the agent's task completion quality, while still enabling timely evaluation of completed tasks. The large language model evaluation service refers to the service that evaluates the task completion quality of completed tasks based on a large language model.

[0101] In this embodiment, by utilizing the task execution records of the agent, the task execution records of each of the agent's multiple completed tasks can be determined. Then, for each completed task, the task completion quality can be evaluated using the task execution records of the completed task, and a task completion quality score can be obtained. Thus, by combining the task completion quality scores of multiple completed tasks, a task completion quality factor can be determined, and an accurate evaluation of the quality of the agent's task completion can be achieved using multiple completed tasks.

[0102] In an exemplary embodiment, for each completed task, based on the task execution record of the completed task, a task completion quality assessment is performed on the completed task to obtain a task completion quality score, including:

[0103] For each completed task, determine the task objective and the agent's output from the task execution record of the completed task;

[0104] Using a large language model, based on the task objective, agent output, and task completion quality assessment prompts, the completed task is evaluated to obtain a task completion quality score.

[0105] The task objective, in the form of natural language instructions, clearly informs the agent of the specific tasks to be completed, the expected final state, and possible constraints. It serves as the original starting point and ultimate evaluation criterion for all the agent's planning, decision-making, and actions. The agent's output refers to the result produced after executing the task, i.e., the result of the agent's task processing. The large language model refers to a large-scale artificial intelligence model with natural language understanding and generation capabilities; it is the core engine driving the agent's task planning and decision-making.

[0106] Among them, task completion quality assessment prompts refer to predefined prompts used to assess the quality of completed tasks, which can be configured according to actual application scenarios. For example, task completion quality assessment prompts may specifically include role definitions, input parameter requirements, assessment dimensions and standards, and output requirements.

[0107] For example, for each completed task, the server can determine the task objective and the agent's output from the task execution record of the completed task. Then, it can use a large language model to evaluate the task completion quality based on the task objective, the agent's output, and task completion quality assessment prompts to obtain the task completion quality score of the completed task.

[0108] In practical applications, the task completion quality assessment prompts include role definitions, input parameter requirements, assessment dimensions and standards, and output requirements. The server can input the task objective and the agent's output based on the input parameter requirements, allowing the large language model to assess the completed task's completion quality based on the role definition and assessment dimensions and standards, and output the task completion quality assessment result according to the output requirements. In a specific application, this task completion quality assessment result can be the total quality assessment score obtained after assessing the completed task based on multiple assessment dimensions and standards. To eliminate the incomparability caused by different factors due to differences in data units (dimensions) and numerical ranges (order of magnitude), and to ensure that all factors can be analyzed and modeled on a fair benchmark, the server will standardize the total quality assessment score and then use the standardized total score as the task completion quality score. The standardization method can be configured according to the actual application scenario. For example, the standardization method can be normalized to the 0-1 interval, i.e., the task completion quality score, which can be the ratio of the total quality assessment score to the full score.

[0109] In a specific application, role definitions and evaluation dimensions and standards can be configured according to the actual application scenario. For example, the specific prompts for task completion quality evaluation can be as follows: Figure 5As shown, the role definition can be: You are a rigorous quality assessment expert. Please score the task objectives and the agent's actual output from the following three dimensions. The assessment dimensions can include relevance, completeness, and usability. Each dimension can be scored from 1 to 5 points. The scoring criteria for relevance are: Does the output closely relate to the task theme? Is it irrelevant to the question? The scoring criteria for completeness are: Does the output fully cover the key points of the task requirements? The scoring criteria for usability are: Is the output clearly structured, logically coherent, and directly usable?

[0110] In this embodiment, by determining the task objective and agent output from the task execution records of completed tasks, a large language model can be used to accurately assess the task completion quality of completed tasks based on the task objective, agent output, and task completion quality assessment prompts, thereby obtaining a task completion quality score for the completed tasks.

[0111] In one exemplary embodiment, a task completion quality factor is determined based on the task completion quality scores of each of the multiple completed tasks, including:

[0112] Based on the task completion quality scores of multiple completed tasks, determine the agent's first completion quality score in the current quality assessment period, and obtain the agent's second completion quality score in the historical quality assessment period.

[0113] Determine the first weighting coefficient for the current quality assessment cycle and the second weighting coefficient for historical quality assessment cycles;

[0114] The first completion quality score and the second completion quality score are weighted according to the first weighting coefficient and the second weighting coefficient to obtain the task completion quality factor.

[0115] The quality assessment cycle refers to a predefined, fixed timeframe for periodically assessing the quality of an agent, which can be configured according to the actual application scenario. For example, the quality assessment cycle can be 12 hours, 24 hours, etc. The current quality assessment cycle refers to the quality assessment cycle that is currently being executed; it is the quality assessment cycle that is "occurring" in the time dimension, i.e., the most recently started quality assessment cycle. The historical quality assessment cycle refers to the quality assessment cycles that have already been completed in the past; it is the quality assessment cycle that has "ended" in the time dimension.

[0116] For example, based on the task completion quality scores of multiple completed tasks, the server can determine the agent's first completion quality score in the current quality assessment period, obtain the agent's second completion quality score in a historical quality assessment period, determine the first weighting coefficient for the current quality assessment period and the second weighting coefficient for the historical quality assessment period, and weight the first completion quality score and the second completion quality score according to the first weighting coefficient and the second weighting coefficient to obtain the task completion quality factor.

[0117] In practical applications, the server can first obtain a task quantity window, and then select multiple target tasks from the multiple completed tasks according to the task quantity window. Based on the task completion quality scores of each of the multiple target tasks, the first completion quality score of the agent in the current quality assessment period is determined. When obtaining the second completion quality score of the agent in historical quality assessment periods, the number of historical quality assessment periods obtained can be configured according to the actual application scenario. For example, the server can obtain only the second completion quality score of the most recent historical quality assessment period, or it can obtain the second completion quality scores of the most recent two historical quality assessment periods.

[0118] In practical applications, the first weighting coefficient for the current quality assessment period and the second weighting coefficient for historical quality assessment periods can be configured according to the actual application scenario. In this case, the first weighting coefficient is greater than the second weighting coefficient to give higher weight to recent task performance. These first and second weighting coefficients can also be calculated based on a task quantity window. Specifically, based on the task quantity window, a quality assessment smoothing factor can be determined, and then the first and second weighting coefficients can be allocated based on this smoothing factor. It is understandable that when obtaining the second completion quality scores for the two most recent historical quality assessment periods, the second weighting coefficients for different historical quality assessment periods can be the same or different.

[0119] In practical applications, when performing weighting, the task completion quality factor can be obtained by weighted summation or by weighted averaging. Taking the historical quality assessment period as a single period and obtaining the task completion quality factor by weighted summation as an example, the specific formula for calculating the task completion quality factor can be: Task completion quality factor A = α × first completion quality score + (1-α) × second completion quality score, where α is the first weight coefficient and 1-α is the second weight coefficient, that is, the sum of the first weight coefficient and the second weight coefficient is 1.

[0120] In this embodiment, by utilizing the task completion quality scores of multiple completed tasks, the first completion quality score of the agent in the current quality assessment period can be determined. Then, based on obtaining the second completion quality score of the agent in the historical quality assessment period and determining the first weight coefficient of the current quality assessment period and the second weight coefficient of the historical quality assessment period, the task completion quality can be accurately assessed by combining the task completion quality of the current quality assessment period and the task completion quality of the historical quality assessment period, thus obtaining the task completion quality factor.

[0121] In one exemplary embodiment, determining the agent's first completion quality score in the current quality assessment period based on the task completion quality scores of multiple completed tasks includes:

[0122] The number of times the agent is used in the current quality assessment cycle is counted, and the task quantity window is determined based on the number of times the agent is used;

[0123] Select multiple target tasks from a list of completed tasks, based on the task quantity window.

[0124] Based on the task completion quality scores of each of the multiple target tasks, the agent's first completion quality score in the current quality assessment period is determined.

[0125] For example, the server counts the number of times the agent is used in the current quality assessment period, and determines the task quantity window based on the number of times the agent is used. According to the task quantity window, the server selects the most recently completed target tasks from the multiple completed tasks, and determines the first completion quality score of the agent in the current quality assessment period based on the task completion quality scores of the multiple target tasks.

[0126] In specific applications, when determining the task quantity window based on the number of times the agent is used, the server can use the number of times the agent is used to query a predefined usage-task quantity window table to determine the task quantity window corresponding to the number of times the agent is used. This usage-task quantity window table has a predefined correspondence between different usage times and task quantity windows, and this correspondence can be configured according to the actual application scenario.

[0127] Understandably, when evaluating the quality of task completion, it is usually necessary to minimize the task repetition rate between different quality evaluation periods in order to improve the accuracy of task completion quality evaluation. Therefore, the server can use the number of times the agent is used to determine the task quantity window. By assigning a smaller task quantity window to an agent that is used less often, the task repetition rate of that agent between different quality evaluation periods can be effectively reduced, thereby improving the accuracy of task completion quality evaluation and avoiding the use of the same completed tasks for evaluation each time.

[0128] In specific applications, such as Figure 6 As shown, the server can calculate the average score of the task completion quality scores for each of the multiple target tasks, and use this average score as the agent's first completion quality score in the current quality evaluation period. The server can also calculate the median of the task completion quality scores for each of the multiple target tasks, and use this median as the agent's first completion quality score in the current quality evaluation period. Furthermore, the server can calculate the mode of the task completion quality scores for each of the multiple target tasks, and use this mode as the agent's first completion quality score in the current quality evaluation period.

[0129] In this embodiment, by counting the number of times the agent is used in the current quality assessment period, the number of times the agent is used can be used to determine the task quantity window. This can effectively reduce the task duplication rate of the agent between different quality assessment periods, thereby improving the accuracy of task completion quality assessment. By selecting multiple target tasks from multiple completed tasks according to the task quantity window, the task completion quality scores of each of the multiple target tasks can be combined to accurately determine the first completion quality score of the agent in the current quality assessment period.

[0130] In an exemplary embodiment, determining a first weighting coefficient for the current quality assessment period and a second weighting coefficient for historical quality assessment periods includes:

[0131] The number of times the agent is used in the current quality assessment cycle is counted, and the task quantity window is determined based on the number of times the agent is used;

[0132] Determine the quality assessment smoothing factor based on the task quantity window;

[0133] The quality assessment smoothing factor is determined as the first weighting coefficient for the current quality assessment period, and the second weighting coefficient for the historical quality assessment period is determined based on the first weighting coefficient and the predefined weighting coefficient.

[0134] The quality assessment smoothing factor determines the weight of the first completion quality score in calculating the overall task completion quality score. It acts as an adjustment knob controlling the weighting of old and new information, quantifying the trade-off between prioritizing the current quality assessment cycle and relying on historical quality assessment cycles in trend analysis. The predefined weight coefficient sum refers to the sum of the predefined first and second weight coefficients, which can be configured according to the actual application scenario. For example, the predefined weight coefficient sum can specifically be 1.

[0135] For example, the server counts the number of times the agent is used in the current quality assessment period, and determines the task quantity window based on the number of agent usages. Based on the task quantity window, a quality assessment smoothing factor is calculated, and the quality assessment smoothing factor is determined as the first weighting coefficient for the current quality assessment period. The server then determines the second weighting coefficient for the historical quality assessment period based on the first weighting coefficient and the predefined weighting coefficient.

[0136] In practical applications, the task quantity window can be viewed as an equivalent time window, i.e., the period length. The server can then calculate the quality assessment smoothing factor using a predefined conversion relationship between the period length and the smoothing factor. This conversion relationship can be configured according to the actual application scenario. For example, the specific conversion relationship could be: Quality assessment smoothing factor α = 2 / (N+1), where N is the task quantity window. Calculating the quality assessment smoothing factor in this way reflects both long-term stability and rapid response to recent performance fluctuations.

[0137] In practical applications, the server calculates the difference between a predefined weighting coefficient and the first weighting coefficient, and then determines the second weighting coefficient for each historical quality assessment period based on this difference. If there is only one historical quality assessment period, the server can use this difference as the second weighting coefficient for that period. If there are multiple historical quality assessment periods, the server can use the average of these differences as the second weighting coefficient for each period, or it can determine the second weighting coefficient for each period by proportionally sampling.

[0138] Understandably, when sampling proportionally, the proportion of the historical quality assessment period most recent to the current quality assessment period should be larger. The proportions of different historical quality assessment periods can be configured according to the actual application scenario. Taking two historical quality assessment periods as an example, including the first historical quality assessment period and the second historical quality assessment period, with a difference of 0.6, the proportion of the first historical quality assessment period most recent to the current quality assessment period can be 0.7, and the second weighting coefficient of this first historical quality assessment period can be 0.42. The proportion of the second historical quality assessment period, which is further away from the current quality assessment period, can be 0.3, and the second weighting coefficient of this first historical quality assessment period can be 0.18.

[0139] In this embodiment, by statistically analyzing the number of times the agent is used in the current quality assessment period, the number of agent usages can be used to determine the task quantity window. This effectively reduces the task duplication rate between different quality assessment periods, thereby improving the accuracy of task completion quality assessment. By determining the quality assessment smoothing factor based on the task quantity window, the periodic characteristics of the task quantity window can be used to accurately convert the quality assessment smoothing factor. Furthermore, the quality assessment smoothing factor can be used to determine the first weighting coefficient for the current quality assessment period and the second weighting coefficient for historical quality assessment periods, which can reflect long-term stability and quickly respond to recent performance fluctuations.

[0140] In one exemplary embodiment, based on the agent's user preference operation records, a user preference evaluation is performed on the agent to determine user preference factors, including:

[0141] Based on the user preference operation records of the intelligent agent, count the number of times the intelligent agent is activated and the number of times the intelligent agent is deactivated;

[0142] Based on the number of times the agent is activated and deactivated, user preference is evaluated to determine the user preference factor.

[0143] The number of times an agent is activated refers to the number of times the agent is activated by the user within the preference operation statistics period. The number of times an agent is deactivated refers to the number of times the agent is deactivated by the user within the preference operation statistics period. The preference operation statistics period is a predefined fixed time frame for statistical analysis of preference operations on agents, which can be configured according to the actual application scenario. For example, the preference operation statistics period can be the past week.

[0144] For example, the server can count the number of times the agent is enabled and disabled based on the agent's user preference operation records, and then evaluate the agent's user preferences based on the number of times the agent is enabled and disabled to determine the user preference factor.

[0145] In practical applications, when evaluating user preferences for an agent, the server calculates the total number of times the agent is enabled and disabled, and also calculates the difference between the number of times the agent is enabled and disabled. The user preference factor is then calculated using the difference and the total number of times the agent is disabled.

[0146] In a specific application, the server can calculate the ratio of the difference in the number of occurrences to the total number of occurrences, and use this ratio as the user preference factor. When the number of times the agent is activated and deactivated is relatively small, to reduce extreme estimations, the server introduces a smoothing constant to calculate the user preference factor. In this case, the server first calculates the sum of the total number of occurrences and the value of the smoothing constant, then calculates the ratio of the difference in the number of occurrences to this sum, and uses this ratio as the user preference factor. This smoothing constant can be configured according to the actual application scenario.

[0147] For example, taking the introduction of a smoothing constant as an example, the specific formula for calculating the user preference factor can be: User preference factor B = (Number of times the agent is activated - Number of times the agent is deactivated) / (Number of times the agent is activated + Number of times the agent is deactivated + Smoothing constant). This formula normalizes the net increase in the number of activations, avoids deviations in absolute values, and the value range tends to be [-1, 1].

[0148] In this embodiment, based on the user preference operation records of the agent, the number of times the agent is enabled and disabled can be counted. Then, the number of times the agent is enabled and disabled can be used to accurately evaluate the user preference of the agent and determine the user preference factor.

[0149] In an exemplary embodiment, the agent is quantitatively evaluated based on a task completion quality factor and a user preference factor to obtain the quantitative evaluation result of the agent, including:

[0150] Based on the agent's usage records, the agent's usage trend is evaluated to determine the usage trend factors;

[0151] Based on task completion quality factors, user preference factors, and usage trend factors, the agent is quantitatively evaluated to obtain the quantitative evaluation results of the agent.

[0152] The usage trend factor refers to an indicator used to evaluate agents based on their usage trends. For example, a usage trend factor could be a usage trend score. In essence, analyzing usage trends helps determine the growth momentum of agents, identifying potential high-potential agents and allowing newly launched agents to gain attention in the market. Taking the usage trend score as an example, a higher score indicates a gradually increasing number of users, while a lower score indicates a gradually decreasing number of users. This allows for the analysis of agents with recently increasing user numbers and a growth trend.

[0153] For example, when performing a quantitative evaluation of an agent, the server will evaluate the agent's usage trend based on the agent's usage records, determine the usage trend factor, and then combine the task completion quality factor, user preference factor, and usage trend factor to perform a quantitative evaluation of the agent and obtain the quantitative evaluation result of the agent.

[0154] In practical applications, the server can determine the number of times the agent is used in each usage statistical period based on the agent's usage records. This usage frequency can then be used to evaluate the agent's usage trend and determine usage trend factors. Based on these trend factors, the server can quantitatively evaluate the agent by weighting task completion quality factors, user preference factors, and usage trend factors, thus obtaining a quantitative evaluation result. The weighting coefficients for the first factor (task completion quality), the second factor (user preference), and the third factor (usage trend) can be configured according to the specific application scenario.

[0155] In this embodiment, by evaluating the usage trend of the agent based on its usage records, and determining the usage trend factor, a quality assessment from the usage trend dimension can be achieved. Furthermore, by combining the task completion quality factor, user preference factor, and usage trend factor, an accurate quantitative assessment of the agent can be achieved, resulting in a quantitative assessment result of the agent.

[0156] In one exemplary embodiment, based on the agent's agent usage records, a usage trend assessment is performed on the agent to determine usage trend factors, including:

[0157] Based on the agent's usage records, determine the agent's first usage count in the current usage statistical period and the agent's second usage count in the historical usage statistical period;

[0158] Based on the first and second usage counts, the usage trend of the agent is evaluated to determine the usage trend factor.

[0159] The usage statistics period refers to a predefined, fixed time frame for periodically counting the number of times an agent is used. This can be configured according to the actual application scenario. For example, the usage statistics period could be the past week. The current usage statistics period refers to the usage statistics period currently in progress; it is the usage statistics period that is "occurring" in the time dimension, i.e., the most recently started usage statistics period. The historical usage statistics period refers to the usage statistics periods that have already been completed in the past; it is the usage statistics period that has "ended" in the time dimension.

[0160] For example, based on the agent's usage records, the server can determine the first number of times the agent is used in the current usage statistics period and the second number of times the agent is used in the historical usage statistics period. Then, based on the first and second usage counts, the server can evaluate the agent's usage trend and determine the usage trend factor.

[0161] In practical applications, the server can calculate the logarithmic difference between the first and second usage counts, using this difference as a usage trend factor. Understandably, using the logarithmic difference to measure usage growth rate can smoothly handle usage counts of different magnitudes. The number of historical usage statistical periods can be configured according to the actual application scenario. For example, the server can determine only the second usage count for the most recent historical usage statistical period, or it can determine the second usage count for the two most recent historical usage statistical periods.

[0162] In a specific application, when there are multiple historical usage periods, the server needs to calculate the logarithmic difference between the first usage count and multiple second usage counts to obtain the usage trend factor. Taking a single historical usage period as an example, the formula for calculating the usage trend factor can be: Usage Trend Factor C = log(First Usage Count + 1) - log(Second Usage Count + 1). Here, the +1 is mainly to make the logarithmic value meaningful, preventing undefined log0 when the usage count is 0, and to mitigate the drastic changes in the input of the log function when the usage count is very low, making the converted value more stable.

[0163] In this embodiment, by utilizing the agent's usage records, it is possible to determine the first number of times the agent is used in the current usage statistical period and the second number of times the agent is used in the historical usage statistical period. Then, the first and second usage counts can be used to evaluate the agent's usage trend, determine the usage trend factor, and achieve quality evaluation from the dimension of usage trend.

[0164] In an exemplary embodiment, the agent is quantitatively evaluated based on task completion quality factors, user preference factors, and usage trend factors to obtain the quantitative evaluation results of the agent, including:

[0165] Based on the release time of the intelligent agent in the intelligent agent market, the timeliness of the intelligent agent is evaluated, and the timeliness factor is determined;

[0166] Based on task completion quality factors, user preference factors, usage trend factors, and timeliness factors, the agent is quantitatively evaluated to obtain the quantitative evaluation results of the agent.

[0167] Timeliness assessment refers to the process of evaluating the freshness, timeliness, and effective value of an agent within a specific time window. A timeliness factor is an indicator used to assess the timeliness of an agent. For example, a timeliness factor can be a timeliness score; a higher score indicates greater freshness. For instance, if the timeliness score is a normalized score, a newly released agent could have a timeliness score of 1.

[0168] For example, the server can determine the number of days since the agent was released in the agent market based on the agent's release time. Then, it can use the number of days since the agent was released to evaluate the timeliness of the agent, determine the timeliness factor, and then combine the task completion quality factor, user preference factor, usage trend factor and timeliness factor to conduct a quantitative evaluation of the agent and obtain the quantitative evaluation result of the agent.

[0169] In practical applications, the server can calculate the timeliness factor based on the number of days since the agent was released and a predefined timeliness evaluation formula. The formula for calculating the timeliness factor is: Timeliness factor D = D = e^(-λ * number of days since release). Where λ is the decay coefficient, which can be configured according to the actual application scenario. When a new release is made, the number of days since release = 0, then D≈1, and it decays exponentially over time.

[0170] In practical applications, based on a determined timeliness factor, the server can quantitatively evaluate the agent by weighting task completion quality, user preference, usage trend, and timeliness factors to obtain the agent's quantitative evaluation result. The weight coefficients for the first factor (task completion quality), the second factor (user preference), the third factor (usage trend), and the third factor (timeliness) can be configured according to the actual application scenario. For example, the first factor weight coefficient could be 0.5, the second factor weight coefficient 0.3, the third factor weight coefficient (usage trend) 0.15, and the third factor weight coefficient (timeliness) 0.05.

[0171] In this embodiment, by evaluating the timeliness of the agent based on its release time in the agent market, a timeliness factor is determined, enabling quality assessment from the timeliness dimension. Furthermore, by combining task completion quality factors, user preference factors, usage trend factors, and timeliness factors, an accurate quantitative assessment of the agent can be achieved, resulting in a quantitative assessment result of the agent.

[0172] In an exemplary embodiment, the agent is quantitatively evaluated based on a task completion quality factor and a user preference factor to obtain the quantitative evaluation result of the agent, including:

[0173] Based on the release time of the intelligent agent in the intelligent agent market, the timeliness of the intelligent agent is evaluated, and the timeliness factor is determined;

[0174] Based on task completion quality factors, user preference factors, and timeliness factors, the agent is quantitatively evaluated to obtain the quantitative evaluation results of the agent.

[0175] For example, the server can determine the number of days since the agent was released based on its release time in the agent market. Then, it can use the number of days since the agent was released to evaluate the timeliness of the agent, determine the timeliness factor, and then combine the task completion quality factor, user preference factor and timeliness factor to conduct a quantitative evaluation of the agent and obtain the quantitative evaluation result of the agent.

[0176] In practical applications, based on the determined timeliness factor, the server can quantitatively evaluate the agent by weighting the task completion quality factor, user preference factor, and timeliness factor to obtain the quantitative evaluation result of the agent. The weight coefficients of the first factor of the task completion quality factor, the second factor of the user preference factor, and the third factor of the timeliness factor can be configured according to the actual application scenario.

[0177] In this embodiment, by evaluating the timeliness of the agent based on its release time in the agent market, a timeliness factor is determined, enabling quality assessment from the timeliness dimension. Furthermore, by combining task completion quality factor, user preference factor, and timeliness factor, an accurate quantitative assessment of the agent can be achieved, resulting in a quantitative assessment result of the agent.

[0178] In one exemplary embodiment, the intelligent agent quality assessment method of this application is applied to an intelligent ranking system in the intelligent agent market as an example to illustrate the intelligent agent quality assessment method of this application. Specifically, three application scenarios of the intelligent agent quality assessment method of this application are provided.

[0179] In Scenario 1, the agent quality assessment method proposed in this application can achieve efficient discovery. Specifically, user Xiao Wang needs an "application best-selling copywriting generation agent." He no longer needs to browse dozens of pages of lists or rely on unreliable advertisements because, in the "Highly Recommended" section on the agent marketplace homepage, the intelligent sorting system automatically recommends several recognized high-performing agents based on high task completion rates and positive user reviews. Xiao Wang successfully completed the task after trying it.

[0180] In Scenario 2, the agent quality assessment method proposed in this application can be used to help emerging agents. Specifically, user Xiao Li is a tech enthusiast who enjoys trying new tools. He discovered a newly released "AI conversation prompt word optimization agent" in the "Potential Rising Stars" section of the agent marketplace. After trying it, he was amazed by its performance and immediately enabled it. His activation directly contributed positive feedback of "user preference factor" to the agent, helping it rise in the overall ranking and be seen by more users more quickly.

[0181] In Scenario 3, the agent quality assessment method proposed in this application can help purify the market. User Xiao Zhao discovered that an "industry data query agent" he had enabled six months ago was frequently malfunctioning and providing outdated information. He promptly shut it down. This shutdown caused the agent's score to drop and its ranking to fall significantly, effectively alerting other users and prompting the developer to either update the product or be naturally eliminated by the market.

[0182] Understandably, applying the agent quality assessment method of this application to the agent market will allow users to perceive an agent market that is "intelligent and vibrant." The homepage recommendations and search rankings of the agent market will no longer be static "hot lists," but rather "reputation lists" and "rising star lists" that dynamically change based on user collective choices and the actual capabilities of agents. Users will find high-quality agents much more efficiently, and each choice they make (enabling or disabling) is like voting for the "purification" of the entire market, collectively shaping an increasingly better tool ecosystem.

[0183] With the explosive growth of the AI ​​agent ecosystem, the AI ​​agent market faces two core challenges: "information overload" and "discovery efficiency." Related technologies primarily employ sorting by release time or by total usage / download. Sorting by release time means arranging AI agents in reverse chronological order of release or update date, with the newest agents appearing at the top. While this provides exposure for all new AI agents, it completely fails to differentiate quality. Users are overwhelmed by a massive influx of unfiltered new AI agents, resulting in extremely high screening costs, and high-quality older AI agents quickly disappear from the rankings. Sorting by total usage / download means ranking AI agents in descending order of their historical total usage or total downloads, with the most used AI agents at the top. This easily leads to a "Matthew effect" or "traffic monopoly." Early-released AI agents, even those with mediocre capabilities, can maintain their dominance due to cumulative effects, making it difficult for new, high-quality AI agents to surpass them. Furthermore, usage volume does not directly equate to quality; some entertaining but impractical AI agents may also rank highly.

[0184] In summary, the relevant technologies suffer from the following drawbacks: First, ranking is disconnected from quality: the ranking criteria (release time, total usage, etc.) are unrelated to the actual problem-solving capabilities of the agents, making it impossible to distinguish the true capabilities of agents and preventing high-quality agents from being showcased. Second, high cold-start barriers: newly released high-quality agents struggle to overcome early traffic barriers, lacking effective and sustainable initial exposure opportunities and evaluation mechanisms, thus hindering their launch and stifling innovation. The cold-start problem refers to the initial stage where newly released or newly entered agents are at a significant disadvantage in ranking and discovery due to a lack of historical usage data, user feedback, and initial exposure. Third, the ecosystem is prone to stagnation: early agents dominate the rankings due to their first-mover advantage, easily forming a fixed top-tier structure, suppressing innovation, causing the market to lose vitality, and hindering the healthy development of the ecosystem. Fourth, a lack of real-time feedback and adaptability: slow ranking updates prevent rapid responses to real user feedback (such as enabling a useful agent or disabling a poor one) and dynamic market changes. These problems severely hinder users from finding the best tools and also affect developers' motivation to optimize their products.

[0185] Based on this, this application proposes a method for evaluating the quality of intelligent agents, which includes a competitive elimination algorithm for the intelligent agent market based on multi-dimensional dynamic evaluation and user feedback. Its core idea is to construct a "data-driven, quality-first, positive feedback loop" intelligent ranking system, shifting market dominance from "seniority" and "traffic" back to "ability" and "user satisfaction." The competitive elimination algorithm refers to a set of algorithms and logic based on a dynamic scoring model that continuously updates the ranking of intelligent agents in the market, giving high-quality agents priority exposure while gradually marginalizing low-quality agents, thereby achieving optimal allocation of market resources.

[0186] The key technical aspects of the intelligent agent quality assessment method in this application mainly include the following:

[0187] First, a scientific capability assessment system: Introducing a large language model as an automated and quantifiable "quality judge," through carefully designed assessment prompts and multi-dimensional scoring standards, the vague concept of "easy to use" is transformed into a precise "task completion quality" score, providing an objective and reliable core basis for ranking and ensuring that the ranking is highly correlated with the actual capabilities of the agent.

[0188] Second, real-time user feedback integration: Every user's action of activating or deactivating an agent is regarded as a key signal, and a reasonable mathematical model (such as normalized net increase in activations) is designed to incorporate it into the comprehensive score, so that the ranking results can quickly reflect the true preferences of collective intelligence.

[0189] Third, establish a fair cold start and growth channel: provide initial traffic to new intelligent agents through "freshness factor (i.e., timeliness factor)" and exclusive exposure positions, and identify and promote potential intelligent agents through "growth momentum factor (i.e. usage trend factor)".

[0190] Fourth, a dynamic scoring model that balances fairness and efficiency: A multi-factor weighted scoring formula is designed, assigning the highest weight to "task completion quality" to ensure quality orientation. It also incorporates factors such as "growth momentum (i.e., usage trend)" and "release freshness (i.e., timeliness)" to provide a fair starting point for new agents and identify agents with potential for growth. This dynamic scoring model can be used to periodically calculate the comprehensive competitiveness score of each agent, i.e., the quantitative evaluation result.

[0191] Fifth, a closed-loop competitive elimination process: Establish a complete automated process from "data collection -> factor calculation -> dynamic scoring -> ranking update -> impact exposure". Higher rankings lead to higher exposure, higher exposure generates more data, which in turn optimizes the scoring, forming a self-reinforcing closed loop of survival of the fittest.

[0192] In one exemplary embodiment, the agent quality evaluation method of this application is applied to the competitive elimination algorithm engine of an intelligent ranking system. It can function as an independent background service, continuously processing real-time data streams from user preference operations and task execution, dynamically calculating and updating agent rankings. The core architecture of this competitive elimination algorithm engine can be specifically as follows: Figure 7 As shown, it includes a data collection and processing module, a factor calculation module, a large language model evaluation service, a dynamic scoring model, and a ranking and updating module. The following section will combine these modules... Figure 7 The functions and technical implementation details of each module are explained.

[0193] In practical applications, the primary responsibility of the data collection and processing module is to subscribe to and receive event messages in real time from the front-end interface and the task execution platform, namely, user preference operation record streams and task execution record streams. The front-end interface is the user interface, where users can initiate user preference operations such as enabling or disabling the agent; the resulting user preference operation records are acquired by the data collection and processing module. After the agent completes a task, the task execution platform generates a task execution record, which is also acquired by the data collection and processing module. In other words, the input sources for the data collection and processing module are task execution records and user preference operation records, and the output is cleaned and formatted standard data records, stored in a temporary message queue or database for subsequent computation.

[0194] The specific format of the task execution record can be: {task_id (task identifier), user_id (user identifier), agent_id (agent identifier), task_description (task objective), agent_output (agent output result), timestamp (timestamp), status (task status)}. The specific format of the user preference operation record can be: {user_id (user identifier), agent_id (agent identifier), action: "enable / disable" (preference operation: enable / disable), timestamp (timestamp)}.

[0195] In practical applications, the main responsibility of the factor calculation module is to periodically calculate a series of quantifiable evaluation factors for each agent, based on the data output from the data collection and preprocessing module and the basic information stored in the agent meta-database (which can be used to identify agents). Specifically, the evaluation factors may include task completion quality factors, user preference factors, usage trend factors, and timeliness factors.

[0196] Among them, the task completion quality factor is the core quality indicator. In the calculation, for each completed task, it is necessary to first call the large language model evaluation service to obtain a multi-dimensional standardized total score (such as normalized to the 0-1 range), that is, the task completion quality score. Then, based on the task completion quality scores of multiple completed tasks, the task completion quality factor is calculated.

[0197] In a specific application, the exponential moving average method can be used to calculate the task completion quality factor, giving higher weight to recent task performance, thereby agilely reflecting dynamic changes in the agent's capabilities. Specifically, the EMA (Exponential Moving Average) value of the agent's score within the most recent equivalent time window (i.e., a window of task quantity, such as N=50 tasks) can be taken, as follows: Figure 8 As shown, the calculation formula for the task completion quality factor can be: Task completion quality factor A = α × the agent's first completion quality score in the current quality assessment period + (1-α) × the second completion quality score in the historical quality assessment period, where α is the first weight coefficient and 1-α is the second weight coefficient, that is, the sum of the first weight coefficient and the second weight coefficient is 1, and the historical quality assessment period can be the previous quality assessment period.

[0198] Among them, the user preference factor is a direct reflection of user preferences. For example... Figure 8 As shown, the formula for calculating the user preference factor is: User preference factor B = (Number of times agent is enabled - Number of times agent is disabled) / (Number of times agent is enabled + Number of times agent is disabled + Smoothing constant). This formula normalizes the net increase in the number of enabled agents, avoiding deviations in absolute values, and the value range tends to be [-1, 1].

[0199] Among these, trend factors are primarily used to identify potential stocks. For example... Figure 8 As shown, the formula for calculating the trend factor is: Trend factor C = log(first usage count + 1) - log(second usage count + 1), where the usage count + 1 is mainly to make the logarithmic value meaningful, to prevent the occurrence of undefined log0 when the usage count is 0, and to mitigate the drastic changes in the input of the log function when the usage count is very low, so that the converted value is more stable.

[0200] Among them, the timeliness factor is mainly used to give new intelligent agents initial exposure. For example, Figure 8 As shown, the formula for calculating the timeliness factor is: Timeliness factor D = D = e^(-λ * number of days since release). Where λ is the decay coefficient, which can be configured according to the actual application scenario. When a new release is made, the number of days since release = 0, then D≈1, and it decays exponentially over time.

[0201] In practical applications, the main responsibility of the large language model evaluation service is to automatically and objectively evaluate the task completion quality of the agent. The specific task completion quality evaluation prompts can be as follows: Figure 5As shown, based on the task completion quality assessment prompts, a task completion quality assessment can be performed on the completed task, resulting in a total quality assessment score. To eliminate the incomparability caused by differences in data units (dimensions) and numerical ranges (order of magnitude) among different factors, and to ensure that all factors can be analyzed and modeled on a fair benchmark, the large language model assessment service standardizes the total quality assessment score and then uses the standardized total score as the task completion quality score. The standardization method can be configured according to the actual application scenario. For example, the standardization method can be normalized to the 0-1 range, i.e., the task completion quality score can be the ratio of the total quality assessment score to the full score.

[0202] In practical applications, the main responsibility of the dynamic scoring model is to synthesize multiple factors into a comprehensive score, that is, to quantitatively evaluate the agent based on multiple factors and obtain a quantitative score result for the agent. Specifically, the formula for calculating the quantitative score result can be: Quantitative score result = w1 * A + w2 * B + w3 * C + w4 * D, where w1 is the weight of the first factor, A is the task completion quality factor, w2 is the weight of the second factor, B is the user preference factor, w3 is the weight of the third factor of the usage trend factor, C is the usage trend factor, w4 is the weight of the third factor of the timeliness factor, and D is the timeliness factor.

[0203] Understandably, factor weights can be configured according to the actual application scenario. For example, if task completion quality is prioritized, the factor weights could be: the first factor weight coefficient could be 0.5, the second factor weight coefficient could be 0.3, the third factor weight coefficient using the trend factor could be 0.15, and the third factor weight coefficient using the timeliness factor could be 0.05.

[0204] In practical applications, the main responsibility of the sorting and updating module is to rank all agents based on their comprehensive scores and update the market display order. The mechanism involves performing batch calculations and sorting at fixed intervals (e.g., N hours, where N is a positive integer and can be configured according to the actual application scenario). The sorting results (a list of agent identifiers) are written to the ranking results database. The output is that the agent market front-end application periodically or in real-time pulls the latest rankings from the database for interface display.

[0205] In an exemplary embodiment, a flowchart is used to illustrate the complete closed-loop logic of the agent quality evaluation method of this application from event triggering to ranking update, particularly the connection between cold start and the main loop, as shown below. Figure 9 As shown, the specific steps include:

[0206] 1. Cold Start Phase: After a new agent is released, it first enters the cold start pool. The system provides it with initial exposure opportunities through specific lists such as "Newest Agents" / "Potential Rising Stars", thereby obtaining the first batch of user and task data.

[0207] 2. Event-driven: The main loop of the algorithm is driven by two types of core events: a) user enable / disable actions; b) agent task completion.

[0208] Specifically, in the case of an event type indicating a user's enable / disable action, the user's preference operation record will be recorded, and the calculation cycle will be awaited. In the case of an event type indicating an agent's task completion, the task execution record of the completed task will be recorded.

[0209] 3. Asynchronous evaluation: For completed tasks, the system will asynchronously call the large language model evaluation service and use the returned score to update the agent's "task completion quality factor".

[0210] 4. Periodic Calculation: The system does not recalculate the global ranking after each event, but waits for a calculation period (such as 1 hour) to arrive before batch processing all the data accumulated within that period.

[0211] 5. Factor Calculation and Scoring: At each calculation cycle point, the system calculates the latest value of each factor for all agents and substitutes it into the weighted formula to obtain the comprehensive score, i.e., the quantitative evaluation result.

[0212] 6. Sorting and Updating: All agents are globally sorted (specifically, in descending order) based on the new comprehensive scores to obtain the agent quality evaluation results, which are then persisted to the database. The market front-end is then updated and displayed accordingly.

[0213] 7. Continuous iteration: After the update is completed, the system continues to detect new events and enters the next cycle, thereby achieving continuous and dynamic optimization of the ranking.

[0214] In a specific application, taking an agent market comprising Agent 1, Agent 2, Agent 3, Agent 4, Agent 5, and Agent 6, with Agent 5 and Agent 6 being newly released agents, as an example, ... Figure 10 As shown, the front-end display interface of the intelligent agent market can be specifically as follows: Figure 10 As shown, the list includes a ranked list of agents and a list of the latest agents. Agents 5 and 6, being newly released agents, are ranked last in the agent ranking list but are recommended in the list of the latest agents. If agents 5 and 6 are discovered and used by more users, and agent 4 is disabled by more users, the agent ranking list may change after a new computation cycle. Figure 11As shown, Agent 5 and Agent 6 are listed before Agent 4. Agent 5 is listed before Agent 6 because it is used by more users and performs better tasks. In this way, users can promptly discover the newly released Agent 5 and Agent 6 in the agent market and gradually abandon Agent 4, which has poor task performance.

[0215] It should be noted that the agent quality evaluation method of this application, by introducing a scientific, dynamic, and multi-dimensional competitive elimination algorithm, produces the following significant beneficial effects compared to related technologies:

[0216] First, it fundamentally improves market discovery efficiency and user experience: users can quickly and accurately discover truly efficient and reliable intelligent agents, greatly reducing search and trial-and-error costs, thereby increasing user satisfaction and loyalty to the platform.

[0217] Second, it activates the vitality of ecological innovation and establishes a level playing field: providing a clear and fair upward path for high-quality new intelligent agents, effectively solving the cold start problem. This incentivizes developers to continuously invest in product optimization, forming a powerful positive cycle of "high quality -> high rating -> high exposure -> more feedback -> even higher quality," preventing the market from being monopolized by early, low-quality intelligent agents.

[0218] Third, it enables automated and intelligent selection of the fittest: through data-driven algorithms, valuable exposure resources are automatically allocated to high-performing intelligent agents, while low-quality and outdated intelligent agents are allowed to naturally disappear, greatly reducing the cost of manual operation of the platform and maintaining the health and competitiveness of the entire ecosystem.

[0219] Fourth, we will build a dynamic platform that gathers collective wisdom: transforming each user's behavior into a "micro-contribution" to optimizing market order, so that the final ranking results embody crowdsourced intelligence, making the market "smarter" and more accurate with use.

[0220] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0221] Based on the same inventive concept, this application also provides an agent quality assessment device for implementing the agent quality assessment method described above. The solution provided by this device is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more embodiments of the agent quality assessment device provided below can be found in the limitations of the agent quality assessment method described above, and will not be repeated here.

[0222] In one exemplary embodiment, such as Figure 12 As shown, an agent quality assessment device is provided, comprising: an agent determination module 1202, a task quality assessment module 1204, a user preference assessment module 1206, a quantitative assessment module 1208, and a ranking module 1210, wherein:

[0223] The agent determination module 1202 is used to determine multiple agents from the agent market;

[0224] The task quality assessment module 1204 is used to assess the quality of the task completed by each agent based on the agent's task execution record and determine the task completion quality factor.

[0225] The user preference assessment module 1206 is used to assess the user preferences of the agent based on the agent's user preference operation records and determine the user preference factors.

[0226] The quantitative evaluation module 1208 is used to quantitatively evaluate the agent based on the task completion quality factor and the user preference factor, and obtain the quantitative evaluation result of the agent.

[0227] The sorting module 1210 is used to sort multiple agents based on their respective quantitative evaluation results to obtain agent quality evaluation results; the agent quality evaluation results are used to update the order of multiple agents in the agent market.

[0228] The aforementioned agent quality assessment device, based on identifying multiple agents from the agent market, evaluates the quality of task completion for each agent by analyzing its task execution records and determining a task completion quality factor. This enables quality assessment from the task quality dimension. Furthermore, by analyzing the agent's user preference operation records and determining a user preference factor, it achieves quality assessment from the user preference dimension. By combining the task completion quality factor and the user preference factor, and considering both dimensions, it achieves accurate quantitative assessment of the agent, yielding a quantitative assessment result. Based on these quantitative assessment results, the device can then rank the multiple agents to obtain the final agent quality assessment result, thus improving the accuracy of agent quality assessment.

[0229] In an exemplary embodiment, the task quality assessment module is further configured to, for each agent, determine the task execution records of each of the agent's multiple completed tasks based on the agent's task execution records, perform task completion quality assessment on each completed task based on the task execution records of the completed task, obtain a task completion quality score for the completed task, and determine a task completion quality factor based on the task completion quality scores of each of the multiple completed tasks.

[0230] In an exemplary embodiment, the task quality assessment module is further configured to, for each completed task, determine the task objective and agent output from the task execution record of the completed task, and, through a large language model, assess the task completion quality of the completed task based on the task objective, agent output, and task completion quality assessment prompts, thereby obtaining a task completion quality score for the completed task.

[0231] In an exemplary embodiment, the task quality assessment module is further configured to determine the first completion quality score of the agent in the current quality assessment period based on the task completion quality scores of each of the multiple completed tasks, and obtain the second completion quality score of the agent in the historical quality assessment period, determine the first weighting coefficient of the current quality assessment period and the second weighting coefficient of the historical quality assessment period, and weight the first completion quality score and the second completion quality score according to the first weighting coefficient and the second weighting coefficient to obtain the task completion quality factor.

[0232] In an exemplary embodiment, the task quality assessment module is further configured to count the number of times the agent is used in the current quality assessment period, and based on the number of times the agent is used, determine a task quantity window, select multiple target tasks from multiple completed tasks according to the task quantity window, and determine the first completion quality score of the agent in the current quality assessment period based on the task completion quality scores of the multiple target tasks.

[0233] In an exemplary embodiment, the task quality assessment module is further configured to count the number of times the agent is used in the current quality assessment period, and determine the task quantity window based on the number of times the agent is used, determine the quality assessment smoothing factor based on the task quantity window, determine the quality assessment smoothing factor as the first weight coefficient of the current quality assessment period, and determine the second weight coefficient of the historical quality assessment period based on the first weight coefficient and the predefined weight coefficient.

[0234] In an exemplary embodiment, the user preference evaluation module is further configured to count the number of times the agent is enabled and the number of times the agent is disabled based on the agent's user preference operation records, and to evaluate the agent's user preferences based on the number of times the agent is enabled and the number of times the agent is disabled, thereby determining the user preference factor.

[0235] In an exemplary embodiment, the quantitative evaluation module is further configured to evaluate the usage trend of the agent based on the agent's usage records, determine the usage trend factor, and perform quantitative evaluation of the agent based on the task completion quality factor, user preference factor, and usage trend factor to obtain the quantitative evaluation result of the agent.

[0236] In an exemplary embodiment, the quantitative evaluation module is further configured to determine the first number of times the agent is used in the current usage statistical period and the second number of times the agent is used in the historical usage statistical period based on the agent's agent usage records, and to evaluate the usage trend of the agent based on the first number of times and the second number of times, and to determine the usage trend factor.

[0237] In an exemplary embodiment, the quantitative evaluation module is further configured to evaluate the timeliness of the agent based on its release time in the agent market, determine the timeliness factor, and perform a quantitative evaluation of the agent based on the task completion quality factor, user preference factor, usage trend factor, and timeliness factor to obtain the quantitative evaluation result of the agent.

[0238] In an exemplary embodiment, the quantitative evaluation module is further configured to evaluate the timeliness of the agent based on the agent's release time in the agent market, determine the timeliness factor, and perform quantitative evaluation of the agent based on the task completion quality factor, user preference factor, and timeliness factor to obtain the quantitative evaluation result of the agent.

[0239] Each module in the aforementioned intelligent agent quality assessment device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0240] In one exemplary embodiment, a computer device is provided, which can be a server or a terminal. Taking the computer device as a server as an example, its internal structure diagram can be as follows: Figure 13As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data such as task execution records and user preference operation records of the intelligent agent. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements an intelligent agent quality evaluation method.

[0241] Those skilled in the art will understand that Figure 13 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0242] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0243] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0244] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0245] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0246] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0247] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0248] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for evaluating the quality of an intelligent agent, characterized in that, The method includes: Identify multiple agents from the agent market; For each agent, the quality of the agent's task completion is evaluated based on the agent's task execution record, and a task completion quality factor is determined. Based on the user preference operation records of the intelligent agent, the user preference is evaluated to determine the user preference factor; The agent is quantitatively evaluated based on the task completion quality factor and the user preference factor to obtain the quantitative evaluation result of the agent. Based on the quantitative evaluation results of each of the multiple agents, the multiple agents are ranked to obtain the agent quality evaluation results; the agent quality evaluation results are used to update the ranking order of the multiple agents in the agent market.

2. The method according to claim 1, characterized in that, For each agent, the quality of task completion is evaluated based on the agent's task execution record to determine a task completion quality factor, including: For each of the intelligent agents, based on the task execution records of the intelligent agent, determine the task execution records of each of the multiple completed tasks of the intelligent agent; For each completed task, based on the task execution record of the completed task, the task completion quality is evaluated to obtain the task completion quality score of the completed task; Based on the task completion quality scores of the multiple completed tasks, a task completion quality factor is determined.

3. The method according to claim 2, characterized in that, For each completed task, based on the task execution record of the completed task, a task completion quality assessment is performed on the completed task to obtain a task completion quality score, including: For each completed task, the task objective and the agent's output result are determined from the task execution record of the completed task; Using a large language model, based on the task objective, the agent's output, and task completion quality assessment prompts, the completed task is evaluated to obtain a task completion quality score.

4. The method according to claim 2, characterized in that, The process of determining the task completion quality factor based on the individual task completion quality scores of the multiple completed tasks includes: Based on the task completion quality scores of each of the multiple completed tasks, the agent's first completion quality score in the current quality assessment period is determined, and the agent's second completion quality score in the historical quality assessment period is obtained. Determine the first weighting coefficient for the current quality assessment period and the second weighting coefficient for the historical quality assessment periods; Based on the first weighting coefficient and the second weighting coefficient, the first completion quality score and the second completion quality score are weighted to obtain the task completion quality factor.

5. The method according to claim 4, characterized in that, The step of determining the agent's first completion quality score in the current quality assessment period based on the task completion quality scores of the plurality of completed tasks includes: The number of times the agent is used in the current quality assessment period is counted, and the task quantity window is determined based on the number of times the agent is used; According to the task quantity window, select multiple target tasks from the multiple completed tasks; Based on the task completion quality scores of the multiple target tasks, the agent's first completion quality score in the current quality assessment period is determined.

6. The method according to claim 4, characterized in that, Determining the first weighting coefficient for the current quality assessment period and the second weighting coefficient for the historical quality assessment periods includes: The number of times the agent is used in the current quality assessment period is counted, and the task quantity window is determined based on the number of times the agent is used; Based on the task quantity window, determine the quality assessment smoothing factor; The quality assessment smoothing factor is determined as the first weighting coefficient for the current quality assessment period, and the second weighting coefficient for the historical quality assessment period is determined based on the first weighting coefficient and the predefined weighting coefficient.

7. The method according to claim 1, characterized in that, The step of evaluating the user preferences of the intelligent agent based on its user preference operation records and determining user preference factors includes: Based on the user preference operation records of the intelligent agent, the number of times the intelligent agent was activated and the number of times the intelligent agent was deactivated were counted. Based on the number of times the agent is activated and the number of times it is deactivated, user preference is evaluated for the agent to determine the user preference factor.

8. The method according to any one of claims 1 to 7, characterized in that, The step of quantitatively evaluating the agent based on the task completion quality factor and the user preference factor to obtain the quantitative evaluation result of the agent includes: Based on the agent's agent usage records, the agent's usage trend is evaluated to determine the usage trend factor; Based on the task completion quality factor, the user preference factor, and the usage trend factor, the agent is quantitatively evaluated to obtain the quantitative evaluation result of the agent.

9. The method according to claim 8, characterized in that, The step of evaluating the usage trend of the agent based on its usage records and determining usage trend factors includes: Based on the agent's agent usage records, determine the first number of times the agent is used in the current usage statistical period and the second number of times the agent is used in the historical usage statistical period; Based on the first number of uses and the second number of uses, the usage trend of the agent is evaluated to determine the usage trend factor.

10. The method according to claim 8, characterized in that, The process of quantitatively evaluating the agent based on the task completion quality factor, the user preference factor, and the usage trend factor to obtain the quantitative evaluation result of the agent includes: Based on the release time of the intelligent agent in the intelligent agent market, the timeliness of the intelligent agent is evaluated, and a timeliness factor is determined; Based on the task completion quality factor, the user preference factor, the usage trend factor, and the timeliness factor, the agent is quantitatively evaluated to obtain the quantitative evaluation result of the agent.

11. The method according to any one of claims 1 to 7, characterized in that, The step of quantitatively evaluating the agent based on the task completion quality factor and the user preference factor to obtain the quantitative evaluation result of the agent includes: Based on the release time of the intelligent agent in the intelligent agent market, the timeliness of the intelligent agent is evaluated, and a timeliness factor is determined; Based on the task completion quality factor, the user preference factor, and the timeliness factor, the agent is quantitatively evaluated to obtain the quantitative evaluation result of the agent.

12. A device for evaluating the quality of an intelligent agent, characterized in that, The device includes: The agent identification module is used to identify multiple agents from the agent market; The task quality assessment module is used to assess the quality of the task completed by each agent based on the agent's task execution record, and determine the task completion quality factor. The user preference evaluation module is used to evaluate the user preferences of the intelligent agent based on the user preference operation records of the intelligent agent and determine the user preference factor. The quantitative evaluation module is used to perform quantitative evaluation on the agent based on the task completion quality factor and the user preference factor, and obtain the quantitative evaluation result of the agent. The sorting module is used to sort the multiple agents based on their respective quantitative evaluation results to obtain agent quality evaluation results; the agent quality evaluation results are used to update the order of the multiple agents in the agent market.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.