Online agent ranking scoring method, system, equipment and medium

By combining objective and subjective scoring methods and introducing an intelligent agent scoring system with referee level and agent level scores, the problem of incomplete scoring in existing technologies is solved, a more comprehensive evaluation and fair ranking mechanism is achieved, and the rapid improvement of intelligent agent technology is promoted.

CN120655362APending Publication Date: 2025-09-16TIANFU JIANGXI LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510801573.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing intelligent agent scoring system lacks a comprehensive integrated scoring mechanism, making it difficult to reflect the differences in individual abilities. The scoring results lack fairness, and the score differences in most scoring systems are small, making it difficult to force individual technical improvements.

Method used

A scoring method combining objective and subjective evaluation is adopted. The comprehensive evaluation score is calculated through the initial comprehensive score, objective evaluation, subjective evaluation comparison win rate and other influencing factors. The ranking is updated based on the comprehensive evaluation score, and the referee level and intelligent agent level score are introduced for weighted scoring.

Benefits of technology

It has achieved a more comprehensive evaluation dimension, making the ranking more convincing, motivating low-level newcomers to improve quickly, increasing the competitive pressure at the high level, improving the fairness of scoring, accelerating the survival of the fittest, and increasing the participation of intelligent agents and the enthusiasm for technical improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655362A_ABST
    Figure CN120655362A_ABST
Patent Text Reader

Abstract

The invention discloses an online agent ranking scoring method, system and device and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the assignment of an initial comprehensive score for a new to-be-evaluated agent; performing objective evaluation on the to-be-evaluated intelligent agent, and updating an initial objective score in the initial comprehensive scores into an objective evaluation score obtained through objective evaluation; based on the initial subjective score, evaluating and comparing the to-be-evaluated agent with each agent in a preset agent library, and obtaining a comparison winning rate of the to-be-evaluated agent relative to each agent in the preset agent library; performing calculation according to other influence factors of subjective evaluation to obtain a subjective evaluation score of the to-be-evaluated agent; calculating to obtain a comprehensive evaluation score of the to-be-evaluated agent, and updating the initial ranking based on the comprehensive evaluation score; the problems that evaluation is not comprehensive enough and ranking is difficult to reflect ability differences among individuals can be solved, and meanwhile the problems that public scoring is not fair and stable enough are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to an online intelligent agent ranking and scoring method, system, device and medium. Background Art

[0002] Artificial intelligence model technology has advanced rapidly in recent years, from deep learning to neural networks to large language models. The release of Manus in early 2025 marked the entry of artificial intelligence into the era of intelligent agents. An AI agent is an intelligent system that can autonomously perceive its environment, analyze information, make decisions, and execute actions. Its core goal is to complete a specific task or achieve a long-term goal. It can be a software program (such as a chatbot), a physical entity (such as a robot), or a hybrid system (such as a self-driving car). It uses a model (or multiple models) to perceive the environment, make decisions, and execute actions to achieve a specific goal.

[0003] From the perspective of agent scoring systems, they are primarily divided into objective and subjective evaluations. Objective evaluation measures model functionality based on objective datasets and clear evaluation metrics. Quantitative metrics are selected and calculated based on the specific task and model type. Subjective evaluation relies on subjective judgment to assess model performance and effectiveness. Judges for subjective evaluations can be professional evaluators, agents, large models, or ordinary users.

[0004] Currently, the agent scoring system has the following main flaws: 1. Lack of a comprehensive agent scoring system. Most agent evaluation benchmarks focus on specific functions, lacking comprehensive and integrated ratings and rankings. This makes it difficult for users to assess the overall performance of an agent from a single functional perspective. For example, the Berkeley Function Calling Leaderboard (BFCL) ranks agents based on their function calling capabilities, while GAIA focuses on ranking agents based on reasoning, multimodal understanding, and web navigation. Therefore, it's difficult to judge an agent's overall quality based on a single rating benchmark.

[0005] 2. Most scoring systems use a maximum score of 100 to score the evaluation objects. The difference in scores between rankings is often relatively small, making it difficult to reflect the ability differences between individuals, and thus it is difficult to force individual technical improvements.

[0006] 3. Online subjective evaluations are mainly participated by the public, and the judges’ abilities vary, resulting in a lack of fairness in the scoring results.

[0007] In view of this, this application is hereby made. Summary of the Invention

[0008] The purpose of the present invention is to provide an online agent ranking and scoring method, system, device and medium to solve the problems existing in the above-mentioned background technology.

[0009] The above technical objectives of the present invention are achieved through the following technical solutions: In a first aspect, the present application provides an online agent ranking and scoring method, comprising the following specific steps: Assign an initial comprehensive score to the new agent to be evaluated. The initial comprehensive score includes the initial objective score and the initial subjective score. Conduct an objective evaluation on the agent to be evaluated, and update the initial objective score in the initial comprehensive score to the objective evaluation score obtained through the objective evaluation; Based on the initial subjective score, the agent to be evaluated is compared with a random single agent in the pre-set agent library, and the winning rate of the agent to be evaluated relative to the target agent in the pre-set agent library is obtained; Based on the comparison of win rates and other influencing factors of subjective evaluation, the subjective evaluation score of the agent to be evaluated is calculated; The subjective evaluation score and the objective evaluation score are used to calculate the comprehensive evaluation score of the agent to be evaluated, and the initial ranking is updated based on the comprehensive evaluation score.

[0010] On the basis of the above technical solution, the present invention can also be improved as follows.

[0011] Furthermore, the above comprehensive evaluation scores are as follows: ; in: , ;and ; Where, Indicates the comprehensive evaluation score. Indicates the objective evaluation score, Indicates the subjective evaluation score, represents the weight of the objective evaluation score, Indicates the weight of the subjective evaluation score.

[0012] Furthermore, other influencing factors of the above subjective evaluation include: referee level weight, agent subjective evaluation comparison number factor and agent level score.

[0013] Furthermore, the above agent level scores are specifically:

[0014] Where, represents the agent level score, represents the ranking percentage of the agent to be evaluated obtained through the initial ranking, Represents a constant.

[0015] Furthermore, the above agent subjective evaluation comparison factor is specifically:

[0016] Where, is the agent’s subjective evaluation comparison factor, Indicates the number of subjective evaluation comparisons of the agent.

[0017] Furthermore, the winning rate of the above-mentioned agent to be evaluated compared with the target agent in the preset agent library is as follows:

[0018] Where, Indicates that the agent A to be evaluated is relative to the benchmark agent in the preset agent library. The comparative win rate, represents the subjective evaluation score of the agent to be evaluated before the evaluation, Represents the subjective evaluation score of agent B before evaluation.

[0019] Furthermore, after a round of competition between the agent A and the benchmark agent B, the subjective evaluation score of the agent A is as follows:

[0020] Where, is the updated subjective evaluation score of the agent to be evaluated, is the subjective score of the agent to be evaluated before evaluation, is the agent level score, represents the number of subjective evaluation comparisons of the agent, Indicates the referee level weight, Indicates the comparison result of this evaluation. If the agent A wins, is 1; if the result is a tie, is 0.5; if A loses to B, is 0; Indicates that the agent A to be evaluated is relative to the agents in the preset agent library. Comparative win rate.

[0021] In a second aspect, the present application provides an online agent ranking and scoring system, which is applied to an online agent ranking and scoring method according to any one of the first aspects, including: The initial setting module is used to assign an initial comprehensive score to the new agent to be evaluated. The initial comprehensive score includes the initial objective score and the initial subjective score; The objective evaluation module is used to objectively evaluate the agent to be evaluated and update the initial objective score in the initial comprehensive score to the objective evaluation score obtained through the objective evaluation; The evaluation and comparison module is used to evaluate and compare the agent to be evaluated with a random single agent in the pre-set agent library based on the initial subjective score, and obtain the comparison win rate of the agent to be evaluated relative to the target agent in the pre-set agent library in this round; The subjective evaluation module is used to calculate the subjective evaluation score of the agent to be evaluated based on the comparison of win rates and other influencing factors of subjective evaluation; The comprehensive scoring module is used to calculate the comprehensive evaluation score of the agent to be evaluated using the subjective evaluation score and the objective evaluation score, and to update the initial ranking based on the comprehensive evaluation score.

[0022] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the methods in the first aspect when executing the computer program.

[0023] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable a computer to execute any one of the methods in the first aspect.

[0024] Compared with the prior art, the present invention has at least the following beneficial effects: 1. Combining objective and subjective evaluations makes the evaluation dimensions more comprehensive, and the rankings based on this are more convincing. At the same time, the rankings can be updated based on the results of individual evaluations, eliminating the need for repeated evaluations.

[0025] 2. By introducing a subjective + objective scoring system, the scoring dimensions are fully covered. Specifically, by introducing comparative scoring, the scores can be flexibly added or subtracted without an upper limit, which encourages low-level newcomers to improve quickly and increases the competitive pressure at high levels, thereby making the ranking more fluid, which can both attract participation and more clearly reflect the ranking gap. In the subjective scoring link, external influencing factors such as grade, referee, and ranking are added for weighting, so that the score is closer to the actual level of the participants.

[0026] 3. The subjective scoring mechanism fully considers the impact of the judges' level on the fairness of the scoring. Since it is a public Dianping system, the judges may be professionals, large models, or ordinary people. To address the uneven level of judges, the judges' professionalism is divided, so that judges with higher levels have more say and can influence the ranking of intelligent entities. At the same time, the judges are managed by level, which also eliminates online water army and ranking manipulation.

[0027] 4. The subjective scoring mechanism accelerates the survival of the fittest, allowing new and excellent intelligent agents to improve their rankings more quickly and allowing inferior intelligent agents to be eliminated more quickly, thereby greatly increasing the enthusiasm of intelligent agents to participate in the evaluation and driving the improvement of intelligent agent technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings: Figure 1 A flowchart of a scoring method according to an embodiment of the present invention; Figure 2 This is a scoring flow chart for the comprehensive evaluation score in an embodiment of the present invention; Figure 3 This is a connection diagram of a scoring system according to an embodiment of the present invention; Figure 4 Schematic diagram of the connection of electronic equipment in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0030] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0031] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0032] In the description of the embodiments of the present invention, "a plurality of" means at least two.

[0033] Example 1: In order to overcome the problem that the evaluation is not comprehensive enough and the ranking is difficult to reflect the difference in ability between individuals, and to solve the problem that the public rating is not fair and stable, this example provides an online agent ranking and scoring method, such as Figure 1 As shown, the following specific steps are included: S1, assigns an initial comprehensive score to the new agent to be evaluated. The initial comprehensive score includes the initial objective score and the initial subjective score.

[0034] S2, objectively evaluate the agent to be evaluated, and update the initial objective score in the initial comprehensive score to the objective evaluation score obtained through objective evaluation.

[0035] Optional: The above comprehensive evaluation scores are as follows: ; in: , ;and ; Where, Indicates the comprehensive evaluation score. Indicates the objective evaluation score, Indicates the subjective evaluation score, represents the weight of the objective evaluation score, Indicates the weight of the subjective evaluation score.

[0036] Objective evaluation refers to the process of quantitatively assessing the performance, effectiveness, and reliability of AI models through a systematic, data-driven, and unbiased approach. Its core goal is to objectively reflect the true capabilities of intelligent agents based on clear indicators and reproducible processes, providing a scientific basis for their optimization, selection, and application. This can be achieved through the following steps: 1. Use standardized evaluation indicators (such as precision, recall, and F1 value) for numerical measurement.

[0037] 2. Classify the agent's capabilities and assign corresponding indicators to each capability. For example, a customer service agent based on a large language model needs to be evaluated on multiple aspects, including its multi-round conversational capabilities, language comprehension, domain expertise, robustness, performance, and scenario application. Each capability is assigned a corresponding indicator and weight.

[0038] 3. Calculate the objective assessment score based on the scores of individual abilities and the weights.

[0039]

[0040] Where, is the total number of competency items assessed, is the weight of each ability; is the raw score for this ability. Score for objective evaluation.

[0041] Specifically, the subjective evaluation is conducted by the referee, who judges the answer results of the evaluation agent and the comparison agent, selects the better one, and adds or subtracts points from the winner and loser based on the comparison results to obtain a new subjective evaluation score and ranking, as shown below.

[0042] S3, based on the initial subjective score, evaluates and compares the agent to be evaluated with a random single agent in the preset agent library, and obtains the comparative win rate of the agent to be evaluated relative to the target agent in the preset agent library for this round.

[0043] The winning rate of the above-mentioned agent to be evaluated compared with the benchmark agent in the preset agent library is as follows:

[0044] Where, Indicates that the agent A to be evaluated is relative to the benchmark agent in the preset agent library. The comparative win rate, represents the subjective evaluation score of the agent to be evaluated before the evaluation, Represents the subjective evaluation score of agent B before evaluation.

[0045] S4, based on the comparison win rate and other influencing factors of the subjective evaluation, calculates the subjective evaluation score of the agent to be evaluated.

[0046] Among them, other influencing factors of the above-mentioned subjective evaluation include: referee level weight, agent subjective evaluation comparison times factor and agent level score.

[0047] Optionally, the agent level score is:

[0048] Where, represents the agent level score, represents the ranking percentage of the agent to be evaluated obtained through the initial ranking, Represents a constant.

[0049] Optionally, the agent subjective evaluation comparison factor is:

[0050] Where, is the agent’s subjective evaluation comparison factor, Indicates the number of subjective evaluation comparisons of the agent.

[0051] S5, using the subjective evaluation score and the objective evaluation score, calculates the comprehensive evaluation score of the agent to be evaluated, and updates the initial ranking based on the comprehensive evaluation score.

[0052] Optionally, after the aforementioned agent A competes against the benchmark agent B for a round, the subjective evaluation score of the agent A is as follows:

[0053] Where, is the updated subjective evaluation score of the agent to be evaluated, is the subjective score of the agent to be evaluated before evaluation, is the agent level score, represents the number of subjective evaluation comparisons of the agent, Indicates the referee level weight, Indicates the comparison result of this evaluation. If the agent A wins, is 1; if the result is a tie, is 0.5; if A loses to B, is 0; Indicates that the agent A to be evaluated is relative to the agents in the preset agent library. Comparative win rate.

[0054] Example 2: This example provides an online agent ranking and scoring method, such as Figure 2 As shown, first determine whether it is a new agent. If it is not a new agent, directly obtain the current comprehensive score and ranking of the agent in the agent library. When it is determined to be a new agent, the specific steps include the following: 1. Comprehensive evaluation score 1. The composition of the comprehensive score, which includes objective and subjective scores: (Formula 1) The weights of objective and subjective scores are calculated as follows: (Formula 2) (Formula 3) (Formula 4) In the above formula, represents the objective score, and the corresponding weight is , Represents the subjective evaluation score, and the corresponding weight is , It is the comprehensive evaluation score.

[0055] 2. Initial score: A single agent must pass a complete round of objective and subjective evaluations and obtain the first comprehensive score before it can be included in the ranking; assuming the initial comprehensive score is , the initial ranking is .

[0056] 3. Update of comprehensive scores and rankings After obtaining the initial ranking, a single agent is evaluated again, and the corresponding domain score and the corresponding weight in the comprehensive score are updated; for example, for a certain agent After an objective evaluation, the objective evaluation score is obtained , the updated comprehensive score is: .....(Formula 5) .....(Formula 6) 1 ..... (Formula 7) According to the updated comprehensive score of the agent A to be evaluated , re-rank and get a new ranking.

[0057] 2. Subjective Evaluation Scores Among them, the subjective evaluation score is determined by the referee's evaluation of the answer results of the evaluated intelligent agent and the comparison intelligent agent, and the better party is selected. The winner and loser are added or subtracted points according to the comparison results to obtain a new subjective evaluation score and ranking.

[0058] 1. The following factors need to be considered when calculating subjective evaluation scores 1) The influence of referee professionalism and proficiency. A higher referee level is considered fairer in scoring, and the referee's influence on the outcome is less.

[0059] 2) The impact of the agent's adequacy in the evaluation. The more agents an agent has played against, the more thorough its comparison is, and the closer the comparison results are to the actual situation.

[0060] 3) The impact of agent level. The higher the agent's level, the lower the score added or subtracted after comparison with the same agent; the lower the agent's level, the higher the score added or subtracted after comparison with the same agent. When a lower-level agent defeats a higher-level agent, the higher the score; when a higher-level agent loses to a lower-level agent, the more points are deducted. In other words, this intensifies competition at the top of the rankings and accelerates the elimination of the weaker ones at the bottom.

[0061] 2. Basic principles of score calculation 1) Since the fairness of subjective evaluation is often affected by external factors, we assume that the actual ability level of an agent at a certain point in time is normally distributed: ; As the basic score, Indicates the deviation between actual ability and basic score. The smaller it is, the closer the actual ability is to the score.

[0062] 2) Predicting Win Rate For the two agents A and B participating in the comparison, the winning rate of A relative to B is: .....(Formula 8) in, is the subjective evaluation score of the current agent A to be evaluated, is the subjective evaluation score of agent B.

[0063] 3) Updated A's score The score of the last round Add (subtract) the score of this round.

[0064] = .....(Formula 9) In the above formula, is the result of this game. If A wins, then A win is 1, a loss is 0, and a draw is 0.5; Used to control the speed at which the score changes in each round, which is affected by many factors.

[0065] 3. Calculation 3.1 Agent ranking factor, The scoring basis is the ranking grade point K.

[0066] .....(Formula 10) Where M is a constant, and the K value is divided into two levels: M and 2 / M; is the ranking percentage of the agent. For the top 50% of agents, K is only half of that for the bottom 50%.

[0067] 1) After playing against each other at the same level, the scores of top-ranked agents change relatively little, intensifying competition for the top players. The scores of lower-ranked agents change relatively much, intensifying the survival of the fittest.

[0068] 2) After challenging at different levels, if the one with lower level wins, it can accelerate the upgrade of the one with lower level.

[0069] 3.2 Number of battles factor To ensure that the more times an agent participates in battles, the smaller the deviation between its ability and score. In other words, the more battles an agent has, the harder it is to gain points from its opponent after winning. This is achieved by dividing the rating K by a factor related to the number of battles.

[0070] .....(Formula 11) Where N is the number of battles, It is a positive number used to control the impact of the number of battles on the K value.

[0071] 3.3 Referee Level Factor Referees are divided into three levels: novice, intermediate and senior. The scoring weight of senior referees is set to 1, and the score weight of intermediate and novice referees decreases accordingly.

[0072] 4. Final adjustment of the subjective evaluation score calculation formula = .....(Formula 11) In the above formula, is the subjective evaluation score of the current agent A to be evaluated, K is the agent level score, N is the number of times the agent participates in the battle, Used to adjust the impact of the number of battles on the score. R is the referee level weight, ranging from (0,1]. The higher the level, the closer the value is to 1.

[0073] The scoring system process is shown in Figure 2 As shown, the details are as follows: 1. Select the agent to be evaluated. The backend system obtains information such as the agent type and description to match the evaluation tool and dataset.

[0074] 2. Set the evaluation type. If it is a new agent, the evaluation type is "ALL"; if it is not a new agent, the evaluation type is specified by the user.

[0075] 3. Objective evaluation: Automated tools obtain the test environment, test tools, and data sets, automatically perform evaluation tasks, and directly score based on objective indicator scoring standards.

[0076] 4. Subjective evaluation: 1) If it is a new agent, set the initial subjective evaluation score and initial grade for it.

[0077] 2) Anonymously and randomly assign competing agents and obtain their level and score information.

[0078] 3) Both competing agents complete the evaluation task; 4) The judges evaluate the agents’ task completion results and select the winning agent of this round.

[0079] 5) Calculate the agent's new subjective evaluation score based on the results of this round of evaluation (see Formula 7-11).

[0080] 6) Based on the new subjective evaluation points, the subjective evaluation ranking and subjective evaluation level will be refreshed.

[0081] 5. Update the agent's comprehensive score (see Formula 4-6).

[0082] 6. Update the overall ranking of intelligent entities.

[0083] 7. End.

[0084] In this embodiment, the proposed evaluation method has the following advantages over existing model or agent scoring systems: 1. Combining objective and subjective evaluations makes the evaluation dimensions more comprehensive, and the ranking based on this is more convincing. At the same time, the ranking can be updated based on the results of individual evaluations, eliminating the need for repeated evaluations.

[0085] 2. The subjective scoring mechanism fully considers the impact of judges' expertise on scoring fairness. Since the Dianping system is open, judges can range from professionals to large-scale models to the general public. To address the varying levels of judges' expertise, a tiered approach is implemented, ensuring that higher-ranked judges have greater influence and influence on agent rankings. Furthermore, tiered management of judges prevents online scams and ranking manipulation.

[0086] 3. The subjective scoring mechanism accelerates the survival of the fittest, allowing new, outstanding agents to rise in the rankings more quickly and eliminating inferior agents more quickly. This significantly increases agents' enthusiasm for participating in the evaluation, driving improvements in agent technology.

[0087] 4. The subjective scoring mechanism takes into account the impact of battle sufficiency on scores and rankings, so that the more fully engaged an agent is, the more stable its score tends to be, which better reflects the agent's true level and ranking.

[0088] Example 3: This embodiment of the present application provides an online agent ranking and scoring system, which is applied to an online agent ranking and scoring method of Example 1, such as Figure 3 Shown, including: The initial setting module is used to assign an initial comprehensive score to the new agent to be evaluated. The initial comprehensive score includes the initial objective score and the initial subjective score; The objective evaluation module is used to objectively evaluate the agent to be evaluated and update the initial objective score in the initial comprehensive score to the objective evaluation score obtained through the objective evaluation; The evaluation and comparison module is used to evaluate and compare the agent to be evaluated with a random single agent in the pre-set agent library based on the initial subjective score, and obtain the comparison win rate of the agent to be evaluated relative to the target agent in the pre-set agent library in this round; A subjective evaluation module, configured to calculate a subjective evaluation score of the agent to be evaluated based on the comparison win rate and other influencing factors of the subjective evaluation, and to count the number of subjective evaluations; A comprehensive scoring module, configured to calculate a comprehensive evaluation score of the agent to be evaluated using the subjective evaluation score and the objective evaluation score, and update the initial ranking based on the comprehensive evaluation score; The ranking module is used to rank the agents from high to low according to their comprehensive evaluation scores, and to divide the levels of the agents according to the comprehensive rankings.

[0089] Example 4: This embodiment of the present application provides an electronic device, such as Figure 4 As shown, it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method of embodiment 1 is implemented.

[0090] Example 5: The embodiment of the present application provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable the computer to execute the method of Example 1.

[0091] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0092] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0093] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0095] Those skilled in the art will understand that all or part of the steps in implementing the above facts and methods can be completed by instructing relevant hardware through a program, and the program involved or the program can be stored in a computer-readable storage medium. When the program is executed, it includes the following steps: the corresponding method steps are then brought out, and the storage medium can be ROM / RAM, a disk, an optical disk, etc.

[0096] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An online agent ranking and scoring method, characterized in that: The specific steps include: Assign an initial comprehensive score to the new agent to be evaluated. The initial comprehensive score includes the initial objective score and the initial subjective score. Performing an objective evaluation on the agent to be evaluated, and updating the initial objective score in the initial comprehensive score to the objective evaluation score obtained through the objective evaluation; Based on the initial subjective score, the agent to be evaluated is evaluated and compared with a random single agent in the preset agent library, and the winning rate of the agent to be evaluated relative to the target agent in the preset agent library in this round is obtained; Based on the comparison win rate and other influencing factors of the subjective evaluation, a subjective evaluation score of the subject to be evaluated is calculated; The subjective evaluation score and the objective evaluation score are used to calculate the comprehensive evaluation score of the agent to be evaluated, and the initial ranking is updated based on the comprehensive evaluation score.

2. An online agent ranking and scoring method according to claim 1, characterized in that: The comprehensive evaluation scores are specifically: ; in: , ;and ; Where, Indicates the comprehensive evaluation score. Indicates the objective evaluation score, Indicates the subjective evaluation score, represents the weight of the objective evaluation score, Indicates the weight of the subjective evaluation score.

3. The online agent ranking and scoring method according to claim 1, characterized in that: Other influencing factors of the subjective evaluation include: referee level weight, agent subjective evaluation comparison times factor and agent level score.

4. The online agent ranking and scoring method according to claim 3, wherein: The agent level score is specifically: Where, represents the agent level score, Indicates the ranking percentage of the agent to be evaluated obtained through ranking, Represents a constant.

5. The online agent ranking and scoring method according to claim 3, wherein: The agent subjective evaluation comparison factor is specifically: Where, is the agent’s subjective evaluation comparison factor, Indicates the number of subjective evaluation comparisons of the agent.

6. The online agent ranking and scoring method according to claim 1, characterized in that: The winning rate of the agent to be evaluated compared with the target agent in the preset agent library is specifically: Where, Indicates that the agent A to be evaluated is relative to the benchmark agent in the preset agent library. The comparative win rate, represents the subjective evaluation score of the agent to be evaluated before the evaluation, Represents the subjective evaluation score of agent B before evaluation.

7. The online agent ranking and scoring method according to claim 1, characterized in that: After a round of competition between agent A and benchmark agent B, the subjective evaluation score of agent A is as follows: Where, is the updated subjective evaluation score of the agent to be evaluated, is the subjective score of the agent to be evaluated before evaluation, is the agent level score, represents the number of subjective evaluation comparisons of the agent, Indicates the referee level weight, Indicates the comparison result of this evaluation. If the agent A wins, is 1; if the result is a tie, is 0.5; if A loses to B, is 0; Indicates that the agent A to be evaluated is relative to the agents in the preset agent library. Comparative win rate.

8. An online agent ranking and scoring system, characterized in that: include: The initial setting module is used to assign an initial comprehensive score to the new agent to be evaluated. The initial comprehensive score includes the initial objective score and the initial subjective score; An objective evaluation module, configured to perform an objective evaluation on the agent to be evaluated and update the initial objective score in the initial comprehensive score to the objective evaluation score obtained through the objective evaluation; An evaluation and comparison module is configured to evaluate and compare the agent to be evaluated with a random single agent in a preset agent library based on the initial subjective score, and obtain a comparison win rate of the agent to be evaluated relative to the target agent in the preset agent library for that round; A subjective evaluation module, configured to calculate a subjective evaluation score of the agent to be evaluated based on the comparison win rate and other influencing factors of the subjective evaluation; The comprehensive scoring module is used to calculate the comprehensive evaluation score of the agent to be evaluated using the subjective evaluation score and the objective evaluation score, and update the initial ranking based on the comprehensive evaluation score.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method according to any one of claims 1 to 7 is implemented when the processor executes the computer program.

10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, which enable a computer to execute the method of any one of claims 1 to 7.