Assessment method and device based on intelligent agent, medium, equipment and program product
By employing multi-round agent evaluation and conflict resolution algorithms, the problems of time-consuming, inefficient, and inaccurate evaluation of agent-generated content by manual evaluation are solved, thereby improving the efficiency and accuracy of agent evaluation.
Patent Information
- Application Number
- CN202511882536.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-02-17
AI Technical Summary
In existing technologies, content evaluation methods that rely on human evaluation agents are time-consuming and inefficient, and the evaluation criteria are easily affected by human subjectivity, which affects the accuracy of the evaluation results.
The target content is processed in multiple rounds by multiple intelligent agents. The first intelligent agent evaluates the content, the second intelligent agent reviews it to obtain improvement suggestions, and the third intelligent agent determines the evaluation result of the target content based on the final round of processing results of multiple intelligent agents. Conflict resolution algorithms such as voting mechanisms, consensus determination or weighted merging algorithms are used to improve the evaluation accuracy.
It has gradually improved the accuracy of intelligent agents in evaluating target content, provided a more accurate data foundation for evaluating the content generation capabilities of target objects, reduced the influence of human subjectivity, and improved evaluation efficiency.
Smart Images

Figure CN121543625A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of intelligent agents and computer technology, and more specifically, to an evaluation method, apparatus, medium, device, and program product based on intelligent agents. Background Technology
[0002] With the development of artificial intelligence technology, the industry has provided the ability to automatically generate content. For example, intelligent agents are used to automatically generate content for enterprises to meet their needs. The evaluation results of the content can support enterprises in achieving other goals. For example, the evaluation results of the content can reflect the performance of the intelligent agent. Therefore, enterprises can iterate on the intelligent agent based on the evaluation results. Therefore, in order to help enterprises better achieve other goals, it is necessary to improve the accuracy of content and intelligent agent evaluation. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, this disclosure provides an agent-based evaluation method, including: Obtain the target content generated by the target object; The target content is processed in multiple rounds by multiple agents to obtain at least the evaluation results of each agent on the target content for the target evaluation index in each round of processing. Each agent includes a first agent and a second agent. Each round of processing includes: evaluating the target content for the target evaluation index by the first agent to obtain an evaluation result, and reviewing the evaluation result by the second agent to obtain improvement suggestions. The evaluation result and the improvement suggestions obtained in the current round of processing are used at least by the first agent to evaluate the target content in the next round of processing. A third agent determines a first target evaluation result of the target content based on the evaluation results obtained by multiple agents in the last round of processing. The first target evaluation result is used to evaluate the content generation capability of the target object.
[0005] Secondly, this disclosure provides an agent-based evaluation device, comprising: The acquisition module is used to acquire the target content generated by the target object; A processing module is configured to perform multiple rounds of processing on the target content through multiple intelligent agents, so as to obtain at least the evaluation results of each intelligent agent pair on the target content for the target evaluation index in each round of processing. Each intelligent agent pair includes a first intelligent agent and a second intelligent agent. Each round of processing includes: evaluating the target content on the target evaluation index by the first intelligent agent to obtain an evaluation result, and reviewing the evaluation result by the second intelligent agent to obtain improvement suggestions. The evaluation result and the improvement suggestions obtained in the current round of processing are at least used by the first intelligent agent to evaluate the target content in the next round of processing. The determination module is used to determine a first target evaluation result of the target content by a third agent based on the evaluation results obtained by multiple agents in the last round of processing. The first target evaluation result is used to evaluate the content generation capability of the target object.
[0006] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0007] Fourthly, this disclosure provides an electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0008] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, includes the steps of the method described in the first aspect.
[0009] Through the above technical solution, the first agent in each agent pair evaluates the target content, and the second agent in each agent pair reviews the evaluation result of the first agent to obtain improvement suggestions. These improvement suggestions can be used to optimize the first agent's evaluation of the target content in the next round, thereby gradually improving the accuracy of the first agent's evaluation of the target content. In addition, the third agent uses the evaluation results obtained by multiple agent pairs in the last round of processing to determine the first target evaluation result of the target content, thereby further improving the accuracy of the evaluation. This provides a data foundation for accurately evaluating the content generation capability of the target object based on the first target evaluation result.
[0010] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart illustrating an agent-based evaluation method according to an embodiment of the present disclosure; Figure 2 This is a schematic diagram illustrating the process by which each intelligent agent independently completes an evaluation, according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram illustrating the current round processing of a pair of intelligent agents according to an embodiment of the present disclosure; Figure 4 This is a block diagram illustrating an agent-based evaluation device according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0012] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0013] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0016] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0021] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0022] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] Intelligent agents are a core concept in the field of artificial intelligence. They refer to intelligent entities that can autonomously perceive their environment, make decisions, and perform actions to achieve specific goals (such as content generation). They can be software or hardware, and they can rely on models to achieve specific goals.
[0024] In related technologies, relying on human evaluation of content generated by intelligent agents is time-consuming, has low evaluation efficiency, and the evaluation criteria are easily affected by human subjectivity, which affects the accuracy of the evaluation results.
[0025] In view of this, embodiments of the present disclosure provide an evaluation method, apparatus, medium, device, and program product based on intelligent agents.
[0026] The embodiments of this disclosure will be explained and described below with reference to the accompanying drawings.
[0027] Figure 1 This is a flowchart illustrating an agent-based evaluation method according to an embodiment of the present disclosure. This agent-based evaluation method can be applied to electronic devices. Furthermore, the agent-based evaluation method can be executed by an agent-based evaluation device, which can be implemented by software and / or hardware, and the software and / or hardware can be configured in the electronic device. (Refer to...) Figure 1 The agent-based evaluation method may include steps 110, 120 and 130.
[0028] In step 110, the target content generated by the target object is obtained.
[0029] The target content can be automatically retrieved from the target object, or it can be generated by a user-specified target object. This target object can be an intelligent agent, hereinafter referred to as the sixth intelligent agent.
[0030] In step 120, multiple agents perform multiple rounds of processing on the target content to obtain at least the evaluation results of each agent on the target content for the target evaluation index in each round of processing. Each agent pair includes a first agent and a second agent. Each round of processing includes: the first agent evaluating the target content for the target evaluation index to obtain an evaluation result, and the second agent reviewing the evaluation result to obtain improvement suggestions. The evaluation result and improvement suggestions obtained in the current round of processing are used at least by the first agent to evaluate the target content in the next round of processing.
[0031] The target evaluation metrics can be configured based on preset prompt templates. For example, the target evaluation metrics could be the integrity, innovativeness, robustness, and illusion of the sixth agent, etc.
[0032] In each agent pair, the first agent and the second agent exist in pairs. In this way, the target content can be distributed to the first agent in each agent pair in parallel, thereby accelerating the evaluation of the target content.
[0033] The first agent can evaluate the target content based on data related to the target content. This data can be generated by the sixth agent during the content generation process. This data could include user-provided questions, the content generation plan, SQL (Structured Query Language), the SQL execution results, the execution code, and the execution results of the code. The plan refers to the various tasks required to generate the target content, including the generation and execution of SQL and execution code. In each round of processing, the first agent can evaluate the target content based on this data.
[0034] In each agent pair, the first agent relies on at least one of the following: the model, the evaluation parameters, and the prompt word template. Specifically, the model relied on by the first agent can be determined based on the type of the target evaluation metric; for example, a more powerful model can be chosen for a complex type of target evaluation metric. The evaluation parameters can be check items and evaluation criteria used to check whether the check items are qualified. Differences in prompt word templates can be, for example, differences in language style or different methods of guiding the evaluation approach.
[0035] In each round of processing, the first agent in the agent pair evaluates the target content based on the target evaluation indicators to obtain an evaluation result. The second agent in the agent pair then reviews the evaluation result to obtain improvement suggestions. That is, the second agent can simulate a human to identify deviations, omissions, and potential problems in the process of the first agent obtaining the evaluation result, thereby judging the rationality and accuracy of the first agent's evaluation and obtaining improvement suggestions to feed back to the first agent. In this way, the evaluation result and improvement suggestions obtained in the previous round of processing can be used by the first agent to evaluate the target content in the next round of processing. Thus, in multiple rounds of processing, the first agent can continuously optimize its evaluation, thereby improving the accuracy of the evaluation result.
[0036] Furthermore, the evaluation results and improvement suggestions obtained from the current round of processing can also be used to evaluate the target content of the first agent in other agent pairs. For specific implementation methods, please refer to the following related embodiments, which will not be elaborated upon here. Therefore, each agent pair in this disclosure can complete the evaluation independently, i.e., there is no interaction between agent pairs. Alternatively, agent pairs can also complete the evaluation through interaction. For implementation methods of both evaluation methods, please refer to the following related embodiments, which will not be elaborated upon here.
[0037] In step 130, a third agent determines a first target evaluation result of the target content based on the evaluation results obtained by multiple agents in the last round of processing. The first target evaluation result is used to evaluate the content generation capability of the target object.
[0038] Among them, content generation capability corresponds to the target evaluation index. For example, the target evaluation index can be innovativeness. Therefore, content generation capability refers to the innovativeness of the target object.
[0039] Since the evaluation results obtained by each agent in the last round of processing are the evaluation results obtained by the first agent in the first round of processing through multiple rounds of processing, it is possible to accelerate the determination of the first objective evaluation result and reduce the interference of the evaluation results in other rounds of processing on the first objective evaluation result by relying solely on the evaluation results obtained in the last round of processing.
[0040] In this disclosure, the third intelligent agent can use a conflict resolution algorithm to determine the evaluation result of the first objective. The conflict resolution algorithm may include, for example, a voting mechanism, a consensus determination, or a weighted merging algorithm. The selection of the conflict resolution algorithm can be determined based on the evaluation scenario. The implementation method for selecting the conflict resolution algorithm based on the evaluation scenario can be referred to in the following related embodiments.
[0041] The voting mechanism determines the first objective evaluation result based on the principle of "majority rule," and is applicable to evaluation scenarios where each agent has the same weight for its corresponding evaluation result. For example, if agent A is the evaluation result for 1 in the last round of processing, agent A is the evaluation result for 2 in the last round of processing, and agent B is the evaluation result for 3 in the last round of processing, then the first objective evaluation result is A.
[0042] Consistency determination requires all evaluation results to reach a consensus. Only when all evaluation results are consistent can a valid result be output, which is suitable for evaluation scenarios requiring high accuracy. For example, if the evaluation result obtained by agent 1 in the last round of processing is A, the evaluation result obtained by agent 2 in the last round of processing is A, and the evaluation result obtained by agent 3 in the last round of processing is A, then the evaluation result of the first objective is A.
[0043] Weighted merging refers to assigning different weights to each evaluation result based on its "confidence," and finally obtaining the first target evaluation result through weighted calculation. It is suitable for evaluation scenarios where different agents have different levels of importance for their corresponding evaluation results. For example, if an agent has a higher confidence in the model, it will be given a higher weight, so that its evaluation can occupy a major position in the first target evaluation result.
[0044] Through the above technical solution, the first agent in each agent pair evaluates the target content, and the second agent in each agent pair reviews the evaluation result of the first agent to obtain improvement suggestions. These improvement suggestions can be used to optimize the first agent's evaluation of the target content in the next round, thereby gradually improving the accuracy of the first agent's evaluation of the target content. In addition, the third agent uses the evaluation results obtained by multiple agent pairs in the last round of processing to determine the first target evaluation result of the target content, thereby further improving the accuracy of the evaluation. This provides a data foundation for accurately evaluating the content generation capability of the target object based on the first target evaluation result.
[0045] In some embodiments, a seventh agent may uniformly receive the evaluation task for the target content. The seventh agent distributes the evaluation task of the target content to the first agent in different agent pairs in parallel according to the evaluation requirements. It is understood that the evaluation task carries the target content and data related to the target content for the purpose of evaluation.
[0046] The following examples illustrate both independent evaluation by each agent pair and evaluation completed through interaction between agent pairs.
[0047] For the method where each agent completes the evaluation independently, the above-mentioned step of processing the target content in multiple rounds by multiple agents to obtain at least the evaluation results of each agent on the target evaluation index in each round of processing can be implemented in the following way: processing the target content in multiple rounds by multiple agents respectively to obtain at least the evaluation results of each agent on the target evaluation index in each round of processing.
[0048] In each round of processing of the agent pair, the first agent in the agent pair evaluates the target content according to the target evaluation index to obtain the evaluation result, and the evaluation result is reviewed by the second agent in the agent pair to obtain improvement suggestions. The evaluation result and improvement suggestions obtained by the agent pair in the current round of processing are used by the first agent in the agent pair to optimize the evaluation of the target content in the next round of processing.
[0049] In this embodiment, the agent stops processing multiple rounds when a stopping condition is met. The stopping condition could be, for example, reaching a preset round threshold. For instance, if the preset round threshold is set to 3 rounds, then after the 3rd round, the multi-round processing ends, and the evaluation results obtained by each agent in the 3rd round are transmitted to a third agent to determine the first target evaluation result.
[0050] For example, the stopping condition could be that the evaluation results obtained by each agent in the current round of processing are consistent, that is, if the evaluation results output by each agent are the same, the multi-round processing ends, and the evaluation results obtained by each agent in the last round are passed to the third agent to determine the first target evaluation result.
[0051] Figure 2 This is a schematic diagram illustrating the process of each agent pair independently completing an evaluation according to an embodiment of this disclosure. The seventh agent distributes the evaluation task in parallel to agent pair 1, agent pair 2, and agent pair 3, so that agent pair 1, agent pair 2, and agent pair 3 can independently process the target content in multiple rounds. Taking agent pair 1 as an example, the first agent in agent pair 1 evaluates the target content carried in the evaluation task, thereby obtaining an evaluation result. The second agent in agent pair 1 reviews the evaluation result, identifies deviations, omissions, or potential problems in the evaluation result, and generates targeted improvement suggestions. These improvement suggestions are fed back to the first agent in agent pair 1 so that the first agent in agent pair 1 can optimize its own evaluation result in the next round of processing.
[0052] As an example, the prompt word template provided for the first agent could be as follows: To ensure the quality of the evaluation of the "factual accuracy" indicator, two professional evaluators were invited to conduct a comprehensive evaluation of the target content generated by the target object: Evaluator 1 and Evaluator 2. You are Evaluator 1. Please conduct an accurate evaluation of the target content based on the "factual accuracy" indicator.
[0053] Evaluation criteria: {XXX} Evaluation Item: {XXX} Information for the assessment task: {XXX}.
[0054] As an example, the prompt word template provided for the second agent could be as follows: "Please carefully review and verify the evaluation results obtained by the first intelligent agent, and check for any omissions or inappropriateness to ensure that the evaluation results are accurate and consistent with the data facts. Please pay special attention to data conclusions and potential errors, including numerical accuracy issues, numerical calculations, indicator statistics, and discrepancies between data trends and actual values. Generate targeted improvement suggestions and provide negative feedback to the first intelligent agent."
[0055] In this embodiment, the implementation method of determining the first target evaluation result of the target content by a third agent based on the evaluation results obtained by multiple agents in the last round of processing can refer to the above-mentioned related embodiments, and will not be repeated here.
[0056] The above scheme utilizes different agents to independently evaluate the target content, and finally combines the evaluation results of multiple agents in the last round of processing to obtain the first target evaluation result for evaluating the content generation capability of the target object.
[0057] For the evaluation method where each intelligent agent pair completes the evaluation through interaction, in each round of processing, multiple intelligent agent pairs sequentially evaluate and review the target content to obtain the evaluation results and improvement suggestions of each intelligent agent pair for the target content to be evaluated in each round of processing for the target evaluation indicators. The improvement suggestions and evaluation results obtained by the first intelligent agent pair in the current round of processing are used to support the second intelligent agent pair in evaluating the target content in the current round of processing. The first intelligent agent pair performs the evaluation and review processing of the target content before the second intelligent agent pair. In each round of processing, based on the evaluation results and improvement suggestions obtained by all intelligent agent pairs in the current round of processing by the fourth intelligent agent pair, a summary processing is performed to obtain a summary result. The summary result obtained in the current round of processing is used to support each intelligent agent pair in evaluating the target content in the next round of processing.
[0058] In this embodiment, the first agent pair performs the aforementioned evaluation and review process on the agent pair preceding the second agent pair. Since this embodiment employs an iterative mechanism, the evaluation of the second agent pair only relies on the evaluation results and improvement suggestions from the first agent pair's evaluation and review process, thereby reducing the amount of data the agents need to process.
[0059] In some embodiments, the summary results obtained in the current round of processing are used to support each agent pair in evaluating the target content in the next round of processing. There is no need to pass the evaluation results and improvement suggestions of each agent pair in the current round of processing to the next round. Instead, the summary results are used to replace them, which can ensure that the agent pair can rely on the information of the previous round in the next round of processing, and can also reduce the amount of data that the agent pair needs to process in the next round.
[0060] In this embodiment, the multi-round processing is stopped when a stopping condition is met. The stopping condition may be, for example, that the number of rounds reaches a preset round threshold. For example, if the preset round threshold is set to 3 rounds, then after the 3rd round of processing, the multi-round processing ends, and the evaluation results obtained by each agent in the 3rd round are transmitted to the third agent to determine the first target evaluation result.
[0061] For example, the stopping condition could be that all the evaluation results obtained by each agent in the current round of processing are consistent, that is, if the evaluation results of each agent's output are the same, the multi-round processing ends, and the evaluation results obtained by each agent in the last round are passed to the third agent to determine the first target evaluation result.
[0062] As an example, the prompt word template provided for the first agent in a second agent pair could be as follows: "Hello, evaluator 2, the above evaluation task needs to be completed by you and evaluator 1 together. You both need to conduct a comprehensive evaluation of the target content generated by the target object from the perspective of the 'factual accuracy' indicator." Your evaluation criteria: {XXX}; Assessor 1's assessment of "factual accuracy": {XXX}; Evaluation officer 1's suggestions for improving the assessment results on "factual accuracy": {XXX}; Do you have any other differing opinions regarding the "factual accuracy" metric? Please provide your assessment results.
[0063] As an example, the prompt word template provided for the fourth agent could be as follows: "The following are the evaluation results and improvement suggestions from the two evaluators in this round:" Evaluator 1's evaluation result: {XXX}; Evaluator 1's improvement suggestions: {XXX}; Evaluator 2's evaluation result: {XXX}; Evaluator 2's improvement suggestions: {XXX}; As the summarizer, please provide a comprehensive summary of this round of evaluation from the perspective of "factual accuracy." If all evaluators disagree on a certain issue, give the evaluation you believe to be correct, and obtain the summary results to guide the evaluators to reach a consensus in the next round of evaluation and minimize disagreements.
[0064] Figure 3 This is a schematic diagram of the current round processing of multiple agent pairs according to an embodiment of the present disclosure. In the current round processing, the first agent pair first performs evaluation and review processing, that is, the first agent 1 in the first agent pair evaluates the target content and obtains evaluation result 1, and the second agent 1 in the first agent pair reviews the evaluation result 1 and obtains improvement suggestion 1. Next, the second intelligent agent performs the evaluation and review process. That is, the first intelligent agent 2 in the second intelligent agent pair evaluates the target content to obtain evaluation result 2, and the second intelligent agent 2 in the second intelligent agent pair reviews the evaluation result 2 to obtain improvement suggestion 2. In the process of obtaining evaluation result 2, the evaluation result 1 and improvement suggestion 1 obtained by the first intelligent agent pair are combined. Next, the fourth agent performs a summary process based on evaluation result 1, improvement suggestion 1, evaluation result 2, and improvement suggestion 2 to obtain the summary result corresponding to the current round of processing. The summary result obtained in the current round of processing is used to support the evaluation of the first agent pair and the second agent pair in the next round of processing, thereby minimizing the discrepancies between the first agent pair and the second agent pair in the next round.
[0065] Furthermore, after obtaining the summary results corresponding to the current round of processing, the system determines whether to stop the multi-round processing or continue to execute the next round of processing based on the stopping conditions.
[0066] In this embodiment, the implementation method of determining the first target evaluation result of the target content by a third agent based on the evaluation results obtained by multiple agents in the last round of processing can refer to the above-mentioned related embodiments, and will not be repeated here.
[0067] Based on the above scheme, an interactive debate mechanism is proposed. Through the interaction between different agent pairs, the evaluation results and improvement suggestions of one agent pair are fed back to another agent pair, thereby solving the problem of conflicting evaluations of the same target object by multiple agent pairs and quickly achieving the goal of reaching a consensus among different agent pairs.
[0068] In some embodiments, each round of processing further includes: after the first agent in the agent pair obtains the evaluation result, the fifth agent determines whether the evaluation result needs to be reviewed; if the second agent in the agent pair determines that the evaluation result needs to be reviewed, the second agent reviews the evaluation result to obtain improvement suggestions.
[0069] Therefore, by using an intelligent agent to simulate a human, specifically to determine whether the evaluation results obtained by the first intelligent agent need to be reviewed, the objectivity of the overall solution is improved.
[0070] In some embodiments, each agent obtains a corresponding output based on a thought chain, which is used to characterize the steps that the agent needs to perform to obtain the output.
[0071] CoT (Chain of Thought) is a reasoning strategy used to improve the ability to solve complex problems. Its core is to enable intelligent agents to "think step by step" like humans, breaking down problems, deducing step by step, and finally obtaining the corresponding output.
[0072] In this embodiment, the corresponding outputs for each agent used in the evaluation can be obtained based on thought chains. Specifically, a thought chain is incorporated into the prompt word template for each agent. For example, the prompt words provided to the fourth agent could be as follows: "The following are the evaluation results and improvement suggestions from the two evaluators in this round:" Evaluator 1's evaluation result: {XXX}; Evaluator 1's improvement suggestions: {XXX}; Evaluator 2's evaluation result: {XXX}; Evaluator 2's improvement suggestions: {XXX}; As the summarizer, please provide a comprehensive summary of this round of evaluation from the perspective of "factual accuracy." If all evaluators disagree on a particular issue, provide the evaluation you believe to be correct, and obtain the summary result. This will guide the evaluators to reach a consensus in the next round of evaluation and minimize disagreements. Please provide the summary result according to the following steps: 1. Review the assessment results of the two assessors and identify all factual statements; 2. For each statement of fact, determine its accuracy and point out any omissions or misunderstandings in the report; 3. Analyze and summarize the differences between the two evaluators, and provide authoritative factual evidence and suggestions for improvement; 4. Provide a comprehensive summary of the accuracy of the facts. For example, in some embodiments, when the target evaluation index is complex, the target evaluation index can be broken down into multiple sub-indicators, and the agent can be informed to evaluate each sub-indicator separately through a thought chain. Finally, the evaluation result of the target evaluation index is determined based on the sub-evaluation results corresponding to each sub-indicator.
[0073] Therefore, utilizing thought chains can ensure the traceability of the overall logic and improve the accuracy of evaluation results. In some embodiments, the above-described agent-based evaluation method may further include the following steps: transforming the evaluation result of the first target to obtain the evaluation result of the second target, wherein the evaluation result of the second target is structured data; and outputting the evaluation result of the second target.
[0074] The evaluation results for the second objective can be JSON data or other structured data. Taking JSON data as an example, the data structure for the evaluation results of the second objective is defined, including fields such as indicator name, agent's thought process, evaluation score, and evaluation error list. It is understandable that, depending on the indicator, the fields can be configured to suit the evaluation of that indicator system.
[0075] By using the above method to output the evaluation results of the second objective in a structured form, the readability of the evaluation results for users can be improved.
[0076] Figure 4 This is a block diagram of an agent-based evaluation device according to an embodiment of the present disclosure, with reference to... Figure 4 The agent-based evaluation device 400 includes: Module 401 is used to obtain the target content generated by the target object; Processing module 402 is configured to perform multiple rounds of processing on the target content through multiple intelligent agents, so as to obtain at least the evaluation results of each intelligent agent pair on the target content for the target evaluation index in each round of processing. Each intelligent agent pair includes a first intelligent agent and a second intelligent agent. Each round of processing includes: evaluating the target content on the target evaluation index by the first intelligent agent to obtain an evaluation result, and reviewing the evaluation result by the second intelligent agent to obtain improvement suggestions. The evaluation result and the improvement suggestions obtained in the current round of processing are at least used by the first intelligent agent to evaluate the target content in the next round of processing. The determining module 403 is used to determine a first target evaluation result of the target content by a third agent based on the evaluation results obtained by multiple agents in the last round of processing. The first target evaluation result is used to evaluate the content generation capability of the target object.
[0077] Optionally, in each round of processing, the target content is evaluated and reviewed sequentially by the multiple agent pairs to obtain the evaluation results and improvement suggestions of each agent pair for the target content in each round of processing for the target evaluation indicators. The improvement suggestions and evaluation results obtained by the first agent pair in the current round of processing are used to support the second agent pair in evaluating the target content in the current round of processing. The first agent pair performs the evaluation and review processing on the target content before the second agent pair. Furthermore, in each round of processing, the fourth agent summarizes the evaluation results and improvement suggestions obtained by all agent pairs in the current round of processing to obtain a summary result. The summary result obtained in the current round of processing is used to support each agent pair in evaluating the target content in the next round of processing.
[0078] Optionally, the processing module 402 is further configured to: perform multiple rounds of processing on the target content through multiple intelligent agents, so as to obtain at least the evaluation results of each intelligent agent on the target content for the target evaluation index in each round of processing.
[0079] Optionally, the agent-based evaluation device 400 further includes: The judgment module is used to determine, after the first agent in the agent pair obtains the evaluation result, whether the evaluation result needs to be reviewed by the fifth agent. If the second agent in the agent pair determines that the evaluation result needs to be reviewed, it is used to review the evaluation result to obtain improvement suggestions.
[0080] Optionally, each agent obtains a corresponding output based on a thought chain, wherein the thought chain is used to characterize the steps required to obtain the output.
[0081] Optionally, the agent-based evaluation device 400 further includes: A stop module is configured to stop the multi-round processing when a stop condition is met, wherein the stop condition includes at least one of the following: The number of rounds has reached the preset round threshold; Each of the aforementioned agents maintains consistency with the evaluation results obtained in the current round of processing.
[0082] Optionally, the agent-based evaluation device 400 further includes: The conversion module is used to convert the first target evaluation result to obtain the second target evaluation result, wherein the second target evaluation result is structured data; The output module is used to output the evaluation result of the second target.
[0083] The implementation methods of each module in the above-mentioned agent-based evaluation device 400 can refer to the above-mentioned related embodiments, and will not be repeated here.
[0084] This disclosure also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described agent-based evaluation method.
[0085] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described agent-based evaluation method.
[0086] This disclosure also provides an electronic device, including: A storage device on which computer programs are stored; A processing device is used to execute the computer program in the storage device to implement the steps of the agent-based evaluation method described above.
[0087] The following is for reference. Figure 5 This diagram illustrates a structural schematic of an electronic device 500 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0088] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. RAM 503 also stores various programs and data required for the operation of electronic device 500. Processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0089] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0090] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0091] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0092] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0093] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0094] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire target content generated by a target object; perform multiple rounds of processing on the target content through multiple intelligent agents to obtain at least the evaluation results of each intelligent agent pair for the target content in each round of processing for a target evaluation index, wherein each intelligent agent pair includes a first intelligent agent and a second intelligent agent, and each round of processing includes: evaluating the target content for the target evaluation index by the first intelligent agent to obtain an evaluation result, and reviewing the evaluation result by the second intelligent agent to obtain improvement suggestions, wherein the evaluation result and the improvement suggestions obtained in the current round of processing are at least used by the first intelligent agent to evaluate the target content in the next round of processing; and determining a first target evaluation result of the target content by a third intelligent agent based on the evaluation result obtained by the multiple intelligent agent pairs in the last round of processing, wherein the first target evaluation result is used to evaluate the content generation capability of the target object.
[0095] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0097] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a module does not necessarily limit the module itself; for example, an acquisition module can also be described as "a module for acquiring target content generated by a target object".
[0098] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0099] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0100] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0101] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0102] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. An agent-based evaluation method, characterized in that, The method comprises: obtaining target content generated by a target object; performing multi-round processing on the target content by a plurality of agent pairs to obtain at least evaluation results of each agent pair for the target content in each round of processing with respect to a target evaluation index, each agent pair comprising a first agent and a second agent, each round of processing comprising: performing evaluation on the target content by the first agent with respect to the target evaluation index to obtain an evaluation result, and performing review on the evaluation result by the second agent to obtain improvement suggestions, and the evaluation result and the improvement suggestions obtained in the current round of processing are used at least for the first agent to evaluate the target content in the next round of processing; determining a first target evaluation result of the target content based on the evaluation results obtained by the plurality of agent pairs in the last round of processing by a third agent, the first target evaluation result being used to evaluate content generation capability of the target object.
2. The method of claim 1, wherein, In each round of processing, the target content is sequentially evaluated and reviewed by the plurality of agent pairs to obtain the evaluation result and the improvement suggestions of each agent pair for the target content in each round of processing with respect to the target evaluation index, the evaluation result and the improvement suggestions obtained by the first agent in the current round of processing are used to support the second agent to evaluate the target content in the current round of processing, and the first agent performs the evaluation and review on the target content before the second agent; and in each round of processing, based on a fourth agent, the evaluation results and the improvement suggestions obtained by all the agent pairs in the current round of processing are summarized to obtain a summary result, and the summary result obtained in the current round of processing is used to support each agent pair to evaluate the target content in the next round of processing.
3. The method of claim 1, wherein, The method further comprises: performing multi-round processing on the target content by a plurality of agent pairs to obtain at least evaluation results of each agent pair for the target content in each round of processing with respect to a target evaluation index.
4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: after the first agent in the agent pair obtains the evaluation result, determining whether the evaluation result needs to be reviewed by a fifth agent, and the second agent in the agent pair is used to review the evaluation result to obtain improvement suggestions in a case where it is determined that the evaluation result needs to be reviewed.
5. The method according to any one of claims 1 to 3, characterized in that, Each agent obtains a corresponding output based on a thought chain, and the thought chain is used to represent steps that need to be performed to obtain the output.
6. The method of any one of claims 1-3, wherein, The method further comprises: stopping the multi-round processing in a case where a stop condition is met, the stop condition comprising at least one of the following: the number of rounds reaches a preset round threshold; the evaluation result obtained by each agent pair in the current round of processing meets consistency.
7. The method according to claims 1-3, characterized by, The method further comprises: The first target evaluation result is converted to obtain a second target evaluation result, and the second target evaluation result is structured data. The second target evaluation result is output.
8. An agent-based evaluation apparatus, characterized by comprising: The method comprises the following steps: An acquisition module is configured to acquire target content generated by a target object; A processing module is configured to perform multi-round processing on the target content by a plurality of agents to obtain at least an evaluation result of each agent pair for a target evaluation index in each round of processing, each agent pair comprising a first agent and a second agent, and each round of processing comprising: performing evaluation on the target content by the first agent for the target evaluation index to obtain an evaluation result, and performing review on the evaluation result by the second agent to obtain improvement suggestions, and the evaluation result and the improvement suggestions obtained in the current round of processing are used at least for the first agent to evaluate the target content in the next round of processing; A determination module is configured to determine, by a third agent, a first target evaluation result of the target content based on the evaluation results obtained by the plurality of agent pairs in the last round of processing, and the first target evaluation result is used to evaluate the content generation capability of the target object.
9. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by a processing device to implement the steps of the method of any one of claims 1-7.
10. An electronic device, comprising: The method comprises the following steps: A storage device having a computer program stored thereon; A processing device configured to execute the computer program in the storage device to implement the steps of the method of any one of claims 1-7.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to implement the steps of the method of any one of claims 1-7.