Evaluation methods, devices, media, and electronic equipment for intelligent agents and large language models
By creating multi-dimensional evaluation tasks and displaying multiple evaluators, the problem of the single evaluation method for intelligent agents and large language models in existing technologies is solved, realizing the intuitiveness and flexibility of comparative evaluation, and improving evaluation efficiency and effectiveness.
Patent Information
- Application Number
- CN202510984372.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-07-16
AI Technical Summary
In existing technologies, it is difficult to intuitively compare the dialogue capabilities of different intelligent agents and large language models, and the evaluation methods are too simplistic to meet diverse evaluation needs.
This paper provides an evaluation method for intelligent agents and large language models. By creating multi-dimensional evaluation tasks, it displays different types of evaluators and allows users to customize configurations, enabling visualization of comparative evaluation results and flexible use of multiple evaluators.
It enables intuitive comparison and evaluation of multiple agents and large language models in the same domain, improving evaluation efficiency and flexibility. It supports custom configuration of multiple evaluation dimensions, making it convenient for users to analyze the effects of agent construction with different configurations.
Smart Images

Figure CN120508507B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of large model technology and intelligent agent technology, specifically to an evaluation method, apparatus, medium, and electronic device for intelligent agents and large language models. Background Technology
[0002] With the rapid development of artificial intelligence technology, it has been applied to many fields. In practical applications, it is necessary to evaluate intelligent objects based on artificial intelligence technology in order to optimize the ability of intelligent objects to handle downstream tasks in practical applications based on the evaluation results. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] Firstly, this disclosure provides an evaluation method for intelligent agents and large language models, including:
[0005] An evaluation task is created, wherein the evaluation task is used to evaluate at least one dialogue of multiple objects to be evaluated in a dialogue process from at least one evaluation dimension, the multiple objects to be evaluated including at least an agent, and the multiple objects to be evaluated are used to implement a dialogue task in the same domain.
[0006] The evaluation task is performed to obtain evaluation results, wherein the evaluation results are used to compare the dialogue capabilities among the plurality of objects to be evaluated;
[0007] The evaluation results are displayed on the first interface;
[0008] The evaluation method also includes:
[0009] The second interface displays different types of evaluators, including at least one of the following: prompt word-based evaluator, code-based evaluator, basic matching rule-based evaluator, retrieval enhancement-based evaluator, human-evaluated evaluator, and natural language processing-based evaluator.
[0010] In response to the configuration operation of the evaluator for the target type in the second interface, a configured evaluator is obtained, wherein the configured evaluator is used to evaluate the dialogue from the corresponding evaluation dimension.
[0011] Secondly, this disclosure provides an evaluation apparatus for intelligent agents and large language models, comprising:
[0012] A creation module is used to create an evaluation task, wherein the evaluation task is used to evaluate at least one dialogue of multiple objects to be evaluated in a dialogue process from at least one evaluation dimension, the multiple objects to be evaluated include at least an agent, and the multiple objects to be evaluated are used to implement a dialogue task in the same domain.
[0013] An evaluation module is used to perform the evaluation task to obtain evaluation results, wherein the evaluation results are used to compare the dialogue capabilities among the multiple objects to be evaluated;
[0014] A first display module is used to display the evaluation results on a first interface;
[0015] The evaluation device for the intelligent agent and the large language model also includes:
[0016] The second display module is used to display a second interface, wherein the second interface displays different types of evaluators, including at least one of prompt word-based evaluators, code-based evaluators, basic matching rule-based evaluators, retrieval enhancement-based evaluators, human-based evaluators, and natural language processing-based evaluators.
[0017] A configuration module is used to obtain a configured evaluator in response to a configuration operation for an evaluator of a target type in the second interface, wherein the configured evaluator is used to evaluate the dialogue from the corresponding evaluation dimension.
[0018] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the evaluation method described in the first aspect.
[0019] Fourthly, this disclosure provides an electronic device, comprising:
[0020] A storage device on which computer programs are stored;
[0021] A processing device for executing the computer program in the storage device to implement the steps of the evaluation method described in the first aspect.
[0022] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the evaluation method described in the first aspect.
[0023] The above technical solution allows for the comparative evaluation of multiple agents that perform dialogue tasks in the same domain. The evaluation results displayed on the first interface provide an intuitive comparison of the dialogue capabilities between different agents, facilitating user analysis of the effectiveness of building agents with different configurations. Furthermore, custom configurations can be performed based on configuration operations to obtain evaluators with different evaluation dimensions.
[0024] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0025] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0026] Figure 1 This is a flowchart illustrating an evaluation method for an intelligent agent and a large language model according to embodiments of this disclosure;
[0027] Figure 2 This is a schematic diagram illustrating an embodiment of the present disclosure for determining the evaluation results of a dialogue across all evaluation dimensions;
[0028] Figure 3 This is a schematic diagram of a first interface shown according to an embodiment of the present disclosure;
[0029] Figure 4 This is a schematic diagram illustrating the configuration of a second interface according to an embodiment of the present disclosure;
[0030] Figure 5 This is a schematic diagram illustrating a configuration interface according to an embodiment of the present disclosure;
[0031] Figure 6 This is a schematic diagram illustrating a third interface according to an embodiment of the present disclosure;
[0032] Figure 7 This is a schematic diagram illustrating a process for creating an evaluation task according to an embodiment of the present disclosure;
[0033] Figure 8 This is a schematic diagram of a fifth interface according to an embodiment of the present disclosure;
[0034] Figure 9 This is a block diagram illustrating an evaluation apparatus for an intelligent agent and a large language model according to embodiments of this disclosure;
[0035] Figure 10 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0036] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0037] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0038] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0039] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0040] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0041] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0042] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0043] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0044] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0045] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0046] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0047] Artificial intelligence-based intelligent objects, such as large language models and agents, have demonstrated outstanding performance in multiple fields. However, the methods used to evaluate LLMs and agents in these technologies are relatively simplistic; the evaluation results are typically presented only for a single large language model or agent in the visualization interface, making it difficult to intuitively compare the evaluation results of different intelligent objects.
[0048] In view of this, the present disclosure provides an evaluation method for intelligent agents and large language models. The embodiments of the present disclosure will be explained and described below with reference to the accompanying drawings.
[0049] Figure 1 This is a flowchart illustrating an evaluation method for an intelligent agent and a large language model according to embodiments of this disclosure. This evaluation method for intelligent agents and large language models can be applied to electronic devices. Furthermore, this evaluation method for intelligent agents and large language models can be executed by an evaluation device for intelligent agents and large language models, wherein the evaluation device for intelligent agents and large language models can be implemented by software and / or hardware, and the software and / or hardware can be configured in an electronic device. (Refer to...) Figure 1 The evaluation method for the intelligent agent and the large language model may include steps 110, 120, 130, 140 and 150.
[0050] In step 110, an evaluation task is created, wherein the evaluation task is used to evaluate at least one dialogue of multiple objects to be evaluated in the dialogue process from at least one evaluation dimension, the multiple objects to be evaluated include at least one agent, and the multiple objects to be evaluated are used to implement a dialogue task in the same domain.
[0051] The types of multiple objects to be evaluated can be singular. As an example, multiple objects to be evaluated are all agents. It is understood that different agents have different configurations, such as relying on different large language models, thereby enabling the comparative evaluation of multiple agents.
[0052] The types of multiple objects to be evaluated can be diverse. As an example, in addition to the agent, the multiple objects to be evaluated can also include a large language model, specifically a first large language model, and the agent is built upon at least a second large language model. The first and second large language models may be different or the same, and the dialogue task implemented by the first large language model is the same as the dialogue task implemented by the agent. For example, a dialogue task refers to providing an input question to an intelligent object and obtaining a corresponding response from the intelligent object; an input question and a response constitute a dialogue. The following explanation and illustration will use the example of multiple objects to be evaluated including an agent and a large language model to illustrate this disclosure.
[0053] As can be seen from the above, the intelligent agent is built upon at least the second largest language model. That is, the intelligent agent is a system built by incorporating the large language model as a component and combining it with other components. For example, it can be configured with plugins that have search capabilities; it can also be configured with knowledge bases containing specialized domain knowledge; it can also be configured with plugins that can generate images from text, and so on.
[0054] As can be seen from the above, the first and second language models can be different, and the performance of the first language model is better than that of the second language model. Based on this, the evaluation results of the agent based on the first language model and the agent that depends on the second language model are used to evaluate whether the agent obtained by configuring additional plugins on the second language model with poor performance can reach or exceed the performance of the first language model. This is suitable for business needs that want to build agents using models with higher cost performance.
[0055] As can be seen from the above, the first and second language models can be the same, that is, the bare model and the agent are evaluated, so as to evaluate the effect of the agent construction.
[0056] The evaluation dimensions can be manually configured or default. It's worth noting that default evaluation dimensions are mandatory. Examples of default evaluation dimensions include the usage of tokens in the dialogue and the time taken for smart objects to generate responses. Manually configured evaluation dimensions can be found in the following related embodiments, which will not be elaborated upon here.
[0057] In step 120, an evaluation task is performed to obtain evaluation results, which are used to compare the dialogue capabilities of multiple objects to be evaluated.
[0058] Since the evaluation task may involve evaluation of multiple evaluation dimensions, the evaluation of different evaluation dimensions can be performed in parallel, thereby improving the evaluation efficiency.
[0059] In step 130, the evaluation results are displayed on the first interface.
[0060] The evaluation results may include at least one of the first evaluation results, the second evaluation results, and the third evaluation results for each object to be evaluated.
[0061] The first evaluation result is used to characterize the evaluation results of each dialogue corresponding to the object to be evaluated under each evaluation dimension.
[0062] The second evaluation result represents the fusion result of the first objective evaluation result, which represents the evaluation results of all dialogues corresponding to the evaluated object under the same evaluation dimension. The second evaluation result allows for the selection of evaluation dimensions to be displayed, showing only the second evaluation results corresponding to the selected dimensions. The fusion result of the first objective evaluation result can be based on weights. Due to different focuses in the dialogues, the weights of each dialogue under different evaluation dimensions may be different. These weights can be manually configured by the user or automatically recommended by the model. As an example, in the case of automatic model recommendation, the model can automatically generate recommended weights for each dialogue in each evaluation dimension based on the dialogue and evaluation dimensions. For example, if the dialogue is about weather, and the evaluation dimension is accuracy, a higher weight can be set for accuracy; if the evaluation dimension is creativity, a lower weight can be set for creativity, thereby improving the accuracy of the evaluation.
[0063] The third evaluation result is used to characterize the fusion result of the second objective evaluation result. The second objective evaluation result is used to characterize the evaluation results of all dialogues corresponding to the object to be evaluated under all evaluation dimensions. Specifically, for the third evaluation result, taking the evaluation dimensions associated with the evaluation task as including RAG (Retrieval-Augmented Generation) dimension and other dimensions (such as accuracy dimension) as an example, since RAG capabilities are not necessarily invoked during inference, the third evaluation result can be determined in the following way:
[0064] First, calculate the evaluation results of the dialogue in other dimensions;
[0065] Next, if the trace data shows that the object to be evaluated invoked RAG capabilities during the reasoning process in response to the dialogue, the evaluation results from other dimensions and the RAG dimension can be fused to obtain the evaluation results of the dialogue across all evaluation dimensions. Figure 2 This is a schematic diagram illustrating how to determine the evaluation results of a dialogue across all evaluation dimensions, according to an embodiment of this disclosure. (Refer to...) Figure 2 The object to be evaluated performs an evaluation task, completes reasoning on the input question, and obtains a dialogue. It is determined whether the RAG capability is invoked during the reasoning process. If the RAG capability is invoked, the retrieved context is extracted from the Trace data. The RAG evaluator determines the evaluation result of the dialogue in the RAG dimension based on the context, and further integrates the evaluation results of the dialogue in the RAG dimension and other dimensions to obtain the evaluation result of the dialogue in all evaluation dimensions, which is the second target evaluation result of the dialogue. If the RAG capability is not invoked, the evaluation results of the dialogue in other evaluation dimensions are integrated to obtain the evaluation result of the dialogue in all evaluation dimensions, which is the second target evaluation result of the dialogue.
[0066] Finally, the evaluation results of the second objective corresponding to all dialogues are merged to obtain the third evaluation result.
[0067] As an example, step 130 may include: displaying the first evaluation result of each object to be evaluated in a first area of the first interface; displaying the second evaluation result of each object to be evaluated in a second area of the first interface through a chart and / or table; and displaying the third evaluation result of each object to be evaluated in a third area of the first interface. The first area, the second area, and the third area are different from each other.
[0068] Tables can display specific numerical values, while charts are used to visually reflect the differences between different objects being evaluated. For example, charts can be radar charts or bar charts. Taking a radar chart as an example, it includes N category axes, each representing an evaluation dimension, where N is an integer greater than 1. Furthermore, when the radar chart includes second evaluation results for different objects, different legends can be used to identify each object. Charts also support scaling; that is, the chart can be enlarged or reduced by performing a scaling operation.
[0069] Figure 3 This is a schematic diagram of a first interface shown according to an embodiment of the present disclosure. It is worth noting that... Figure 3 The evaluation set includes the input questions provided to the object to be evaluated (i.e., Figure 3 The questions shown in the figure) and manually configured reference content, which can be manually configured reference responses to the input questions, can assist in the evaluation of the object to be evaluated.
[0070] Continue to refer to Figure 3The first area 301 of the first interface displays the first evaluation results of each dialogue of different objects to be evaluated in each evaluation dimension, for example... Figure 3 The dimensions shown are 1, 2, 3, 4, and 5. Dimensions 1, 2, and 5 are dimensions obtained based on manual configuration, while dimensions 3 and 4 are the default evaluation dimensions, which are automatically added to the evaluation task when it is configured.
[0071] Continue to refer to Figure 3 The second evaluation result is displayed in the second area 302 of the first interface. The evaluation dimensions corresponding to the second evaluation result can include dimension 1, dimension 2, dimension 3 and dimension 4. The second evaluation results of all objects to be evaluated in dimension 1, dimension 2, dimension 3 and dimension 4 are displayed in the form of radar chart and table.
[0072] Continue to refer to Figure 3 The first screen can also display information on whether the conversation was successful.
[0073] In step 140, a second interface is displayed, which shows different types of evaluators.
[0074] In step 150, in response to the configuration operation of the evaluator for the target type in the second interface, a configured evaluator is obtained, wherein the configured evaluator is used to evaluate the dialogue from the corresponding configured evaluation dimensions.
[0075] The above evaluation can be implemented based on evaluators, each corresponding to an evaluation dimension. Users can pre-configure evaluators and select the appropriate evaluator when configuring an evaluation task to evaluate the dialogue based on the corresponding evaluation dimension.
[0076] In this embodiment, different types of evaluators may include at least one of prompt-based evaluators, code-based evaluators, basic matching rule-based evaluators, retrieval-enhanced generation-based evaluators, human-based evaluators, and natural language processing-based evaluators. These evaluators can be regarded as internal evaluators. These internal evaluators can pre-configure basic content, which may be content shared by different evaluation dimensions, in order to improve configuration efficiency.
[0077] In addition, when configuring evaluation tasks, external evaluators can be selected. These external evaluators are provided by third parties and can be accessed via gateways or other means to meet high-level and complex evaluation needs.
[0078] The above technical solution allows for the comparative evaluation of multiple agents performing dialogue tasks within the same domain. The evaluation results displayed on the first interface provide a clear visual comparison of the dialogue capabilities of different agents, facilitating user analysis of the effectiveness of agent configurations. Furthermore, it enables the evaluation of two types of objects: a large language model and an agent. The evaluation results for various objects are visualized, allowing evaluators to perform comparative analysis and review. When the large language models corresponding to the two types of objects are identical, it facilitates the evaluation of the agent's performance. Conversely, when the large language models are inconsistent, it facilitates the evaluation of whether an agent built based on the second large language model can achieve or surpass the performance of the bare model. Additionally, custom configurations can be performed based on configuration operations, resulting in evaluators with different evaluation dimensions.
[0079] Figure 4 This is a schematic diagram illustrating the configuration of a second interface according to an embodiment of this disclosure. (Refer to...) Figure 4 The second interface displays different types of evaluators, with the evaluator based on the prompt word corresponding to... Figure 4 The Prompt type in [the dataset]. A cue-based evaluator can be implemented using an LLM (Local Management Model). The LLM scores the output of the object to be evaluated based on predefined cue words. Cue words can be preset or user-defined, and different cue words correspond to different evaluation dimensions, such as correctness, creativity, depth, balance, and accuracy.
[0080] Code-based evaluator correspondence Figure 4 The Code type in the code is used. Code-based evaluators are implemented based on code, and different codes can correspond to different evaluation rules. Different evaluation rules correspond to different evaluation dimensions. For example, it supports configuring evaluation rules by writing code such as Python and Java.
[0081] Evaluator based on basic matching rules Figure 4 The basic types in the algorithm. The evaluator based on basic matching rules is implemented based on rules such as regular expressions and fuzzy matching. Users can configure rules such as regular expressions and fuzzy matching to perform evaluations from evaluation dimensions such as detecting specific keywords or format errors.
[0082] The evaluator based on retrieval enhancement corresponds to Figure 4The RAG type in [the context of search enhancement]. The evaluator based on search enhancement can be implemented using LLM, evaluating search relevance, response accuracy, and knowledge coverage. Search relevance describes the cosine similarity between the input question and the recalled documents; response accuracy describes the similarity (e.g., ROUGE-L) and entity matching degree between the reasoning result (i.e., the response) of the evaluated object and the retrieved documents; knowledge coverage describes whether the reasoning result of the evaluated object covers the key information of the retrieved documents, and knowledge coverage can be achieved using TF-IDF (Term Frequency-Inverse Document Frequency).
[0083] Evaluator based on human evaluation Figure 4 The human-based rule type. Human-based evaluators allow users to define evaluation criteria via text, scores, or options, and then perform the evaluation manually.
[0084] The evaluator based on natural language processing corresponds to Figure 4 The NLP (Natural Language Processing) type in this context refers to an evaluator based on natural language processing. This evaluator utilizes natural language processing techniques and its evaluation dimensions include, for example, user satisfaction and conversational coherence.
[0085] As can be seen from the above, an evaluator of a certain type can be configured with multiple evaluation logics, and each evaluation logic corresponds to an evaluation dimension. Therefore, an evaluator that evaluates from each evaluation dimension can be obtained based on manually customized configuration.
[0086] The configuration operations of the evaluator may include at least one of the following: configuration operations for the prompt words on which the evaluator depends, configuration operations for the variables associated with the prompt words on which the evaluator depends, configuration operations for the evaluation model on which the evaluator depends, configuration operations for the model parameters of the evaluation model, and configuration operations for the evaluation logic of the evaluator.
[0087] For pre-built evaluators (such as cue-based evaluators), the cue words can be modified according to actual needs.
[0088] For objects to be evaluated with multiple configured variables, the default evaluators (such as prompt-based evaluators) cannot support adjustments to the evaluator's input fields or their format based on changing variables. Therefore, the evaluator cannot adapt to dynamically variable objects, leading to errors and unrecognizable results, thus affecting the accuracy of the evaluation and the judgment of the evaluators. To solve this problem, configuration operations for variables associated with the prompt words upon which the default evaluator depends are supported, allowing adjustments to the evaluator's input parameters based on actual conditions. This adjustment can include adding variables or changing their format.
[0089] For evaluators that are bound to a specific underlying model, the inability to switch the dependent model results in evaluation performance being limited by the model's capabilities and exhibiting low flexibility in adapting to different scenarios. Therefore, it is necessary to provide configuration options for the evaluation model that the evaluator depends on, as well as the model parameters of the evaluation model.
[0090] For rule-based or code-based evaluators, the evaluation logic can be configured according to actual needs to adapt to different scenarios.
[0091] Figure 5 This is a schematic diagram illustrating a configuration interface according to an embodiment of the present disclosure. This configuration interface can be displayed as a floating layer on top of a second interface. (In conjunction with...) Figure 4 After selecting a target type, you can configure the evaluator to obtain an evaluator that belongs to the target type and can evaluate the dialogue from the corresponding configured evaluation dimensions. When the configuration of the evaluator for the target type is triggered, it will be displayed. Figure 5 The configuration interface shown. Continue referring to... Figure 5 By configuring parameters for the evaluator, such as the model and prompts, the configured evaluator can evaluate the dialogue from the perspective of correctness.
[0092] For code-based evaluators, testing capabilities can also be provided. Therefore, in some embodiments, the evaluation method for the aforementioned intelligent agent and large language model may further include: in response to a debugging operation on a configured evaluator, displaying a third interface, wherein the third interface displays the evaluator's program code and an input parameter editing area, the program code being used to characterize the evaluator's evaluation logic, and the input parameter editing area being used to edit the input parameters of the test program code; in response to a test request for the program code, running the program code based on the input parameters to obtain test results; and displaying the test results in the test result area of the third interface.
[0093] Figure 6 This is a schematic diagram illustrating a third interface according to an embodiment of the present disclosure. (Refer to...) Figure 6The system displays the evaluator's program code and supports online editing of the code. The input parameter editing area 601 is used to edit the input parameters of the test program code. When a test request is initiated against the test control 602, the program code runs based on the input parameters edited in the input parameter editing area 601, and the test results are displayed in the test result area 603. Users can adjust the program code based on the test results to configure the evaluator. This enhances the ability to customize the evaluator's configuration.
[0094] Since users may configure different evaluators according to different needs, and some pre-built evaluators depend on different model architectures, cross-framework communication and resource scheduling issues arise. Furthermore, related technologies do not support dynamic parsing of dynamic input parameters for evaluators. Therefore, this paper proposes encapsulating each configured evaluator as an independent plugin and managing each plugin through a predefined unified interface.
[0095] The predefined unified interface can be either an RPC (Remote Procedure Call) interface or an HTTP (Hypertext Transfer Protocol) interface. Communication with plugins is achieved through this predefined unified interface, resolving issues related to cross-framework communication and resource scheduling, and the lack of support for dynamic parsing of dynamic input parameters for the evaluator. The interface's input parameters can include the dialog of the object to be evaluated, plugin metadata, and configuration parameters. Plugin metadata is used to identify the plugin being invoked, i.e., the evaluator being invoked; hot updates of plugins can be implemented based on configuration parameters. The interface's output parameters include the evaluation results and scoring rationale, etc.
[0096] In some embodiments, each plugin can run as an independent microservice. As an example, plugin runtime deployment can be achieved using the Docker container engine.
[0097] In some embodiments, when multiple plugins are invoked, parallel invocation of multiple plugins can be supported to improve evaluation efficiency.
[0098] In some embodiments, the evaluation task is further used to instruct multiple objects to be evaluated to reason about an input question in a selected target dataset to generate at least one dialogue. The steps of creating the evaluation task described above can be implemented as follows: displaying a fourth interface for creating the evaluation task; determining a target dataset in response to a selection operation for a dataset in the fourth interface; determining multiple objects to be evaluated in response to a selection operation for candidate objects to be evaluated in the fourth interface; determining an evaluator in response to a selection operation for a configured candidate evaluator and / or a group of candidate evaluators in the fourth interface, the evaluator being used to evaluate at least one dialogue from a corresponding evaluation dimension; and creating the evaluation task based on the evaluator, the multiple objects to be evaluated, and the target dataset in response to a creation operation in the fourth interface.
[0099] It is understood that the evaluation task in this embodiment is not only used to instruct each object to be evaluated to reason about the input question in the selected target dataset to generate at least one dialogue, but also to evaluate at least one dialogue of each object to be evaluated in the dialogue process from at least one evaluation dimension, and the evaluation is implemented by calling the evaluator.
[0100] Candidate evaluator groups can be created in advance. During creation, the group name and corresponding group description of the candidate evaluator group can be configured, as well as the evaluators in the candidate evaluator group can be configured.
[0101] Figure 7 This is a schematic diagram illustrating a process for creating an evaluation task according to an embodiment of this disclosure. The target dataset includes an input question and reference content. Candidate evaluation objects include those belonging to a large language model and those belonging to an agent. A selection is made from the candidate evaluation objects to obtain the evaluation object. In this embodiment, when executing the configured evaluation task, the input question is first reasoned using the evaluation object to obtain a response, thus obtaining the corresponding dialogue. Then, the evaluator is invoked to evaluate the dialogue from the corresponding evaluation dimensions to obtain the evaluation result.
[0102] Continue to refer to Figure 7 When selecting an evaluator to evaluate at least one dialogue from at least one evaluation dimension, a pre-configured set of candidate evaluators can be directly reused, avoiding the need for repeated configuration of evaluators for periodically conducted evaluation tasks.
[0103] In other embodiments, when creating an evaluation task, an evaluator can be selected in real time from a provided evaluator 1, evaluator 2, ..., evaluator M to determine the evaluator used to evaluate at least one dialogue from at least one evaluation dimension. It is understood that M is an integer greater than 2.
[0104] In other embodiments, a pre-configured set of candidate evaluators can be directly reused and fine-tuned to obtain an evaluator for evaluating at least one dialogue from at least one evaluation dimension.
[0105] As can be seen from the above, the evaluation dimensions can include not only manually configured dimensions, but also mandatory evaluation dimensions. Therefore, when creating an evaluation task, the evaluators associated with the evaluation task include not only the evaluators selected by the user from the candidate evaluators, but also the evaluators corresponding to the mandatory evaluation dimensions.
[0106] In some embodiments, the evaluation method for the above-described agent and large language model may further include: displaying a fifth interface in response to a selection operation on a target dialogue in a first interface, wherein the fifth interface is at least used to display the evaluation results of the target dialogue under each evaluation object and under each evaluation dimension; and displaying the scoring reason in response to a hover operation on the evaluation results of the target dialogue in the fifth interface, wherein the evaluation results are obtained by an evaluator based on prompt words.
[0107] Among them, the target dialogue is a set of dialogues in a round of dialogue performed by the object to be evaluated. A round of dialogue may include multiple sets of dialogues.
[0108] Figure 8 This is a schematic diagram illustrating a fifth interface according to an embodiment of the present disclosure. (Refer to...) Figure 8 The fifth interface 801 can be displayed on top of the first interface as a floating layer. The content displayed on the fifth interface 801 can include the target dialogue (including the input question and the reasoning results of the intelligent object for the dialogue), the evaluation results of the target dialogue under various evaluation dimensions, the deep thinking process of the large language model, and so on.
[0109] Continue to refer to Figure 8 The scoring reason 802 can be displayed as a floating layer on the fifth interface 801. By displaying the scoring reason, the evaluator can intuitively understand the evaluation logic of the evaluator.
[0110] Based on the same concept, embodiments of this disclosure provide an evaluation apparatus for intelligent agents and large language models. Figure 9 This is a block diagram illustrating an evaluation apparatus for an intelligent agent and a large language model according to embodiments of this disclosure, with reference to... Figure 9 The evaluation device 900 for the intelligent agent and the large language model includes:
[0111] A creation module 901 is used to create an evaluation task, wherein the evaluation task is used to evaluate at least one dialogue of multiple objects to be evaluated in a dialogue process from at least one evaluation dimension, the multiple objects to be evaluated include at least an agent, and the multiple objects to be evaluated are used to implement a dialogue task in the same domain.
[0112] Evaluation module 902 is used to perform the evaluation task to obtain evaluation results, wherein the evaluation results are used to compare the dialogue capabilities among the plurality of objects to be evaluated;
[0113] The first display module 903 is used to display the evaluation results on the first interface;
[0114] The evaluation device 900 for the intelligent agent and the large language model also includes:
[0115] The second display module 904 is used to display a second interface, wherein the second interface displays different types of evaluators, including at least one of prompt word-based evaluators, code-based evaluators, basic matching rule-based evaluators, retrieval enhancement-based evaluators, human-based evaluators, and natural language processing-based evaluators.
[0116] The configuration module 905 is used to obtain a configured evaluator in response to the configuration operation of the evaluator for the target type in the second interface, wherein the configured evaluator is used to evaluate the dialogue from the corresponding evaluation dimension.
[0117] Optionally, the configuration operation includes at least one of the following:
[0118] Configuration operations for the prompt words that the evaluator depends on;
[0119] Configuration operations for variables associated with the prompt words on which the evaluator depends;
[0120] Configuration operations for the evaluation model on which the evaluator depends;
[0121] Configuration operations for the model parameters of the evaluation model;
[0122] Configuration operations for the evaluation logic of the evaluator.
[0123] Optionally, the evaluation device 900 for the intelligent agent and the large language model further includes:
[0124] The third display module is used to display a third interface in response to a debugging operation on a configured evaluator. The third interface displays the program code of the evaluator and an input parameter editing area. The program code is used to characterize the evaluation logic of the evaluator, and the input parameter editing area is used to edit the input parameters for testing the program code.
[0125] The testing module is used to respond to a test request for the program code, run the program code based on the input parameters, and obtain test results;
[0126] The fourth display module is used to display the test results in the test result area of the third interface.
[0127] Optionally, each configured evaluator is encapsulated as an independent plugin, and each plugin is managed through a predefined unified interface.
[0128] Optionally, the evaluation task is further configured to instruct the plurality of objects to be evaluated to reason about an input question in a selected target dataset to generate the at least one dialogue, and the creation module 901 is further configured to:
[0129] A fourth interface is displayed, wherein the fourth interface is used to create an evaluation task;
[0130] In response to a selection operation for a dataset in the fourth interface, the target dataset is determined;
[0131] In response to the selection operation for candidate objects to be evaluated in the fourth interface, the plurality of objects to be evaluated are determined;
[0132] In response to the selection operation for the configured candidate evaluator and / or candidate evaluator group in the fourth interface, an evaluator is determined, wherein the evaluator is used to evaluate the at least one dialogue from the corresponding evaluation dimension;
[0133] In response to the creation operation in the fourth interface, the evaluation task is created based on the evaluator, the plurality of objects to be evaluated, and the target dataset.
[0134] Optionally, the first display module 903 is further configured to:
[0135] The first evaluation result of each of the objects to be evaluated is displayed in the first area of the first interface, wherein the first evaluation result is used to characterize the evaluation result of each dialogue corresponding to the object to be evaluated under each of the evaluation dimensions;
[0136] In the second area of the first interface, the second evaluation results of each of the objects to be evaluated in each of the evaluation dimensions are displayed by charts or tables. The second evaluation results are used to characterize the fusion result of the first target evaluation results. The first target evaluation results are used to characterize the evaluation results of all the dialogues corresponding to the objects to be evaluated under the same evaluation dimension.
[0137] The third evaluation result of each of the objects to be evaluated is displayed in the third area of the first interface. The third evaluation result is used to characterize the fusion result of the second target evaluation result. The second target evaluation result is used to characterize the evaluation results of all the dialogues corresponding to the objects to be evaluated under all the evaluation dimensions.
[0138] Optionally, the chart includes a radar chart, which identifies each of the objects to be evaluated by different legends. The radar chart includes N category axes, each category axis representing one of the evaluation dimensions, where N is an integer greater than 1.
[0139] Optionally, the evaluation device 900 for the intelligent agent and the large language model further includes:
[0140] The fifth display module is used to display a fifth interface in response to a selection operation on the target dialogue in the first interface, wherein the fifth interface is used to display the evaluation results of the target dialogue under each of the objects to be evaluated and under each of the evaluation dimensions.
[0141] The sixth display module is used to display the scoring reason in response to a hover operation on the fifth interface for the evaluation result of the target dialogue, wherein the evaluation result is obtained by an evaluator based on prompt words.
[0142] Optionally, the plurality of objects to be evaluated also includes a first major language model, which is used to implement the dialogue task, and the agent is built based on at least a second major language model, wherein the first major language model and the second major language model are different or the same.
[0143] The implementation methods of each module in the evaluation device 900 for the intelligent agent and the large language model can refer to the above-mentioned related embodiments, and will not be repeated here.
[0144] Based on the same concept, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described evaluation method for intelligent agents and large language models.
[0145] Based on the same concept, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described evaluation method for intelligent agents and large language models.
[0146] Based on the same concept, embodiments of this disclosure also provide an electronic device, including:
[0147] A storage device on which computer programs are stored;
[0148] A processing device for executing the computer program in the storage device to implement the steps of the above-described evaluation method for intelligent agents and large language models.
[0149] The following is for reference. Figure 10The diagram illustrates a structural schematic of an electronic device 1000 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0150] like Figure 10 As shown, the electronic device 1000 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for the operation of the electronic device 1000. The processing unit 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0151] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic device 1000 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 1000 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0152] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1008, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.
[0153] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0154] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0155] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0156] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: create an evaluation task, wherein the evaluation task is used to evaluate at least one dialogue of a plurality of objects to be evaluated in a dialogue process from at least one evaluation dimension, the plurality of objects to be evaluated including at least an agent, the plurality of objects to be evaluated being used to implement a dialogue task in the same domain; execute the evaluation task to obtain an evaluation result, wherein the evaluation result is used to compare the dialogue capabilities among the plurality of objects to be evaluated; display the evaluation result on a first interface; the evaluation method further includes: displaying a second interface, wherein the second interface displays different types of evaluators, the different types of evaluators including at least one of prompt-based evaluators, code-based evaluators, basic matching rule-based evaluators, retrieval-enhanced generation evaluators, human-evaluated evaluators, and natural language processing-based evaluators; and, in response to a configuration operation for an evaluator of a target type on the second interface, obtain a configured evaluator, wherein the configured evaluator is used to evaluate the dialogue from the corresponding evaluation dimension.
[0157] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0159] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0160] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0162] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0163] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0164] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. An evaluation method for intelligent agents and large language models, characterized in that, The evaluation method includes: An evaluation task is created, wherein the evaluation task is used to evaluate at least one dialogue of multiple objects to be evaluated in a dialogue process from at least one evaluation dimension, the multiple objects to be evaluated including at least an agent, and the multiple objects to be evaluated are used to implement a dialogue task in the same domain. The evaluation task is performed to obtain evaluation results, wherein the evaluation results are used to compare the dialogue capabilities among the plurality of objects to be evaluated; The evaluation results are displayed on the first interface; The evaluation method also includes: The second interface displays different types of evaluators, including at least one of the following: prompt word-based evaluator, code-based evaluator, basic matching rule-based evaluator, retrieval enhancement-based evaluator, human-evaluated evaluator, and natural language processing-based evaluator. In response to the configuration operation of the evaluator for the target type in the second interface, a configured evaluator is obtained, wherein the configured evaluator is used to evaluate the dialogue from the corresponding evaluation dimension; The evaluation task is further configured to instruct the plurality of objects to be evaluated to reason about an input question in a selected target dataset to generate the at least one dialogue. Creating the evaluation task includes: displaying a fourth interface, wherein the fourth interface is used to create the evaluation task; determining the target dataset in response to a selection operation for a dataset in the fourth interface; determining the plurality of objects to be evaluated in response to a selection operation for candidate objects to be evaluated in the fourth interface; determining an evaluator in response to a selection operation for a configured candidate evaluator and / or a group of candidate evaluators in the fourth interface, wherein the evaluator is used to evaluate the at least one dialogue from the corresponding evaluation dimension; and creating the evaluation task based on the evaluator, the plurality of objects to be evaluated, and the target dataset in response to a creation operation in the fourth interface. The display of the evaluation results on the first interface includes: The first evaluation result of each of the objects to be evaluated is displayed in the first area of the first interface, wherein the first evaluation result is used to characterize the evaluation result of each dialogue corresponding to the object to be evaluated under each of the evaluation dimensions; In the second area of the first interface, the second evaluation results of each of the objects to be evaluated in each of the evaluation dimensions are displayed by charts or tables. The second evaluation results are used to characterize the fusion result of the first target evaluation results. The first target evaluation results are used to characterize the evaluation results of all the dialogues corresponding to the objects to be evaluated under the same evaluation dimension. The third evaluation result of each of the objects to be evaluated is displayed in the third area of the first interface. The third evaluation result is used to characterize the fusion result of the second target evaluation result. The second target evaluation result is used to characterize the evaluation results of all the dialogues corresponding to the objects to be evaluated under all the evaluation dimensions.
2. The evaluation method according to claim 1, characterized in that, The configuration operation includes at least one of the following: Configuration operations for the prompt words that the evaluator depends on; Configuration operations for variables associated with the prompt words on which the evaluator depends; Configuration operations for the evaluation model on which the evaluator depends; Configuration operations for the model parameters of the evaluation model; Configuration operations for the evaluation logic of the evaluator.
3. The evaluation method according to claim 1, characterized in that, The evaluation method also includes: In response to a debugging operation on a configured evaluator, a third interface is displayed, wherein the third interface displays the program code of the evaluator and an input parameter editing area, the program code being used to characterize the evaluation logic of the evaluator, and the input parameter editing area being used to edit the input parameters for testing the program code; In response to a test request for the program code, the program code is run based on the input parameters to obtain test results; The test results are displayed in the test results area of the third interface.
4. The evaluation method according to claim 1, characterized in that, Each configured evaluator is encapsulated as an independent plugin, and each plugin is managed through a predefined unified interface.
5. The evaluation method according to claim 1, characterized in that, The chart includes a radar chart, which identifies each of the objects to be evaluated by different legends. The radar chart includes N category axes, each of which is used to represent one of the evaluation dimensions, where N is an integer greater than 1.
6. The evaluation method according to claim 1, characterized in that, The evaluation method also includes: In response to a selection operation for a target dialogue in the first interface, a fifth interface is displayed, wherein the fifth interface is at least used to display the evaluation results of the target dialogue under each of the objects to be evaluated and under each of the evaluation dimensions; In response to a hover operation on the fifth interface for the evaluation result of the target dialogue, a rating reason is displayed, wherein the evaluation result is obtained by an evaluator based on prompt words.
7. The evaluation method according to any one of claims 1-6, characterized in that, The plurality of objects to be evaluated also includes a first major language model, which is used to implement the dialogue task. The agent is built based on at least a second major language model, and the first major language model and the second major language model may be different or the same.
8. An evaluation device for intelligent agents and large language models, characterized in that, include: A creation module is used to create an evaluation task, wherein the evaluation task is used to evaluate at least one dialogue of multiple objects to be evaluated in a dialogue process from at least one evaluation dimension, the multiple objects to be evaluated include at least an agent, and the multiple objects to be evaluated are used to implement a dialogue task in the same domain. An evaluation module is used to perform the evaluation task to obtain evaluation results, wherein the evaluation results are used to compare the dialogue capabilities among the multiple objects to be evaluated; A first display module is used to display the evaluation results on a first interface; The evaluation device for the intelligent agent and the large language model also includes: The second display module is used to display a second interface, wherein the second interface displays different types of evaluators, including at least one of prompt word-based evaluators, code-based evaluators, basic matching rule-based evaluators, retrieval enhancement-based evaluators, human-based evaluators, and natural language processing-based evaluators. A configuration module is used to respond to the configuration operation of the evaluator for the target type in the second interface to obtain a configured evaluator, wherein the configured evaluator is used to evaluate the dialogue from the corresponding evaluation dimension. The evaluation task is also used to instruct the plurality of objects to be evaluated to reason about input questions in a selected target dataset to generate the at least one dialogue, and the creation module is further used to: A fourth interface is displayed, wherein the fourth interface is used to create an evaluation task; In response to a selection operation for a dataset in the fourth interface, the target dataset is determined; In response to the selection operation for candidate objects to be evaluated in the fourth interface, the plurality of objects to be evaluated are determined; In response to the selection operation for the configured candidate evaluator and / or candidate evaluator group in the fourth interface, an evaluator is determined, wherein the evaluator is used to evaluate the at least one dialogue from the corresponding evaluation dimension; In response to the creation operation in the fourth interface, the evaluation task is created based on the evaluator, the plurality of objects to be evaluated, and the target dataset. The first display module is further configured to: The first evaluation result of each of the objects to be evaluated is displayed in the first area of the first interface, wherein the first evaluation result is used to characterize the evaluation result of each dialogue corresponding to the object to be evaluated under each of the evaluation dimensions; In the second area of the first interface, the second evaluation results of each of the objects to be evaluated in each of the evaluation dimensions are displayed by charts or tables. The second evaluation results are used to characterize the fusion result of the first target evaluation results. The first target evaluation results are used to characterize the evaluation results of all the dialogues corresponding to the objects to be evaluated under the same evaluation dimension. The third evaluation result of each of the objects to be evaluated is displayed in the third area of the first interface. The third evaluation result is used to characterize the fusion result of the second target evaluation result. The second target evaluation result is used to characterize the evaluation results of all the dialogues corresponding to the objects to be evaluated under all the evaluation dimensions.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program implements the steps of the evaluation method according to any one of claims 1-7.
10. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the evaluation method according to any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the evaluation method according to any one of claims 1-7.
Citation Information
Patent Citations
Psychological counseling dialogue evaluation system based on large language model
CN118983091A
Intelligent agent evaluation method and device, electronic equipment and storage medium
CN119127648A