Intelligent agent or large model evaluation method, device, medium, equipment and product
By creating target tasks and acquiring dialogue data in real time, intelligent agents or large models are evaluated, which solves the problems of low evaluation iteration frequency and insufficient real-time performance in existing technologies. This achieves an efficient and flexible evaluation method that can adapt to the evaluation needs of different fields.
Patent Information
- Application Number
- CN202511151587.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-08-15
AI Technical Summary
In existing technologies, the evaluation methods for intelligent agents or large models have low iteration frequency and low real-time performance, which cannot adapt to rapid business iteration. Furthermore, the application scenarios in different fields vary greatly, the coverage of general evaluation sets is not strong, the construction cost is high, and the quality is difficult to guarantee.
By creating target tasks, the system evaluates agents or large models using dialogue data that meets the filtering criteria. It acquires dialogue data in real time and displays the evaluation results through an interface. It supports flexible configuration of evaluation dimensions and conditions, achieving both realism and real-time evaluation.
It reduces the cost of building evaluation sets, improves the authenticity and real-time performance of evaluations, adapts to the evaluation needs of different fields, and meets the requirements of rapid business iteration.
Smart Images

Figure CN120723608B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular, to an agent or large model evaluation method and device, medium, equipment and product. BACKGROUND
[0002] With the development of artificial intelligence technology, agents and large models are gradually applied to more fields and application scenarios. In order to ensure the performance of agents or large models in actual application, it is usually necessary to systematically evaluate them, so as to optimize the agents or large models based on the evaluation results. Therefore, the evaluation of agents or large models is of more and more important significance. SUMMARY
[0003] This summary is provided to introduce a selection of concepts, which will be described with greater specificity in the detailed description section. This summary does not intend to identify key or essential features of the claimed technology, nor is it intended for use in determining the scope of the claimed technology.
[0004] In a first aspect, the present disclosure provides an agent or large model evaluation method, which comprises:
[0005] creating a target task, the target task being used to evaluate an evaluation object from at least one evaluation dimension by using dialogue data meeting a screening condition, the dialogue data being a question input to the evaluation object and an answer output by the evaluation object for the question in an interaction process of the evaluation object, the evaluation object comprising an agent or a large model;
[0006] continuously acquiring dialogue data of the evaluation object in a process of executing the target task;
[0007] in response to acquiring target dialogue data meeting the screening condition, determining a target evaluation result corresponding to the target dialogue data;
[0008] displaying the target dialogue data and the target evaluation result through a first interface.
[0009] In a second aspect, the present disclosure provides an agent or large model evaluation device, which comprises:
[0010] a task creation module, configured to create a target task, the target task being used to evaluate an evaluation object from at least one evaluation dimension by using dialogue data meeting a screening condition, the dialogue data being a question input to the evaluation object and an answer output by the evaluation object for the question in an interaction process of the evaluation object, the evaluation object comprising an agent or a large model;
[0011] an acquisition module, configured to continuously acquire dialogue data of the evaluation object in a process of executing the target task;
[0012] a first determination module, configured to determine a target evaluation result corresponding to the target dialogue data in response to acquisition of the target dialogue data meeting the screening condition;
[0013] a first display module, configured to display the target dialogue data and the target evaluation result through a first interface.
[0014] In a third aspect, the present disclosure provides a computer readable medium, having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method of the first aspect of the present disclosure.
[0015] In a fourth aspect, the present disclosure provides an electronic device, comprising:
[0016] a storage device, having a computer program stored thereon;
[0017] a processing device, configured to execute the computer program in the storage device to implement the steps of the method of the first aspect of the present disclosure.
[0018] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method of the first aspect of the present disclosure.
[0019] Through the above technical solution, by creating a target task for evaluating an agent or a large model, real dialogue data of the agent or the large model is captured in real time in the process of executing the target task and applied to the evaluation of the agent or the large model, and then the target evaluation result obtained is displayed through a first interface. Thus, the authenticity and real-time performance of the evaluation set can be ensured, and the cost of constructing the evaluation set can be greatly reduced.
[0020] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0021] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings in which:
[0022] Figure 1 is a flowchart of an evaluation method of an agent or a large model according to an embodiment of the present disclosure;
[0023] Figure 2is an exemplary schematic diagram of a second interface in a method for evaluating an agent or a large model according to the present disclosure;
[0024] Figure 3 is an exemplary schematic diagram of an interface for triggering selection of an evaluator in a method for evaluating an agent or a large model according to the present disclosure;
[0025] Figure 4 is an exemplary schematic diagram of a link details interface in a method for evaluating an agent or a large model according to the present disclosure;
[0026] Figure 5 is an exemplary schematic diagram of a first interface in a method for evaluating an agent or a large model according to the present disclosure;
[0027] Figure 6 is a block diagram of an apparatus for evaluating an agent or a large model according to an embodiment of the present disclosure;
[0028] Figure 7 A structural schematic of an electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0029] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0030] It should be understood that the various steps of the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.
[0031] The term "comprising" and variations thereof as used herein are open-ended, that is, "comprising but not limited to." The term "based on" is "based, at least in part, on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments." Related terms are defined in the description that follows.
[0032] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0033] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0034] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0035] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.
[0036] For example, in response to receiving the active request of the user, the prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0037] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of pop-up window, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0038] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0039] At the same time, it can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the present technical solution should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0040] As described in the background, the evaluation of the intelligent agent or large model is of great significance to its performance. In the related art, the intelligent agent or large model is usually evaluated in an offline manner through a preset evaluation set. This evaluation method has a low iteration frequency and a relatively fixed evaluation dimension. In the case where the business iteration frequency is much faster than the intelligent agent or large model application, this evaluation method has the problems of insufficient agility and low real-time performance. In addition, application scenarios in different fields have great differences, and it is difficult to cover them using a general evaluation set. Building an evaluation set for each field separately has the problems of high construction cost, difficult quality guarantee, and weak coverage.
[0041] To solve the above problems in the prior art, the present disclosure provides an intelligent agent or large model evaluation method, device, medium, equipment and product.
[0042] Figure 1 The flowchart of the intelligent agent or large model evaluation method according to an embodiment of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the method provided by the present disclosure can include steps 11 to 14. Figure 1
[0043] In step 11, a target task is created, which is used to evaluate the evaluation object from at least one evaluation dimension using dialog data that meets the filtering condition. The dialog data is the question input to the evaluation object and the answer output by the evaluation object during the interaction process. The evaluation object includes an intelligent agent or a large model.
[0044] Generally, during the interaction with the intelligent agent or large model, a question is input to the intelligent agent or large model to obtain the answer output by the intelligent agent or large model to the question. This question and answer can be used as a dialog data.
[0045] The evaluation dimension can be a demand-configured evaluation dimension or a default evaluation dimension. Generally, a plurality of evaluators can be preconfigured, and each evaluator can correspond to an evaluation dimension. By selecting different evaluators, the demand-configured evaluation dimension can be achieved. In addition, one or more default evaluators can be set to use the default evaluation dimension by using the default evaluator. For example, the evaluation dimension can include, but is not limited to, correctness, first token (token) time consumption, single-round dialog time consumption, token consumption, etc.
[0046] In the present disclosure, different evaluation tasks can be created for different evaluation objects, and different evaluation tasks are independent of and do not affect each other. The present disclosure will be described with respect to the target task. In actual application, when an evaluation task is created for an evaluation object, it can be used as a target task, and a series of steps provided by the present disclosure can be performed.
[0047] The following first illustrates the relevant steps involved in the creation scenario of the target task.
[0048] In a possible implementation, the method provided by the present disclosure can further include the following steps:
[0049] In response to receiving the first instruction for creating the target task for the evaluation object, a second interface is displayed, the second interface including a first configuration area for configuring the screening condition, the first configuration area displaying a first trigger option for adding the screening condition;
[0050] In response to a trigger operation on the first trigger option, a first configuration item for configuring a property field to be screened, a second configuration item for configuring reference content, and a third configuration item for configuring comparison logic between the property field and the reference content are displayed in the first configuration area; wherein the property field includes at least one of a property field in link tracking data associated with the dialogue data, a content field of the dialogue data, and a label field associated with the dialogue data;
[0051] In response to a configuration operation on the first trigger option, the first configuration item, the second configuration item, and the third configuration item, the screening condition corresponding to the target task is determined.
[0052] Optionally, the user can select the evaluation object and create an evaluation task, i.e., the target task, for the evaluation object. In response to receiving the first instruction for creating the target task, a second interface for configuring the target task can be displayed. The second interface can include a first configuration area for configuring the screening condition, the first configuration area displaying a first trigger option for adding the screening condition. For example, the first configuration area can be as shown in A1 of Figure 2 , the first trigger option can be as shown in B1 of Figure 2 By clicking the first trigger option, a screening condition can be added.
[0053] When the first trigger option is not triggered, that is, if the target task is created in a state where the first trigger option is not triggered, it means that the target task is not set with a screening condition. In actual application, the dialogue data can be directly used to evaluate the evaluation object without screening the dialogue data. Alternatively, a default screening condition can be set, and in the case where the user does not set any screening condition, the default screening condition is directly called to screen the obtained dialogue data, and the dialogue data meeting the default screening condition is used to evaluate the evaluation object.
[0054] When the first trigger option is triggered, in response to the triggering operation for the first trigger option, a first configuration item for configuring a screened attribute field, a second configuration item for configuring reference content, and a third configuration item for configuring comparison logic between the attribute field and the reference content can be displayed in the first configuration area. The attribute field can include at least one of an attribute field in link tracking data associated with the dialogue data, a content field of the dialogue data, and a label field associated with the dialogue data.
[0055] The first configuration item is used to configure the screened attribute field. When the dialogue data is screened, the field values of the dialogue data corresponding to the attribute fields can be obtained and screened.
[0056] The attribute field can include an attribute field in link tracking data associated with the dialogue data. Generally, each time an intelligent agent or a large model has an effective interaction (one question and one answer) with a user, a link tracking data (i.e., trace data) corresponding to a unique identifier (TraceID) is generated, and the link tracking data records relevant information at the time of interaction, corresponding to a series of attribute fields. Optionally, the attribute fields in the link tracking data (i.e., trace data) can include, but are not limited to, time consumption, Token quantity, first character reply time consumption, first character reply time, start time, end time, node type, node name, call type, node state, input, and output. For example, the time consumption and the first character reply time consumption can be in units of milliseconds (ms). For another example, the first character reply time, the start time, and the end time can include year, month, day, and time point information. For another example, the node type can be user input. For another example, the call type can be user input. For another example, the node state can be success or failure. For another example, the input and the output can be the question and the actual output answer, respectively, at the time of interaction.
[0057] Optionally, the content of the first configuration item for configuring the screened attribute field can be manually input by the user, such as directly inputting the first character reply time consumption as the configuration content, or a single selection option can be provided for the user to select directly according to all the attribute fields contained in the link tracking data.
[0058] The attribute field can also include a content field of the dialogue data. Optionally, the content field of the dialogue data can be a question (input) or an answer (output) in the dialogue data. For example, options can be provided for the user to select, such as input content field and output content field.
[0059] The attribute field can further include a label field associated with the dialogue data. Optionally, the label field associated with the dialogue data can be set on demand in advance. For example, different dimensions of labels such as quality labels and the like can be set. For example, after each interaction of the intelligent agent or the large model, an evaluation feedback option (e.g., like or dislike) can be provided at the same time as the output answer, and feedback of the interactive object can be obtained, which is associated with the quality label, and the label is marked as good (like) or bad (dislike). Accordingly, the label field can provide the quality label, and the value of the quality label can be good or bad.
[0060] The second configuration item can be used to configure reference content, which is generally used to compare with the field value of the attribute field of the dialogue data. Optionally, the second configuration item can be provided as an input box for direct input by the user.
[0061] The third configuration item can be used to configure the comparison logic between the attribute field and the reference content. Optionally, a plurality of different options for representing the comparison logic can be provided for the user to select when configuring, and the comparison logic can be flexibly set on demand. For example, the comparison logic can include but is not limited to greater than or equal to, less than or equal to, greater than, less than, equal to, contains, does not contain, and the like.
[0062] For example, the first configuration item can be as shown in C1 of Figure 2 , the second configuration item can be as shown in C3 of Figure 2 , and the third configuration item can be as shown in C2 of Figure 3 .
[0063] Therefore, the user can configure the screening condition in the first configuration area according to the demand. In response to the configuration operation on the first trigger option, the first configuration item, the second configuration item, and the third configuration item, the screening condition corresponding to the target task can be determined.
[0064] Optionally, as shown on the right side of the first configuration area in Figure 2 , the first configuration area can further provide a component C4 for deleting the screening condition, and the screening condition corresponding to the component C4 can be deleted. In combination with the first trigger option and the component for deleting the screening condition, the screening condition can be flexibly added or deleted, so that the screening condition can be flexibly adjusted.
[0065] The first configuration item, the second configuration item, and the third configuration item are associated with each other and collectively constitute a screening rule. For example, the attribute field of the first configuration item can be configured as delay, the second configuration item can be configured as 1000 ms, and the third configuration item can be configured as greater than or equal to, so as to generate a screening rule that the delay duration is greater than or equal to 1000 ms. When the screening rule takes effect, if the corresponding delay duration of the dialogue data is greater than or equal to 1000 ms, it is considered that the dialogue data meets the screening rule. For another example, the attribute field of the first configuration item can be configured as input, the second configuration item can be configured as a string S, and the third configuration item can be configured as containing, so as to generate a screening rule that the input (i.e., the question) contains the string S. When the input in the dialogue data contains the string S, it is considered that the dialogue data meets the screening rule.
[0066] Correspondingly, the user can add multiple groups of the first configuration item, the second configuration item, and the third configuration item by repeatedly triggering the first trigger option, and perform configuration operations respectively to obtain a screening condition finally containing multiple screening rules. Based on this, when the dialogue data meets each screening rule in the screening condition, it is considered that the dialogue data meets the screening condition.
[0067] In another possible implementation, the second interface can further include a fourth configuration item for configuring a sampling rate. Correspondingly, the target task is used to sample the dialogue data meeting the screening condition according to the sampling rate for evaluation.
[0068] Based on this, the method provided by the present disclosure can further include the following steps:
[0069] In response to the configuration operation on the fourth configuration item, the sampling rate corresponding to the target task is determined.
[0070] Optionally, the fourth configuration item can be provided as an input box for the user to input a value representing the sampling rate.
[0071] Optionally, the fourth configuration item can be provided as an adjustment component, which can include a slider and a numerical value display area corresponding to the slider. For example, as shown in FIG. A2, the user can adjust the numerical value of the numerical value display area by sliding the slider, and the corresponding numerical value is the sampling rate. Figure 2
[0072] If the user has performed a configuration operation on the fourth configuration item and determined the sampling rate, when the target task is actually executed, the dialogue data meeting the screening condition will be sampled according to the sampling rate. For example, if the sampling rate is set to 50%, and there are 1000 dialogue data meeting the screening condition, 500 dialogue data (e.g., randomly sampled) will be sampled from the 1000 dialogue data, and then the 500 sampled dialogue data are used to evaluate the evaluation object.
[0073] In another possible implementation, the second interface can further include a fifth configuration item for configuring a number of evaluations, and accordingly, the target task is configured to stop when the number of evaluations of the evaluation object using the dialogue data reaches the number of evaluations.
[0074] Based on this, the method provided by the present disclosure can further include the following steps:
[0075] In response to a configuration operation on the fifth configuration item, the number of evaluations corresponding to the target task is determined.
[0076] Optionally, the fifth configuration item can be provided as an input box for a user to input a value representing the number of evaluations. For example, the fifth configuration item can be as shown in FIG. A3. Figure 2
[0077] For example, if the fifth configuration item sets the number of evaluations as 1000, when the evaluation object is evaluated using 1000 dialogue data, the number of evaluations reaches the number of evaluations 1000, and the target task can be ended, and no new dialogue data is obtained and the evaluation object is no longer evaluated.
[0078] In another possible implementation, the second interface can further include a second configuration area for configuring an application range of the screening rule, and the second configuration area displays a second trigger option for starting the configuration of the application range. Accordingly, the method provided by the present disclosure can further include the following steps:
[0079] In response to a trigger operation on the second trigger option, a sixth configuration item for configuring a time window of the target task is displayed in the second configuration area;
[0080] In response to a configuration operation on the sixth configuration item, the application range corresponding to the target task is determined, and the target task is configured to be executed within a time period indicated by the time window.
[0081] Optionally, the application range of the screening rule can be used to indicate a working time period of the target task and / or a time range in which the screening rule takes effect. The time window of the target task can be configured through the sixth configuration item to indicate the working time period of the target task.
[0082] For example, the second configuration area can be as shown in FIG. A4, the second trigger option can be as shown in FIG. B2, and by triggering the second trigger option to the starting state, the sixth configuration item B3 can be displayed in the second configuration area. For example, the sixth configuration item can provide a calendar option for a user to directly select a start time (year 1-month 1-day 1 hour 1: minute 1: second 1) and an end time (year 2-month 2-day 2 hour 2: minute 2: second 2). Figure 2 Figure 2
[0083] By the configuration operation on the sixth configuration item, the application range corresponding to the target task can be determined, so that the target task is executed within the time period indicated by the sixth configuration item, that is, the dialogue data generated in the actual interaction process of the evaluation object within the time period is obtained within the time period indicated by the sixth configuration item.
[0084] Optionally, the target task can end when the current time reaches the end time corresponding to the time window, or the number of times of evaluating the evaluation object by using the obtained dialogue data meeting the screening condition reaches the evaluation number.
[0085] In another possible implementation, in response to the triggering operation on the second trigger option, the second configuration area can further display a seventh configuration item for configuring the effective period of the target task. Accordingly, the method provided by the present disclosure can further include the following steps:
[0086] In response to the configuration operation on the seventh configuration item, the effective period corresponding to the target task is determined, and the target task is used to execute within the time period indicated by the time window and to determine the dialogue data meeting the screening condition within the effective period.
[0087] The effective period corresponding to the target task can be configured by the seventh configuration item to indicate the time point or time range at which the screening condition takes effect. Optionally, the seventh configuration item can be provided with a first component for selecting the effective date range and a second component for selecting the time period. For example, the first component can provide date selection options such as weekdays, holidays, Mondays, etc., and the second component can provide time setting functions for users to select a start time point (hour 3: minute 3: second 3) and an end time point (hour 4: minute 4: second 4).
[0088] By the configuration operation on the sixth configuration item and the seventh configuration item, the target task is determined to be executed within the time range indicated by the sixth configuration item, and the screening condition is applied to the dialogue data within the time range set by the seventh configuration item. For example, the seventh configuration item can be as shown in FIG. B4, where the left side is the first component of the previous example, and the right side is the second component in the previous example. Figure 2
[0089] In the above manner, when creating the target task, the target task can be flexibly configured and configured on demand in combination with the above-mentioned various configuration items. The target task configured with the above-mentioned configuration items will be executed according to the configuration content.
[0090] When creating the target task, in addition to setting the related settings for screening the dialogue data, it is also necessary to set from which evaluation dimension the dialogue data meeting the screening condition should be evaluated after being screened out.
[0091] Therefore, the method provided by the present disclosure can further include the following steps:
[0092] In response to receiving the second instruction for selecting the evaluator, a third interface is displayed, the third interface including a plurality of evaluators, and each of the evaluators corresponding to an evaluation dimension;
[0093] In response to a selection operation on the evaluator, at least one selected evaluator is determined, and the target task is configured to use the selected evaluator to evaluate the evaluation object from at least one evaluation dimension using the dialog data meeting the screening condition.
[0094] In the process of creating the target task, the second instruction for selecting the evaluator can be triggered. For example, the second instruction can be triggered after the screening condition is configured, or the second instruction can be triggered before the screening condition is configured, and the present disclosure does not make strict limitations in this regard.
[0095] Optionally, a configuration entry for configuring the evaluator can be provided, and the user can trigger the generation of the second instruction by triggering the configuration entry.
[0096] In the case of receiving the second instruction for selecting the evaluator, in response to the second instruction, a third interface for selecting the evaluator can be displayed. The third interface can provide a plurality of evaluators, each of which corresponds to an evaluation dimension. Optionally, the evaluators provided in the third interface can be already configured and can be directly used, and the built-in evaluation logic can be directly used for evaluation of the evaluation object.
[0097] For example, the interface for triggering the selection of the evaluator can be as shown in Figure 3 The button B5 for triggering the second instruction is provided, and by triggering the button B5, the third interface can be displayed, and the third interface can display a plurality of selectable evaluators, and the user can select one or more of them as the evaluator used by the target task.
[0098] For example, a plurality of evaluator types can be provided, and the evaluator types can include but are not limited to Prompt (prompt word) type, RAG (Retrieval-Augmented Generation, retrieval-augmented generation) type, Code (code) type, basic type, NLP (Natural Language Processing, natural language processing) type, artificial rule type, etc. Each evaluator type can include a plurality of evaluators, and each evaluator corresponds to an evaluation dimension.
[0099] The user can select the evaluator according to actual needs in the third interface, and in response to a selection operation on the evaluator, at least one selected evaluator of the user can be determined, and at least one evaluation dimension of the target task is determined.
[0100] After the user selects the evaluator, the selected evaluator can be displayed in the A5 area as shown in Figure 3 It can be seen that the current task has selected 4 evaluators.
[0101] In this way, when creating a target task, the evaluator can be flexibly selected, and the evaluation object can be evaluated according to the required evaluation dimension in the subsequent process.
[0102] Returning to Figure 1 In step 12, in the process of executing the target task, the dialog data corresponding to the actual interaction of the evaluation object is continuously obtained.
[0103] According to the start execution time indicated by the target task (that is, the start time in the time window described above), the target task can be executed, and in the process of executing the target task, the dialog data generated by the evaluation object in the actual interaction during the period needs to be continuously obtained. That is, each time the evaluation object performs an effective interaction and obtains a question and answer dialog data, it can be used as candidate dialog data, which is screened according to the screening condition to determine whether the current dialog data can be used for evaluation of the target task. Each time a dialog data is determined to meet the screening condition, the evaluation object can be evaluated using the dialog data.
[0104] Each time the evaluation object has an effective interaction, a link tracking data is generated, which is uniquely identified by TraceID. For example, the link details interface of the link tracking data can be as shown in Figure 4 Through the link tracking data, the relevant information of the dialog data can be obtained for comparison with the screening condition.
[0105] In step 13, in response to obtaining the target dialog data meeting the screening condition, the target evaluation result corresponding to the target dialog data is determined.
[0106] Each time the dialog data is obtained through step 12, it can be screened using the screening condition, and whether the dialog data meets the screening condition is determined according to the attribute field, reference content and comparison logic indicated by the screening condition. Further, the target dialog data meeting the screening condition can be determined, and each time a target dialog data is determined, the evaluation object can be evaluated using the target dialog data and the selected evaluator, and each selected evaluator is used to evaluate the evaluation object to obtain the target evaluation result corresponding to the target dialog data. Therefore, the target evaluation result includes the evaluation result corresponding to each evaluation dimension of the target dialog data.
[0107] In step 14, the target dialog data and the target evaluation result are displayed through the first interface.
[0108] After obtaining the target evaluation result, the target evaluation result can be displayed through the first interface, and the corresponding target dialogue data is also displayed.
[0109] In a possible implementation, step 14 can include the following steps:
[0110] In the first interface, the target dialogue data and the target evaluation result are displayed as a row in a table, wherein the target evaluation result includes sub-evaluation results corresponding to each evaluation dimension of the target task, and each sub-evaluation result is displayed in a column.
[0111] That is, the target dialogue data and the corresponding target evaluation result are displayed as a row in the table of the evaluation report, the questions and answers in the target dialogue data are displayed in two columns, and the target evaluation result is displayed in columns according to the evaluation dimensions, and each column is used to display the sub-evaluation result of the evaluation dimension corresponding to the column.
[0112] For example, the first interface can be as shown in Figure 5 wherein the original report is the evaluation report, and one row corresponds to one target dialogue data. The overview data is displayed at the top, which is used to display the evaluation dimensions and the comprehensive results corresponding to each evaluation dimension.
[0113] Through the above technical solution, the target task for evaluating the agent or the large model is created, and the real dialogue data of the agent or the large model is captured in real time during the execution of the target task and applied to the evaluation of the agent or the large model, and then the target evaluation result is displayed through the first interface. Therefore, the authenticity and real-time performance of the evaluation set can be ensured, and the cost of constructing the evaluation set can be greatly reduced.
[0114] Figure 6 is a block diagram of an agent or a large model evaluation device according to an embodiment of the present disclosure. As Figure 6 shown, the device 60 can include:
[0115] The task creation module 61 is configured to create a target task, wherein the target task is used to evaluate an evaluation object from at least one evaluation dimension by using dialogue data meeting a screening condition, the dialogue data is a question input to the evaluation object and an answer output by the evaluation object during interaction, and the evaluation object includes an agent or a large model.
[0116] The acquisition module 62 is configured to continuously acquire dialogue data corresponding to actual interaction of the evaluation object during execution of the target task.
[0117] The first determination module 63 is configured to determine a target evaluation result corresponding to the target dialogue data in response to acquisition of the target dialogue data meeting the screening condition.
[0118] The first display module 64 is configured to display the target dialogue data and the target evaluation result through the first interface.
[0119] Optionally, the device 60 further comprises:
[0120] The second display module is configured to display a second interface in response to receiving a first instruction for creating a target task for an evaluation object, the second interface comprising a first configuration area for configuring the screening condition, the first configuration area displaying a first trigger option for adding a screening condition;
[0121] The third display module is configured to display, in the first configuration area, a first configuration item for configuring a screened attribute field, a second configuration item for configuring reference content, and a third configuration item for configuring comparison logic between the attribute field and the reference content in response to a trigger operation on the first trigger option; wherein the attribute field comprises at least one of an attribute field in link tracking data associated with dialogue data, a content field of dialogue data, and a label field associated with dialogue data;
[0122] The second determination module is configured to determine the screening condition corresponding to the target task in response to a configuration operation on the first trigger option, the first configuration item, the second configuration item, and the third configuration item.
[0123] Optionally, the second interface further comprises a fourth configuration item for configuring a sampling rate, the target task being configured to sample dialogue data meeting the screening condition at the sampling rate for evaluation;
[0124] The device 60 further comprises:
[0125] The third determination module is configured to determine the sampling rate corresponding to the target task in response to a configuration operation on the fourth configuration item.
[0126] Optionally, the second interface further comprises a fifth configuration item for configuring an evaluation number, the target task being configured to stop when the number of times of evaluating the evaluation object using dialogue data reaches the evaluation number.
[0127] The device 60 further comprises:
[0128] The fourth determination module is configured to determine the evaluation number corresponding to the target task in response to a configuration operation on the fifth configuration item.
[0129] Optionally, the second interface further comprises a second configuration area for configuring an application range of the screening rule, the second configuration area displaying a second trigger option for starting to configure the application range.
[0130] The apparatus 60 further includes:
[0131] The fourth display module is configured to display, in response to a triggering operation on the second triggering option, a sixth configuration item for configuring a time window of the target task in the second configuration region.
[0132] The fifth determination module is configured to determine, in response to a configuration operation on the sixth configuration item, an application range corresponding to the target task, the target task being configured to be executed in a time period indicated by the time window.
[0133] Optionally, the second configuration region further displays a seventh configuration item for configuring an effective period of the target task.
[0134] The apparatus 60 further includes:
[0135] The sixth determination module is configured to determine, in response to a configuration operation on the seventh configuration item, an effective period corresponding to the target task, the target task being configured to be executed in a time period indicated by the time window and to determine the dialog data meeting the screening condition in the effective period.
[0136] Optionally, the apparatus 60 further includes:
[0137] The fifth display module is configured to display, in response to receiving a second instruction for selecting an evaluator, a third interface, the third interface including a plurality of evaluators, and each evaluator corresponding to an evaluation dimension.
[0138] In response to a selection operation on the evaluator, at least one selected evaluator is determined, and the target task is configured to evaluate the evaluation object from at least one evaluation dimension using the dialog data meeting the screening condition by using the selected evaluator.
[0139] Optionally, the first display module 64 includes:
[0140] The display sub-module is configured to display, in the first interface, the target dialog data and the target evaluation result as a row in a table, wherein the target evaluation result includes a sub-evaluation result corresponding to each evaluation dimension of the target task, and each sub-evaluation result is displayed by column.
[0141] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described in detail here.
[0142] Based on the same inventive concept, the present disclosure further provides a computer readable medium having stored thereon a computer program which, when executed by a processing apparatus, implements the steps of the agent or large model evaluation method described above.
[0143] Based on the same inventive concept, the present disclosure further provides an electronic device comprising:
[0144] a storage device having stored thereon a computer program;
[0145] a processing apparatus configured to execute the computer program in the storage device to implement the steps of the agent or large model evaluation method described above.
[0146] Based on the same inventive concept, the present disclosure further provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the agent or large model evaluation method described above.
[0147] Reference will now be made to the drawings, in which Figure 7 shows a structural diagram of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), car terminals (e.g., car navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0148] As shown in Figure 7 , the electronic device 600 can include a processing apparatus (e.g., a central processing unit, a graphics processing unit, etc.) 601 which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing apparatus 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0149] In general, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 608 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 609. The communication devices 609 can allow the electronic device 600 to communicate wirelessly or wired with other devices to exchange data. Although Figure 7 The electronic device 600 is shown with various devices, but it is understood that all of the illustrated devices are not required to implement or be present. More or less devices can alternatively be implemented or present.
[0150] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 609, or installed from the storage devices 608, or installed from the ROM 602. When the computer program is executed by the processing devices 601, the above-described functions defined in the methods of embodiments of the present disclosure are performed.
[0151] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable medium or a carrier wave transmits, propagates, or transfers a program used by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.
[0152] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0153] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device and not be assembled into the electronic device.
[0154] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: create a target task for evaluating an evaluation object from at least one evaluation dimension using dialogue data meeting a screening condition, the dialogue data being a question input to the evaluation object and an answer output by the evaluation object for the question in an interaction process of the evaluation object, the evaluation object including an agent or a large model; in the process of executing the target task, continuously acquire dialogue data corresponding to an actual interaction of the evaluation object; in response to acquiring target dialogue data meeting the screening condition, determine a target evaluation result corresponding to the target dialogue data; and display the target dialogue data and the target evaluation result through a first interface.
[0155] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages, including object oriented programming languages, such as Java, Smalltalk, C++, or conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0156] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the block can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0157] The modules described in the embodiments of the present disclosure can be implemented in the form of software or in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself, for example, the task creation module can also be described as a "module for creating a target task".
[0158] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0159] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0160] According to one or more embodiments of the present disclosure, a method for evaluating an agent or a large model is provided, the method comprising:
[0161] creating a target task, the target task being configured to evaluate an evaluation object from at least one evaluation dimension using dialog data meeting a filtering condition, the dialog data being a question input to the evaluation object and an answer output by the evaluation object in response to the question during an interaction between the evaluation object and the evaluation object, the evaluation object including an agent or a large model;
[0162] During execution of the target task, continuously obtaining dialog data corresponding to an actual interaction of the evaluation object;
[0163] In response to obtaining target dialog data meeting the filtering condition, determining a target evaluation result corresponding to the target dialog data;
[0164] Displaying the target dialog data and the target evaluation result through a first interface.
[0165] According to one or more embodiments of the present disclosure, a method for evaluating an agent or a large model is provided, and the method further comprises:
[0166] In response to receiving a first instruction for creating a target task for an evaluation object, a second interface is displayed, the second interface comprising a first configuration area for configuring the screening condition, the first configuration area displaying a first trigger option for adding a screening condition;
[0167] In response to a trigger operation on the first trigger option, a first configuration item for configuring a screened attribute field, a second configuration item for configuring reference content, and a third configuration item for configuring comparison logic between the attribute field and the reference content are displayed in the first configuration area; wherein the attribute field comprises at least one of an attribute field in link tracking data associated with dialogue data, a content field of dialogue data, and a label field associated with dialogue data;
[0168] In response to a configuration operation on the first trigger option, the first configuration item, the second configuration item, and the third configuration item, a screening condition corresponding to the target task is determined.
[0169] According to one or more embodiments of the present disclosure, a method for evaluating an agent or a large model is provided, and the second interface further comprises a fourth configuration item for configuring a sampling rate, and the target task is used to sample in dialogue data meeting the screening condition according to the sampling rate for evaluation;
[0170] The method further comprises:
[0171] In response to a configuration operation on the fourth configuration item, a sampling rate corresponding to the target task is determined.
[0172] According to one or more embodiments of the present disclosure, a method for evaluating an agent or a large model is provided, and the second interface further comprises a fifth configuration item for configuring an evaluation number, and the target task is used to stop when the number of times of evaluating the evaluation object with dialogue data reaches the evaluation number;
[0173] The method further comprises:
[0174] In response to a configuration operation on the fifth configuration item, an evaluation number corresponding to the target task is determined.
[0175] According to one or more embodiments of the present disclosure, a method for evaluating an agent or a large model is provided, and the second interface further comprises a second configuration area for configuring an application range of the screening rule, and the second configuration area displays a second trigger option for starting to configure the application range;
[0176] The method further includes:
[0177] In response to a triggering operation on the second trigger option, a sixth configuration item for configuring a time window of the target task is displayed in the second configuration area.
[0178] In response to a configuration operation on the sixth configuration item, an application range corresponding to the target task is determined, and the target task is used to execute in a time period indicated by the time window.
[0179] According to one or more embodiments of the present disclosure, a method for evaluating an agent or a large model is provided, and the second configuration area further displays a seventh configuration item for configuring an effective period of the target task.
[0180] The method further includes:
[0181] In response to a configuration operation on the seventh configuration item, an effective period corresponding to the target task is determined, and the target task is used to execute in a time period indicated by the time window and determine conversation data meeting a screening condition in the effective period.
[0182] According to one or more embodiments of the present disclosure, a method for evaluating an agent or a large model is provided, and the method further includes:
[0183] In response to receiving a second instruction for selecting an evaluator, a third interface is displayed, the third interface includes a plurality of evaluators, and each evaluator corresponds to an evaluation dimension;
[0184] In response to a selection operation on the evaluator, at least one selected evaluator is determined, and the target task is used to evaluate the evaluation object from at least one evaluation dimension using conversation data meeting a screening condition by using the selected evaluator.
[0185] According to one or more embodiments of the present disclosure, a method for evaluating an agent or a large model is provided, and the target conversation data and the target evaluation result are displayed through a first interface, including:
[0186] In the first interface, the target conversation data and the target evaluation result are displayed as a row in a table, wherein the target evaluation result includes sub-evaluation results corresponding to each evaluation dimension of the target task, and each sub-evaluation result is displayed by column.
[0187] According to one or more embodiments of the present disclosure, an evaluation device for an agent or a large model is provided, and the device includes:
[0188] The task creation module is configured to create a target task, the target task being configured to evaluate an evaluation object from at least one evaluation dimension by using dialogue data meeting a screening condition, the dialogue data being a question input by the evaluation object during interaction with the evaluation object and an answer output by the evaluation object to the question, and the evaluation object including an agent or a large model.
[0189] The acquisition module is configured to continuously acquire dialogue data corresponding to actual interaction of the evaluation object during execution of the target task.
[0190] The first determination module is configured to determine a target evaluation result corresponding to the target dialogue data in response to acquisition of the target dialogue data meeting the screening condition.
[0191] The first display module is configured to display the target dialogue data and the target evaluation result through a first interface.
[0192] According to one or more embodiments of the present disclosure, a computer readable medium is provided, and the computer readable medium has a computer program stored thereon, the computer program being executed by a processing device to implement steps of the evaluation method of the agent or the large model according to any of the embodiments of the present disclosure.
[0193] According to one or more embodiments of the present disclosure, an electronic device is provided, and the electronic device includes:
[0194] A storage device has a computer program stored thereon.
[0195] A processing device is configured to execute the computer program in the storage device to implement steps of the evaluation method of the agent or the large model according to any of the embodiments of the present disclosure.
[0196] According to one or more embodiments of the present disclosure, a computer program product is provided, and the computer program product includes a computer program, the computer program being executed by a processor to implement steps of the evaluation method of the agent or the large model according to any of the embodiments of the present disclosure.
[0197] The above description is merely preferred embodiments of the present disclosure and a description of principles of applied technologies. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the disclosed concept. For example, the above features can be replaced with technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0198] Moreover, while operations have been depicted in a particular order, this should not be understood as requiring such an order nor limiting it to only those operations shown and described. One of ordinary skill in the art will recognize that many of the operations can be performed in a differing order, or be performed concurrently, that some operations can be performed in any order or omitted, and that some operations can be performed in parallel. Similarly, while several specific implementation details have been discussed in the context of the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0199] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.
Claims
1. A method for evaluating intelligent agents or large models, characterized in that, The method includes: Create a target task, which is used to evaluate an evaluation object from at least one evaluation dimension using dialogue data that meets the screening criteria. The dialogue data consists of the question input to the evaluation object during interaction with the evaluation object and the answer output by the evaluation object to the question. The evaluation object includes an agent or a large model. During the execution of the target task, dialogue data corresponding to the actual interaction of the evaluation object is continuously acquired; In response to obtaining target dialogue data that meets the filtering conditions, a target evaluation result corresponding to the target dialogue data is determined; The target dialogue data and the target evaluation results are displayed on the first interface; The method further includes: In response to receiving a first instruction for creating a target task for the evaluation object, a second interface is displayed, the second interface including a first configuration area for configuring the filtering conditions, the first configuration area displaying a first trigger option for adding filtering conditions; In response to a trigger operation for the first trigger option, a first configuration item for configuring the filtered attribute field, a second configuration item for configuring the reference content, and a third configuration item for configuring the comparison logic between the attribute field and the reference content are displayed in the first configuration area; wherein, the attribute field includes at least one of the attribute field in the link tracing data associated with the dialogue data, the content field of the dialogue data, and the tag field associated with the dialogue data. In response to configuration operations for the first trigger option, the first configuration item, the second configuration item, and the third configuration item, the filtering conditions corresponding to the target task are determined.
2. The method according to claim 1, characterized in that, The second interface also includes a fourth configuration item for configuring the sampling rate, wherein the target task is used to sample dialogue data that meets the filtering conditions according to the sampling rate for evaluation; The method further includes: In response to the configuration operation for the fourth configuration item, the sampling rate corresponding to the target task is determined.
3. The method according to claim 1, characterized in that, The second interface also includes a fifth configuration item for configuring the number of evaluations, wherein the target task is used to stop when the number of times the evaluation object is evaluated using dialogue data reaches the number of evaluations; The method further includes: In response to the configuration operation for the fifth configuration item, the number of evaluations corresponding to the target task is determined.
4. The method according to claim 1, characterized in that, The second interface also includes a second configuration area for configuring the application scope of the filtering conditions, and the second configuration area displays a second trigger option for enabling the configuration of the application scope; The method further includes: In response to a trigger operation for the second trigger option, a sixth configuration item for configuring the time window of the target task is displayed in the second configuration area; In response to the configuration operation for the sixth configuration item, the application scope corresponding to the target task is determined, and the target task is to be executed within the time period indicated by the time window.
5. The method according to claim 4, characterized in that, The second configuration area also displays a seventh configuration item for configuring the effective period of the target task; The method further includes: In response to the configuration operation for the seventh configuration item, the effective period corresponding to the target task is determined. The target task is used to execute within the period indicated by the time window and to determine the dialogue data that meets the filtering conditions within the effective period.
6. The method according to claim 1, characterized in that, The method further includes: In response to receiving a second instruction for selecting an evaluator, a third interface is displayed, the third interface including multiple evaluators, and each evaluator corresponds to an evaluation dimension; In response to the selection operation for the evaluator, at least one selected evaluator is determined, and the target task is to use the selected evaluator to evaluate the evaluation object from at least one evaluation dimension using dialogue data that meets the screening criteria.
7. The method according to claim 1, characterized in that, The step of displaying the target dialogue data and the target evaluation results through the first interface includes: In the first interface, the target dialogue data and the target evaluation results are displayed as a row in a table. The target evaluation results include sub-evaluation results corresponding to each evaluation dimension of the target task, and each sub-evaluation result is displayed in a column.
8. An evaluation device for intelligent agents or large models, characterized in that, The device includes: The task creation module is used to create a target task, which is used to evaluate an evaluation object from at least one evaluation dimension using dialogue data that meets the screening criteria. The dialogue data consists of the question input by the evaluation object during the interaction with the evaluation object and the answer output by the evaluation object for the question. The evaluation object includes an intelligent agent or a large model. The acquisition module is used to continuously acquire the dialogue data of the evaluation object during the execution of the target task; The first determining module is used to determine the target evaluation result corresponding to the target dialogue data in response to obtaining target dialogue data that meets the filtering conditions; The first display module is used to display the target dialogue data and the target evaluation results through a first interface; The device further includes: The second display module is used to display a second interface in response to receiving a first instruction for creating a target task for the evaluation object. The second interface includes a first configuration area for configuring the filtering conditions. The first configuration area displays a first trigger option for adding filtering conditions. The third display module is used to respond to a trigger operation for the first trigger option by displaying a first configuration item for configuring the filtered attribute field, a second configuration item for configuring the reference content, and a third configuration item for configuring the comparison logic between the attribute field and the reference content in the first configuration area; wherein, the attribute field includes at least one of the attribute field in the link tracing data associated with the dialogue data, the content field of the dialogue data, and the tag field associated with the dialogue data. The second determining module is used to determine the filtering conditions corresponding to the target task in response to the configuration operation for the first triggering option, the first configuration item, the second configuration item and the third configuration item.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program performs the steps of the method according to any one of claims 1-7.
10. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Intelligent agent evaluation method and device based on multiple rounds of dialogues, electronic equipment and medium
CN119829964A