Method, apparatus, electronic device and computer program product for evaluating a data analysis agent built based on a large model

By splitting the agent processing process into multiple subtasks and automatically evaluating each subtask, a comprehensive evaluation is generated, and the problem of incomplete evaluation in the prior art is solved, and a more accurate and efficient agent performance evaluation is achieved.

CN118708921BActive Publication Date: 2025-09-02BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410831618.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2025-09-02
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

The existing agent evaluation methods cannot comprehensively evaluate their reasoning process, resulting in inaccurate and comprehensive evaluation results, and are prone to false positive problems.

Method used

The process of the agent is split into multiple subtasks, and the output of each subtask is automatically evaluated, and a comprehensive evaluation is generated by summarizing the evaluation results of each subtask, including knowledge recall, code writing and conclusion generation.

Benefits of technology

A more comprehensive and accurate agent evaluation is achieved, which can specifically measure its performance in each link, identify advantages and disadvantages, improve overall performance, and save evaluation resources and shorten evaluation cycles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118708921B_ABST
    Figure CN118708921B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to methods, devices, electronic devices, and computer program products for evaluation. The method includes generating multiple outputs corresponding to multiple subtasks by an intelligent agent performing multiple subtasks according to user input. The method also includes determining a sub-evaluation of each of the multiple outputs based on the multiple outputs. In addition, the method also includes determining a comprehensive evaluation of the intelligent agent based on the sub-evaluation of each of the multiple outputs. Through the evaluation method of the embodiment of the present disclosure, not only can more comprehensive and accurate evaluation results be generated, but also the performance of the intelligent agent in each link can be specifically measured. This refined evaluation method helps to identify the advantages and disadvantages of the intelligent agent in different links, so as to carry out targeted optimization and adjustment and improve the overall performance of the intelligent agent. In addition, automatic evaluation of the output of each subtask can significantly save the resources required for evaluation and shorten the evaluation cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more particularly to a method, device, electronic device, and computer program product for evaluation. Background Art

[0002] An agent is an intelligent entity capable of perceiving its environment, making decisions, and executing actions. Equipped with memory, logical analysis, task decomposition, and problem-solving capabilities, agents are capable of comprehensively processing and solving complex problems. As a supporting technology for large-scale models, agents will play an increasingly important role in various fields, including data analysis, intelligent customer service, and intelligent assistants.

[0003] With the rapid development of large models, more and more research and applications are focusing on the integration of large models with big data. Data analysis agents based on large models have become an important tool in the field of big data analysis. Through automated data processing and analysis, they greatly improve the efficiency and accuracy of data analysis and have received widespread attention and research in academia and industry. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device, computer program product, and medium for evaluation.

[0005] According to a first aspect of the present disclosure, an evaluation method is provided. The method includes generating, by an agent, a plurality of subtasks according to user input, a plurality of outputs corresponding to the plurality of subtasks. The method also includes determining, based on the plurality of outputs, a sub-evaluation for each of the plurality of outputs. Furthermore, the method also includes determining a comprehensive evaluation of the agent based on the sub-evaluation for each of the plurality of outputs.

[0006] According to a second aspect of the present disclosure, an evaluation device is provided. The device includes an output generation module configured to generate multiple outputs corresponding to multiple subtasks by having an agent perform multiple subtasks based on user input. The device also includes a sub-evaluation determination module configured to determine a sub-evaluation for each of the multiple outputs based on the multiple outputs. In addition, the device also includes a comprehensive evaluation determination module configured to determine a comprehensive evaluation for the agent based on the sub-evaluation of each of the multiple outputs.

[0007] According to a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory coupled to the processor, wherein the memory has instructions stored therein, and when the instructions are executed by the processor, the electronic device executes the method according to the first aspect.

[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions that, when executed, cause a computer to perform the steps of the method of the first aspect of the present disclosure.

[0009] In a fifth aspect of the present disclosure, a computer-readable storage medium is provided, wherein one or more computer instructions are stored on the computer-readable storage medium, wherein the one or more computer instructions are executed by a processor to implement the method according to the first aspect.

[0010] This summary is intended to introduce a selection of concepts in a simplified form that are further described below in the detailed description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0012] Figure 1 A schematic diagram illustrating an example environment in which devices and / or methods according to embodiments of the present disclosure may be implemented;

[0013] Figure 2 A flowchart showing a process of an evaluation method according to an embodiment of the present disclosure;

[0014] Figure 3 A flowchart showing a process of a data analysis agent performing a data analysis task according to an embodiment of the present disclosure is shown;

[0015] Figure 4 A flowchart of a process for performing automatic evaluation of knowledge recall according to an embodiment of the present disclosure is shown;

[0016] Figure 5 A flowchart of a process for performing automatic evaluation of analytical reasoning according to an embodiment of the present disclosure is shown;

[0017] Figure 6A A flowchart of a process for performing automatic code writing evaluation according to an embodiment of the present disclosure is shown;

[0018] Figure 6B-6D A schematic diagram illustrating a maximum common match of tabular data according to an embodiment of the present disclosure is shown;

[0019] Figure 7A flowchart of a process for performing automatic evaluation of conclusion generation according to an embodiment of the present disclosure is shown;

[0020] Figure 8 A flowchart illustrating a process of performing task completion aggregation according to an embodiment of the present disclosure is shown;

[0021] Figure 9 A block diagram illustrating an evaluation device according to an embodiment of the present disclosure; and

[0022] Figure 10 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0023] Throughout the drawings, the same or similar reference numbers denote the same or similar elements. DETAILED DESCRIPTION

[0024] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information (such as voice) involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0025] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0026] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or the same objects, unless explicitly stated otherwise. Other explicit and implicit definitions may also be included below.

[0027] As mentioned earlier, intelligent agents are widely valued and applied in various fields. Therefore, quickly and accurately evaluating the capabilities and effectiveness of intelligent agents has become a research priority. Currently, relevant evaluation methods often only assess the final execution results, which is prone to false positives and fails to fully evaluate the agent's reasoning process, resulting in incomplete and inaccurate evaluation results.

[0028] In order to solve the above problems, an embodiment of the present disclosure provides an evaluation scheme, which splits the processing process of the intelligent agent into a workflow of multiple subtasks, automatically evaluates the output of each subtask, and generates a corresponding sub-evaluation. By summarizing the sub-evaluation results of each subtask, a comprehensive evaluation of the intelligent agent is obtained. Through this comprehensive evaluation method, not only can more comprehensive and accurate evaluation results be generated, but also the performance of the intelligent agent in each link can be specifically measured. This detailed evaluation method helps to identify the advantages and disadvantages of the intelligent agent in different links, so as to carry out targeted optimization and adjustment and improve the overall performance of the intelligent agent. In addition, automatic evaluation of the output of each subtask can significantly save the resources required for evaluation and shorten the evaluation cycle.

[0029] Figure 1 Schematic diagram of an example environment in which devices and / or methods according to embodiments of the present disclosure may be implemented. Figure 1 As shown, example environment 100 may include a computing device 110, which may be a user terminal, mobile device, computer, etc., or a computing system, a single server, a distributed server, or a cloud-based server. Computing device 110 may receive user input 120. For example, user input 120 may be a user-entered data query and analysis requirement, including text, voice, images, videos, data tables, etc. For example, a text-based user input may be "What are the three cities with the highest temperatures in the region this week and what are their respective temperatures?"

[0030] The agent 130 can process the user input 120 and process it by splitting the processing process into multiple subtasks. For example, the subtask 140 can be a knowledge recall task. The agent 130 can search the database and knowledge base based on the user input 120 to generate a corresponding subtask output 142, which can include data tables, indicators, dimensions, dimension values, domain knowledge, etc. For example, for the user input 120, it is necessary to obtain the temperature of the city from the data table, and it is necessary to obtain some indicators from each data table, etc. Then, the computing device 100 can compare the subtask output with the benchmark output to generate a sub-evaluation 144 to measure the effectiveness of the agent 130 in performing knowledge recall.

[0031] Similarly, subtask 150 may be a code writing task that may generate subtask output 152 including a data table, and computing device 100 may generate a sub-evaluation 154 based on subtask output 152. For example, sub-output 152 may be SQL code and Python code for obtaining the city's temperature this week from a data table, and sub-evaluation 154 may be a metric measuring the correctness of the code. Similarly, subtask 160 may be a conclusion generation task that may generate subtask output 162 for answering the final conclusion of user input 120, and computing device 100 may generate a sub-evaluation 164 based on subtask output 162. It should be understood that the three subtasks are merely examples, and the intelligent agent in embodiments of the present disclosure may include fewer or more subtasks.

[0032] Continue to refer Figure 1 Computing device 110 may aggregate sub-evaluations 144, 154, and 164 to generate a comprehensive evaluation 170 for agent 130. As previously described, sub-evaluation 144 evaluates agent 130's knowledge recall ability, sub-evaluation 154 evaluates agent 130's code writing ability, and sub-evaluation 164 evaluates agent 130's conclusion generation ability. Comprehensive evaluation 170 allows for a comprehensive and accurate assessment of the agent's overall performance and the effectiveness of each subtask.

[0033] It should be understood that the architecture and functions in the example environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure. The embodiments of the present disclosure may also be applied to other environments with different structures and / or functions.

[0034] The following will be combined Figures 2 to 10 The process according to the embodiment of the present disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0035] Figure 2 FIG2 shows a flow chart of an evaluation method according to an embodiment of the present disclosure. At block 202, an agent may perform multiple subtasks according to user input to generate multiple outputs corresponding to the multiple subtasks. For example, referring to FIG2 Figure 1 , computing device 110 can execute subtask 140, subtask 150, and subtask 160 according to user input through agent 130, and generate subtask output 142, subtask output 152, and subtask output 162, respectively. It should be understood that the three subtasks are only used as examples, and the agent in the embodiments of the present disclosure may include fewer or more subtasks.

[0036] At block 204, a sub-assessment for each of the plurality of outputs may be determined based on the plurality of outputs. Figure 1 , computing device 110 may determine corresponding sub-assessments 144 , 154 , and 164 based on sub-task output 142 , sub-task output 152 , and sub-task output 162 .

[0037] At block 206, a comprehensive evaluation for the agent may be determined based on the sub-evaluations for each of the plurality of outputs. Figure 1 The computing device 110 may determine a comprehensive assessment 170 for the agent 130 based on the sub-assessment 144 , the sub-assessment 154 , and the sub-assessment 164 .

[0038] Thus, method 200 according to the embodiment of the present disclosure not only generates more comprehensive and accurate evaluation results, but also specifically measures the agent's performance at each stage. This refined evaluation approach helps identify the strengths and weaknesses of the agent in different stages, enabling targeted optimization and adjustments to improve the agent's overall performance. Furthermore, automatically evaluating the output of each subtask can significantly reduce evaluation resources and shorten the evaluation cycle.

[0039] Figure 3 The flowchart of the process 300 of the data analysis agent performing the data analysis task according to the embodiment of the present disclosure is shown. The data analysis agent is the agent that the user uses to perform the data analysis task. For the problem to be solved, the problem-solving steps of the data analysis agent can be decomposed. For example, Figure 3 In the example, the execution process of the data analysis agent is broken down into subtasks such as knowledge recall, analytical reasoning, code writing, sandbox execution, and conclusion generation. In other embodiments of the present disclosure, different execution subtasks and / or workflows will be broken down according to specific problems. Figure 3 As shown, at box 302, the data analysis agent can obtain data demands. For example, in the problem-solving process of the data analysis agent, data demands are the core requirements input by the user, and the data analysis agent needs to perform subsequent knowledge recall, analytical reasoning, code writing, sandbox execution and conclusion generation steps based on this requirement. Data demands can be expressed in a variety of forms, including but not limited to text, voice, pictures, videos, data tables and other forms. In some embodiments, users can describe the problem to be solved or the information to be obtained in natural language. For example, in an e-commerce data analysis scenario, the user may enter the following data demand: "The three cities with the highest temperature in the region this week and what are their respective temperatures." It should be understood that the user can also enter the data demand in other forms, such as by voice.

[0040] At box 304, the data analysis agent can perform knowledge recall based on the data demands. For example, when performing knowledge recall, the data analysis agent can retrieve relevant data tables, indicators, dimensions and business knowledge from the database and knowledge base based on the user's data demands to support subsequent analytical reasoning and code writing. In some embodiments, the database and knowledge base are pre-configured. For example, the data table can be a table of a database such as Hive, Clickhouse, Mysql, etc. The knowledge base can be knowledge mined from the business and stored in the vector library in the form of vectors. In addition, dimensions and indicators can be fields in the data table. Dimensions can be attributes or characteristics of data used to classify, group or slice data, such as geographic location, time, etc. Indicators can be specific numerical values ​​used to measure and evaluate dimensions, such as visits, conversion rates, etc.

[0041] At block 306, the data analysis agent may perform analytical reasoning. Analytical reasoning is the process by which the data analysis agent proposes analytical ideas and specific steps to solve a problem based on the user's data demands and the recalled knowledge. For example, for the aforementioned data demands example, the corresponding analytical reasoning results may be:

[0042] (1) First, we need to determine the time range of this week, that is, from the first day of the week to today.

[0043] (2) Then, we need to filter out this week’s data from the [Weather Data Warehouse], filter out the temperature data for each city, group them by city ID and city name, and calculate the maximum temperature for each city.

[0044] (3) Finally, we sort the cities in descending order by maximum temperature and take out the top three cities with the highest temperatures and their temperatures.

[0045] At block 308, the data analysis agent may execute code writing. For example, when executing code writing, the data analysis agent may convert the analytical reasoning steps into executable code. The coding language may include, but is not limited to, SQL and Python. The specific language used depends on the data source and analysis requirements. The embodiments of the present disclosure do not limit the type of coding language. In some embodiments, the data analysis agent may write corresponding Python program code and perform Hive queries using SQL.

[0046] At block 310, the code written by the data analysis agent can be executed in a sandbox environment. Sandbox execution involves running the code generated by the data analysis agent in a controlled and isolated computing environment that simulates a production environment without impacting actual production data or systems. For example, through sandbox execution, the data analysis agent can obtain intermediate and final results of the code execution.

[0047] At block 308, the data analysis agent may perform conclusion generation. For example, when performing conclusion generation, the data analysis agent may present the results of the data analysis to the user in a clear, accurate, and easy-to-understand manner. In embodiments of the present disclosure, the generated conclusions may be in various forms, such as text, tables, charts, voice, or video. For example, for the aforementioned data appeal example, the corresponding generated conclusions may be as shown in Table 1:

[0048] Table 1 Conclusion generation

[0049] City ID City Name temperature 1003 A1 City 43.1 1001 A2 City 42.8 1002 A3 City 42.5

[0050] Furthermore, the generated conclusions can be presented in a variety of formats to meet the needs of different users and application scenarios. In addition to tabular form, the conclusions can also be presented in a variety of formats, such as text, charts, voice, and video, which are not limited in the embodiments of the present disclosure. In some embodiments, the user can specify the format of the generated conclusions.

[0051] Figure 4 A flowchart of a process 400 for performing automatic evaluation of knowledge recall according to an embodiment of the present disclosure is shown. At box 402, the recall results of the knowledge recall and the corresponding benchmark recall results can be obtained. For example, it is first necessary to obtain the recall knowledge obtained by the data analysis agent when processing user data demands. The recall knowledge may include the table names in the database, relevant indicators and dimensions, application filtering conditions, and domain knowledge, etc. In addition, it is also necessary to obtain the benchmark recall results as a comparison object, and the benchmark recall results contain all correct knowledge recall items. At box 404, the data table in the recall result can be compared with the data table in the benchmark recall result to evaluate the correctness of the data table selection. For example, this step can evaluate whether the data analysis agent has correctly selected the data table for analysis. In some embodiments, the table name recalled by the data analysis agent can be compared with the table name in the benchmark recall result to check whether they are consistent. If they are consistent, the data table selection is correct, otherwise the data table selection is incorrect.

[0052] At box 406, the indicators, dimensions, filtering conditions and domain knowledge in the recall result can be compared with the baseline recall result to evaluate the grounding correctness. In some embodiments, the indicators, dimensions, filtering conditions and domain knowledge in the baseline recall result can be traversed to determine whether they are included in the recall result. At box 408, the evaluation result of the knowledge recall can be determined based on the table selection correctness and the calibration correctness. In some embodiments, the evaluation results of the first two steps can be combined to determine the evaluation result of the knowledge recall. For example, the knowledge recall can be scored according to the following example rules: if the table selection is correct and the indicators, dimensions, filtering conditions and domain knowledge are all recalled, then a score of 1 is given; if the table selection is wrong, or any of the indicators, dimensions, filtering conditions and domain knowledge are not recalled, then a score of 0 is given. It should be understood that this scoring rule is only an example, and the embodiments of the present disclosure can also use other rules to score knowledge recall.

[0053] Figure 5 A flowchart of a process 500 for performing automatic evaluation of analytical reasoning according to an embodiment of the present disclosure is shown. At box 502, the reasoning results of the analytical reasoning and the benchmark reasoning results can be obtained. For example, first, the reasoning results generated by the data analysis agent performing analytical reasoning can be obtained, which include the analysis ideas and specific steps of the data analysis agent for the user's data demands. In addition, a benchmark reasoning process can be obtained as a comparison object, and the benchmark reasoning process should include the correct steps to solve the user's data demands. At box 504, the prompt content (prompt) for evaluation can be determined based on the reasoning results and the benchmark reasoning results. In some embodiments, the example prompt can be:

[0054] Benchmark data analysis and reasoning process:

[0055] ```

[0056] First, we need to determine the time range for this week, which is from the first day of the week to today.

[0057] Then, we need to filter out this week's data from the [Meteorological Data Warehouse], filter out the temperature data of each city, group them by city ID and city name, and calculate the maximum temperature of each city.

[0058] Finally, we sort by the highest temperature in descending order and take out the top 3 cities with the highest temperatures

[0059] and its temperature.

[0060] ```

[0061] Data analysis and reasoning process of agent prediction:

[0062] ```

[0063] First, we need to determine the time range for this week, which is from the first day of the week to today.

[0064] Next, we need to filter this week's data from the Weather Data Warehouse, filter out the temperature data for each city, group them by city name, and calculate the maximum temperature for each city.

[0065] Finally, we sort by the highest temperature in descending order and take out the top 3 cities with the highest temperatures

[0066] and its temperature.

[0067] ```

[0068] If the data analysis and reasoning process predicted by the agent meets the benchmark data analysis and reasoning process

[0069] 1 point if successful, 0 points if not

[0070] Output json format example:

[0071] ```

[0072] {"score":1}

[0073] ```

[0074] The prompt in the reference example includes the baseline inference result, the agent's predicted inference result, the evaluation criteria, and the format of the output data. In some embodiments, additional evaluation criteria and output data formats can be defined. In addition, the prompt can also include other content, which is not limited by this disclosure.

[0075] At box 506, the evaluation result of the analytical reasoning can be determined based on the prompt using the analytical reasoning evaluation model. For example, the analytical reasoning evaluation model can be a language model that has been fine-tuned in a supervised manner on a pre-trained language model. For example, fine-tuning using labeled training data on an existing pre-trained model can enable the analytical reasoning evaluation model to better perform the evaluation task. In some embodiments, the format of the evaluation result is the format specified in the prompt. For example, if it is specified in the prompt that "if the data analysis and reasoning process predicted by the intelligent agent meets the benchmark data analysis and reasoning process, score 1, otherwise score 0", then the evaluation result is 1 or 0. In some embodiments, the evaluation result can be between 0 and 1 to indicate the degree of compliance of the reasoning result with the benchmark reasoning result.

[0076] Figure 6A FIG. 6 is a flowchart showing a process 600 for performing automatic code writing evaluation according to an embodiment of the present disclosure. Figure 6B-6DA schematic diagram of the maximum common matching of table data according to an embodiment of the present disclosure is shown. Figures 6A-6D To describe the process of writing evaluation code in the embodiment of the present disclosure. Figure 6A As shown, at block 602, the correctness of the SQL code can be determined using an abstract syntax tree. For example, the SQL code can be parsed into an abstract syntax tree, which can reflect the grammatical structure and logical relationships of the SQL code. The abstract syntax tree can then be used to calculate the similarity between the SQL code generated by the data analysis agent and the benchmark SQL code, which serves as a metric for measuring the correctness of the SQL code.

[0077] At box 604, the correctness of the Python code can be determined based on the code execution result. For example, the Python code generated by the data analysis agent and the benchmark Python code can be executed simultaneously, and the execution results (for example, Python tables) can be obtained from the sandbox to determine the correctness of the Python code. In some embodiments, the code for data query, data processing, and data analysis can be regarded as the state flow of the code execution result through a table fuzzy matching algorithm, thereby measuring the intermediate execution results of the code to measure the correctness of the Python code. For example, the intermediate execution result can be the value of the program variable after each step of the code execution. In some embodiments, the execution result is tabular data, and the similarity of the data can be calculated by finding the maximum common matching part between the tables and scoring the intersection over union (IoU). For example, assuming that the tabular data generated by the Python generated by the data analysis agent is Table 2 and the benchmark tabular data is Table 3, the maximum common matching part of Table 2 and Table 3 can be as follows Figure 6B-6D shown.

[0078] Table 2 Tabular data generated by the agent

[0079] 1 5 11 1 1 6 11 1 3 8 11 0 4 9 11 9 11 11 12 12

[0080] Table 3 Benchmark table data

[0081] 1 5 9 1 1 6 10 1 3 7 11 1 4 9 12 1

[0082] like Figure 6B As shown, the rectangular boxes show the matching parts of Table 2 and Table 3, and the dotted lines show the corresponding matching parts. Figure 6B-6D Three different maximum common matching parts of Table 2 and Table 3 are shown. Then, the similarity between Table 2 and Table 3 can be calculated by IoU, and the calculation formula is:

[0083] IoU = Overlap Area / Union Area (1)

[0084] The overlapping area is the area of ​​the largest common matching part, such as Figure 6B-6DAlthough Tables 2 and 3 have three different maximum common matching parts, the overlapping area is 6. The union area is the union of the elements in Tables 2 and 3. Therefore, by finding the maximum common matching part of the tabular data and calculating the IoU similarity, the correctness of the Python code generated by the data analysis agent can be effectively evaluated. It also provides a quantitative similarity score, making the evaluation results more objective and accurate.

[0085] In some embodiments, the timestamps before and after the sandbox execution code is called can be recorded separately, and the difference between the two timestamps can be recorded as the sandbox time. Sandbox time can be used to measure the efficiency of code execution. By monitoring and optimizing the sandbox time, performance problems in the code can be discovered and solved. In some embodiments, the correctness of the tool use of the written code can be evaluated to evaluate the correctness of tool selection and parameter passing. In some embodiments, the self-debugging (Self-Debug) repair of the data analysis agent can be evaluated. For example, the number of Self-Debug times can be evaluated to evaluate the agent's ability to repair code.

[0086] Figure 7 A flowchart of a process 700 for performing automatic evaluation of conclusion generation according to an embodiment of the present disclosure is shown. At block 702, the generated conclusion and the baseline conclusion of the agent can be obtained. For example, the generated conclusion obtained by the data analysis agent performing the conclusion generation subtask can be obtained first. In addition, the baseline conclusion can be obtained as a comparison object, and the baseline conclusion includes the correct analysis results and conclusions. At block 704, a prompt for evaluation can be determined based on the generated conclusion and the baseline conclusion. In some embodiments, the example prompt can be:

[0087] Benchmark data analysis conclusions:

[0088] ```

[0089] The city with the highest temperature in the region this week is City X, with city ID 12345 and a temperature of 43.1 degrees.

[0090] ```

[0091] Conclusions from the data analysis generated by the agent:

[0092] ```

[0093] The city with the highest temperature in the region this week is City X, with city ID 12345 and a temperature of 43.1 degrees.

[0094] ```

[0095] If the data analysis conclusion generated by the agent is consistent with the data analysis conclusion of the benchmark, a score of 1 is given; otherwise, a score of 0 is given.

[0096] Output json format example:

[0097] ```

[0098] {"score":1}

[0099] ```

[0100] The prompt in the reference example includes the baseline data analysis conclusion, the agent-generated data analysis conclusion, the evaluation criteria, and the format of the output data. In some embodiments, additional evaluation criteria and output data formats can be defined. In addition, the prompt can also include other content, which is not limited by this disclosure.

[0101] At box 706, the evaluation result of the conclusion generation can be determined based on the prompt using the conclusion evaluation model. For example, the conclusion evaluation model can be a language model that has been fine-tuned in a supervised manner on a pre-trained language model. For example, fine-tuning using labeled training data on an existing pre-trained model can enable the conclusion evaluation model to better perform the evaluation task. In some embodiments, the format of the evaluation result is the format specified in the prompt. For example, if it is specified in the prompt that "if the data analysis conclusion generated by the agent meets the benchmark data analysis conclusion, score 1, otherwise score 0", then the evaluation result is 1 or 0. In some embodiments, the evaluation result can be between 0 and 1 to indicate the degree of compliance of the generated conclusion with the benchmark conclusion.

[0102] Figure 8 FIG. 8 is a flow chart illustrating a process 800 for performing task completion aggregation according to an embodiment of the present disclosure. At block 802, a knowledge recall automatic evaluation algorithm may be executed. As previously described, Figure 4 The process 400 shown is an example process of the automatic evaluation algorithm for knowledge recall, which can evaluate the knowledge recall subtask by selecting the correctness of the table and the correctness of the calibration. At block 804, the automatic evaluation algorithm for analytical reasoning can be executed. As previously described, Figure 5 The process 500 shown is an example process of the analytical reasoning automatic evaluation algorithm, which can use the analytical reasoning evaluation model to evaluate the analytical reasoning subtask. At box 806, the code correctness automatic evaluation algorithm can be executed. As mentioned above, the process 600 shown in Figure 6 is an example process of the code writing automatic evaluation algorithm, which can evaluate the code writing subtask by the correctness of the SQL code and the correctness of the Python code. At box 808, the conclusion generation automatic evaluation algorithm can be executed. As mentioned above, Figure 7 The process 700 shown is an example process of an automatic evaluation algorithm for conclusion generation, and a conclusion evaluation model can be used to evaluate the conclusion generation subtask.

[0103] At block 810, a task completion summary may be performed. By integrating the evaluation results of each subtask, the overall performance of the data analysis agent in the task is determined. The task completion summary may also be referred to as a comprehensive evaluation of the data analysis agent. For example, an example of a comprehensive evaluation may be: comprehensive evaluation = table selection correctness + calibration correctness + analysis and reasoning correctness + SQL correctness + Python correctness + conclusion correctness. In some embodiments, the comprehensive evaluation result may be determined based on predetermined rules. For example, the example rules may be as shown in Table 4:

[0104] Table 4 Example rules for task completion summary

[0105]

[0106] It should be understood that embodiments of the present disclosure may also utilize other rules to perform task completion aggregation. Thus, by performing task completion aggregation, each subtask can be automatically evaluated, generating a more comprehensive, integrated assessment. Furthermore, targeted optimization and adjustments can be performed based on the evaluation of each subtask.

[0107] Figure 9 1 shows a block diagram of an evaluation device 900 according to some embodiments of the present disclosure. Figure 9 As shown, apparatus 900 includes an output generation module 902 configured to generate multiple outputs corresponding to the multiple subtasks by having an agent perform multiple subtasks based on user input. Apparatus 900 also includes a sub-evaluation determination module 904 configured to determine a sub-evaluation for each of the multiple outputs based on the multiple outputs. Furthermore, apparatus 900 also includes a comprehensive evaluation determination module 906 configured to determine a comprehensive evaluation for the agent based on the sub-evaluation for each of the multiple outputs.

[0108] Figure 10 A block diagram of an electronic device 1000 is shown, in accordance with certain embodiments of the present disclosure. Figure 10 FIG1 shows a block diagram of an electronic device 1000 according to some embodiments of the present disclosure. The device 1000 may be a device or apparatus described in an embodiment of the present disclosure. Figure 10As shown, the device 1000 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 1001, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 1002 or computer program instructions loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The CPU / GPU 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004. Although not shown in FIG. Figure 10 As shown in FIG, device 1000 may further include a co-processor.

[0109] Various components in device 1000 are connected to I / O interface 1005, including: an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0110] The various methods or processes described above may be performed by the CPU / GPU 1001. For example, in some embodiments, the methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the CPU / GPU 1001, one or more steps or actions in the methods or processes described above may be performed.

[0111] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.

[0112] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.

[0113] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0114] The computer program instructions for performing the disclosed operation can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or the object code written in any combination of one or more programming languages, programming languages ​​include object-oriented programming languages, and conventional procedural programming languages.Computer-readable program instructions can be performed completely on a user's computer, partially on a user's computer, performed as an independent software package, partly on a user's computer and partly on a remote computer, or performed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network-including local area network (LAN) or wide area network (WAN), or can be connected to an external computer (such as utilizing an internet service provider to connect by the internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to carry out personalized customization electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLA), this electronic circuit can perform computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0115] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0116] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0117] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented by a special hardware-based system that performs the prescribed function or action, or can be implemented by a combination of special hardware and computer instructions.

[0118] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements to existing technologies, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

[0119] Some example implementations of the present disclosure are listed below.

[0120] Example 1. An evaluation method comprising:

[0121] The agent performs a plurality of subtasks according to user input to generate a plurality of outputs corresponding to the plurality of subtasks;

[0122] determining a sub-assessment for each of the plurality of outputs based on the plurality of outputs; and

[0123] Based on the sub-evaluations of each of the plurality of outputs, a composite evaluation for the agent is determined.

[0124] Example 2. The method of example 1, wherein the plurality of subtasks comprises a knowledge recall task, and determining the sub-evaluation of each of the plurality of outputs comprises:

[0125] Obtaining a knowledge recall output corresponding to the knowledge recall task from the plurality of outputs;

[0126] An evaluation of the knowledge recall task is determined based on the knowledge recall output and a baseline knowledge recall corresponding to the user input.

[0127] Example 3. The method of any of Examples 1-2, wherein determining the knowledge recall assessment corresponding to the knowledge recall task comprises:

[0128] Determining the correctness of data table selection by comparing the knowledge recall output with the data table in the benchmark knowledge recall;

[0129] Determining calibration accuracy by comparing the knowledge recall output with the indicators, dimensions, screening conditions, and domain knowledge in the benchmark knowledge recall; and

[0130] An evaluation of the knowledge recall task is determined based on the data table selection correctness and the calibration correctness.

[0131] Example 4. The method of any of Examples 1-3, wherein the plurality of subtasks comprises an analytical reasoning task, and determining the sub-evaluation of each of the plurality of outputs comprises:

[0132] Obtaining analysis and reasoning output corresponding to the analysis and reasoning task;

[0133] generating prompt content based on the analytical reasoning output and the benchmark analytical reasoning;

[0134] The prompt content is processed using a reasoning evaluation model to determine an evaluation of the analytical reasoning task.

[0135] Example 5. The method of any of Examples 1-4, wherein the plurality of subtasks includes a code writing task, and determining the sub-assessment for each of the plurality of outputs includes:

[0136] Obtaining a code writing output corresponding to the code writing task;

[0137] An evaluation of the code writing task is determined based on the code writing output and a benchmark code.

[0138] Example 6. The method of any of Examples 1-5, wherein determining the assessment of the code writing task comprises:

[0139] Determining the structural similarity between the coding output and the benchmark code using an abstract syntax tree;

[0140] Determining execution correctness based on the coding output and the data results generated by the benchmark code; and

[0141] The evaluation of the coding task is determined based on the structural similarity and the execution correctness.

[0142] Example 7. The method of any of Examples 1-6, wherein determining the execution correctness of the coding output comprises:

[0143] determining the maximum common match and union of the prediction data table and the benchmark data table; and

[0144] Based on the maximum common matching and the union, an intersection-over-union ratio is determined as the execution correctness.

[0145] Example 8. The method of any of Examples 1-7, wherein the plurality of subtasks comprises a conclusion generation task, and determining the sub-evaluation of each of the plurality of outputs comprises:

[0146] Obtaining a conclusion generation output corresponding to the conclusion generation task;

[0147] Generate output and benchmark conclusions based on the conclusions and generate prompt content

[0148] The prompt content is processed using a conclusion evaluation model to determine an evaluation of the conclusion generation task.

[0149] Example 9. The method of any of Examples 1-8, wherein determining a comprehensive assessment for the agent comprises:

[0150] combining the sub-evaluations based on the sub-evaluations of each of the plurality of outputs using a predetermined strategy; and

[0151] By combining the sub-evaluations, the comprehensive evaluation for the agent is determined.

[0152] Example 10. An evaluation device comprising:

[0153] an output generation module configured to generate a plurality of outputs corresponding to the plurality of subtasks by executing the plurality of subtasks according to user input by the agent;

[0154] a sub-assessment determination module configured to determine a sub-assessment for each of the plurality of outputs based on the plurality of outputs; and

[0155] A comprehensive evaluation determination module is configured to determine a comprehensive evaluation for the agent based on the sub-evaluation of each of the plurality of outputs.

[0156] Example 11. The apparatus of Example 10, wherein the plurality of subtasks includes a knowledge recall task, and the recall assessment determination module includes:

[0157] a recall output acquisition module, configured to acquire a knowledge recall output corresponding to the knowledge recall task from a plurality of outputs;

[0158] The recall evaluation determination module is configured to determine an evaluation of the knowledge recall task based on the knowledge recall output and a baseline knowledge recall corresponding to the user input.

[0159] Example 12. The apparatus of any of Examples 10-11, wherein determining the knowledge recall assessment corresponding to the knowledge recall task comprises:

[0160] a table selection determination module configured to determine the correctness of data table selection by comparing the knowledge recall output with the data table in the benchmark knowledge recall;

[0161] a calibration correctness determination module configured to determine calibration correctness by comparing the knowledge recall output with the indicators, dimensions, screening conditions, and domain knowledge in the benchmark knowledge recall; and

[0162] The recall evaluation second determination module is configured to determine the evaluation of the knowledge recall task based on the data table selection correctness and the calibration correctness.

[0163] Example 13. The apparatus of any of Examples 10-12, wherein the plurality of subtasks include an analytical reasoning task, and the recall assessment determination module includes:

[0164] A reasoning output determination module, configured to obtain an analysis and reasoning output corresponding to the analysis and reasoning task;

[0165] a first prompt content generating module configured to generate prompt content based on the analysis and reasoning output and the benchmark analysis and reasoning;

[0166] The reasoning evaluation determination module is configured to process the prompt content using a reasoning evaluation model to determine the evaluation of the analytical reasoning task.

[0167] Example 14. The apparatus of any of Examples 10-13, wherein the plurality of subtasks includes a code writing task, and the recall assessment determination module includes:

[0168] a code output determination module, configured to obtain a code writing output corresponding to the code writing task;

[0169] The code evaluation determination module is configured to determine an evaluation of the code writing task based on the code writing output and a benchmark code.

[0170] Example 15. The apparatus of any of Examples 10-14, wherein determining the assessment of the code writing task comprises:

[0171] a structural similarity determination module configured to determine the structural similarity between the coding output and the benchmark code using an abstract syntax tree;

[0172] an execution correctness determination module configured to determine execution correctness based on the code writing output and data results generated by the benchmark code; and

[0173] A second code evaluation determination module is configured to determine the evaluation of the code writing task based on the structural similarity and the execution correctness.

[0174] Example 16. The apparatus of any of Examples 10-15, wherein determining the execution correctness of the coding output comprises:

[0175] a data table intersection and union determination module configured to determine a maximum common match and union between the prediction data table and the benchmark data table; and

[0176] The data table intersection-to-union ratio determination module is configured to determine the intersection-to-union ratio as the execution correctness based on the maximum common matching and the union.

[0177] Example 17. The apparatus of any of Examples 10-16, wherein the plurality of subtasks includes a conclusion generation task, and the recall assessment determination module includes:

[0178] a conclusion output determination module, configured to obtain a conclusion generation output corresponding to the conclusion generation task;

[0179] A second prompt content determination module is configured to generate prompt content based on the conclusion generation output and the benchmark conclusion;

[0180] The conclusion evaluation determination module is configured to process the prompt content using a conclusion evaluation model to determine the evaluation of the conclusion generation task.

[0181] Example 18. The apparatus of any of Examples 10-17, wherein determining a comprehensive assessment for the agent comprises:

[0182] a sub-assessment combining module configured to combine the sub-assessments using a predetermined strategy based on the sub-assessments of each of the plurality of outputs; and

[0183] The second comprehensive evaluation determination module is configured to determine the comprehensive evaluation for the agent by combining the sub-evaluations.

[0184] Example 19. An electronic device comprising:

[0185] processor; and

[0186] A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device performs actions, the actions comprising:

[0187] The agent performs a plurality of subtasks according to user input to generate a plurality of outputs corresponding to the plurality of subtasks;

[0188] determining a sub-assessment for each of the plurality of outputs based on the plurality of outputs; and

[0189] Based on the sub-evaluations of each of the plurality of outputs, a composite evaluation for the agent is determined.

[0190] Example 20. The electronic device of Example 19, wherein the plurality of subtasks comprises a knowledge recall task, and determining the sub-evaluation of each of the plurality of outputs comprises:

[0191] Obtaining a knowledge recall output corresponding to the knowledge recall task from the plurality of outputs;

[0192] An evaluation of the knowledge recall task is determined based on the knowledge recall output and a baseline knowledge recall corresponding to the user input.

[0193] Example 21. The electronic device of any of Examples 19-20, wherein determining the knowledge recall assessment corresponding to the knowledge recall task comprises:

[0194] Determining the correctness of data table selection by comparing the knowledge recall output with the data table in the benchmark knowledge recall;

[0195] Determining calibration accuracy by comparing the knowledge recall output with the indicators, dimensions, screening conditions, and domain knowledge in the benchmark knowledge recall; and

[0196] An evaluation of the knowledge recall task is determined based on the data table selection correctness and the calibration correctness.

[0197] Example 22. The electronic device of any of Examples 19-21, wherein the plurality of subtasks comprises an analytical reasoning task, and determining the sub-evaluation of each of the plurality of outputs comprises:

[0198] Obtaining analysis and reasoning output corresponding to the analysis and reasoning task;

[0199] generating prompt content based on the analytical reasoning output and the benchmark analytical reasoning;

[0200] The prompt content is processed using a reasoning evaluation model to determine an evaluation of the analytical reasoning task.

[0201] Example 23. The electronic device of any of Examples 19-22, wherein the plurality of subtasks comprises a code writing task, and determining the sub-evaluation of each of the plurality of outputs comprises:

[0202] Obtaining a code writing output corresponding to the code writing task;

[0203] An evaluation of the code writing task is determined based on the code writing output and a benchmark code.

[0204] Example 24. The electronic device of any of Examples 19-23, wherein determining the assessment of the code writing task comprises:

[0205] Determining the structural similarity between the coding output and the benchmark code using an abstract syntax tree;

[0206] Determining execution correctness based on the coding output and the data results generated by the benchmark code; and

[0207] The evaluation of the coding task is determined based on the structural similarity and the execution correctness.

[0208] Example 25. The electronic device of any of Examples 19-24, wherein determining the execution correctness of the coding output comprises:

[0209] determining the maximum common match and union of the prediction data table and the benchmark data table; and

[0210] Based on the maximum common matching and the union, an intersection-over-union ratio is determined as the execution correctness.

[0211] Example 26. The electronic device of any of Examples 19-25, wherein the plurality of subtasks comprises a conclusion generation task, and determining the sub-evaluation of each of the plurality of outputs comprises:

[0212] Obtaining a conclusion generation output corresponding to the conclusion generation task;

[0213] Generate output and benchmark conclusions based on the conclusions and generate prompt content

[0214] The prompt content is processed using a conclusion evaluation model to determine an evaluation of the conclusion generation task.

[0215] Example 27. The electronic device of any of Examples 19-26, wherein determining a comprehensive assessment for the agent comprises:

[0216] combining the sub-evaluations based on the sub-evaluations of each of the plurality of outputs using a predetermined strategy; and

[0217] By combining the sub-evaluations, the comprehensive evaluation for the agent is determined.

[0218] Example 28. A computer-readable storage medium having one or more computer instructions stored thereon, wherein the one or more computer instructions are executed by a processor to implement the method according to any one of Examples 1 to 9.

[0219] Example 29. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any one of Examples 1 to 9.

[0220] Although the present disclosure has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An evaluation method comprising: A data analysis agent constructed based on a large model performs a data analysis task according to a user input in a natural language format, wherein the data analysis task is specified in the user input and is divided into a plurality of subtasks, including a knowledge recall task, an analytical reasoning task, a code writing task, and a conclusion generation task; The data analysis agent executes the plurality of subtasks to generate a plurality of outputs corresponding to the plurality of subtasks; determining a sub-evaluation for each of a plurality of sub-tasks based on the plurality of outputs; as well as determining a comprehensive evaluation for the agent based on the sub-evaluations of each of the plurality of subtasks, Wherein determining the sub-assessment of the code writing task among the plurality of sub-tasks comprises: Obtaining a code writing output corresponding to the code writing task; and determining an evaluation of the coding task based on the coding output and a benchmark code, Wherein determining the assessment of the coding task comprises: Determining execution correctness based on the coding output and the data results generated by the benchmark code; and determining the evaluation of the coding task based on the structural similarity and the execution correctness between the coding output and the benchmark code, Determining the execution correctness of the coding output includes: determining the maximum common match and union of the prediction data table and the benchmark data table; and Based on the maximum common matching and the union, an intersection-over-union ratio is determined as the execution correctness.

2. The method of claim 1 , wherein determining the sub-assessment of the knowledge recall task among the plurality of sub-tasks comprises: Obtaining a knowledge recall output corresponding to the knowledge recall task from the plurality of outputs; as well as An evaluation of the knowledge recall task is determined based on the knowledge recall output and a baseline knowledge recall corresponding to the user input.

3. The method of claim 2, wherein determining the knowledge recall assessment corresponding to the knowledge recall task comprises: Determining the correctness of data table selection by comparing the knowledge recall output with the data table in the benchmark knowledge recall; Determine calibration accuracy by comparing the knowledge recall output with the indicators, dimensions, screening conditions, and domain knowledge in the benchmark knowledge recall; as well as An evaluation of the knowledge recall task is determined based on the data table selection correctness and the calibration correctness.

4. The method of claim 1 , wherein determining the sub-assessment of the analytical reasoning task among the plurality of sub-tasks comprises: Obtaining analysis and reasoning output corresponding to the analysis and reasoning task; generating prompt content based on the analytical reasoning output and the benchmark analytical reasoning; as well as The prompt content is processed using a reasoning evaluation model to determine an evaluation of the analytical reasoning task.

5. The method according to claim 1, further comprising: An abstract syntax tree is used to determine the structural similarity between the coding output and the benchmark code.

6. The method of claim 1 , wherein determining the sub-evaluation of the conclusion-generating task among the plurality of sub-tasks comprises: Obtaining a conclusion generation output corresponding to the conclusion generation task; generating output and benchmark conclusions based on the conclusions, and generating prompt content; as well as The prompt content is processed using a conclusion evaluation model to determine an evaluation of the conclusion generation task.

7. The method of claim 1 , wherein determining a comprehensive assessment for the agent comprises: Based on the sub-evaluations of each of the plurality of sub-tasks, combining the sub-evaluations using a predetermined strategy; as well as By combining the sub-evaluations, the comprehensive evaluation for the agent is determined.

8. An evaluation device comprising: an output generation module configured to execute a data analysis task according to a user input in a natural language form through a data analysis agent constructed based on the large model, wherein the data analysis task is specified in the user input and is divided into a plurality of subtasks, the plurality of subtasks including a knowledge recall task, an analytical reasoning task, a code writing task, and a conclusion generation task; The data analysis agent executes the plurality of subtasks to generate a plurality of outputs corresponding to the plurality of subtasks; a sub-assessment determination module configured to determine a sub-assessment of each of the plurality of sub-tasks based on the plurality of outputs; a comprehensive evaluation determination module configured to determine a comprehensive evaluation for the agent based on the sub-evaluation of each of the plurality of subtasks; a code output determination module, configured to obtain a code writing output corresponding to the code writing task; a code evaluation determination module configured to determine an evaluation of the code writing task based on the code writing output and a benchmark code; an execution correctness determination module configured to determine execution correctness based on the code writing output and data results generated by the benchmark code; a code evaluation second determination module configured to determine the evaluation of the code writing task based on the structural similarity and the execution correctness between the code writing output and the benchmark code; a data table intersection and union determination module configured to determine a maximum common match and union between the prediction data table and the benchmark data table; as well as The data table intersection-to-union ratio determination module is configured to determine the intersection-to-union ratio as the execution correctness based on the maximum common matching and the union.

9. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 7.

10. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising computer-executable instructions for performing the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Automatic evaluation method for Python drawing program questions

    CN110765014A

  • File processing method and device, electronic equipment and storage medium

    CN116702720A