Self-commenting ability assessment method of large language model
By building a data set for error tool calls and a fine-grained evaluation system, the self-criticism ability of large language models in tool calling tasks is evaluated, and the problem of bias in the existing technology of model error detection and mitigation ability evaluation is solved, and a more accurate evaluation of the performance of model tool usage is achieved.
Patent Information
- Application Number
- CN202510190720.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to deeply evaluate the ability of large language models to detect and mitigate errors in tool calling tasks, resulting in bias in the evaluation of model tool usage performance.
By constructing a data set covering error tool calls and designing a fine-grained evaluation system, the model can detect the performance of errors after self-reflection to better evaluate the self-criticism ability of large language models in tool calling tasks.
This has achieved in-depth evaluation of the self-critical ability of large language models in complex error scenarios, providing more accurate tool usage performance evaluation, and helping to understand the specific behavior of the model when facing different errors.
Smart Images

Figure CN120045901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method for evaluating the self-criticism ability of large language models. Background Art
[0002] As a breakthrough in the field of artificial intelligence, large language models (LLMs) have demonstrated extraordinary capabilities in various tasks. The interaction between large language models and external tools enables them to solve more complex tasks, making large language models more adaptable to dynamic real-world environments. Therefore, the evaluation of tool calls in large language models is a topic that requires in-depth research.
[0003] Existing technologies are usually limited to the use scenarios of single tools or comparing the execution results with predefined benchmark answers. However, real-world applications often involve complex and multi-step tool call behaviors. Due to their own capacity limitations or external unstable factors, large language models may cause errors in complex intermediate processes. Due to the complexity of the external environment and the high difficulty of the tool usage tasks themselves, ignoring the process status of tool calls may lead to biases in the evaluation of the model's tool call capabilities. Current benchmark tests mainly solve these problems by filtering out incorrect data or treating errors as sub-optimal nodes to expand the tool answer search space. Therefore, these methods cannot deeply understand how large language models detect and mitigate errors during tool calls, and thus cannot fully evaluate their tool usage capabilities.
[0004] Given that the sources of errors are diverse and the reasonable solutions for different errors are not the same, benchmarks that ignore error recovery in large language models cannot accurately evaluate the actual tool usage performance of the models.
[0005] How to evaluate the ability of large language models to handle different errors during tool call tasks is the main technical problem to be solved by the present invention. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a method for evaluating the self-criticism ability of large language models. Although the current evaluation benchmarks for large language model tool calls can accurately characterize the tool call capabilities of different large language models, they usually filter out incorrect data or only compare the prediction results with predefined benchmark answers, ignoring the possible error corrections during the process. Considering the diversity and complexity of error types, this result-oriented evaluation method does not make a targeted fine-grained design for the error correction phenomenon. Therefore, the present invention aims to construct a dataset covering incorrect tool calls through an innovative method and design a targeted fine-grained evaluation system to detect the performance of the model in correcting errors after self-reflection, so as to better quantitatively evaluate the ability of large language models to handle different errors when performing tool call tasks.
[0007] To solve the above technical problems, the present invention adopts the following technical solutions:
[0008] A method for evaluating the self-criticism ability of large language models, comprising:
[0009] Dataset construction stage:
[0010] Collect tool call traces and tool service interface documents from multiple tool call datasets;
[0011] Classify and expand errors; errors include model-driven errors and external environmental errors. Model-driven errors are expanded through simulation by the large language model, and external environmental errors are generated by repeated interface calls or expanded through simulation by the large language model;
[0012] Construct a tool service interface cache system and collect the responses of tools to model-driven errors according to the accessible status of the interfaces;
[0013] Evolve the constructed basic dataset and perform data verification;
[0014] Evaluation metric design stage:
[0015] Introduce step-level fine-grained evaluation tasks;
[0016] Design multi-dimensional evaluation metrics, including an error reflection metric for evaluating whether the large language model can correctly identify model-driven errors in the tool call trace, an error correction metric for evaluating whether the large language model can take correct actions to correct errors after identifying model-driven errors, an error retry metric for evaluating whether the large language model can retry failed tool call behaviors when encountering external environmental errors, and a skip / complete metric for evaluating whether the large language model can break the infinite retry loop and correctly execute subsequent subtasks or end the tool call behavior when it cannot solve external environmental errors; by weighted summing the multi-dimensional evaluation metrics, the overall self-criticism ability score of the large language model is obtained.
[0017] In one embodiment, collecting tool call traces and tool service interface documents from multiple tool call datasets specifically includes:
[0018] Collecting the interaction traces between predefined tool call behaviors and corresponding tool responses in multiple tool call datasets and the tool service interface documents used; reviewing the collected interaction traces to filter out data containing errors; improving the descriptions of the tool service interface documents to ensure the clarity and accuracy of the documents; standardizing the formats of the collected interaction traces and tool service interface documents.
[0019] In one embodiment, the model-driven errors are extended through simulation by a large language model, and the external environment errors are generated through repeated interface calls or extended through simulation by a large language model, specifically including:
[0020] For model-driven errors, using a large language model as an error simulator to simulate errors in tool calls; for external environment errors, when the interface is accessible, making repeated calls to collect errors that occur when the external environment is unstable, and when the interface is inaccessible, using a large language model as a tool service interface simulator to collect error responses.
[0021] In one embodiment, building a tool service interface cache system and collecting the responses of tools to model-driven errors according to the accessible state of the interface specifically includes:
[0022] Building a tool service interface cache system and collecting the tool service interface responses corresponding to model-driven errors according to the accessible state of the interface; the tool service interface cache system collects the parameters in the tool call traces and the corresponding tool service interface responses, and uses a large language model as an interface simulator to handle the situation where the interface is inaccessible.
[0023] In one embodiment, the basic dataset before the evolution of the dataset includes:
[0024] Tool call tasks, which record at least the task objectives that the user expects to solve;
[0025] Tool service interface documents, which describe at least the functions, parameters, and usage methods of each tool;
[0026] Tool call traces, which record at least the tool call behavior steps and responses of some tasks;
[0027] Error data, which record at least the model-driven error behaviors or external environment error responses that occur during tool calls;
[0028] Response data, which records at least the response of the tool service interface to the model-driven error during the tool call process.
[0029] In one embodiment, evolving the constructed basic data set and performing data verification specifically include:
[0030] Evolving the basic data set through long context injection, additional tool addition, fuzzy instruction generation, and difficult tool transformation to make the data closer to real-world scenarios; and performing data verification so that the evolved data and the basic data have the same answers after being input into the large language model.
[0031] In one embodiment, the error reflection metrics for evaluating whether the large language model can correctly identify model-driven errors in the tool call trajectory specifically include:
[0032] Each tool call trajectory includes multiple steps of tool call behavior. According to the correctness of the k-th step of tool call behavior a k to determine whether to generate a criticism c pred ; if there is no error in a k , and the large language model can correctly predict no error, the error reflection metric is full marks; if there is an error in a k , compare c pred with the benchmark answer c gt . If the large language model can successfully identify and accurately classify the error type, the error reflection metric is full marks.
[0033] In one embodiment, the error correction metrics for evaluating whether the large language model can take correct actions to correct errors after identifying model-driven errors specifically include:
[0034] Each tool call trajectory includes multiple steps of tool call behavior. The large language model generates a corrective action k for the error of the k-th step of tool call behavior a and compare with the benchmark answer to evaluate whether the large language model can take correct actions to correct errors after identifying the errors and generate error correction metrics.
[0035] In one embodiment, the error retry metrics for evaluating whether the large language model can retry failed tool call behaviors when encountering external environment errors specifically include:
[0036] Each tool call trajectory includes multiple steps of tool call behavior. When it is found that there is any error caused by external environment instability in the response r k generated by the k-th step of tool call behavior, the large language model generates a repeated tool call behavior Compare with the standard answer that is the same as the tool call behavior a in the k-th step k to evaluate whether the large language model can retry the failed tool call behavior when encountering external environment errors, and generate an error retry metric.
[0037] In one embodiment, the skip / complete metric for evaluating whether the large language model can break the infinite retry loop and correctly execute subsequent subtasks or end the tool call behavior when it cannot solve external environment errors specifically includes:
[0038] If the external environment error cannot be solved within the retry count limit, the large language model generates a skip behavior to continue to the next feasible subtask or ends the tool call behavior, which is uniformly recorded as By comparing the skip behavior with the standard behavior of the next subtask or by checking whether it is an end tool call behavior, to evaluate whether the large language model can break the infinite retry loop and correctly execute subsequent subtasks or end the tool call behavior when it cannot solve external environment errors, and generate a skip / complete metric.
[0039] Compared with the prior art, the beneficial technical effects of the present invention are:
[0040] 1. Novel application scenario: The present invention proposes a novel application scenario, that is, to evaluate the self-criticism ability of large language models (LLMs) in the tool call error scenario. The present invention proposes a self-criticism ability evaluation system for large language model tool learning (hereinafter referred to as the CRITICTOOL system), which evaluates the performance of large language models in these complex scenarios by simulating various tool use errors that may occur in the real world, including in-model drive errors and external environment errors. This evaluation method not only covers various error modes that large language models may encounter in tool call tasks. In a preferred embodiment, the present invention also enhances the realism and challenge of the evaluation method through data evolution strategies (such as long text context, additional tools, noisy queries, and more difficult tools). This novel application scenario provides deeper insights into the tool learning ability of large language models and provides valuable guidance for the development of future tool call agents.
[0041] 2. High-quality dataset: The CRITICTOOL system constructs a high-quality dataset. Based on the evaluation benchmarks of several high-quality tool calls, through a systematic collection process, it can gather 733 high-quality tool call traces from tools in different fields. Then, it uses the collected traces for error diversification processing to ensure the diversity and coverage of the dataset. In one embodiment, the basic data of the CRITICTOOL system altogether contains 1490 test cases, specifically 1316 in-model drive error test cases and 174 out-of-environment error test cases. Moreover, the CRITICTOOL system also adopts an extensible and robust self-evolving hybrid (SRM) strategy. By introducing long text contexts, additional tools, noisy queries, and more difficult tools, it increases the complexity and authenticity of the dataset, obtaining 1250 evolved data cases, including 1000 in-model drive error test cases and 250 out-of-environment error test cases. This high-quality dataset provides a solid foundation for evaluating the self-criticism ability of large language models in tool call tasks.
[0042] 3. Fine-grained evaluation metrics: The CRITICTOOL system adopts fine-grained evaluation metrics to comprehensively evaluate the self-criticism ability of large language models. These metrics include Reflect, Correct, Retry, and Skip / Finish, which respectively correspond to the large language model's behaviors in identifying errors, analyzing error causes, taking corrective measures, and coping strategies when facing unsolvable errors. These evaluation metrics not only focus on the overall performance of large language models in tool call tasks but also delve into each step to evaluate the specific behaviors of large language models when facing different error patterns. Through this fine-grained evaluation method, the CRITICTOOL system can provide a more in-depth analysis of the tool usage ability of large language models, reveal the main bottlenecks in current large language models' tool learning, and provide new directions for the future development of large language models in the field of tool learning. Brief Description of the Drawings
[0043] Figure 1 It is the flowchart of the method in the embodiment of the present invention.
[0044] Figure 2 It is the schematic diagram of the innovation of the method in the embodiment of the present invention.
[0045] Figure 3 It is the schematic diagram of the superiority of data evolution in the embodiment of the present invention.
[0046] Figure 4 It is the schematic diagram of the evaluation system in the embodiment of the present invention. Detailed Embodiment
[0047] The following is a detailed description of a preferred embodiment of the present invention in conjunction with the accompanying drawings.
[0048] The present invention mainly evaluates the ability of a large language model to handle different errors when performing tool call tasks by comparing the scores of different metrics of the large language model on a constructed dataset. Therefore, the method of the present invention mainly includes two key parts: dataset construction and evaluation metric design.
[0049] As Figure 1 shown, a method for evaluating the self-criticism ability of a large language model in the present invention includes the following steps:
[0050] A1. Dataset construction stage, including:
[0051] A11. Collect tool call traces and tool service interface documents from multiple tool call datasets;
[0052] A12. Classify and expand errors; errors include model-driven errors and external environment errors. Model-driven errors are expanded by simulating with a large language model, and external environment errors are generated by repeated interface calls or expanded by simulating with a large language model;
[0053] A13. Build a tool service interface cache system, and collect the responses of tools to model-driven errors according to the accessible status of the interfaces;
[0054] A14. Evolve the constructed basic dataset and perform data verification;
[0055] A2. Evaluation metric design stage, including:
[0056] A21. Introduce step-level fine-grained evaluation tasks;
[0057] A22. Design multi-dimensional evaluation metrics, including an error reflection metric for evaluating whether the large language model can correctly identify model-driven errors in the tool call trace, an error correction metric for evaluating whether the large language model can take correct actions to correct errors after identifying model-driven errors, an error retry metric for evaluating whether the large language model can retry failed tool call behaviors when encountering external environment errors, and a skip / complete metric for evaluating whether the large language model can break the infinite retry loop and correctly execute subsequent subtasks or end the tool call behavior when it cannot solve external environment errors; by weighted summing the multi-dimensional evaluation metrics, the overall self-criticism ability score of the large language model is obtained.
[0058] The CRITICTOOL system in the present invention provides an effective method for evaluating the self-criticism ability of large language models in the context of tool call errors through novel application scenarios, high-quality datasets, and fine-grained evaluation metrics. It not only provides in-depth insights into the performance of models in complex real-world tasks but also enhances the practical applicability of the evaluation method by simulating diverse error patterns. These advantages together promote a more comprehensive understanding of the tool learning ability of large language models and provide new perspectives and improvement directions for the future development of tool call agents.
[0059] As Figure 2 shown, the dataset construction stage includes four key processes: data collection, error expansion, response processing, and data evolution.
[0060] In one embodiment, collecting the tool call traces and tool service interface documents from multiple tool call datasets in step A11 specifically includes:
[0061] Collecting the interaction traces between the predefined tool call behaviors and the corresponding tool responses in multiple tool call datasets and the tool service interface documents used; strictly reviewing the collected interaction traces and filtering out the data containing errors; improving the description of the tool service interface documents to ensure the clarity and accuracy of the documents; and standardizing the formats of the collected interaction traces and tool service interface documents.
[0062] During the data collection process, the present invention was extensively tested on four high-quality and widely used large language model tool call datasets (BFCL, T-Eval, API-Bank, NESTFUL), and the errors that occurred during the tool call process were analyzed. The errors were divided into two major categories according to the source of the errors: in-model-driven errors and external environment errors. Among them, the in-model-driven errors include four sub-types of errors: tool selection error, tool hallucination error, parameter key error, and parameter value error. During the testing process, it was found that the data quality was uneven, some error situations were caused by unreasonable data, and the errors found during the testing process were not sufficient to comprehensively cover the scenarios where errors occurred. Therefore, the present invention selected two datasets with more suitable complexity, diversity, and data quality - the BFCL dataset and the T-EVAL dataset, and further performed data cleaning to facilitate the subsequent generation of error data. The present invention collected the interaction traces between the predefined tool call behaviors of the above datasets and the corresponding tool responses. To ensure the quality and reliability of the dataset, the collected data needs to be strictly reviewed, and any data containing errors (such as incorrect annotations or failed tool calls) is manually filtered out. Then, the present invention extracts the tool service interface (API) documentation and improves any ambiguous or insufficient descriptions to ensure the clarity and accuracy of the documentation, thereby minimizing potential misunderstandings. To further improve consistency, the present invention standardizes all tool call traces and API descriptions.
[0063] In one of the embodiments, the in-model-driven errors in step A12 are extended by large language model simulation, and the external environment errors are generated by repeated interface calls or extended by large language model simulation, specifically including:
[0064] For in-model-driven errors, an advanced large language model is used as an error simulator to simulate errors in tool calls; for external environment errors, when the interface is accessible, repeated calls are made to collect errors that occur when the external environment is unstable, and when the interface is inaccessible, an advanced large language model is used as a tool service interface simulator to collect error responses.
[0065] The error expansion process is to systematically increase the diversity of errors obtained during the data collection process, so that the dataset of the present invention can more comprehensively cover potential error scenarios. Observations during the data collection process show that although interacting with different tools, large language models tend to exhibit similar behaviors in specific categories of errors. This similarity enables the present invention to easily obtain more errors that occur during tool calls and expand the diversity of the dataset of the present invention. For model-driven errors, the present invention uses the large language model GPT-4o as an error simulator to simulate the error-prone behaviors of large language models during tool calls. In the prompt, examples of previously collected error patterns are used as few-shot examples to let the error simulator generate various error instances in a wider range of tools and tasks. For external environment errors, the present invention implements repeated API call and API simulation strategies. Due to the inherent instability of the external environment, responses containing errors may occur during calls at certain times. Therefore, the present invention makes repeated calls to accessible APIs to collect external environment errors. For inaccessible APIs, the present invention uses GPT-4o as an API simulator to collect external environment errors.
[0066] In one embodiment, building a tool service interface cache system in step A13 and collecting the response of the tool to model-driven errors according to the accessibility status of the interface specifically includes:
[0067] A tool service interface cache system is built, and the tool service interface response corresponding to the model-driven error is collected according to the accessibility status of the interface; the tool service interface cache system collects the parameters in the tool call trace and the corresponding tool service interface response, and uses a large language model as an interface simulator to handle the situation where the interface is inaccessible.
[0068] In one embodiment, search the cache in the cache system to check whether the parameters of the currently called tool have been cached. If so, use the cached tool response as the tool service interface response corresponding to the current driven error; if not, verify the accessibility of the corresponding interface. When the interface is accessible, perform the tool call behavior and use the actual tool service interface response. When the interface is inaccessible, use a large language model as an interface simulator to generate the tool service interface response of the tool call behavior in the current driven error.
[0069] During the response processing, since the responses received by large language models from the environment during tool calls are crucial for their self-criticism and error correction, the present invention performs response processing on the in-model drive errors extended. However, factors such as unstable network environments and restricted access make it so that not all collected APIs are executable. The present invention constructs an API caching system to collect environmental responses based on the accessible status of the APIs. This caching system stores the parameters in the collected tool call traces and the corresponding tool responses, and uses GPT-4o as an API simulator to handle situations where tool execution is unavailable. There are three sub-processes for obtaining response processing from the caching system: cache retrieval, API execution, and simulator response. The present invention first searches the cache to check if the tools and parameters used in the current call have been cached. If a match is found, the cached response is used as the environmental response for the current tool call. If there is no match in the cache, the present invention verifies the accessibility of the API. When the API is available, the present invention executes the tool call and uses the actual API response. When there is no cache and the API is unavailable, the present invention uses GPT-4o as an API simulator to ensure that the tool call assistant can still receive feedback on its current operation.
[0070] In one embodiment, the constructed basic dataset is evolved and data-verified in step A14. The basic dataset before dataset evolution includes:
[0071] Tool call tasks, which at least record the task objectives expected to be solved by the user;
[0072] Tool service interface documents, which at least describe the functions, parameters, and usage methods of each tool;
[0073] Tool call traces, which at least record the tool call behavior steps and responses of part of the tasks;
[0074] Error data, which at least record the in-model drive error behaviors or external environment error responses that occur during tool calls;
[0075] Response data, which at least record the responses of the tool service interface to the in-model drive errors during tool calls.
[0076] The constructed basic dataset is evolved and data-verified, specifically including:
[0077] The basic dataset is evolved through long context injection, additional tool addition, fuzzy instruction generation, and difficult tool transformation to make the data closer to real-world scenarios; and data verification is performed so that the evolved data and the basic data have the same answers after being input into the large language model.
[0078] Considering that real-world tool calls usually involve complex contexts, complex tools, and ambiguous user queries, in order to more realistically evaluate the performance of large language models in tool call tasks, the present invention conducts data evolution from two dimensions: scalability and robustness. The data evolution specifically includes four sub-strategies: long context, additional tools, ambiguous instructions, and difficult tools. The long context randomly selects conversations with 1000 to 3000 tokens from the LongBench dataset as the context. Additional tools refer to randomly selecting 4 to 8 additional APIs from the total tool library and adding them to the tool list. Ambiguous instructions use prompts to make the original user instructions verbose and complex using GPT-4o and add typos. Difficult tools use prompts to make the original API documentation verbose and complex using GPT-4o and add typos. After data evolution, the present invention conducts data verification to ensure that the reference answers of the evolved data remain unchanged. The present invention uses prompts to use GPT-4o to check whether the modifications or additions made during the evolution process have a significant impact on the tool usage task. Through the above evolution strategies, the present invention constructs an evolved dataset based on the basic dataset.
[0079] In one embodiment, the introduction of step-level fine-grained evaluation tasks in step A21 specifically includes:
[0080] According to the characteristics of error correction and self-reflection, the present invention constructs a fine-grained evaluation system for the above error dataset. The present invention first defines the evaluation task. In the CRITICTOOL system, each tool usage task is defined as a tuple (Q, L, T), where Q is the query related to the tool call task, and L represents the list of APIs that the large language model assistant can call. The present invention defines the tool call trajectory T as a series of tool-response pairs {(a i , r i )}, which captures the interaction between the tool call behavior a of the large language model assistant at the i-th step and the corresponding tool service interface response r. The action a is regarded as (goal, tool, args) or (tool, args) according to whether the Chain of Thought (CoT) strategy is applied.
[0081] The complex interaction between the large language model and the environment may lead to potential errors at any step. Therefore, it is necessary to evaluate the self-criticism ability of the large language model at the step level. The test data includes the first k steps of the tool call trajectory of each task, where k is randomly selected, and errors may be introduced at step k. When the tool call behavior a k at step k contains an error, the present invention defines the solution as where c represents the reflection on the error, and It is a corrective action for errors; if a k There is no error, then or
[0082] When evaluating the self-criticism ability of the evaluation model for endogenous errors, the CRITICTOOL system jointly uses correct and incorrect data to ensure fairness and robustness. The present invention evaluates the prediction result of the (k + 1)-th step and decomposes the self-criticism process into two dimensions. When the large language model assistant makes a tool call, it should first identify whether an error has occurred in the previous tool call and determine its specific category. This process is called reflection, which is the primary step of the model's self-criticism. Based on the result of reflection, the model needs to take corrective measures to recover from the error. The present invention defines this process as correction, highlighting the model's ability to improve and adapt its behavior. For correct data, the present invention only focuses on whether the model reflection process is correct and does not involve the evaluation of error correction; for incorrect data, the present invention evaluates the model's abilities in both the reflection and correction dimensions.
[0083] For tasks involving external environmental errors, the present invention hopes that the large language model assistant can properly handle these responses containing error signals in subsequent steps. The present invention encourages the model to retry failed tool calls within a limited number of times to avoid accidental errors caused by environmental instability. If the problem still exists after multiple retries, the large language model assistant should skip the problematic step and continue to execute the remaining viable subtasks, or end the tool call behavior and notify the user that further guidance is needed.
[0084] In one embodiment, the error reflection metric for evaluating whether the large language model can correctly identify endogenous errors in the tool call trajectory in step A22 specifically includes:
[0085] Each tool call trajectory includes multiple steps of tool call behavior. According to the correctness of the k-th step of tool call behavior a k to determine whether to generate criticism c pred ; if there is no error in a k , and the large language model can correctly predict no error, then the error reflection metric is full marks; if there is an error in a k , compare c pred with the reference answer c gt . If the large language model can successfully identify and accurately classify the error type, then the error reflection metric is full marks.
[0086] The error reflection (Reflect) evaluator requires the large language model assistant to determine whether to generate criticism c k according to the correctness of the tool call behavior a pred . If ak There are errors in pred Compare with the reference answer c gt This dimension focuses on evaluating whether the large language model can accurately identify errors in the tool call trace. In error-free tasks, if the large language model can correctly predict no errors, it gets a full score; in error-containing tasks, if the large language model can successfully identify and accurately classify the error types, it also gets a full score.
[0087]
[0088] Among them,
[0089]
[0090] Among them, ReflectScore is the score of the error reflection index, detectscore (detection) indicates whether the model correctly identifies whether an error occurs during the reflection process, and categoryscore (classification) indicates whether the model can correctly classify the error during the reflection process (tool selection error, tool hallucination error, parameter key error, parameter value error).
[0091] In one of the embodiments, the error correction index used to evaluate whether the large language model can take correct actions to correct errors after identifying internal drive errors in the model in step A22 specifically includes:
[0092] Each tool call trace includes multiple steps of tool call behavior. The large language model generates a corrective action for the error in k the k-th step of tool call behavior a and compares with the reference answer to evaluate whether the large language model can take correct actions to correct errors after identifying the errors and generate an error correction index.
[0093] The error correction (Correct) evaluator requires the large language model assistant to generate a corrective action for the error in k the detected tool call action a and compares with the reference answer This dimension evaluates whether the model can take correct actions to correct errors after identifying the errors. This includes tool prediction, parameter correction, and the accuracy of the thinking process under the CoT strategy.
[0094]
[0095] Among them,
[0096]
[0097] thoughtscore = cosinesimilarity(thought pred , thought gt );
[0098] CorrectScore represents the error correction index score, toolscore (tool) represents whether the model can call the correct tool service interface to obtain information for the current task during the correction process, argsscore (parameters) represents whether the model can pass the correct parameters when calling the tool service interface during the correction process, thoughtscore (target) represents whether the model has the correct target for the current task during the correction process, tool pred represents the tool service interface called by the model during the evaluation process, tool gt represents the tool service interface that the model should call in the standard answer, represents the parameters passed by the model to the tool service interface during the evaluation process, represents the parameters that the model should pass to the tool service interface in the standard answer, n represents the number of parameters, thought pred represents the model's target for the current step during the evaluation process, thought gt represents the model's target for the current step in the standard answer, cosinestmilarity represents cosine similarity.
[0099] In one of the embodiments, the error retry metric for evaluating whether the large language model can retry failed tool call behaviors when encountering external environment errors in step A22 specifically includes:
[0100] Each tool call trace includes multiple steps of tool call behaviors. When it is found that there are any errors caused by unstable external environments in the response r k generated by the k-th step of tool call behavior, the large language model generates repeated tool call behaviors will be compared with the same standard answer as the k-th step of tool call behavior a k to evaluate whether the large language model can retry failed tool call behaviors when encountering external environment errors and generate an error retry metric.
[0101] The error retry (Retry) evaluator requires the large language model assistant to generate repeated tool calls when any error signals are found in r k k The evaluator will and the same standard answer as a k Compare. This dimension evaluates whether the model can appropriately retry failed tool calls when encountering external environment errors.
[0102]
[0103] Retry Score represents the error retry metric score.
[0104] In one embodiment, the skip / finish metric in step A22 for evaluating whether the large language model can break the infinite retry loop and correctly execute subsequent subtasks or end the tool call behavior when it cannot solve external environment errors specifically includes:
[0105] If the external environment error cannot be resolved within the retry limit, the large language model generates a skip behavior to continue the next feasible subtask or ends the tool call behavior, which is uniformly recorded as By comparing the skip behavior with the standard behavior of the next subtask or by checking whether it is the end tool call behavior, to evaluate whether the large language model can break the infinite retry loop and correctly execute subsequent subtasks or end the tool call behavior when it cannot solve external environment errors, and generate the skip / finish metric.
[0106] If the environment error cannot be resolved within the retry limit, the Skip / Finish evaluator requires the large language model assistant to skip and continue the next feasible subtask or end the tool call behavior. The evaluator compares the skip behavior with the standard behavior of the next subtask or checks whether it is the end tool call behavior. This dimension evaluates whether the model can break the infinite retry loop, correctly execute subsequent subtasks or end the tool call behavior when it cannot solve the environment error.
[0107]
[0108] Among them,
[0109]
[0110] Skip / FinishScore is the skip / finish metric, breakscore (break) indicates whether the model can break the infinite retry loop, action score (behavior) indicates whether the model can correctly execute subsequent subtasks or end the tool call behavior, and is also split into tool call and parameter passing evaluations, indicating the tool call behavior of the model after breaking the infinite retry loop or the behavior of ending the tool call during the evaluation process, It indicates the behavior of the model retrying the previously failed tool call in the standard answer, and FinishAction indicates the behavior of the model ending the tool call.
[0111] In one embodiment, the overall self-criticism ability score of the large language model is obtained by weighted summing of the multi-dimensional evaluation indicators in step A22, which specifically includes:
[0112] The evaluation dimensions of error reflection and error correction are applied to model-driven errors, and the evaluation dimensions of error retry and skip / completion are applied to environmental external errors. Finally, the present invention weights and sums the scores of the above dimensions according to their importance in completing the tool call task to obtain the overall self-criticism ability score of the model:
[0113] Overall Score=0.2×Reflect Score+0.3×Correct Score+0.05×RetryScore+0.45×Skip / Finish Score.
[0114] This overall self-critical ability score provides a comprehensive perspective to evaluate the self-critical ability of large language models in tool invocation tasks. Through this detailed evaluation method, the present invention can have a deeper understanding of the model's behavior when facing different error types, and provide new perspectives and improvement directions for future developments in the field of tool learning.
[0115] Table 1, the evaluation results of the present invention on various models.
[0116]
[0117]
[0118] In Table 1, GPT-4o leads in self-criticism for tool usage error scenarios, achieving an impressive total score of 69.01. Large open-source models LLaMA3.1-70B and Qwen2.5-72B follow closely behind, achieving quite competitive scores and demonstrating strong self-criticism capabilities. For model-driven errors, closed-source models GPT-4o and Claude3.5 provide similar top-level performance, but Claude3.5 performs slightly worse in misclassification. In contrast, the self-criticism performance of open-source models varies greatly. While most open-source models lag significantly behind closed-source models, LLaMA3.1 and Qwen2.5 are exceptions, and their performance is not only close to that of closed-source models, but even exceeds that of closed-source models in some aspects. Models fine-tuned on tool-invoked data perform poorly in handling internal errors. Except for AgentLM-8B, other models have almost no instruction following or self-criticism capabilities, which can be attributed to the damage to their generalization ability caused by fine-tuning on specific data. For external environment errors, most models can identify errors and avoid infinite failed call retries, except for Claude3.5 and Minstral-8B, which are weak in this regard, and some models fine-tuned with tool call data completely lack this ability. In terms of continuing to perform subsequent tasks or stopping tool call behavior to handle errors, GPT-4o performs better than other models, and some large open source models also achieve quite strong performance.
[0119] CRITICTOOL provides an effective method to evaluate the self-criticism ability of large language models in the context of tool call errors through novel application scenarios, high-quality datasets, and fine-grained evaluation metrics. It not only provides deep insights into the performance of the model in complex real-world tasks, but also enhances the practical applicability of the evaluation by simulating diverse error patterns. These advantages together promote a more comprehensive understanding of the learning ability of large language model tools, and provide new perspectives and improvement directions for the future development of tool call agents.
[0120] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.
[0121] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment contains only one independent technical solution. This narrative style of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for evaluating the self-criticism ability of a large language model, characterized in that: include: Dataset construction phase: Collect tool call traces and tool service interface documents from multiple tool call datasets; Categorize and expand on the errors; Errors include model-driven errors and environment-external errors. Model-driven errors are extended through large language model simulations, while environment-external errors are generated through repeated interface calls or extended through large language model simulations. Build a tool service interface cache system and collect tool responses to model-driven errors based on the accessible status of the interface; Perform data evolution and data verification on the constructed basic data set; Evaluation indicator design phase: Introducing step-level fine-grained evaluation tasks; Design multi-dimensional evaluation indicators, including error reflection indicators for evaluating whether the large language model can correctly identify model-driven errors in tool call traces, error correction indicators for evaluating whether the large language model can take correct actions to correct errors after identifying model-driven errors, error retry indicators for evaluating whether the large language model can retry failed tool call behaviors when encountering external errors in the environment, and skip / completion indicators for evaluating whether the large language model can break the infinite retry cycle and correctly execute subsequent subtasks or end tool call behaviors when the external errors in the environment cannot be resolved; the overall self-criticism ability score of the large language model is obtained by weighted summation of the multi-dimensional evaluation indicators.
2. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The collecting of tool call traces and tool service interface documents from multiple tool call data sets specifically includes: Collect the interaction traces between predefined tool call behaviors and corresponding tool responses in multiple tool call data sets and the tool service interface documents used; review the collected interaction traces and filter out data containing errors; improve the description of the tool service interface documents to ensure the clarity and accuracy of the documents; standardize the formats of the collected interaction traces and tool service interface documents.
3. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The model internal drive error is extended through the large language model simulation, and the environment external error is generated through repeated interface calls or extended through the large language model simulation, specifically including: For model-driven errors, a large language model is used as an error simulator to simulate errors in tool calls. For external errors, when the interface is accessible, repeated calls are made to collect errors that occur when the external environment is unstable. When the interface is inaccessible, a large language model is used as a tool service interface simulator to collect error responses.
4. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The construction tool service interface cache system collects the tool's response to the model internal drive error according to the accessible state of the interface, specifically including: A tool service interface caching system was built, and the tool service interface responses corresponding to model-driven errors were collected based on the accessible status of the interface. The tool service interface caching system collected the parameters in the tool call trace and the corresponding tool service interface responses, and used a large language model as an interface simulator to handle situations where the interface is inaccessible.
5. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The basic data sets before data set evolution include: Tool call tasks at least record the task objectives that users expect to solve; Tool service interface documentation, which at least describes the functions, parameters, and usage of each tool; Tool call traces, which record the tool call behavior steps and responses for at least some tasks; Error data, which at least records the internal error behavior or external environment error response that occurs during the tool call process; The response data at least records the response of the tool service interface to the model internal drive error during the tool call process.
6. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The data evolution and data verification of the constructed basic data set specifically include: The basic data set is evolved through long context injection, additional tool addition, fuzzy instruction generation and difficult tool transformation to make the data closer to real-world scenarios; and data verification is performed so that the evolved data and the basic data have the same answer after being input into the large language model.
7. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The error reflection indicators used to evaluate whether the large language model can correctly identify model-driven errors in tool call traces specifically include: Each tool call trajectory includes multiple steps of tool call behavior. k The correctness of the pred ; if a k There are no errors in , and the large language model can correctly predict that there are no errors, then the error reflection index is full score; if a k There is an error in c pred With the benchmark answer c gt For comparison, if the large language model can successfully identify and accurately classify the error type, the error reflection indicator will be full score.
8. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The error correction indicators used to evaluate whether the large language model can take correct actions to correct the errors after identifying the model-driven errors specifically include: Each tool call trace includes multiple steps of tool call behavior. The large language model is based on the k-th step tool call behavior a k Errors generate corrective actions And compare Answer with benchmark To evaluate whether the large language model can take correct actions to correct errors after identifying them, and generate error correction indicators.
9. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The error retry indicators used to evaluate whether a large language model can retry a failed tool call behavior when encountering an external error in the environment specifically include: Each tool call trajectory includes multiple steps of tool call behavior. When the response r generated by the k-th step tool call behavior is found, k When there is any error caused by instability in the external environment, the large language model generates repeated tool call behavior Will With the k-th step tool call behavior a k Same standard answer By comparison, we evaluate whether the large language model can retry failed tool call behaviors when encountering external errors in the environment and generate error retry indicators.
10. The method for evaluating the self-criticism ability of a large language model according to claim 1, characterized in that: The skip / completion indicators used to evaluate whether the large language model can break the infinite retry cycle and correctly execute subsequent subtasks or end the tool call behavior when it cannot resolve external errors in the environment include: If the external error cannot be resolved within the retry limit, the large language model generates a skipping behavior to continue the next feasible subtask or terminate the tool call behavior, which is uniformly recorded as By adding the skip behavior Standard behavior with the next subtask Compare, or by looking at Whether it is the end tool call behavior, to evaluate whether the large language model can break the infinite retry cycle and correctly execute subsequent subtasks or end the tool call behavior when it cannot solve the external error of the environment, and generate skip / completion indicators.
Citation Information
Cited By
Assessment method and device of intelligent agent and large language model, medium and electronic equipment
CN120508507A