Task response method based on large model, and large model fine tuning method and device

By introducing a knowledge boundary evaluation mechanism into the large language model, the over-dependence or over-confidence problem of the model in response to tasks is solved, and more efficient and flexible tool call decisions are achieved, which enhances the robustness of the model.

CN120218181APending Publication Date: 2025-06-27AISPEECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510312068.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Large language models are prone to over-reliance or over-confident rejection of external tools when responding to tasks, resulting in increased computational costs and time delays, and the prior art has failed to effectively solve this problem.

Method used

By introducing a decision-making mechanism based on knowledge boundary evaluation, the initial output data of the big model for task requests is obtained and knowledge boundary evaluation is carried out, and whether to call matching target tools to respond.

Benefits of technology

It effectively avoids large models over-rely relying on external tools when they can complete tasks independently, or calling tools in time when tasks exceed the boundaries of model knowledge, which enhances the robustness of the model and task completion efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218181A_ABST
    Figure CN120218181A_ABST
Patent Text Reader

Abstract

The invention discloses a task response method based on a large model, and a large model fine tuning method and device, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining first preliminary output data of the large model for a task request; based on the first preliminary output data, determining a knowledge boundary evaluation result of the large model for the task request; the knowledge boundary evaluation result is used for indicating the probability that the task request exceeds or does not exceed the knowledge boundary of the large model; and according to a knowledge boundary evaluation result, deciding whether to call a matched target tool to respond to the task request. Therefore, a decision-making mechanism based on knowledge boundary evaluation is introduced, so that the large model can perform judgment according to the actual demand of the task and does not depend on a tool blindly or excessively confident any more, and the robustness of the large model in different environments and task scenes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a task response method based on a large model, a large model fine-tuning method, and a device. Background Art

[0002] The goal of tool learning is to enable a large language model (LLM) to effectively utilize external tools, thereby enhancing its performance in various downstream tasks. Tools can be regarded as an extension of the knowledge or ability boundary of the LLM. By invoking tools, the model can complete tasks beyond its knowledge boundary and even access information from different modalities. Specifically, by enabling the large language model to learn how to effectively utilize external tools such as calculators, search engines, inference models, etc., the model can call these tools during task execution to obtain information beyond its own knowledge boundary or complete specific functions, thus better completing the tasks.

[0003] Although tool learning can improve the performance of the model, the large model may over-rely on tools and call tools even when it can complete tasks independently, resulting in unnecessary computational costs and time delays. In addition, the large model may be overly confident in the use of tools and refuse to use them when tool invocation is required, thus affecting the success rate of the task. In particular, current tool learning technologies mainly focus on how to enable the model to effectively use tools to improve performance, but do not fully consider the decision-making mechanism of the large model for tool invocation during design, resulting in the large model may over-rely on tools or overly confidently refuse to use tools in some cases, weakening the tool intelligence of the large model and increasing the cost of task completion in real-world scenarios.

[0004] In response to the above problems, the industry has not yet proposed a better solution. Summary of the Invention

[0005] This application provides a task response method based on a large model, a large model fine-tuning method, an electronic device, a storage medium, and a computer software program product, which are used to at least solve the problem that the large model of tool learning in the current related technologies is overly confident or dependent on tools.

[0006] In a first aspect, an embodiment of this application provides a task response method based on a large model, including: obtaining first preliminary output data of the large model for a task request; based on the first preliminary output data, determining a knowledge boundary evaluation result of the large model for the task request; the knowledge boundary evaluation result is used to indicate the probability that the task request exceeds or does not exceed the knowledge boundary of the large model; according to the knowledge boundary evaluation result, making a decision on whether to invoke a matching target tool to respond to the task request.

[0007] Second aspect, an embodiment of the present application provides a large model fine-tuning method, including: obtaining second preliminary output data of the large model for a first task request sample; the first task request sample is any sample in the first data sample set; based on the second preliminary output data, determining a knowledge boundary evaluation result of the large model for the first task request sample; the knowledge boundary evaluation result is used to indicate the probability that the first task request sample exceeds or does not exceed the knowledge boundary of the large model; and performing modeling fine-tuning on the large model according to the knowledge boundary evaluation results corresponding to each of the first task request samples in the first data sample set.

[0008] Third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the task response method or the large model fine-tuning method based on the large model according to any embodiment of the present application.

[0009] Fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, and characterized in that when the program is executed by a processor, the steps of the task response method or the large model fine-tuning method based on the large model according to any embodiment of the present application are implemented.

[0010] The beneficial effects of the embodiments of the present application are as follows: By introducing a decision-making mechanism based on knowledge boundary evaluation, obtaining the preliminary output data of the large model for the task request and performing knowledge boundary evaluation, it can effectively prevent the large model from relying too much on external tools when it can complete the task independently, or making a decision to call external tools in a timely manner when the task request exceeds the knowledge boundary of the model, enabling the large model to make judgments according to the actual needs of the task, no longer blindly relying on tools or being overconfident, and enhancing the robustness of the large model in different environments and task scenarios. Description of the Drawings

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 Shows a flowchart of an example of the task response method based on the large model according to the embodiment of the present application; Figure 2Shows a flowchart of an example of the large model fine-tuning method according to an embodiment of the present application; Figure 3 Shows a schematic diagram of comparison effects of an example of comparing the automatic tool invocation method with the efficient tool invocation method according to an embodiment of the present application; Figure 4 Shows a flowchart of an example of modeling the knowledge boundary of a large model according to an embodiment of the present application; Figure 5 Shows a schematic diagram of simulation effects of an example of the relationship between the SFT data ratio and the comprehensive ratio of overconfidence and over-reliance on tools; Figure 6 Shows a schematic diagram of simulation effects of an example of the trade-off relationship between inference time and performance; Figure 7 Shows a schematic diagram of simulation effects of an example of how the data ratio affects the total utility of a large model; Figure 8 Shows a schematic diagram of simulation effects of an example of comparing tool invocation strategies between explicit modeling and an uncertainty-based benchmark method; Figure 9 Shows a schematic diagram of simulation effects of an example of tool usage at different accuracy levels; Figure 10 Is a schematic structural diagram of an embodiment of the electronic device of the present application. Detailed implementation manners

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0014] It should be noted that the current progress of tool learning in related research enables large language models to integrate external tools, thereby expanding their knowledge boundaries and enhancing their task performance in various downstream tasks. However, relying on tools often introduces trade-offs among performance, speed, and cost, and LLMs sometimes show over-reliance and overconfidence in tool usage.

[0015] In addition, solving tasks through tool calls usually requires more steps, longer completion times, and additional tool call costs. For example, in a question-and-answer scenario involving a search tool, the model must first generate a query to retrieve the tool, wait for the search results, and then process these results to generate the final answer. In contrast, a direct answer only requires generating a response, resulting in a trade-off between performance and speed. Unfortunately, recent research has shown that LLMs like O1 struggle to find a balance between the two: overthinking in simple reasoning tasks and underthinking in more complex tasks. Similar problems also occur in tool usage scenarios. Current LLMs exhibit over-reliance on tools, calling tools even when the task can be completed independently, and at the same time showing overconfidence when necessary and refusing to use tools. This inconsistency reflects the challenges faced by models like O1, weakening the model's tool intelligence and increasing the cost of task completion in real-world scenarios.

[0016] In some current related technology research, LLM alignment techniques have also been proposed. These mainly train language models to conform to user intentions, enhance the model's instruction-following ability, usefulness, harmlessness, and honesty, etc., through methods such as supervised fine-tuning, direct preference optimization, or reinforcement learning from human feedback. Through LLM alignment techniques, it is aimed to make the behavior of the language model consistent with human intentions, preferences, and values. Through different training methods, the model can better understand and follow instructions, provide more helpful, safer, and more honest answers, etc.

[0017] However, most current LLM alignment techniques focus on enhancing the model's instruction-following ability and behavior norms, etc., but pay less attention to the issue of how the model reasonably decides whether to call a tool based on its own knowledge boundaries. In addition, these alignment techniques usually assume that the model's knowledge boundaries are binary (i.e., "known" or "unknown"), while ignoring the uncertainty of the model's knowledge and the balance between the cost and benefit of tool calls. Moreover, the binary knowledge boundary assumption oversimplifies the model's cognitive process and cannot accurately reflect the model's knowledge state in actual applications.

[0018] It should be noted that in the current research on related technologies, attempts are usually made to solve these problems by optimizing tool call strategies, increasing conditional restrictions on tool calls, or further improving LLM alignment techniques. For example, more complex rules may be set to control tool calls, or attempts may be made to improve the model's decision-making ability for tool calls through more training data. However, these methods often only solve the problems to a certain extent and are difficult to fundamentally solve the irrationality of the model's tool call decisions.

[0019] It should be understood that the purpose of the above description of the current related technology is only to facilitate the public's better understanding of the inventive spirit and motivation of the present application, and is not regarded as a limitation of the present application. In addition, the technical solutions described in the above current related technology are not prior art, and may also be unpublished technical solutions, such as those under research or in the laboratory stage.

[0020] In view of the deficiencies in the above-mentioned current related technology under study, in the embodiments of the present application, it is aimed to improve how the LLM decides when and how to use external tools to complete tasks. The main challenge lies in aligning the behavior of the model with its knowledge boundaries so that it can judge when to call a tool based on its own confidence.

[0021] In view of this, Figure 1 The flowchart of an example of the large model-based task response method according to the embodiments of the present application is shown.

[0022] In the embodiments of the present application, instead of simply regarding the knowledge of the model as "known" or "unknown", the uncertainty of the model knowledge is fully considered. By identifying an "uncertain region" where the model makes a probabilistic estimate of its knowledge, a better decision-making balance between task success and tool usage cost is achieved.

[0023] As Figure 1 shown, in step S110, the first preliminary output data of the large model for the task request is obtained.

[0024] It should be understood that the types of task requests can be diverse, such as mathematical calculations, knowledge answering, translation, etc., to support responses to tasks in different fields, and no limitations should be made here.

[0025] In some embodiments, the task request can be a natural language question, instruction or query, which may come from a user, other systems or an automated process, and this request may contain structured or unstructured data. In addition, for the received task request, the large language model understands the user's needs through natural language processing technology, uses advanced natural language processing algorithms to semantically parse the text input by the user, and extracts key information and requirements. Furthermore, a preliminary output is generated, which may be a direct answer, explanation, solution or preliminary step of the reasoning process, and may include the need for tool calls. Thus, accurately understanding the user input and generating a preliminary output provides accurate information for subsequent tool call decisions, improving the response speed and accuracy of the system.

[0026] In step S120, based on the first preliminary output data, the knowledge boundary evaluation result of the large model for the task request is determined, and the knowledge boundary evaluation result is used to indicate the probability that the task request exceeds or does not exceed the knowledge boundary of the large model.

[0027] In the embodiments of the present application, different from the blind confidence of directly outputting answers or the over-reliance on directly invoking tools, by integrating the tool decision-making stage and pre-analyzing the knowledge boundaries of the output data, the knowledge boundary estimation is integrated into the tool decision-making process of the model, thereby achieving more efficient tool invocation.

[0028] In some embodiments, the large model can analyze the first preliminary output data to evaluate whether the content involved in the task request exceeds the knowledge base or reasoning ability of the model. The evaluation result will show a higher probability of "exceeding the knowledge boundary"; if the content of the task request belongs to the known field of the large model or within the scope of the model's training data, the evaluation result will show a higher probability of "not exceeding the knowledge boundary".

[0029] It should be noted that the methods for calculating the knowledge boundary evaluation results can be diverse and unrestricted. For example, by comparing the task request and the model output with the known knowledge base to evaluate whether the task request exceeds the model's knowledge boundary, or based on the complexity of the model output and the depth of logical reasoning to determine whether the model can solve the problem independently, or identify whether the task request involves multi-modal information, etc. In addition, the specific details of more novel analysis methods will be elaborated in combination with other examples below.

[0030] Through the embodiments of the present application, by evaluating the knowledge boundary, the large model can, when processing complex tasks, evaluate its capabilities in real time, accurately judge whether it has the ability to complete the task independently, or whether external tool assistance is required, and can dynamically adjust its tool usage strategy according to the complexity and difficulty of the task request, avoiding over-reliance on or wrongly rejecting tools.

[0031] In step S130, according to the knowledge boundary evaluation result, decide whether to invoke a matching target tool to respond to the task request.

[0032] Specifically, on the one hand, if the knowledge boundary evaluation result indicates that the task request exceeds the knowledge boundary of the large model (i.e., the probability of exceeding the knowledge boundary is relatively high), the corresponding external tool will be automatically selected for invocation according to the type and requirements of the task through the tool matching mechanism, such as a calculator, a search engine, or an inference model, etc. On the other hand, if the knowledge boundary evaluation result indicates that the task request does not exceed the knowledge boundary of the model, the large model may not need to invoke an external tool and can directly give an answer or solution based on the first preliminary output data.

[0033] Through the embodiments of the present application, dynamic adjustment of the model in tool invocation decision-making is achieved, and whether to invoke a tool and which tool to invoke are flexibly selected according to the specific situation of the task and cost considerations, thereby reducing unnecessary tool invocation costs while ensuring the success rate of the task.

[0034] In some embodiments, when the system decides to invoke a tool, it sends a request to the corresponding tool and obtains the output result of the tool. Then, the output result of the tool is integrated into the final output of the model to form a complete task solution. Exemplarily, by interacting with external tools, sending accurate request information, and receiving the results returned by the tools, the output results of the tools are integrated with the output of the model itself to ensure the integrity and consistency of the final output. Thus, through effective tool invocation and result integration, the large model can make full use of the capabilities of external tools to improve the quality and efficiency of task completion.

[0035] It should be noted that the form of the output result of the large model and the user interaction method can be diversified. In some embodiments, the system outputs the final result and optimizes the tool invocation decision-making mechanism of the model according to the user feedback and the task execution result to improve the execution effect of future tasks. Exemplarily, the final result is presented to the user in a user-friendly manner, and a user feedback mechanism is established to collect the evaluation and suggestions of the user on the generated result. Finally, according to the feedback and the task execution result, the tool invocation decision-making mechanism of the model is optimized to improve the performance of the system. Thus, through user-friendly result output and feedback mechanism, the user experience is improved, and at the same time, the model is continuously optimized according to the user feedback and the task execution result to ensure that the system can adapt to the changing user needs and improve the overall performance and reliability.

[0036] Through the embodiments of the present application, by combining the preliminary output of the large model with the knowledge boundary evaluation, the system can effectively decide whether tool assistance is needed, avoiding the situation of blindly invoking tools or wrongly rejecting tools. Through accurate knowledge boundary evaluation, the situation of the large model over-invoking external tools when it can complete tasks independently is reduced, thus effectively reducing unnecessary computational overhead and time delay. In addition, according to different task requirements, task complexities, and knowledge boundary evaluation results, whether to invoke tools is flexibly adjusted, enabling the system to have a stronger intelligent response ability, adapt to diverse task scenarios, and improve the performance of the system in complex task scenarios.

[0037] Regarding the details of the knowledge boundary evaluation and analysis in step S120, in some examples of the embodiments of the present application, it can be achieved through consistency estimation. Specifically, the first preliminary output data includes the output feedback information of multiple input samplings of the large model for the task request. Calculate the output result consistency according to each output feedback information in the first preliminary output data. For example, calculate the vector variance corresponding to each output feedback information as the output result consistency, and then determine the knowledge boundary evaluation result of the large model for the task request according to the output result consistency. Thus, by sampling the output of the model multiple times and calculating the degree of consistency of the output results, the higher the consistency, the more dynamically the model can judge its confidence in the task, avoiding blind confidence or over-reliance on external tools. For example, if the output is highly consistent, the model can confidently handle the task; if the consistency is low, the model can recognize its limitations and thus avoid misjudgment.

[0038] As a further optimized implementation method, it can also improve the accuracy of the knowledge boundary evaluation result by integrating the absolute estimation method. In the absolute estimation method, the output of the model is compared with external real data to calculate the accuracy of the model output, so as to obtain the true mastery degree of the model for this task. Specifically, obtain external real data related to the task request, and calculate the model output accuracy corresponding to the first preliminary output data according to the external real data. Here, the sources of external real data can be diverse, such as domain expert databases, open knowledge bases, or task data sets with answer annotations, etc. Verify the first preliminary output data through external real data to obtain the true mastery degree of the model's knowledge of this task request. Then, determine the knowledge boundary evaluation result of the large model for the task request according to the model output accuracy and the output result consistency. In one example, the model output accuracy and the output result consistency are fused to determine the corresponding knowledge boundary evaluation result. In another example, the model output accuracy and the output result consistency are analyzed separately, and the two analysis results are synthesized to obtain the knowledge boundary evaluation result of the large model for the task request.

[0039] In the embodiments of the present application, the knowledge boundary of the model is estimated by two methods: consistency estimation and absolute estimation. Consistency estimation judges the confidence of the model by evaluating the output consistency of multiple samplings of the same task by the model; absolute estimation evaluates the accuracy of the model by comparing it with external real data. By introducing the absolute estimation method and comparing the output of the large model with external real data, the degree of the model's mastery of the task can be measured more precisely, avoiding the bias that may be brought by relying only on internal consistency estimation. Especially when facing more complex or difficult tasks, it can provide a more accurate evaluation. Thus, the model can more accurately evaluate its own mastery of the task, providing a more reliable basis for subsequent tool call decisions.

[0040] Regarding the implementation details of step S130, in some examples of the embodiments of the present application, the knowledge boundary evaluation result is compared with a preset first probability threshold, and whether to call a matching target tool to respond to the task request is decided according to the comparison result. In this way, according to the result of the knowledge boundary estimation, the system dynamically decides whether to call a tool and which tool to call. If the confidence of the model in the task is low and the estimated accuracy is not high, it tends to call a tool; if the confidence is high and the accuracy is also high, it gives priority to completing the task independently.

[0041] Specifically, according to different application scenarios and cost sensitivities, corresponding decision rules are formulated, such as setting thresholds for confidence and accuracy, and calling a tool when the estimated result of the model is lower than the threshold. In the case of deciding to call a tool, the most suitable tool is selected according to the type and requirements of the task, such as a calculator, a search engine, or an inference model, etc. Thus, the dynamic adjustment of the model in tool call decision-making is realized, and whether to call a tool and which tool to call are flexibly selected according to the specific situation of the task and cost considerations, so as to reduce unnecessary tool call costs while ensuring the success rate of the task.

[0042] Through the embodiments of the present application, the multi-object alignment framework and algorithm solve the deficiencies of the current related technologies in the tool call decision-making of large language models by combining probability knowledge boundary estimation and dynamic decision-making. First, through the probability knowledge boundary estimation method (including consistency estimation and absolute estimation), the large model can more accurately evaluate its own mastery of the task, so as to more reasonably balance performance, speed, and cost when calling a tool. Second, through dynamic decision-making, the large model can flexibly adjust the tool call strategy according to different application scenarios and cost sensitivities. Thus, the tool call behavior of the model can be more finely controlled, and an optimized response to the task that comprehensively considers the knowledge uncertainty of the model and the tool call cost is realized.

[0043] Figure 2 The flowchart of an example of the large model fine-tuning method according to the embodiments of the present application is shown.

[0044] As Figure 2 shown, in step S210, second preliminary output data of the large model for the first task request sample is obtained, and the task request sample is any sample in the first data sample set.

[0045] In some embodiments, the first data sample set can be for specific tasks or diverse tasks, such as classification, regression, question answering, translation, etc., and each sample corresponds to a specific task request (for example, a piece of text or a question), which is used to test the large model's mastery of the task knowledge.

[0046] In step S220, based on the second preliminary output data, determine the knowledge boundary evaluation result of the large model for the first task request sample, where the knowledge boundary evaluation result is used to indicate the probability that the first task request sample exceeds or does not exceed the knowledge boundary of the large model.

[0047] Regarding the relevant details of the knowledge boundary evaluation result, reference can be made to the description in combination with other examples in the above text, and thus it will not be elaborated here.

[0048] In step S230, according to the knowledge boundary evaluation results corresponding to each first task request sample in the first data sample set, perform modeling fine-tuning on the large model.

[0049] In an example of the embodiment of the present application, samples whose evaluation results indicate that the task request does not exceed the knowledge boundary can be selected and used as training samples to perform fine-tuning training on the large model, enhancing the performance of the large model when directly outputting answers for the corresponding task. In another example of the embodiment of the present application, samples whose evaluation results indicate that the task request exceeds the knowledge boundary can be selected and used as training samples, and fine-tuning can be performed by invoking the knowledge of external tools as auxiliary information, which helps the large model to recognize its knowledge boundary.

[0050] Through the embodiment of the present application, by judging whether each task request sample in the data sample set exceeds the knowledge boundary of the large model, the model can better understand its own ability range in actual applications. Through the fine-tuning process, it can dynamically adjust its tool decision-making behavior according to the different complexities of the task requests. The fine-tuned model can more effectively identify when external tools are needed, reducing unnecessary computational overhead. When the task does not exceed the knowledge boundary, the model can directly return the result to improve efficiency, and only when the task exceeds the knowledge boundary, external tools are called to ensure the response effect.

[0051] Regarding the implementation details of step S230, in some examples of the embodiment of the present application, the knowledge boundary of the large model can be modeled through an explicit modeling method. Specifically, the knowledge boundary evaluation result can be regarded as a knowledge confidence score, reflecting the degree of mastery of the large model for the current task. For example, a higher confidence score means that the model believes that the task does not exceed its knowledge boundary, and vice versa means that the task may exceed the understanding range of the model. Specifically, for each first task request sample in the first data sample set, compare the knowledge boundary evaluation result corresponding to the first task request sample with a preset second probability threshold, and configure a corresponding tool call label for the first task request sample according to the comparison result, where the tool call label is used to indicate whether the large model calls a tool or directly answers. Then, perform modeling fine-tuning on the large model according to each first task request sample configured with a tool call label.

[0052] Through the embodiments of the present application, for samples with high confidence scores and tasks within the knowledge boundary, the large model should learn how to improve the prediction accuracy for these tasks. For samples with low confidence scores and tasks beyond the knowledge boundary, the large model should learn how to correctly select appropriate external tools to assist in task completion. In addition, through confidence comparison, it can be better applied in various tasks. When facing different task requests, the large model can automatically determine whether to call tools based on the difficulty of the task (i.e., the confidence score), enabling the large model to more intelligently adapt to the requirements of various tasks.

[0053] As a further optimization of the embodiments of the present application, the explicit modeling method and the implicit modeling method can also be combined to model the knowledge boundary of the large model, which is particularly helpful for the large model to learn and optimize the knowledge boundary of the target task. In the implicit modeling method, the large model directly outputs actions (i.e., directly answers or calls tools) according to predefined decision rules. Specifically, a second data sample set is obtained, and each second task request sample in the second data sample set has an estimated knowledge score for the target task. Compared with the first data sample set, the task request samples in the second data sample set focus more on the target task, that is, knowledge boundary evaluation is needed to help the large model determine how to efficiently complete the target task, and the estimated knowledge score corresponding to each task sample can effectively indicate the degree of mastery of the task request sample within the knowledge scope of the target task. It should be understood that the calculation method of the estimated knowledge score can be diversified, and an estimated knowledge score can be assigned to each task request sample through preprocessing and analysis of the task samples in the second data sample set, such as preliminary judgment of the task samples, comparison of historical data and external data, etc. Then, according to the estimated knowledge score, a corresponding task tool call label is configured for the second task request sample, and the task tool call label is used to indicate that the large model calls tools or directly answers under the target task. Furthermore, based on each second task request sample and each first task request sample configured with a tool call label, fine-tuning of the large model in terms of the target task is performed.

[0054] It should be noted that implicit modeling has a faster inference speed (single response generation), but it needs to be trained separately for the target tasks of different task types to enhance the directional fine-tuning for different tasks to ensure the accuracy of model responses. In addition, explicit modeling has higher flexibility, does not need to be trained separately for different types of tasks, and supports multi-task optimization learning.

[0055] In the embodiments of the present application, by combining explicit modeling and implicit modeling, the learning ability of the large model for the knowledge boundary of the target task is enhanced. By estimating the configuration of the knowledge score and the task tool call label, the large model can accurately determine whether a task exceeds its knowledge boundary during the multi-task learning process, thereby optimizing the knowledge mastery and task completion strategies.

[0056] In some examples of the embodiments of the present application, according to the response feedback results of the large model for each first data sample, the overall response accuracy and the tool usage ratio are statistically calculated. Specifically, for each first data sample, the large model generates a preliminary response according to the task request, which may be a direct answer or the result of completing the task through an external tool, and the response feedback results of the model on each task request can be recorded, such as whether the result returned by the model is correct and whether an external tool is called, so as to calculate the corresponding overall response accuracy and tool usage ratio.

[0057] The overall response accuracy and the tool usage ratio are weighted and summed to determine the total utility value. In order to find a balance between the call efficiency (tool usage ratio) and the response accuracy, these two metrics are weighted and summed. By assigning weights to the accuracy and the tool usage ratio, the priority of the model's performance in different tasks can be adjusted according to the importance and requirements of the tasks.

[0058] As a further preferred implementation, considering that the call costs corresponding to different tools are different, for example, the call cost corresponding to a calculator should be less than that of a search engine, and different weighting coefficients can be set for different types of tasks in the system to meet the personalized needs of different tool learning scenarios. Specifically, the weighting coefficient corresponding to the target task is obtained, and the overall response accuracy and the tool usage ratio are weighted and summed according to the obtained weighting coefficient to determine the total utility value. Thus, by assigning personalized weighting coefficients to each task, the system can learn the differential call costs between different tools and can make flexible decisions according to the costs of the tools and the task complexity during the task execution process.

[0059] Furthermore, the large model is fine-tuned according to the total utility value to optimize the balance between the call efficiency and the response accuracy of the large model. Exemplarily, the total utility value is used to evaluate whether the fine-tuning of the large model converges, optimizing the balance between the task accuracy and the tool call efficiency of the model, and enabling it to maintain a high success rate in various tasks.

[0060] In the embodiments of the present application, by optimizing the tool invocation decision-making mechanism of the large language model, the performance and efficiency of the model in various tasks can be directly improved, and the unnecessary tool invocation cost can be reduced. Specifically, the model can more reasonably determine when to invoke a tool and which tool to invoke, thereby reducing the waste of computing resources while ensuring the success rate of the task. The chain reactions that can be brought about by this effect include improving the user experience. The optimization of the model in tool invocation enables users to obtain more efficient and accurate services, thereby enhancing the user's satisfaction and trust in the system; reducing the operating cost. Reducing unnecessary tool invocations can reduce the computing cost and time delay of the system, thereby reducing the operating cost of the enterprise and increasing the commercial value of the system; promoting the wide application of the model. A more efficient tool invocation decision-making mechanism enables the large language model to better adapt to various complex tasks and application scenarios, thereby promoting its wide application in more fields and driving the development and popularization of AI technology; driving technological innovation. This solution provides a new idea and method for the tool invocation decision of the large language model, and is expected to stimulate more technological innovation and research, further improving the performance and intelligent level of the AI system.

[0061] In some examples of the embodiments of the present application, an alignment framework and algorithm for efficient tool invocation of large models are also proposed.

[0062] 1. Introduction Aiming at the alignment problem of the knowledge boundary of the LLM, it aims to make more intelligent decisions when invoking tools. In this paper, a multi-objective alignment framework is proposed, which combines probabilistic knowledge boundary estimation and dynamic decision-making, enabling the LLM to better evaluate when to invoke tools based on its own confidence. This framework includes two knowledge boundary estimation methods - consistency-based estimation and absolute estimation, as well as two training strategies to incorporate these estimation results into the model's decision-making process. Experimental results show that in various tool invocation scenarios, this framework can effectively improve the tool usage efficiency and significantly reduce unnecessary tool invocations.

[0063] By introducing an alignment framework for efficient tool invocation, which combines probabilistic knowledge boundary estimation and dynamic decision-making, mainly including two main components: 1) Knowledge boundary estimation. Specifically, it includes two methods for evaluating the knowledge boundary of the model: using the consistency-based estimation method to sample the output of the model multiple times, calculating the degree of consistency of the output results, and fusing and utilizing external real data through absolute estimation to evaluate the accuracy of the model output.

[0064] 2) Knowledge boundary modeling: Different data is constructed to demonstrate implicit modeling, where the model makes decisions based on a predefined knowledge confidence threshold, and explicit modeling, where the model outputs an answer and a confidence score. This framework helps the model use tools more efficiently, calling tools only when necessary, thus improving performance and reducing costs. It has been evaluated in multiple tool usage scenarios, and the experimental results show a significant reduction in unnecessary tool calls and an improvement in the overall efficiency of the tools.

[0065] Figure 3 The schematic diagram of the comparison effect showing an example of comparing the automatic tool call method with the efficient tool call method according to the embodiments of the present application is shown.

[0066] As Figure 3 shown, the method herein can effectively enable the LLM to switch between independent answering and tool calling (upper part), thereby reducing the over-reliance and over-confidence of the model on tools (lower part).

[0067] Through the multi-objective alignment framework proposed herein, efficient tool calling can be achieved, corresponding evaluation metrics are provided, a tool alignment algorithm and corresponding data generation methods are proposed, and extensive experiments have been conducted in multiple tool calling scenarios to verify the effectiveness of the method herein.

[0068] 2 Related Work 2.1 LLM Alignment LLM alignment aims to train language models to conform to user intentions, and the methods adopted include supervised fine-tuning, direct preference optimization (DPO), or reinforcement learning from human feedback (RLHF). Most work focuses on improving aspects such as the instruction-following ability, usefulness, harmlessness, and honesty of LLMs. In addition, some studies have proposed aligning the model with its knowledge boundary by training the LLM to reject unknown questions. However, these methods assume that the knowledge boundary of the model is binary - that is, the model either knows the answer or does not. Different from this, the work herein believes that the knowledge boundary is more complex, and there is a gray area where the behavior of the model in this ambiguous area can be dynamically determined according to specific application scenarios.

[0069] 2.2 Tool Learning Recent progress in tool learning has enabled LLMs to effectively integrate external tools, enhancing real-time knowledge retrieval, multimodal capabilities, and domain-specific knowledge. Methods include using in-context learning for tool description and demonstration and explicit training on tool-augmented datasets. Some studies have also explored how to complete tasks within a limited number of tool calls and how to call tools more reliably. However, previous research on tool calls has mostly ignored the relationship between tool use and the model's knowledge boundary. In addition, no unified evaluation metrics have been proposed to evaluate efficient tool calls.

[0070] 3 Problem Description 3.1 LLM Alignment With the rapid development of large language models, ensuring their alignment with human instructions, preferences, and values has become a key research area. Alignment methods aim to optimize model responses based on predefined objectives such as usefulness, truthfulness, and safety. Specifically, given an input prompt x i and the alignment objective - usefulness, the following scoring principle is adopted to represent the alignment objective: , Equation (1) where, and represent useful and useless responses respectively. The preference order can be determined by human annotation or a scoring model trained using human preference data. The collected preference data can be further used to train a reward model or fine-tune the LLM policy, thus improving the alignment with human expectations.

[0071] 3.2 Multi-Objective Alignment for Efficient Tool Calls Although alignment with usefulness is crucial, efficient tool calls pose additional alignment challenges. A well-aligned LLM should not only provide useful responses but also minimize unnecessary tool use, as excessive tool calls increase inference latency and computational costs. Therefore, a multi-objective alignment framework is proposed to balance usefulness and tool cost.

[0072] First, the alignment objectives for usefulness and tool cost are defined separately. For the alignment objective of usefulness, it is defined as follows: , Equation (2) where, represents the correct response, represents the wrong response. At the same time, for tool cost, it is defined as: , Equation (3) where, represents the response without using tools, Represents the response using the tool. Combining these two objectives, the final alignment formula is: , Equation (4) where respectively represent the correct response without using the tool, the correct response using the tool, the incorrect response without using the tool, and the incorrect response using the tool. This ordering reflects that an ideal LLM should solve problems as independently as possible, only calling the tool when necessary, while avoiding wrong answers and unnecessary tool calls.

[0073] 3.3 Evaluation Method for Efficient Tool Invocation To quantify the trade - off between usefulness and tool cost, a benefit - cost utility function is defined as follows: , Equation (5) where and respectively represent that when the response is the correct answer or contains a tool call, its value is 1, represents the cost associated with tool use. Then, the overall utility of the model on a dataset containing N samples is calculated as: , Equation (6) where Acc and TR represent the overall accuracy and the tool use ratio on the dataset respectively.

[0074] The parameter is very crucial because it determines the relative penalty for tool use. If is large, it means being more sensitive to cost and having a greater penalty for calling the tool. If is set too high, the model may completely avoid using the tool when necessary; conversely, if is set too low, the model may over - use the tool. Therefore, choosing a moderate value can ensure a balance between efficiency and effectiveness.

[0075] In addition, the cost of tool use varies between different tasks and tools. To account for these differences, can be dynamically adjusted according to the specific tool used. In this study, values of 0.2, 0.4, and 0.6 are assigned to the calculator, search engine, and external LLM inference respectively for , and these different values reflect the increased computational cost and inference latency brought by these tools.

[0076] 4 Methodology 4.1 Efficient Tool Learning Framework Figure 4 The flowchart shows an example of modeling the knowledge boundary of a large model according to an embodiment of the present application.

[0077] The key to achieving efficient tool invocation lies in aligning the LLM with its own knowledge boundary. Different from the way of simply dichotomizing knowledge into "known" and "unknown", human cognition - and thus also the LLM - operates based on a continuous spectrum. As Figure 4 shown on the left, there is a large "uncertainty region" within which the model can only probabilistically estimate its knowledge. Some previous methods that enforced strict binary classification failed to capture this nuanced understanding, leading to inaccurate estimates and suboptimal tool invocation strategies.

[0078] To enable effective tool use, the model must first understand its own knowledge boundary and then use this understanding to adjust its decision-making process. This perspective is consistent with the efficiency goal discussed earlier: a model that dichotomizes knowledge will have difficulty adjusting its behavior under different cost considerations (represented by ). If the model only classifies knowledge as "known" or "unknown", it will either always invoke a tool in uncertain situations or always directly answer, ignoring cost-sensitive optimization.

[0079] Here, a solution is proposed to let the model learn to estimate the uncertainty of its knowledge in a probabilistic manner rather than performing binary classification, which will provide greater flexibility for tool invocation. According to different values (representing different real-world tool costs), the model can be trained to dynamically adjust its behavior. This can be achieved implicitly through a controlled training data distribution or by having the model output confidence estimates and performing thresholding during inference to decide whether to invoke a tool.

[0080] 4.2 Knowledge Boundary Estimation Here, two knowledge boundary estimation methods are proposed, as Figure 4 shown in the middle: Consistency-Based Estimation This method relies on self-consistency. If the model's outputs for a given question are highly consistent across multiple samples, it indicates a stronger grasp of the underlying knowledge. To achieve this, the variance of the model's sampled responses is measured and used as an indicator of knowledge certainty. Higher consistency means the model has more confidence in its knowledge.

[0081] Absolute Estimation Based on Ground Truth Although the consistency-based estimation method is useful, it does not directly utilize external validation. To address this issue, an absolute estimation method based on the correctness of the ground truth is introduced. The model responses are sampled multiple times for the same problem, and the average accuracy is calculated using the ground truth, which provides a measure of model knowledge for external validation and corrects potential biases in self-estimation.

[0082] 4.3 Training Method To incorporate knowledge boundary estimation into the behavior of the model, two SFT strategies are adopted, as Figure 4 shown on the right: implicit modeling and explicit modeling.

[0083] Implicit Modeling In this method, the model directly outputs an action (i.e., a direct answer or calls a tool) according to predefined decision rules. Specifically, all training samples are sorted based on the estimated knowledge scores of the samples, and a threshold is set: samples exceeding the threshold are labeled as direct answers, while samples below the threshold are labeled as calling tools. Since different values correspond to different tool usage preferences, separate SFT models are trained for different thresholds to adapt to different scenarios. This method is efficient during inference because the model only needs to generate one response for each query. However, it requires multiple rounds of training to adapt to different values.

[0084] Explicit Modeling Different from implicit modeling, explicit modeling trains the model to output both an answer and a related knowledge confidence score simultaneously. This allows for dynamic adjustment of the tool call decision during inference without training separate SFT models for different values. During the inference process, a threshold is set for the confidence score: if the confidence score exceeds the threshold, the model answers directly; otherwise, the model calls a tool. This method does not require retraining, but it introduces additional inference latency because both an answer and an uncertainty estimate need to be generated for each query.

[0085] Each method has its advantages and disadvantages. Implicit modeling has a faster inference speed (single response generation), but requires multiple trainings to adapt to different values. Explicit modeling is more flexible during inference (threshold adjustment without retraining), but has a slower inference speed due to the need for a two-step generation process. In the experiments of this paper, both methods were evaluated to determine the best strategy for efficient tool calls.

[0086] 5 Experiments 5.1 Experimental Setup 5.1.1 Task Scenario The method in this paper was evaluated in three scenarios, each of which required the use of specific external tools: symbolic calculation through a calculator, fact retrieval using a retrieval-augmented generation system, and complex reasoning using a strong inference model.

[0087] Arithmetic calculation (calculator) To evaluate numerical reasoning ability, an arithmetic dataset was constructed with input numbers sampled from a logarithmic scale to ensure different orders of magnitude and avoid duplication as much as possible. This paper combined hundreds of instruction templates generated by ChatGPT to increase language diversity. The calculation was performed through a symbolic calculator that achieved mathematical evaluation through code execution.

[0088] Knowledge-based question answering (retrieval-augmented generation) To evaluate the retrieval ability of factual knowledge, this paper used the TriviaQA dataset, a widely used question answering dataset. 10,000 instances were sampled for training and the development set of 11,313 instances was used for evaluation because the true labels of the official test set were not available. To enhance factual accuracy, a retrieval system was integrated using Pyserini, a reproducible information retrieval Python toolkit for sparse and dense representations.

[0089] Complex reasoning (reasoning model) To evaluate multi-step reasoning tasks, the MATH dataset was used. The original training-test split was used. Given the inherent complexity of mathematical reasoning, DeepSeek-R1 was used, which provided strong reasoning ability. However, it also brought a trade-off: higher computational cost and slower reasoning speed.

[0090] 5.1.2 Baseline methods Baseline methods are divided into two major categories: prompt-based methods and uncertainty-based methods.

[0091] Prompt-based methods Prompt-based methods determine how the model interacts with external tools and determines its tool-using behavior.

[0092] - Baseline (w / o tool): This method allows the model to answer questions relying entirely on its own internal knowledge without using any tools.

[0093] - Baseline (all tool): This method forces the model to always call tools.

[0094] - Auto tool: This method allows the model to decide when to use tools based on its estimated confidence.

[0095] - ICL tool (10-shot): This method provides the model with 10 example interactions (5 correct and 5 incorrect) to better guide the model in deciding whether to answer directly or use the tool.

[0096] Uncertainty-based methods Uncertainty-based methods estimate the confidence of the model's generated answers, use these confidences to determine the best utility, and optimize tool calls by searching for the best confidence threshold. Four methods are explored: - Raw logits - P(True) - Verbalized Confidence - Agreement (Self-Consistency) Each method provides a different way to evaluate the model's confidence.

[0097] 5.1.3 Training details Two baseline models were used: LLAMA-3.1-8B-INSTRUCT and QWEN-2.5-7B-INSTRUCT. To be consistent with the experimental settings in this paper, the DeepSpeed-Chat framework was customized. The learning rate used during training was 5×10 -5 , the batch size was 128, and other training parameters were set to the default values of DeepSpeed-Chat. By default, 10,000 samples were used for supervised fine-tuning, and all models were trained for 2 rounds on A800 GPUs.

[0098] 5.2 Main results Table 1: Performance comparison under three tool call scenarios. Utility is a comprehensive evaluation metric of accuracy and tool usage rate. A larger value indicates higher cost sensitivity and a greater penalty for calling the tool.

[0099]

[0100] Table 1 compares the performance of all evaluation methods. The method in this paper achieved the highest utility scores on all three datasets, demonstrating its effectiveness in balancing task success and tool efficiency. In the method of this paper, the performance based on absolute knowledge boundary estimation is better than that based on consistency estimation because the external supervision by true labels makes the boundary estimation more accurate, thus making better tool call decisions. The proposed scheme in this paper reduces the tool usage by nearly 50% while maintaining an accuracy comparable to the best model, significantly reducing the dependence on external tools and the computational cost compared with the fully automated baseline method. The training-based method further improves the efficiency, achieving better performance than Auto Tool while reducing the tool usage. This validates the effectiveness of optimizing tool calls by aligning the internal knowledge boundaries of the model.

[0101] 5.3 Overconfidence and Over-reliance on Tools Specifically, it analyzes how explicit modeling affects the trade-off between overconfidence and over-reliance on tools by adjusting the SFT data ratio. The SFT data ratio represents the proportion of training samples containing tool calls. As this ratio increases, the confidence estimation of the model and the dependence on external tools change. Figure 5 Fig. shows an example of the effect simulation diagram of the relationship between the SFT data ratio and the combined ratio of overconfidence and over-reliance on tools, which expresses the trade-off between overconfidence and over-reliance on tools under different SFT data ratios. A higher tool usage rate increases the dependence on external tools, leading to over-invocation of tools, while overconfidence gradually decreases as the model relies more on tools. Each dataset has an optimal SFT data ratio at which the combined ratio of overconfidence and over-reliance on tools is minimized, achieving a balance between model confidence and tool dependence. Figure 5 This turning point in provides guidance for optimal model selection. At this ratio, the model maintains well-calibrated knowledge boundaries while minimizing unnecessary tool usage, ensuring both efficiency and accuracy.

[0102] 5.4 Inference Time Since tool calls increase the computational overhead, the inference cost is evaluated by measuring the actual execution time. Using VLLM on NVIDIA A800 GPUs, the inference time for each sample was calculated and the total running time on the dataset was aggregated. Figure 6 Fig. shows an example of the effect simulation diagram of the trade-off relationship between inference time and performance. In Figure 6The location method in the upper left corner area achieves a more favorable balance. The method in this paper always performs superiorly in terms of efficiency, being able to achieve higher performance under the same inference time or reduce latency while maintaining accuracy. By optimizing tool usage, the method in this paper minimizes costs without sacrificing performance, ensuring efficient practical deployment and being applicable to practical applications.

[0103] 5.5 Ablation Experiments 5.5.1 Implicit Modeling Method To understand how implicit modeling affects the results, this paper conducted an ablation study to observe how different supervised fine-tuning (SFT) data ratios affect the behavior of the model. The data ratio refers to the proportion of training samples in which the model uses tools rather than answering questions independently. By keeping the total size of the dataset constant but changing this ratio and observing how it affects the model's preference for using tools or answering independently, it helps to find the optimal balance based on cost. When the cost of using tools is low, a higher ratio makes the model use tools more frequently, thereby improving accuracy by using external resources. On the other hand, if the cost of tool usage is high, a lower ratio makes the model more inclined to answer questions independently, thus reducing costs. The key is to find the appropriate balance so that the model can efficiently decide when to use tools according to the situation. Figure 7 A simulation effect schematic diagram showing an example of how the data ratio affects the total utility of the large model is presented, which shows the impact of the SFT data ratio on utility. The ratio represents the proportion of training samples in which the model calls tools rather than answering directly. Initially, as the ratio increases, the utility increases and then decreases after reaching the peak. The optimal ratio varies depending on the dataset and depends on the cost of the tools. If the tool cost is high, the optimal ratio is low. This indicates that the implicit modeling method helps the model make informed choices based on task costs and achieve a balance between accuracy and efficiency.

[0104] 5.5.2 Explicit Modeling Method Different from the implicit method, explicit modeling enables the model to directly output confidence scores and use them together with the predictions to make tool invocation decisions based on thresholds. To further evaluate its effectiveness, explicit modeling was compared with the uncertainty-based method because both methods essentially rely on confidence estimation to determine the knowledge boundary. To ensure a fair comparison, the confidence threshold was adjusted to control the tool invocation ratio, and the performance of the model at different tool usage levels was evaluated by systematically varying the threshold. Figure 8 A simulation effect schematic diagram showing an example of the comparison of tool invocation strategies between explicit modeling and the uncertainty-based benchmark method is presented. As Figure 8As shown, the relationship between the tool invocation rate and model performance is presented. Explicit modeling consistently outperforms the uncertainty-based baseline method at all invocation ratios, demonstrating its ability to provide more reliable knowledge boundary estimates. The performance gap remains stable, highlighting the robustness of explicit confidence modeling. By leveraging these confidence scores, the method in this paper can more finely control tool invocations, optimizing task success while reducing unnecessary computational overhead.

[0105] 5.6 Aligned Knowledge Boundary Distribution To examine whether the model has learned the knowledge boundary, it is compared with the auto_tool method through the tool invocation distribution. Figure 9 A schematic diagram of the simulation effect showing an example of tool usage at different accuracy levels is presented. Higher accuracy reflects a better understanding of the problem. An ideal model should rely on tools in difficult situations while minimizing tool usage for confidently answering questions. However, auto_tool shows an almost uniform tool invocation pattern, indicating its lack of awareness of the knowledge boundary. In contrast, the method in this paper shows a gradual decrease in tool usage as the accuracy increases, indicating adaptive tool invocation based on knowledge confidence. In this paper, the situation of over-reliance on tools is also analyzed, that is, the model unnecessarily uses tools even when it can correctly answer questions. Figure 7 shows that the baseline method exhibits a gradually increasing phenomenon of over-reliance on tools as the accuracy increases, resulting in unnecessary computational overhead. On the contrary, the method in this paper reduces the situation of over-reliance on tools, achieving a more intelligent invocation.

[0106] It should be noted that in the alignment framework for efficient tool invocation proposed in this paper, it is evaluated through experiments on three datasets. On the one hand, the number of tools used in the experiments is limited, and three representative tools are selected: a math calculator, a search engine, and an external large model. The motivation for this selection is that most tools have highly specific knowledge. For example, a tool for retrieving weather information for a certain day contains knowledge that the model does not possess, so the model needs to invoke this tool to complete the task. On the other hand, different models and knowledge sources can also be regarded as tools, which means that the discussion on modeling knowledge boundaries in this study still has high value.

[0107] 6 Conclusion In the research of this paper, a novel method is proposed to improve the decision-making process of LLMs on when and how to use external tools. By introducing the concept of "uncertainty region" and probabilistic knowledge boundary estimation, the framework of this paper achieves more informed and efficient tool use. Through a large number of experiments, it is proved that this method can reduce unnecessary tool calls, thereby improving performance and cost-effectiveness. By combining implicit and explicit modeling techniques, greater real-time decision-making flexibility is provided for the model, advancing the tool intelligence of LLMs and making tool calls more cautious and efficient. Future work can further explore improvements and broader applications.

[0108] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0109] In some embodiments, the embodiments of this application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to be used to execute any one of the methods for training a large language model in the above of this application.

[0110] In some embodiments, the embodiments of this application also provide a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is made to execute any one of the methods for training a large language model described above.

[0111] In some embodiments, the embodiments of this application also provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for training a large language model.

[0112] Figure 10 is a schematic hardware structure diagram of an electronic device for executing the method for training a large language model provided by another embodiment of this application, as Figure 10 shown, the device includes: One or more processors 1010 and a memory 1020, Figure 10 Taking one processor 1010 as an example.

[0113] The device for executing the method for training a large language model may further include: an input device 1030 and an output device 1040.

[0114] The processor 1010, the memory 1020, the input device 1030, and the output device 1040 may be connected by a bus or other means, Figure 10 Taking connection by a bus as an example.

[0115] The memory 1020, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for training a large language model in the embodiments of the present application. The processor 1010 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 1020, that is, implements the method for training a large language model in the above method embodiments.

[0116] The memory 1020 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 1020 may optionally include a memory remotely provided relative to the processor 1010, and these remote memories may be connected to the electronic device through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0117] The input device 1030 may receive input digital or character information, and generate signals related to the user settings and function controls of the electronic device. The output device 1040 may include a display device such as a display screen.

[0118] The one or more modules are stored in the memory 1020, and when executed by the one or more processors 1010, execute the method for training a large language model in any of the above method embodiments.

[0119] The above product can execute the method provided in the embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.

[0120] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0121] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.

[0122] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.

[0123] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or rather the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A task response method based on a large model, comprising: Obtaining first preliminary output data of the large model in response to the task request; Based on the first preliminary output data, determining a knowledge boundary evaluation result of the large model for the task request; The knowledge boundary evaluation result is used to indicate the probability that the task request exceeds or does not exceed the knowledge boundary of the large model; According to the knowledge boundary evaluation result, a decision is made whether to call a matching target tool to respond to the task request.

2. The method according to claim 1, wherein: The first preliminary output data includes output feedback information of the large model for multiple input samples requested by the task. The step of determining, based on the first preliminary output data, a knowledge boundary evaluation result of the large model for the task request includes: Calculating the consistency of the output result according to each piece of output feedback information in the first preliminary output data; According to the consistency of the output results, a knowledge boundary evaluation result of the large model for the task request is determined.

3. The method according to claim 1, wherein: Determining the knowledge boundary evaluation result of the large model for the task request according to the consistency of the output result includes: Acquire external real data related to the task request, and calculate the model output accuracy corresponding to the first preliminary output data according to the external real data; The knowledge boundary evaluation result of the large model for the task request is determined according to the model output accuracy and the output result consistency.

4. The method according to claim 1, wherein: The step of deciding whether to call a matching target tool to respond to the task request according to the knowledge boundary evaluation result includes: The knowledge boundary evaluation result is compared with a preset first probability threshold, and a decision is made based on the comparison result whether to call a matching target tool to respond to the task request.

5. A large model fine-tuning method, comprising: Obtain second preliminary output data of the large model for the first task request sample; The first task request sample is any sample in the first data sample set; Based on the second preliminary output data, determining a knowledge boundary evaluation result of the large model for the first task request sample; The knowledge boundary evaluation result is used to indicate the probability that the first task request sample exceeds or does not exceed the knowledge boundary of the large model; The large model is fine-tuned according to the knowledge boundary evaluation results corresponding to each of the first task request samples in the first data sample set.

6. The method according to claim 5, wherein: The step of fine-tuning the large model according to the knowledge boundary evaluation results corresponding to each of the first task request samples in the first data sample set includes: For each first task request sample in the first data sample set, compare the knowledge boundary evaluation result corresponding to the first task request sample with a preset second probability threshold, and configure a corresponding tool call tag for the first task request sample according to the comparison result; the tool call tag is used to instruct the large model to call the tool or directly answer; The large model is fine-tuned according to each first task request sample configured with a tool call tag.

7. The method according to claim 6, further comprising: Acquire a second data sample set, wherein each second task request sample in the second data sample set has an estimated knowledge score for the target task; According to the estimated knowledge score, a corresponding task tool calling tag is configured for the second task request sample; the task tool calling tag is used to instruct the large model to call or directly answer the tool under the target task; The step of fine-tuning the large model according to each first task request sample configured with a tool call tag includes: According to each of the second task request samples and each of the first task request samples configured with a tool call tag, the large model is fine-tuned in terms of the target task.

8. The method according to claim 5, wherein: The step of fine-tuning the large model according to the knowledge boundary evaluation results corresponding to each of the first task request samples in the first data sample set includes: According to the response feedback results of the big model for each of the first data samples, the overall response accuracy and tool usage ratio are counted; The overall accuracy of the response and the proportion of tool use are weighted and summed to determine a total utility value; The large model is fine-tuned according to the total utility value to optimize the balance between the calling efficiency and the response accuracy of the large model.

9. The method according to claim 8, wherein: The weighted sum of the overall accuracy of the response and the tool usage ratio to determine the total utility value includes: The weighting coefficient corresponding to the target task is obtained, and the overall accuracy of the response and the proportion of tool use are weighted and summed according to the obtained weighting coefficient to determine the total utility value.

10. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 9.

Citation Information

Cited By

  • Method and device for training model, storage medium and electronic equipment

    CN120523957A

  • Method, device, storage medium and electronic device for training a model

    CN120523957B

  • Intelligent agent tool selection method with self-reflecting and caching mechanism

    CN121525735A