Tool calling method and device based on large language model

Through the large language model filtering and calling tools that best match task information, a single large language model is solved, and the problem that it is difficult for a single large language model to cope with complex task requirements is realized, efficient resource utilization and dynamic task scheduling are achieved, and the adaptability and scalability of the model is improved.

CN120104216AActive Publication Date: 2025-06-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510582863.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

A single large language model is difficult to provide accurate and comprehensive answers to complex task requirements, especially when multi-platform tool resource occupancy.

Method used

By obtaining the problem text, matching the candidate tool collection in the tool library, and filtering the target tools that best match the task information based on the large language model, generating tool call instructions to achieve task execution.

Benefits of technology

It realizes efficient utilization of resources and dynamic scheduling of task tools, improves task response capabilities, improves model adaptability and scalability, and achieves the best balance of performance and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104216A_ABST
    Figure CN120104216A_ABST
Patent Text Reader

Abstract

The invention discloses a tool calling method based on a large language model, and the method comprises the steps: obtaining a question text which comprises task information; obtaining a matched first candidate tool set in the question text and a preset tool library; screening the first candidate tool set on the basis of locally available current hardware resources and / or available current online resources of a cloud application program interface (API) to obtain a second candidate tool set; determining a target tool matched with the task information from a second candidate tool set based on a large language model, and generating a tool calling instruction based on the target tool; selecting a called target tool based on the tool calling instruction, and obtaining an output result of the target tool; based on the question text and the output result, reply text corresponding to the large language model is obtained. According to the method, the tool with a higher matching degree can be obtained in combination with the resource occupation conditions of the tools of multiple platforms, and efficient utilization of resources and dynamic scheduling of task tools are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of large language models, and in particular to a tool calling method and device based on a large language model. Background Art

[0002] With the rapid development of artificial intelligence technology, it has demonstrated powerful text generation capabilities in the field of natural language processing, and can provide answers to questions or queries. However, faced with complex task requirements, a single platform or a single-modal large language model often finds it difficult to provide accurate and comprehensive answers. Summary of the invention

[0003] The purpose of this application is to provide a tool calling method based on a large language model, combining the resource usage of tools on multiple platforms to obtain tools with higher matching degrees, achieve efficient use of resources and dynamic scheduling of task tools.

[0004] In a first aspect, the present application provides a tool calling method based on a large language model, the method comprising: Obtaining a question text, wherein the question text includes task information; Obtaining a first candidate tool set matching the question text with a preset tool library, wherein the tool library includes a plurality of tools; Based on currently available local hardware resources and / or currently available online resources of a cloud application programming interface (API), the first candidate tool set is screened to obtain a second candidate tool set; determining a target tool matching the task information from the second candidate tool set based on the large language model, and generating a tool calling instruction based on the target tool; Selecting a target tool to be called based on the tool calling instruction so that the target tool executes the task corresponding to the task information and obtains the output result of the target tool; Based on the question text and the output results, the large language model generates the answer text corresponding to the question text.

[0005] Optionally, the method further includes a training method for a large language model, and the training method includes: Constructing a first problem sample, and determining a plurality of first tools that need to be called in the first problem sample; Split the first problem sample according to graph logic and / or tree logic to obtain multiple sub-problems in the first problem sample; Logically associating the plurality of first tools according to the plurality of sub-problems to obtain a tool chain; The tool chain is disassembled in a defective manner to obtain a first tool sample set with multiple permutations and combinations; The large language model is trained with the first question sample and a plurality of first tool sample sets.

[0006] Optionally, the method further includes a training method for a large language model, and the training method includes: Constructing a second problem sample, and determining a plurality of second tools that need to be called in the second problem sample; Performing a plurality of permutations and combinations of the plurality of second tools in a defective manner, and / or adding a plurality of tools irrelevant to the plurality of second tools, to obtain a plurality of second tool sample sets; The large language model is trained with the second question sample and a plurality of second tool sample sets.

[0007] Optionally, the tools include local tools and online tools corresponding to the cloud application programming interface (API). Based on the current hardware resources available locally and / or the current online resources available to the cloud API, the first candidate tool set is screened to obtain the second candidate tool set, including: If the currently available local hardware resources are less than the first threshold, the local tools whose hardware resources are greater than the second threshold are removed from the first candidate tool set to obtain a second candidate tool set; and / or, If the current online resources available for the cloud API are less than the third threshold, the cloud APIs corresponding to the online tools whose computing resources in the first candidate tool set are greater than the fourth threshold are eliminated to obtain a second candidate tool set, wherein the current online resources available for the cloud API include computing resources and bandwidth resources.

[0008] Optionally, the tools include local tools and online tools corresponding to the cloud application programming interface (API). Based on the current hardware resources available locally and / or the current online resources available to the cloud API, the first candidate tool set is screened to obtain the second candidate tool set, including: If there are local tools and online tools with the same functions in the first candidate tool set, and if the current locally available hardware resources meet the hardware resources required by the local tools, the cloud API corresponding to the online tool is removed from the first candidate tool set to obtain a second candidate tool set.

[0009] Optionally, the tool library stores tool information of the tools, including type, name, function, tool parameter, description information and requirements, wherein obtaining the first candidate tool set matching the question text with the preset tool library includes: Get the text embedding vector of the question text; Get the tool embedding vector of the tool description of each tool; The similarity between the text embedding vector and the tool embedding vector of each tool is calculated by the retriever, and the K tools with the highest similarity are determined as the first candidate tool set.

[0010] Optionally, the tool library stores tool information of the tool, the tool information including type, name, function, tool parameter, description information and requirement, wherein determining a target tool associated with the task information from the second candidate tool set based on the large language model, and generating a tool call instruction based on the target tool include: Structuring the tool information of each tool to obtain the Json description information corresponding to each tool; Add the Json description information of each tool in the second candidate tool set to the initial prompt word template to obtain an adjusted prompt word template, wherein the initial prompt word template includes a system prompt word and a task type template; Based on the adjusted prompt word template, the information of each tool in the second candidate tool set and the question text, the large language model determines the target tool that matches the task information; A tool calling instruction is generated based on the tool information of the target tool, wherein the tool calling instruction includes a tool name and tool parameters.

[0011] Optionally, selecting a target tool to be called based on the tool calling instruction so that the target tool executes a task corresponding to the task information and obtains an output result of the target tool includes: Generate structured function call instructions based on tool call instructions, and output the structured function call instructions to the FastAPI platform; The FastAPI platform uses asynchronous monitoring to monitor structured function call instructions; The FastAPI platform selects the target tool to call based on the structured function call instructions; The FastAPI platform monitors the execution status of the target tool in real time to obtain the output results of the target tool.

[0012] In a second aspect, the present application provides a tool calling device based on a large language model, the device comprising: An acquisition module, used for acquiring a question text, wherein the question text includes task information; A retrieval module is used to obtain a first candidate tool set matching the question text with a preset tool library, wherein the tool library includes a plurality of tools; A screening module, configured to screen the first candidate tool set based on currently available local hardware resources and / or currently available online resources of a cloud application program interface (API), to obtain a second candidate tool set; A language model module, used to determine a target tool matching the task information from the second candidate tool set based on the large language model, and generate a tool calling instruction based on the target tool; A tool calling module is used to select a target tool to be called based on a tool calling instruction, so that the target tool executes the task corresponding to the task information and obtains the output result of the target tool; The language model module is also used to generate a reply text corresponding to the question text based on the question text and the output result.

[0013] In a second aspect, the present application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the tool calling method based on the large language model as described above are implemented.

[0014] This application uses a tool scheduling strategy of mixed calling of local tools and online tools in the cloud based on the obtained problem text and the matching first candidate tool set, combined with the resource usage of tools on multiple platforms, to further screen the first candidate tool set to obtain a better second candidate tool set, and determines the target tool matching the task information from the second candidate tool set based on the large language model to call the target tool. By screening the tools at multiple levels, not only can the tools that better match the task information be obtained, but also the task tools and resources can be dynamically combined to achieve efficient resource utilization and dynamic scheduling of task tools, thereby achieving the best balance between performance and cost, enhancing the ability to respond to task requirements, and improving the adaptability and scalability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A first flow chart of a tool calling method based on a large language model provided in an embodiment of the present application; Figure 2 A second flow chart of a tool calling method based on a large language model provided in an embodiment of the present application; Figure 3 A third flow chart of a tool calling method based on a large language model provided in an embodiment of the present application; Figure 4 A fourth flow chart of a tool calling method based on a large language model provided in an embodiment of the present application; Figure 5 A fifth flow chart of a tool calling method based on a large language model provided in an embodiment of the present application; Figure 6 A sixth flow chart of a tool calling method based on a large language model provided in an embodiment of the present application; Figure 7 A system block diagram of a tool calling device based on a large language model provided in an embodiment of the present application; Figure 8 A system block diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] The present application will be described in detail below in conjunction with the specific implementation modes shown in the accompanying drawings, but these implementation modes do not limit the present application. Structural, methodological, or functional changes made by ordinary technicians in the field based on these implementation modes are included in the protection scope of the present application.

[0017] Please refer to Figure 1 , an embodiment of the present application provides a tool calling method based on a large language model, the method comprising steps S101-S106.

[0018] S101, obtaining a question text, wherein the question text includes task information.

[0019] Exemplarily, the question text input by the user can be obtained based on human-computer interaction. The question text is used to represent the question. The question text can be information in any one or more modes such as text, image, audio, etc.

[0020] The question text includes task information, which is text information based on natural language description of the task to be performed subsequently. For example, "Please translate the following sentence into English" is a text translation task. In subsequent processing, the model or tool corresponding to the text translation task is retrieved.

[0021] S102, obtaining a first candidate tool set that matches the question text with a preset tool library, wherein the tool library includes a plurality of tools.

[0022] The tool library includes multiple tools. The tools include local tools and online tools corresponding to the cloud application programming interface (API). Tools can also be understood as models, which is not limited here. The cloud application programming interface (API) can be understood as a functional service interface deployed on the cloud platform, usually presented in HTTP / HTTPS style, supporting remote calls of models or systems to complete tasks. For example, tools may include text question-and-answer tools, counter tools, network knowledge retrieval tools, target detection, image recognition, image segmentation, text maps, etc. Tools can be local model encapsulation, or remote API (such as Baidu Encyclopedia API, Open API call, etc.) to meet the needs of different fields or different tasks.

[0023] Based on the question text, a search is performed in the tool library to obtain various tools matching the question text, and the various tools are used to form a first candidate tool set.

[0024] S103, based on currently available local hardware resources and / or currently available online resources of a cloud application programming interface (API), the first candidate tool set is screened to obtain a second candidate tool set.

[0025] Based on the obtained first candidate tool set, the available local and cloud resources are dynamically monitored, and the various tools in the first candidate tool set are further screened according to the current hardware resources available locally and / or the current online resources available through the cloud application programming interface (API) to obtain better tools, so as to form a second candidate tool set and realize efficient utilization of local and cloud resources.

[0026] S104, determining a target tool matching the task information from the second candidate tool set based on the large language model, and generating a tool calling instruction based on the target tool.

[0027] The large language model can be understood as a trained large language model. Based on the large language model, the question text is semantically understood. After determining the second set of candidate tools, since there are still many candidate tools, it is necessary to use the large language model to filter the target tool that matches the question text from the candidate tools. Therefore, based on the semantic understanding of the question text by the large language model, the target tool that matches the task information is determined from the second set of candidate tools, and the tool call instruction is generated based on the tool information of the target tool.

[0028] The number of target tools can be one or more. It is related to the number of tools needed in the question text. Some task information requires a target tool. For example, if the question is "What is the weather like today", the target tool is the weather tool. If the question is "Recognize the text in the image and translate it into English", the target tools are the image recognition tool and the translation tool.

[0029] For example, InternVL2.0 is used as the large language model base, and the autoregressive method is used to achieve semantic understanding. In this way, the multimodal large language model obtained by training the large language model base can predict the next token step by step based on the tokens it has generated in the past, forming a sequential language output.

[0030] Exemplarily, the above-mentioned autoregressive method can adopt the following autoregressive function, so that the large language model base can be trained by the following loss function, so that the output is closer to the natural distribution of real data.

[0031] ; Represents the predicted output (token) at time step t, Represents the output of the previous time step (Token). Represents the model parameters of the large language model, which can be a vector or a set of parameters that need to be learned through training data, representing the mapping from the observed historical sequence to the next output distribution. Given all previous observations Under the conditions of and model parameters θ, it is observed that The conditional probability is to input all observations before time step t into a large language model with parameter θ, and predict the output of the current time step t based on the large language model In the training process of the large language model, by minimizing the loss function , that is, maximizing the log-likelihood of predicting the true target at all time steps, thereby approximating the true data distribution. After the large language model is trained, the trained parameters θ are obtained. During the inference process, based on the trained model parameters θ, the large language model can predict the new input sequence in an autoregressive manner.

[0032] S105, selecting a target tool to be called based on the tool calling instruction, so that the target tool executes the task corresponding to the task information, and obtains an output result of the target tool.

[0033] According to the tool calling instruction, a general target tool is selected, the target tool executes the task information in the problem text, and outputs the result.

[0034] S106, based on the question text and the output result, the large language model generates a reply text corresponding to the question text.

[0035] The large language model integrates the question text and output results, adjusts or supplements the output content, and outputs the answer text corresponding to the question text.

[0036] In an embodiment of the present application, a problem text and a matching first candidate tool set are obtained, and in combination with the resource usage of tools on multiple platforms, a tool scheduling strategy of mixed calling of local tools and online tools in the cloud is adopted to further screen the first candidate tool set to obtain a better second candidate tool set, and a target tool matching the task information is determined from the second candidate tool set based on a large language model to implement the calling of the target tool. By screening tools at multiple levels, not only can tools that better match the task information be obtained, but also task tools and resources can be dynamically combined to achieve efficient utilization of resources and dynamic scheduling of task tools, thereby achieving the best balance between performance and cost.

[0037] An embodiment of the present application, such as Figure 2 As shown, the method also includes a training method for a large language model, and the training method includes the steps of: S201, constructing a first problem sample, and determining a plurality of first tools that need to be called in the first problem sample; S202, splitting the first question sample according to graph logic and / or tree logic to obtain multiple sub-questions of the first question sample; S203, logically associating the multiple first tools according to the multiple sub-problems to obtain a tool chain; S204, disassembling the tool chain in a defective manner to obtain a first tool sample set with multiple permutations and combinations; S205: Training a large language model using the first question sample and a plurality of first tool sample sets.

[0038] In the embodiment of the present application, during the training process of the large language model, the logical order of the tools in the sample tool set is shuffled to eliminate the influence of the model on the tool calling preference, while also avoiding the model's dependence on the order and position of the input sequence. Different tool combinations are generated by random sampling, random shuffling, etc. to enhance the diversity of tool sample combinations, improve the semantic understanding and robustness of the large language model for tools, and increase the model's ability to refuse to answer.

[0039] Exemplarily, a first problem sample is constructed, and the first problem sample is split according to graph logic and / or tree logic to obtain multiple sub-problems of the first problem sample. Each sub-problem corresponds to a tool, and multiple first tools that need to be called in the first problem sample can be determined. Multiple first tools are logically associated according to multiple sub-problems to obtain a tool chain, which represents a combination of multiple tools in which the output of the previous tool is used as the input of the next tool. For example, the first problem sample is "Help me identify this picture and translate it into English", and the tool chain called is: image recognition tool → text translation. The tool chain is disassembled in a defective manner to obtain first tool sample sets with multiple permutations and combinations. For example, a first tool sample set is constructed that includes all the first tools in the tool chain, such as the first tool sample set is {image recognition, text translation, search engine}. After model training, the large language model can output normal tool call results; a first tool sample set includes a part of the first tools in the tool chain, such as the first tool sample set is {image recognition, search engine}. After model training, the large language model can output "the tool is missing in the tool set and cannot be called and returned"; a first tool sample set does not include all the first tools in the tool chain, such as the first tool sample set is {search engine}. After model training, the large language model can output "there is no executable tool for the current task", which can realize the semantic understanding ability of the large language model, improve the accuracy of the large language model, and enhance the large language model's ability to handle tool missing and refuse to answer.

[0040] An embodiment of the present application, such as Figure 3 As shown, the method also includes a training method for a large language model, and the training method includes the steps of: S301, constructing a second problem sample, and determining a plurality of second tools that need to be called in the second problem sample; S302, performing a plurality of permutations and combinations on the plurality of second tools in a defective manner, and / or adding a plurality of tools irrelevant to the plurality of second tools, to obtain a plurality of second tool sample sets; S303: training the large language model using the second question sample and a plurality of second tool sample sets.

[0041] Construct a second problem sample, determine multiple second tools that need to be called in the second problem sample, perform multiple permutations and combinations on the multiple second tools in a missing manner, and / or add several tools that are not related to the multiple second tools to obtain multiple second tool sample sets.

[0042] For example, a second tool sample set is constructed with multiple second tools, a second tool sample set includes multiple second tools and other tools except the multiple second tools, and a second tool sample set lacks at least one second tool. In this embodiment, by constructing multiple tool sample sets, considering the data of combination types such as inclusion, missing and redundancy of tools in the tool set, the large language model is trained to improve the large language model's ability to understand tool semantics, and realize the generalization ability of the large language model for unknown tools.

[0043] An embodiment of the present application, such as Figure 4 As shown, the tool library stores tool information of tools, and the tool information includes type, name, parameter description information and requirements. The first candidate tool set matching the question text with the preset tool library is obtained, including the steps of: S401, obtaining a text embedding vector of the question text; S402, obtaining a tool embedding vector of a tool description of each tool; S403, calculating the similarity between the text embedding vector and the tool embedding vector of each tool through the retriever, and determining the K tools with the highest similarity as the first candidate tool set.

[0044] Get the question text input by the user and determine the text embedding vector of the question text. Text embedding refers to the technology of converting text content (such as words, sentences or paragraphs) into a dense vector of fixed length, which can capture the semantic and contextual information of the text, so that semantically related words are close in distance in the embedding space, and irrelevant words are far away. The semantic relationship between words can be represented by the distance of the vector. Exemplarily, an embedding model, i.e., an embedding model, can be used to express the question text in the form of a vector. For example, the embedding model can be a BGE Embedding model, which uses the BGEEmbedding model to perform vector representation on the question text to obtain the text embedding vector of the question text.

[0045] Exemplarily, an embedding model is used to obtain a tool embedding vector of a tool description of a tool. The embedding model may be a BGE Embedding model.

[0046] The similarity between the text embedding vector and the tool embedding vector of each tool is calculated by the retriever, and the K tools with the highest similarity are determined as the first candidate tool set to determine the tool set that is semantically related to the task information of the question text. For example, the similarity can be cosine similarity. For example, the retriever is a BGE Ranker re-ranking model.

[0047] Assuming that the task-related text of task A includes question text and prompt words, the tool description T of the task-related text is input into the BGE Ranker rearrangement model, and BGE encoding is performed to obtain the text embedding vector. The tool description is input into the BGE Ranker rearrangement model, and BGE encoding is performed to obtain the tool embedding vector. The cosine similarity between the text embedding vector and the tool embedding vector is calculated to determine the Top-K tools with the highest cosine similarity.

[0048] An example is as follows:

[0049]

[0050]

[0051] Top-K Tools

[0052] in, The text embedding vector representing the task-related text T, Characterization Toolset The tool embedding vector corresponding to the tool description. For each of the Top-K tools, a unique hash value is generated for the tool description of the tool, and the hash value is used as the key (key) and the corresponding tool identifier is used as the value (Value). A key-value pair of the tool is constructed and the key-value pair is stored in an efficient search structure hash map (Hash Map) or key-value memory (Key-Value Store). The tool identifier is used to represent the tool name. When the system needs to quickly identify or obtain detailed information about the tool based on the tool description, it can calculate the hash value of the tool description to be queried, and search in the search structure hash table or key-value memory based on the calculated hash value. When the matching hash value is queried, the corresponding tool identifier or tool-related detailed information can be quickly retrieved, so that the tool associated with the task can be retrieved more quickly.

[0053] By calculating the similarity between the semantic vector of task information and the semantic vector of tool information in the tool library, a high-confidence tool set can be generated, which will greatly reduce the context occupancy of the subsequent large language model and increase the diversity of tool types in the context of the large language model.

[0054] Exemplarily, the tool embedding vector of each tool is hash-encoded to achieve more efficient retrieval.

[0055] In one embodiment of the present application, the tools include local tools and online tools corresponding to a cloud application programming interface (API). The first candidate tool set is screened based on the current hardware resources available locally and / or the current online resources available to the cloud API to obtain a second candidate tool set, including: if the current hardware resources available locally are less than a first threshold, local tools in the first candidate tool set whose hardware resources are greater than a second threshold are eliminated to obtain a second candidate tool set.

[0056] Exemplarily, the hardware resources may include image processor resources, and the image processor resources include image processor video memory occupancy, total video memory, image processor usage, etc. The image processor resources may be determined according to the platform, and the first threshold and the second threshold may be determined according to actual conditions, or may be a dynamic threshold. The current image processor resource occupancy of the platform may be obtained through a command, thereby obtaining the available current image processor resources.

[0057] In one embodiment of the present application, the tools include local tools and online tools corresponding to cloud application programming interfaces (APIs), and the first candidate tool set is screened based on the current hardware resources available locally and / or the current online resources available to the cloud APIs to obtain the second candidate tool set, including: if the current online resources available to the cloud APIs are less than the third threshold, the cloud APIs corresponding to the online tools whose computing resources are greater than the fourth threshold in the first candidate tool set are removed to obtain the second candidate tool set, wherein the current online resources available to the cloud APIs include computing resources and bandwidth resources. The third threshold and the fourth threshold can be determined according to actual conditions, or can be a dynamic threshold.

[0058] Exemplarily, the online resources available to the cloud API can be queried through the API status interface provided by the cloud, and computing resources, bandwidth resources, etc. can be obtained. For example, if the computing resources of the online tool corresponding to a certain cloud application programming interface API exceed 20% of the online resources available to the cloud API, the cloud API corresponding to the online tool is removed from the second candidate tool set.

[0059] In one embodiment of the present application, the tools include local tools and online tools corresponding to cloud application programming interfaces (APIs). Based on the current hardware resources available locally and / or the current online resources available to the cloud APIs, the first candidate tool set is screened to obtain the second candidate tool set, including: if there are local tools and online tools with the same functions in the first candidate tool set, and if the current hardware resources available locally meet the hardware resources required by the local tools, the cloud APIs corresponding to the online tools are removed from the first candidate tool set to obtain the second candidate tool set. If two tools in the second candidate tool set can achieve the same function, it is considered a functional conflict, and the local tool is used first to reduce the tool overhead and increase the real-time performance of the tool.

[0060] In this embodiment, based on the local tool resource occupancy and the online tool resource occupancy corresponding to the cloud API, the resource occupancy of tools on multiple platforms is integrated to dynamically select the tool with the best computing resources according to task requirements, so as to achieve the best balance between computing performance and cost, so as to realize efficient utilization of resources and dynamic scheduling of tasks.

[0061] An embodiment of the present application, such as Figure 5 As shown, determining a target tool associated with the task information from the second candidate tool set based on the large language model, and generating a tool calling instruction based on the target tool, including the steps of: S501, structure the tool information of each tool to obtain the Json description information corresponding to each tool; S502, adding the Json description information of each tool in the second candidate tool set to the initial prompt word template to obtain an adjusted prompt word template, wherein the initial prompt word template includes a system prompt word and a task type template; S503, based on the adjusted prompt word template, each tool information in the second candidate tool set and the question text, the large language model determines a target tool that matches the task information; S504: Generate a tool calling instruction based on the tool information of the target tool, wherein the tool calling instruction includes a tool name and tool parameters.

[0062] The tool call instruction can include type type, tool definition function, function includes tool name name, tool description description, parameter parameter, required type required, etc.

[0063] Assuming that the system prompt word in the above S502 can be represented by SP, the task type template is represented by Template(T), Template(T) can set a fixed template according to the task type, and the second candidate tool set is represented by FL, then the adjusted prompt word template P(T) can be represented by the following formula: P(T) =SP+Template(T)+FL.

[0064] Exemplarily, the tool information of each tool is structured to obtain the Json description information corresponding to each tool. The tool information includes type, name, parameter description information and requirements. The tool information is parameterized and described in Json format.

[0065] For example, taking semantic-sam tool and clip tool as examples, the description format of each tool can adopt Json format, including type, function, name, description, parameters and required fields. For example, type indicates the tool function type; function indicates the function of the tool function called by the tool, and function can include name and description, where name can be the tool name and description can be the core function description of the tool; parameters indicate the parameters contained in the tool, assuming that the parameters include image and boxes, where: image indicates the path of the input image, and its format is string type (such as local path or URL); boxes indicates the coordinates of the target area, and supports string format or JSON array (such as [[x1, y1, x2, y2], ...]); required indicates that the image and boxes parameters must be provided when calling.

[0066] For example, the description is "According to the coordinate information in the input boxes parameter, the image is segmented, and then different fine-grained segmentation results are output.". For another example, the description is: "The path to the image used to pass to the model as input.". The Json description information of each tool in the second candidate tool set is added to the initial prompt word template to obtain an adjusted prompt word template, wherein the initial prompt word template includes a system prompt word and a task type template. Based on the adjusted prompt word template, the information of each tool in the second candidate tool set and the question text, the large language model determines the target tool associated with the task information, and generates a tool call instruction based on the tool information of the target tool, and the tool call instruction may include a tool name and tool parameters.

[0067] In this embodiment, the Json description information of the tool is added to the system prompt words. Under the guidance of the system prompt words, the large language model can better understand the tool semantic information, so as to obtain the target tool that better matches the task information.

[0068] An embodiment of the present application, such as Figure 6 As shown, based on the tool calling instruction, the target tool to be called is selected so that the target tool performs the task corresponding to the task information, and the output result of the target tool is obtained, including: S601, generating a structured function call instruction based on the tool call instruction, and outputting the structured function call instruction to the FastAPI platform; S602, the FastAPI platform uses an asynchronous monitoring method to monitor the structured function call instruction; S603, the FastAPI platform selects a target tool to be called according to the structured function call instruction; S604, the FastAPI platform monitors the execution status of the target tool in real time to obtain the output result of the target tool.

[0069] In this embodiment, the FastAPI framework is used as the monitoring mechanism for tool calls. The local model is suspended in the background through the framework, and the tool call is realized through the port access and callback mechanism. For the cloud API, it is accessed through http and other forms. Based on the FastAPI framework, standardized tool interfaces, real-time status monitoring (supporting WebSocket and polling), dynamic tool calls and result returns, tool registration and discovery, error handling and retry mechanisms, as well as performance monitoring and logging can be implemented to realize the mechanism of efficient tool calls and status management driven by task requirements. The local tools are deployed in the framework, and the local tools are called through asynchronous monitoring.

[0070] Exemplarily, the above asynchronous monitoring method can be expressed by the following formula: , Callback ; Among them, D i Represents the i-th tool, Callback represents the callback function, and AsyncListen() represents the asynchronous listening function.

[0071] Exemplarily, a structured function call instruction is generated based on the tool call instruction, and the structured function call instruction is output to the FastAPI platform. A structured function call instruction in the function calling format may be used, and the structured function call instruction may include API_name and API_params, where API_name may be the tool name and API_params may be the tool parameter name. API_params may also be in the form of a dictionary containing tool parameter names and values. Function calling is a method in which a large language model generates an executable function description in a specific format (structured (Json)). Through this mechanism, the large language model can call external tools, APIs, etc. to complete tasks that it cannot complete through text. As the tool call backend in the function calling mechanism of the large language model, the FastAPI platform parses the structured tool request, dynamically calls the registered tool function, and returns the output result of the tool call. Based on the question text and the output result, the large language model generates a reply text corresponding to the question text.

[0072] For example, the user asks: Please tell me what animals are in the picture? The tool called is grounding_dino. The structured function call instruction in the function calling format is {"API_Name": "grounding_dino", "API_Name": {"image_path": " / data / cats.jpg", "prompt": "a cat on the table"}}. The result of the grounding_dino call returns: A cat was found. The large language model forms a reply based on the call result: According to the detection result, the picture contains a cat, and the cat is on the table.

[0073] Based on the same inventive concept, an embodiment of the present application also provides a tool calling device based on a large language model. The implementation solution for solving the problem provided by this device is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in the embodiments of one or more tool calling devices based on a large language model provided below can be found in the above limitations on the tool calling method based on a large language model, which will not be repeated here.

[0074] like Figure 7 As shown, the present application provides a tool calling device based on a large language model, the device comprising: An acquisition module 701 is used to acquire a question text, wherein the question text includes task information; A search module 702 is used to obtain a first candidate tool set that matches the question text with a preset tool library, wherein the tool library includes a plurality of tools; A screening module 703 is used to screen the first candidate tool set based on currently available local hardware resources and / or currently available online resources of a cloud application program interface API to obtain a second candidate tool set; A language model module 704 is used to determine a target tool matching the task information from the second candidate tool set based on the large language model, and generate a tool calling instruction based on the target tool; A tool calling module 705 is used to select a target tool to be called based on the tool calling instruction, so that the target tool executes the task corresponding to the task information and obtains the output result of the target tool; The language model module 704 is also used to generate a reply text corresponding to the question text based on the question text and the output result.

[0075] Optionally, the language model module 704 is specifically used for: Constructing a first problem sample, and determining a plurality of first tools that need to be called in the first problem sample; Split the first problem sample according to graph logic and / or tree logic to obtain multiple sub-problems in the first problem sample; Logically associating the plurality of first tools according to the plurality of sub-problems to obtain a tool chain; The tool chain is disassembled in a defective manner to obtain a first tool sample set with multiple permutations and combinations; The large language model is trained with the first question sample and a plurality of first tool sample sets.

[0076] Optionally, the language model module 704 is specifically used for: Constructing a second problem sample, and determining a plurality of second tools that need to be called in the second problem sample; Performing a plurality of permutations and combinations on the plurality of second tools in a defective manner, and / or adding a plurality of tools irrelevant to the plurality of second tools, to obtain a plurality of second tool sample sets; The large language model is trained with the second question sample and a plurality of second tool sample sets.

[0077] Optionally, the tool includes a local tool and an online tool corresponding to a cloud application programming interface (API). The screening module 703 is specifically used for: If the currently available local hardware resources are less than the first threshold, the local tools whose hardware resources are greater than the second threshold are removed from the first candidate tool set to obtain a second candidate tool set; and / or, If the current online resources available for the cloud API are less than the third threshold, the cloud APIs corresponding to the online tools whose computing resources in the first candidate tool set are greater than the fourth threshold are eliminated to obtain a second candidate tool set, wherein the current online resources available for the cloud API include computing resources and bandwidth resources.

[0078] Optionally, the tool includes a local tool and an online tool corresponding to a cloud application programming interface (API). The screening module 703 is specifically used for: If there are local tools and online tools with the same functions in the first candidate tool set, and if the current locally available hardware resources meet the hardware resources required by the local tools, the cloud API corresponding to the online tool is removed from the first candidate tool set to obtain a second candidate tool set.

[0079] Optionally, the tool library stores tool information of the tool, the tool information including type, name, function, tool parameter, description information and requirements, and the retrieval module 702 is specifically used to: Get the text embedding vector of the question text; Get the tool embedding vector of the tool description of each tool; The similarity between the text embedding vector and the tool embedding vector of each tool is calculated by the retriever, and the K tools with the highest similarity are determined as the first candidate tool set.

[0080] Optionally, the tool library stores tool information of the tools, including type, name, function, tool parameters, description information and requirements. The language model module 704 is specifically used for: Structuring the tool information of each tool to obtain the Json description information corresponding to each tool; Add the Json description information of each tool in the second candidate tool set to the initial prompt word template to obtain an adjusted prompt word template, wherein the initial prompt word template includes a system prompt word and a task type template; Based on the adjusted prompt word template, the information of each tool in the second candidate tool set and the question text, the large language model determines the target tool that matches the task information; A tool calling instruction is generated based on the tool information of the target tool, wherein the tool calling instruction includes a tool name and tool parameters.

[0081] Optionally, the tool calling module 705 is specifically used for: Generate structured function call instructions based on tool call instructions, and output the structured function call instructions to the FastAPI platform; The FastAPI platform uses asynchronous monitoring to monitor structured function call instructions; The FastAPI platform selects the target tool to call based on the structured function call instructions; The FastAPI platform monitors the execution status of the target tool in real time to obtain the output results of the target tool.

[0082] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned tool calling methods based on a large language model is implemented.

[0083] Figure 8 It is a schematic diagram of the hardware structure of a computer device provided in an embodiment of the present application. Figure 8 The computer device shown includes: a processor 801, a communication interface 802, a memory 803 and a communication bus 804. The processor 801, the communication interface 802 and the memory 803 communicate with each other via the communication bus 804. Figure 8 The connection method between the processor 801 , the communication interface 802 , and the memory 803 shown is merely exemplary. During implementation, the processor 801 , the communication interface 802 , and the memory 803 may also be communicatively connected to each other using other connection methods besides the communication bus 804 .

[0084] The memory 803 can be used to store computer programs, which may include instructions and data to implement the steps of any of the above tool calling methods based on a large language model. In an embodiment of the present application, the memory 803 may be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical storage, and registers. The memory 803 may include a hard disk and / or a memory.

[0085] The processor 801 may be a general-purpose processor, which may be a processor that performs specific steps and / or operations by reading and executing a computer program (e.g., a computer program) stored in a memory (e.g., a memory 803). The general-purpose processor may use data stored in the memory (e.g., the memory 803) in the process of performing the steps and / or operations. The general-purpose processor may be, for example, but not limited to, a central processing unit (CPU). In addition, the processor 801 may also be a dedicated processor, which may be a processor specially designed to perform specific steps and / or operations. The dedicated processor may be, for example, but not limited to, an ASIC and an FPGA. In addition, the processor 801 may also be a combination of multiple processors, such as a multi-core processor.

[0086] The communication interface 802 may include an input / output (I / O) interface, a physical interface, and a logical interface for interconnecting devices within the network device, as well as an interface for interconnecting the network device with other devices (e.g., network devices). The communication network may be Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 802 may be a module, a circuit, a transceiver, or any device capable of implementing communication.

[0087] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 801 or an instruction in the form of software. The method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in the processor for execution. The software module can be located in a mature storage medium in the field such as a random access memory flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 803, and the processor 801 reads the information in the memory 803 and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.

[0088] Although the preferred embodiments of the present application have been disclosed for illustrative purposes, those skilled in the art will appreciate that various modifications, additions and substitutions are possible, without departing from the scope and spirit of the application as disclosed in the accompanying claims.

Claims

1. A tool calling method based on a large language model, characterized in that: The method comprises: Obtaining a question text, wherein the question text includes task information; Acquire a first candidate tool set that matches the question text with a preset tool library, wherein the tool library includes a plurality of tools; Based on currently available local hardware resources and / or currently available online resources of a cloud application programming interface (API), the first candidate tool set is screened to obtain a second candidate tool set; Determining a target tool matching the task information from the second candidate tool set based on the large language model, and generating a tool calling instruction based on the target tool; Selecting a target tool to be called based on the tool calling instruction, so that the target tool executes the task corresponding to the task information, and obtaining an output result of the target tool; Based on the question text and the output result, the large language model generates a reply text corresponding to the question text.

2. The tool calling method based on a large language model according to claim 1, characterized in that: The method also includes a training method for the large language model, and the training method includes: Constructing a first problem sample, and determining a plurality of first tools that need to be called in the first problem sample; Split the first question sample according to graph logic and / or tree logic to obtain multiple sub-questions in the first question sample; Logically associating the plurality of first tools according to the plurality of sub-problems to obtain a tool chain; Disassembling the tool chain in a defective manner to obtain a first tool sample set in multiple permutations and combinations; The large language model is trained using the first question sample and a plurality of first tool sample sets.

3. The tool calling method based on a large language model according to claim 1, characterized in that: The method also includes a training method for the large language model, and the training method includes: Constructing a second problem sample, and determining a plurality of second tools that need to be called in the second problem sample; Performing a plurality of permutations and combinations on the plurality of second tools in a defective manner, and / or adding a plurality of tools irrelevant to the plurality of second tools, to obtain a plurality of second tool sample sets; The large language model is trained using the second question sample and a plurality of second tool sample sets.

4. The tool calling method based on a large language model according to claim 1, characterized in that: The tools include local tools and online tools corresponding to the cloud application programming interface (API). Based on the current hardware resources available locally and / or the current online resources available to the cloud API, the first candidate tool set is screened to obtain a second candidate tool set, including: If the currently available local hardware resources are less than the first threshold, remove the local tools whose hardware resources are greater than the second threshold from the first candidate tool set to obtain a second candidate tool set; and / or, If the current online resources available for the cloud API are less than the third threshold, the cloud APIs corresponding to the online tools whose computing resources in the first candidate tool set are greater than the fourth threshold are eliminated to obtain a second candidate tool set, wherein the current online resources available for the cloud API include computing resources and bandwidth resources.

5. The tool calling method based on a large language model according to claim 1, characterized in that: The tools include local tools and online tools corresponding to the cloud application programming interface (API). Based on the current hardware resources available locally and / or the current online resources available to the cloud API, the first candidate tool set is screened to obtain a second candidate tool set, including: If there are local tools and online tools with the same functions in the first candidate tool set, and if the current locally available hardware resources meet the hardware resources required by the local tools, the cloud API corresponding to the online tool is removed from the first candidate tool set to obtain a second candidate tool set.

6. The tool calling method based on a large language model according to claim 1, characterized in that: The tool library stores tool information of the tool, wherein the tool information includes type, name, function, tool parameter, description information and requirement, wherein obtaining the first candidate tool set matching the question text with the preset tool library includes: Obtaining a text embedding vector of the question text; Get the tool embedding vector of the tool description of each tool; The similarity between the text embedding vector and the tool embedding vector of each tool is calculated by a retriever, and K tools with the highest similarity are determined as the first candidate tool set.

7. The tool calling method based on a large language model according to claim 1, characterized in that: The tool library stores tool information of the tool, the tool information including type, name, function, tool parameter, description information and requirement, wherein determining a target tool associated with the task information from the second candidate tool set based on the large language model, and generating a tool call instruction based on the target tool comprises: Structuring the tool information of each tool to obtain Json description information corresponding to each tool; Adding the Json description information of each tool in the second candidate tool set to the initial prompt word template to obtain an adjusted prompt word template, wherein the initial prompt word template includes a system prompt word and a task type template; Based on the adjusted prompt word template, each tool information in the second candidate tool set and the question text, the large language model determines a target tool that matches the task information; A tool calling instruction is generated based on the tool information of the target tool, wherein the tool calling instruction includes a tool name and tool parameters.

8. The tool calling method based on a large language model according to claim 7, characterized in that: Selecting a target tool to be called based on the tool calling instruction so that the target tool executes the task corresponding to the task information and obtaining an output result of the target tool includes: Generate a structured function call instruction based on the tool call instruction, and output the structured function call instruction to the FastAPI platform; The FastAPI platform monitors the structured function call instruction in an asynchronous monitoring manner; The FastAPI platform selects a target tool to call according to the structured function call instruction; The FastAPI platform monitors the execution status of the target tool in real time to obtain the output results of the target tool.

9. A tool calling device based on a large language model, characterized in that: The device comprises: An input module, used for obtaining a question text, wherein the question text includes task information; A retrieval module, used to obtain a first candidate tool set matching the question text with a preset tool library, wherein the tool library includes a plurality of tools; A screening module, configured to screen the first candidate tool set based on currently available local hardware resources and / or currently available online resources of a cloud application programming interface (API) to obtain a second candidate tool set; A language model module, configured to determine a target tool matching the task information from the second candidate tool set based on the large language model, and generate a tool calling instruction based on the target tool; A tool calling module, used for selecting a target tool to be called based on the tool calling instruction, so that the target tool executes the task corresponding to the task information and obtains the output result of the target tool; The language model module is also used to generate a reply text corresponding to the question text based on the question text and the output result.

10. A computer device comprising a memory and a processor, characterized in that: The memory stores a computer program, wherein the processor implements the steps of the tool calling method based on a large language model according to any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Task solution-oriented training method and use method of generative large language model

    CN116756564A

  • Plug-in retrieval method and device, storage medium and equipment

    CN117828024A

  • Language model tool calling method and device, computer equipment and storage medium

    CN118690853A

  • Dynamic tool selection and optimization system and method for large model external tool calling

    CN119166318A

  • Task scale scheduling method and device based on large model

    CN119201392A

Cited By

  • Task execution method, device and system, electronic device and storage medium

    CN120723411A

  • Function display method and device based on large model, equipment and medium

    CN120851190A

  • Hierarchical decision architecture model and training reasoning system and method thereof

    CN120875023A

  • Tool calling method and device for large language model, medium and equipment

    CN121541945A

  • Hierarchical verifiable function trusted scheduling and checking method and system for large model inference

    CN122595043A