A tool invocation method, apparatus, device and storage medium
Patent Information
- Application Number
- CN202611000657.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]有鉴于此,本发明的目的在于提供一种工具调用方法、装置、设备及存储介质,能够解决智能体工具调度中任务规划可靠性不足、工具误调用及执行效率低下的问题
[0015]可见,本申请在接收到用户指令时,基于预设工具库中各工具的能力摘要信息将所述用户指令分解为包括至少一个子步骤的结构化执行计划,并为每一子步骤绑定对应的工具标识;解析所述结构化执行计划中各子步骤对应工具的输入输出数据结构,以得到工具间依赖关系,并基于所述工具间依赖关系生成有向无环图执行路径;按照所述有向无环图执行路径执行每一子步骤时,基于每一子步骤绑定的工具标识从所述预设工具库中提取完整的工具定义信息,并为每一子步骤构建仅包括对应工具定义信息的原子执行环境,以在所述原子执行环境中执行每一子步骤,然后将执行结果按照所述有向无环图执行路径传递至下游子步骤。
Smart Images

Figure CN122816807A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of scheduling technology, and in particular to a tool invocation method, apparatus, device, and storage medium. Background Technology
[0002] When intelligent agents handle user problems, they typically need to analyze the problem, plan multiple executable steps, and sequentially call tools to obtain intermediate results. Current mainstream methods for tool invocation by intelligent agents often employ full prompt injection or keyword matching, injecting descriptions of all available tools into the context window of a large language model at once, allowing the model to make linear decisions. This method has the following problems: the actual capability boundaries and input / output constraints of the tools are not fully considered during task planning, resulting in insufficient reliability of the generated execution plan; when the tool library is large, full injection is prone to context window overflow, and excessive interference information can cause the model to misinterpret, mistakenly calling non-existent tools or incorrectly selecting tools with high similarity.
[0003] In addition, traditional execution modes often use serial execution for each step. The system cannot automatically identify the dependencies between steps, nor can it convert steps without dependencies into parallel operations. This results in the overall response time increasing linearly with the task complexity, making it difficult to meet real-time requirements.
[0004] Therefore, the aforementioned technical problems urgently need to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a tool invocation method, apparatus, device, and storage medium that can solve the problems of insufficient reliability of task planning, erroneous tool invocation, and low execution efficiency in intelligent agent tool scheduling. The specific solution is as follows: The first aspect of this application provides a tool invocation method, including: Upon receiving a user instruction, the user instruction is decomposed into a structured execution plan including at least one sub-step based on the capability summary information of each tool in the preset tool library, and a corresponding tool identifier is bound to each sub-step. The input and output data structures of the tools corresponding to each sub-step in the structured execution plan are parsed to obtain the dependencies between the tools, and a directed acyclic graph execution path is generated based on the dependencies between the tools. When executing each sub-step according to the directed acyclic graph execution path, complete tool definition information is extracted from the preset tool library based on the tool identifier bound to each sub-step, and an atomic execution environment containing only the corresponding tool definition information is constructed for each sub-step to execute each sub-step in the atomic execution environment. Then, the execution result is passed to the downstream sub-step according to the directed acyclic graph execution path.
[0006] Optionally, the capability summary information includes the tool name, function description, and core input / output parameters, and the tool definition information includes the tool name, function description, complete input / output parameters, parameter validation rules, usage precautions, and reference examples.
[0007] Optionally, the step of binding a corresponding tool identifier to each sub-step includes: When the same sub-step corresponds to multiple candidate tools, named entity recognition is performed on the input text of the current sub-step to extract the entity type; The entity type is matched with the input parameter definitions of each candidate tool, and a matching score is calculated. The target tool is determined from the plurality of candidate tools based on the matching score, and the tool identifier of the target tool is bound to the current sub-step.
[0008] Optionally, determining the target tool from the plurality of candidate tools based on the matching score includes: If multiple candidate tools have the same matching score, the multiple candidate tools are sorted according to a preset tool type sorting strategy to obtain a sorting result; wherein, the tool type sorting strategy is that resource acquisition tools have a higher priority than data processing tools, and data processing tools have a higher priority than data analysis tools. The tool that ranks first in the sorting results is identified as the target tool.
[0009] Optionally, generating the directed acyclic graph execution path based on the inter-tool dependencies includes: If a semantic match is detected between the output parameters of the first tool and the input parameters of the second tool, a directed edge is established from the first tool to the second tool. Each sub-step in the structured execution plan is treated as a node, and a directed acyclic graph is generated based on the established directed edges; wherein, the attributes of each node include the tool identifier bound to each sub-step; The topological sorting algorithm is used to divide the nodes in the directed acyclic graph into layers to generate the execution path of the directed acyclic graph; the nodes without prerequisite dependencies are located in the first layer, and the nodes that depend on the output results of other tools are located in the subsequent layers.
[0010] Optionally, the tool invocation method further includes: Traverse all nodes in the execution path of the directed acyclic graph. If the input parameters of the first node do not depend on the output of the second node, and the input parameters of the second node do not depend on the output of the first node, then it is determined that the first node and the second node have the conditions for parallel execution. During the execution phase, multiple nodes that meet the conditions for parallel execution are simultaneously initiated with call requests, and data aggregation is performed after each node returns the execution results.
[0011] Optionally, the tool invocation method further includes: For isolated nodes in a directed acyclic graph that have no dependencies, these isolated nodes are placed in a separate execution thread pool and executed in parallel with the other nodes.
[0012] A second aspect of this application provides a tool invocation apparatus, comprising: The execution plan generation module is used to decompose the user instruction into a structured execution plan including at least one sub-step based on the capability summary information of each tool in the preset tool library when a user instruction is received, and to bind a corresponding tool identifier to each sub-step. The execution path generation module is used to parse the input and output data structures of the tools corresponding to each sub-step in the structured execution plan to obtain the dependencies between tools, and generate a directed acyclic graph execution path based on the dependencies between tools. The execution module is used to extract complete tool definition information from the preset tool library based on the tool identifier bound to each sub-step when executing each sub-step according to the directed acyclic graph execution path, and to construct an atomic execution environment for each sub-step that only includes the corresponding tool definition information, so as to execute each sub-step in the atomic execution environment, and then pass the execution result to the downstream sub-step according to the directed acyclic graph execution path.
[0013] A third aspect of this application provides an electronic device including a processor and a memory; wherein the memory is used to store a computer program, which is loaded and executed by the processor to implement the aforementioned tool invocation method.
[0014] A fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions, which, when loaded and executed by a processor, implement the aforementioned tool invocation method.
[0015] As can be seen, when this application receives a user instruction, it decomposes the user instruction into a structured execution plan including at least one sub-step based on the capability summary information of each tool in the preset tool library, and binds a corresponding tool identifier to each sub-step; it parses the input and output data structure of the tool corresponding to each sub-step in the structured execution plan to obtain the inter-tool dependency relationship, and generates a directed acyclic graph execution path based on the inter-tool dependency relationship; when executing each sub-step according to the directed acyclic graph execution path, it extracts complete tool definition information from the preset tool library based on the tool identifier bound to each sub-step, and constructs an atomic execution environment for each sub-step that only includes the corresponding tool definition information, so as to execute each sub-step in the atomic execution environment, and then passes the execution result to the downstream sub-step according to the directed acyclic graph execution path.
[0016] Beneficial Effects: This application generates structured execution plans with bound tool identifiers based on tool capability summaries, enabling task planning to have a global tool perspective, ensuring precise matching between execution steps and target tools, and improving the reliability and executability of task planning. Furthermore, by parsing the tool input / output data structure, the dependencies between tools are obtained, and a directed acyclic graph (DAG) execution path is constructed based on this. This clarifies the dependencies between tools and optimizes the execution process, allowing tools with dependencies to automatically chain together, and tools without dependencies to execute in parallel according to the structure. Finally, when executing each sub-step according to the DAG execution path, this application constructs an atomic execution environment for each step containing only the corresponding tool definition information, thereby achieving lightweight and isolated execution context, shielding irrelevant tool interference, reducing context redundancy, avoiding erroneous tool calls, and significantly improving tool call accuracy. Simultaneously, by orderly transmitting execution results according to the DAG path, the overall accuracy, stability, and execution efficiency of agent tool scheduling are improved. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 This is a flowchart of a tool invocation method disclosed in this application; Figure 2 This is an execution flowchart of a priority scheduling mechanism disclosed in this application; Figure 3 This is a flowchart of a specific tool invocation method disclosed in this application; Figure 4This is a flowchart illustrating the execution of a tool registration mechanism disclosed in this application; Figure 5 This is a schematic diagram of the structure of a tool calling device disclosed in this application; Figure 6 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] In existing technologies, mainstream intelligent agent tool invocation methods mostly employ full prompt injection or keyword matching, which injects the descriptions of all available tools into the context window of a large language model at once, allowing the model to make linear decisions. This method suffers from the following problems: the actual capability boundaries and input / output constraints of the tools are not fully considered during task planning, resulting in insufficient reliability of the generated execution plan; when the tool library is large, full injection is prone to context window overflow, and excessive interference information can cause the model to misinterpret, mistakenly invoking non-existent tools or incorrectly selecting tools with high similarity. Furthermore, traditional execution modes often employ sequential execution, and the system cannot automatically identify dependencies between steps, nor can it convert steps without dependencies into parallel operations, resulting in overall response time increasing linearly with task complexity, making it difficult to meet real-time requirements. Therefore, this application discloses a tool invocation method, apparatus, device, and storage medium that can solve the problems of insufficient reliability in task planning, mis-invocation of tools, and low execution efficiency in intelligent agent tool scheduling.
[0021] Figure 1 A flowchart illustrating a tool invocation method provided in an embodiment of this application. See also... Figure 1 As shown, the tool's invocation methods include: Step S11: Upon receiving a user instruction, the user instruction is decomposed into a structured execution plan including at least one sub-step based on the capability summary information of each tool in the preset tool library, and a corresponding tool identifier is bound to each sub-step.
[0022] In this embodiment, upon receiving a user instruction, the task planning node is first initiated, thus entering the task planning phase. During this phase, the system does not send the user instruction along with all tool information to the large model. Instead, it extracts the capability summary information of each tool from a pre-defined tool library and injects it into the task planner's context window. The task planner then uses this capability summary information to decompose the user instruction into a structured execution plan containing multiple ordered sub-steps, clearly defining the tool identifier to be invoked in each sub-step. In other words, a structured execution plan with clearly defined tool bindings is generated.
[0023] It's important to note that the capability summary information only includes the tool name, functional description, and core input / output parameters. This ensures that the tool's context is not excessively long while still allowing the model to understand the tool's capabilities. In other words, this method minimizes the information in the task planner's context window, providing only a global view that lets the model know what tools are available, what they can do, and what the key inputs and outputs are, without loading complete parameter validation logic or usage examples. This gives the task planner an omniscient perspective, enabling it to explicitly specify the specific tool name and expected deliverables for each step when decomposing sub-steps. This generates a structured execution plan with tool bindings, ensuring precise matching between execution steps and target tools, thus improving the reliability and executability of task planning.
[0024] In a specific implementation, binding a corresponding tool identifier to each sub-step includes: when the same sub-step corresponds to multiple candidate tools, performing named entity recognition on the input text of the current sub-step to extract the entity type; matching the entity type with the input parameter definitions of each candidate tool to calculate a matching score; determining the target tool from the multiple candidate tools based on the matching score, and binding the tool identifier of the target tool to the current sub-step.
[0025] That is, such as Figure 2As shown in the illustration, this application discloses a priority scheduling mechanism for handling situations where multiple candidate tools correspond to the same sub-step. When multiple candidate tools correspond to the same sub-step, Named Entity Recognition (NER) is first performed on the input text (user input / context) of the current sub-step to extract entity types. For example, for the input "analyze the traffic of IP 192.168.1.1", the system will identify the entity type as [IP_Address]. Then, an affinity scoring engine is used to match the extracted entity types with the input parameters (i.e., input parameter schema) of each candidate tool to calculate the matching score. For example, the input parameter of tool A (a general search tool) might be a generic query: string, while the input parameter of tool B (an IP threat query tool) is explicitly ip: ip_address. Because the identified entity type [IP_Address] strongly matches the parameter type of tool B, tool B will receive a higher affinity weighting. Finally, based on the matching score, an ordered list of tools is output in descending order of score to determine the target tool from multiple candidate tools. Generally, the tool with the highest score is selected as the target tool, and the tool identifier of the target tool is bound to the current sub-step.
[0026] It should be noted that the step of determining the target tool from the plurality of candidate tools based on the matching score includes: if the matching scores of the plurality of candidate tools are the same, then the plurality of candidate tools are sorted according to a preset tool type sorting strategy to obtain a sorting result; wherein, the tool type sorting strategy is that resource acquisition tools have a higher priority than data processing tools, and data processing tools have a higher priority than data analysis tools; the tool that ranks first in the sorting result is determined as the target tool.
[0027] In other words, this application also discloses a processing strategy when multiple candidate tools have the same matching score. Specifically, a secondary sorting is performed based on a preset tool type sorting strategy. The tool type sorting strategy is as follows: resource acquisition tools have higher priority than data processing tools, and data processing tools have higher priority than data analysis tools, i.e., Resource_Fetch (acquisition) > Data_Process (processing) > Analysis (analysis). That is, when multiple tools have similar or identical scores, they are forcibly rearranged according to this sorting strategy to ensure that data (acquisition) precedes conclusions (analysis), thus ensuring the correctness of the execution logic. That is, data is acquired first (resource acquisition), then data is processed (data processing), and finally a conclusion is drawn (data analysis). After sorting, the tool ranked first is determined as the target tool, and its tool identifier is bound to the current sub-step.
[0028] Step S12: Parse the input and output data structures of the tools corresponding to each sub-step in the structured execution plan to obtain the dependencies between tools, and generate a directed acyclic graph execution path based on the dependencies between tools.
[0029] In this embodiment, after obtaining the structured execution plan, it is necessary to further transform the plan into an efficiently executable graph structure. The core of this is to construct a directed acyclic graph execution path by statically analyzing the data dependencies between tools, thus laying the foundation for subsequent parallel scheduling.
[0030] In this specific implementation, the sub-steps in the structured execution plan are not executed sequentially. Instead, the input and output data structures of the tools corresponding to each sub-step are parsed to obtain the dependencies between tools, i.e., to detect whether there are data transfer relationships between different tools. Specifically, the system extracts the input parameter definitions (i.e., what format of data the tool needs to receive) and output parameter definitions (i.e., what format of data the tool will return after execution). These data structures are usually described in the form of a schema, specifying the parameter name, type (such as string, integer, IP address type, etc.), and whether they are required. Through this parsing process, the system can obtain the basic information necessary for subsequently building dependencies, i.e., what each tool needs and what it produces. Then, a directed acyclic graph execution path is constructed based on the analyzed tool dependencies, thereby clarifying the tool dependencies and optimizing the execution process, enabling tools with dependencies to be automatically linked together, and tools without dependencies to be executed in parallel according to the structure.
[0031] Step S13: When executing each sub-step according to the directed acyclic graph execution path, extract complete tool definition information from the preset tool library based on the tool identifier bound to each sub-step, and construct an atomic execution environment for each sub-step that only includes the corresponding tool definition information, so as to execute each sub-step in the atomic execution environment, and then pass the execution result to the downstream sub-step according to the directed acyclic graph execution path.
[0032] In this embodiment, when entering a specific step execution stage, the system uses an interceptor to read the previously generated directed acyclic graph execution path. When executing each sub-step according to the directed acyclic graph execution path, it constructs an atomic execution environment for each step that only contains the corresponding tool definition information. This achieves lightweighting and isolation of the execution context, shields against interference from irrelevant tools, reduces context redundancy, avoids incorrect tool calls, and significantly improves the accuracy of tool calls. At the same time, by orderly transmitting execution results according to the directed acyclic graph path, the overall accuracy, stability, and execution efficiency of agent tool scheduling are improved.
[0033] It should be noted that the tool definition information includes the tool name, function description, complete input and output parameters, parameter validation rules, usage precautions, and reference examples. Understandably, during the execution phase, the simplified capability summary information is no longer used. Instead, based on the tool identifier bound to the current sub-step, complete tool definition information is extracted from a pre-defined tool library. This includes, but is not limited to, the tool name, function description, complete input and output parameters, parameter validation rules, usage precautions, and reference examples. Then, using this complete information, an atomic execution environment containing only the tool definition information is constructed for the current sub-step. In this environment, the execution agent's context only contains the tools necessary for the current step; all other irrelevant tools are forcibly hidden. This eliminates the need for the execution agent to choose from dozens or even hundreds of tools, achieving single-channel precise invocation. This fundamentally eliminates model illusions and mis-invoking problems caused by context redundancy, significantly improving accuracy.
[0034] Finally, within the constructed atomic execution environment, the system calls the corresponding tool to execute the current sub-step. After execution, the system passes the execution result to the downstream sub-steps according to the directed acyclic graph execution path. Specifically, when a sub-step is completed, the system automatically parses its result data and distributes the result to all dependent next-level nodes according to the directed acyclic graph, using it as input parameters for these nodes.
[0035] As can be seen, this application generates a structured execution plan based on tool capability summaries and binds tool identifiers, enabling task planning to have a global tool perspective, ensuring precise matching between execution steps and target tools, and improving the reliability and executability of task planning. Furthermore, by parsing the tool input and output data structures, the dependencies between tools are obtained, and a directed acyclic graph execution path is constructed based on this. This clarifies the dependencies between tools and optimizes the execution process, allowing tools with dependencies to be automatically linked and tools without dependencies to be executed in parallel according to the structure. Finally, when executing each sub-step according to the directed acyclic graph execution path, this application constructs an atomic execution environment for each step containing only the corresponding tool definition information, thereby achieving lightweight and isolated execution context, shielding irrelevant tool interference, reducing context redundancy, avoiding erroneous tool calls, and significantly improving tool call accuracy. Simultaneously, by orderly transmitting execution results according to the directed acyclic graph path, the overall accuracy, stability, and execution efficiency of agent tool scheduling are improved.
[0036] Figure 3 A flowchart illustrating a specific tool invocation method provided in this application embodiment. See also... Figure 3 As shown, the tool's invocation methods include: Step S21: Upon receiving a user instruction, the user instruction is decomposed into a structured execution plan including at least one sub-step based on the capability summary information of each tool in the preset tool library, and a corresponding tool identifier is bound to each sub-step.
[0037] In this embodiment, as Figure 4 As shown, the process by which the system extracts only the tool's capability summary information can be called forward registration. In the forward registration phase, the system only registers the tool's metadata, specifically including: tool name, functional description, and core input / output parameters. This phase does not include complex parameter validation logic and examples; its core purpose is to minimize token consumption, ensuring that the task planner's context window only contains a global view that allows the model to understand what tools are available, what they can do, and what the key inputs and outputs are, without loading redundant detailed information. This gives the task planner an omniscient perspective, enabling it to decompose tasks based on a full understanding of the tool's capability boundaries.
[0038] Step S22: Parse the input and output data structures of the tools corresponding to each sub-step in the structured execution plan to obtain the dependencies between tools.
[0039] Step S23: If a semantic match is detected between the output parameters of the first tool and the input parameters of the second tool, then a directed edge is established from the first tool to the second tool.
[0040] In this embodiment, when establishing directed edges, the main focus is on detecting whether there is a semantic matching relationship between the input and output parameters of different tools. Specifically, this embodiment compares the input and output parameters of each tool one by one. When a semantic match is detected between the output parameter of the first tool and the input parameter of the second tool, a directed edge is established from the first tool to the second tool. The semantic matching can be determined in several ways: the system can detect whether there is an alias relationship between the names of the two parameters (e.g., both user_id and uid represent user identifiers), or it can detect whether the data types of the parameters are compatible (e.g., the output parameter type is ip_address, the input parameter type is string, and both semantically represent IP addresses). Furthermore, a semantic alignment algorithm can be used to calculate the semantic similarity of the parameter descriptions to determine whether a matching relationship exists.
[0041] For example, the structured definition of a system static analysis tool yields: Tool A outputs: {user_id: str, last_login: timestamp}, Tool B inputs: {uid: str}, and Tool C inputs: {date: timestamp}. Then, if a semantic mapping is detected between the input uid of Tool B and the output user_id of Tool A (determined through schema aliases or semantic alignment), a directed edge A->B is established, and similarly, a directed edge A->C is established. The physical meaning of this directed edge is: the execution of the second tool depends on the output of the first tool; that is, the first tool must be executed before the second tool, and the output of the first tool needs to be used as an input parameter for the second tool.
[0042] Step S24: Treat each sub-step in the structured execution plan as a node, and generate a directed acyclic graph based on the established directed edges; wherein, the attributes of each node include the tool identifier bound to each sub-step.
[0043] In this embodiment, each sub-step in the structured execution plan is mapped to a node in the graph. This means that the number of nodes in the graph equals the number of sub-steps in the execution plan, and each node represents a specific task unit to be executed. Simultaneously with node mapping, attributes are defined for each node, including the tool identifier bound to that sub-step. This attribute allows each node in the graph to be associated with a specific tool capability, providing a clear basis for invocation in subsequent execution phases.
[0044] Furthermore, based on the aforementioned directed edges, these nodes are connected to form a complete directed acyclic graph. The direction of the directed edges represents the constraints on data flow and execution order. Since the dependencies established in this system are based on the actual input and output requirements of the tools, and there are no circular dependencies (a tool cannot depend on its own output), the generated graph is necessarily acyclic.
[0045] Step S25: Use the topological sorting algorithm to divide the nodes in the directed acyclic graph into layers to generate the execution path of the directed acyclic graph; where nodes without prerequisite dependencies are located in the first layer, and nodes that depend on the output results of other tools are located in subsequent layers.
[0046] In this embodiment, a topological sorting algorithm is further used to sort and stratify all nodes in the directed acyclic graph. Topological sorting is a sorting algorithm for directed acyclic graphs. Its core idea is to arrange the nodes in the graph into a linear sequence such that for each directed edge u→v, node u is placed before node v.
[0047] The specific layering rules are as follows: Nodes with no prerequisite dependencies, i.e., nodes with no directed edges pointing to them (in-degree of 0), are placed in the first layer. These nodes do not depend on the output of any other tools and can therefore be executed immediately; while nodes that depend on the output of other tools, i.e., nodes with at least one directed edge pointing to them (in-degree greater than 0), are placed in subsequent layers. The execution of these nodes must wait for their prerequisite nodes to complete execution and provide the corresponding output.
[0048] Through this layered processing, the original directed acyclic graph is transformed into a hierarchical execution path. Nodes in the first layer can execute in parallel. After each node completes its execution, its result is distributed to the next layer, triggering the execution of that next layer. In the example above, the layered result is: Layer 1: Tool A (no prerequisites, executes immediately); Layer 2: Tools B and C (dependent on the output of A).
[0049] Step S26: When executing each sub-step according to the directed acyclic graph execution path, extract complete tool definition information from the preset tool library based on the tool identifier bound to each sub-step, and construct an atomic execution environment for each sub-step that only includes the corresponding tool definition information, so as to execute each sub-step in the atomic execution environment, and then pass the execution result to the downstream sub-step according to the directed acyclic graph execution path.
[0050] In this embodiment, during the execution phase, the system no longer uses the simplified capability summary, but instead extracts complete tool definition information from a preset tool library based on the tool identifier bound to the current sub-step. This process can be called backward registration. Figure 4 As shown. In the backward registration phase, the system registers the complete schema of the tools, specifically including: tool name, description, input parameters, output parameters, usage notes, and reference examples. Unlike the metadata registration in the forward registration phase, the core purpose of the backward registration phase is to maximize the accuracy of the parameters generated by the model. By providing a complete parameter structure, validation rules, and calling examples, the execution agent can accurately understand the parameter requirements and return format of each tool, thereby generating accurate tool calling instructions. The system utilizes the complete tool definition information to construct an atomic execution environment for the current sub-step, containing only the tool definition information. In this environment, the execution agent's context only contains the tools necessary for the current step; all other irrelevant tools are forcibly hidden.
[0051] Furthermore, the above method also includes: traversing all nodes in the execution path of the directed acyclic graph; if the input parameters of the first node do not depend on the output of the second node, and the input parameters of the second node do not depend on the output of the first node, then it is determined that the first node and the second node have the conditions for parallel execution; during the execution phase, multiple nodes that have the conditions for parallel execution are simultaneously initiated with call requests, and after each node returns the execution result, data aggregation is performed.
[0052] In other words, the system also possesses parallel execution optimization capabilities. Specifically, the system traverses all nodes in the execution path of the directed acyclic graph and performs the following parallelism determination: if the input parameters of the first node do not depend on the output of the second node, and the input parameters of the second node do not depend on the output of the first node, then the first and second nodes are deemed to have the conditions for parallel execution. The essence of this determination is to detect whether there is a bidirectional data dependency between the two nodes. Only when the two nodes do not depend on each other's output can they safely execute in parallel. Correspondingly, during the execution phase, for all nodes determined to have the conditions for parallel execution, the system simultaneously initiates a call request, which means that the tools corresponding to these nodes are executed concurrently, rather than the serial waiting in the traditional mode. After all parallel nodes return execution results, the system uniformly performs data aggregation. The aggregated data is used on the one hand for preparing the input parameters of subsequent dependent nodes, and on the other hand for synthesizing the final answer or subsequent processing steps. Through this parallel execution mechanism, the system can maximize the utilization of computing resources and significantly shorten the response time of complex multi-tool tasks while ensuring the correctness of data dependencies.
[0053] Furthermore, the above method also includes: for isolated nodes in a directed acyclic graph that have no dependencies, these isolated nodes are placed in an independent execution thread pool and executed in parallel with the other nodes. In a directed acyclic graph, there may be some isolated nodes that have no dependencies. It should be noted that isolated nodes are characterized by having neither outgoing edges pointing to other nodes nor incoming edges pointed to by other nodes (i.e., both in-degree and out-degree are 0). They represent independent task units whose execution does not depend on the output of any other tool, and their output is not used by other tools.
[0054] In this embodiment, for isolated nodes in a directed acyclic graph that have no dependencies, the system places them in an independent execution thread pool for parallel execution with the other nodes. Since there are no data dependencies between isolated nodes, they can also execute in parallel without waiting for each other. This mechanism further improves the system's execution efficiency. Isolated nodes are not subject to the order constraints of the main execution path but are processed as independent parallel tasks, avoiding idle waiting time caused by the execution order of the main path.
[0055] In this embodiment, the specific processes of steps S21 and 22 can be referred to the corresponding contents disclosed in the foregoing embodiments, and will not be repeated here.
[0056] As can be seen, in the planning phase, this application minimizes the context token occupation by registering only lightweight metadata of the tools in advance, while giving the task planner a global tool view, significantly improving the reliability of task decomposition and tool binding. In the execution phase, by registering the complete schema of the tools on demand in the backward phase and building an atomic execution environment for each sub-step, the context interference of irrelevant tools is eliminated, fundamentally solving the problems of model illusion and tool mis-invocation, and improving the invocation accuracy. Furthermore, by constructing a directed acyclic graph by parsing the semantic dependencies between the input and output of the tools, and by using topological sorting and hierarchical execution combined with parallelism determination and independent thread pool processing of isolated nodes, the traditional serial execution mode is transformed into an efficient parallel execution flow, which greatly shortens the response time of complex multi-tool tasks, while ensuring the correct transmission and aggregation of data dependencies.
[0057] See Figure 5 As shown in the figure, this application also discloses a tool invocation device, including: The execution plan generation module 11 is used to decompose the user instruction into a structured execution plan including at least one sub-step based on the capability summary information of each tool in the preset tool library when the user instruction is received, and to bind a corresponding tool identifier to each sub-step. The execution path generation module 12 is used to parse the input and output data structures of the tools corresponding to each sub-step in the structured execution plan to obtain the dependencies between tools, and generate a directed acyclic graph execution path based on the dependencies between tools. Execution module 13 is used to extract complete tool definition information from the preset tool library based on the tool identifier bound to each sub-step when executing each sub-step according to the directed acyclic graph execution path, and to construct an atomic execution environment for each sub-step that only includes the corresponding tool definition information, so as to execute each sub-step in the atomic execution environment, and then pass the execution result to the downstream sub-step according to the directed acyclic graph execution path.
[0058] As can be seen, this application generates a structured execution plan based on tool capability summaries and binds tool identifiers, enabling task planning to have a global tool perspective, ensuring precise matching between execution steps and target tools, and improving the reliability and executability of task planning. Furthermore, by parsing the tool input and output data structures, the dependencies between tools are obtained, and a directed acyclic graph execution path is constructed based on this. This clarifies the dependencies between tools and optimizes the execution process, allowing tools with dependencies to be automatically linked and tools without dependencies to be executed in parallel according to the structure. Finally, when executing each sub-step according to the directed acyclic graph execution path, this application constructs an atomic execution environment for each step containing only the corresponding tool definition information, thereby achieving lightweight and isolated execution context, shielding irrelevant tool interference, reducing context redundancy, avoiding erroneous tool calls, and significantly improving tool call accuracy. Simultaneously, by orderly transmitting execution results according to the directed acyclic graph path, the overall accuracy, stability, and execution efficiency of agent tool scheduling are improved.
[0059] Furthermore, embodiments of this application also provide an electronic device. Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0060] Figure 6 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the tool invocation method disclosed in any of the foregoing embodiments.
[0061] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0062] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.
[0063] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. It can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including computer programs capable of performing the tool invocation methods executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include image quality control strategies collected by the electronic device 20, etc.
[0064] Furthermore, this application also discloses a storage medium storing a computer program, which, when loaded and executed by a processor, implements the tool invocation method steps disclosed in any of the foregoing embodiments.
[0065] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0066] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0067] The above provides a detailed description of the tool invocation method, apparatus, device, and storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for invoking a tool, characterized in that, include: Upon receiving a user instruction, the user instruction is decomposed into a structured execution plan including at least one sub-step based on the capability summary information of each tool in the preset tool library, and a corresponding tool identifier is bound to each sub-step. The input and output data structures of the tools corresponding to each sub-step in the structured execution plan are parsed to obtain the dependencies between the tools, and a directed acyclic graph execution path is generated based on the dependencies between the tools. When executing each sub-step according to the directed acyclic graph execution path, complete tool definition information is extracted from the preset tool library based on the tool identifier bound to each sub-step, and an atomic execution environment containing only the corresponding tool definition information is constructed for each sub-step, so as to execute each sub-step in the atomic execution environment, and then the execution result is passed to the downstream sub-step according to the directed acyclic graph execution path.
2. The tool invocation method according to claim 1, characterized in that, The capability summary information includes the tool name, function description, and core input / output parameters. The tool definition information includes the tool name, function description, complete input / output parameters, parameter validation rules, usage precautions, and reference examples.
3. The tool invocation method according to claim 1, characterized in that, The process of binding a corresponding tool identifier to each sub-step includes: When the same sub-step corresponds to multiple candidate tools, named entity recognition is performed on the input text of the current sub-step to extract the entity type; The entity type is matched with the input parameter definitions of each candidate tool, and a matching score is calculated. The target tool is determined from the plurality of candidate tools based on the matching score, and the tool identifier of the target tool is bound to the current sub-step.
4. The tool invocation method according to claim 3, characterized in that, The step of determining the target tool from the plurality of candidate tools based on the matching score includes: If multiple candidate tools have the same matching score, the multiple candidate tools are sorted according to a preset tool type sorting strategy to obtain a sorting result; wherein, the tool type sorting strategy is that resource acquisition tools have a higher priority than data processing tools, and data processing tools have a higher priority than data analysis tools. The tool that ranks first in the sorting results is identified as the target tool.
5. The tool invocation method according to any one of claims 1 to 4, characterized in that, The generation of directed acyclic graph execution paths based on the inter-tool dependencies includes: If a semantic match is detected between the output parameters of the first tool and the input parameters of the second tool, a directed edge is established from the first tool to the second tool. Each sub-step in the structured execution plan is treated as a node, and a directed acyclic graph is generated based on the established directed edges; wherein, the attributes of each node include the tool identifier bound to each sub-step; The topological sorting algorithm is used to divide the nodes in the directed acyclic graph into layers to generate the execution path of the directed acyclic graph; nodes without prerequisite dependencies are located in the first layer, and nodes that depend on the output results of other tools are located in subsequent layers.
6. The tool invocation method according to claim 5, characterized in that, Also includes: Traverse all nodes in the execution path of the directed acyclic graph. If the input parameters of the first node do not depend on the output of the second node, and the input parameters of the second node do not depend on the output of the first node, then it is determined that the first node and the second node have the conditions for parallel execution. During the execution phase, multiple nodes that meet the conditions for parallel execution are simultaneously initiated with call requests, and data aggregation is performed after each node returns the execution results.
7. The tool invocation method according to claim 6, characterized in that, Also includes: For isolated nodes in a directed acyclic graph that have no dependencies, these isolated nodes are placed in a separate execution thread pool and executed in parallel with the other nodes.
8. A tool calling device, characterized in that, include: The execution plan generation module is used to decompose the user instruction into a structured execution plan including at least one sub-step based on the capability summary information of each tool in the preset tool library when a user instruction is received, and to bind a corresponding tool identifier to each sub-step. The execution path generation module is used to parse the input and output data structures of the tools corresponding to each sub-step in the structured execution plan to obtain the dependencies between tools, and generate a directed acyclic graph execution path based on the dependencies between tools. The execution module is used to extract complete tool definition information from the preset tool library based on the tool identifier bound to each sub-step when executing each sub-step according to the directed acyclic graph execution path, and to construct an atomic execution environment for each sub-step that only includes the corresponding tool definition information, so as to execute each sub-step in the atomic execution environment, and then pass the execution result to the downstream sub-step according to the directed acyclic graph execution path.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein: The memory is used to store computer programs; The computer program is loaded and executed by the processor to implement the tool invocation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store computer-executable instructions, which, when loaded and executed by a processor, implement the tool invocation method as described in any one of claims 1 to 7.