Tool invocation method and apparatus
Patent Information
- Application Number
- CN202611007800.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0010]本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。
Smart Images

Figure CN122816722A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of large language model tool invocation, intelligent agent workflow execution control, and protocol adaptation technology. Background Technology
[0002] With the widespread application of Large Language Model (LLM) intelligent agent workflows, the use of tool calls is becoming increasingly common in the multi-stage intelligent agent workflow pipeline (intent classification → parameter collection → content generation → design → production) of AIGC (Artificial Intelligence Generated Content) multi-stage creation platforms. The granularity of tool information requirements varies significantly across different stages.
[0003] Currently, the industry mainly implements tool invocation in several ways: First, register the complete parameter schema (data structure) of all tools at once in the system prompt, and all requests share the same tool list; Second, the tool description includes two granularities: name, description overview and schema details; Third, wrap external service calls with code to mainly implement post-processing functions. Summary of the Invention
[0004] This disclosure provides a tool invocation method and apparatus.
[0005] In a first aspect, embodiments of this disclosure propose a tool invocation method, comprising: receiving a user request and initiating a multi-stage intelligent agent workflow; constructing tool hint information required for invoking a corresponding large language model based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow; constructing business parameters based on the tool hint information using the large language model; in response to detecting the output of the tool invocation intent from the large language model, triggering a pre-hook before tool execution, performing infrastructure parameter injection, protocol format conversion, and tool type routing judgment based on the business parameters, and executing the tool to generate tool execution results; and triggering a post-hook after tool execution, extracting protocol fields from the tool execution results, and storing the protocol fields.
[0006] Secondly, embodiments of this disclosure propose a tool invocation device, comprising: a workflow initiation module configured to receive a user request and initiate a multi-stage intelligent agent workflow; a prompt information construction module configured to construct tool prompt information required for invoking a corresponding large language model based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow; a business parameter construction module configured to construct business parameters based on the tool prompt information using the large language model; a pre-hook triggering module configured to trigger a pre-hook before tool execution in response to detecting the output tool invocation intent of the large language model, perform infrastructure parameter injection, protocol format conversion, and tool type routing judgment based on the business parameters, and execute the tool to generate tool execution results; and a post-hook triggering module configured to trigger a post-hook after tool execution is completed, extract protocol fields from the tool execution results, and store the protocol fields.
[0007] Thirdly, embodiments of this disclosure provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described in the first aspect.
[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described in the first aspect.
[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0010] The key or essential features of the embodiments disclosed herein are not intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. Wherein: Figure 1 This is a flowchart of one embodiment of the tool invocation method according to this disclosure; Figure 2 This is a flowchart of yet another embodiment of the tool invocation method according to this disclosure; Figure 3 This is a three-layer architecture diagram for optimizing the calling of large language model tools; Figure 4 This is the data flow diagram for the pre-hook and post-hook; Figure 5 This is a flowchart of the node strategy differentiation injection decision process; Figure 6 It is a flowchart for the closed-loop process of accumulating experience from failures; Figure 7 This is a schematic diagram of the structure of one embodiment of the tool calling device according to the present disclosure; Figure 8 This is a block diagram of an electronic device used to implement the tool invocation method of the embodiments of this disclosure. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0014] Figure 1 A flow 100 of an embodiment of a tool invocation method according to the present disclosure is shown. The tool invocation method includes the following steps: Step 101: Receive user request and start multi-stage agent workflow.
[0015] In this embodiment, the execution entity of the tool invocation method can receive user requests and initiate a multi-stage intelligent agent workflow.
[0016] User requests can refer to natural language task instructions input by the user, such as "Help me generate a short technology video." A multi-stage intelligent agent workflow can be a pipeline that sequentially executes complex tasks by classifying them according to intent, collecting parameters, generating content, designing, and producing. Each node can be independently configured with tool invocation strategies, making it suitable for scenarios such as AIGC multi-stage creation platforms.
[0017] Step 102: Based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow, construct the tool hint information required for the corresponding large language model call.
[0018] In this embodiment, the aforementioned execution entity can construct the tool hint information required for calling the corresponding large language model based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow.
[0019] The workflow execution engine can invoke the dynamic tooltip construction function based on the tool injection strategy configuration of the current execution node to construct the tooltip information for this large language model invocation. The tool injection strategy can be a configuration item on the workflow node, including three modes: no injection, overview injection, and full document injection, used to control the granularity of tool information injected into the large language model. The tooltip information can be a tool description context provided to the large language model to guide the model to output a legitimate and accurate tool invocation intent.
[0020] In some embodiments, during the initialization phase, tool information is registered to the tool registry. The tool information includes the tool name, a description overview, full documentation, and a tool call error correction suggestion set.
[0021] The initialization phase can refer to the system startup or service loading phase. The tool registry can be a configuration center or database storing all tool metadata. Tool information can be a complete description of the tool. This complete documentation includes parameter schemas, calling formats, return value descriptions, etc. A tool call error correction set can be used to store historical failure experiences to prevent the model from repeating the same mistakes.
[0022] In some embodiments, a dynamic tooltip build function is invoked to traverse the active tools in the tool registry and selectively inject at least one of the tool name, description overview, and full documentation into the invocation context of the large language model, according to the tool registration policy configuration.
[0023] The dynamic tooltip constructor can iterate through the active toolset and selectively inject tool names, descriptions, and complete parameter definitions according to strategy levels (no injection / overview / full documentation). The dynamic tooltip constructor can be a function that automatically concatenates tool information based on node strategies. Active tools refer to the set of tools allowed to be invoked by the current node. Selective injection allows for flexible control of the tooltip length based on node requirements, reducing token consumption, avoiding information overload, and improving model understanding efficiency.
[0024] In some embodiments, in response to the tool registration strategy being configured as a full document injection strategy and the number of active tools being greater than 1, the tool name and description overview are injected in the first round. After receiving the intermediate results of the tool selected by the large language model, the full document is injected in subsequent rounds.
[0025] Under a complete documentation injection strategy, if multiple tools are available, a progressive disclosure approach can be adopted to avoid excessively long prompts and difficulties in model selection: in the first round, only the tool name and function overview are provided, and the complete parameter structure is supplemented after the model selects the target tool. This can significantly reduce token consumption and improve the accuracy of model selection.
[0026] Taking multi-node differentiated injection as an example, if the system registers 50 tools and a user inputs "Help me make a short technology video," compared to the full document mode, the total token consumption is reduced. Detailed execution steps are as follows: First, when the workflow engine builds a large language model request, it calls the `buildToolContext(nodeId, activeTools)` function, passing in the current node identifier and the active toolset. `buildToolContext(nodeId, activeTools)` is the tool context function, used to read the tool injection strategy based on the node identifier, traverse the active toolset, dynamically build and return the tool information context required by the large language model according to the differentiated strategy, supporting three injection modes: no injection, overview, and full documentation, as well as progressive disclosure in multi-tool scenarios.
[0027] Then, the function reads the injection strategy field from the node configuration table (enumerated values: NONE (no injection mode) / OVERVIEW (overview mode) / FULL_DOC (full document mode)), assembles the tool description list according to the strategy, and then appends the system prompt words. <tools>Tag area.
[0028] Finally, in full document mode and when the number of active tools is greater than 1, the first round of calls only writes the tool name and overview description. After receiving the intermediate results of the tool selected by the large language model, the second round of calls appends the complete parameter JSON (JavaScript Object Notation) schema of the tool, ensuring that the large language model fills in the parameter fields only when there is sufficient information.
[0029] Step 103: Construct business parameters based on tooltips using a large language model.
[0030] In this embodiment, the aforementioned execution entity can utilize a large language model to construct business parameters based on tooltip information.
[0031] Business parameters can be the input parameters required by the tool to execute business logic, such as the theme, duration, style, and resolution of the generated video. The large language model can understand the tool parameter specifications based on the injected tool hints, transforming the user's natural language requirements into structured business parameter objects, preparing for subsequent tool calls.
[0032] Step 104: In response to the detection of the large language model output tool call intent, a pre-hook is triggered before the tool is executed. Based on the business parameters, infrastructure parameter injection, protocol format conversion and tool type routing judgment are performed, and the tool is executed to generate the tool execution result.
[0033] In this embodiment, in response to detecting the intent to call the large language model output tool, the aforementioned execution entity can trigger a pre-hook before the tool is executed, perform infrastructure parameter injection, protocol format conversion, and tool type routing judgment based on business parameters, and then execute the tool to generate the tool execution result.
[0034] The tool invocation intent can be a structured invocation command output by a large language model, containing the tool name and business parameters. Pre-hooks can be protocol processing logic that runs before tool execution, used to convert the general business parameters output by the model into a format recognizable by external services. Infrastructure parameters can include callback addresses, tracing IDs, service identifiers, authentication information, etc. Protocol format conversion can include JSON to Base64, private protocol encapsulation, etc. Tool type routing can be used to select the corresponding executor or service interface based on the tool type, ultimately invoking the external tool and returning the execution result.
[0035] Step 105: After the tool finishes execution, trigger the post-hook to extract the protocol field from the tool execution result and store the protocol field.
[0036] In this embodiment, the aforementioned execution entity can trigger a post-hook after the tool is executed, extract the protocol field from the tool execution result, and store the protocol field.
[0037] Post-hooks can be result processing logic executed after the tool completes its execution, triggered regardless of success or failure. Protocol fields can include status codes, extended fields (ext), task identifiers, execution messages, etc. Storage purposes can include workflow status tracking, asynchronous task recovery, failure analysis, and experience accumulation.
[0038] In some embodiments, the session identifier, tool name, and call sequence number are combined to generate a storage key; the storage key and protocol field are written to persistent storage in the form of key-value pairs.
[0039] The session identifier (sessionId) uniquely identifies a user request. The tool name (toolName) identifies the currently invoked tool. The call sequence number (callSeq) distinguishes multiple calls to the same tool. The key combination format is sessionId:toolName:callSeq, ensuring global uniqueness and facilitating quick task location during subsequent callbacks, polling, and status queries.
[0040] In some embodiments, a post-hook is used to extract extended parameter fields from the tool execution result and identify the extended parameter fields to determine the execution mode of the current tool call; in response to the execution mode being synchronous, the next execution node of the multi-stage agent workflow continues to be executed according to the tool execution result; in response to the execution mode being asynchronous, an asynchronous task identifier is extracted from the tool execution result, the execution of the multi-stage agent workflow is paused until a callback result or polling result corresponding to the asynchronous task identifier is received, and then the execution of the multi-stage agent workflow is resumed.
[0041] Extended fields (ext) can be custom fields in the tool's returned results used to identify the execution mode. Synchronous mode means that the tool returns the final result immediately after being invoked, and the workflow can directly proceed to the next node. Asynchronous mode means that the tool returns a task ID, and execution needs to wait for the callback or polling to complete before resuming, which is suitable for time-consuming tasks such as video generation and image rendering.
[0042] Taking the protocol adaptation of pre-hooks and post-hooks as an example, an external content production platform requires a proprietary protocol format. The large language model only outputs standard business parameters ({"topic": "Technology Trends", "style": "Professional", "duration": 60}). The pre-hook automatically performs Base64 encoding, injects service identifiers and tracing headers, constructs the proprietary protocol format, and passes it to the tool's execution function. The large language model's prompts do not require any protocol details. The detailed execution steps are as follows: First, after the execution framework detects that the large language model produces a tool_call structure, it sequentially calls the pre-hook chain before the actual execution tool.
[0043] Then, the pre-hook receives the raw parameter object of the large language model, executes Base64.encode(JSON.stringify(bizParams)), generates the data field, and simultaneously reads the service identifier (serviceId) from the environment configuration, the trace identifier (traceId) and the segment identifier (spanId) from the link context to inject the request headers, and outputs the private format {serviceId, data, headers}.
[0044] Finally, after the tool completes its execution, the post-hook extracts the asynchronous task identifier from the ext field of the response body and writes it to the session-level task persistent storage (key: sessionId:toolName:callSeq) for lazy recovery detection queries.
[0045] The tool invocation method provided in this disclosure significantly reduces token consumption through node strategy differential injection, decouples the protocol and prompt words through pre- and post-hooks, and ensures stable workflow operation through adaptive execution mode. It solves problems such as prompt word expansion, protocol coupling, unstable invocation, and inability to accumulate experience in the prior art.
[0046] Figure 2 A flow 200 of yet another embodiment of the tool invocation method according to this disclosure is shown. The tool invocation method includes the following steps: Step 201: Receive user request and start multi-stage agent workflow.
[0047] Step 202: Based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow, construct the tool hint information required for the corresponding large language model call.
[0048] Step 203: Construct business parameters based on tooltip information using a large language model.
[0049] Step 204: In response to the detection of the large language model output tool call intent, a pre-hook is triggered before the tool is executed. Based on the business parameters, infrastructure parameter injection, protocol format conversion and tool type routing judgment are performed, and the tool is executed to generate the tool execution result.
[0050] Step 205: After the tool finishes execution, trigger the post-hook to extract the protocol field from the tool execution result and store the protocol field.
[0051] In this embodiment, the specific operations of steps 201-205 can be referred to Figure 1 The relevant descriptions of steps 101-105 in the corresponding embodiments will not be repeated here.
[0052] Step 206: In response to the tool call failure, extract the failure context.
[0053] In this embodiment, in response to a tool call failure, the aforementioned execution entity can extract the failure context.
[0054] Tool call failures can include incorrect parameters, network errors, service error codes, authentication failures, etc. The failure context can include the tool name, error code, error message, large language model input parameters (after anonymization), timestamp, etc., for subsequent analysis of the failure cause.
[0055] Step 207: Based on the failure context, generate error correction prompts for candidate tool calls.
[0056] In this embodiment, the aforementioned execution entity can generate error correction prompts for candidate tool calls based on the failure context.
[0057] When a tool call fails, the failure experience collector can automatically extract the failure context, generate candidate tool call error correction prompts, and mark them as pending review. Specifically, candidate error correction prompts can be automatically generated by the failure experience collector, such as "the parameter image_count cannot be 0, the minimum value is 1," to guide the large language model to avoid similar errors in subsequent calls.
[0058] Step 208: Determine the target tool call error correction prompt from the candidate tool call error correction prompts and add it to the tool call error correction prompt set.
[0059] In this embodiment, the aforementioned execution entity can generate error correction prompts for candidate tool calls based on the failure context.
[0060] Once the error correction prompts for the target tool call are manually reviewed and confirmed to be effective, they will be marked as ACTIVE and written into the corresponding error correction prompt set of the tool, forming a reusable failure experience library and enabling the continuous evolution of tool call capabilities.
[0061] Step 209: In response to the subsequent large language model call to the corresponding tool in the multi-stage intelligent agent workflow, inject the tool call error correction prompt set into the corresponding tool prompt information.
[0062] In this embodiment, in response to the subsequent large language model call to the corresponding tool in the multi-stage intelligent agent workflow, the aforementioned execution entity can inject the tool call error correction prompt set into the corresponding tool prompt information.
[0063] The next time the same tool is called, the dynamic tooltip constructor can automatically append error correction suggestions to the end of the tool description, allowing the large language model to construct parameters based on historical experience, significantly improving the success rate of the call and avoiding repeated mistakes.
[0064] Taking the failure experience loop as an example, on the first call to the image generation tool, the large language model transmits 0 images, and the server returns an error. The failure experience collector automatically extracts candidate tool call error correction prompts, which are then manually reviewed and added to the tool call error correction prompt set. Subsequent calls by the large language model automatically refer to these prompts to avoid repeating the same error. The detailed execution steps are as follows: First, when the post-hook detects that the tool returns a status code other than 2xx or a business error code, it triggers the failure experience collector.
[0065] Then, the collector extracts the tool name, error code (such as ERR_PARAM_INVALID), and large language model input parameters from the tool call records (the field name and type are retained after desensitizing sensitive fields), automatically generates candidate tool call error correction prompts (such as "Note: the image_count parameter is not allowed to be 0, the minimum value is 1"), writes them to the review queue and sets status=PENDING (pending review status).
[0066] Finally, after the user clicks "Approved" in the management backend, the status is updated to ACTIVE. The next time the buildToolContext calls the tool, all ACTIVE status tool call error correction prompts will be automatically appended to the end of the tool description with newlines and injected into the large language model context.
[0067] The tool invocation method provided in this disclosure constructs a complete closed loop of experience accumulation, namely "failure capture → automatic generation → manual review → automatic injection", which enables tool invocation to evolve itself and continuously improve accuracy and stability.
[0068] The tool invocation method provided in this disclosure can be applied not only to AIGC multi-stage creation platforms, but also extended to all scenarios of multi-stage intelligent agents invoking external services, such as AI (Artificial Intelligence) code generation (invoking compilation / testing / deployment toolchains), AI customer service (invoking order / logistics interfaces), and AI data analysis (invoking SQL (Structured Query Language) / chart generation services).
[0069] Taking AI code generation as an example, the detailed execution steps are as follows: First, register three types of tools in the tool registry: compile_tool (compilation tool), unit_test_tool (unit testing tool), and deploy_tool (deployment tool). Configure the node strategies as follows: Requirements analysis node (no injection) → Solution generation node (overview) → Code generation node (complete documentation, only inject compile_tool) → Test node (complete documentation, only inject unit_test_tool).
[0070] Then, the pre-hook is responsible for injecting CI / CD (Continuous Integration / Continuous Deployment) authentication token and build environment parameters, and the post-hook extracts the compilation / test result code and writes it to the state storage.
[0071] Finally, if compile_tool returns a compilation error, the failure experience loop automatically generates candidate pitfall avoidance tips such as "Do not use relative paths in import statements", which are then reviewed and incorporated into the context constraints for the next code generation.
[0072] Figure 3 This diagram illustrates a three-layer architecture for optimizing tool invocation within a large language model. The three layers are: Visual Layer 301, Protocol Layer 302, and Evolution Layer 303. Visual Layer 301 enables differentiated node injection strategies, including three modes: no injection, overview injection, and full document injection, controlling the scope of tool information perceptible to the large language model. Protocol Layer 302 achieves protocol adaptation through pre-hooks and post-hooks, completing infrastructure parameter injection, format conversion, result extraction, and execution mode determination. Evolution Layer 303 enables an automatic closed-loop accumulation of failure experience, including automatic extraction, manual review, and automatic injection, forming a continuous optimization mechanism. This three-layer architecture comprehensively optimizes tool invocation performance from three dimensions: information visibility, protocol adaptability, and self-evolutionary capability.
[0073] It should be noted that the visual layer 301 can also be replaced by using natural language to restrict the tool; the protocol layer 302 can also be replaced by processing the protocol within the tool; and the evolution layer 303 can also be replaced by manually updating the tool documentation periodically, without any limitations here.
[0074] Figure 4 The data flow diagram for the pre-hooks and post-hooks is shown. The data flow for the pre-hooks and post-hooks is as follows: Large language model output parameters 401 → Pre-hook 402 → Tool execution 403 → Post-hook 404 → State update 405, sequentially passing through the semantic layer, protocol layer, business logic layer, protocol layer, and persistence layer. Among them, the large language model output parameters 401 belongs to the semantic layer and only contains business parameters. Pre-hook 402 belongs to the protocol layer and is responsible for parameter enhancement, protocol encapsulation, and routing judgment. Tool execution 403 completes the business logic of external service calls. Post-hook 404 belongs to the protocol layer and is responsible for result parsing, field extraction, execution pattern recognition, and state processing. State update 405 writes task information to persistent storage to support the continued execution of the workflow.
[0075] Figure 5 The flowchart illustrates the node strategy differential injection decision process. Step 501 involves a user request arriving at the workflow node. Step 502 involves reading the node's tool injection strategy configuration. Step 503 is the injection strategy judgment step, dividing the process into three execution branches: no injection, overview injection, and full document injection. In the no injection branch, step 504 is executed, leaving the tool field empty. In the overview injection branch, step 505 is executed, iterating through active tools and writing their names and descriptions. In the full document injection branch, step 506 is executed, judging the multi-tool scenario. If the judgment result is negative, step 507 is executed, directly injecting the full schema; if the judgment result is positive, steps 508 (first round only overview injection), 509 (large language model tool selection), and 510 (subsequent rounds supplementing with full schema injection) are executed sequentially. Step 511 constructs the large language model call context; the execution results of the no injection, overview injection, and full document injection branches are all incorporated into this step, ultimately completing the construction of tooltips adapted to the current node.
[0076] Figure 6 A flowchart illustrating the closed-loop process for accumulating failure experience is shown. The closed-loop process for accumulating failure experience may include the following steps: Step 601, tool call failed.
[0077] This failure can include various call exception scenarios such as incorrect parameter format, abnormal parameter value, service authentication failure, network error, and error code returned by external service.
[0078] Step 602: The post-hook captures the failure context.
[0079] The failure context can include key information such as the tool name, failure error code / error message, de-identified parameters transmitted by the large language model, and timestamp, which are used to fully reconstruct the failure scenario of this call.
[0080] Step 603: The failure experience collector generates candidate pitfall avoidance tips.
[0081] Among them, the failure experience collector can automatically generate candidate pitfall avoidance suggestions based on the captured failure context, which can be used to guide large language models to avoid similar calling errors.
[0082] Step 604: Write to the pending review queue (status = PENDING).
[0083] The generated candidate pitfall avoidance tips are written to a pending review queue with a status of PENDING.
[0084] Step 605, manual review.
[0085] If the review is approved, the candidate pitfall avoidance tip will be added to the tool call error correction tip set and marked as ACTIVE. If the review is rejected, the candidate pitfall avoidance tip will be deleted and marked as REJECTED.
[0086] Step 606, the next time the large language model calls this tool.
[0087] This refers to the trigger scenario where the large language model subsequently calls the corresponding tool again.
[0088] Step 607: The dynamic tooltip construction function automatically reads the pitfall avoidance hint set.
[0089] Among them, the dynamic tooltip constructor can automatically read the pitfall avoidance tips that have already taken effect in the tool call error correction prompt set.
[0090] Step 608: Concatenate the pitfall avoidance tips to the end of the tool description and inject them into the large language model.
[0091] Specifically, the pitfall avoidance tips read are appended to the end of the tool description and injected into the calling context of the large language model.
[0092] Step 609: Use the large language model reference to avoid pitfalls in the constructor to prevent similar errors.
[0093] Among them, the pitfall avoidance prompts of the large language model reference injection construct business parameters to avoid similar call errors, thereby forming a complete self-evolving closed loop of tool call "failure capture - experience generation - manual review - automatic reuse - error avoidance".
[0094] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a tool invocation device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0095] like Figure 7 As shown, the tool invocation device 700 of this embodiment may include: a workflow initiation module 701, a prompt information construction module 702, a business parameter construction module 703, a pre-hook triggering module 704, and a post-hook triggering module 705. The workflow initiation module 701 is configured to receive a user request and initiate a multi-stage intelligent agent workflow; the prompt information construction module 702 is configured to construct the tool prompt information required for the corresponding large language model invocation based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow; the business parameter construction module 703 is configured to construct business parameters based on the tool prompt information using the large language model; the pre-hook triggering module 704 is configured to trigger a pre-hook before tool execution in response to detecting the large language model outputting a tool invocation intent, perform infrastructure parameter injection, protocol format conversion, and tool type routing judgment based on the business parameters, and execute the tool to generate the tool execution result; the post-hook triggering module 705 is configured to trigger a post-hook after tool execution is completed, extract protocol fields from the tool execution result, and store the protocol fields.
[0096] In this embodiment, the specific processing and technical effects of the workflow initiation module 701, prompt information construction module 702, business parameter construction module 703, pre-hook trigger module 704, and post-hook trigger module 705 in the tool invocation device 700 can be referred to respectively. Figure 1 The relevant descriptions of steps 101-105 in the corresponding embodiments will not be repeated here.
[0097] In some optional implementations of this embodiment, the tooltip building module 702 is further configured to: call the dynamic tooltip building function to traverse the active tools in the tool registry, and selectively inject at least one of the tool name, description overview and full documentation into the calling context of the large language model according to the tool registration strategy.
[0098] In some optional implementations of this embodiment, the prompt information construction module 702 is further configured to: in response to the tool registration strategy being configured as a full document injection strategy and the number of active tools being greater than 1, to inject the tool name and description overview in the first round, and after receiving the intermediate results of the tool selected by the large language model, to inject the full document in subsequent rounds.
[0099] In some optional implementations of this embodiment, the post-hook trigger module 705 is further configured to: combine the session identifier, tool name and call sequence number to generate a storage key; and write the storage key and protocol field into persistent storage in the form of key-value pairs.
[0100] In some optional implementations of this embodiment, the tool invocation device 700 further includes: a tool information registration module, configured to register tool information to a tool registry during the initialization phase. The tool information includes the tool name, a description overview, complete documentation, and a tool invocation error correction prompt set.
[0101] In some optional implementations of this embodiment, the tool invocation device 700 further includes: an error correction prompt generation module, configured to extract the failure context in response to a tool invocation failure; generate candidate tool invocation error correction prompts based on the failure context; determine the target tool invocation error correction prompt from the candidate tool invocation error correction prompts, and add it to the tool invocation error correction prompt set.
[0102] In some optional implementations of this embodiment, the tool invocation device 700 further includes: an error correction prompt injection module, configured to inject the tool invocation error correction prompt set into the corresponding tool prompt information in response to subsequent large language model invocations of the multi-stage intelligent agent workflow.
[0103] In some optional implementations of this embodiment, the tool invocation device 700 further includes: a subsequent execution module, configured to use a post-hook to extract extended parameter fields from the tool execution result and identify the extended parameter fields to determine the execution mode of the current tool invocation; in response to the execution mode being synchronous execution mode, to continue executing the next execution node of the multi-stage intelligent agent workflow according to the tool execution result; in response to the execution mode being asynchronous execution mode, to extract an asynchronous task identifier from the tool execution result, to pause the execution of the multi-stage intelligent agent workflow until a callback result or polling result corresponding to the asynchronous task identifier is received, and then resume the execution of the multi-stage intelligent agent workflow.
[0104] The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.
[0105] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0106] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0107] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0108] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0109] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the tool invocation method. For example, in some embodiments, the tool invocation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the tool invocation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the tool invocation method by any other suitable means (e.g., by means of firmware).
[0110] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0111] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0112] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0114] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0115] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0116] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.
[0117] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.< / tools>
Claims
1. A tool invocation method, comprising: Receive user requests and initiate a multi-stage intelligent agent workflow; Based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow, construct the tool hint information required for calling the corresponding large language model; Using the large language model, business parameters are constructed based on the tooltip information; In response to the detection of the large language model output tool invocation intent, a pre-hook is triggered before the tool is executed, and infrastructure parameter injection, protocol format conversion and tool type routing judgment are performed based on the business parameters, and the tool is executed to generate the tool execution result; After the tool completes execution, a post-hook is triggered to extract the protocol field from the tool execution result and store the protocol field.
2. The method according to claim 1, wherein, The step of constructing the tool hint information required for calling the corresponding large language model based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow includes: The dynamic tooltip constructor iterates through the active tools in the tool registry and selectively injects at least one of the tool name, description overview, and full documentation into the calling context of the large language model according to the tool registration strategy configuration.
3. The method according to claim 2, wherein, The step of constructing the tool hint information required for calling the corresponding large language model based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow also includes: In response to the tool registration strategy being configured as a full document injection strategy and the number of active tools being greater than 1, the first round of calls injects the tool name and description overview. After receiving the intermediate results of the tool selected by the large language model, the full document is injected in subsequent rounds.
4. The method according to claim 1, wherein, The storage of the protocol fields includes: Combine the session identifier, tool name, and call sequence number to generate a storage key; The storage key and the protocol field are written to persistent storage in the form of key-value pairs.
5. The method according to claim 1, wherein, The method further includes: During the initialization phase, tool information is registered to the tool registry. The tool information includes the tool name, a description overview, complete documentation, and a tool call error correction prompt set.
6. The method according to claim 5, wherein, The method further includes: In response to a tool call failure, extract the failure context; Based on the failure context, generate error correction prompts for candidate tool calls; The target tool call error correction prompt is determined from the candidate tool call error correction prompts and added to the tool call error correction prompt set.
7. The method according to claim 6, wherein, The method further includes: In response to subsequent large language model calls to corresponding tools in the multi-stage agent workflow, the tool call error correction prompt set is injected into the corresponding tool prompt information.
8. The method according to claim 1, wherein, The method further includes: Using the post-hook, the extended parameter field is extracted from the tool execution result, and the extended parameter field is identified to determine the execution mode of the current tool call; In response to the execution mode being synchronous execution mode, the next execution node of the multi-stage intelligent agent workflow continues to be executed based on the execution result of the tool; In response to the execution mode being asynchronous, the asynchronous task identifier is extracted from the tool's execution result, the execution of the multi-stage agent workflow is paused until a callback result or polling result corresponding to the asynchronous task identifier is received, and then the execution of the multi-stage agent workflow is resumed.
9. A tool calling device, comprising: The workflow initiation module is configured to receive user requests and initiate multi-stage agent workflows; The tooltip construction module is configured to construct the tooltip information required for calling the corresponding large language model based on the tool injection strategy configuration of the current execution node of the multi-stage intelligent agent workflow. The business parameter construction module is configured to construct business parameters based on the tooltip information using the large language model; The pre-hook triggering module is configured to respond to the detection of the large language model output tool call intention, trigger the pre-hook before the tool is executed, perform infrastructure parameter injection, protocol format conversion and tool type routing judgment based on the business parameters, and execute the tool to generate the tool execution result; The post-hook trigger module is configured to trigger the post-hook after the tool is executed, extract the protocol field from the tool execution result, and store the protocol field.
10. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
11. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method of any one of claims 1-8.
12. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.