Service end intelligent agent external tool asynchronous execution and context backfill method and system

CN122820141APending Publication Date: 2026-09-25CHONGQING NUOYUAN IND SOFTWARE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611239211.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

该方式将上下文衔接逻辑暴露给客户端,服务端无法统一管理工具状态、任务边界、消息顺序和将执行结果转换为可被大语言模型处理的工具返回内容并替换至内存上下文副本对应位置的过程,容易导致会话历史不一致

Benefits of technology

[0010]上述说明仅是本申请技术方案的概述,为了能够更清楚了解本申请的技术手段,而可依照说明书的内容予以实施,并且为了让本申请的上述和其它目的、特征和优点能够更明显易懂,以下特举本申请的具体实施方式。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820141A_ABST
    Figure CN122820141A_ABST
Patent Text Reader

Abstract

The application discloses a method and system for asynchronous execution and context backfilling of external tools of a server agent, and relates to the fields of agents and large models. The method comprises the following steps: configuring a list of external tools in the server agent in advance; in the case of a tool calling request generated by a large language model, if the tool to be called belongs to the list of external tools, the server generates a tool result identifier, and extracts calling parameters from the tool calling request; based on the tool result identifier and the calling parameters, writing the result into a result storage, and sending an external tool dispatch event to an external organizer through a streaming event; submitting an execution result through a result backfilling interface provided by the server, and updating the session state and result content in the result storage; the tool result identifier queries the session state in the result storage, and if the session state is completed, the complete result is materialized into tool return content. The application can use the tool result asynchronously generated by the external system of the agent in the same reasoning task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to intelligent agents and large model technology, and in particular to a method and system for asynchronous execution and context backfilling of external tools for server-side intelligent agents. Background Technology

[0002] With the application of Large Language Model (LLM) agents in enterprise systems, agents typically access external capabilities such as databases, search services, knowledge graphs, code executors, and business APIs through tool invocation mechanisms. Existing agent tool invocation patterns generally require the tool to execute synchronously within the agent's running process, and the tool's returned results are immediately added to the session context as tool messages for the large language model to continue inference.

[0003] However, in enterprise-level server-side intelligent agent scenarios, some tools are not suitable for synchronous execution directly within the Agent process. For example, graph reasoning, complex retrieval, simulation computing, heterogeneous computing tasks, proprietary business inference engines, and systems requiring independent authentication or resource isolation are often executed by systems outside the Agent. These external tools are characterized by long execution times, independent execution environments, asynchronous result generation, and cross-service call chains.

[0004] Existing technologies mainly include the following categories: The first type is the synchronous tool invocation scheme. After the model invokes the tool, the Agent server immediately executes the tool and writes the complete result into the context after waiting for the tool to return. This method is simple to implement, but it requires that the tool execution must be within the controllable range of the Agent server, making it difficult to integrate tools that are time-consuming or deployed on external systems. Simple task submission, simple callbacks, and simple polling cannot form a complete closed loop. External systems need to know the execution parameters and result identifiers, and the Agent server needs to know when to wait, which result to wait for, and how to inject the result into the model context.

[0005] The second type is the asynchronous task submission scheme. The agent submits the task to an external system and immediately returns the task number, allowing the user or subsequent processes to query the task result. This method is suitable for background tasks, but the external execution result cannot naturally enter the same large language model inference chain. The model in the current context typically only sees that the task has been submitted and cannot continue inference based on the result. For tools executed by the agent from an external system, existing asynchronous task schemes usually only let the model know that the task has been submitted, and cannot allow the model to wait for and use the external execution result within the same server-side inference process.

[0006] The third type is the client-side orchestration scheme. In this scheme, the client, upon receiving the model tool's invocation intent, calls the external system itself and then sends the result as new user input or supplementary information to the Agent. This approach exposes the context connection logic to the client, making it impossible for the server to uniformly manage tool states, task boundaries, message order, and the process of converting execution results into tool return content that can be processed by the large language model and replacing it in the corresponding position of the memory context copy. This can easily lead to inconsistencies in session history. In the client-side orchestration scheme, the client needs to understand the tool invocation protocol, external system invocation methods, session message formats, and context completion logic, increasing system complexity and reducing the server's control over session consistency. Summary of the Invention

[0007] This application provides a method and system for asynchronous execution and context backfilling of external tools by a server-side intelligent agent. This enables the intelligent agent to dispatch tool invocation intent to an external system, which then executes and backfills the results asynchronously. The server waits before the large language model reads the context and converts the complete result into tool return content for processing by the large language model, replacing the original tool message at the corresponding position in the memory context copy.

[0008] This application provides a method for asynchronous execution and context backfilling of external tools for server-side intelligent agents, including the following steps: A list of external asynchronous tools is pre-configured in the server-side intelligent agent. The list of external asynchronous tools is used to record the names of tools that are executed asynchronously by external systems independently of the server-side intelligent agent. When using a large language model to generate a tool call request, if the tool to be called belongs to the list of external asynchronous tools, the server generates a globally unique tool result identifier and extracts call parameters from the tool call request. The tool result identifier is used to associate the tool call of the large language model, the execution task of the external system, and the result storage record of the server. Based on the tool result identifier and call parameters, write them into the result storage, and send external tool dispatch events to the external orchestrator through streaming events. The streaming event is a structured data message containing a session identifier, tool result identifier, tool name and call parameters. The external orchestrator is a message middleware or server component that receives the streaming event and schedules external systems to execute tasks. The result storage also contains the corresponding session identifier and tool call status. The execution result is submitted through the result feedback interface provided by the server, and the tool call status and result content in the result storage are updated. The execution result is the result of the external system executing the tool logic based on the received dispatch event. In the same agent inference task, before reading the current task context and constructing the input of the large language model, the server queries the tool call status in the result storage according to the tool result identifier. If the tool call status is completed, the complete result is converted into the tool return content used for processing by the large language model, and the original tool message at the corresponding position in the memory context copy is replaced. The conversion and replacement operation only applies to the memory context copy sent to the large language model and does not modify the persistent session message chain.

[0009] This application also provides a server-side intelligent agent system, including: An external tool identification module is used to determine whether a tool call is an external tool based on a pre-configured list; an identifier generation module is used to generate a globally unique tool result identifier for an external tool call, wherein the external asynchronous tool list is used to record the names of tools executed asynchronously by external systems, independent of the server-side intelligent agent; The results storage module is used to maintain the execution status and results of tool calls; The streaming event dispatch module is used to send tool execution requests to external systems via streaming channels; The result backfilling interface module is used to receive the execution results from external systems and update the result storage. The context construction engine is used to identify the tool call status in the query result storage based on the tool result before model inference. If the tool call status is completed, the complete result is converted into tool return content for processing by the large language model and the original tool message at the corresponding position in the memory context copy is replaced. The conversion and replacement operation only applies to the memory context copy sent to the large language model and does not modify the persistent session message chain.

[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0011] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a basic flowchart illustrating the asynchronous execution and context backfilling method of the server-side intelligent agent external tool in an embodiment of this application. Detailed Implementation

[0012] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0013] This application provides a method for asynchronous execution and context backfilling of external tools for server-side intelligent agents. The method in this application involves the following parties: The server-side intelligent agent interacts with the large language model, for example, through conversations, and maintains the conversation message chain and result storage.

[0014] An external orchestrator is used to receive streaming events and schedule external systems to perform tasks.

[0015] External systems are third-party systems or distributed computing clusters that actually execute the tool logic.

[0016] Result storage is used to store tool result identifiers and their corresponding execution status and result content.

[0017] like Figure 1 As shown, the method in this application embodiment includes the following steps: In step S100, a list of external asynchronous tools is pre-configured in the server-side agent, which is used to distinguish between ordinary tools and external asynchronous tools. In a specific example, the list of external asynchronous tools is used to record the names of tools that are not suitable for synchronous execution within the server-side agent process and need to be executed asynchronously by an external system independent of the server-side agent.

[0018] In step S101, when a tool call request is generated using a large language model, if the tool to be called belongs to the list of external asynchronous tools, the server generates a globally unique tool result identifier and extracts call parameters from the tool call request. The tool result identifier is used to associate the tool call of the large language model, the execution task of the external system, and the result storage record of the server. For example, the large language model in this application can be an LLM, a self-developed model such as Llama-3.1-70B fine-tuned on an industry corpus using LoRA, or an existing large model such as GPT-4 or Qwen. No specific limitation is made. The server-side agent, upon receiving the tool call request generated using the large language model, queries the list of external asynchronous tools. If the tool name exists in the list, it is determined to be an external (asynchronous) tool and enters the asynchronous execution process; otherwise, it is executed synchronously within the process as a normal tool.

[0019] After entering the asynchronous execution process, for external tools, the server generates a globally unique tool result identifier based on the following fields: session identifier, message location identifier, tool call identifier, current timestamp and hash operation result. In the specific example, the tool result identifier is used to uniquely associate: the tool call action of the large language model, the execution task of the external system, and the record in the result storage.

[0020] Meanwhile, the server extracts call parameters from the tool message, such as query conditions, entity identifiers, and inference types.

[0021] In step S102, the tool result identifier and call parameters are written to the result storage, and an external tool dispatch event is sent to an external orchestrator via a streaming event. The streaming event is a structured data message containing a session identifier, tool result identifier, tool name, and call parameters. The external orchestrator is a message middleware or server component that receives the streaming event and schedules an external system to execute a task. The result storage also contains the corresponding session identifier and tool call status. Some embodiments, as shown in Table 1, specifically include the following records in the result storage: Table 1 Result Storage Record Table

[0022] The tool call status is used to describe the session status, and its status fields (as shown in the tool call status fields in Table 1) include at least the pending, completed, failed, and timeout statuses.

[0023] Furthermore, the server sends structured external tool dispatch events to the external orchestrator via streaming event channels such as SSE, WebSocket, or message queues. These external tool dispatch events include: session identifier, tool result identifier, tool name, tool invocation identifier, and invocation parameters. This event is only used to trigger execution by the external system and is not written to the session history.

[0024] The external system listens for streaming events, parses the call parameters, and executes the actual business logic. After execution, in step S103, the execution result is submitted through the result feedback interface provided by the server, and the tool call status and result content in the result storage are updated. The execution result is the result of the external system executing the tool logic based on the received dispatch event. For example, the result feedback interface can be an independent interface, and the feedback execution result includes: tool result identifier, execution status such as success / failure, complete execution result, or failure reason.

[0025] In some embodiments, before updating the tool invocation status and result content in the result storage, the method further includes: performing session ownership verification, tool result identifier existence verification, and verification that the current status is pending on the backfilled execution result, and allowing the update of result content and tool invocation status if the verification passes.

[0026] In step S104, in the same agent inference task, before the server reads the current task context and constructs the input of the large language model, it queries the tool call status in the result storage according to the tool result identifier. If the tool call status is completed, the complete result is converted into the tool return content used for processing by the large language model, and the original tool message at the corresponding position in the memory context copy is replaced. The conversion and replacement operation only applies to the memory context copy sent to the large language model and does not modify the persistent session message chain.

[0027] The method of this application allows the external system execution results to no longer be limited to background task results or the next user input. Instead, they can be entered into the current Agent inference chain through a mechanism that waits for the server and converts the complete results into tool return content for use in large language model processing and replaces the original tool message at the corresponding position in the memory context copy. This enables the server agent to use the tool results asynchronously generated by the external system in the same inference task.

[0028] In some embodiments, writing the result storage based on the tool result identifier and call parameters further includes writing the original tool message containing the call parameters and tool result identifier into the session message chain, without immediately replacing it with the execution result. For example, the original tool message contains the call parameters and tool result identifier, but does not contain the execution result. In a specific example, in an agent system based on a Large Language Model (LLM), the session message chain refers to a sequence of message objects organized and persistently stored in chronological order within the same session, with each message carrying a role, content, timestamp, and metadata.

[0029] In this embodiment, once the session message chain is written, its content remains unchanged regardless of the execution and backfilling of external tools. Even after the tool has finished executing and the result has been generated, the original tool message (including the calling parameters and tool result identifier) ​​remains intact.

[0030] When the tool invocation status is updated to complete, the process of converting the complete result into tool return content for use in large language model processing and replacing the memory context copy only applies to the memory context copy fed into the large language model and does not modify the session message chain; that is, the persistent session message chain remains unchanged. For example, consider a user querying information: Use the large language model output tool to call the request query_logistics(id="123"); The server determines that the tool is an external asynchronous tool and generates a tool result identifier res_abc; The original tool message is written to the session, and the tool call status is set to pending completion. The streaming event is sent to the system, which then queries the third-party interface; the query details are then submitted back to the interface. The process by which the model reads the completed state before the next inference step, converts the complete result of the query information details into tool return content that can be processed by a large language model, and replaces the memory context copy; The model then generates a natural language response. Throughout the process, the conversation message chain remains stable, and no data inconsistency issues arise due to asynchronous execution.

[0031] In some embodiments, identifying the tool call status in the query result storage based on the tool result identifier further includes: if the tool call status is pending completion, polling and waiting within a preset timeout period, and after successful waiting, performing the operation of converting the complete result into tool return content for use in large language model processing and replacing the memory context copy.

[0032] Before constructing the input context for the large language model, the server first detects external asynchronous tool messages within the scope of the current task. For example, for a specific external asynchronous tool message, the server using the method of this application executes the following processing flow: Set the maximum waiting timeout to Tmax = 10 seconds; An exponential backoff polling strategy is adopted, with an initial polling interval of 200ms. The interval doubles after each polling failure until the maximum interval of 2 seconds is reached. During each polling cycle, the server queries the status field in the results storage based on the tool ID. For example, after the third poll, the external system has completed the information query and written the result to the results storage through the backfill interface, updating the status to "completed".

[0033] The process of dynamically converting the data into content that can be processed by a large language model is described. The server then reads the complete execution result from the result store, such as JSON. "id":"202x-001"; "status":"xx"; "logistics":"xx"; ...}

[0034] Subsequently, the server dynamically transforms the above results into the returned content of the tool and replaces the content at the corresponding position in the memory context copy, forming a context fragment that is sent into the large language model.

[0035] The complete results described above are not written to the session message chain. The session message chain still only stores the original tool messages, thus ensuring the immutability of the persistent data.

[0036] If the tool call status is timed out or failed, an error placeholder is generated and the original tool message at the corresponding position in the memory context copy is replaced.

[0037] In some embodiments, polling and waiting within a preset timeout period includes: The server periodically queries the storage status of the results using either exponential backoff or a fixed interval; that is, it gradually increases the waiting time between each polling interval to reduce system load.

[0038] If the status changes to "completed" within the timeout threshold, the waiting is immediately terminated and the process proceeds to the step of converting the complete result into content returned by the tool used for processing the large language model.

[0039] If the process is not completed within the threshold, the status will be marked as timed out, and corresponding error placeholder content will be generated. Based on the previous example, if the execution is not completed within 10 seconds, and the server still cannot find a completed status after multiple polls, it will determine that the waiting process has timed out, and update the status in the result storage from pending completion to timed out.

[0040] This error placeholder is used to inform the large language model that the tool was invoked, execution failed to complete on time, the model should perform downgraded inference accordingly, or the situation should be explained to the user.

[0041] Example of generating error placeholder content when the status is "failed": The external system failed to execute. Suppose that an exception occurred during the execution of the external system, such as invalid parameters or insufficient permissions. The result is submitted through the backfill interface. After the server verifies the result, the result storage status is updated to "failed".

[0042] When the server queries this state during context construction, it directly generates a placeholder error for the failure type in the memory context.

[0043] In some embodiments, for historical external tool messages that are not within the scope of the current inference task, the method further includes rendering them as lightweight placeholder content. This lightweight placeholder content includes at least the tool name, tool result identifier, tool invocation status, and an on-demand loading prompt. When the large language model determines that the complete content of historical external tool results is needed, it can invoke the on-demand loading tool to read the complete results from the result storage via the tool result identifier.

[0044] This application's method introduces a mechanism in the pending state that involves polling and waiting, converting the complete result into tool return content for use in large language model processing, and replacing the original tool message at the corresponding position in the memory context copy. Combined with an error placeholder strategy for timeout and failure scenarios, and through preset timeout periods and state machine constraints, it prevents external system anomalies from causing the entire inference chain to freeze. Even if the tool fails to execute, the model can still understand the reason for failure through the error placeholder content, thus enabling an appropriate response.

[0045] In some embodiments, the method further includes: when the large language model determines that the complete content of the historical external tool results needs to be used, invoking the on-demand loading tool and passing the tool result identifier to the server; The server retrieves the complete execution result from the result store based on this identifier and returns it to the model. For example, when the server detects that a tool message is not within the scope of the current inference task and its status is "completed" or "failed," it constructs a memory context copy using a lazy rendering strategy. In the current inference round, the large language model autonomously determines whether to rely on historical tool results based on the user's question or inference needs.

[0046] For example, when dealing with questions in an interactive session, after analyzing the context using a large language model, it identifies relevant tool calls in the history, but the current context only contains placeholder information, making a direct answer impossible. This can be addressed by, for instance, setting the `load_tool_result` field to load tools on demand, with parameters including the aforementioned tool result identifier. Upon receiving the `load_tool_result` call, the server executes the following verification process: Session ownership verification: Verify whether the tool result identifier, such as tool_result_id, belongs to the current session to prevent unauthorized access across sessions.

[0047] Identifier validity verification: Verify whether the identifier exists in the result storage; Status validity check: The status of the corresponding record in the query results storage.

[0048] If all the above checks pass, the server reads the complete execution result from the result storage based on tool_result_id.

[0049] After obtaining the complete historical results, the large language model incorporates them into the current reasoning context and generates an answer to the user's question accordingly.

[0050] Throughout the entire on-demand loading process of this application, the original historical tool messages in the session message chain are never modified. The server only temporarily injects the complete result into the memory context of this inference, realizing on-demand expansion and effectively improving the efficiency of context management.

[0051] This application also proposes a server-side intelligent agent system, including: An external tool identification module is used to determine whether a tool call is an external tool based on a pre-configured list; an identifier generation module is used to generate a globally unique tool result identifier for an external tool call, wherein the external asynchronous tool list is used to record the names of tools executed asynchronously by external systems, independent of the server-side intelligent agent; The results storage module is used to maintain the execution status and results of tool calls; The streaming event dispatch module is used to send tool execution requests to external systems via streaming channels; The result backfilling interface module is used to receive the execution results from external systems and update the result storage. The context construction engine is used to identify the tool call status in the query result storage based on the tool result before model inference. If the tool call status is completed, the complete result is converted into tool return content for processing by the large language model and the original tool message at the corresponding position in the memory context copy is replaced. The conversion and replacement operation only applies to the memory context copy sent to the large language model and does not modify the persistent session message chain.

[0052] This application also proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned asynchronous execution and context backfilling method for server-side intelligent external tools.

[0053] It should be noted that, in the embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0054] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0055] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0056] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims. All of these forms are within the protection scope of this application.

Claims

1. A method for asynchronous execution and context backfilling of external tools for server-side intelligent agents, characterized in that, Includes the following steps: A list of external asynchronous tools is pre-configured in the server-side intelligent agent. The list of external asynchronous tools is used to record the names of tools that are executed asynchronously by external systems independently of the server-side intelligent agent. When using a large language model to generate a tool call request, if the tool to be called belongs to the list of external asynchronous tools, the server generates a globally unique tool result identifier and extracts call parameters from the tool call request. The tool result identifier is used to associate the tool call of the large language model, the execution task of the external system, and the result storage record of the server. Based on the tool result identifier and call parameters, write them into the result storage, and send external tool dispatch events to the external orchestrator through streaming events. The streaming event is a structured data message containing a session identifier, tool result identifier, tool name and call parameters. The external orchestrator is a message middleware or server component that receives the streaming event and schedules external systems to execute tasks. The result storage also contains the corresponding session identifier and tool call status. The execution result is submitted through the result feedback interface provided by the server, and the tool call status and result content in the result storage are updated. The execution result is the result of the external system executing the tool logic based on the received dispatch event. In the same agent inference task, before reading the current task context and constructing the input of the large language model, the server queries the tool call status in the result storage according to the tool result identifier. If the tool call status is completed, the complete result is converted into the tool return content used for processing by the large language model, and the original tool message at the corresponding position in the memory context copy is replaced. The conversion and replacement operation only applies to the memory context copy sent to the large language model and does not modify the persistent session message chain.

2. The asynchronous execution and context backfilling method for server-side intelligent agents as described in claim 1, characterized in that, The process of writing the result storage based on the tool result identifier and calling parameters also includes: The original tool message, containing the call parameters and tool result identifier, is written into the session message chain; the original tool message does not contain the execution result of the external system; and, When the tool invocation status is updated to complete, the process of converting the complete result into the tool return content used for processing the large language model and replacing the memory context copy only applies to the memory context copy fed into the large language model and does not modify the persistent session message chain.

3. The asynchronous execution and context backfilling method for server-side intelligent agents as described in claim 2, characterized in that, The tool call status in the query result storage, based on the tool result identifier, also includes: If the tool call status is pending, it will poll and wait for a preset timeout period. After the wait is successful, it will perform the operation of converting the complete result into the tool return content used for large language model processing and replacing the memory context copy. If the tool call status is timed out or failed, an error placeholder is generated and the original tool message at the corresponding position in the memory context copy is replaced.

4. The asynchronous execution and context backfilling method for server-side intelligent agents as described in claim 3, characterized in that, Polling wait within the preset timeout period includes: The server periodically queries the storage status of the results using either exponential backoff or a fixed interval. If the status changes to completed within the timeout threshold, the waiting is terminated and the process proceeds to the step of converting the complete result into content returned by the tool used for processing the large language model. If the process is not completed within the specified timeframe, the status will be marked as timed out, and corresponding error placeholder content will be generated.

5. The asynchronous execution and context backfilling method for server-side intelligent agents as described in claim 1, characterized in that, The records in the result storage specifically include: Tool result identifier, session identifier, tool name, tool call identifier, call parameters, tool call status, and creation time; The tool invocation status includes at least the statuses of pending completion, completed, failed, and timeout.

6. The asynchronous execution and context backfilling method for server-side intelligent agents as described in claim 1, characterized in that, The streaming events are transmitted via server-side push event SSE, WebSocket long connection, or message queue topic subscription; The streaming events are only used to trigger external systems to execute tasks and are not written to the session history.

7. The asynchronous execution and context backfilling method for server-side intelligent agents as described in claim 1, characterized in that, Before updating the session state and result content in the result store, the following steps are also included: The execution results of the backfill are verified for session ownership, existence of tool result identifier, and current status as pending. If the verification passes, the result content and tool call status can be updated.

8. The method for asynchronous execution and context backfilling of external tools for server-side intelligent agents as described in claim 1, characterized in that, It also includes rendering historical external tool messages that are not within the scope of the current inference task as lightweight placeholder content. The lightweight placeholder content includes at least the tool name, tool result identifier, tool call status, and on-demand loading prompt.

9. The asynchronous execution and context backfilling method for server-side intelligent agents as described in claim 8, characterized in that, Also includes: When the large language model determines that the complete content of the historical external tool results needs to be used, it calls the on-demand loading tool and passes the tool result identifier to the server. The server reads the complete execution result from the result storage based on this identifier and returns it to the model.

10. A server-side intelligent agent system, used to implement the asynchronous execution and context backfilling method for server-side intelligent agent external tools as described in any one of claims 1-9, characterized in that, include: The external tool identification module is used to determine whether a tool call is an external tool based on a pre-configured list; The identifier generation module is used to generate globally unique tool result identifiers for external tool calls. The list of external asynchronous tools is used to record the names of tools executed asynchronously by external systems, independent of the server-side intelligent agent. The results storage module is used to maintain the execution status and results of tool calls; The streaming event dispatch module is used to send tool execution requests to external systems via streaming channels; The result backfilling interface module is used to receive the execution results from external systems and update the result storage. The context construction engine is used to identify the tool call status in the query result storage based on the tool result before model inference. If the tool call status is completed, the complete result is converted into tool return content for processing by the large language model and the original tool message at the corresponding position in the memory context copy is replaced. The conversion and replacement operation only applies to the memory context copy sent to the large language model and does not modify the persistent session message chain.