Large language model agent tool call failure self-repair method and system thereof
Patent Information
- Application Number
- CN202610982387.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]针对现有技术中大语言模型Agent工具调用失败处理框架因将根因诊断与修复策略合并交由大语言模型单一反思过程处理而导致的修复策略错配与多步骤任务整体重启的核心瓶颈,本发明提供大语言模型Agent工具调用失败自修复方法及其系统,通过将失败根因诊断从大语言模型反思过程中剥离为独立的离散分类器并以诊断结果硬性约束修复策略路由、修复成功后从失败步骤断点续行的诊断—修复因果耦合架构,在不修改大语言模型本身且最少额外推理开销的前提下,从概率推理与因果分类的范畴分立原理层面上实现大语言模型Agent在复杂自动化任务场景中的高可靠稳定执行
[0010]本发明的有益效果包括三方面。其一,本发明通过错误类型分类模型将工具调用失败按根因离散分类为参数格式错误、权限认证失效、工具服务不可用与返回值解析失败四者中的一者,并将所述根因类别作为诊断结果硬性附加至Agent上下文作为修复策略生成的输入条件,使修复策略路由实现根因驱动的策略级跳转,在多工具Agent执行日志验证集上的根因类别识别准确率达到94.3%;其机理在于将本质属于离散分类范畴的失败诊断从大语言模型连续表征空间的反思过程中剥离出来,由独立分类器以分类损失为优化目标完成,规避了大语言模型范畴错位所致的策略错配,相较所述现有技术的通用三分类异常处理实现了在参数格式错误场景下避免盲目重试、在权限认证失效场景下避免参数微调浪费的本质改进。其二,本发明的差异化修复模块依据所述诊断结果从修复策略路由表中选取对应的修复策略,参数格式错误触发工具文档重读与参数重构、权限认证失效触发凭证刷新工具调用、工具服务不可用触发备用工具替换搜索、返回值解析失败触发输出格式重新推断,各修复策略以最少额外推理开销为约束设计,相较现有技术中将错误信息回传大语言模型由其自由反思的方案,单次修复推理token消耗降低约28%;其机理在于诊断结果对修复策略生成施加了离散硬约束,将原先的连续推理空间收窄为四类预定义策略路径,避免了大语言模型在所有可能修复路径上的开放搜索。其三,本发明的断点续行机制在所述修复策略执行成功后从所述失败步骤序号处对所述多步骤任务续行并保留已完成步骤的中间执行结果,使多步骤Agent任务在引入自修复机制后的端到端成功完成率从基线的63.4%提升至91.2%;其机理在于将失败步骤之前已完成步骤的算力固化为可复用的中间执行结果而非废弃,从而打破了多步骤任务因任意单步失败累积概率而整体崩塌的链路困局。所述根因诊断、所述差异化修复与所述断点续行三者通过Agent上下文形成因果耦合闭环,前者为后者提供硬性输入条件,三者协同效应远高于单一机制独立部署的简单线性叠加,构成本发明在复杂自动化任务场景下大语言模型Agent高可靠稳定执行的整体性技术贡献。
Smart Images

Figure CN122838147A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a self-repair method and system for large language model agent tool call failures. Background Technology
[0002] With the continuous breakthroughs in the natural language understanding and generation capabilities of large language models, the Agent architecture, which uses large language models as the decision-making core and completes automated tasks by calling external tools, has become the mainstream paradigm for the application of artificial intelligence. These external tools typically include RESTful API interfaces, SQL database query interfaces, Python code execution sandboxes, and vector retrieval interfaces. After receiving a user's multi-step task request, the Agent decomposes it into sub-tasks using the large language model and generates tool call requests. These requests are then sent to the external tools for execution, and the next action is determined based on the execution results, thus completing a complex automated process that originally required manual intervention. However, when the external tools return failure signals due to parameter format mismatches, expired credentials, service rate limiting, or abnormal data structure, the multi-step tasks often suffer from reliability issues such as link interruption, wasted inference resources, and wasted computing power of completed steps, becoming a key bottleneck for the large-scale deployment of Agent technology.
[0003] Chinese patent application CN118963869A discloses a method and apparatus for executing large-scale model tasks based on knowledge graphs. This method constructs a target knowledge graph using function information from API information as nodes and function call dependencies as edges. A function registry is then built based on the target knowledge graph, and an execution agent is established based on the target knowledge graph and the function registry. After the large model generates function calls, the execution agent verifies whether the function name and parameters match the function registry. If they match, a call chain is generated and executed; otherwise, an error message is fed back to the large model. The execution agent further incorporates an exception handling and retry mechanism, determining whether to retry, skip, or terminate the call chain based on the exception type. However, the exception handling in this scheme only uses a general three-category approach—retry, skip, and abort—to determine subsequent actions, without differentiating based on the root cause of tool call failure. In this scheme, parameter format errors and authentication failures may both be categorized into retry branches and subjected to the same retry strategy, leading to blind retries in parameter format error scenarios that repeatedly consume inference resources. At the same time, after an exception occurs, the scheme uses fallback handling such as aborting the call chain or adjusting the call order based on the context, without disclosing an execution stack merging mechanism that resumes the multi-step task from the failed step and retains the intermediate execution results of completed steps. When a multi-step task fails in a single step, the computing power of completed steps is also discarded, limiting the overall reliability improvement.
[0004] Chinese patent application CN120633639A discloses an intelligent agent system and interaction method based on a large language model. The system includes a data acquisition module, a natural language processing module, a knowledge reasoning module, an output module, and an adaptation module, enhancing the user interaction experience through sentiment analysis and personalized response technologies. The patent application explicitly points out in its background section that in most current intelligent agent systems, insufficient accuracy in parameter generation during the tool invocation process frequently leads to tool execution failures, and the system lacks effective anomaly recovery capabilities, often resulting in the interruption of the entire task flow once an error occurs. This background statement confirms the real-world pain point regarding the robustness of tool invocation by large language model agents. However, the proposed solution does not provide root cause diagnosis and differentiated repair mechanisms for tool invocation failures at the claim level. Its anomaly handling still relies on manual transfer and multimodal fallback, failing to fundamentally solve the bottleneck problems of mismatched repair strategies and overall task restart after tool invocation failures.
[0005] Further analysis of existing large language model agent tool call failure handling paths reveals that the commonly adopted technical approach in this field is to send the error message text returned by the tool call back to the large language model, which then reflects on and repairs based on the error message in the next round of inference. The fundamental limitation of this approach is that the next-word inference of a large language model is essentially a nearest-neighbor search within a continuous representation space. Reflection on error messages tends to be parameter fine-tuning repair, failing to make policy-level transitions between different failure root causes. The same error message may originate from four completely different root causes in different contexts: incorrect parameter format, invalid authentication, unavailable tool service, or failed return value parsing. The repair strategies required for these four root causes are fundamentally different: incorrect parameter format requires rereading the tool documentation and reconstructing parameters; invalid authentication requires credential refresh tool calls; unavailable tool service requires replacement with a backup tool; and failed return value parsing requires re-inferring the output format. Combining root cause diagnosis and repair strategies into a single reflection process of the large language model can lead to mismatched repair strategies due to the misalignment of the large language model's scope. When a multi-step agent task relies on multiple tool calls and the failure probability accumulates along the steps, the overall task restart triggered by a single-step failure discards all intermediate execution results of completed steps, severely limiting the end-to-end success rate.
[0006] Therefore, the existing framework for handling failures in large language model agent tools lacks a discrete classification and policy-level jump mechanism for failure root causes independent of the large language model's reflection process, and it also lacks an execution stack merging mechanism to resume execution from the breakpoint of the failed step after successful repair. As a result, it cannot meet the high reliability and stability requirements of large language model agents in complex automated task scenarios, and a new technical solution is urgently needed to fundamentally break through the above bottlenecks. Summary of the Invention
[0007] To address the core bottlenecks of existing large language model agent tool call failure handling frameworks, which suffer from mismatched repair strategies and overall task restarts due to the merging of root cause diagnosis and repair strategies into a single reflection process of the large language model, this invention provides a self-repair method and system for large language model agent tool call failures. By separating the root cause diagnosis of failure from the large language model reflection process into an independent discrete classifier and using the diagnosis results to rigidly constrain the routing of repair strategies, and resuming from the breakpoint of the failed step after successful repair, this diagnosis-repair causal coupling architecture achieves highly reliable and stable execution of large language model agents in complex automated task scenarios without modifying the large language model itself and with minimal additional inference overhead, based on the principle of separation of probabilistic reasoning and causal classification.
[0008] The technical solution of this invention is: a self-repair method for tool call failures using a large language model agent, comprising the following steps: receiving a tool call request sent by the agent to an external tool during the execution of a multi-step task, and monitoring the execution return result of the tool call request; recording a tool call failure event when the execution return result hits a failure signal, wherein the tool call failure event includes a failure error message text and a failure step number; inputting the failure error message text into an error type classification model, wherein the error type classification model outputs the root cause category of the tool call failure, wherein the root cause category is parameter format error, authentication failure, tool service unavailable, and return value parsing failure. One of the four: The root cause category is forcibly appended to the Agent context as a diagnostic result, and the diagnostic result serves as the input condition for generating the repair strategy; Based on the diagnostic result, the corresponding repair strategy is selected from the repair strategy routing table and executed. The repair strategy routing table includes: tool document rereading and parameter reconstruction for parameter format errors, credential refresh tool invocation for invalid permission authentication, backup tool replacement search for unavailable tool services, and output format re-inference for failed return value parsing; After the repair strategy is executed successfully, the multi-step task is resumed from the failure step number, retaining the intermediate execution results of steps completed before the failure step number.
[0009] This invention also provides a self-repairing system for failed tool calls by a large language model agent, comprising: a tool call monitoring module, used to receive tool call requests sent by the agent to an external tool during the execution of a multi-step task, and to monitor the execution return result of the tool call request, recording a tool call failure event when the execution return result hits a failure signal, the tool call failure event including a failure error message text and a failure step number; a failure diagnosis module, used to input the failure error message text into an error type classification model, the error type classification model outputting the root cause category of the tool call failure, the root cause category being one of four: parameter format error, authentication failure, tool service unavailable, and return value parsing failure; and a top-bottom ... The document attachment module is used to forcibly attach the root cause category as a diagnostic result to the Agent context, and the diagnostic result serves as the input condition for generating the repair strategy. The differential repair module is used to select and execute the corresponding repair strategy from the repair strategy routing table based on the diagnostic result. The repair strategy routing table includes: rereading and reconstructing tool documents for parameter format errors, calling credential refresh tools for invalid permission authentication, searching for alternative tools for unavailable tool services, and re-inferring the output format for failed return value parsing. The breakpoint continuation execution module is used to resume the multi-step task from the breakpoint at the failed step number after the repair strategy is successfully executed, retaining the intermediate execution results of the steps completed before the failed step number.
[0010] The beneficial effects of this invention include three aspects. First, this invention uses an error type classification model to discretly classify tool call failures into four root causes: parameter format error, authentication failure, tool service unavailability, and return value parsing failure. The root cause category is then forcibly appended to the Agent context as a diagnostic result and used as input conditions for generating repair strategies. This enables root cause-driven strategy-level jumps in repair strategy routing, achieving a root cause category identification accuracy of 94.3% on a multi-tool Agent execution log verification set. The mechanism lies in separating the failure diagnosis, which essentially belongs to the discrete classification category, from the reflection process of the continuous representation space of the large language model. An independent classifier completes this process with classification loss as the optimization objective, avoiding strategy mismatch caused by misalignment in the large language model's categories. Compared to the general three-class anomaly handling of existing technologies, this invention achieves a fundamental improvement by avoiding blind retries in parameter format error scenarios and avoiding wasted parameter fine-tuning in authentication failure scenarios. Secondly, the differentiated repair module of this invention selects the corresponding repair strategy from the repair strategy routing table based on the diagnostic results. Parameter format errors trigger tool document rereading and parameter reconstruction, invalid permission authentication triggers credential refresh tool invocation, unavailable tool services trigger backup tool replacement search, and failed return value parsing triggers output format re-inference. Each repair strategy is designed with minimal additional inference overhead as a constraint. Compared with the existing technology that sends error information back to the large language model for free reflection, the token consumption for a single repair inference is reduced by approximately 28%. The mechanism is that the diagnostic results impose discrete hard constraints on the generation of repair strategies, narrowing the original continuous inference space to four types of predefined strategy paths, avoiding open search of the large language model on all possible repair paths. Third, the breakpoint continuation mechanism of this invention, after the repair strategy is successfully executed, resumes the multi-step task from the failed step number and retains the intermediate execution results of the completed steps. This increases the end-to-end success rate of the multi-step agent task from the baseline of 63.4% to 91.2% after the introduction of the self-repair mechanism. Its mechanism is to solidify the computing power of the completed steps before the failed step into reusable intermediate execution results instead of discarding them, thereby breaking the link dilemma of the multi-step task collapsing as a whole due to the cumulative probability of any single-step failure. The root cause diagnosis, the differentiated repair, and the breakpoint continuation form a causal coupling closed loop through the agent context. The former provides hard input conditions for the latter, and the synergistic effect of the three is far greater than the simple linear superposition of a single mechanism deployed independently. This constitutes the overall technical contribution of this invention to the highly reliable and stable execution of large language model agents in complex automated task scenarios. Attached Figure Description
[0011] Figure 1 This is a flowchart of the self-repair method for failed calls to the large language model Agent tool described in this invention;
[0012] Figure 2 This is an architecture diagram of the self-repair system for failed calls to the large language model Agent tool described in this invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0014] refer to Figure 1 The self-repair method for failed calls to the large language model Agent tool described in this invention includes five core steps, from S1 to S5. Each step is explained in detail below.
[0015] Step S1: Tool Invocation Listening and Failure Event Logging. During the Agent's execution of multi-step tasks, the large language model acts as the decision-making core, splitting user requests into multiple sub-tasks according to a preset task decomposition mechanism, and generating corresponding tool invocation requests in each sub-task. The tool invocation requests are described using a structured data format, and each request includes at least three fields: tool identifier, request parameters, and request time. The external tools are at least one of the following: RESTful API interface, SQL database query interface, Python code execution sandbox, and vector retrieval interface. This embodiment uses a hybrid tool pool that simultaneously mounts all four types of external tools as an example.
[0016] The tool call monitoring is implemented by inserting a transparent interception layer between the Agent's main inference loop and the actual entry point of the external tool. This transparent interception layer receives the tool call request generated by the large language model, forwards it to the corresponding external tool for execution, and synchronously captures the execution result when the external tool returns it. The transparent interception layer does not modify the content of the tool call request; it only records the call sequence and execution result, remaining completely unaware of the large language model's inference process.
[0017] The execution return result includes two parts: a success flag and return data. The success flag hit failure signal determination logic is as follows: the HTTP status code of the execution return result is greater than or equal to 400; or the JSON payload of the execution return result contains any of the preset error key fields, namely the error field, the err_code field, or the exception field; or the external tool call process throws a runtime exception, and the exception class of the runtime exception inherits from the Exception base class; or the waiting time for receiving the execution return result exceeds the preset tool call timeout threshold, which is set to 30s in this embodiment.
[0018] When the execution return result hits the failure signal, the tool call listener records the tool call failure event. The tool call failure event is stored as a structured data object in the failure event queue of the Agent context. The tool call failure event includes two core fields: the failure error message text and the failure step number. The failure error message text is the original error description string returned by the external tool, including the values of the key error fields, the exception class name and exception message of the runtime exception, and any diagnostic hints that may be carried in the execution return result. The failure step number is the step count value of the multi-step task that has been executed when the tool call failure event occurs, maintained by the Agent's step counter, and increments continuously from the first step of the multi-step task. The failure event queue adopts a first-in-first-out circular buffer structure, with a maximum capacity of 100 entries in this embodiment. Old failure events are automatically dequeued when new failure events are enqueued, avoiding memory bloat under long-running tasks.
[0019] After the tool call failure event is enqueued, the tool call listener transfers control to the failure diagnosis module and proceeds to step S2. At the same time, it freezes the current execution stack of the multi-step task and persists the intermediate execution results of all completed steps from step 1 to the step number minus one step in the failure step to the Agent's execution stack snapshot storage area after serialization with Pickle protocol version 4, ensuring that the intermediate execution results can be restored as is when resuming from a breakpoint.
[0020] Step S2: Input the failure error message text into the error type classification model and output the root cause category. The failure error message text is extracted from the tool call failure event in Step S1, and together with the tool call context summary, constitutes the input of the error type classification model in the pre-processing stage. The tool call context summary includes three fields: the tool identifier of the called tool, the request parameters of the tool call request, and the failure step number. These three fields are concatenated into a single string using a preset separator, and then combined with the failure error message text to form a binary input to the error type classification model.
[0021] The error type classification model is a multi-classification model obtained by fine-tuning a pre-trained encoder. In this embodiment, the pre-trained encoder is a bidirectional encoder based on a Transformer architecture, with 12 encoder layers and a hidden layer dimension of 768. A classification head is appended to the output of the pre-trained encoder. This classification head consists of a fully connected layer and softmax activation, with an output dimension of 4, corresponding to the four members of the root cause category. After fine-tuning on a multi-tool agent execution log validation set, the error type classification model achieves a root cause category identification accuracy of no less than 94.3%.
[0022] The training loss function of the error type classification model is cross-entropy loss, and its calculation formula is as follows: ,in: The training loss function value of the error type classification model is a scalar with a value range of [0, +∞) and a dimensionless unit. It is calculated by this formula for each training batch and is used to characterize the degree of deviation between the root cause category prediction distribution and the true distribution of the error type classification model in that training batch. The smaller the loss value, the better the model classification performance. The total number of training samples in the training batch is a positive integer scalar with a value range of [1, +∞). The unit is dimensionless and is determined by the preset training hyperparameters. In this embodiment, the value is 32. The index of the training sample is a positive integer scalar with a value range of [1, ..., ...]. The unit is dimensionless and is used to accumulate all training samples in summation operations; The total number of the root cause categories is a positive integer scalar with a fixed value of 4 and a dimensionless unit, corresponding to the four members of the root cause categories. The index of the root cause category is a positive integer scalar with a value range of [1, ..., ...]. The unit is dimensionless; For the first The training sample belongs to the first... The true labels of each root cause category are scalars, taking values of 0 or 1, and are dimensionless. They are obtained from manual or semi-automatic annotation of the training samples in the multi-tool Agent execution log verification set. The error type classification model for the first The training sample at the th ... The output logit value for each root cause category is a scalar with a value range of (-∞, +∞) and a dimensionless unit. It is obtained from the original output of the fully connected layer of the classification head before softmax activation. The index of the root cause category in softmax normalization is a positive integer scalar with a value range of [1, ..., ... The unit is dimensionless, and... Functionally equivalent but used for decoupling inner and outer summations. Left side of the equation. Since it is dimensionless, the right side of the equation Dimensionless Dimensionless When applied to a dimensionless exponent ratio, it remains dimensionless, with both sides being dimensionless quantities and having the same dimensions.
[0023] The error type classification model calculates the logit vector from the input tuple of the failure error message text and the tool call context summary during the inference phase, and then takes the root cause category corresponding to the largest logit value as the final output, that is:
[0024] ,in: The root cause category index output by the error type classification model for the current tool call failure event is a positive integer scalar with a value of 1, 2, 3 or 4, and the unit is dimensionless. It is calculated by this formula during the inference stage and corresponds to one of the four root cause categories: parameter format error, authorization authentication failure, tool service unavailable, and return value parsing failure. and The definition is the same as the previous formula, where For the error type classification model in the inference phase, the current sample is in the first... The output logit values for each root cause category are dimensionless. (Left side of the equation) For dimensionless integers, the right side of the equation The operator, when applied to a dimensionless logit value, returns an index that is a dimensionless integer with consistent dimensions.
[0025] To address the issue of input dimension drift in the error type classification model caused by the unbounded step number space of the multi-step task, the tool invokes the failed step numbers from the context summary, which undergo positional encoding normalization before being input into the error type classification model. This positional encoding normalization process is performed in two stages. The first stage is step number normalization mapping:
[0026] ,in: The normalized step number is a scalar with a value range of (0,1] and a dimensionless unit. It is calculated by this formula and represents the relative position of the failed step number in the total length scale of the multi-step task. The original integer value of the failure step number is a positive integer scalar, with a value range of [1, ..., ...]. The unit is dimensionless and is obtained by the step counter when the tool call failure event occurs; The maximum number of steps for the multi-step task is a positive integer scalar, ranging from [1, +∞), with a dimensionless unit. It is estimated by the Agent task planner during task initialization based on task complexity; in this embodiment, it is set to 50 by default. (Left side of the equation) The numerator on the right side of the equation is dimensionless. With denominator Both are dimensionless, and the result of dividing them is dimensionless, but they have the same dimension.
[0027] The second stage is sinusoidal position encoding:
[0028] ,in: For the position encoding vector in the th The values are taken in an even-numbered dimension, are scalars, and range from [-1, 1]. The unit is dimensionless and is calculated from the sine term of this formula. For the position encoding vector in the th The values are taken in an odd number of dimensions, scalars, with a range of [-1, 1], and are dimensionless. They are calculated from the cosine term of this formula. For the position-encoded dimension, the index is a non-negative integer scalar with a value range of [0, ...]. -1], the unit is dimensionless, used to control the frequency attenuation of different dimensions; The preset dimension of the position encoding vector is a positive even scalar with a value range of [2, +∞) and a dimensionless unit. It is determined by the preset input embedding dimension of the error type classification model, and in this embodiment, the value is 64. The definition is the same as the formula mentioned above. Both sides of the equation are dimensionless and have the same dimensions.
[0029] The positional encoding vector and the text embedding vector output by the pre-trained encoder are added bitwise to each other and then input into the classification head to obtain the final root cause category output. This positional encoding normalization process enables the error type classification model to maintain a root cause category identification accuracy of no less than 94.3% when deployed to multi-step agent tasks with up to 50 steps after fine-tuning on samples with 5 to 30 training steps.
[0030] To further address the issue of out-of-generalization failure in the error type classification model caused by the heterogeneity of error message text formats returned by different external tools, this embodiment performs error message heterogeneity normalization processing before inputting the failed error message text into the pre-trained encoder. This error message heterogeneity normalization processing maps the failed error message text to a unified seven-tuple of error messages using a set of error pattern regular expression extractors. The set of error pattern regular expression extractors includes three sub-extractors: an API error regular expression extractor, a database error regular expression extractor, and a code execution error regular expression extractor. The seven-tuple of error messages includes seven components: error code, error phrase, error level, suspicious parameters, number of prompts, HTTP status, and exception class name. The mapping function for the error message heterogeneity normalization processing is:
[0031] ,in: The error information seven-tuple is a 7-dimensional vector. The value range of each component varies depending on the component type. The unit is dimensionless. It is obtained by mapping the failure error information text through the error pattern regularity extractor set using this formula and serves as the unified format input for the error type classification model. The error code component is a string type, filled with the error code field extracted from the failure error message text by the error pattern regular expression extractor, such as E_PARAM_INVALID or 401; The error phrase component is a string with a maximum length of 50 characters, obtained by extracting a brief description of the error. For error level components, positive integer scalars, with values ranging from [1,5], are obtained by mapping error codes to a preset level table, representing five levels: client error, server error, network error, protocol error, and unknown error. The suspicious parameter component is a string type, obtained by regular expression matching of parameter-related fields from the error message, and represents the name of the request parameter that may be involved in the failure of the tool call; The number of indicators is a non-negative integer scalar with a value range of [0, 10], obtained by counting the number of diagnostic indicators contained in the error messages; This is an HTTP status component, an integer scalar with a value range of [100, 599] or 0, extracted from the HTTP response header of the execution result. 0 indicates a tool call that is not an HTTP channel. The exception class name component is a string type, extracted from the type attribute of the exception object of the runtime exception; The mapping function for the set of error pattern regular expression extractors consists of three sub-extractors connected in parallel: API error regular expression extractor, database error regular expression extractor, and code execution error regular expression extractor. The corresponding sub-extractor is activated according to the source tool type of the failure error message text. The failure error message text is extracted from the tool call failure event in step S1; , , These are API error regular expression extractors, database error regular expression extractors, and code execution error regular expression extractors, designed respectively for error formats of RESTful API interfaces and vector retrieval interfaces, SQLSTATE format of SQL database query interfaces, and Traceback format of Python code execution sandboxes; This is the index variable of the regular expression extractor, and its value is the union of the three sub-extractors of the error pattern regular expression extractor set. This represents the union operation of the set of outputs of all applicable sub-extractors. The left side of the equation... All components are dimensionless, and the right side of the equation... The output is a dimensionless septum with the same dimensions on both sides.
[0032] The error information seven-tuple is then mapped to a vector representation of uniform dimension through a preset embedding layer. This vector is then concatenated bit-by-bit with the positional encoding vector and the original embedding vector of the failure error information text, and input into the pre-trained encoder of the error type classification model. This error information heterogeneity normalization processing ensures that the error type classification model does not need to be retrained when a new tool is launched. Only the corresponding regularization extractor needs to be configured for the new tool to maintain a root cause category identification accuracy of over 94.3%.
[0033] Step S3: The diagnostic result is forcibly appended to the Agent context as input for generating the repair strategy. The root cause category output in Step S2, i.e., the diagnostic result, is forcibly appended to the Agent context. This forcible appending is achieved by inserting a pre-formatted diagnostic result declaration statement at the system-level prompt location in the Agent context. The diagnostic result declaration statement uses a structured template, the template content of which is the diagnostic result of the tool call failure event: the root cause category is {root cause category Chinese name}, the corresponding repair strategy is {repair strategy Chinese name}, and the next round of tool calls must be strictly followed; deviation from this strategy path is prohibited. The root cause category Chinese name is mapped according to the index value output by the error type classification model: 1 corresponds to parameter format error, 2 corresponds to authentication failure, 3 corresponds to tool service unavailable, and 4 corresponds to return value parsing failure. The repair strategy Chinese name is mapped according to the repair strategy routing table: parameter format error is mapped to tool document rereading and parameter reconstruction, authentication failure is mapped to credential refresh tool call, tool service unavailable is mapped to backup tool replacement search, and return value parsing failure is mapped to output format re-inference.
[0034] The diagnostic result declaration statement is inserted into the Agent context at a fixed anchor position between the system-level prompt and the user's current request. This anchor position is reserved as a diagnostic anchor point in the Agent's dialogue template. The diagnostic anchor point uses the same role-labeled weight as the system-level prompt, ensuring that the large language model cannot ignore the instruction at this position when generating the next round of tool calls. This achieves a hard input condition constraint on the generation of the repair strategy by the diagnostic result. The core mechanism of this hard constraint is that the diagnostic result is injected into the Agent context with system-level weights rather than ordinary dialogue history weights. This causes a significant shift in the conditional probability axis of the root cause category in the generation distribution of the large language model, preventing the large language model from deviating from the correct repair strategy under the free reflection path.
[0035] After the diagnostic results are hard-attached, the context attachment module further selectively simplifies the dialogue history before failure in the Agent context, retaining the tool call and execution return results of the most recent 5 rounds before the failure step number. The simplified Agent context is used as the input for step S4.
[0036] Step S4: Select and execute the corresponding repair strategy from the repair strategy routing table based on the diagnostic results. After receiving the diagnostic results from the Agent context, the differentiated repair module immediately queries the repair strategy routing table, matches the corresponding repair strategy based on the root cause category, and executes it. The repair strategy routing table is a predefined mapping structure stored in the Agent's internal policy register and loaded from the configuration file at startup. The four mapping relationships of the repair strategy routing table are detailed below.
[0037] Parameter format errors are mapped to tool documentation rereading and parameter reconstruction. The differential repair module uses the suspicious parameters extracted from the failure error message text as search keywords to retrieve relevant parameter definition paragraphs from the official documentation corpus of the external tool. The parameter definition paragraphs are then injected into the next round of prompts in the large language model to trigger the reconstruction of the request parameters. The reconstructed tool call request is then resent to the external tool for execution.
[0038] The authentication failure is mapped to a credential refresh tool call. The differentiated repair module identifies authentication failure signals in the failure error message text, such as key phrases like Unauthorized, 401, and tokenexpired. Upon a match, it calls a pre-registered credential refresh tool to request a new access token from the identity authentication service based on the OAuth 2.0 protocol or a custom credential refresh protocol. After updating the Agent's credential register, it attaches the new token to the original tool call request and resends it to the external tool for execution.
[0039] If a tool service is unavailable, a replacement tool is mapped to be searched. The differentiated repair module matches and retrieves functionally equivalent backup tools from the Agent's tool pool according to predefined capability tags. It then selects the first backup tool in descending order of its historical availability. The request parameters of the original tool call request are remapped according to the backup tool's parameter schema and sent to the backup tool for execution. The historical availability rate is the ratio of the number of successful calls in the tool's most recent 100 calls to the total number of calls.
[0040] If the return value parsing fails, it is mapped to the output format for re-inference. The differential repair module concatenates the original data sample actually returned by the external tool with the original expected parsing format description and inputs it into the large language model. This triggers the large language model to infer the schema of the actual returned data. The inferred new schema is used to update the configuration of the tool return value parser inside the Agent to re-parse the returned data. This process does not re-call the external tool and completes the repair only at the client parsing layer.
[0041] To prevent the repair strategy from consuming excessive inference resources and causing cascading crashes in repeated failure scenarios, this embodiment sets a repair inference budget cap for the repair strategy corresponding to each root cause category, and dynamically constrains it using historical success rate weights. The dynamic constraint formula for the repair inference budget cap is:
[0042] ,in: Root cause category At any moment The cumulative number of inference tokens consumed is a non-negative integer scalar with a value range of [0, +∞). The unit is the number of tokens, which is accumulated in real time by the Agent's token meter after each inference call of the large language model. Define the corresponding symbol in step S2 for the index subscript of the root cause category, with a value range of [1,4] and a dimensionless unit; The timestamp is a non-negative real scalar with a value range of [0,+∞) and a unit of seconds (s), which is read from the Agent system clock. Root cause category The corresponding base value for the repair inference budget upper limit is a positive integer scalar, ranging from [100, 10000], in units of tokens, and is determined by the Agent configuration parameters. In this embodiment, the parameter format is incorrect. A value of 2000 indicates that authentication has failed. Value 500 indicates the tool service is unavailable. The value is 3000 and the return value failed to be parsed. The value is 1500; Root cause category The historical success rate weight is a scalar with a value range of (0,1] and a dimensionless unit. It is dynamically calculated by the subsequent historical success rate weight sub-formula of this formula. This is a scalar multiplication operator. The left side of the equation represents the number of tokens, and the right side represents... For the number of tokens, The product is dimensionless, and the result of multiplication is the number of tokens. The dimensions of the left and right sides are the same.
[0043] The specific formula for calculating the historical success rate weight is as follows:
[0044] ,in: The definition is the same as the formula mentioned above; Root cause category The historical number of successful repairs is a non-negative integer scalar with a value range of [0, +∞), and the unit is times. It is obtained by accumulating the Agent's repair history log after each repair strategy is executed, and represents the cumulative number of successful repairs under the root cause category. Root cause category The historical repair failure count is a non-negative integer scalar with a value range of [0, +∞), and the unit is times. It is obtained by accumulating the repair history log of the Agent after each repair strategy execution failure, and represents the cumulative number of failures of the repair strategy under this root cause category. To prevent the exclusion of zero decimals, a positive real scalar is used, with a value range of (0, 0.01], and the unit is dimensionless. In this embodiment, the value is fixed at 0.001, representing the prevention of... and This is also a numerical stability protection term where the denominator is zero when the condition is met. (Left side of the equation) Since the terms are dimensionless, the numerator on the right side of the equation is a dimensionless integer (the fraction representing the order of magnitude), and the denominator is a dimensionless integer (the fraction representing the order of magnitude) added together. After implicit dimensionless transformation, it is still a dimensionless integer. The division of the numerator and denominator is dimensionless, and the dimensions on both sides are consistent.
[0045] when touch If the repair fails even after reaching the upper limit, the differentiated repair module stops repair attempts for the current root cause category, marks the tool call failure event as a repair budget exhaustion, and reports it to the Agent control layer. The Agent control layer then decides whether to switch to manual intervention or terminate the entire multi-step task. This dynamic constraint mechanism allows root cause categories with high historical success rates to receive a more lenient budget limit to support sufficient repair attempts, while root cause categories with low historical success rates are quickly triggered with a tighter budget limit for fallback processing. This avoids ineffective, repeated repairs that consume overall inference resources, reducing overall inference resource consumption by approximately 35% in long-term tasks compared to solutions without budget constraints.
[0046] The criterion for determining the success of the repair strategy is that the resent tool call request or the re-parsed return data returns a success flag and no longer hits the failure signal. After the repair strategy is successfully executed, the differentiated repair module updates the repair history log, records the repair result in the n_{succ}(c) statistical count, and then transfers control to the breakpoint continuation execution module to proceed to step S5.
[0047] Step S5: Resume the multi-step task from the breakpoint at the failed step number. The core mechanism of this breakpoint resumption is execution stack merging driven by the execution stack state flags. The execution stack state flags reflect the state change status of the step corresponding to the failed step number. The execution stack state flags are one of three: a partially written flag, a rolled-back flag, and a committed flag. The partially written flag indicates that the step corresponding to the failed step number has made partial state changes to external resources but has not yet completed the final commit. The rolled-back flag indicates that the step corresponding to the failed step number has undone all state changes through the transaction rollback mechanism. The committed flag indicates that the step corresponding to the failed step number has completed all state changes and has been confirmed by the final commit.
[0048] The execution stack status flag is maintained in real time by the Agent's execution stack status tracker during the execution of the step corresponding to the failed step number. The execution stack status tracker inserts status hooks before each tool call request is issued and after the execution return result is received. The status hooks update the execution stack status flags according to the transaction attributes of the tool call request and the content of the execution return result.
[0049] To achieve execution stack merging for breakpoint continuation, this embodiment introduces a continuation condition function. The continuation condition function determines the continuation path based on the value of the execution stack state flag, and its calculation formula is as follows:
[0050] ,in: The output of the continuation condition function is an integer scalar with a value of 0 or 1 and a dimensionless unit. It is calculated by this formula based on the value of the execution stack state flag. 1 indicates that direct breakpoint continuation is allowed, and 0 indicates that a lightweight rollback needs to be triggered before continuation. It represents the path selection signal of the breakpoint continuation execution module. The value of the execution stack state flag is an enumeration type, and its value range is { , , The unit is dimensionless and is maintained in real time by the execution stack state tracker during the execution of the step corresponding to the failed step number, representing the state change situation of the step corresponding to the failed step number. The enumerated value of the committed marker is a dimensionless symbolic constant with a fixed value of the string "committed". The enumerated value of the rolled-back mark is a dimensionless symbolic constant with a fixed value of the string "rolled_back". The enumerated value for writing the marker to the aforementioned part is a dimensionless symbolic constant, with a fixed value of the string "partial". (Left side of the equation) The values 1 and 0 on the right side of the equation are dimensionless integers, and their dimensions are consistent.
[0051] when When the failure step number is reached, the breakpoint continuation execution module directly resumes the multi-step task from the failure step number, loads the intermediate execution result persisted by the execution stack snapshot storage area in step S1, restores the intermediate execution result to the Agent's execution stack, and the large language model generates the next action based on the restored execution stack and the repaired tool call result generated by the repair strategy; the step corresponding to the failure step number is marked as repaired and skipped, and the remaining steps of the multi-step task continue to be executed from the failure step number plus one step.
[0052] when When the execution stack state is marked as partially written, the breakpoint resume execution module first triggers a lightweight rollback and then resumes execution. The lightweight rollback process involves identifying each partial state change corresponding to the failed step number and calling the corresponding reverse operation to restore the state to its pre-change state. After the reverse operation is completed, the execution stack state is updated to the rolled-back mark. The breakpoint resume execution module then re-enters the resume condition function with the updated rolled-back mark. Enter the direct breakpoint continuation path. The lightweight rollback consumes only about 5% to 10% of the computing power of the traditional full task restart. This embodiment achieves efficient self-repair in state pollution scenarios through this mechanism.
[0053] The breakpoint continuation module continuously monitors the remaining steps of the multi-step task after continuation. If the tool call fails again in the remaining steps, the entire process from steps S1 to S5 is repeated. If all remaining steps are successfully executed, the multi-step task ends in a self-repairing completion state and returns the final task result. The actual test results on the publicly available multi-step agent task benchmark set show that after introducing the self-repairing method described in this invention, the end-to-end success rate of the multi-step agent task increases from 63.4% before the introduction to 91.2% after the introduction, and the average additional inference token consumption for a single failure repair is reduced by approximately 28% compared to the baseline solution.
[0054] The self-repair method for failed calls to the large language model Agent tool described in this embodiment has been verified in the following three typical automated task scenarios.
[0055] The first type of scenario is the automation of financial transactions. The multi-step task includes continuous steps such as account inquiry, asset valuation, transaction strategy generation, order submission and transaction confirmation. The failure of permission authentication and the unavailability of tool services are the root causes of high frequency failures. By introducing the self-repair mechanism described in this invention, the end-to-end success rate of the backend is increased from 58.7% to 93.1%.
[0056] The second type of scenario is the automated task of medical record query. The multi-step task includes continuous steps such as patient identity verification, electronic medical record retrieval, image data loading, test result summary and diagnosis assistance generation. Parameter format errors and return value parsing failures are high-frequency failure root causes. By introducing the self-repair mechanism described in this invention, the end-to-end success rate is increased from 65.2% to 89.8%.
[0057] The third scenario is the automation of enterprise data dashboard tasks. The multi-step tasks include continuous steps such as data source connection, SQL query generation, query execution, result aggregation, visualization chart generation and report export. The differentiated repair module improves the end-to-end success rate from 66.3% to 90.7% through full coverage routing of four types of repair strategies.
[0058] To further verify the beneficial effects of the present invention, a baseline scheme is set as a comparative example in this embodiment. The baseline scheme is a prior art approach that directly sends the failed error message text back to the large language model for free reflection, without introducing the error type classification model and the repair strategy routing table, and without setting the breakpoint continuation mechanism. After a tool call fails, the multi-step task is restarted as a fallback. Comparative tests were conducted on three application scenarios identical to this embodiment and on 1000 randomly generated multi-step Agent tasks. The baseline scheme achieved an end-to-end success rate of 63.4%, an average inference token consumption of 3850 tokens per repair, and an average task completion time of 42.7 seconds. The scheme in this embodiment achieved an end-to-end success rate of 91.2%, an average inference token consumption of 2772 tokens per repair, and an average task completion time of 28.4 seconds. The comparison of these three core indicators verifies the overall technical contribution of the diagnostic-repair causal coupling architecture formed by the error type classification model, the differentiated repair module, and the breakpoint continuation mechanism.
[0059] refer to Figure 2The self-repair system for failed tool calls of the large language model Agent described in this invention comprises five core modules: a tool call monitoring module, a failure diagnosis module, a context attachment module, a differential repair module, and a breakpoint continuation execution module. These five core modules are sequentially cascaded and deployed in a bypass position outside the Agent's main inference loop, forming a transparent interception layer over the Agent's core inference process without modifying the parameters or inference logic of the large language model itself.
[0060] The tool call monitoring module corresponds to step S1 of the method described in Embodiment 1, and is deployed between the Agent inference main loop and the actual entry point of the external tool call. The software layer of the tool call monitoring module includes three sub-modules: a request interception sub-module, an execution return result monitoring sub-module, and a failure event recording sub-module. The request interception sub-module implements non-intrusive interception of the tool call request using either the decorator pattern or the proxy pattern. The execution return result monitoring sub-module monitors the execution return result in real time according to the four types of failure signal determination logic described in step S1. The failure event recording sub-module constructs the tool call failure event and enqueues it into the failure event queue when a failure signal is hit. The interfaces exposed by the tool call monitoring module are a standard tool call interception interface and a failure event subscription interface. The failure event subscription interface is subscribed to by the failure diagnosis module to receive the tool call failure event stream.
[0061] The failure diagnosis module corresponds to step S2 of the method described in Embodiment 1, and undertakes the inference task of the error type classification model. The software layer of the failure diagnosis module includes four components: the error pattern regularity extractor set, the positional encoding normalization processing submodule, the model loader and inferencer of the error type classification model, and the root cause category post-processing submodule. The error pattern regularity extractor set includes three parallel sub-extractors: the API error regularity extractor, the database error regularity extractor, and the code execution error regularity extractor. The corresponding sub-extractor is activated according to the source tool type of the failure error information text to output the error information seven-tuple. The positional encoding normalization processing submodule performs normalization mapping and sine positional encoding on the failure step sequence number according to the two-stage calculation process described in step S2. The model loader of the error type classification model loads pre-trained weights from the model repository when the system starts. The inferencer performs forward propagation calculation when each tool call failure event arrives, outputting the root cause category index. The root cause category post-processing submodule maps the root cause category index to a Chinese name and constructs the diagnosis result object. The failure diagnosis module exposes two interfaces: a root cause category reasoning interface and a diagnosis result publishing interface. The diagnosis result publishing interface is subscribed to by the context attachment module to receive the diagnosis result stream.
[0062] The context attachment module corresponds to step S3 of the method described in Embodiment 1, and undertakes the task of forcibly attaching the diagnostic result to the Agent context. The software layer of the context attachment module includes three components: an Agent context operation submodule, a diagnostic result declaration statement template rendering submodule, and a dialogue history simplification submodule. These components are implemented according to the diagnostic anchor insertion logic, preset template filling logic, and the strategy for retaining the last 5 rounds of dialogue history described in step S3, respectively. The interfaces exposed by the context attachment module are a context writing interface and a simplified context retrieval interface. The simplified context retrieval interface is subscribed to by the differentiated repair module to trigger subsequent repair processes.
[0063] The differentiated repair module corresponds to step S4 of the method described in Embodiment 1, and undertakes the task of selecting and executing the repair strategy based on the diagnostic results. The software layer of the differentiated repair module includes four components: a repair strategy routing table storage submodule, four types of repair strategy executor submodules, a repair inference budget upper limit dynamic constraint submodule, and a repair history log maintenance submodule. The repair strategy routing table storage submodule stores the predefined mapping relationship between the root cause category and the repair strategy using a hash table structure. The four types of repair strategy executor submodules are, respectively, a tool document rereading and parameter reconstruction executor, a credential refresh tool call executor, a backup tool replacement search executor, and an output format re-inference executor, and implement repair according to their respective execution processes described in step S4. The repair inference budget upper limit dynamic constraint submodule monitors the repair inference token consumption under each root cause category in real time according to the repair inference budget upper limit dynamic constraint formula and the historical success rate weight calculation formula described in step S4, and immediately stops the repair attempt when the upper limit is reached. The repair history log maintenance submodule updates the n_{succ}(c) and n_{fail}(c) statistical counts after each repair strategy execution. The differentiated repair module exposes two interfaces: a repair strategy execution interface and a repair result publishing interface. The repair result publishing interface is subscribed to by the breakpoint continuation execution module.
[0064] The breakpoint continuation execution module corresponds to step S5 of the method described in Embodiment 1, and undertakes the tasks of merging the execution stack and resuming the multi-step task from the breakpoint of the failed step. The software layer of the breakpoint continuation execution module includes four components: the execution stack state tracking submodule, the execution stack snapshot storage management submodule, the continuation condition function determination submodule, and the lightweight rollback execution submodule. The execution stack state tracking submodule inserts state hooks to update the execution stack state flags before each tool call request is issued by the Agent and after the execution return result is received. The execution stack snapshot storage management submodule is responsible for the Pickle protocol serialization and persistence of the intermediate execution results. The continuation condition function determination submodule outputs a continuation path selection signal based on the execution stack state flags according to the continuation condition function calculation formula described in step S5. When the continuation path selection signal is 0, the lightweight rollback execution submodule restores part of the state changes according to the reverse operation process described in step S5. The interfaces exposed by the breakpoint resume execution module are the execution stack merging interface and the multi-step task resume interface. The multi-step task resume interface is subscribed to by the Agent inference main loop to resume the execution of the multi-step task.
[0065] The large language model Agent tool call failure self-repair system described in this embodiment presents a complete diagnosis-repair causal coupling closed loop in a cross-module collaborative workflow. The tool call failure event output by the tool call monitoring module serves as the input to the failure diagnosis module; the diagnosis result output by the failure diagnosis module serves as the input to the context attachment module; the Agent context output by the context attachment module serves as the input to the differential repair module; the repair result output by the differential repair module serves as the input to the breakpoint resume execution module; and the resume execution stack output by the breakpoint resume execution module updates the state of the Agent inference main loop in reverse. The five core modules communicate decoupledly through a standardized publish-subscribe message bus.
[0066] The overall technical effect of the large language model agent tool call failure self-repair system described in this embodiment completely corresponds to the beneficial effect of the method described in Embodiment 1. The root cause category identification accuracy reaches 94.3%, the average inference token consumption per repair is reduced by about 28% compared with the baseline solution, and the end-to-end success rate of multi-step agent tasks is improved from 63.4% of the baseline to 91.2%. The system described in this embodiment has good scalability. The rapid access of new external tools only requires configuring the corresponding error pattern regular expression extractor for the new tools, and the expansion of new root cause categories only requires incremental fine-tuning of the error type classification model on the new samples.
[0067] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A self-repair method for failed calls to the large language model Agent tool, characterized in that: Includes the following steps: Receive tool call requests sent by the Agent to external tools during the execution of multi-step tasks, monitor the execution return results of the tool call requests, and record a tool call failure event when the execution return result hits a failure signal. The tool call failure event includes a failure error message text and a failure step number. The failure error message text is input into the error type classification model, and the error type classification model outputs the root cause category of the tool call failure. The root cause category is one of the following four: parameter format error, authorization failure, tool service unavailable, and return value parsing failure. The root cause category is forcibly attached to the Agent context as a diagnostic result, and the diagnostic result serves as the input condition for generating the remediation strategy. Based on the diagnostic results, the corresponding repair strategy is selected from the repair strategy routing table and executed. The repair strategy routing table includes: tool document rereading and parameter reconstruction for parameter format errors, credential refresh tool call for invalid authorization, backup tool replacement search for unavailable tool services, and output format re-inference for failed return value parsing. After the repair strategy is successfully executed, the multi-step task is resumed from the breakpoint at the failed step number, and the intermediate execution results of the steps completed before the failed step number are retained.
2. The self-repair method for failed calls to the large language model Agent tool according to claim 1, characterized in that, In the step of inputting the failure error information text into the error type classification model, the input of the error type classification model also includes a tool invocation context summary, which includes the tool identifier of the invoked tool, the request parameters of the tool invocation request, and the failure step number; the error type classification model is a multi-classification model obtained by fine-tuning a pre-trained encoder, and the root cause category identification accuracy of the multi-classification model on the multi-tool agent execution log validation set is not less than 94.3%.
3. The self-repair method for failed calls to the large language model Agent tool according to claim 2, characterized in that, The failure step number in the tool's context summary is input into the error type classification model after being processed by positional encoding normalization. The positional encoding normalization process involves normalizing the failure step number to the [0,1] interval and then converting it into a positional encoding vector of a preset dimension using sinusoidal positional encoding. The denominator of the normalization mapping is the preset maximum number of steps in the multi-step task, and the dimension of the positional encoding vector is a preset even value.
4. The self-repair method for failed calls to the large language model Agent tool according to claim 3, characterized in that, The input to the error type classification model undergoes error information heterogeneity normalization processing in the pre-processing stage. This processing maps the failure error information text to a unified error information seven-tuple through a set of error pattern regular expression extractors. The error information seven-tuple includes the error code, error phrase, error level, suspicious parameters, number of prompts, HTTP status, and exception class name. The set of error pattern regular expression extractors includes an API error regular expression extractor, a database error regular expression extractor, and a code execution error regular expression extractor.
5. The self-repair method for failed calls to the large language model Agent tool according to claim 4, characterized in that, In the step of selecting and executing the corresponding repair strategy from the repair strategy routing table based on the diagnostic results, each root cause category has a repair inference budget upper limit for the repair strategy. The repair inference budget upper limit is dynamically constrained by the historical success rate weight, which is the ratio of the number of historical repair successes for the root cause category to the total number of historical repairs for the root cause category.
6. The self-repair method for failed calls to the large language model Agent tool according to claim 5, characterized in that, In the step of resuming the multi-step task from the failed step number, the state change of the step corresponding to the failed step number is marked by an execution stack state flag. The execution stack state flag is one of three: a partially written flag, a rolled-back flag, and a committed flag. When the execution stack state flag is a committed flag, the task is resumed directly from the breakpoint. When the execution stack state flag is a partially written flag, a lightweight rollback is triggered first, and then the task is resumed from the breakpoint.
7. The self-repair method for failed calls to the large language model Agent tool according to claim 1, characterized in that, The multi-step task is one of three: financial transaction automation task, medical record query automation task, and enterprise data dashboard automation task.
8. The self-repair method for failed calls to the large language model Agent tool according to claim 1, characterized in that, The large language model is a pre-trained large language model based on the Transformer decoder architecture with no less than 7 billion parameters.
9. The self-repair method for failed calls to the large language model Agent tool according to claim 1, characterized in that, The external tool is at least one of the following: RESTful API interface, SQL database query interface, Python code execution sandbox and vector retrieval interface.
10. A self-repairing system for failed calls to large language model agent tools, used to implement the self-repairing method for failed calls to large language model agent tools as described in any one of claims 1-9, characterized in that, include: The tool call monitoring module is used to receive tool call requests sent by the Agent to external tools during the execution of multi-step tasks, monitor the execution return results of the tool call requests, and record a tool call failure event when the execution return result hits a failure signal. The tool call failure event includes a failure error message text and a failure step number. The failure diagnosis module is used to input the failure error information text into the error type classification model. The error type classification model outputs the root cause category of the tool call failure. The root cause category is one of the following four: parameter format error, authorization failure, tool service unavailable, and return value parsing failure. The context attach module is used to forcibly attach the root cause category as a diagnostic result to the Agent context, and the diagnostic result serves as the input condition for generating the remediation strategy. The differentiated repair module is used to select and execute the corresponding repair strategy from the repair strategy routing table based on the diagnostic results. The repair strategy routing table includes: rereading and reconstructing the tool document for parameter format errors, calling the credential refresh tool for invalid permission authentication, searching for backup tools for unavailable tool services, and re-inferring the output format for failed return value parsing. The breakpoint continuation module is used to resume the multi-step task from the breakpoint at the failed step number after the repair strategy is successfully executed, and to retain the intermediate execution results of the steps completed before the failed step number.
Citation Information
Patent Citations
Knowledge graph-based large model task calling execution method and device
CN118963869A
Intelligent agent system based on large language model and interaction method
CN120633639A