A Method for Agent Response Rewriting and Reflective Intervention Based on Task Completion Monitoring

CN122840100APending Publication Date: 2026-09-29CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611035545.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,实际部署中该模式面临两类突出的可靠性问题:其一,模型可能在任务清单中仍存在未完成项时即输出结束信号,造成需求遗漏;其二,模型可能直接生成最终答复而根本未将请求拆解为结构化任务清单,使系统无法辨别任务是“真实完成”还是“根本未被拆解”

Benefits of technology

[0040]通过在工具调用循环中实时拦截模型响应并读取结束原因、任务清单及反思状态标记,实现对智能体运行时“提前终止”与“空任务清单结束”两类异常情形的精确检测;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840100A_ABST
    Figure CN122840100A_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent agent control technology, and provides a method for intelligent agent response rewriting and reflection intervention based on task completion monitoring. The method includes: determining whether the intervention trigger condition is met by combining the termination reason in the read original response, the completion degree of the task list, and the value of the reflection state flag; when the intervention trigger condition is met, rewriting the termination reason in the original response from a state representing normal stopping to a first state representing the need to call a tool, injecting a tool call object into the original response, and generating a rewritten response; generating a state update command, which is returned to the runtime environment along with the rewritten response; executing the reflection tool in the tool call loop, obtaining the returned preset structured guidance prompts, and driving the large language model to select the corresponding branch from multiple preset execution branches to continue execution based on the current task completion status. This invention can improve the controllability and task completion rate of intelligent agent execution in complex task scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent agent control technology, and in particular to an intelligent agent response rewriting and reflection intervention method based on task completion monitoring. Background Technology

[0002] Intelligent agent systems based on large language models have been widely applied in complex multi-step reasoning and external tool invocation scenarios. Their typical operating mode is a tool invocation loop: after receiving a user request, the model generates a structured task list, iterates through multiple rounds to call external tools to obtain information or perform operations, and finally returns a response. However, in actual deployment, this mode faces two prominent reliability issues: first, the model may output a termination signal when there are still incomplete items in the task list, resulting in missed requests; second, the model may directly generate a final response without breaking down the request into a structured task list, making it impossible for the system to distinguish whether the task is "truly completed" or "not broken down at all." Prompt word engineering alone cannot fundamentally solve these problems; a runtime intervention mechanism is urgently needed to ensure the integrity and controllability of the agent's execution.

[0003] Existing intervention techniques have the following shortcomings in practical applications:

[0004] 1. The reflexive mechanism is limited to the offline phase and cannot address premature termination of the agent during runtime. Reflexive mechanisms all occur during the offline training data generation or prior knowledge construction phase; when the agent terminates prematurely in the tool call loop ("the model gives finish_reason="stop" but there are still unfinished items in the task list"), there is a lack of runtime intervention. 2. Task list completion information is not used for response layer judgment. Although a task dependency graph is constructed and scheduling is based on it, the task list is only used for scheduling matching and not as a basis for determining "whether the model should terminate"; when the model ignores the task list and terminates directly, there is no mechanism to intervene in the response based on the task list completion. 3. The lack of joint intervention between the response layer and the state layer easily leads to infinite loops. If only the response is rewritten without updating the state, the model may still give finish_reason="stop" in the next iteration, resulting in an infinite loop; existing middleware architectures generally only perform single-time interception and do not disclose a joint intervention mechanism of "response rewriting + state command backfeedback". 4. There is a lack of handling for "task not broken down" branches. When the task list is empty and the model ends directly, it is impossible to determine whether this means "the task is truly completed" or "the task was never broken down in the first place".

[0005] Chinese patent CN121659990A discloses an agent reinforcement method. It uses an evaluator to perform a success / failure binary judgment on the agent's complete execution trajectory in an Alfworld environment. Only after a trajectory is judged as a failure is the Self-Reflection module triggered to generate reflection text offline. High-quality reflections are then filtered by a hierarchical validator and stored in a vector experience base for reuse in subsequent tasks, thus achieving non-parametric iterative optimization of the agent's decision-making ability. However, this mechanism is entirely limited to the offline stage, meaning it only generates and stores reflection texts after a failed trajectory has been completed. It cannot monitor or intervene in the model's response in real time during agent runtime, and especially cannot address the issues of immediate response rewriting and infinite loop prevention when the model terminates prematurely due to an output termination signal in a tool call loop, but there are still unfinished items in the task list.

[0006] Therefore, how to provide an intervention method that can ensure the integrity and controllability of intelligent agent execution at runtime has become an urgent technical problem to be solved. Summary of the Invention

[0007] In view of this, in order to overcome the shortcomings of the prior art, the present invention aims to provide a method for rewriting and reflecting on the response of an intelligent agent based on task completion monitoring.

[0008] This invention provides a method for agent response rewriting and reflection intervention based on task completion monitoring, the method comprising the following steps:

[0009] Step S1: In the tool call loop of the agent based on the large language model, intercept the original response returned by the large language model, read the termination reason in the original response and the task list and reflection status flag in the runtime state; make a combined judgment based on the termination reason, the completion degree of the task list and the value of the reflection status flag to determine whether the intervention trigger condition is met.

[0010] Step S2: When the intervention triggering condition is met, the termination reason in the original response is rewritten from the state representing normal cessation to the first state representing the need to call the tool. A tool call object containing the name of the reflection tool is injected into the original response to generate the rewritten response.

[0011] Step S3: Generate a state update command carrying a first preset value. By returning the state update command and the rewritten response to the runtime environment, the intervention under the same conditions in the current tool call loop is prevented from being triggered repeatedly.

[0012] Step S4: The tool call loop executes the reflection tool, obtains the preset structured guidance prompts returned by the reflection tool, and drives the large language model to select the corresponding branch from multiple preset execution branches according to the current task completion status to continue execution until all tasks are completed or terminated normally.

[0013] Optionally, in the agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, step S1 involves a combined determination as follows:

[0014] When the termination reason is characterized as normal stop, the reflection status flag is set to the second preset value representing no intervention, and the task list is either empty or contains at least one of the pending items in a non-complete state, it is determined that intervention needs to be triggered.

[0015] When the termination reason is normal stop, the reflection status flag is set to the second preset value, the task list is not empty and all pending items are in the completed state, it is determined that no intervention is required.

[0016] When the termination reason is not characterized as normal stop, or the reflection status flag value is not the second preset value, it is determined that no intervention is required.

[0017] Optionally, in the agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, in step S1, the completion rate of the task list is determined in the following manner:

[0018] The task list contains at least one to-do item, and each to-do item has a status field.

[0019] When the status field value of all pending items in the task list is equal to the preset string indicating the completed status, the task list is considered to be fully completed.

[0020] If the status field value of any pending item in the task list is not equal to the preset string indicating the completed status, it is determined that there are incomplete items in the task list;

[0021] The task list is considered empty if it does not contain any to-do items.

[0022] Optionally, in the agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, in step S2, the rewritten response is generated in the following manner:

[0023] Locate the first response data in the original response, and rewrite the end reason field value in the metadata field of the first response data from a string indicating normal termination to a string indicating that a tool needs to be invoked;

[0024] Generate a globally unique call identifier, use the call identifier as the identifier, the preset reflection tool name as the tool name, and the empty parameter object as the parameter to construct a tool call object, and mark the type of the tool call object as the tool call type;

[0025] Write the constructed tool call object into the tool call field of the first response data to form the rewritten first response data.

[0026] Optionally, in the intelligent agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, in step S2, a globally unique call identifier is generated in the following manner: a random string composed of hexadecimal random characters is generated with a preset fixed string as a prefix, and the prefix and the random string are concatenated to form a globally unique call identifier within the current tool call loop; the hexadecimal random string contains at least 16 hexadecimal characters.

[0027] Optionally, in the intelligent agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, after forming the rewritten response in step S2, message synchronization is further performed in the following manner: the rewritten first response data is appended to the message queue of the current request and to the message storage field in the runtime state of the intelligent agent.

[0028] Optionally, in the intelligent agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, in step S3, a first preset value is used to update the reflection state marker to represent the intervened state, and the first preset value is an integer.

[0029] Optionally, in the agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, step S3 prevents the repeated triggering of interventions under the same conditions in the current tool call loop in the following manner:

[0030] Construct a command object containing a state update instruction, which is used to set the value of the reflection state flag field in the runtime state to a first preset value;

[0031] The command object and the rewritten response are encapsulated into the same return result and fed back into the agent's graph state machine;

[0032] When it is determined that no intervention is needed and a normal response is returned, another command object containing a state update instruction is constructed, and the value of the reflection state flag field is reset to a second preset value. The second preset value is an integer used to represent the no-intervention state.

[0033] Optionally, in the intelligent agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, in step S4, the multiple preset execution branches include at least the following task list reconstruction branch: when the intervention triggering reason is a branch where the task list is empty, the structured guidance prompt prioritizes guiding the large language model to call the preset task list reconstruction tool; after the task list reconstruction tool is called, the task list is regenerated according to the current user request, and the reconstructed task list is written into the runtime state as the basis for subsequent completion determination.

[0034] Optionally, in the agent response rewriting and reflection intervention method based on task completion monitoring of the present invention, in step S4, the preset structured guidance prompt returned by the reflection tool is predefined text, which includes guidance content in the following four branches:

[0035] The first branch, when all tasks have been completed, guides the large language model to output a confirmation message that all tasks have been completed;

[0036] The second branch guides the large language model to output a request for further information when the user needs to provide additional information.

[0037] The third branch guides the large language model to continue executing unfinished tasks when there are still unfinished tasks to be done.

[0038] The fourth branch, when the task list is empty, guides the large language model to first call the task list reconstruction tool to create the task list before continuing execution.

[0039] The agent response rewriting and reflection intervention method based on task completion monitoring of the present invention has the following beneficial technical effects:

[0040] By intercepting the model response in real time during the tool call loop and reading the termination reason, task list, and reflection status flag, accurate detection of two abnormal situations during agent runtime is achieved: "early termination" and "end of empty task list".

[0041] By rewriting responses in place and injecting reflection tools, responses that were originally terminated are transformed into tool calls, allowing the model to continue executing unfinished tasks; the state command backfeed mechanism effectively avoids repeated triggering of interventions in the same loop.

[0042] Through the structured guidance of reflection tools, the model is driven to select a reasonable path from multiple preset branches to continue execution, thereby significantly improving the task completion rate and execution reliability of the agent in multi-step task scenarios without modifying the underlying model parameters. Attached Figure Description

[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart illustrating the agent response rewriting and reflection intervention method based on task completion monitoring according to an exemplary embodiment 1 of the present invention.

[0045] Figure 2 A flowchart illustrating the combined judgment logic of the agent response rewriting and reflection intervention method based on task completion monitoring according to an exemplary embodiment 1 of the present invention;

[0046] Figure 3 This is a flowchart illustrating the process of determining the completion of a task list according to the intelligent agent response rewriting and reflection intervention method based on task completion monitoring according to an exemplary embodiment 1 of the present invention.

[0047] Figure 4 This is a flowchart illustrating the task list reconstruction branch of the intelligent agent response rewriting and reflection intervention method based on task completion monitoring according to an exemplary embodiment 1 of the present invention.

[0048] Figure 5 This is a schematic diagram of a four-branch structured guidance prompt for an agent response rewriting and reflection intervention method based on task completion monitoring according to an exemplary embodiment 1 of the present invention. Detailed Implementation

[0049] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0050] It should be noted that, in the absence of conflict, the following embodiments and features can be combined with each other; and, based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0051] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0052] Example 1

[0053] Exemplary embodiment 1 of the present invention provides a method for agent response rewriting and reflection intervention based on task completion monitoring. Figure 1 This is a flowchart illustrating the agent response rewriting and reflection intervention method based on task completion monitoring according to an exemplary embodiment 1 of the present invention. Figure 1 As shown, in this embodiment, the method of the present invention is implemented in the following manner:

[0054] Step S1: In the tool call loop of the agent based on the large language model, intercept the original response returned by the large language model, read the termination reason in the original response and the task list and reflection status flag in the runtime state; make a combined judgment based on the termination reason, the completion degree of the task list and the value of the reflection status flag to determine whether the intervention trigger condition is met.

[0055] Figure 2 This is a flowchart illustrating the combined judgment logic of the agent response rewriting and reflection intervention method based on task completion monitoring according to an exemplary embodiment 1 of the present invention, as follows: Figure 2 As shown, in this embodiment, the combination determination is performed in the following manner:

[0056] When the termination reason is characterized as normal stop, the reflection status flag is set to the second preset value indicating no intervention, and the task list is either empty or contains at least one to-do item in a non-complete state, intervention is determined to be required. When the termination reason is characterized as normal stop, the reflection status flag is set to the second preset value, the task list is not empty, and all to-do items are in a completed state, intervention is determined to be unnecessary. When the termination reason is not characterized as normal stop, or the reflection status flag is not set to the second preset value, intervention is determined to be unnecessary.

[0057] This embodiment defines the specific logic for determining intervention by combining the termination reason, the reflection status flag value, and the task list completion degree. It clarifies the boundary conditions for triggering intervention and not needing intervention, making the intervention decision-making process deterministic and reproducible. It avoids the instability of heuristic or probabilistic judgments, provides clear and strict judgment rules, and ensures the interpretability and testability of the intervention logic.

[0058] Figure 3 This is a flowchart illustrating the process of determining task list completion according to the agent response rewriting and reflection intervention method based on task completion monitoring according to Exemplary Embodiment 1 of the present invention. Figure 3 As shown, in this embodiment, the completion rate of the task list is determined in the following manner:

[0059] The task list contains at least one to-do item, and each to-do item has a status field.

[0060] When the status field value of all pending items in the task list is equal to the preset string indicating the completed status, the task list is considered to be fully completed.

[0061] If the status field value of any pending item in the task list is not equal to the preset string indicating the completed status, it is determined that there are incomplete items in the task list;

[0062] The task list is considered empty if it does not contain any to-do items.

[0063] This embodiment establishes a unified and objective quantitative indicator for the completion of a task list by defining specific standards for the determination of the completion rate. This allows for an accurate distinction between the two states of "task not completed" and "task truly completed," providing a reliable input basis for subsequent intervention and judgment, and avoiding misjudgments or omissions caused by ambiguous definitions of completion.

[0064] Step S2: When the intervention trigger condition is met, the termination reason in the original response is rewritten from the state representing normal cessation to the first state representing the need to call the tool. A tool call object containing the name of the reflection tool is injected into the original response to generate the rewritten response.

[0065] In this embodiment, the rewritten response is generated as follows: Locate the first response data in the original response, and rewrite the end reason field value in the metadata field of the first response data from a string representing normal stopping to a string representing the need to call a tool; generate a globally unique call identifier, use the call identifier as the identifier, a preset reflection tool name as the tool name, and an empty parameter object as the parameter to construct a tool call object, and mark the type of the tool call object as the tool call type; write the constructed tool call object into the tool call field of the first response data to form the rewritten first response data.

[0066] This embodiment provides a standardized, non-intrusive response modification scheme by limiting the specific operation steps of response rewriting. It simulates the tool call behavior that the model should output at the response level, so that subsequent tool call loops can be seamlessly connected without modifying the core scheduling logic of the agent, thus reducing the difficulty of engineering implementation and compatibility risks.

[0067] For example, the method in this embodiment generates a globally unique call identifier as follows: A random string composed of hexadecimal random characters is generated using a preset fixed string as a prefix; the prefix is ​​concatenated with the random string to form a globally unique call identifier within the current tool call loop; the hexadecimal random string contains at least 16 hexadecimal characters. This embodiment ensures the global uniqueness of the call identifier within a single tool call loop by defining the rules for generating the call identifier, avoiding call tracking confusion or tool execution abnormalities caused by identifier conflicts between multiple tool calls, and improving stability in concurrent or long-running scenarios.

[0068] It should be noted that after the rewritten response is generated, this embodiment also includes message synchronization in the following manner: the rewritten first response data is appended to the message queue of the current request and to the message storage field in the runtime state of the agent.

[0069] This embodiment limits the message synchronization steps, and simultaneously appends the rewritten response to the request message queue and the runtime state message storage, ensuring the complete visibility and consistency of the rewritten response in the subsequent context. This allows subsequent model calls to be aware of the injected tool call events, thereby maintaining the logical continuity of the dialogue state and the tool call chain.

[0070] Step S3: Generate a state update command carrying a first preset value. By returning the state update command and the rewritten response to the runtime environment, the intervention under the same conditions in the current tool call loop is prevented from being triggered repeatedly.

[0071] In this embodiment, interventions under the same conditions in the current tool call loop are prevented from being repeatedly triggered in the following manner:

[0072] A command object containing a state update instruction is constructed. This state update instruction is used to set the value of the reflection state flag field in the runtime state to a first preset value. The first preset value is used to update the reflection state flag to represent the intervened state. The first preset value is an integer, which provides a direct state control means for the anti-infinite loop mechanism, ensuring that the intervention action can be recorded once it is triggered, so that it can be effectively shielded in the subsequent judgment in the same round. This design is simple and efficient, requiring only an integer field to realize state memory without adding extra system complexity.

[0073] The command object and the rewritten response are encapsulated into the same return result and fed back into the agent's graph state machine. When it is determined that no intervention is needed and a normal response is returned, another command object containing a state update instruction is constructed, and the value of the reflection state flag field is reset to a second preset value. The second preset value is an integer used to represent the uninterrupted state. This ensures that the intervention in this round is triggered only once, and also ensures that the intervention capability of the next new task is restored normally, fundamentally eliminating the risk of infinite loops caused by intervention logic, while ensuring the continuity of multi-round task processing.

[0074] Step S4: The tool call loop executes the reflection tool, obtains the preset structured guidance prompts returned by the reflection tool, and drives the large language model to select the corresponding branch from multiple preset execution branches according to the current task completion status to continue execution until all tasks are completed or terminated normally.

[0075] Figure 4 This is a flowchart illustrating the task list reconstruction branch of the agent response rewriting and reflection intervention method based on task completion monitoring according to Exemplary Embodiment 1 of the present invention. Figure 4 As shown, in this embodiment, the multiple preset execution branches include at least the following task list reconstruction branches: when the intervention triggering reason is a task list empty branch, the structured guidance prompt will first guide the large language model to call the preset task list reconstruction tool; after the task list reconstruction tool is called, the task list will be regenerated according to the current user request, and the reconstructed task list will be written into the runtime state as the basis for subsequent completion determination.

[0076] This embodiment limits the task list reconstruction branch. When the intervention trigger reason is "task list is empty", it prioritizes guiding the model to call the task list reconstruction tool to achieve forced remedy for the situation of "tasks not being broken down". This enables the model to actively rebuild the structured task list at runtime, fundamentally avoiding the omission of key tasks due to model laziness or misunderstanding.

[0077] Figure 5 This is a schematic diagram of the four-branch structured guidance prompt of the agent response rewriting and reflection intervention method based on task completion monitoring according to Exemplary Embodiment 1 of the present invention, as shown below. Figure 5 As shown, in this embodiment, the preset structured guidance prompts returned by the reflection tool are predefined text, which includes guidance content in the following four branches:

[0078] The first branch, when all tasks have been completed, guides the large language model to output a confirmation message that all tasks have been completed;

[0079] The second branch guides the large language model to output a request for further information when the user needs to provide additional information.

[0080] The third branch guides the large language model to continue executing unfinished tasks when there are still unfinished tasks to be done.

[0081] The fourth branch, when the task list is empty, guides the large language model to first call the task list reconstruction tool to create the task list before continuing execution.

[0082] This embodiment limits the pre-defined structured guidance prompts returned by the reflection tool to include four clear branches, providing the model with a clear and limited decision space, enabling the model to quickly select the correct subsequent path based on the current context; at the same time, the prompts are predefined text, which does not introduce additional model inference overhead, thus improving controllability while maintaining efficiency, achieving a reasonable balance between guidance effect and computational cost.

[0083] Example 2

[0084] Exemplary Example 2 of the present invention provides a method for rewriting and reflecting on the response of an agent based on task completion monitoring. In this embodiment, the method of this embodiment is further described in detail in a specific scenario.

[0085] The method in this embodiment is implemented in an architecture that includes the following modules:

[0086] The termination reason and completion determination module is located in the response rewriting middleware layer and is called after each model call. It reads the termination reason (finish_reason) from the model response and the task list (todos) and reflection status flag (reflect_stop) from the runtime state. It then determines whether intervention is needed based on the following combined logic:

[0087] If finish_reason == "stop" and reflect_stop == 0:

[0088] If todos is not empty and not all todos have been completed, or

[0089] If todos is empty, it is determined that "intervention is needed";

[0090] All other cases are deemed "no intervention required".

[0091] The response in-situ rewriting module: When it is determined that "intervention is needed", the response_metadata["finish_reason"] in the first line of the response is rewritten from "stop" to "tool_calls"; and a tool call object {args: {}, id: <newly generated call ID>, name: <reflection tool name>, type: "tool_call"} is constructed and written into the tool_calls field of the first line of the response; at the same time, the rewritten first line of the response is appended to request.messages and state["messages"].

[0092] The Reflection Tool Injection Module is responsible for generating a globally unique call ID (e.g., prefixed with "call_" followed by a random hexadecimal string) and binding this ID to the injected tool call object along with the reflection tool name.

[0093] The status command feedback module constructs a status update command (update={"reflect_stop": 1}) and returns it along with the rewritten response. When the model calls the response rewriting middleware layer again, since reflect_stop is already 1, interventions under the same conditions will not be triggered again, thus avoiding infinite loops. When it is determined that "no intervention is needed", a command (update={"reflect_stop": 0}) is constructed to reset the flag, ensuring that the next round of new tasks is triggered normally.

[0094] Reflection Tool: After the injected tool call is executed in the next tool loop, a structured guidance prompt is returned, requiring the model to choose one of four branches for output:

[0095] If all tasks have been completed, output "After reflection and confirmation, all tasks have been completed";

[0096] If the user needs to provide additional information, output "Further information required";

[0097] If there are still unfinished todos, continue execution;

[0098] If the tasks are not broken down as required, the task list rebuilding tool (write_todos) will be called first to create the todos, and then execution will continue.

[0099] Task list reconstruction tool: When todos are empty and the model ends directly, the task list reconstruction tool is triggered by the reflection tool, and the model reconstructs todos according to the user's request; the reconstructed todos are written into the runtime state as the basis for subsequent completion determination.

[0100] Example 3

[0101] Exemplary Example 3 of the present invention provides a method for rewriting and reflecting on agent responses based on task completion monitoring. In this embodiment, the method of the present invention is implemented in the following manner:

[0102] The model call request is preprocessed by several middleware components before calling the underlying large language model to obtain the original response. The response enters the response rewriting middleware, which reads finish_reason from the response and todos and reflect_stop from the runtime state.

[0103] If finish_reason != "stop" or reflect_stop != 0, it is determined that "no intervention is needed" and a normal response is returned; at the same time, reflect_stop is reset to 0 through the status command to ensure that the next round of new tasks can be triggered normally.

[0104] If `finish_reason == "stop"` and `reflect_stop == 0`, further determine `todos`: if `todos` is not empty and all are completed, the task is considered truly completed and determined as "no intervention required". If `todos` is not empty and not all are completed, or if `todos` is empty, it is determined as "intervention required". The rewrite steps include:

[0105] Rewrite the response_metadata["finish_reason"] in the first response from "stop" to "tool_calls"; generate a globally unique call ID ("call_" + hexadecimal random string); construct a tool call object {args:{}, id: <call ID>, name: <reflect tool name>, type: "tool_call"}, and write it into the tool_calls field of the first response; append the rewritten first response to request.messages and state["messages"]; construct the state update command Command(update={"reflect_stop": 1}), and return it together with the rewritten response. If the first response is empty (result.result[0] does not exist), skip the intervention and return the unmodified response to avoid incorrect rewriting of abnormal responses.

[0106] In the next tool loop, the injected reflexive tool call is executed, returning a structured guidance prompt. The model selects a branch based on the guidance prompt:

[0107] Branch A: All tasks have been completed; output confirmation.

[0108] Branch B: Requires user to provide additional information; output a request for further information.

[0109] Branch C: There are still unfinished todos; continue execution.

[0110] Branch D: Tasks were not broken down as required. First, the task list reconstruction tool (write_todos) is called to create todos, and then execution continues.

[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0112] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for agent response rewriting and reflection intervention based on task completion monitoring, characterized in that, Includes the following steps: Step S1: In the tool call loop of the agent based on the large language model, intercept the original response returned by the large language model, read the termination reason in the original response and the task list and reflection status flag in the runtime state; make a combined judgment based on the termination reason, the completion degree of the task list and the value of the reflection status flag to determine whether the intervention trigger condition is met. Step S2: When the intervention triggering condition is met, the termination reason in the original response is rewritten from the state representing normal cessation to the first state representing the need to call the tool. A tool call object containing the name of the reflection tool is injected into the original response to generate the rewritten response. Step S3: Generate a state update command carrying a first preset value. By returning the state update command and the rewritten response to the runtime environment, the intervention under the same conditions in the current tool call loop is prevented from being triggered repeatedly. Step S4: The tool call loop executes the reflection tool, obtains the preset structured guidance prompts returned by the reflection tool, and drives the large language model to select the corresponding branch from multiple preset execution branches according to the current task completion status to continue execution until all tasks are completed or terminated normally.

2. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 1, characterized in that, In step S1, the combination determination is performed as follows: When the termination reason is characterized as normal stop, the reflection status flag is set to the second preset value representing no intervention, and the task list is either empty or contains at least one of the pending items in a non-complete state, it is determined that intervention needs to be triggered. When the termination reason is normal stop, the reflection status flag is set to the second preset value, the task list is not empty and all pending items are in the completed state, it is determined that no intervention is required. If the termination reason is not characterized as normal stopping, or the reflection status flag value is not the second preset value, it is determined that no intervention is required.

3. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 1, characterized in that, In step S1, the completion rate of the task list is determined as follows: The task list contains at least one to-do item, and each to-do item has a status field. When the status field value of all pending items in the task list is equal to the preset string indicating the completed status, the task list is considered to be fully completed. If the status field value of any pending item in the task list is not equal to the preset string indicating the completed status, it is determined that there are incomplete items in the task list; The task list is considered empty if it does not contain any to-do items.

4. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 1, characterized in that, In step S2, the rewritten response is generated as follows: Locate the first response data in the original response, and rewrite the end reason field value in the metadata field of the first response data from a string indicating normal termination to a string indicating that a tool needs to be invoked; Generate a globally unique call identifier, use the call identifier as the identifier, the preset reflection tool name as the tool name, and the empty parameter object as the parameter to construct a tool call object, and mark the type of the tool call object as the tool call type; Write the constructed tool call object into the tool call field of the first response data to form the rewritten first response data.

5. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 4, characterized in that, In step S2, a globally unique call identifier is generated as follows: a random string consisting of hexadecimal random characters is generated using a preset fixed string as a prefix, and the prefix is ​​concatenated with the random string to form a globally unique call identifier within the current tool call loop; the hexadecimal random string contains at least 16 hexadecimal characters.

6. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 4, characterized in that, In step S2, after the rewritten response is generated, message synchronization is performed as follows: the rewritten first response data is appended to the message queue of the current request and to the message storage field in the runtime state of the agent.

7. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 1, characterized in that, In step S3, the first preset value is used to update the reflection state marker to represent the intervened state, and the first preset value is an integer.

8. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 1, characterized in that, In step S3, prevent interventions under the same conditions from being repeatedly triggered in the current tool call loop in the following manner: Construct a command object containing a state update instruction, which is used to set the value of the reflection state flag field in the runtime state to a first preset value; The command object and the rewritten response are encapsulated into the same return result and fed back into the agent's graph state machine; When it is determined that no intervention is needed and a normal response is returned, another command object containing a state update instruction is constructed, and the value of the reflection state flag field is reset to a second preset value. The second preset value is an integer used to represent the no-intervention state.

9. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 1, characterized in that, In step S4, the multiple preset execution branches include at least the following task list reconstruction branch: when the intervention trigger reason is a branch where the task list is empty, the structured guidance prompt will first guide the large language model to call the preset task list reconstruction tool; after the task list reconstruction tool is called, the task list will be regenerated according to the current user request, and the reconstructed task list will be written into the runtime state as the basis for subsequent completion determination.

10. The agent response rewriting and reflection intervention method based on task completion monitoring according to claim 1, characterized in that, In step S4, the preset structured guidance prompts returned by the reflection tool are predefined text, which includes guidance content for the following four branches: The first branch, when all tasks have been completed, guides the large language model to output a confirmation message that all tasks have been completed; The second branch guides the large language model to output a request for further information when the user needs to provide additional information. The third branch guides the large language model to continue executing unfinished tasks when there are still unfinished tasks to be done. The fourth branch, when the task list is empty, guides the large language model to first call the task list reconstruction tool to create the task list before continuing execution.

Citation Information

Patent Citations

  • Intelligent agent strengthening method

    CN121659990A