IT system fault handling method based on large model agent ReACt circulation

By introducing a human-computer collaboration ReACt loop in IT system fault handling, and using the collaborative work of the m module and the h module, the problem of insufficient efficiency and accuracy caused by relying solely on the LLM agent is solved, and efficient and accurate fault response is achieved.

CN120371577APending Publication Date: 2025-07-25NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510290909.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In IT system failure handling, it is difficult to achieve full automation by relying solely on large language model (LLM) agents, resulting in insufficient fault response efficiency and accuracy, and a combination of human-machine methods is required to improve efficiency and accuracy.

Method used

The human is virtualized into an H module, and works in conjunction with the machine software module m module, and troubleshooting is handled through the reason and act steps of the ReACt loop, and the m reflection and h reflection mechanism are introduced, the degree of automation is flexibly adjusted, and the h module is introduced to complete the task only when necessary.

Benefits of technology

It improves the efficiency and accuracy of fault handling, reduces labor costs, adapts to the fault handling needs of different scenarios, gives full play to the advantages of LLM agents, and uses human experience and judgment capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371577A_ABST
    Figure CN120371577A_ABST
Patent Text Reader

Abstract

The invention discloses an IT system fault handling method based on large model agent ReACt circulation. The IT system fault handling method comprises the steps that (1) machine software modules in an IT system are marked as m modules; creating a virtual module for simulating interaction between a human and the m module; recording the virtual module as an h module; (2) when a fault event occurs in the IT system, four links of confirmation, troubleshooting, disposal and redisk are executed in sequence, and a plurality of ReACt cycles are executed in each link; the ReACt cycle of each round comprises two types of steps, namely a reaction step and an act step; m reflection and h reflection are introduced in the reaction step, m reflection refers to reflection performed by an m module, and h reflection refers to reflection performed by an h module; and then calling the act step to execute a tool determined by the reflection result of the react step. According to the invention, the efficiency and accuracy of fault response are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology and relates to an IT system fault handling method based on a large model intelligent agent ReACt cycle. Background Art

[0002] With the development of artificial intelligence technology, the application of large language model (LLM) agents in the field of IT operation and maintenance is becoming more and more widespread. LLM agents can use various tool modules, ReACT (Reason-ACT, thinking-action cycle) framework and FC (Function Calling, function call reasoning) mechanism to complete task planning and tool calling. The relevant cases are as follows:

[0003] 1. BabyAGI

[0004] BabyAGI is an AI-based task management system that uses OpenAI and Pinecone API to create, prioritize, and execute tasks. Here’s how it works:

[0005] Task generation Agent: Generates formatted tasks through templated prompt word engineering.

[0006] Execute Agent: Execute the defined function.

[0007] Vector database retrieval agent: returns RAG (retrieve-augmented generation) results to provide context information for task execution.

[0008] Prioritization Agent: Sends the task-ID mapping table to LLM, and LLM returns the re-ordered task ID sequence after making a decision.

[0009] Loop execution: extract the first unfinished task from the task list, send it to the execution agent to complete the task, organize the results and store them in the vector database, create a new task based on the goal and the result of the previous task, reorder the task list, and then repeat the above process.

[0010] 2. HuggingGPT

[0011] HuggingGPT uses ChatGPT as a task planner, selects a model on the HuggingFace platform based on the task description, executes the results and generates a response. The specific process is:

[0012] The user enters a task description.

[0013] ChatGPT, as a task planner, analyzes task requirements and determines the models on the HuggingFace platform that need to be invoked.

[0014] Invoke the corresponding model to execute the task and obtain the results.

[0015] ChatGPT generates the final response based on the results of the model execution.

[0016] 3. API-Bank

[0017] The API-Bank benchmark is designed to evaluate the ability of large language models to use tools in real-world applications. By simulating real scenarios, the model is required to make decisions at multiple levels, including whether to call an API, select the correct API, process the results returned by the API, and plan complex tasks. For example:

[0018] Call an API: Given a description of an API, the model needs to determine whether to call the given API, call it correctly, and make an appropriate response to the API return.

[0019] Retrieve an API: The model needs to search for APIs that may solve the user's needs and learn how to use them by reading the documentation.

[0020] Plan API calls: Faced with ambiguous user requests (such as arranging a group meeting, booking flights / hotels / restaurants for a trip), the model may need to make multiple API calls to solve the problem.

[0021] 4. Agent built with Langchain

[0022] A simple Agent can be built through Langchain to call the tool module to complete tasks. For example:

[0023] Environment configuration: Install the langchain library, set the OpenAI API key, and initialize the model.

[0024] Create a tool module: Create a tool multiply that can calculate the product of two integers.

[0025] Task execution: When the user inputs a task to calculate the product of two integers, the Agent calls the multiply tool to complete the calculation and output the result.

[0026] These cases demonstrate that large language model LLM agents, with the help of the ReACT framework and the FC mechanism, complete complex tasks through task planning and tool invocation.

[0027] However, in the actual process of IT system fault handling, it is difficult to achieve full automation relying solely on LLM agents, and a human-machine combination approach is still needed to improve the efficiency and accuracy of fault response. Summary of the Invention

[0028] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide an IT system fault handling method based on the ReACt cycle of large model agents, aiming to improve the efficiency and accuracy of fault response through human-machine collaboration. The present invention virtualizes a human (human) as a module (h module) and collaborates with a traditional machine software module (m module) to jointly complete the fault handling task.

[0029] The human-machine combination method for IT system fault handling based on the ReACt cycle of large model agents proposed by the present invention realizes an efficient and accurate fault handling process by virtualizing a human as an h module and collaborating with an m module. While ensuring the degree of automation, this method effectively solves the problems that may be brought about by simply relying on LLM agents through a flexible human-machine collaboration mechanism, providing a new solution for the intelligentization in the field of IT operation and maintenance.

[0030] The technical solution of the present invention is as follows:

[0031] An IT system fault handling method based on the ReACt cycle of large model agents, the steps of which include:

[0032] 1) Denote the machine software module in the IT system as the m module; create a virtual module for simulating the interaction between a human and the m module; denote the virtual module as the h module;

[0033] 2) When a fault event occurs in the IT system, successively execute four links: confirmation, troubleshooting, handling, and review, and perform several rounds of ReACt cycles for each link; each round of ReACt cycle includes two types of steps: reason step and act step;

[0034] Introduce m reflection and h reflection in the reason step, where m reflection refers to the reflection by the m module, and h reflection refers to the reflection by the h module; then call the act step to execute the tool determined by the reflection result of the reason step.

[0035] Furthermore, the act step includes multiple act type tasks, and different act type tasks are completed by different act type modules; if an act type task cannot be completed by the m module, then introduce the h module to complete this act type task.

[0036] Furthermore, the formal expression of the act step is: a ← λa m +(1 - λ)a h; where a m is the result output by the act type task completed by the m module, and a h is the result output by the act type task completed by the h module. λ represents the degree of automation, and the value of λ is determined by the reason step.

[0037] Furthermore, in the reason step, serial execution of m reflection and h reflection is introduced: r ← H(M(a)); H(M(a)) means that first the m module analyzes and reflects on the result a to obtain M(a), and then the h module analyzes and reflects on M(a) to obtain the result r. r represents the reflection result of the reason step, and a represents the output result of the previous act step.

[0038] Furthermore, in the reason step, parallel execution of m reflection and h reflection is introduced: r ← H(a) + M(a); H(a) + M(a) means that the h module and the m module simultaneously analyze and reflect on the result a to obtain the corresponding reflection results H(a) and M(a), and then the reflection results are aggregated to obtain the reflection result r. r represents the reflection result of the reason step, and a represents the output result of the previous act step.

[0039] Furthermore, in the confirmation step: First, the m module is called through the reason step to perform serial execution of m reflection and h reflection, and an alarm confirmation scheme is formed based on the fault context information.

[0040] Furthermore, in the troubleshooting step: First, the m module is called through the reason step to perform parallel execution of m reflection and h reflection, and an alarm confirmation scheme is formed based on the fault context information.

[0041] Furthermore, the m module includes a large model.

[0042] A server, characterized in that it includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the above method.

[0043] A computer-readable storage medium, on which a computer program is stored, characterized in that the computer program implements the above method when executed by a processor.

[0044] The advantages of the present invention are as follows:

[0045] 1. Improve the efficiency of fault handling: Through human-machine collaboration, give full play to the advantages of the LLM agent, and at the same time utilize human experience and judgment ability.

[0046] 2. Enhance the accuracy of handling: Introduce the h reflection mechanism to reduce possible errors generated by the LLM.

[0047] 3. Reduced labor costs: The h module is introduced only when necessary to maximize the degree of automation.

[0048] 4. Flexibly adapt to different scenarios: By adjusting the reflection method and the degree of automation, it can adapt to different fault handling links. Description of the Drawings

[0049] Figure 1 It is the flowchart of the method of the present invention.

[0050] Figure 2 It is the confirmation flowchart.

[0051] Figure 3 It is the troubleshooting flowchart. Detailed Implementation Manner

[0052] The present invention will be further described in detail below with reference to the drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0053] The present invention includes the following aspects:

[0054] I. Define two types of modules:

[0055] h module: A virtual module representing a person; when handling a fault, the h module simulates a person to interact with the IT system through traditional operation and maintenance tools.

[0056] m module: A machine software module, including a large model.

[0057] II. ReACt loop

[0058] The fault handling workflow is divided into four links: confirmation, troubleshooting, handling, and review. Each link is carried out through one or more rounds of ReAct loops. ReAct gives an idea, and the present invention designs a ReAct loop according to the idea provided by ReAct. Each ReAct loop contains two types of steps: reason (r) and act (a).

[0059] III. Human - machine collaboration mechanism

[0060] 1. reason step: Introduce the m - reflection and h - reflection mechanisms. Among them, m - reflection refers to the reflection carried out by the m module, and h - reflection refers to the reflection carried out by the h module. r represents the reflection result of the reason step (i.e., the r step), and a represents the output result of the previous act step (i.e., the a step); reflection is to observe and analyze the result a of the previous action. The reason (r) step can be carried out in a serial or parallel manner:

[0061] – Serial reflection: r←H(M(a)); H(M(a)) means that first, the m module conducts an analysis and reflection, and then the h module further analyzes and reflects.

[0062] – Parallel reflection: r ← H(a) + M(a); H(a) + M(a) means that the h module and the m module analyze simultaneously and then summarize the results; there is no restriction on the summarization process, which can be addition or other methods.

[0063] 2. act step: The reason step will determine which modules to call to execute the act step. The act step is divided into multiple tasks, and different tasks are completed by different act class modules (i.e., the shells in the tool wrapping mechanism), and different act class modules complete different act class tasks. Some tasks cannot be completed by the m module, and the h module needs to be introduced to make up for the lack of capabilities of the m module to jointly complete the act step. This phenomenon can be formally expressed as: a ← λa m +(1 - λ)a h where a m is the result output by the act class task completed by the m module, a h is the result output by the act class task completed by the h module, and λ represents the degree of automation, which is determined by the reason link: λ ← λ(r). That is to say, the output of the reason link will include λ.

[0064] 3. Function Calling improvement: Introduce the hFC mechanism. When the confidence in tool selection is lower than the threshold, use the h module as the selected tool. In the FC mechanism, the m module (such as a large model) selects external tools through planning and reasoning (reason) and then calls the selected tools. In the present invention, tools generally refer to various tools used for alarm confirmation, troubleshooting, and handling in the actual operation and maintenance process. The m module itself does not have the ability to interact with the link and needs to call the selected tool to complete the a step.

[0065] IV. Workflow (as Figure 1 shown)

[0066] Initialization: The occurrence event of a fault triggers the workflow of this system. When this system starts, the initial a step is executed: obtain the fault context information. Usually, this information is provided by an external system when triggering this system.

[0067] The confirmation link is as Figure 2 shown:

[0068] 1. Read the alarm context and set it as the initial value of a, $a\gets a_0$;

[0069] 2. r step: Form an alarm confirmation plan, using serial reflection $r\gets H(M(a))$;

[0070] 3. a step: Execute alarm confirmation $a\gets a(r)$;

[0071] 4. r step: Confirm that the alarm confirmation is successful, and use serial reflection: $r\gets H(M(a))$;

[0072] 5. If successful, exit. If the loop threshold is exceeded, also exit; otherwise, go to step 2.

[0073] The troubleshooting process is as Figure 3 shown:

[0074] r step: Formulate the solution for this step, and use serial reflection $r←H(M(a))$, where the value of a is inherited from the previous step;

[0075] a step: Execute the solution for this step and obtain the context information $a←a(r)$;

[0076] r step: Analyze the execution result of the solution for this step and determine whether it has been successfully completed, using parallel reflection: $r←H(a)+M(a)$;

[0077] If successful, exit. If the loop threshold is exceeded, also exit; otherwise, go to step 1.

[0078] The process of the handling and review phase is the same as that of the troubleshooting phase, except for the different task objectives stated in the prompt.

[0079] V. System Architecture

[0080] The system architecture of the present invention includes the following main components: 1. LLM agent: As the core m module, it is responsible for the main execution of the reason and act steps. 2. Tool module library: Contains various IT operation and maintenance tools for the LLM agent to call. 3. Human-computer interaction interface: Used for the h module (human) to participate in the workflow. 4. Workflow engine: Coordinates the execution of each step and manages the ReACt loop. 5. Data storage: Stores fault information, handling records, etc.

[0081] 1. ReACt Loop Implementation

[0082] The execution process of each ReACt loop is as follows:

[0083] 1) reason step:

[0084]

[0085] 2) act step:

[0086] 3) hFC implementation:

[0087]

[0088] 2. Workflow implementation

[0089] Take the troubleshooting phase as an example:

[0090]

[0091]

[0092] Example 1: Disposal of network device failures

[0093] Suppose the core switch in a data center fails, causing some services to be interrupted. The system automatically creates a trouble ticket and starts the trouble - shooting workflow.

[0094] 1. Confirmation phase:

[0095] –r step: The LLM proposes a ping - test plan

[0096] –h reflection: The operation and maintenance personnel confirm that the plan is feasible

[0097] –a step: Execute the ping - test and confirm the existence of the failure

[0098] 2. Troubleshooting phase:

[0099] –r step: The LLM suggests checking the switch logs

[0100] –a step: Automatically obtain the switch logs

[0101] –r step: The LLM analyzes the logs and suspects a configuration error

[0102] –h reflection: The operation and maintenance personnel agree with this judgment

[0103] 3. Disposal phase:

[0104] –r step: The LLM proposes a configuration correction plan

[0105] –h reflection: The operation and maintenance personnel review and approve the plan

[0106] –a step: Since it involves core equipment, introduce the h module to execute the configuration modification

[0107] 4. Review phase:

[0108] –r step: The LLM generates a preliminary review report

[0109] –h reflection: The operation and maintenance personnel supplement the experience summary

[0110] –a step: Automatically integrate the report and archive it

[0111] Through this human-machine combination method, the fast analysis ability of the LLM is fully utilized, and human supervision is ensured, improving the efficiency and accuracy of fault handling.

[0112] Regarding the h module and the m module:

[0113] The m module represents the original automation module, which is divided into four categories: confirmation, troubleshooting, handling, and review. Each major category is further divided into two sub-categories: reason and act. Each sub-category contains multiple different specific m module instances (i.e., calling specific tools, which are tool call codes) for performing different work tasks. For example, in the troubleshooting category, different m modules are responsible for performing different troubleshooting tasks. For example, the m1 module is responsible for reading and writing the operation and maintenance interface A, the m2 module is responsible for accessing the operation and maintenance interface B, the m3 module is responsible for remotely logging in to certain servers and executing commands, and the m4 is responsible for reading certain database tables, etc.

[0114] The characteristics of the M module and the h module are that there is a one-to-one correspondence between the m module and the h module, that is, a dual relationship. A pair of dual h modules and m modules follow the same interface, interface, and context interaction relationship, and the difference lies only in the different internal implementation mechanisms. This is a principle of separating the interface from the implementation.

[0115] The characteristic of the m module is pure automation implementation, and its specific implementation may adopt expert systems, machine learning algorithms, large language models, etc. The characteristic of the h module is that it must be executed with the help of humans.

[0116] For example, a pair of dual m modules and h modules are responsible for remotely logging in to a server. The m module can automatically input the login security credentials and complete the login. The h module then pops up a prompt to the human user (operation and maintenance personnel), asking the latter to input the login password, and then completes the login.

[0117] When the m module fails to execute, the dual h module can be automatically called for execution. For example, an m module is responsible for parsing the system error log to determine the cause of the error. If this module fails to find the cause of the error from the error log, the corresponding h module is called. The h module asks for help from the human expert user. After the expert inputs the cause of the error, the h module returns control to the m module, and the m module continues to execute the subsequent business logic.

[0118] Although specific embodiments of the present invention are disclosed for illustrative purposes, the purpose is to help understand the content of the present invention and implement it accordingly. Those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection required by the present invention is defined by the scope defined in the claims.

Claims

1. An IT system fault handling method based on the ReACt loop of large model agents, the steps of which include: 1) Denote the machine software modules in the IT system as m modules; Create a virtual module for simulating the interaction between humans and the m modules; Denote the virtual module as h module; 2) When a fault event occurs in the IT system, four links of confirmation, troubleshooting, handling, and review are sequentially executed, and several rounds of ReACt loops are executed in each link; each round of the ReACt loop includes two types of steps: the reason step and the act step; m reflection and h reflection are introduced in the reason step, where m reflection refers to the reflection by the m module, and h reflection refers to the reflection by the h module; then call the act step to execute the tool determined by the reflection result of the reason step.

2. The method according to claim 1, wherein The act step includes multiple act type tasks, and different act type tasks are completed by different act type modules; if an act type task cannot be completed by the m module, the h module is introduced to complete the act type task.

3. The method according to claim 2, wherein The formal expression of the act step is: a ← λa m +(1 - λ)a h ; Among them, a m is the result output by the act class task completed by the m module, a h is the result output by the act class task completed by the h module, λ represents the degree of automation, and the value of λ is determined by the reason link.

4. The method according to claim 1 or 2 or 3, characterized in that Serial execution of m reflection and h reflection is introduced in the reason step: r←H(M(a)); H(M(a)) means that first the m module analyzes and reflects on the result a to obtain M(a), and then the h module analyzes and reflects on M(a) to obtain the result r, r represents the reflection result of the reason step, and a represents the output result of the previous act step.

5. The method according to claim 1 or 2 or 3, characterized in that Parallel execution of m reflection and h reflection is introduced in the reason step: r←H(a)+M(a); H(a)+M(a) means that the h module and the m module simultaneously analyze and reflect on the result a to obtain the corresponding reflection results H(a) and M(a), and then summarize the reflection results to obtain the reflection result r, r represents the reflection result of the reason step, and a represents the output result of the previous act step.

6. The method according to claim 3, wherein In the confirmation link: First, call the m module through the reason step to adopt serial execution of m reflection and h reflection, and form an alarm confirmation plan according to the fault context information.

7. The method according to claim 5, wherein In the troubleshooting link: First, call the m module through the reason step to adopt parallel execution of m reflection and h reflection, and form an alarm confirmation plan according to the fault context information.

8. The method according to claim 1 or 2 or 3, characterized in that The m module includes a large model.

9. A server, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing any one of the methods according to claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements any one of the methods according to claims 1 to 8.

Citation Information

Cited By

  • Multi-agent power grid project intelligent monitoring, control and evaluation system and method

    CN121146283A