Operation and maintenance root cause analysis system and method, medium and equipment

Through the operation and maintenance root cause analysis system, using large language models and AI-Agent automated operation and maintenance troubleshooting, the problem of traditional operation and maintenance troubleshooting is solved, and the effect of rapid fault location and improvement of operation and maintenance efficiency is achieved.

CN120179443APending Publication Date: 2025-06-20CHANJET INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510135201.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The traditional operation and maintenance troubleshooting process takes a long time, relies on expert experience, long analysis chains and dispersed tools, resulting in low fault positioning efficiency and poor accuracy.

Method used

Design an operation and maintenance root cause analysis system, combines a large language model and AI-Agent, trains a large language model through the prompt word module, and calls multiple data analysis tools to realize automated operation and maintenance root cause analysis.

Benefits of technology

Quickly locate faults, shorten troubleshooting time, improve operation and maintenance efficiency, reduce manual intervention, promote the solidification and transmission of expert experience, and improve the accuracy of troubleshooting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179443A_ABST
    Figure CN120179443A_ABST
Patent Text Reader

Abstract

The invention provides an operation and maintenance root cause analysis system and method, a medium and equipment, and relates to the field of information technology operation and maintenance. The system comprises a cue word module which comprises task description, analysis logic, result constraint and behavior limitation; the training module is used for AI-Agent to train a large language model through the cue word module; the tool set comprises a plurality of data analysis tools including domain name monitoring, business change checking, server resource bottleneck detection and a log query API (Application Program Interface), and the API is registered in the AI-Agent; the monitoring module is used for monitoring the running state of the IT system; the analysis module is used for receiving the alarm information sent by the monitoring module, calling the tool set through the large language model and carrying out operation and maintenance root cause analysis; and the report distribution module receives the operation and maintenance root cause analysis result and distributes the result through a work order system or an instant messaging tool. According to the invention, the efficiency and accuracy of troubleshooting can be improved, manual intervention is reduced, and the solidification and transmission of expert experience are promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology operation and maintenance, and particularly relates to an operation and maintenance root cause analysis system, method, medium and device. Background Art

[0002] In the traditional fault troubleshooting process, operation and maintenance personnel need to manually check a series of dashboards, execute a series of data analysis scripts or query relevant logs, which is a relatively time-consuming process. There are many factors affecting the length of this process. It requires both rich experience and excellent professional knowledge, and at the same time, it requires a stress-resistant mentality, because the situation when a fault occurs is extremely urgent, and the inability to quickly find the cause and stop the loss will have a very large impact on online customers.

[0003] The defects and deficiencies of traditional technical means are as follows:

[0004] 1) When a fault occurs, the analysis chain for locating the problem is long, resulting in the problem being difficult to be quickly located, which is unacceptable to users;

[0005] 2) Problem location depends on expert experience, and the troubleshooting time for the same problem may vary for different people;

[0006] 3) Expert experience is difficult to be solidified and effectively transmitted;

[0007] 4) Various dashboards, tools and materials are scattered everywhere, making it inconvenient to consult. Summary of the Invention

[0008] Therefore, the technical problem to be solved by the present invention is to provide an operation and maintenance root cause analysis system, method, medium and device, which can improve the efficiency and accuracy of fault troubleshooting, reduce manual intervention, and promote the solidification and transmission of expert experience.

[0009] In a first aspect, the present invention discloses an operation and maintenance root cause analysis system, including:

[0010] A prompt word module, including task description, analysis logic, result constraint and behavior restriction, and the prompt word is structured text.

[0011] A training module for training a large language model by an AI-Agent through the prompt word module;

[0012] A tool set, including multiple data analysis tools for the large language model to call. The data analysis tools include domain name monitoring, business change inspection, server resource bottleneck detection and log query API, and the API is registered in the AI-Agent;

[0013] A monitoring module for monitoring the running state of the IT system, and when an alarm occurs, sending the alarm information to the analysis module;

[0014] An analysis module, configured to receive the alarm information sent by the monitoring module, call the tool set through the large language model, perform operation and maintenance root cause analysis, and obtain the operation and maintenance root cause analysis result;

[0015] A report distribution module, which receives the operation and maintenance root cause analysis result sent by the analysis module and distributes it through a work order system or an instant messaging tool in a set format.

[0016] Further, the prompt word module specifically includes:

[0017] A task description, which tells the large language model that the task is to perform alarm analysis, respond according to the expected analysis result, describe the form of the input alarm information, and the meaning of each field in the alarm information;

[0018] Analysis logic, which classifies according to different alarm information and business types, and executes corresponding analysis steps for each type of alarm according to the corresponding business type;

[0019] Result constraint, which constraints the operation and maintenance root cause analysis result, including the form, content and meaning of the operation and maintenance root cause analysis result;

[0020] Behavior restriction, which is used to restrict the behavior of the large language model when it starts to have hallucinations.

[0021] Further, in the prompt word module, the step of executing corresponding analysis steps for each type of alarm according to the corresponding business type is to execute corresponding analysis steps for domain name monitoring alarms, specifically including:

[0022] 1) Check whether there are any changes in the online business in a recent period of time;

[0023] 2) Check whether the errors or increased response time of the domain name are concentrated in a single user or interface;

[0024] 3) Check whether the anomalies of the domain name are concentrated in an upstream node of a single backend;

[0025] 4) Check whether there are resource bottlenecks in this upstream node;

[0026] 5) Check whether there are other alarms in this business line for a period of time.

[0027] Further, the form of the operation and maintenance root cause analysis result of the analysis module is JSON structured data.

[0028] Further, the report distribution module has a built-in message template for formatting the distributed content, and the message template supports the markdown structure; the json-structured data in the operation and maintenance root cause analysis result includes a key-value pair, which is a markdown string, and the report distribution module distributes the markdown string.

[0029] In a second aspect, the present invention also provides an operation and maintenance root cause analysis method, and the analysis method includes:

[0030] Configuring prompt words, including task description, analysis logic, result constraints, and behavior restrictions, and the prompt words are structured text;

[0031] Model training, where the Agent trains a large language model through the prompt words;

[0032] Deploying a toolset, including multiple data analysis tools for the large language model to call. The data analysis tools include domain name monitoring, business change inspection, server resource bottleneck detection, and log query API, and the API is registered in the AI-Agent;

[0033] System monitoring, which is used to monitor the running status of the IT system, and when an alarm occurs, sends the alarm information to the large language model;

[0034] Operation and maintenance root cause analysis, which is used to receive the alarm information sent by the system monitoring, call the toolset through the large language model, and perform operation and maintenance root cause analysis to obtain the operation and maintenance root cause analysis result;

[0035] Report distribution, which receives the operation and maintenance root cause analysis result and distributes it through a work order system or an instant messaging tool in a set format.

[0036] In a third aspect, the present invention also discloses a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the foregoing operation and maintenance root cause analysis method is implemented.

[0037] In a fourth aspect, the present invention also discloses an electronic device, including a memory, a processor, and a computer program. The computer program is stored on the memory and can run on the processor, and when the processor executes the computer program, the foregoing operation and maintenance root cause analysis method is implemented.

[0038] Beneficial effects:

[0039] The operation and maintenance root cause analysis system, method, medium, and electronic device disclosed by the present invention, by combining the logical reasoning ability of the large language model and the toolset orchestration ability of the AI-Agent, achieve the following technical effects:

[0040] Quick positioning: Shorten the troubleshooting time and improve the operation and maintenance efficiency;

[0041] Knowledge transfer and solidification: Through prompt learning, effectively transfer and solidify expert experience;

[0042] Reduce labor costs: Reduce the manual operations of operation and maintenance personnel and lower labor costs;

[0043] Improve accuracy: Reduce human errors and improve the accuracy of troubleshooting;

[0044] Scalability: The modular design of prompt words makes the system easy to expand and adapt to new business scenarios and analysis steps. Brief Description of the Drawings

[0045] Figure 1 It is a schematic diagram of the interrelationships of each module of the operation and maintenance root cause analysis system according to Embodiment 1 of the present invention;

[0046] Figure 2 It is the scheme architecture diagram of Embodiment 1 of the present invention;

[0047] Figure 3 It is the test effect of the page in the specific embodiment;

[0048] Figure 4 It is the markdown result of the DingTalk group in the specific embodiment;

[0049] Figure 5 It is the alarm analysis result of the sudden increase in the status code of domain name 302 in the specific embodiment;

[0050] Figure 6 It is the application display result of the specific embodiment;

[0051] Figure 7 It is the flowchart of the operation and maintenance root cause analysis method according to Embodiment 2 of the present invention. Detailed Embodiment

[0052] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. The principles and features of the present invention are described below. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The embodiments cited are only used to explain the present invention and are not used to limit the scope of the present invention.

[0053] The traditional IT operation and maintenance troubleshooting process is cumbersome and time-consuming, relying on the experience, professional knowledge and stress resistance of operation and maintenance personnel. Existing technical means have problems such as long analysis chains, reliance on expert experience, difficulty in solidifying and transferring knowledge, and scattered tools.

[0054] Large language model technology is good at processing unstructured text and logical reasoning. However, at the same time, in the context of existing technologies, large models still have many limitations. For example, they are not good at big data operations, there are limitations on the number of tokens during the conversation process, etc. Of course, we can train on the basis of general large models to solve the problem of insufficient tokens carried during the conversation to a certain extent, but this means high costs and is not a general solution. The AI-Agent solution can, on the one hand, provide a toolset for the large model (such as a calculator, etc., to make up for its computing power), enabling the large model to focus on understanding the business and logical reasoning, and thus performing the orchestration and scheduling of tools to achieve true usability. On the other hand, within the acceptable range of existing tokens, the experience and knowledge of the business can be provided to the large model in the form of prompts, enabling it to have knowledge in the professional field and reducing the possibility of making mistakes and having hallucinations.

[0055] The present invention is implemented through large model technology and AI-Agent technology. After the Agent learns the root cause analysis experience of operation and maintenance experts through prompts, it automatically selects appropriate root cause analysis steps through reasoning, calls relevant analysis tools, and finally integrates and outputs the results, improving the efficiency of operation and maintenance fault troubleshooting and reducing risks.

[0056] Embodiment 1

[0057] As Figure 1 shown, this embodiment is an operation and maintenance root cause analysis system, including:

[0058] The prompt module, including task description, analysis logic, result constraint, and behavior restriction.

[0059] The task description describes that the task of the large language model is to perform alarm analysis and respond according to the expected analysis results, describes the form of the input alarm information, and the meanings of each field in the alarm information;

[0060] The analysis logic classifies according to different alarm information and business types, and executes corresponding analysis steps for each type of alarm according to the corresponding business type;

[0061] The result constraint restricts the operation and maintenance root cause analysis results, including the form, content, and meaning of the operation and maintenance root cause analysis results;

[0062] The behavior restriction is used to restrict the behavior of the large language model when it starts to have hallucinations or various situations.

[0063] Figure 2This is the schematic diagram of the solution architecture of the embodiments of the present invention. Starting from the analysis of a certain business line, through research, it is found that the vast majority of alarms in this business line are caused by individual user access. Some are due to the overly large volume of the accounting set, some are malicious attacks or interface brushing, corresponding to different loss prevention operations. When the domain name monitoring of this business alarms, the analysis process of the operation and maintenance experts is divided into several steps:

[0064] 1. Check whether there are any changes online within the last 30 minutes;

[0065] 2. Check whether the increase in domain name errors or response time is concentrated in a single user or interface;

[0066] 3. Check whether the anomalies of the domain name are concentrated in an upstream node of a single backend;

[0067] 4. Check whether there are resource bottlenecks in this upstream node;

[0068] 5. Check whether there are other alarms in this business line within 10 minutes.

[0069] After that, start to compile the prompt words to let the AI troubleshoot problems according to the expert experience.

[0070] The composition of the prompt words is also structured and is divided into the following parts:

[0071] The first part is to let the AI understand what it is going to do and what its input is. We need to inform the AI that the task to be executed is alarm analysis and respond according to the expected results. Here, it is necessary to tell in what form the alarm will be provided and what the meanings of the fields are.

[0072] The second part is to let the AI know how to do it and what the logic is. Classify according to different alarms and business types. Each type of alarm has a different analysis method. We first define that for the above-mentioned business line, the five analysis steps should be executed successively.

[0073] The third part is to impose constraints on the return results of the AI, including what content must be included and what the meanings are.

[0074] The fourth part is the restrictions on the behavior of the AI. When the AI starts to have hallucinations or various situations, it can be remedied here and gradually improved.

[0075] Because the AI always shows its thinking and reasoning process, sometimes it even directly returns in English, and sometimes it beautifies the returned JSON string by itself, resulting in the backend being unable to recognize it. Therefore, some behavioral constraints need to be imposed on the AI.

[0076] The benefits of modularizing the prompt words include, but are not limited to:

[0077] 1. It has extremely strong scalability. When there are new scenarios or new analysis steps, they can be directly inserted at the corresponding positions.

[0078] 2. Highly structured text is more understandable for both humans and AI and is less likely to suffer from logical confusion.

[0079] A training module for the Agent to train a large language model through the prompt word module. An Agent is a computer program or entity that can act autonomously in a specific environment, perceive the environment, make decisions, and interact with other Agents or humans. They possess characteristics such as autonomy, reactivity, sociality, and adaptability, and can adjust their behavior according to environmental changes to achieve preset goals. The LLM large model is an important tool for the Agent to perform task planning and knowledge reasoning. It has powerful language processing and knowledge reasoning capabilities through learning a large amount of text data.

[0080] A toolset including multiple data analysis tools for the large language model to call. Since general large models do not possess business knowledge and are also difficult to perform big data computing tasks, it is necessary to provide single-point data analysis tools for the large model to call, such as checking business changes and viewing whether there are bottlenecks in server resources. By registering the APIs of these tools in the Agent, the large model can "understand" the functions and call methods of these APIs.

[0081] A monitoring module for monitoring the operating status of the IT system. When an alarm occurs, it sends the alarm information to the large language model; when an alarm occurs, the monitoring system directly sends the alarm to the AI. After the AI completes the root cause analysis, it sends the results to the report distribution module.

[0082] An analysis module for receiving the alarm information sent by the monitoring module, calling the toolset through the large language model, and performing operation and maintenance root cause analysis to obtain the operation and maintenance root cause analysis results.

[0083] A report distribution module that receives the operation and maintenance root cause analysis results sent by the analysis module and distributes them through the work order system or instant messaging tool in a set format.

[0084] In a preferred embodiment, the prompt word module requires that AI must return a json structure. The report distribution module completes the distribution of work orders and DingTalk groups. The report distribution module is directly connected to the AI, which means that the AI's response must be highly structured data (json) and cannot be wrong. On the other hand, the content sent to the DingTalk group will give people a very bad impression if it is not organized. This means that the report distribution module needs to have a built-in message template (DingTalk supports markdown structure) to structure the messages sent to DingTalk, which will cause us to need to modify this part of the code for subsequent changes in the message content. So we added a kv to the returned json key-value pair, which is the expected markdown string, so that the report distribution module can directly return this string to DingTalk.

[0085] After the prompt word configuration is completed, the previously developed tools are also configured, and then you can test it. The test effect of the page is shown in Figure 3 .

[0086] In less than 20 seconds, AI completed four tool calls and returned the results in the expected format. Manual verification found that the analysis process was in line with expectations and the correct conclusion was obtained. The markdown results of the DingTalk group can be seen in Figure 4 .

[0087] Figure 5 For the alarm analysis result of a sudden increase in the 302 status code of a domain name, AI used four tools in succession through logical reasoning and finally determined that the root cause was the self-healing of the memory of a backend server. The result was in line with the expectations of the operation and maintenance experts.

[0088] like Figure 6 As shown, this system has been called more than 20,000 times in a month, and has realized the analysis of all online alarms.

[0089] When an alarm occurs, AI will analyze the alarm based on manual experience and generate an analysis report based on the analysis results. The whole process takes about 20 seconds. Basically, when the operation and maintenance personnel receive the alarm notification, AI has completed the analysis and pointed out the content that may need to be paid attention to in the next step, or even the solution. In the past, for experienced operation and maintenance experts, this action required opening the relevant dashboards, log systems, etc. one by one, which took several minutes. For inexperienced oncall students, it may take longer or even they don’t know how to analyze. Therefore, AI not only realizes the automated analysis of alarms, but also allows operation and maintenance experts to instill alarm analysis experience into AI through natural language, completing the effective transmission and transfer of knowledge.

[0090] Before the system was officially launched, multiple rounds of tests with manual questions were conducted. The results showed that for different scenarios, the analysis process of the AI was 100% in line with our expectations.

[0091] After the official launch, all online alarm AIs will first conduct an analysis. For the scenarios that we have fully defined (i.e., the scenarios where expert experience is incorporated), the AI can analyze the root cause and achieve a closed-loop. For the other undefined scenarios, the AI can also perform basic analysis operations, which are inevitable when manually analyzing without AI. Therefore, the efficiency improvement brought by this scenario is very obvious.

[0092] As of now, 30% of the online alarm AIs can be fully automatically processed in a closed-loop, and the recall rate of the root cause analysis of these alarms is 100%.

[0093] The AI is like a student with excellent computer knowledge. When we continuously pass on experience to it and provide corresponding tools, it can quickly and tirelessly help us complete the work.

[0094] Embodiment 2

[0095] The present invention also provides an operation and maintenance root cause analysis method, as Figure 7 shown, the analysis method includes:

[0096] Configure prompt words, including task description, analysis logic, result constraints, and behavior restrictions. The prompt words are structured text;

[0097] Model training, where the AI-Agent trains the large language model through the prompt words;

[0098] Deploy a tool set, including multiple data analysis tools for the large language model to call. The data analysis tools include domain name monitoring, business change inspection, server resource bottleneck detection, and log query API. The API is registered in the AI-Agent;

[0099] System monitoring, which is used to monitor the running status of the IT system. When an alarm occurs, the alarm information is sent to the large language model;

[0100] Operation and maintenance root cause analysis, which is used to receive the alarm information sent by the monitoring module, call the tool set through the large language model, and conduct operation and maintenance root cause analysis to obtain the operation and maintenance root cause analysis result;

[0101] Report distribution, which receives the operation and maintenance root cause analysis result and distributes it in a set format through the work order system or instant messaging tool.

[0102] Embodiment 3

[0103] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the above-mentioned operation and maintenance root cause analysis method.

[0104] Embodiment 4

[0105] The present invention also provides an electronic device, including a memory, a processor, and a computer program. The computer program is stored on the memory and can run on the processor. When the processor executes the computer program, the above-mentioned operation and maintenance root cause analysis method is implemented.

[0106] Obviously, the above embodiments are merely examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.

Claims

1. An operation and maintenance root cause analysis system, characterized in that: include: The prompt word module includes task description, analysis logic, result constraints and behavior restrictions, and the prompt words are structured texts; A training module, used for AI-Agent to train a large language model through the prompt word module; A tool set, including multiple data analysis tools for the large language model to call, the data analysis tools include domain name monitoring, business change checking, server resource bottleneck detection and log query API, the API is registered in the AI-Agent; The monitoring module is used to monitor the operation status of the IT system and send the alarm information to the analysis module when an alarm occurs; An analysis module, used for receiving the alarm information sent by the monitoring module, calling the tool set through the large language model, performing operation and maintenance root cause analysis, and obtaining an operation and maintenance root cause analysis result; The report distribution module receives the operation and maintenance root cause analysis result sent by the analysis module, and distributes it through a work order system or an instant messaging tool in a set format.

2. The system according to claim 1, characterized in that The prompt word module specifically includes: The task description informs the large language model that the task is to perform alarm analysis and respond according to the expected analysis results, and describes the input alarm information format and the meaning of each field in the alarm information; The analysis logic is classified according to different alarm information and business types, and corresponding analysis steps are performed for each alarm according to the corresponding business type; The result constraint constrains the operation and maintenance root cause analysis result, including the form, content and meaning of the operation and maintenance root cause analysis result; The behavior restriction is used to constrain the behavior of the large language model when hallucination begins to occur.

3. The system according to claim 2, characterized in that In the prompt word module, the corresponding analysis steps are performed for each alarm according to the corresponding business type, which is the analysis steps performed for the domain name monitoring alarm, specifically including: 1) Check whether there are any changes in online business in the near future; 2) Check whether domain name errors or increased time consumption are concentrated on a single user or interface; 3) Check whether the domain name anomalies are concentrated in a backend upstream node; 4) Check whether there is a resource bottleneck on this upstream node; 5) Check whether there are other alarms for this business line over a period of time.

4. The system according to claim 1, characterized in that The operation and maintenance root cause analysis result of the analysis module is in the form of json structure data.

5. The system according to claim 4, characterized in that The report distribution module has a built-in message template for formatting the distributed content, and the message template supports the markdown structure; the json structure data in the operation and maintenance root cause analysis result includes a key-value pair, which is a markdown string, and the report distribution module distributes the markdown string.

6. A method for root cause analysis of operation and maintenance, characterized in that: The analysis method comprises: Configuration prompt words, including task description, analysis logic, result constraints and behavior restrictions, the prompt words are structured text; Model training: the agent trains a large language model using the prompt words; A deployment tool set, including multiple data analysis tools for the large language model to call, the data analysis tools include domain name monitoring, business change checking, server resource bottleneck detection, and log query API, the API is registered in the AI-Agent; System monitoring, used to monitor the operation status of the IT system and send the alarm information to the analysis module when an alarm occurs; Operation and maintenance root cause analysis, used to receive the alarm information sent by the system monitoring, call the tool set through the large language model, perform operation and maintenance root cause analysis, and obtain the operation and maintenance root cause analysis result; Report distribution: receiving the operation and maintenance root cause analysis results and distributing them in a set format through a work order system or instant messaging tool.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the operation and maintenance root cause analysis method described in claim 6 is implemented.

8. An electronic device comprising a memory, a processor and a computer program, wherein the computer program is stored in the memory and can be run on the processor, wherein: When the processor executes the computer program, the operation and maintenance root cause analysis method as described in claim 6 is implemented.

Citation Information

Patent Citations

  • Log analysis method and device, electronic equipment and storage medium

    CN118132711A

  • Log analysis positioning method and system based on large model technology

    CN118227421A

  • Operation and maintenance alarm analysis method, device and equipment based on large language model

    CN119065917A