Fault detection agent scheme based on large language model

Through the fault detection agent solution based on the large language model, using RAG technology and automation tool API, the automated detection and processing of faults of large network clusters is realized, the high cost and low efficiency problems of traditional manual dependence are solved, the operation and maintenance efficiency is improved and the potential to replace manual operation and maintenance is possessed.

CN120011545AInactive Publication Date: 2025-05-16BEIJING YUSUN NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510120595.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The fault detection and processing of large network clusters rely on manual labor, resulting in high labor costs, inefficiency and high experience dependence. Traditional AI technologies have poor generalization capabilities when facing highly domain-based knowledge and lack comprehensive judgment capabilities.

Method used

The fault detection agent scheme based on a large language model is adopted. By receiving fault descriptions or real-time alarm data, RAG technology is used to retrieve the fault knowledge base, make inference decisions, and call the automation tool API for fault analysis and processing.

Benefits of technology

The automated root cause detection of faults has been realized, which significantly reduces the time and labor cost of fault detection, improves operation and maintenance efficiency, and has the potential to replace manual operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011545A_ABST
    Figure CN120011545A_ABST
Patent Text Reader

Abstract

The invention discloses a fault detection agent scheme based on a large language model. The fault detection agent scheme comprises the following steps of receiving clues and performing knowledge retrieval; thinking and deciding; executing the operation; and comprehensive analysis: completing an agent fault detection reasoning framework based on a ReAct format, and carrying out fault analysis. According to the method, the RCAgent can automatically extract structured data from a large number of operation and maintenance documents and is used for professional knowledge base construction and training set construction; the RCAgent has an API-based tool set and a tool interaction module, and can interact with a system environment in real time to obtain more comprehensive and detailed fault information; the RCAgent takes a large language model which is subjected to customized fine tuning of a professional field training set as a core, and supports iterative multi-step reasoning oriented to a complex fault root cause detection task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent network operation and maintenance, and in particular to a fault detection intelligent agent solution based on a large language model. Background Art

[0002] Currently, traditional operation and maintenance of large network clusters mainly relies on manual monitoring and processing. This model has huge challenges:

[0003] (1) High labor costs: Large network clusters contain massive nodes and complex business scenarios, requiring a large number of operation and maintenance personnel (On-Call Engineers) to monitor and handle faults day and night.

[0004] (2) Inefficiency: It may take hours or even days to discover, locate, and resolve a large-scale network cluster failure, resulting in low business efficiency, delayed business progress, and huge economic losses.

[0005] (3) High reliance on experience: The effectiveness of operation and maintenance often depends on the experience and knowledge reserves of the operation and maintenance personnel. This experiential knowledge is highly specialized and difficult to be generalized and used.

[0006] Even though traditional AI technology has improved the level of automation to a certain extent, it is still highly dependent on manually labeled data, and traditional models show poor generalization ability and lack of comprehensive judgment ability when faced with highly domain-specific knowledge. Summary of the invention

[0007] The purpose of the present invention is to provide a fault detection intelligent agent solution based on a large language model to solve the problems raised in the above background technology.

[0008] To achieve the above object, the present invention provides the following technical solution: a fault detection agent solution based on a large language model, comprising the following steps:

[0009] Step S1: Receiving clues and knowledge retrieval: The model accepts manually input fault descriptions or real-time alarm data as original alarm information, performs RAG on the constructed 2,000 fault knowledge fragments according to the alarm information, and retrieves relevant historical data and processing flows from the vector library. The retrieved knowledge fragments and original alarm information together form the reasoning input for the large model of the intelligent agent to think and make decisions;

[0010] Step S2: Thinking and decision-making: Based on the alarm data, the retrieved fault knowledge base, and the completed fault detection trajectory, consider the tool to be called next or output the root cause result;

[0011] Step S3: Execute the operation: The agent calls the automation tool API and divides the tools into low-risk and high-risk. When the agent selects a low-risk tool, it automatically interacts with the environment to obtain Observation. When a high-risk tool is selected, the agent outputs a high-risk prompt to the operation and maintenance personnel, and calls the tool after the operation and maintenance personnel agree.

[0012] Step S4: Comprehensive analysis: Complete the intelligent agent fault detection reasoning framework based on the ReAct format and perform fault analysis.

[0013] Preferably, there are two comprehensive analysis methods in step S4, including the following:

[0014] 1) Decision similarity evaluation: The ReAct dataset is divided into training set, validation set, and test set in a ratio of 6:1:1. In the test set, the beginning part of the prompt and the observation part of the operation are fixed, and the model is asked to output thoughts and actions. The model output is compared with the reference output to calculate the semantic similarity;

[0015] 2) Simulation environment testing: Based on the virtual network environment built by Mininet+FRR, typical failure scenarios are reproduced, operation and maintenance tools are implemented, and the end-to-end evaluation model is used to determine whether the correct root cause can be located, while ignoring the intermediate process of location.

[0016] Preferably, the RAG workflow in step S1 is:

[0017] S11. Preprocessing: First, the large-scale corpus is preprocessed, including word segmentation, stop word removal, and vocabulary construction steps;

[0018] S12. Retrieval: During the generation process, RAG technology will retrieve relevant text fragments in the corpus based on the current context information;

[0019] S13, Generation: After obtaining the search results, RAG technology will use the generative model to generate new text. The generation process will comprehensively consider the current context information, search results and the knowledge base of the generative model itself, so as to generate more accurate and diverse texts;

[0020] S14, Post-processing: Finally, the generated text is post-processed, including removing duplicates and correcting grammatical errors to improve the quality of the generated results.

[0021] Preferably, the model in step S1 performs format processing on a network fault document set in multiple formats such as PDF and Word, extracts valid text information, and further converts it into a data set that can be directly used for knowledge retrieval and training based on GPT-4 prompt design and rule checking. The data set is further checked for length, semantic similarity, and format to ensure retrieval and training effects.

[0022] Preferably, step S3 includes a fault system environment interaction module, which is used to API-ize each tool in the tool set, and inform the model of the name, definition, calling method, and risk level of these APIs through prompts, to ensure that the large language model can understand the tools and call appropriate tools according to reasoning requirements.

[0023] Preferably, in order to stimulate the model's reasoning ability in the professional field of network fault operation and maintenance, step S4 constructs a customized model training method based on the aforementioned ReAct training data set, and designs semantic similarity and fault location end-to-end evaluation indicators for the evaluation of the agent's reasoning ability.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] Based on the high intelligence and reasoning ability of the big model, the present invention designs a series of data automation processing modules to generate the operation and maintenance expert knowledge base and ReAct format training set, and uses the training method adapted to the ReAct data set to improve the model intelligence; further, we have developed a series of fault detection tool sets, designed the interaction module between the big model and the knowledge base, and the fault detection tool set, and finally formed a fault detection intelligent agent, which provides a more intelligent and efficient solution for fault management.

[0026] Compared with traditional intelligent models, large language models have more significant reasoning capabilities. They have the ability to perform multi-step analysis, reflection, and planning when faced with complex problems. They are adapted to the distributed detection process of the root causes of network system failures and have great potential to replace manual operation and maintenance.

[0027] Since network fault root cause detection is a professional task, in order to enable the large language model to have sufficient knowledge background to carry out complex reasoning, the present invention designs a high-quality data set automatic generation module, and trains the model through a customized fine-tuning solution to enable it to have basic knowledge of network fault operation and maintenance; after receiving accurate prompt requirements and fault information input, the intelligent agent iteratively plans tasks according to the ReAct framework through the model, and interacts with the network environment in real time through the interactive module, and reasons according to the actual situation until the root cause of the fault is located, thereby automatically completing the root cause detection work. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a framework diagram of the fault detection intelligent agent of the present invention;

[0029] Figure 2 This is a flow chart of the fault detection agent training of the present invention;

[0030] Figure 3 A diagram of the fault detection agent evaluation method of the present invention. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] See also Figure 1-3 The present invention provides a technical solution: a fault detection agent solution based on a large language model, comprising the following steps:

[0033] Step S1: Receiving clues and knowledge retrieval: The model accepts manually input fault descriptions or real-time alarm data as original alarm information, performs RAG on the constructed 2,000 fault knowledge fragments according to the alarm information, and retrieves relevant historical data and processing flows from the vector library. The retrieved knowledge fragments and original alarm information together form the reasoning input for the large model of the intelligent agent to think and make decisions;

[0034] Step S2: Thinking and decision-making: Based on the alarm data, the retrieved fault knowledge base, and the completed fault detection trajectory, consider the tool to be called next or output the root cause result;

[0035] Step S3: Execute the operation: The agent calls the automation tool API and divides the tools into low-risk and high-risk. When the agent selects a low-risk tool, it automatically interacts with the environment to obtain Observation. When a high-risk tool is selected, the agent outputs a high-risk prompt to the operation and maintenance personnel, and calls the tool after the operation and maintenance personnel agree.

[0036] Step S4: Comprehensive analysis: Complete the intelligent agent fault detection reasoning framework based on the ReAct format and perform fault analysis.

[0037] In the present invention, there are two comprehensive analysis methods in step S4, including the following:

[0038] 1) Decision similarity evaluation: The ReAct dataset is divided into training set, validation set, and test set in a ratio of 6:1:1. In the test set, the beginning part of the prompt and the observation part of the operation are fixed, and the model is asked to output thoughts and actions. The model output is compared with the reference output, and the semantic similarity is calculated. After training, the average semantic similarity of the model is improved by 0.12, which is 0.11 higher than GPT-4.

[0039] 2) Simulation environment testing: Based on the virtual network environment built by Mininet+FRR, typical fault scenarios are reproduced, operation and maintenance tools are implemented, and the model is evaluated end-to-end to determine whether it can locate the correct root cause, while ignoring the intermediate process of locating. At present, the simulation environment evaluation process has been opened up based on simple fault scenarios and operation and maintenance tool sets. Through simulation analysis, it is found that the intelligent agent can shorten the single-node fault diagnosis time from the traditional 1 hour to 5 minutes, saving 92% of the time. Taking a network with 1,000 nodes as an example, about 701 man-days of labor costs can be saved in one year.

[0040] In the present invention, the RAG workflow in step S1 is:

[0041] S11. Preprocessing: First, the large-scale corpus is preprocessed, including word segmentation, stop word removal, and vocabulary construction steps;

[0042] S12. Retrieval: During the generation process, RAG technology will retrieve relevant text fragments in the corpus based on the current context information;

[0043] S13, Generation: After obtaining the search results, RAG technology will use the generative model to generate new text. The generation process will comprehensively consider the current context information, search results and the knowledge base of the generative model itself, so as to generate more accurate and diverse texts;

[0044] S14, Post-processing: Finally, the generated text is post-processed, including removing duplicates and correcting grammatical errors to improve the quality of the generated results.

[0045] In the present invention, the model in step S1 performs format processing on a network fault document set in multiple formats such as PDF and Word, extracts valid text information, and further converts it into a dataset that can be directly used for knowledge retrieval and training based on GPT-4 prompt design and rule checking. The dataset is further checked for length, semantic similarity, and format to ensure retrieval and training effects.

[0046] In the present invention, step S3 includes a fault system environment interaction module, which is used to API each tool in the tool set, and inform the model of the name, definition, calling method, and risk level of these APIs through prompts to ensure that the large language model can understand the tools and call the appropriate tools according to the reasoning requirements.

[0047] In the present invention, in order to stimulate the model's reasoning ability in the professional field of network fault operation and maintenance, step S4 constructs a customized model training method based on the aforementioned ReAct training data set, and designs semantic similarity and fault location end-to-end evaluation indicators for the evaluation of the agent's reasoning ability.

[0048] The present invention designs RCAgent, a fault root cause detection and analysis intelligent agent framework based on a large language model. The framework can automatically extract high-quality ReAct data sets from a large number of expert operation and maintenance documents, and efficiently fine-tune the parameters of the large language model on the data set. The framework is centered on the large language model with knowledge enhancement after fine-tuning. The intelligent agent interacts with the fault environment through the tool interaction module based on the model reasoning results, and iteratively reasons and calls tools through environmental feedback until the root cause corresponding to the fault is determined.

[0049] Compared with traditional intelligent models, the reasoning ability of the large language model of the present invention is more significant. When faced with complex problems, it has the ability of multi-step analysis, reflection, and planning. It is adapted to the distributed detection process of the root cause of network system failures and has a high potential to replace manual operation and maintenance. Since network failure root cause detection is a professional field task, in order to enable the large language model to have sufficient knowledge background to carry out complex reasoning, the present invention designs a high-quality data set automatic generation module, and trains the model through a customized fine-tuning scheme to enable it to have basic knowledge of network failure operation and maintenance; after receiving accurate prompt requirements and fault information input, the intelligent agent iteratively plans tasks according to the ReAct framework through the model, and interacts with the network environment in real time through the interactive module, and performs reasoning according to the actual situation until the root cause of the failure is located, thereby automatically completing the root cause detection of the failure.

[0050] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in the field. Although the embodiments of the present invention have been shown and described, it is understood by ordinary technicians in the field that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the attached claims and their equivalents.

Claims

1. A fault detection agent solution based on a large language model, characterized by: The steps include: Step S1: Receiving clues and knowledge retrieval: The model accepts manually input fault descriptions or real-time alarm data as original alarm information, performs RAG on the constructed 2,000 fault knowledge fragments according to the alarm information, and retrieves relevant historical data and processing flows from the vector library. The retrieved knowledge fragments and original alarm information together form the reasoning input for the large model of the intelligent agent to think and make decisions; Step S2: Thinking and decision-making: Based on the alarm data, the retrieved fault knowledge base, and the completed fault detection trajectory, consider the tool to be called next or output the root cause result; Step S3: Execute the operation: The agent calls the automation tool API and divides the tools into low-risk and high-risk. When the agent selects a low-risk tool, it automatically interacts with the environment to obtain Observation. When a high-risk tool is selected, the agent outputs a high-risk prompt to the operation and maintenance personnel, and calls the tool after the operation and maintenance personnel agree. Step S4: Comprehensive analysis: Complete the intelligent agent fault detection reasoning framework based on the ReAct format and perform fault analysis.

2. The fault detection agent solution based on a large language model according to claim 1 is characterized in that: There are two comprehensive analysis methods in step S4, including the following: 1) Decision similarity evaluation: The ReAct dataset is divided into training set, validation set, and test set in a ratio of 6:1:

1. In the test set, the beginning part of the prompt and the observation part of the operation are fixed, and the model is asked to output thoughts and actions. The model output is compared with the reference output to calculate the semantic similarity; 2) Simulation environment testing: Based on the virtual network environment built by Mininet+FRR, typical failure scenarios are reproduced, operation and maintenance tools are implemented, and the end-to-end evaluation model is used to determine whether the correct root cause can be located, while ignoring the intermediate process of location.

3. The fault detection agent solution based on a large language model according to claim 1 is characterized in that: The RAG workflow in step S1 is as follows: S11. Preprocessing: First, the large-scale corpus is preprocessed, including word segmentation, stop word removal, and vocabulary construction steps; S12. Retrieval: During the generation process, RAG technology will retrieve relevant text fragments in the corpus based on the current context information; S13, Generation: After obtaining the search results, RAG technology will use the generative model to generate new text. The generation process will comprehensively consider the current context information, search results and the knowledge base of the generative model itself, so as to generate more accurate and diverse texts; S14, Post-processing: Finally, the generated text is post-processed, including removing duplicates and correcting grammatical errors to improve the quality of the generated results.

4. The fault detection agent solution based on a large language model according to claim 1 is characterized in that: In step S1, the model performs format processing on a network fault document set in multiple formats such as PDF and Word, extracts valid text information, and further converts it into a dataset that can be directly used for knowledge retrieval and training based on GPT-4 prompt design and rule checking. The dataset is further checked for length, semantic similarity, and format to ensure retrieval and training effects.

5. The fault detection agent solution based on a large language model according to claim 1 is characterized in that: The step S3 includes a fault system environment interaction module, which is used to API each tool in the tool set, and inform the model of the name, definition, calling method, and risk level of these APIs through prompts to ensure that the large language model can understand the tools and call the appropriate tools according to reasoning requirements.

6. The fault detection agent solution based on a large language model according to claim 1 is characterized in that: In order to stimulate the model's reasoning ability in the professional field of network fault operation and maintenance, step S4 constructs a customized model training method based on the aforementioned ReAct training data set, and designs semantic similarity and fault location end-to-end evaluation indicators for the evaluation of the agent's reasoning ability.

Citation Information

Cited By

  • Industrial system automatic fault diagnosis method based on large language model

    CN120371587A

  • An automated fault diagnosis method for industrial systems based on a large language model

    CN120371587B

  • Composite fault detection and diagnosis system and method based on cloud-edge-end collaborative AI intelligent agent

    CN121143288A

  • Domain large model application method and device based on DSL (Digital Subscriber Line) and format checker

    CN121960456A