Automatic fault processing method and device based on AI large model

Through the automatic fault handling method based on AI large model, the problem of coordinated processing time waste and difficult to archive historical experience in IT-related early warning problems is solved, and efficient fault handling and rapid acquisition of important early warnings are achieved.

CN120029807APending Publication Date: 2025-05-23SHANGHAI DONGPU INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510148995.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, IT-related early warning issues need to be coordinated and handled upstream and downstream, resulting in wasted time, affecting processing speed, and may expand the scope of failure impact, and historical experience is difficult to archive and reference in a timely manner.

Method used

The fault automatic handling method based on AI large model is adopted, and fault problems are identified by obtaining the operating status of IT resources, merging and pushing alarm information, using RAG technology and large-scale LLM to generate solutions, and automatically performing the fault processing process, generating and automatically archive fault reports.

Benefits of technology

It improves the efficiency and accuracy of fault handling, reduces the decline in development attention, promptly pushes important fault warnings, reduces the time to coordinate upstream and downstream personnel, and timely archives and quotes historical experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029807A_ABST
    Figure CN120029807A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic fault processing, and discloses an automatic fault processing method based on an AI large model. Comprising the steps that the running state of IT resources is obtained, important fault problems are recognized through an AI large model, the fault problems are combined, and meanwhile alarm information is pushed; a historical knowledge base is retrieved by using an RAG technology, a fault problem is preliminarily analyzed, and a suggestion and a solution are generated by using a large-scale LLM; a solution capable of being automatically executed is confirmed, and the fault processing process is automatically executed through the AI large model; generating a fault report according to the fault processing process, automatically filing the fault report, and learning the fault report through an AI large model; the method has the advantages that faults can be automatically processed, and fault solutions can be archived in time and accurately called at any time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic fault handling, and in particular, to a method and device for automatically handling faults based on an AI large model. Background Art

[0002] At present, in large enterprises, there are many IT resources such as application systems and server hosts. During IT monitoring, a large number of low-level warnings will be issued, which will not only lead to a decrease in the attention of development and processing personnel, resulting in the omission of important fault warning information, but also may cause faults and affect user usage.

[0003] For IT-related warning problems, it is usually necessary to coordinate and handle them upstream and downstream, wasting precious time in finding relevant personnel, affecting the processing speed, and likely expanding the scope of influence of the fault. After the relevant problems are processed, it is difficult to precipitate, and it is not easy to find references in a timely manner when using historical experience for reference.

[0004] Therefore, a method and device for automatically handling faults based on an AI large model are provided to solve the above problems. Summary of the Invention

[0005] The main purpose of the present invention is to solve the IT-related warning problems in the prior art. Usually, it is necessary to coordinate and handle them upstream and downstream, wasting precious time in finding relevant personnel, affecting the processing speed, and likely expanding the scope of influence of the fault. After the relevant problems are processed, it is difficult to precipitate, and it is not easy to find references in a timely manner when using historical experience for reference.

[0006] The first aspect of the present invention provides a method for automatically handling faults based on an AI large model. The method for automatically handling faults based on an AI large model includes: Obtain the running status of IT resources, identify fault problems through an AI large model, merge the fault problems and push alarm information at the same time; Use the RAG technology to retrieve the historical knowledge base, conduct a preliminary analysis of the fault problems, and use a large-scale LLM to generate suggestions and solutions; Confirm the automatically executable solutions, and automatically execute the fault handling process through an AI large model; Generate a fault report according to the fault handling process, automatically archive the fault report, and learn from the fault report through an AI large model.

[0007] Further, the obtaining the running status of IT resources, identifying important fault problems through an AI large model, merging the fault problems and pushing alarm information at the same time includes: Regularly collect the running status information of IT resources through an IT monitoring tool. The IT monitoring tool includes Zabbix and Nagios, and the IT resources include application systems and servers; Analyze the operating status information of IT resources through AI big models to identify abnormal value information and potential fault problems; Determine whether the number of potential fault problems is greater than 1 within the preset detection time. If so, analyze the potential fault problems, obtain the root cause information of the fault problems, merge the fault problems with the same root cause information, generate a problem description, generate push alarm information based on the problem description, and send the push alarm information to the operation and maintenance client.

[0008] Furthermore, if the number of potential fault problems is 1 within a preset detection time, push alarm information is directly generated according to the potential fault problem, and the push alarm information is sent to the operation and maintenance client.

[0009] Furthermore, the “using RAG technology to search the historical knowledge base, conduct preliminary analysis of the fault problem, and use large-scale LLM to generate suggestions and solutions” specifically includes: Use information retrieval algorithms to retrieve documents and records related to the current fault from the historical knowledge base; LLM generates a preliminary analysis report of the fault based on the documents and records related to the current fault retrieved from the historical knowledge base; Organize and filter the relevant document contents retrieved from the historical knowledge base, extract key information, and form a complete context text together with the current fault problem description; According to the API requirements, the constructed context text is formatted and encapsulated, and then sent to the API of the large-scale language model, requesting that fault suggestions and solutions be generated based on the formatted and encapsulated context text; After the backend receives the content generated by the model, it parses and processes it and generates a fault handling decision.

[0010] Furthermore, the “confirming a solution that can be automatically executed and automatically executing the fault handling process through the AI ​​big model” includes: Obtain the fault suggestions and solutions generated by the API for analysis to determine whether there are parts in the fault suggestions and solutions that can be automatically executed by writing automation scripts or using existing automation tools. If so, include them in the list of automatic execution solutions.

[0011] Furthermore, it also includes: Specify the parameters of the scheme design in the automatic execution scheme list and set the parameter legal conditions; the parameter legal conditions include: the parameter is not a null value and the parameter is not an abnormal value; Perform parameter validity verification and eliminate parameters that do not meet parameter validity conditions. The parameters include the file path, parameter value information, and service name that need to be modified.

[0012] Further, the "generating a fault report according to the fault handling process, automatically archiving the fault report, and learning the fault report through the AI large model" includes: Generating a fault report according to the fault handling process, where the fault report includes detailed information about the fault, the repair solution adopted, the repair time, and the repair result, and automatically archiving the fault report; Learning the fault report through the AI large model, extracting the fault mode and repair strategy in the fault report, and storing the fault mode and repair strategy correspondingly in the historical knowledge base.

[0013] The second aspect of the present invention provides a device for automatically handling faults based on an AI large model, and the device for automatically handling faults based on an AI large model includes: A fault identification module, configured to obtain the operating status of IT resources, identify important fault problems through the AI large model, merge the fault problems, and push alarm information; A solution generation module, configured to retrieve the historical knowledge base using the RAG technology, perform a preliminary analysis on the fault problems, and generate suggestions and solutions using a large-scale LLM; A fault handling execution module, configured to confirm the solutions that can be automatically executed and automatically execute the fault handling process through the AI large model; A fault report generation and learning module, configured to generate a fault report according to the fault handling process, automatically archive the fault report, and learn the fault report through the AI large model.

[0014] The third aspect of the present invention provides an electronic device, and the electronic device includes a memory and at least one processor, and instructions and data are stored in the memory; The at least one processor calls the instructions and data in the memory so that the electronic device executes each step of the method for automatically handling faults based on an AI large model as described in any one of the above.

[0015] The fourth aspect of the present invention provides a readable storage medium, and instructions and data are stored on the readable storage medium, and it is characterized in that when the instructions are executed by a processor, each step of the method for automatically handling faults based on an AI large model as described in any one of the above is implemented.

[0016] The present invention uses a dedicated IT monitoring tool, AI big model, to obtain the operating status of enterprise IT-related machines, systems, etc., issue early warnings, automatically pull up a dedicated fault handling group for important fault problems, push relevant alarms, recommend problem solutions based on the AI ​​big model, and extract historical fault handling experience to assist in quickly advancing problem / solutions, extract historical communication records within the group, use the AI ​​big model to automatically summarize them as fault event reports for storage, and automatically extract historical experience solutions the next time a problem occurs to assist in quickly resolving the fault; solves the problem of a large number of invalid alarms in IT monitoring affecting development attention, and facilitates timely acquisition of important fault problem early warnings and timely processing; solves the problem of difficulty and slow coordination between upstream and downstream related personnel when a fault problem occurs, and quickly collaborates with relevant experts to deal with it; solves the problem of valuable historical experience not being able to be archived and saved in time, and the problem of being able to archive it in time and use it accurately at any time. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart of a method for automatically handling faults based on an AI big model provided in an embodiment of the present invention; Figure 2 A schematic diagram of the structure of an automatic fault handling device based on an AI large model provided in an embodiment of the present invention; Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The embodiment of the present invention provides a method for automatic fault handling based on an AI big model, including obtaining the operating status of IT resources, identifying important fault problems through the AI ​​big model, merging the fault problems and pushing alarm information at the same time; using RAG technology to retrieve the historical knowledge base, performing a preliminary analysis of the fault problems, and using large-scale LLM to generate suggestions and solutions; confirming solutions that can be automatically executed, and automatically executing the fault handling process through the AI ​​big model; generating a fault report according to the fault handling process, automatically archiving the fault report, and learning the fault report through the AI ​​big model. The main purpose of the present invention is to solve the IT-related early warning problems in the prior art, which usually require upstream and downstream coordination and processing, waste precious time when looking for relevant personnel, affect the processing speed, and are likely to expand the scope of influence of the fault. After the relevant problems are processed, it is difficult to settle, and it is not easy to find the reference problem in time when using reference historical experience.

[0019] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0020] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 The first embodiment of the method for automatically handling faults based on the AI ​​big model of the present invention includes: Obtain the operating status of IT resources, identify faults through AI big models, merge faults and push alarm information simultaneously; Specifically, it includes: regularly collecting the operating status information of IT resources through IT monitoring tools, the IT monitoring tools include Zabbix and Nagios, and the IT resources include application systems and servers; in this step, the CPU usage, memory usage, disk space, network traffic and other indicators of the server, as well as the response time, throughput and other indicators of the application system are regularly collected through commonly used IT monitoring tools such as Zabbix and Nagios.

[0021] Analyze the operating status information of IT resources through AI big models to identify abnormal value information and potential fault problems; Determine whether the number of potential fault problems is greater than 1 within the preset detection time. If so, analyze the potential fault problems, obtain the root cause information of the fault problems, merge the fault problems with the same root cause information, generate a problem description, generate push alarm information based on the problem description, and send the push alarm information to the operation and maintenance client.

[0022] If the number of potential fault problems is 1 within the preset detection time, a push alarm message is directly generated based on the potential fault problem and sent to the operation and maintenance client. If the number of potential fault problems is 0 within the preset detection time, no push alarm message is generated and the next detection time period is entered; the fault problems can be merged or alarm messages can be directly generated based on the number of fault problems, which can improve the efficiency of fault problem solving and prevent repeated solutions to the same type of problems. It solves the problem of a large number of invalid alarms from IT monitoring affecting development attention, and facilitates timely acquisition of important fault problem warnings and timely processing.

[0023] Use the RAG technology to retrieve the historical knowledge base, conduct a preliminary analysis of the fault problem, and use a large-scale LLM to generate suggestions and solutions; Specifically, it includes: using information retrieval algorithms such as inverted index and vector space model to retrieve relevant documents and records related to the current fault from the historical knowledge base; The LLM generates a preliminary analysis report of the fault based on the documents and records retrieved from the historical knowledge base related to the current fault. For example, for the fault of high server memory usage, the LLM can generate an analysis report indicating that it may be caused by a memory leak in a certain application, the affected scope is all applications on the server, and the recommended solution is to restart the application or optimize its memory usage; Organize and filter the relevant document content retrieved from the historical knowledge base, extract key information, and form a complete context text together with the description of the current fault problem; According to the API requirements, format and encapsulate the constructed context text, and send it to the API of the large-scale language model to request the generation of fault suggestions and solutions based on the formatted and encapsulated context text; After receiving the input, the model generates a reply similar to the following based on the knowledge it has learned and the given context: "Based on historical cases and the current fault description, it is recommended to first check whether the configuration related to the login interface in the configuration file of the application system app1 is correct, and focus on checking whether the interface path is consistent with the expectation. If the configuration is correct, you can further check the network routing settings to see if there is a situation where requests are wrongly forwarded resulting in 404 errors. In terms of solutions, if it is determined to be a configuration problem, correct the configuration in a timely manner and restart the relevant service modules; if it is a network routing problem, contact the network operation and maintenance team to assist in troubleshooting and adjusting the routing table." After the backend receives the content generated by the model, it parses and processes it to prepare for the fault handling decision After the backend receives the content generated by the model, it parses and processes it to generate a fault handling decision; Confirm the solutions that can be automatically executed, and automatically execute the fault handling process through the AI large model; Specifically, it includes: obtaining the fault suggestions and solutions generated by the API for analysis, and judging whether there is a part in the fault suggestions and solutions that can be automatically executed by writing automation scripts or using existing automation tools. If so, list it in the automatic execution plan list; It also includes: clarifying the parameters designed in the plans in the automatic execution plan list and setting the legal conditions for the parameters; the legal conditions for the parameters include: the parameters are not null values and the parameters are not abnormal values; Perform parameter validity verification and eliminate parameters that do not meet parameter validity conditions. The parameters include file paths that need to be modified, parameter value information, and service names; for example, file paths that need to be modified that are controls are eliminated, and parameter value information that is abnormal is eliminated. In general, parameter validity verification is a method commonly used by technicians in this field to remove abnormal parameter values. Verifying parameters and eliminating parameters that do not meet the conditions can improve the accuracy of subsequent execution.

[0024] Generate fault reports based on the fault handling process, automatically archive the fault reports, and learn the fault reports through the AI ​​big model; Specifically, it includes: generating a fault report according to the fault handling process, wherein the fault report includes detailed information of the fault, the repair plan adopted, the repair time and the repair result, and automatically archiving the fault report; Through the AI ​​big model to learn fault reports, extract the fault modes and repair strategies in the fault reports, and store the fault modes and repair strategies correspondingly in the historical knowledge base.

[0025] See also Figure 1 The second embodiment of the method for automatically handling faults based on the AI ​​big model of the present invention includes: Obtain the operating status of IT resources, identify faults through AI big models, merge faults and push alarm information simultaneously; Specifically, it includes: regularly collecting the operating status information of IT resources through IT monitoring tools, the IT monitoring tools include Zabbix and Nagios, and the IT resources include application systems and servers; in this step, the CPU usage, memory usage, disk space, network traffic and other indicators of the server, as well as the response time, throughput and other indicators of the application system are regularly collected through commonly used IT monitoring tools such as Zabbix and Nagios.

[0026] Analyze the operating status information of IT resources through AI big models to identify abnormal value information and potential fault problems; A time period is used as a preset detection time, for example, one hour, and it is determined whether the number of potential fault problems is greater than one within the one-hour detection time. If so, the potential fault problem is analyzed to obtain the root cause information of the fault problem, and the fault problems with the same root cause information are merged to generate a problem description. According to the problem description, a push alarm information is generated and sent to the operation and maintenance client; If the number of potential fault problems is 1 within the detection time of 1 hour, a push alarm information is directly generated based on the potential fault problem and sent to the operation and maintenance client; If the number of potential fault problems is 0 within 1 hour of detection time, no push alarm information will be generated and the next detection period will be entered; fault problems can be merged or alarm information can be directly generated according to the number of fault problems, which can improve the efficiency of fault problem solving and prevent repeated solutions to the same type of problems. It solves the problem that a large number of invalid alarms from IT monitoring affect development attention, and facilitates timely acquisition of important fault problem warnings and timely processing.

[0027] Use RAG technology to search the historical knowledge base, conduct preliminary analysis of the fault problem, and use large-scale LLM to generate suggestions and solutions; Specifically, it includes: using information retrieval algorithms, such as inverted index and vector space model, to retrieve documents and records related to the current fault from the historical knowledge base; LLM generates a preliminary analysis report of the fault based on documents and records related to the current fault retrieved from the historical knowledge base. For example, for a fault of excessive memory usage on a server, LLM can generate an analysis report indicating that it may be caused by a memory leak in a certain application, affecting all applications on the server. The recommended solution is to restart the application or optimize its memory usage. Organize and filter the relevant document contents retrieved from the historical knowledge base, extract key information, and form a complete context text together with the current fault problem description; According to the API requirements, the constructed context text is formatted and encapsulated, and then sent to the API of the large-scale language model, requesting that fault suggestions and solutions be generated based on the formatted and encapsulated context text; After receiving the input, the model generates a response similar to the following based on the knowledge it has learned and the given context: "Based on historical cases and the current fault description, it is recommended to first check whether the login interface configuration in the configuration file of the application system app1 is correct, and focus on whether the interface path is consistent with expectations. If the configuration is correct, you can further check the network routing settings to see if there is a 404 error caused by incorrect request forwarding. In terms of solutions, if it is determined to be a configuration problem, correct the configuration in time and restart the relevant service modules; if it is a network routing problem, contact the network operation and maintenance team to assist in troubleshooting and adjusting the routing table." After the backend receives the content generated by the model, it parses and processes it to prepare for fault handling decisions. After receiving the content generated by the model, the backend parses and processes it and generates a fault handling decision; Identify solutions that can be automatically executed and automatically execute the troubleshooting process through AI big models; Specifically, it includes: obtaining the fault suggestions and solutions generated by the API for analysis, determining whether there are parts in the fault suggestions and solutions that can be automatically executed by writing automation scripts or using existing automation tools, and if so, including them in the list of automatic execution solutions; It also includes: clarifying the parameters of the scheme design in the automatic execution scheme list and setting the parameter legal conditions; the parameter legal conditions include: the parameter is not a null value and the parameter is not an abnormal value; Perform parameter validity verification and eliminate parameters that do not meet parameter validity conditions. The parameters include file paths that need to be modified, parameter value information, and service names; for example, file paths that need to be modified that are controls are eliminated, and parameter value information that is abnormal is eliminated. In general, parameter validity verification is a method commonly used by technicians in this field to remove abnormal parameter values. Verifying parameters and eliminating parameters that do not meet the conditions can improve the accuracy of subsequent execution.

[0028] Generate fault reports based on the fault handling process, automatically archive the fault reports, and learn the fault reports through the AI ​​big model; Specifically, it includes: generating a fault report according to the fault handling process, wherein the fault report includes detailed information of the fault, the repair plan adopted, the repair time and the repair result, and automatically archiving the fault report; Through AI big model learning fault reports, the fault modes and repair strategies in the fault reports are extracted, and the fault modes and repair strategies are stored in the historical knowledge base. For example, when the login function of the application system app1 has a 404 error, it is found that the login interface path in the configuration file has been modified by mistake. The corresponding solution is to restore the correct configuration to solve the problem and store it in the historical knowledge base. This solves the problem that valuable historical experience cannot be archived and used in time, and can be archived in time and accurately called at any time.

[0029] See also Figure 1 The third embodiment of the method for automatically handling faults based on the AI ​​big model of the present invention includes: Obtain the operating status of IT resources, identify important fault problems through AI big models, merge fault problems and push alarm information at the same time; Specifically, they include: Regularly collect the operation status information of IT resources through IT monitoring tools, wherein the IT monitoring tools include Zabbix and Nagios, and the IT resources include application systems and servers; Analyze the operating status information of IT resources through AI big models to identify abnormal value information and potential fault problems; Determine whether the number of potential fault problems is greater than or equal to 1 within the first preset detection time. If so, analyze the potential fault problems, obtain the root cause information of the fault problems, merge the fault problems with the same root cause information, generate a problem description, generate push alarm information according to the problem description, and send the push alarm information to the operation and maintenance client; if the number of potential fault problems is 1 within the first preset detection time, directly generate push alarm information according to the potential fault problems, and send the push alarm information to the operation and maintenance client; In the specific implementation process, common IT monitoring tools such as Zabbix and Nagios are used to regularly collect indicators such as server CPU usage, memory usage, disk space, network traffic, and application system response time and throughput. AI big models can use deep learning algorithms such as convolutional neural networks (CNN) or recurrent neural networks (RNN) to analyze these indicators and identify outliers and potential faults. When multiple faults are identified, clustering algorithms are used to merge similar faults.

[0030] If multiple related fault problems are detected in a short period of time (for example, multiple different modules of the same application system have response timeouts), in order to avoid causing too much scattered alarm information interference to the operation and maintenance personnel, these problems need to be merged. By analyzing the root cause of the fault problem (for example, whether it is caused by the same server resource bottleneck), the scope of impact (the application system modules involved), and other factors, similar fault problems can be integrated into a more general problem description. For example, there were originally three problems about the response timeout of different interfaces of the application system app1. After merging, the description is "multiple interfaces of the application system app1 have response timeouts, suspected to be caused by insufficient server resources"; according to the merged fault problems, the alarm information content is constructed, including the fault title, detailed description, fault occurrence time, resources involved (server name, application system name, etc.), and then the appropriate alarm channel is selected for push.

[0031] Use RAG technology to search the historical knowledge base, conduct preliminary analysis of the fault problem, and use large-scale LLM to generate suggestions and solutions; Specifically, they include: Use information retrieval algorithms to retrieve documents and records related to the current fault from the historical knowledge base; LLM generates a preliminary analysis report of the fault based on the input information; Organize and filter the relevant document contents retrieved from the historical knowledge base, extract key information, and form a complete context text together with the current fault problem description; According to the API requirements, the constructed context text is formatted and encapsulated, and then sent to the API of the large-scale language model, requesting that fault suggestions and solutions be generated based on the above information; After receiving the content generated by the model, the backend parses and processes it and prepares fault handling decisions; In the specific use process, RAG technology first uses information retrieval algorithms, such as inverted index or vector space model, to retrieve documents and records related to the current fault from the historical knowledge base. Then, the retrieved information is input into a large-scale LLM, such as GPT-4 or BERT. LLM can generate a preliminary analysis report of the fault based on the input information, including possible causes, scope of impact, and recommended solutions. For example, for a fault with excessive memory usage on a server, LLM can generate an analysis report indicating that it may be caused by a memory leak in a certain application, and the scope of impact is all applications on the server. The recommended solution is to restart the application or optimize its memory usage.

[0032] The relevant document content retrieved from the historical knowledge base is sorted and screened, key information is extracted, and a complete context text is constructed together with the current fault problem description. For example, the cause analysis of the previous similar app1 application system login function failure, the attempted solutions, etc., are spliced ​​with the current "404 error occurred in the login function of the application system app1, and there is no obvious abnormality on the server" fault description to form a text similar to the following format: "Historical cases show that when the login function of the application system app1 previously had a 404 error, it was found that the login interface path in the configuration file was mistakenly modified, and the problem was solved by restoring the correct configuration. The current fault is that the login function of the application system app1 has a 404 error, and there is no obvious abnormality on the server." According to the API requirements, the constructed context text is formatted and encapsulated, and sent to the API of the large-scale language model, requesting it to generate fault suggestions and solutions based on this information. After receiving the input, the model generates a response similar to the following based on the knowledge it has learned and the given context: "Based on historical cases and the current fault description, it is recommended to first check whether the login interface configuration in the configuration file of the application system app1 is correct, and focus on whether the interface path is consistent with expectations. If the configuration is correct, you can further check the network routing settings to see if there is a situation where the request is forwarded incorrectly, resulting in a 404 error. In terms of solutions, if it is determined to be a configuration problem, correct the configuration in time and restart the relevant service modules; if it is a network routing problem, contact the network operation and maintenance team to assist in troubleshooting and adjusting the routing table." After receiving the content generated by the model, the backend parses and processes it to prepare for fault handling decisions.

[0033] Identify solutions that can be automatically executed and automatically execute the troubleshooting process through AI big models; Specifically, it includes: obtaining the fault suggestions and solutions generated by the API for analysis, determining whether there are parts in the fault suggestions and solutions that can be automatically executed by writing automation scripts or using existing automation tools, and if so, including them in the list of automatic execution solutions; clarifying the parameters of the solution design in the list of automatic execution solutions and verifying their legitimacy, the parameters including the file path to be modified, the specific parameter value and the name of the service; In the specific use process, for each fault, there may be multiple potential repair solutions. The AI ​​big model can use decision tree algorithms or reinforcement learning algorithms to determine which solutions can be automatically executed. For example, if the fault is that the server disk space is insufficient, the solution that can be automatically executed may be to delete some temporary files or compress some old log files. The AI ​​big model can determine the feasibility and risk of these solutions based on the current status and historical data of the server, and automatically execute them. During the execution process, the AI ​​big model can monitor the status of the server in real time to ensure the effectiveness of the repair solution.

[0034] Analyze the suggestions and solutions generated by LLM to determine which parts can be automatically executed by writing automation scripts or using existing automation tools. For example, for suggestions such as "If it is determined to be a configuration problem, correct the configuration in time and restart the relevant service module", if the configuration modification is a simple text replacement operation and the service restart is supported by the corresponding command line tool, then it can be regarded as an automatically executable part. Judgment can be made based on the rule engine, such as using the rule engine framework such as Drools, and according to predefined rules, such as "If the fault type is a configuration file parameter error and the modification method is simple text replacement, and the corresponding service has a restart command, it is determined to be automatically executable", each suggestion is evaluated and screened in this way to determine the final list of automatic execution solutions. For the determined automatic execution solution, it is necessary to clarify the parameters involved, such as the configuration file path to be modified, the specific parameter value, the name of the service, etc., and perform legality verification. For example, verify whether the configuration file path exists and whether the parameter value conforms to the required format. You can perform these operations by writing corresponding verification functions, such as using the os.path.exists function in Python to verify whether a file path exists, using regular expressions to verify the parameter value format, etc., to ensure that the automatic execution plan does not fail due to problems such as parameter errors during actual execution.

[0035] Generate fault reports based on the fault handling process, automatically archive the fault reports, and learn the fault reports through the AI ​​big model; Specifically, they include: Generate a fault report according to the fault handling process, the fault report includes detailed information of the fault, the repair plan adopted, the repair time and the repair result, and automatically archive the fault report; Use AI big models to learn fault reports and extract fault modes and repair strategies from them; In the specific implementation process, when the fault handling is completed, the system will feed back the results to the administrator and automatically generate a fault report. The fault report can include detailed information of the fault, the repair plan taken, the repair time and results, etc. Then, the fault report is automatically archived for subsequent query and analysis. The AI ​​big model can use machine learning algorithms such as supervised learning or unsupervised learning to analyze the fault report and extract useful information such as fault mode, repair strategy, etc. This information can be used to improve the fault identification and repair capabilities of the AI ​​big model to form a closed loop. For example, the AI ​​big model can learn that certain types of faults are usually caused by specific reasons by analyzing a large number of fault reports, so that these faults can be identified and resolved faster in the future. It solves the problem that valuable historical experience cannot be archived and used in time, and can be archived in time and accurately called at any time.

[0036] The above describes the method for automatically handling faults based on the AI ​​big model in the embodiment of the present invention. The following describes the device for automatically handling faults based on the AI ​​big model in the embodiment of the present invention. Figure 2 In the embodiment of the present invention, the method and device for automatically handling faults based on the AI ​​big model includes the following for the above embodiment: The fault identification module 201 is used to obtain the operating status of IT resources, identify important fault problems through the AI ​​big model, merge the fault problems and push alarm information at the same time; Solution generation module 202, for searching the historical knowledge base using RAG technology, performing preliminary analysis on the fault problem, and generating suggestions and solutions using large-scale LLM; The fault handling execution module 203 is used to confirm the automatically executable solution and automatically execute the fault handling process through the AI ​​big model; The fault report generation learning module 204 is used to generate a fault report according to the fault handling process, automatically archive the fault report, and learn the fault report through the AI ​​big model.

[0037] above Figure 2 The automatic fault handling device based on the AI ​​big model in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The electronic device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0038] Figure 37 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device 700 may have relatively large differences due to different configurations or performances, and may include one or more processors 710 (for example, one or more processors) and a memory 720, and one or more storage media 730 (for example, one or more storage devices, including RAM\FLASH, etc.) storing application programs 733 or data 732. Among them, the memory 720 and the storage medium 730 may be temporary storage or permanent storage. The program stored in the storage medium 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the electronic device 700. Furthermore, the processor 710 may be configured to communicate with the storage medium 730 to execute a series of instruction operations in the storage medium 730 on the electronic device 700.

[0039] The electronic device 700 may also include one or more power supplies 740, one or more input / output interfaces 750, and / or one or more operating systems 731, such as FreeRTOS, Android, etc. It will be appreciated by those skilled in the art that Figure 3 The structure of the electronic device shown does not constitute a limitation on the electronic device, and may include more or less components than shown in the figure, or combine some components, or arrange the components differently.

[0040] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of a method for automatic fault handling based on an AI large model.

[0041] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0042] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, mobile device, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0043] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for automatic fault handling based on an AI large model, characterized in that: The method for automatically handling faults based on the AI ​​big model includes: Obtain the operating status of IT resources, identify faults through AI big models, merge faults and push alarm information simultaneously; Use RAG technology to search the historical knowledge base, conduct preliminary analysis of the fault problem, and use large-scale LLM to generate suggestions and solutions; Identify solutions that can be automatically executed and automatically execute the troubleshooting process through AI big models; Generate a fault report based on the fault handling process, automatically archive the fault report, and learn the fault report through the AI ​​big model.

2. The method for automatic fault handling based on AI big model according to claim 1 is characterized in that: The operation status of IT resources is obtained, important fault problems are identified through the AI ​​big model, and the fault problems are merged and the alarm information is pushed at the same time, including: Regularly collect the operation status information of IT resources through IT monitoring tools, wherein the IT monitoring tools include Zabbix and Nagios, and the IT resources include application systems and servers; Analyze the operating status information of IT resources through AI big models to identify abnormal value information and potential fault problems; Determine whether the number of potential fault problems is greater than 1 within the preset detection time. If so, analyze the potential fault problems, obtain the root cause information of the fault problems, merge the fault problems with the same root cause information, generate a problem description, generate push alarm information based on the problem description, and send the push alarm information to the operation and maintenance client.

3. The method for automatic fault handling based on AI big model according to claim 2 is characterized in that: It also includes that if the number of potential fault problems is 1 within a preset detection time, push alarm information is directly generated according to the potential fault problem, and the push alarm information is sent to the operation and maintenance client.

4. The method for automatic fault handling based on AI big model according to claim 1 is characterized in that: The "using RAG technology to search the historical knowledge base, conduct preliminary analysis of the fault problem, and use large-scale LLM to generate suggestions and solutions" specifically includes: Use information retrieval algorithms to retrieve documents and records related to the current fault from the historical knowledge base; LLM generates a preliminary analysis report of the fault based on the documents and records related to the current fault retrieved from the historical knowledge base; Organize and filter the relevant document contents retrieved from the historical knowledge base, extract key information, and form a complete context text together with the current fault problem description; According to the API requirements, the constructed context text is formatted and encapsulated, and then sent to the API of the large-scale language model, requesting that fault suggestions and solutions be generated based on the formatted and encapsulated context text; After the backend receives the content generated by the model, it parses and processes it and generates a fault handling decision.

5. The method for automatic fault handling based on AI big model according to claim 1 is characterized in that: The "confirmation of solutions that can be automatically executed and automatic execution of the fault handling process through the AI ​​big model" includes: Obtain the fault suggestions and solutions generated by the API for analysis to determine whether there are parts in the fault suggestions and solutions that can be automatically executed by writing automation scripts or using existing automation tools. If so, include them in the list of automatic execution solutions.

6. The method for automatic fault handling based on AI big model according to claim 5 is characterized in that: Also includes: Clarify the parameters of the scheme design in the automatic execution scheme list and set the legal conditions of the parameters; The parameter legal conditions include: the parameter is not a null value and the parameter is not an abnormal value; Perform parameter validity verification and eliminate parameters that do not meet parameter validity conditions. The parameters include the file path, parameter value information, and service name that need to be modified.

7. The method for automatic fault handling based on AI big model according to claim 1 is characterized in that: The “generating a fault report according to the fault handling process, automatically archiving the fault report, and learning the fault report through the AI ​​big model” includes: Generate a fault report according to the fault handling process, the fault report includes detailed information of the fault, the repair plan adopted, the repair time and the repair result, and automatically archive the fault report; Through the AI ​​big model to learn fault reports, extract the fault modes and repair strategies in the fault reports, and store the fault modes and repair strategies correspondingly in the historical knowledge base.

8. An automatic fault handling device based on an AI large model, characterized in that: include: The fault identification module is used to obtain the operating status of IT resources, identify important fault problems through the AI ​​big model, merge the fault problems and push alarm information at the same time; Solution generation module, which uses RAG technology to search the historical knowledge base, conduct preliminary analysis of fault problems, and use large-scale LLM to generate suggestions and solutions; The fault handling execution module is used to confirm the solutions that can be automatically executed and automatically execute the fault handling process through the AI ​​big model; The fault report generation learning module is used to generate fault reports according to the fault handling process, automatically archive the fault reports, and learn the fault reports through the AI ​​big model.

9. An electronic device, comprising a memory and at least one processor, wherein the memory stores instructions and data; The at least one processor calls the instructions and data in the memory so that the electronic device executes the various steps of the method for automatic fault handling based on the AI ​​big model as described in any one of claims 1-7.

10. A readable storage medium having instructions and data stored thereon, characterized in that: When the instructions are executed by the processor, the various steps of the method for automatic fault handling based on the AI ​​big model as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Equipment abnormal fault recovery method and device based on large model, equipment and medium

    CN120469848A

  • LLM-based equipment fault behavior logic model intelligent generation method and device

    CN121809545A

  • Intelligent Generation Method and Device for Equipment Fault Behavior Logic Model Based on LLM

    CN121809545B