System fault analysis and handling method, device, equipment and computer program product

Through the method of combining vector knowledge base and large language model, system failures are automatically analyzed and dealt with, and the problem of inefficient fault handling in the existing technology is solved, and the whole process of intelligent and automated fault self-healing is achieved.

CN119938376APending Publication Date: 2025-05-06中国邮政储蓄银行股份有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411925753.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology relies on manual operations in system fault analysis and handling, and lacks intelligent solutions, resulting in inefficient fault handling.

Method used

By obtaining fault alarm information, using vector knowledge base and large language model for fault analysis, determining the fault solution and calling the corresponding fault handling agent, automatic fault handling is achieved.

Benefits of technology

It improves the accuracy of fault analysis and the efficiency of fault handling, and realizes intelligent management of the entire process and automatic fault self-healing without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938376A_ABST
    Figure CN119938376A_ABST
Patent Text Reader

Abstract

The invention discloses a system fault analysis and disposal method, device and equipment and a computer program product, and the method comprises the steps: obtaining current fault alarm information, and carrying out the retrieval in a vector knowledge base, and obtaining a historical fault retrieval result; determining a current fault solution and a corresponding fault handling agent by utilizing a large language model according to a historical fault retrieval result; according to the historical fault retrieval result and the large language model, calling a corresponding fault handling system by using a fault handling agent to handle the fault; and summarizing fault processing results by using a large language model. According to the invention, the three links of fault alarm, fault analysis and fault disposal are connected in series through the large model, intelligent management and automatic fault self-healing of the whole process are realized, and the large model is effectively connected with an external fault disposal system through an intelligent agent technology, so that the large model not only can accurately analyze the fault, but also has the fault disposal capability, and the fault disposal efficiency is improved. And the fault handling efficiency and the intelligent level are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of fault handling, and in particular to a method, device, equipment and computer program product for analyzing and handling system faults. Background Art

[0002] With the iteration and development of technology, all enterprises have generally achieved a high level of automation in the field of operation and maintenance. Specifically, enterprises have established a bottom-up all-round monitoring system from physical resources to application status, and have also applied an automated operation and maintenance platform that includes management scripts. These automated systems are like a hundred flowers blooming, and have widely penetrated into all aspects of operation and maintenance work.

[0003] System fault analysis and handling are the core of operation and maintenance work. At present, system fault analysis and handling mainly rely on various automation tools, such as monitoring tools, script execution tools, etc., and even need to log in to the physical machine to view the log of the business system to troubleshoot. Although the automation platform significantly reduces the manpower burden in fault analysis, it still requires deep human participation, especially in the fault handling link, which basically relies on full manual operation. In addition, current intelligent research is mainly focused on the fault analysis level, and lacks research on intelligent solutions for fault handling. Summary of the invention

[0004] The embodiments of the present application provide a system fault analysis and handling method, apparatus, device and computer program product to improve the accuracy of fault analysis and the efficiency of fault handling.

[0005] The present application embodiment adopts the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a system failure analysis and handling method, the system failure analysis and handling method comprising:

[0007] Get the current fault alarm information;

[0008] Search the vector knowledge base according to the current fault alarm information to obtain the historical fault retrieval results;

[0009] According to the historical fault retrieval results, a large language model is used to determine a current fault solution and a fault handling agent corresponding to the current fault solution;

[0010] According to the historical fault retrieval results and the large language model, the fault handling agent is used to call a corresponding fault handling system so that the fault handling system handles the current fault alarm information;

[0011] A fault handling result of the fault handling system is received, and the fault handling result is summarized by using the large language model.

[0012] Optionally, the vector knowledge base is constructed in the following manner:

[0013] Obtain historical fault operation and maintenance documents;

[0014] The historical fault operation and maintenance document is vectorized to obtain vectorized historical fault operation and maintenance data, wherein the vectorized historical fault operation and maintenance data includes a fault name, a fault description, a fault solution, and fault handling information.

[0015] Optionally, the determining of the fault handling agent corresponding to the current fault solution includes:

[0016] If the current fault solution contains keyword information associated with script execution, determining that the fault handling agent corresponding to the current fault solution is a script execution agent;

[0017] If the current fault solution contains keyword information associated with database execution, it is determined that the fault handling agent corresponding to the current fault solution is a database agent.

[0018] Optionally, calling a corresponding fault handling system using the fault handling agent according to the historical fault retrieval result and the large language model includes:

[0019] Acquiring the fault handling information required by the fault handling system from the historical fault retrieval results;

[0020] According to the fault handling information and preset prompt words required by the fault handling system, the large language model is used to extract the key fault handling parameters required by the fault handling system;

[0021] According to the key fault handling parameters required by the fault handling system, the corresponding fault handling system is called by the fault handling agent.

[0022] Optionally, the fault handling agent is a script execution agent, the fault handling key parameter is a script execution key parameter, the fault handling system is a script service system, and calling the corresponding fault handling system using the fault handling agent according to the fault handling key parameter required by the fault handling system includes:

[0023] Generate a first call request according to the script execution key parameters required by the script servitization system;

[0024] The script service system is called according to the first call request, so that the script service system executes the script and returns the script execution result.

[0025] Optionally, the fault handling agent is a database agent, the fault handling key parameters are database connection and execution key parameters, the fault handling system is a database system, and calling the corresponding fault handling system using the fault handling agent according to the fault handling key parameters required by the fault handling system includes:

[0026] Generate a second call request according to the database connection and execution key parameters required by the database system;

[0027] The database system is called according to the second call request, so that the database system executes a database statement and returns a database statement execution result.

[0028] Optionally, calling a corresponding fault handling system by using the fault handling agent according to the key fault handling parameters required by the fault handling system includes:

[0029] Obtain risk level information corresponding to the key fault handling parameters;

[0030] Determine whether to call a corresponding fault handling system to handle the fault according to the risk level information corresponding to the key fault handling parameters.

[0031] In a second aspect, an embodiment of the present application further provides a system fault analysis and handling device, the system fault analysis and handling device comprising:

[0032] An acquisition unit, used to acquire current fault alarm information;

[0033] A retrieval unit, used to search in the vector knowledge base according to the current fault alarm information to obtain the historical fault retrieval result;

[0034] A determination unit, configured to determine a current fault solution and a fault handling agent corresponding to the current fault solution by using a large language model according to the historical fault retrieval result;

[0035] A calling unit, configured to call a corresponding fault handling system using the fault handling agent according to the historical fault retrieval result and the large language model, so that the fault handling system handles the current fault alarm information;

[0036] The summarizing unit is used to receive the fault handling result of the fault handling system and summarize the fault handling result by using a large language model.

[0037] In a third aspect, an embodiment of the present application further provides a device, including:

[0038] A processor; and a memory arranged to store computer executable instructions, which, when executed, cause the processor to perform any of the aforementioned system failure analysis and handling methods.

[0039] In a fourth aspect, an embodiment of the present application further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements any of the aforementioned system fault analysis and handling methods.

[0040] At least one of the above technical solutions adopted in the embodiment of the present application can achieve the following beneficial effects: the system fault analysis and handling method of the embodiment of the present application first obtains the current fault alarm information; then searches in the vector knowledge base according to the current fault alarm information to obtain the historical fault retrieval results; then, according to the historical fault retrieval results, the large language model is used to determine the current fault solution and the fault handling agent corresponding to the current fault solution; then, according to the historical fault retrieval results and the large language model, the fault handling agent is used to call the corresponding fault handling system so that the fault handling system handles the current fault alarm information; finally, the fault handling result of the fault handling system is received, and the fault handling result is summarized using the large language model. The system fault analysis and handling method of the embodiment of the present application connects the three links of fault alarm occurrence, fault analysis and fault handling in series through a large model, realizes full-process intelligent management and full-process automated fault self-healing, without manual intervention, and effectively connects the large model with the external fault handling system through the intelligent agent technology, so that the large language model can not only accurately analyze faults, but also has the ability to handle faults, significantly improving the efficiency and intelligence level of fault handling. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0042] Figure 1 A flowchart of a method for analyzing and handling system failures in an embodiment of the present application is provided;

[0043] Figure 2 This is a schematic diagram of the overall process of system failure analysis and handling in an embodiment of the present application;

[0044] Figure 3 A schematic diagram of a fault handling process executed by a script in an embodiment of the present application;

[0045] Figure 4 A schematic diagram of a fault handling process executed by a database in an embodiment of the present application;

[0046] Figure 5 This is a schematic diagram of the structure of a system fault analysis and handling device in an embodiment of the present application;

[0047] Figure 6 This is a schematic diagram of the structure of a device in an embodiment of the present application. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0049] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0050] The main technical terms involved in this application include:

[0051] 1) Large Language Model (LLM): A natural language processing model based on deep learning technology that can understand and generate natural language text.

[0052] 2) Retrieval-Augmented Generation (RAG): is a technology that combines information retrieval and text generation. It enables large language models to refer to authoritative knowledge bases outside the training data source before generating responses.

[0053] 3) Script service system: It is a unique system in this application. As a key component of the automated operation and maintenance system, it can remotely issue system instructions by implanting agents in physical machines, and thus can remotely execute various system operations, commands and scripts.

[0054] 4) Prompt: refers to the input template provided by the user when interacting with the language model, which is used to guide the model to generate specific expected outputs or perform specified tasks.

[0055] 5) Function Call: This means that during the process of the model generating text, it is allowed to call external functions or services.

[0056] 6) AI Agent: refers to a system or entity that has the ability to independently perform tasks, make decisions, learn and adapt to environmental changes.

[0057] The present application embodiment provides a method for analyzing and handling system failures, such as Figure 1 As shown, a flow chart of a method for analyzing and handling system failures in an embodiment of the present application is provided, and the method for analyzing and handling system failures at least includes the following steps S110 to S150:

[0058] Step S110, obtaining current fault alarm information.

[0059] Combination Figure 2 , provides a schematic diagram of the overall process of system fault analysis and handling in an embodiment of the present application. When performing system fault analysis and handling, it is necessary to first obtain the current fault alarm information. The embodiment of the present application not only supports manual input of fault alarm information, but also supports input of fault alarm information through interface calls. The content of the fault alarm information can generally include the fault name, fault occurrence time, and fault description.

[0060] Step S120: searching in the vector knowledge base according to the current fault alarm information to obtain historical fault retrieval results.

[0061] After obtaining the fault alarm information, considering that the content of the original input fault alarm information does not have a unified format, especially the manually configured alarm descriptions are often brief and customized, when processing the fault alarm information, the key fault information can be first extracted through a large language model to improve the accuracy of subsequent fault root cause retrieval.

[0062] After obtaining the processed key fault information, you can use open source vector models such as m3e-base to convert the current key fault information into a fault vector, and then search it in the vector knowledge base. The vector knowledge base here is constructed by vectorizing historical fault knowledge.

[0063] When performing retrieval, a general cosine similarity algorithm can be used as the retrieval basis. By calculating the cosine value between the current fault vector and the historical fault vector, the historical fault results with higher similarity are taken and provided to the large language model as context information for analysis and reasoning.

[0064] The above historical fault retrieval results can be understood as the fault description fragment that is most relevant to the current fault alarm information. This fault description fragment is a whole, for example, it can include basic fault description information, fault causes and solutions, etc. If a large language model is needed to assist in fault handling, it can also include key information such as script execution and database.

[0065] Step S130, based on the historical fault retrieval results, a large language model is used to determine a current fault solution and a fault handling agent corresponding to the current fault solution.

[0066] The historical fault information retrieved in the above steps is used as the input of the large language model. The large language model can accurately analyze the current fault cause and give a corresponding fault solution description by combining the historical fault knowledge and the semantic analysis and understanding ability of the model itself. The specific type of large language model to be used can be flexibly selected by those skilled in the art in combination with the existing technology, and is not specifically limited here.

[0067] Furthermore, the embodiments of the present application set up different fault handling agents for different fault solutions. The fault handling agent can be understood as a functional module with functions such as executing fault handling. Through the agent technology, the large language model is effectively connected with the external fault handling system, thereby meeting different types of fault handling needs.

[0068] Step S140, based on the historical fault retrieval results and the large language model, the fault handling agent is used to call a corresponding fault handling system so that the fault handling system handles the current fault alarm information.

[0069] Combined with the historical fault knowledge retrieved in the above steps and the semantic analysis and understanding capabilities of the large language model itself, the fault handling agent is further used to call the corresponding external fault handling system, and the external fault handling system is used to handle the fault. The external fault handling system supported by the embodiment of the present application may include, for example, a script service system and a database, etc., to meet different types of fault handling requirements.

[0070] Step S150: receiving the fault handling result of the fault handling system, and summarizing the fault handling result by using the large language model.

[0071] According to the fault handling results fed back by the external fault handling system, the large language model can be further used to combine prompt words and context information such as fault causes and fault solutions to perform summary reasoning, and output the final fault handling summary reasoning content such as a fault handling report. Of course, the specific form of reasoning and summarization can be flexibly set by those skilled in the art according to actual needs, and no specific limitation is made here.

[0072] The system fault analysis and handling method of the embodiment of the present application connects the three links of fault alarm occurrence, fault analysis and fault handling in series through a large model, realizing full-process intelligent management and full-process automated fault self-healing without the need for human intervention, and effectively connects the large model with the external fault handling system through intelligent agent technology, so that the large language model can not only accurately analyze faults, but also has the ability to handle faults, significantly improving the efficiency and intelligence level of fault handling.

[0073] In some embodiments of the present application, the vector knowledge base is constructed in the following manner: obtaining historical fault operation and maintenance documents; vectorizing the historical fault operation and maintenance documents to obtain vectorized historical fault operation and maintenance data, wherein the vectorized historical fault operation and maintenance data includes fault name, fault description, fault solution and fault handling information.

[0074] The construction of the vector knowledge base is to vectorize various types of historical fault operation and maintenance documents through vector models such as open source vector models such as m3e-base and enter them into the knowledge base for subsequent knowledge retrieval. The vectorization process mainly involves the following two aspects:

[0075] 1) The vectorization of historical fault operation and maintenance documents involves document paragraph segmentation. The multi-level title indexing method can be used to add multi-level titles to each document block, such as "Title 1-Title 2-Text". In this way, paragraph information is introduced to improve the large language model's understanding of the document context.

[0076] 2) Knowledge items in the knowledge base need to follow a unified format: fault name - fault description - fault solution - fault handling information (such as script information / database information, etc.). Fault name and fault description are used to match fault alarm information. Fault solution is used to guide the large model to analyze the cause of the fault and provide a solution to the current fault. Fault handling information is the information required by different fault handling solutions. For example, script information needs to include script name, machine IP address where the script needs to be executed, script parameters, and other information, which is used to further interact with the large language model to remotely execute the script to solve the fault.

[0077] In some embodiments of the present application, determining the fault handling agent corresponding to the current fault solution includes: if the current fault solution contains keyword information associated with script execution, then determining that the fault handling agent corresponding to the current fault solution is a script execution agent; if the current fault solution contains keyword information associated with database execution, then determining that the fault handling agent corresponding to the current fault solution is a database agent.

[0078] The fault handling agent of the embodiment of the present application may include a script execution agent and a database agent. The script execution agent is essentially a Python function module that integrates HTTP requests. By connecting to the script service system, it has the ability to execute scripts and endows the large language model with this ability. The database agent is similar to the script execution agent and is a Python function module that realizes the ability to connect and operate the database through coding and provides a calling interface for the large language model.

[0079] Specifically, the specific agent to be used for fault handling can be determined based on the information in the fault solution currently given by the large language model. For example, if keyword information related to "script execution" can be extracted from the current fault solution, it means that the fault needs to be solved by script execution, so the script execution agent is determined to be used; if keyword information related to "SQL execution" is extracted from the current fault solution, the database agent is determined to be used.

[0080] The implementation of the agent can include API calls, Python integrated environment, etc. The script execution agent and database agent integrate Python and call the interface of the script service system and connect to the database through the HTTP protocol respectively. Since the agent itself is designed through code, it has strong scalability.

[0081] In some embodiments of the present application, calling the corresponding fault handling system using the fault handling agent based on the historical fault retrieval results and the large language model includes: obtaining fault handling information required by the fault handling system from the historical fault retrieval results; extracting key fault handling parameters required by the fault handling system using the large language model based on the fault handling information and preset prompt words required by the fault handling system; and calling the corresponding fault handling system using the fault handling agent based on the key fault handling parameters required by the fault handling system.

[0082] As mentioned above, the vector knowledge base stores historical fault information, including fault name, fault description, fault solution, and fault handling information, etc. Therefore, after searching and matching in the vector knowledge base based on the current fault alarm information, relevant historical fault information can be matched. The relevant historical fault information contains historical fault handling information, which serves as the basis for subsequent fault handling.

[0083] Since historical fault handling information contains a lot of information, we can use the constructed prompt word Prompt and the large language model to further extract key parameters related to fault handling in the context as the basis for subsequent fault handling execution.

[0084] The embodiments of the present application can more accurately and efficiently utilize historical fault information and a large language model to extract key parameters for fault handling, and call a corresponding fault handling system to handle current fault alarm information, thereby improving fault handling efficiency.

[0085] In some embodiments of the present application, the fault handling agent is a script execution agent, the fault handling key parameters are script execution key parameters, and the fault handling system is a script servitization system. According to the fault handling key parameters required by the fault handling system, calling the corresponding fault handling system using the fault handling agent includes: generating a first call request according to the script execution key parameters required by the script servitization system; calling the script servitization system according to the first call request, so that the script servitization system executes the script and returns the script execution result.

[0086] Combination Figure 3 , a schematic diagram of a script execution fault handling process in an embodiment of the present application is provided. For fault warning information that requires a script execution agent to handle the fault, the script execution agent needs to use an external script service system to handle the fault. This process requires first extracting key parameter information related to the script execution from the retrieved detailed information of the historical script execution, such as the script name, the machine IP that needs to execute the script, script parameters, etc.

[0087] A first call request is generated based on the above script execution parameters, and the script servitization system is called by the first call request. The script servitization system can execute the corresponding script command on the specified machine by parsing the script execution parameters contained in the request, and finally feed back the script execution result to the large language model.

[0088] It should be emphasized here that the "script service system" in this application is a self-owned system, which is essentially a distributed system that manages all physical machines or containers. The script service system can issue scripts or instructions to all managed physical machines or containers. Therefore, as long as there is a script to solve the problem on the corresponding machine, this solution can call the script service system to execute the corresponding script command, thereby giving the large language model the ability to solve the corresponding problem.

[0089] The embodiment of the present application connects the existing script service system to the large language model, so that the large language model has the ability to execute actual operation and maintenance operations such as scripts, thereby improving the fault handling capability.

[0090] In some embodiments of the present application, the fault handling agent is a database agent, the fault handling key parameters are database connection and execution key parameters, the fault handling system is a database system, and calling the corresponding fault handling system using the fault handling agent according to the fault handling key parameters required by the fault handling system includes: generating a second call request according to the database connection and execution key parameters required by the database system; calling the database system according to the second call request, so that the database system executes database statements and returns the database statement execution results.

[0091] Combination Figure 4 , provides a schematic diagram of a fault handling process executed by a database in an embodiment of the present application. For fault warning information that requires a database agent to handle the fault, it is necessary to access the database through the database agent to handle the fault. This process requires first extracting key parameter information related to database connection and SQL statement execution from the retrieved database connection information, such as the database IP address, port, user name and password, and parameters such as the SQL statement to be executed.

[0092] The database agent initiates a second call request to the database based on the above parameters, thereby accessing the database and executing the corresponding SQL statement for fault handling, and finally feeds back the database statement execution results to the large language model.

[0093] The embodiments of the present application enable the large language model to interact efficiently with the database, so that the large language model has the ability to execute actual operation and maintenance operations such as database statements, thereby improving the fault handling capability.

[0094] In some embodiments of the present application, calling the corresponding fault handling system using the fault handling agent according to the key fault handling parameters required by the fault handling system includes: obtaining risk level information corresponding to the key fault handling parameters; and determining whether to call the corresponding fault handling system for fault handling according to the risk level information corresponding to the key fault handling parameters.

[0095] Considering that script execution agents and database agents may involve risky operations during the execution process, such as some script execution commands or SQL statements that affect service stability, a risk identification mechanism can be introduced to ensure the safety of fault handling.

[0096] The fault handling risks involved in the embodiments of the present application are mainly divided into two types: one is the risk of fault handling based on the database, and the other is the risk of fault handling based on the script service system.

[0097] In the case of database-based fault handling, it is possible to directly determine whether there is a risk through SQL statements. For example, the large language model directly determines whether the SQL statement to be executed is a query-type SQL. Query-type SQL generally does not have systemic risks and can meet coarse-grained risk control. Of course, more refined risk intensity control can also be achieved, that is, database risk identification can be achieved through the large language model itself, and risk granularity control can be achieved through prompt words. The specific prompt word content can be flexibly set according to actual needs and is not specifically limited here.

[0098] When it comes to fault handling based on a script service system, the execution of the script may involve too many aspects, and it is difficult to directly use a large language model to make risk assessments on the script execution. Therefore, a separate risk classification label attribute can be made for the script in the script service system. Before the large language model calls the script service to execute the script, it will first obtain the risk label attribute of the script, and then determine whether to further execute the script command for fault handling to ensure the safety of the operation.

[0099] In summary, the system failure analysis and handling method of the present application has achieved at least the following technical effects:

[0100] This application integrates RAG, Prompt, Function Call and other capabilities, fully tapping and leveraging the natural language understanding capabilities of the large language model. It is a relatively complete and systematic application framework for the large language model in the field of fault analysis and handling, and can achieve full process automation and intelligence from fault occurrence to handling without human intervention.

[0101] In the fault handling phase, this application uses intelligent agent technology to achieve two core functions: one is to connect the existing script service system to the big model, so that the big language model has the ability to execute scripts and other actual operation and maintenance operations; the other is to give the big language model the ability to interact efficiently with the database. The realization of these two functions enables the big language model to not only accurately analyze faults, but also has the ability to handle faults, significantly improving the efficiency and intelligence level of fault handling.

[0102] The present application also provides a system fault analysis and handling device 500, such as Figure 5 As shown, a schematic diagram of the structure of a system fault analysis and handling device in an embodiment of the present application is provided, wherein the system fault analysis and handling device 500 includes: an acquisition unit 510, a retrieval unit 520, a determination unit 530, a calling unit 540, and a summarizing unit 550, wherein:

[0103] An acquisition unit 510 is used to acquire current fault alarm information;

[0104] A retrieval unit 520 is used to search the vector knowledge base according to the current fault alarm information to obtain a historical fault retrieval result;

[0105] A determination unit 530, configured to determine a current fault solution and a fault handling agent corresponding to the current fault solution by using a large language model according to the historical fault retrieval result;

[0106] A calling unit 540 is used to call a corresponding fault handling system using the fault handling agent according to the historical fault retrieval result and the large language model, so that the fault handling system handles the current fault alarm information;

[0107] The summarizing unit 550 is used to receive the fault handling result of the fault handling system and summarize the fault handling result by using a large language model.

[0108] In some embodiments of the present application, the vector knowledge base is constructed in the following manner: obtaining historical fault operation and maintenance documents; vectorizing the historical fault operation and maintenance documents to obtain vectorized historical fault operation and maintenance data, wherein the vectorized historical fault operation and maintenance data includes fault name, fault description, fault solution and fault handling information.

[0109] In some embodiments of the present application, the determination unit 530 is specifically used to: if the current fault solution contains keyword information associated with script execution, determine that the fault handling agent corresponding to the current fault solution is a script execution agent; if the current fault solution contains keyword information associated with database execution, determine that the fault handling agent corresponding to the current fault solution is a database agent.

[0110] In some embodiments of the present application, the calling unit 540 is specifically used to: obtain the fault handling information required by the fault handling system from the historical fault retrieval results; extract the fault handling key parameters required by the fault handling system using the large language model according to the fault handling information required by the fault handling system and preset prompt words; and call the corresponding fault handling system using the fault handling agent according to the fault handling key parameters required by the fault handling system.

[0111] In some embodiments of the present application, the fault handling agent is a script execution agent, the fault handling key parameters are script execution key parameters, the fault handling system is a script servitization system, and the calling unit 540 is specifically used to: generate a first calling request according to the script execution key parameters required by the script servitization system; call the script servitization system according to the first calling request, so that the script servitization system executes the script and returns the script execution result.

[0112] In some embodiments of the present application, the fault handling agent is a database agent, the fault handling key parameters are database connection and execution key parameters, the fault handling system is a database system, and the calling unit 540 is specifically used to: generate a second calling request based on the database connection and execution key parameters required by the database system; call the database system according to the second calling request so that the database system executes database statements and returns the database statement execution results.

[0113] In some embodiments of the present application, the calling unit 540 is specifically used to: obtain risk level information corresponding to the key fault handling parameters; and determine whether to call the corresponding fault handling system to perform fault handling according to the risk level information corresponding to the key fault handling parameters.

[0114] It can be understood that the above-mentioned system fault analysis and handling device can implement the various steps of the system fault analysis and handling method provided in the above-mentioned embodiments. The relevant explanations about the system fault analysis and handling method are applicable to the system fault analysis and handling device and will not be repeated here.

[0115] Figure 6 Schematic diagram of the structure of a device in the embodiment of the present application. Figure 6 As shown, the device includes one or more processors (or processing units), may further include one or more memories coupled to the processors, and may further include a communication module coupled to the processors.

[0116] The communication module can be used to communicate with other devices or apparatuses, such as the transmission or reception of data and / or signals. The communication module can have at least one communication module for communication. The communication module can include any interface necessary for communicating with other devices. Exemplarily, the communication module can be a transceiver, a circuit, a bus, a module, or other types of communication modules.

[0117] The processor may include, but is not limited to, at least one of the following: a general-purpose computer, a special-purpose computer, a microcontroller, a digital signal controller (DSP), or one or more of a controller-based multi-core controller architecture. The device may have multiple processors, such as application-specific integrated circuit chips, which are time-dependent and synchronized with a clock of a main processor.

[0118] The memory may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, at least one of the following: read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, hard disk, compact disc (CD), digital video disc (DVD), or other magnetic storage and / or optical storage. Examples of volatile memories include, but are not limited to, at least one of the following: random access memory (RAM), or other volatile memories that do not persist during the duration of a power outage.

[0119] The computer program includes computer executable instructions executed by an associated processor. The program can be stored in ROM. The processor can perform any suitable actions and processes by loading the program into RAM.

[0120] The possible implementation of the present application can be implemented by means of a program, so that the communication device can perform any process discussed in the above embodiments. The possible implementation of the present application can also be implemented by hardware or by a combination of software and hardware.

[0121] In some embodiments, the program may be tangibly contained in a computer-readable storage medium, which may be included in the device (such as in a memory) or other storage device accessible by the device. The program may be loaded from the computer-readable storage medium to the RAM for execution. The computer-readable storage medium may include any type of tangible non-volatile memory, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc.

[0122] The present application embodiment also provides a computer-readable storage medium, on which computer instructions or program codes are stored, and when the processor runs the instructions or the program codes, the processor executes the methods and functions involved in any of the above embodiments. Computer-readable media can be any tangible medium containing or storing programs for or related to instruction execution systems, devices or equipment. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or devices, or any suitable combination thereof. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrations. More detailed examples of computer-readable storage media include electrical connections with one or more wires, magnetic media (e.g., disks, floppy disks, hard disks, tapes, magnetic storage devices), optical media (e.g., optical storage devices, DVDs), semiconductor media (e.g., solid-state hard drives), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), or any suitable combination thereof, etc.

[0123] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The embodiment of the present application also provides at least one computer program product tangibly stored on a non-temporary computer-readable storage medium. The computer program product includes one or more computer executable instructions, such as instructions included in a program module, which are executed in a device on a real or virtual processor of the target to perform the process, method and function involved in any of the above embodiments. When the computer program instruction is loaded and executed on a computer, a process or function according to an embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instruction can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instruction can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.

[0124] The present application embodiment also proposes a computer program product, including a computer program or instruction, when the computer program or instruction is run on a computer, the computer is made to perform the process, method and function in the above-mentioned embodiment. Usually, a program module includes routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or realize specific abstract data types. In various embodiments, the functions of program modules can be combined or divided between program modules as needed. Machine executable instructions for program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.

[0125] In general, various embodiments of the present application may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be performed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of the present disclosure are shown and described as block diagrams, flow charts, or using some other graphical representations, it should be understood that the boxes, devices, systems, techniques, or methods described herein may be implemented as, for example, non-limiting examples, hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0126] It should be noted that although the embodiments of the present application are described above in conjunction with the accompanying drawings, the above embodiments are not independent of each other, and they can also be combined to obtain other embodiments. The division of the modes, situations, categories and embodiments in the embodiments of the present application is only for the convenience of description and should not constitute a special limitation. The features in the various modes, categories, situations and embodiments can be combined with each other in a logical manner. The various implementation methods of the present application can be combined arbitrarily to achieve different technical effects. The embodiments of the present application no longer list various combinations.

[0127] In addition, although the operation of the method of the present disclosure is described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flow chart can change the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution. It should also be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of a device described above can be further divided into being embodied by multiple devices.

[0128] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0129] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for analyzing and handling system failures, characterized in that: The system failure analysis and treatment method includes: Get the current fault alarm information; Search the vector knowledge base according to the current fault alarm information to obtain the historical fault retrieval results; According to the historical fault retrieval results, a large language model is used to determine a current fault solution and a fault handling agent corresponding to the current fault solution; According to the historical fault retrieval results and the large language model, the fault handling agent is used to call a corresponding fault handling system so that the fault handling system handles the current fault alarm information; A fault handling result of the fault handling system is received, and the fault handling result is summarized by using the large language model.

2. The method for analyzing and handling system failures according to claim 1, characterized in that: The vector knowledge base is constructed in the following way: Obtain historical fault operation and maintenance documents; The historical fault operation and maintenance document is vectorized to obtain vectorized historical fault operation and maintenance data, wherein the vectorized historical fault operation and maintenance data includes a fault name, a fault description, a fault solution, and fault handling information.

3. The method for analyzing and handling system failures according to claim 1, characterized in that: The fault handling agent corresponding to the current fault solution is determined to include: If the current fault solution contains keyword information associated with script execution, determining that the fault handling agent corresponding to the current fault solution is a script execution agent; If the current fault solution contains keyword information associated with database execution, it is determined that the fault handling agent corresponding to the current fault solution is a database agent.

4. The method for analyzing and handling system failures according to claim 1, characterized in that: The method of calling a corresponding fault handling system using the fault handling agent according to the historical fault retrieval result and the large language model comprises: Acquiring the fault handling information required by the fault handling system from the historical fault retrieval results; According to the fault handling information and preset prompt words required by the fault handling system, the large language model is used to extract the key fault handling parameters required by the fault handling system; According to the key fault handling parameters required by the fault handling system, the corresponding fault handling system is called by the fault handling agent.

5. The method for analyzing and handling system failures according to claim 4, characterized in that: The fault handling agent is a script execution agent, the fault handling key parameter is a script execution key parameter, the fault handling system is a script service system, and the method of calling the corresponding fault handling system by using the fault handling agent according to the fault handling key parameter required by the fault handling system includes: Generate a first call request according to the script execution key parameters required by the script servitization system; The script service system is called according to the first call request, so that the script service system executes the script and returns the script execution result.

6. The method for analyzing and handling system failures according to claim 4, characterized in that: The fault handling agent is a database agent, the fault handling key parameters are database connection and execution key parameters, the fault handling system is a database system, and the method of calling the corresponding fault handling system using the fault handling agent according to the fault handling key parameters required by the fault handling system includes: Generate a second call request according to the database connection and execution key parameters required by the database system; The database system is called according to the second call request, so that the database system executes a database statement and returns a database statement execution result.

7. The method for analyzing and handling system failures according to claim 4, characterized in that: The method of calling the corresponding fault handling system using the fault handling agent according to the key fault handling parameters required by the fault handling system includes: Obtain risk level information corresponding to the key fault handling parameters; Determine whether to call a corresponding fault handling system to handle the fault according to the risk level information corresponding to the key fault handling parameters.

8. A system failure analysis and handling device, characterized in that: The system failure analysis and treatment device comprises: An acquisition unit, used to acquire current fault alarm information; A retrieval unit, used to search in the vector knowledge base according to the current fault alarm information to obtain the historical fault retrieval result; A determination unit, configured to determine a current fault solution and a fault handling agent corresponding to the current fault solution by using a large language model according to the historical fault retrieval result; A calling unit, configured to call a corresponding fault handling system using the fault handling agent according to the historical fault retrieval result and the large language model, so that the fault handling system handles the current fault alarm information; The summarizing unit is used to receive the fault handling result of the fault handling system and summarize the fault handling result by using a large language model.

9. A device comprising: processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor executes the system fault analysis and handling method according to any one of claims 1 to 7.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the system failure analysis and handling method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Fault processing method and device, storage medium, program product and electronic equipment

    CN120196728A

  • Transaction system operation and maintenance troubleshooting method and system based on agent, and electronic equipment

    CN121542077A