Business system fault processing method, device and electronic equipment

By introducing a multi-party collaborative analysis and processing mechanism into the business system, and utilizing work orders and neural network models, the problem of low efficiency in handling server failures in the business system was solved, and an efficient and reliable failure handling process was achieved, ensuring stable system operation.

CN115080284BActive Publication Date: 2026-02-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110268906.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-12
Publication Date
2026-02-03
Estimated Expiration
2041-05-29

AI Technical Summary

Technical Problem

In existing technologies, troubleshooting server failures in business systems requires a long time and significant manpower costs, resulting in low efficiency and impacting system stability.

Method used

By monitoring business system failures, work orders are sent to operation and maintenance nodes and server supplier nodes for multi-party collaborative analysis and processing, including first failure analysis, second failure analysis and failure handling process, and neural network models are used to improve the accuracy and efficiency of analysis.

Benefits of technology

It has enabled automated handling of business system faults, improved the reliability and efficiency of fault analysis, reduced the workload of each node, and ensured the stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115080284B_ABST
    Figure CN115080284B_ABST
Patent Text Reader

Abstract

The application provides a fault processing method and device of a business system, an electronic device and a computer readable storage medium, relates to monitoring and fault detection technologies in the field of cloud technologies, and the method comprises the following steps: monitoring a server that has a fault in a business system, and sending a work order carrying fault data to an operation and maintenance node, so that the operation and maintenance node performs first fault analysis and processing according to the fault data and obtains a first fault analysis result; according to the first fault analysis result, a work order carrying fault data is sent to a server supplier node, so that the server supplier node performs second fault analysis and processing according to the fault data and obtains a second fault analysis result; the second fault analysis result is sent to the operation and maintenance node, and according to a feedback result of the operation and maintenance node for the second fault analysis result, a fault processing process is performed. Through the application, the fault processing efficiency can be improved, so that the stable operation of the business system is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer technology, and more particularly to a fault handling method, apparatus, electronic device, and computer-readable storage medium for a business system. Background Technology

[0002] With the development of monitoring and fault detection technologies in the field of cloud technology, people are using applications running on smartphones, tablets, and computers more and more frequently in their daily learning and life. The stable operation of the business systems on which these applications rely is a basic guarantee for people's normal learning and life.

[0003] If a server encounters a fault during the operation of a business system, it may generate corresponding alarm information. When the maintenance personnel of the business system receive the corresponding alarm, they need to handle each alarm one by one, which requires a long processing time and a large manpower cost, reducing the efficiency of fault handling and affecting the stable operation of the business system. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and computer-readable storage medium for handling faults in a business system, which can efficiently handle faults in the business system.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides a fault handling method for a business system, including:

[0007] Monitor the server that has failed in the business system, send a work order carrying the failure data to the operation and maintenance node, so that the operation and maintenance node can perform a first failure analysis and processing based on the failure data and obtain a first failure analysis result;

[0008] Based on the first fault analysis result, a work order carrying the fault data is sent to the server supplier node, so that the server supplier node can perform a second fault analysis process based on the fault data to obtain a second fault analysis result.

[0009] The second fault analysis result is sent to the operation and maintenance node, and the fault handling process is executed based on the feedback from the operation and maintenance node regarding the second fault analysis result.

[0010] This application provides a fault handling device for a business system, including:

[0011] The first analysis module is used to monitor the servers that have failed in the business system and send a work order carrying the failure data to the operation and maintenance node so that the operation and maintenance node can perform a first failure analysis process based on the failure data and obtain a first failure analysis result.

[0012] The second analysis module is used to send the work order carrying the fault data to the server supplier node based on the first fault analysis result, so that the server supplier node can perform a second fault analysis process based on the fault data to obtain a second fault analysis result.

[0013] The review module is used to send the second fault analysis result to the operation and maintenance node, and execute the fault handling process based on the feedback result from the operation and maintenance node regarding the second fault analysis result.

[0014] In the above scheme, the first analysis module is further configured to: receive a fault reporting request sent by the business system, the fault reporting request being generated when the business system performs business perception processing and determines the server that has failed; or, receive a fault reporting request sent by the monitoring system, the fault reporting request being generated when the monitoring system performs monitoring processing on the business system and determines the server that has failed; and, in response to the fault reporting request, obtain the reporting information carried in the fault reporting request.

[0015] In the above scheme, the source of the fault reporting request of the work order can be either the business system or the monitoring system; the first analysis module is further configured to: when the source of the fault reporting request is the business system, send a work order carrying fault data to the operation and maintenance node, so that the operation and maintenance node can perform a first fault analysis process based on the fault data and obtain a first fault analysis result.

[0016] In the above scheme, the device further includes a third analysis module, used for: when the source of the fault report request is the monitoring system, sending a work order carrying the fault data to the server supplier node, so that the server supplier node can perform third fault analysis processing to obtain the third fault analysis result; sending the third fault analysis result to the operation and maintenance node, and executing the fault handling process according to the feedback result of the operation and maintenance node on the third fault analysis result.

[0017] In the above scheme, the first analysis module is further configured to: acquire fault data including fault logs and basic information; bind the fault data with the corresponding work order for the fault, and send the work order carrying the fault data to the operation and maintenance node.

[0018] In the above scheme, when the fault data includes a log link corresponding to the fault log and an information link to the basic information, the first analysis module is configured to: send a work order carrying the log link and the information link to the operation and maintenance node, so that the operation and maintenance node performs the following processing: in response to the data acquisition operation for the log link and the information link, acquire the fault log corresponding to the log link and the basic information corresponding to the information link, so as to present the fault log and the basic information; wherein, the fault log and the basic information are used to perform the first fault analysis processing; and perform the first fault analysis processing based on the fault log and the basic information to obtain the first fault analysis result.

[0019] In the above scheme, the first analysis module is further configured to: call the first neural network model and perform the following processing: perform a first fault analysis based on the fault log and the basic information to obtain the first fault analysis result; wherein, the training samples of the first neural network model include fault log samples and basic information samples of the fault samples, and the labeled data of the training samples includes the pre-labeled first fault analysis result of the fault samples.

[0020] In the above scheme, the second analysis module is further configured to: send the work order carrying the fault data to the server supplier node when at least one of the following conditions is met: the first fault analysis result indicates that the cause of the fault is unknown; the first fault analysis result indicates the cause of the fault, and the first fault analysis result meets the verification conditions.

[0021] In the above scheme, the second analysis module is further configured to: when the first fault analysis result characterizes the cause of the fault and the first fault analysis result does not need to be reviewed, execute the fault handling process corresponding to the fault type characterized by the first fault analysis result.

[0022] In the above scheme, the review conditions include at least one of the following conditions: the historical frequency of the fault is less than the frequency threshold; the importance of the fault as represented by the first fault analysis result is greater than the importance threshold; the confidence of the cause of the fault as represented by the first fault analysis result is less than the confidence threshold; and the fault analysis accuracy of the operation and maintenance node is less than the analysis accuracy threshold.

[0023] In the above scheme, the servers in the business system are provided by different suppliers; the second analysis module is further configured to: obtain the supplier identifier of the server that has failed; determine the supplier that provided the server that failed based on the supplier identifier, and send the work order carrying the failure data to the server supplier node of the supplier.

[0024] In the above scheme, before sending the work order carrying the fault data to the server supplier node of the supplier, the second analysis module is further configured to: when the number of server supplier nodes of the supplier is multiple, perform any one of the following operations: based on the work orders of each server supplier node of the supplier, determine the server supplier node for receiving the work order from the multiple server supplier nodes of the supplier; based on the historical work orders of each server supplier node of the supplier, determine the server supplier node for receiving the work order from the multiple server supplier nodes of the supplier.

[0025] In the above scheme, the second analysis module is further configured to: perform the following processing for each server supplier node of the supplier: obtain the assigned work orders of the server supplier node; perform completion time prediction processing on each of the assigned work orders of the server supplier node to obtain the predicted completion time of each assigned work order; accumulate the predicted completion times of multiple assigned work orders to obtain the cumulative completion time of the server supplier node; and determine the server supplier node with the smallest cumulative completion time as the server supplier node for receiving the work orders.

[0026] In the above scheme, the second analysis module is further configured to: perform the following processing for each server supplier node of the supplier: obtain the historical work orders and corresponding performance characteristics of the server supplier node; fuse the performance characteristics of each historical work order of the server supplier node to obtain the comprehensive performance characteristics of the server supplier node; determine the similarity between the comprehensive performance characteristics of the server supplier node and the performance characteristics of the work order; and determine the server supplier node with the highest similarity as the server supplier node for receiving the work order.

[0027] In the above scheme, the audit module is further configured to: execute a fault handling process corresponding to the fault type represented by the second fault analysis result.

[0028] In the above scheme, the second analysis module is further configured to: determine the urgency level of the work order corresponding to the first fault analysis result; query the sending method corresponding to the urgency level; and send the work order carrying the fault data to the server supplier node according to the sending method.

[0029] In the above scheme, before executing the fault handling process based on the feedback result of the operation and maintenance node regarding the second fault analysis result, the audit module is further configured to: receive the feedback result sent by the operation and maintenance node regarding the second fault analysis result; wherein, the feedback result is obtained by the operation and maintenance node performing any one of the following processes: in response to the feedback operation regarding the second fault analysis result, obtaining the feedback result submitted by the feedback operation regarding the second fault analysis result; calling a second neural network model to perform feedback result prediction processing on the second fault analysis result to obtain the feedback result regarding the second fault analysis result; wherein, the training samples of the second neural network model include fault samples, and the labeled data of the training samples includes the pre-labeled second fault analysis result of the fault samples.

[0030] This application provides an electronic device, including:

[0031] Memory, used to store executable instructions;

[0032] A processor, when executing executable instructions stored in the memory, implements the method provided in the embodiments of this application.

[0033] This application provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the fault handling method provided in this application.

[0034] The embodiments of this application have the following beneficial effects:

[0035] By streamlining and coordinating processes between operation and maintenance nodes and server provider nodes, the entire process from fault detection to analysis and handling is integrated, enabling automated fault handling of business systems. Dual fault analysis ensures the reliability of fault analysis, and because faults are distributed to multiple nodes for processing, the workload of each node can be saved, thereby improving the efficiency of fault handling. Attached Figure Description

[0036] Figure 1 This is a flowchart illustrating the fault handling method for a business system provided in an embodiment of this application;

[0037] Figure 2 This is a schematic diagram of the architecture of the fault handling system provided in the embodiments of this application;

[0038] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0039] Figures 4A-4D This is a flowchart illustrating the fault handling method for a business system provided in an embodiment of this application;

[0040] Figure 5 This is an overall architecture diagram of the fault handling system provided in the embodiments of this application;

[0041] Figure 6 This is a data retrieval diagram of the fault handling method provided in the embodiments of this application;

[0042] Figure 7 This is a schematic diagram of the real-time log acquisition page of the fault handling method provided in the embodiments of this application;

[0043] Figure 8 This is a schematic diagram of the real-time log acquisition page of the fault handling method provided in the embodiments of this application;

[0044] Figure 9 This is a schematic diagram of the out-of-band historical log acquisition page of the fault handling method provided in the embodiments of this application;

[0045] Figure 10 This is a schematic diagram of the in-band history log acquisition page of the fault handling method provided in the embodiments of this application;

[0046] Figure 11 This is a schematic diagram illustrating the basic information retrieval of the fault handling method provided in the embodiments of this application;

[0047] Figure 12 This is a schematic diagram illustrating the attachment upload method of the fault handling method provided in the embodiments of this application;

[0048] Figure 13 This is a schematic diagram of the email notification for the fault handling method provided in the embodiments of this application;

[0049] Figure 14 This is a schematic diagram of the server supplier node interface of the fault handling method provided in the embodiments of this application;

[0050] Figure 15 This is an image of the editing interface of the fault handling method provided in the embodiments of this application;

[0051] Figure 16 This is a schematic diagram illustrating the fault resolution method provided in the embodiments of this application;

[0052] Figure 17 This is a schematic diagram of the interface of the fault handling method provided in the embodiments of this application;

[0053] Figure 18 This is a schematic diagram of the processed work order page of the fault handling method provided in the embodiments of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0056] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0058] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0059] 1) In-band logs: In-band logs refer to the logs of the server operating system. In-band logs include, but are not limited to, dmesg logs, mcelog logs, smart logs, and crash logs.

[0060] 2) Out-of-band logs: Out-of-band logs refer to logs managed outside of the band, including SEL logs, SDR logs, register values, and vendor-managed out-of-band one-click logs, etc.

[0061] 3) Nodes: Nodes refer to the electronic devices used by various user roles, such as the electronic devices used by the operations and maintenance team, including the terminals or servers they use, etc.

[0062] In related technologies, when a business system malfunctions, it generates corresponding alarm information, which is received by the electronic devices used by the system's maintenance personnel. The alarms are then handled offline. However, since the cause of the malfunction may be unclear, this process requires a long processing time and significant manpower, reducing the efficiency of fault handling and affecting the stable operation of the business system.

[0063] In implementing the embodiments of this application, the applicant discovered that if the fault handling process is carried out through multi-party interaction, the processing efficiency and reliability can be improved. Multi-party interaction refers to the interaction between the operation and maintenance party, the business party, and the server provider. See [link to relevant documentation]. Figure 1 , Figure 1 This is a flowchart illustrating the fault handling method for a business system provided in this application embodiment. When an ambiguous fault occurs on the server of the business system, a corresponding ambiguous alarm is generated, and an ambiguous work order is created and sent to the business node for preliminary judgment (preliminary analysis of ambiguous faults). The logs are used to determine if a hardware problem exists. If it is determined to be a hardware-related hardware problem, the hardware problem is sent to the operations and maintenance node. The operations and maintenance node collects the server logs for preliminary analysis of the hardware fault. If the location of the hardware problem can be identified, the analysis results are directly fed back to the business node so that the business node can create a hardware fault work order. If the cause of the fault cannot be confirmed, the work order will be sent to the operations and maintenance node. Logs are aggregated and issues are reported to the server provider node for troubleshooting and optimization. After analysis, the server provider node returns the corresponding solution (result feedback) to the operations node. The operations node reviews the solution, and if it passes the review, it sends the solution (result feedback) back to the business node. Upon receiving the solution, the business node handles the fault accordingly. If the server provider node confirms it is a hardware issue, it creates a maintenance work order (i.e., creates a hardware fault order) to replace the faulty component. This interactive process effectively improves fault handling efficiency and ensures the stable operation of the business system.

[0064] In some embodiments, the operation and maintenance node can be specifically the operation and maintenance terminal, the business node can be specifically the business terminal, the server supplier node can be specifically the supplier terminal, and the server that generates an ambiguous alarm is the business server.

[0065] This application provides a fault handling method, apparatus, electronic device, and computer-readable storage medium for a business system, which can improve fault handling efficiency and ensure the normal operation of the business system. The following describes exemplary applications of the electronic device provided in this application, which can be implemented as a server. Specifically, an exemplary application of the electronic device as a diagnostic analysis server in an online diagnostic analysis system will be described below.

[0066] See Figure 2 , Figure 2This is a schematic diagram of the architecture of the fault handling system provided in this application embodiment. The terminal connects to the server through a network, which can be a wide area network, a local area network, or a combination of both. The fault handling system includes a business system 10, a monitoring system 20, an online diagnostic analysis system 30, a business terminal 400-1, an operation and maintenance terminal 400-2, and a server supplier terminal 400-3. The business system includes at least one business server 200-1, the monitoring system 20 includes at least one monitoring server 200-2, and the online diagnostic analysis system 30 includes at least one diagnostic analysis server 200-3. Any business server or a specific business server in the business system can detect faults. The online diagnostic analysis system 30 can provide services in the form of cloud services.

[0067] In some embodiments, after the business server 200-1 or the monitoring server 200-2 detects a fault in a server, it reports the fault to the diagnostic analysis server 200-3 to create a work order. The diagnostic analysis server 200-3 obtains the fault data and sends the fault data and the work order to the maintenance terminal 400-2 so that the maintenance terminal 400-2 can perform an initial analysis to determine the faulty component. In response to the maintenance terminal 400-2 determining the faulty component, the diagnostic analysis server 200-3 creates a corresponding faulty component replacement work order and sends it to the business terminal 400-1 to execute the component replacement process. In response to the maintenance terminal 400-2 not determining the faulty component, the diagnostic analysis server 200-3 sends the fault data along with the work order to... The solution is sent to server provider terminal 400-3 for a second analysis. If server provider terminal 400-3 determines a solution for the work order, it returns the solution to maintenance terminal 400-2 for review. Once approved, the solution is sent to business terminal 400-1 for execution. If server provider terminal 400-3 identifies a faulty component, it returns a solution for that component to maintenance terminal 400-2 for review. Once approved, diagnostic analysis server 200-3 creates a faulty component replacement work order and sends it to business terminal 400-1 to execute the component replacement process.

[0068] In some embodiments, the business server 200-1, monitoring server 200-2, and diagnostic analysis server 200-3 can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the invention.

[0069] See Figure 3 , Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 3 The diagnostic analysis server 200-3 shown includes at least one processor 210, memory 250, and at least one network interface 220. The various components in server 200 are coupled together via bus system 220. It is understood that bus system 240 is used to implement communication between these components. In addition to a data bus, bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 3 The general labeled all buses as Bus System 240.

[0070] Processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0071] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices physically located away from the processor 210.

[0072] The memory 250 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 250 described in this application embodiment is intended to include any suitable type of memory.

[0073] In some embodiments, memory 250 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0074] Operating system 251 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0075] The network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220, such as Bluetooth, WiFi, and Universal Serial Bus (USB).

[0076] In some embodiments, the fault handling device for the business system provided in this application can be implemented in software. Figure 3 A fault handling device 255 for a business system stored in memory 250 is shown. It can be software in the form of programs and plug-ins, including the following software modules: a first analysis module 2551, a second analysis module 2552, an audit module 2553, and a third analysis module 2554. These modules are logical and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.

[0077] The fault handling method for the business system provided in this application embodiment will be described in conjunction with the exemplary application and implementation of the server provided in the embodiments of this application.

[0078] See Figure 4A , Figure 4A This is a flowchart illustrating the fault handling method for a business system provided in this application embodiment, which will be combined with... Figure 4A The steps shown are explained below. The entity executing steps 101-103 is... Figure 2 The diagnostic analysis server 200-3.

[0079] In step 101, the monitoring system sends a work order carrying fault data to the operation and maintenance node for servers that have failed, so that the operation and maintenance node can perform the first fault analysis and processing based on the fault data and obtain the first fault analysis result.

[0080] In step 102, based on the first fault analysis result, a work order carrying fault data is sent to the server supplier node so that the server supplier node can perform a second fault analysis based on the fault data and obtain the second fault analysis result.

[0081] In step 103, the second fault analysis result is sent to the operation and maintenance node, and the fault handling process is executed based on the feedback from the operation and maintenance node regarding the second fault analysis result.

[0082] The following provides a detailed description of the fault handling method for the business system provided in the embodiments of this application. (See also...) Figure 4B , Figure 4B This is a flowchart illustrating the fault handling method for a business system provided in this application embodiment, which will be combined with... Figure 4B The steps shown are explained.

[0083] In step 201, the diagnostic analysis server monitors the servers that have malfunctioned in the business system.

[0084] As an example, a business system can be a server cluster consisting of multiple servers. The size of the server cluster is unlimited; for example, it could be a server cluster consisting of all the servers of an internet company, all the servers of a university, or a server cluster providing services for a specific product. The process of diagnosing and analyzing server failures within a business system can be implemented through the business system itself or a monitoring system.

[0085] In some embodiments, monitoring the server that has failed in the business system in step 201 can be achieved through the following technical solutions: receiving a fault reporting request sent by the business system through a fault monitoring interface, wherein the fault reporting request is generated when the business system performs business perception processing and determines that the server has failed; or, receiving a fault reporting request sent by a monitoring system, wherein the fault reporting request is generated when the monitoring system performs monitoring processing on the business system and determines that the server has failed; in response to the fault reporting request, obtaining the reporting information carried in the fault reporting request; wherein the fields of the reporting information include at least one of the following: server identifier, fault phenomenon, and urgency level.

[0086] As an example, the fault reporting request sent by the business system is received through the fault monitoring interface of the diagnostic analysis server. The fault monitoring interface is called by the business system. The fault reporting request is generated when the business system performs business awareness processing and determines the server that has failed. Business awareness processing includes functional testing and speed testing, etc. The fault reporting request sent by the monitoring system is received through the fault monitoring interface of the diagnostic analysis server. The fault monitoring interface is called by the monitoring system. The fault reporting request is generated when the monitoring system monitors the business system and determines the server that has failed. The monitoring system is composed of third-party servers independent of the business system. In response to the fault reporting request, the monitoring system obtains the reporting information carried in the fault reporting request. The reporting information includes at least one of the following fields: server identifier, fault phenomenon, and urgency level. The fields of the reporting information are specified by the calling protocol of the fault monitoring interface.

[0087] As an example, the fault monitoring interface protocol specifies several fields to improve subsequent analysis efficiency. The fields required by the fault monitoring interface are as follows: 1. Machine fixed asset or network protocol address information, used to identify a unique server and obtain corresponding logs and basic information; 2. Fault phenomenon, used to determine the scope and direction of fault analysis to avoid unnecessary resource waste; 3. Preliminary analysis results of business nodes, used to provide business-level analysis results for reference by operation and maintenance nodes and server supplier nodes (this field is optional); 4. Urgency level of the work order: used to determine the processing method of the work order, the timeliness of transmitting the work order to the server supplier node for processing, and the notification method.

[0088] In step 202, the diagnostic analysis server sends a work order carrying fault data to the operation and maintenance node.

[0089] In some embodiments, sending a work order carrying fault data to the operation and maintenance node in step 202 can be achieved through the following technical solution: obtaining fault data including fault logs and basic information; binding the fault data with the corresponding work order for the fault; and sending the work order carrying the fault data to the operation and maintenance node.

[0090] As an example, fault data can include fault logs (in the form of raw data) and basic information (in the form of raw data), and can also include links to fault logs and basic information. Typically, the amount of data in the basic information is small, so fault data can include links to fault logs and basic information (in the form of raw data).

[0091] In step 203, in response to the work order, the maintenance node performs the first fault analysis based on the fault data to obtain the first fault analysis result.

[0092] In some embodiments, when the fault data includes a log link corresponding to the fault log and an information link to the basic information, a work order carrying the fault data is sent to the operation and maintenance node, so that the operation and maintenance node (in response to the work order) performs a first fault analysis process based on the fault data to obtain a first fault analysis result. This can be achieved through the following technical solution: sending a work order carrying the log link and the information link to the operation and maintenance node, so that the operation and maintenance node performs the following processing: in response to the data acquisition operation for the log link and the information link, acquiring the fault log corresponding to the log link and the basic information corresponding to the information link, so as to present the fault log and the basic information; wherein, the fault log and the basic information are used to perform the first fault analysis process; performing the first fault analysis process based on the fault log and the basic information to obtain the first fault analysis result.

[0093] As an example, a work order carrying a log link and a basic information link is sent to the operations and maintenance node, so that the operations and maintenance node performs the following processing: In response to the data acquisition operation for the log link, the fault log corresponding to the log link is retrieved. When the fault log is a historical log, the data acquisition operation is a trigger operation for the download control. Historical logs are logs that are automatically collected when an anomaly occurs and are stored in a directory on the diagnostic analysis server. In response to the trigger operation for the download control of the link, the fault log corresponding to the link is retrieved from the diagnostic analysis server to the local machine. When the fault log is a real-time log, the data acquisition operation is a trigger operation for the collection control. Real-time logs are logs that are sent to a directory on the diagnostic analysis server only in response to this trigger operation. In response to the trigger operation for the download control, the fault log corresponding to the link is retrieved from the diagnostic analysis server to the local machine for presentation. In the steps of presenting fault logs and basic information, the basic information can be presented before obtaining the fault logs corresponding to the links, or after obtaining the fault logs. The presentation can be paginated or presented on the same page. For the information links of the basic information, the above technical solution is implemented in accordance with the operation of the log links for the logs.

[0094] In some embodiments, when the fault data includes fault logs (raw data) and basic information (raw data), a work order carrying the fault logs is sent to the operation and maintenance node so that the operation and maintenance node performs a first fault analysis process based on the fault data to obtain a first fault analysis result. This can be achieved through the following technical solution: sending a work order carrying the fault logs and basic information to the operation and maintenance node so that the operation and maintenance node performs the following processes: presenting the fault logs and basic information; wherein the fault logs and basic information are used for the first fault analysis process; performing the first fault analysis process to obtain a first fault analysis result.

[0095] As an example, since the work order directly contains fault logs and basic information, the fault logs and basic information can be directly presented when the work order is presented.

[0096] As an example, if the amount of raw data in the fault log exceeds the first data volume threshold, the link to the fault log will be carried on a work order for circulation; otherwise, it will be carried on a work order in the form of raw data. If the amount of raw data in the basic information exceeds the second data volume threshold, the link to the basic information will be carried on a work order for circulation; otherwise, it will be carried on a work order in the form of raw data.

[0097] In some embodiments, before responding to a data retrieval operation for a link, the user initiating the data retrieval operation is authenticated by the operations and maintenance node; once authentication is successful, it is determined whether to continue responding to the data retrieval operation for the link. To increase data confidentiality, authentication is required before retrieving various types of data, and the data retrieval operation can only be performed after successful authentication.

[0098] In some embodiments, the above-mentioned first fault analysis processing based on fault logs and basic information to obtain the first fault analysis result can be achieved by the following technical solution: calling a first neural network model and performing the following processing: performing the first fault analysis processing based on fault logs and basic information to obtain the first fault analysis result; wherein, the training samples of the first neural network model include fault log samples and basic information samples of the fault samples, and the labeled data of the training samples includes the pre-labeled first fault analysis result of the fault samples.

[0099] As an example, training samples are obtained, including fault log samples and basic information samples of fault samples. The labeled data of the training samples includes pre-labeled first fault analysis results of fault samples. The first neural network model is trained based on the training samples. The trained first neural network model is called to perform first fault analysis processing based on fault logs and basic information to obtain the first fault analysis result. The first fault analysis result includes the judgment of the fault type, such as whether the fault is related to hardware, etc. Obtaining the first fault analysis result through the first neural network model is beneficial to improving the efficiency of fault judgment.

[0100] In some embodiments, in response to an input operation for analysis results of a fault, a first fault analysis result is extracted from the input operation.

[0101] As an example, the initiator of the input operation is the user of the operation and maintenance node, such as the operation and maintenance user. The input operation can be an editing operation on the first fault analysis result or a voice input operation on the first fault analysis result. After receiving the input operation, the first fault analysis result is extracted from the input operation. The method of obtaining the first fault analysis result through the input operation is conducive to improving the accuracy of fault diagnosis.

[0102] As an example, the methods of obtaining the first fault analysis result through input operation and the methods of obtaining the first fault analysis result through the first neural network model can be parallel, that is, at least one of the two methods can be executed, or the two methods can be executed in sequence. If the first fault analysis result obtained through input operation does not characterize the cause of the fault, the first neural network model is used for analysis. First, the first fault analysis result is obtained through the first neural network model, and then the confirmation result or modification result of the first fault analysis result is obtained through input operation to update the first fault analysis result obtained by the first neural network model.

[0103] In step 204, the maintenance node returns the first fault analysis result to the diagnostic analysis server.

[0104] As an example, the operation and maintenance node returns the first fault analysis result to the diagnostic analysis server. Step 205 can be that the operation and maintenance node actively sends the first fault analysis result to the diagnostic analysis server, or it can send the first fault analysis result to the diagnostic analysis server in response to the polling request of the diagnostic analysis server.

[0105] In step 205, the diagnostic analysis server sends a work order carrying fault data to the server supplier node based on the first fault analysis result.

[0106] In some embodiments, see Figure 4C , Figure 4C The fault handling method for the business system provided in this application embodiment is as follows: In step 205, a work order carrying fault data is sent to the server supplier node based on the first fault analysis result, which can be achieved through step 2051.

[0107] In step 2051, a work order carrying fault data is sent to the server supplier node when at least one of the following conditions is met: the first fault analysis result indicates that the cause of the fault is unknown; the first fault analysis result indicates the cause of the fault, and the first fault analysis result meets the verification conditions.

[0108] As an example, the review conditions include at least one of the following: The historical frequency of the fault is less than a frequency threshold. For example, if the fault only occurred once, it indicates that the fault is not a common fault and therefore needs to be reviewed through the server provider node to ensure the reliability of fault handling; The first fault analysis result characterizes the importance of the fault. For example, if the severity of the fault is greater than a severity threshold, or if the fault will cause a large-scale impact, it also needs to be reviewed through the server provider node to ensure the reliability of fault handling; The confidence level of the fault cause characterized by the first fault analysis result is less than a confidence level threshold. For example, if the predicted probability (confidence level) corresponding to the fault cause obtained through the first neural network model is less than a confidence level threshold, it is necessary to review it through the server provider node to ensure the reliability of fault handling; The fault analysis accuracy of the operation and maintenance node is less than an analysis accuracy threshold. For example, if the historical operation data of the operation and maintenance node characterizes the accuracy of fault analysis as less than an analysis accuracy threshold, it needs to be reviewed through the server provider node to ensure the reliability of fault handling.

[0109] As an example, when the first fault analysis result indicates that the cause of the fault is unknown, for example, when the first fault analysis result determines the cause of the fault, or when the first fault analysis result indicates that the fault meets the conditions, it indicates that although the first fault analysis result has determined the cause of the fault, a review is still needed to improve the reliability of fault handling.

[0110] In some embodiments, when the first fault analysis result characterizes the cause of the fault and the first fault analysis result does not need to be reviewed, the corresponding type of fault handling process is executed according to the fault type characterized by the first fault analysis result.

[0111] As an example, when the first fault analysis result has a clear fault cause, such as the first fault analysis result indicating that a certain slot of the server has failed and needs to be replaced, the corresponding fault handling process is executed. For example, when the first fault analysis result is a hardware-related repair plan, the fault handling interface is called to enable the hardware maintenance node (or business node) to create the corresponding hardware and carry a repair work order with the repair plan, so as to execute the fault handling process of the corresponding repair plan according to the repair work order; when the first fault analysis result is a repair plan and fault cause unrelated to hardware, the repair plan and fault cause are sent to the business node so that the business node executes the fault handling process of the corresponding repair plan.

[0112] In some embodiments, see Figure 4D , Figure 4D The fault handling method for the business system provided in this application embodiment, in step 205, sends a work order carrying fault data to the server supplier node, which can be achieved through the following steps 2052-2053.

[0113] In step 2052, the supplier identifier of the server that failed is obtained.

[0114] In step 2053, the supplier of the faulty server is identified based on the supplier identifier, and a work order carrying fault data is sent to the supplier's server supplier node.

[0115] As an example, the servers in the business system are provided by different suppliers. For example, the servers in the business system are provided by supplier A and supplier B. Each supplier has a unique supplier identifier. For example, when the diagnostic analysis server sends a work order carrying fault data to the server supplier node, it needs to first determine the server supplier node that will receive the work order. If the server that failed was provided by supplier A, then the work order carrying fault data is sent to the server supplier node of supplier A.

[0116] In some embodiments, before step 2053 sends the work order carrying fault data to the supplier's server supplier nodes, when there are multiple server supplier nodes of the supplier, any one of the following operations is performed: based on the work orders of each server supplier node of the supplier, determine the server supplier node for receiving the work order from the multiple server supplier nodes of the supplier; based on the historical work orders of each server supplier node of the supplier, determine the server supplier node for receiving the work order from the multiple server supplier nodes of the supplier.

[0117] As an example, if supplier A has 3 server supplier nodes, the above technical solution needs to be implemented to determine 1 server supplier node from supplier A's 3 server supplier nodes as the server supplier node that receives the work order.

[0118] In some embodiments, the above-mentioned determination of the server supplier node for receiving the work orders from multiple server supplier nodes based on the work orders of each server supplier node of the supplier can be achieved through the following technical solution: performing the following processing for each server supplier node of the supplier: obtaining the assigned work orders of the server supplier node; performing completion time prediction processing on each assigned work order of the server supplier node to obtain the predicted completion time of each assigned work order; accumulating the predicted completion times of multiple assigned work orders to obtain the cumulative completion time of the server supplier node; and determining the server supplier node with the minimum cumulative completion time as the server supplier node for receiving the work orders.

[0119] As an example, the assigned work orders for server supplier node a are obtained. The completion time of each assigned work order for server supplier node a is predicted. For example, if server supplier node a currently has 4 assigned work orders to complete, the completion time of each assigned work order is predicted. This prediction is based on the difficulty of the assigned work orders or their historical completion times. The predicted completion times of multiple assigned work orders are accumulated to obtain the cumulative completion time of the server supplier node. This is equivalent to predicting the time required for server supplier node a to complete the current assigned work order. Following the same steps, the time required for two other server supplier nodes to complete the current assigned work order is determined. The server supplier node with the smallest cumulative completion time among these three server supplier nodes is selected as the server supplier node to receive work orders, thereby accelerating the processing speed of work orders.

[0120] In some embodiments, the above-mentioned determination of the server supplier node for receiving the work order from multiple server supplier nodes based on the historical work orders of each server supplier node of the supplier can be achieved through the following technical solution: Perform the following processing for each server supplier node of the supplier: obtain the historical work orders of the server supplier node and the corresponding performance characteristics; fuse the performance characteristics of each historical work order of the server supplier node to obtain the comprehensive performance characteristics of the server supplier node; determine the similarity between the comprehensive performance characteristics of the server supplier node and the performance characteristics of the work order; determine the server supplier node with the highest similarity as the server supplier node for receiving the work order.

[0121] As an example, the historical work orders and corresponding performance characteristics of server supplier node a are obtained. For instance, server supplier node a has 3 historical work orders, each with corresponding performance characteristics. The 3 performance characteristics corresponding to these 3 historical work orders are merged to obtain the comprehensive performance characteristics of server supplier node a. The similarity between the comprehensive performance characteristics of server supplier node a and the performance characteristics of the work orders to be assigned is determined. Following the above steps, the comprehensive performance characteristics of the other two server supplier nodes are determined. The server supplier node with the highest similarity among these 3 server supplier nodes is determined as the server supplier node to receive work orders. For example, if a certain server supplier node frequently handles work orders with a certain fault performance, then it is preferred to use this server supplier node to handle work orders with similar fault performance, thereby improving the rationality of work order allocation and thus improving the processing efficiency of work orders.

[0122] In some embodiments, after sending a work order carrying fault data to the server provider node, a timer is started from the time the work order is sent to the server provider node; when the timer exceeds the processing time limit of the work order, a reminder message for the work order is sent to the server provider node.

[0123] In some embodiments, after sending a work order carrying fault data to the operation and maintenance node, a timer can be started from the time the work order is sent to the operation and maintenance node, and when the timer exceeds the processing time limit of the work order, a reminder message for the work order can be sent to the operation and maintenance node.

[0124] As an example, since each work order has its own corresponding urgency level, it is necessary to set a processing time limit for the work orders. The timer starts after the work order is sent to the server provider node / operation node. When the timer exceeds the processing time limit of the work order, a reminder message for the work order is sent to the server provider node / operation node to improve the processing efficiency of the work orders.

[0125] In some embodiments, the step 205 of sending a work order carrying fault data to the server supplier node based on the first fault analysis result can be achieved through the following technical solution: determining the urgency level of the work order corresponding to the first fault analysis result; querying the sending method corresponding to the urgency level; and sending the work order carrying fault data to the server supplier node according to the sending method.

[0126] As an example, if the urgency level of the fault is normal, the server supplier node will be notified via email to check the work order. If the urgency level of the fault is urgent, the server supplier node will be notified via telephone to check the work order, and the corresponding work order will appear on the graphical interface of the corresponding diagnostic analysis server of the server supplier node.

[0127] In step 206, in response to the work order, the server supplier node performs a second fault analysis based on the fault data to obtain the second fault analysis result.

[0128] As an example, the implementation of step 206 can refer to the implementation of step 203, the only difference being that the execution subject is changed from the operation and maintenance node to the server provider node.

[0129] In step 207, the server provider node returns the second fault analysis result to the diagnostic analysis server.

[0130] As an example, the implementation of step 207 can refer to the implementation of step 204, the only difference being that the execution subject is changed from the operation and maintenance node to the server provider node.

[0131] In step 208, the diagnostic analysis server sends the second fault analysis result to the operation and maintenance node.

[0132] As an example, when the diagnostic analysis server sends the second fault analysis result to the operation and maintenance node, it can send only the second fault analysis result to update the work order received by the operation and maintenance node in step 202, or it can directly send a new work order carrying the second fault analysis result to the operation and maintenance node to replace the work order received by the operation and maintenance node in step 202.

[0133] In step 209, the maintenance node determines the feedback result for the second fault analysis result.

[0134] As an example, the feedback result is obtained by the operation and maintenance node performing any of the following processes: in response to the feedback operation for the second fault analysis result, obtain the feedback result for the second fault analysis result submitted by the feedback operation; call the second neural network model to perform feedback result prediction processing on the second fault analysis result to obtain the feedback result for the second fault analysis result; wherein, the training samples of the second neural network model include fault samples, and the labeled data of the training samples include the pre-labeled second fault analysis result of the fault samples.

[0135] As an example, training samples are obtained, including fault log samples and basic information samples of fault samples. The labeled data of the training samples includes pre-labeled second fault analysis results of fault samples. The second neural network model is trained based on the training samples. The trained second neural network model is called to perform feedback result prediction processing on the second fault analysis results to obtain feedback results for the second fault analysis results. The feedback results include whether the review is approved, whether the review is not approved, and the specific reasons, etc.

[0136] As an example, the initiator of the feedback operation is the user of the operation and maintenance node, such as the operation and maintenance user. The feedback operation can be an editing operation on the feedback result or a voice operation on the feedback result. After receiving the feedback operation, the feedback result is extracted from the feedback operation. The method of obtaining the feedback result through the feedback operation is conducive to improving the accuracy of fault diagnosis.

[0137] As an example, the methods of obtaining feedback results through feedback operations and through the second neural network model can be parallel, that is, at least one of the two methods can be executed, or the two methods can be executed in sequence. For example, first obtain the feedback result through the second neural network model, and then obtain the confirmation result or modification result of the feedback result through feedback operations to update the feedback result obtained by the second neural network model, or determine the specific method of obtaining the feedback result according to the pre-configuration.

[0138] In step 210, the maintenance node sends the feedback results to the diagnostic analysis server.

[0139] As an example, the implementation of step 210 can refer to the implementation of step 204.

[0140] In step 211, the diagnostic analysis server executes the fault handling process based on the feedback results from the operation and maintenance nodes regarding the second fault analysis results.

[0141] In some embodiments, the fault handling process in step 211 based on the feedback result of the operation and maintenance node on the second fault analysis result can be implemented by the following technical solution: according to the fault type represented by the second fault analysis result, execute the fault handling process corresponding to the fault type.

[0142] As an example, if the second fault analysis result indicates that a certain slot of the server has failed and needs to be replaced, then the corresponding fault handling process is executed. For example, if the second fault analysis result is a hardware repair plan, the fault handling interface is called to enable the hardware maintenance node (or business node) to create the corresponding hardware and carry a repair work order with the repair plan, so as to execute the fault handling process of the corresponding repair plan according to the repair work order; if the second fault analysis result is a repair plan and fault cause unrelated to hardware, the repair plan and fault cause are sent to the business node so that the business node executes the fault handling process of the corresponding repair plan.

[0143] In some embodiments, the source of the fault reporting request for the work order can be either the business system or the monitoring system. Steps 202 and 203, sending a work order carrying fault data to the operations and maintenance node, responding to the work order, and the operations and maintenance node performing a first fault analysis based on the fault data to obtain a first fault analysis result, can be implemented using the following technical solution: When the source of the fault reporting request is the business system, a work order carrying fault data is sent to the operations and maintenance node so that the operations and maintenance node performs a first fault analysis based on the fault data to obtain a first fault analysis result. When the source of the fault reporting request is the monitoring system, a work order carrying fault data is sent to the server supplier node so that the server supplier node performs a third fault analysis to obtain a third fault analysis result; the third fault analysis result is sent to the operations and maintenance node, and the fault handling process is executed based on the feedback from the operations and maintenance node regarding the third fault analysis result.

[0144] As an example, since fault reporting requests in work orders have two sources, and different sources can represent different fault types—for example, the source could be either a business system (business system awareness) or a monitoring system (monitoring)—a work order carrying fault data is sent to the operations and maintenance (O&M) node. The O&M node responds to the work order and performs a first fault analysis based on the fault data to obtain a first fault analysis result. This can be achieved through the following technical solution: When the fault reporting request originates from the business system, a work order carrying fault data is sent to the O&M node, enabling the O&M node to perform a first fault analysis based on the fault data, obtain a first fault analysis result, and continue executing subsequent processes based on the first fault analysis result. However, when the fault reporting request originates from the business system, a work order carrying fault data is sent to the O&M node, allowing the O&M node to perform a first fault analysis based on the fault data, obtain a first fault analysis result, and continue executing subsequent processes based on the first fault analysis result. When the source of the fault report request is the monitoring system, and the fault reported in the fault report request is only related to hardware, a work order carrying fault data is sent to the server supplier node. This allows the server supplier node to perform third fault analysis and obtain the third fault analysis result. In other words, instead of sending a work order to the operation and maintenance node for first fault analysis, the work order is sent directly to the server supplier node for third fault analysis. Only the third fault analysis result needs to be sent to the operation and maintenance node. Based on the feedback from the operation and maintenance node regarding the third fault analysis result, the fault handling process is executed. Through the above implementation method, the monitoring characteristics of the monitoring system are used to pre-determine the type of fault, thereby simplifying the handling process and improving the handling efficiency.

[0145] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.

[0146] In some embodiments, the fault handling method provided in this application is applied to fault handling scenarios of business systems in large-scale social application software. See [link to relevant documentation]. Figure 5 , Figure 5This is an overall architecture diagram of the fault handling system provided in this application embodiment. The online diagnostic analysis system (diagnostic analysis server) provides an interface to the outside world, providing a way for business systems and monitoring systems to access alarms. Business systems and monitoring systems (e.g., hardware monitoring systems) call the interface to create a work order for a specific fault. After creating a work order for a specific fault, the online diagnostic analysis system pulls logs and basic information related to the faulty server, such as in-band and out-of-band logs, machine model, maintenance records, etc. The pulled information is summarized and sent along with the work order to the operation and maintenance node (electronic device used by server operation and maintenance personnel) so that the operation and maintenance node can perform preliminary analysis and judgment of the fault. If the operation and maintenance node can directly identify the faulty component, the online diagnostic analysis system creates a faulty component replacement work order and sends it to the business node so that the fault handling system can execute the corresponding faulty component replacement process. If the operation and maintenance node cannot directly identify the faulty component, but can only confirm that the current fault belongs to a hardware-related difficult problem, the operation and maintenance node will transfer the pulled information to the maintenance node. After receiving the information, the system, along with the work order, is synchronized to the server supplier node (used by server R&D personnel) through the online diagnostic analysis system. The operations and maintenance node also notifies the supplier via email, telephone, or other means. When the server supplier node determines a hardware-independent solution to the fault, it synchronizes the solution to the operations and maintenance node through the online diagnostic analysis system for review. If the review is successful, the solution is returned to the business node for implementation. If the review fails, the server supplier node is notified again for further analysis of the fault. If the server supplier node confirms that the problem is hardware-related and requires repair, it provides a solution that includes the faulty component. This solution is synchronized to the operations and maintenance node through the online diagnostic analysis system for review. If the review is successful, the operations and maintenance node directly requests the online diagnostic analysis system to create a corresponding faulty component replacement work order via an interface and sends it to the business node for execution of the corresponding faulty component replacement process through the fault handling system.

[0147] In some embodiments, the protocol of the fault monitoring interface for creating work orders specifies multiple fields to improve subsequent analysis efficiency. The fields required by the fault monitoring interface are as follows: 1. Machine fixed asset or network protocol address information, used to identify a unique server and obtain corresponding logs and basic information; 2. Fault phenomenon, used to determine the scope and direction of analysis to avoid unnecessary resource waste; 3. Preliminary analysis results of business nodes: used to provide business-level analysis results for reference by operation and maintenance nodes and server supplier nodes (this field is optional); 4. Urgency level of the work order: used to determine the processing method of the work order, the timeliness of transmission to the server supplier node for processing, and the notification method.

[0148] In some embodiments, see Figure 6 , Figure 6 This is a data retrieval diagram of the fault handling method provided in this application embodiment. When the online diagnostic analysis system retrieves server logs, it retrieves historical logs and real-time collected logs. Historical logs refer to data automatically collected when server anomalies are detected, including in-band and out-of-band one-click logs. After collection, the logs are stored in the corresponding directory according to a specific format for retrieval by the online diagnostic analysis system. Real-time collected logs are collected on demand. Upon receiving a trigger operation for a control, the log collection logic is invoked to collect one-click logs, SDR logs, SEL logs, and message logs in real time, and a batch export function is provided. See also Figure 7 , Figure 7 This is a schematic diagram of the real-time log acquisition page of the fault handling method provided in this application embodiment. The real-time log acquisition page 603 presents different types of real-time in-band logs, real-time out-of-band one-click logs, real-time SDR logs, and real-time SEL logs. Each type of log has a collection control and a download control. Triggering the collection control 601 collects the corresponding logs, and triggering the download control 602 downloads the collected logs to the local machine. The batch collection control 604 triggers the collection of multiple logs upon response to a triggering operation, and the batch export control 605 triggers the download of multiple logs upon response to a triggering operation. See also... Figure 8 , Figure 8 This is a schematic diagram of the real-time log acquisition page of the fault handling method provided in this application embodiment. The real-time log acquisition page 703 presents different types of real-time in-band logs, real-time out-of-band one-click logs, real-time SDR logs, and real-time SEL logs. Each type of log has a collection control and a download control. Triggering the collection control 701 collects the corresponding logs, and triggering the download control 702 downloads the collected logs to the local machine. The batch collection control 704 triggers the collection of multiple logs after a trigger operation, and the batch export control 705 triggers the download of multiple logs after a trigger operation. When collection is successful, a success message is displayed at the corresponding log entry. See also... Figure 9 , Figure 9 This is a schematic diagram of the out-of-band historical log acquisition page of the fault handling method provided in this application embodiment. The out-of-band historical log acquisition page 803 displays multiple out-of-band historical logs, each with a download control (since historical logs are automatically collected when an exception occurs, there is no need to initiate the collection process again). Triggering the download control 802 downloads the corresponding out-of-band historical log to the local machine. The out-of-band historical log acquisition page 803 also includes a collection control to prevent re-collection when historical collection fails. See [link to relevant documentation]. Figure 10 , Figure 10This is a schematic diagram of the in-band history log acquisition page of the fault handling method provided in this application embodiment. The in-band history log acquisition page 903 displays in-band history logs, which have a download control. Triggering the download control 902 downloads the corresponding in-band history logs to the local machine. Triggering the out-of-band log switching control 901 switches to the out-of-band log. Figure 9 The page showing the out-of-band historical log retrieval.

[0149] In some embodiments, see Figure 11 , Figure 11 This is a schematic diagram of the basic information retrieval method provided in the embodiments of this application. The basic information of the server includes maintenance records, model information and configuration information, firmware version information (FW), etc. The operation and maintenance node and the server supplier node can view all the historical records of the server (including the number of failures, hardware replacements and handling suggestions) through the maintenance records so as to have a comprehensive understanding of the failure. The model configuration information refers to the basic information related to the server, such as hardware configuration, model version number, data center information, etc. The firmware version information mainly refers to the basic input / output system version, business management (BMC) version and hard disk version.

[0150] In some embodiments, see Figure 12 , Figure 12 This is a schematic diagram illustrating the attachment upload function of the fault handling method provided in this application embodiment. For servers of older models, the business management does not support the automatic one-click data collection function, and manual download is required through an out-of-band webpage. In addition, the analysis reports and fault analysis reports of the server supplier nodes also need to be promptly fed back to the operation and maintenance nodes. Therefore, the online diagnostic analysis system provides an attachment upload function. Figure 12 An attachment management page 1101 is shown, which displays a file 1102, an attachment selection control 1103, and an upload attachment control 1104. In response to a trigger operation on the attachment selection control 1103, a file selection process is executed, and in response to a trigger operation on the upload attachment control 1104, a file upload process is executed.

[0151] In some embodiments, if a business system sends a fault report work order (request) through an interface, the online diagnostic analysis system, upon receiving the fault report work order, will first assign a work order to the operations and maintenance node. The operations and maintenance node will then conduct a preliminary analysis of the faulty server using logs and fault data. Only if it is determined to be related to a hardware fault will the analysis be escalated to the server supplier node. If the operations and maintenance node confirms that it is not a hardware problem, it will inform the business node to conduct the analysis. If a monitoring system sends a fault report work order through an interface, the online diagnostic analysis system will directly assign a work order to the server supplier node. Since the monitoring system performs fault monitoring based on hardware logs, it will only issue a fault report work order through the monitoring system if it is confirmed to be related to hardware.

[0152] In some embodiments, see Figure 13 , Figure 13 This is a schematic diagram of the email notification method for the fault handling method provided in this application embodiment. When the operation and maintenance node confirms that the fault is a difficult problem (i.e., it cannot be directly confirmed that it is a hardware problem or not a hardware problem), the online diagnostic analysis system will create a corresponding work order and send it to the server supplier node for further analysis. If the urgency level of the fault is normal, the server supplier node will be notified via email. If the urgency level of the fault is urgent, the server supplier node will be notified via telephone, and the corresponding work order will appear on the system graphical interface of the server supplier node. See [link to relevant documentation]. Figure 14 , Figure 14 This is a schematic diagram of the server supplier node interface of the fault handling method provided in the embodiments of this application. Figure 14 Multiple work orders are displayed, but only those belonging to the same supplier are shown because the data display between different suppliers is isolated.

[0153] In some embodiments, see Figure 15 , Figure 15 This is an image of the editing interface of the fault handling method provided in this application embodiment. The server supplier node will conduct a comprehensive analysis based on the fault phenomenon and logs (fault data). After determining the solution, it will edit and record it on the work order editing page 1401. The editing record includes 10 fields such as whether the fault is clear, the analysis process, and the fault classification. After the server supplier node completes the editing process of the above fields, it responds to the trigger operation of the submission control 1402 and sends the solution back to the operation and maintenance node through the online diagnostic analysis system.

[0154] In some embodiments, the rationality and operability of the solution are evaluated by the operation and maintenance node. When the online analysis and diagnosis system is first applied to a server cluster, there are many problems encountered. As the operation and maintenance node and the server supplier node accumulate experience through long-term interaction, the above-mentioned analysis process for the solution can be realized through intelligent review. Based on fault data (e.g., machine configuration, fault phenomenon and maintenance history), the rationality of the solution can be intelligently judged.

[0155] In some embodiments, see Figure 16 , Figure 16 This is a schematic diagram of the fault handling method provided in the embodiment of this application. If the solution provided by the server supplier node is deemed reasonable and hardware repair is required, the online analysis and diagnosis system will automatically connect to the fault handling process and pass the information of the components and slots that need to be replaced to the fault handling process. The fault will be handled according to the normal hardware fault handling process. In addition, if the solution does not require hardware replacement, the solution (e.g., version upgrade or stress test) only needs to be fed back to the business node for processing.

[0156] In some embodiments, see Figure 17 , Figure 17 This is a schematic diagram of the interface of the fault handling method provided in the embodiments of this application. The system used in the fault handling method is an online analysis and diagnosis system, which is typically logged into and used by server provider nodes and operation and maintenance nodes. Figure 17 The system's work order page is shown. Operations nodes can view work orders for all server provider nodes. Information between server provider nodes is isolated. All fields in the work order can be used as query conditions to search for work orders. The work order displays basic information, including server model, version, fault description, preliminary analysis results from the business side, and historical work order information. The work order provides an entry point for log display, including historical logs before and after the anomaly, and real-time collected logs. Out-of-band logs include out-of-band one-click logs, SDR logs, and SEL logs; in-band logs include message logs, dmesg logs, mcelog logs, and crashdump logs. See also... Figure 18 , Figure 18 This is a schematic diagram of the processed work order page of the fault handling method provided in the embodiments of this application. Figure 18 All processed work orders are displayed. Figure 17 After the work order in the system is completed, it will be transferred from... Figure 17 The page will be redirected to Figure 18The system displays the information on the page. The online diagnostic analysis system can automatically collect in-band and out-of-band logs of faulty servers, retrieve basic server information, and provide corresponding pages for display and download. It also notifies relevant personnel (operations nodes and / or server provider nodes) via email and telephone for handling and follow-up, and intelligently processes the analysis results.

[0157] In some embodiments, the fault handling method provided in this application has the following advantages: 1. Labor saving: During the internal testing phase, the system processed 2080 work orders within 3 months, averaging 23 orders per day, saving 4.6 manpower. The manpower saving increases with the number of servers. 2. Efficiency improvement: Before the internal testing, each work order took an average of 5-7 days to process, mainly due to log collection and unordered collaborative processing. After the system participated in the internal testing, each process node had a corresponding processing time, with email and telephone notifications. The systems automatically connected to implement functions such as order initiation, log capture, and work order reminders, improving automation to reduce manpower input and thus improving processing efficiency. The average work order processing time was shortened to 17 hours. 3. Process standardization: The fault analysis process and optimization solutions form a closed loop, helping to clarify the case accumulation process and optimize fault reporting capabilities. The analysis process and work order creation records are all searchable, and the processing results and progress of any node are transparent. Data is stored on the server, providing strong traceability.

[0158] The following description continues to illustrate the exemplary structure of the fault handling device 255 for the business system provided in this application embodiment as a software module. In some embodiments, such as... Figure 3 As shown, the software modules in the fault handling device 255 of the business system stored in the memory 250 may include a first analysis module 2551, used to monitor servers that have failed in the business system, and send work orders carrying fault data to the operation and maintenance nodes so that the operation and maintenance nodes can perform first fault analysis processing based on the fault data to obtain a first fault analysis result; a second analysis module 2552, used to send work orders carrying fault data to the server supplier nodes based on the first fault analysis result so that the server supplier nodes can perform second fault analysis processing based on the fault data to obtain a second fault analysis result; and an audit module 2553, used to send the second fault analysis result to the operation and maintenance nodes and execute the fault handling process based on the feedback from the operation and maintenance nodes regarding the second fault analysis result.

[0159] In some embodiments, the first analysis module 2551 is further configured to: receive a fault reporting request sent by a business system, wherein the fault reporting request is generated when the business system performs business perception processing and determines that a server has failed; or, receive a fault reporting request sent by a monitoring system, wherein the fault reporting request is generated when the monitoring system performs monitoring processing on the business system and determines that a server has failed; and, in response to the fault reporting request, obtain the reporting information carried in the fault reporting request; wherein the fields of the reporting information include at least one of the following: server identifier, fault phenomenon, and urgency level.

[0160] In some embodiments, the source of the fault reporting request for the work order is either the business system or the monitoring system; the first analysis module 2551 is further configured to: when the source of the fault reporting request is the business system, send a work order carrying fault data to the operation and maintenance node, so that the operation and maintenance node can perform a first fault analysis process based on the fault data and obtain a first fault analysis result.

[0161] In some embodiments, the apparatus further includes a third analysis module 2554, configured to: when the source of the fault report request is the monitoring system, send a work order carrying fault data to the server supplier node so that the server supplier node performs third fault analysis processing and obtains third fault analysis results; send the third fault analysis results to the operation and maintenance node, and execute the fault handling process based on the feedback results from the operation and maintenance node regarding the third fault analysis results.

[0162] In some embodiments, the first analysis module 2551 is further configured to: acquire fault data including fault logs and basic information; bind the fault data with the corresponding fault work order; and send the work order carrying the fault data to the operation and maintenance node.

[0163] In some embodiments, when the fault data includes a log link corresponding to the fault log and an information link to the basic information; the first analysis module 2551 is further configured to: send a work order carrying the log link and the information link to the operation and maintenance node, so that the operation and maintenance node performs the following processing: in response to the data acquisition operation for the log link and the information link, acquire the fault log corresponding to the log link and the basic information corresponding to the information link, so as to present the fault log and the basic information; wherein, the fault log and the basic information are used to perform a first fault analysis processing; and perform the first fault analysis processing based on the fault log and the basic information to obtain a first fault analysis result.

[0164] In some embodiments, the first analysis module 2551 is further configured to: invoke the first neural network model and perform the following processing: perform a first fault analysis based on the fault log and basic information to obtain a first fault analysis result; wherein the training samples of the first neural network model include fault log samples and basic information samples of the fault samples, and the labeled data of the training samples includes the pre-labeled first fault analysis result of the fault samples.

[0165] In some embodiments, the second analysis module 2552 is further configured to: send a work order carrying fault data to the server supplier node when at least one of the following conditions is met: the first fault analysis result indicates that the cause of the fault is unknown; the first fault analysis result indicates the cause of the fault, and the first fault analysis result meets the review conditions.

[0166] In some embodiments, the second analysis module 2552 is further configured to: when the first fault analysis result characterizes the cause of the fault and the first fault analysis result does not need to be reviewed, execute a fault handling process of the corresponding type according to the fault type characterized by the first fault analysis result.

[0167] In some embodiments, the review conditions include at least one of the following conditions: the historical frequency of the fault is less than the frequency threshold; the importance of the fault as represented by the first fault analysis result is greater than the importance threshold; the confidence level of the cause of the fault as represented by the first fault analysis result is less than the confidence threshold; and the fault analysis accuracy of the operation and maintenance node is less than the analysis accuracy threshold.

[0168] In some embodiments, the servers in the business system are provided by different suppliers; the second analysis module 2552 is further configured to: obtain the supplier identifier of the server that has failed; determine the supplier that provided the server that failed based on the supplier identifier, and send a work order carrying the failure data to the server supplier node of the supplier.

[0169] In some embodiments, before sending a work order carrying fault data to the supplier's server supplier nodes, the second analysis module 2552 is further configured to: when the number of supplier's server supplier nodes is multiple, perform any one of the following operations: based on the work orders of each supplier's server supplier nodes, determine the server supplier node for receiving the work order from the multiple supplier's server supplier nodes; based on the historical work orders of each supplier's server supplier nodes, determine the server supplier node for receiving the work order from the multiple supplier's server supplier nodes.

[0170] In some embodiments, the second analysis module 2552 is further configured to: perform the following processing for each server supplier node of the supplier: obtain the assigned work orders of the server supplier node; perform completion time prediction processing on each assigned work order of the server supplier node to obtain the predicted completion time of each assigned work order; accumulate the predicted completion times of multiple assigned work orders to obtain the cumulative completion time of the server supplier node; and determine the server supplier node with the minimum cumulative completion time as the server supplier node for receiving work orders.

[0171] In some embodiments, the second analysis module 2552 is further configured to: perform the following processing for each server supplier node of the supplier: obtain historical work orders and corresponding performance characteristics of the server supplier node; fuse the performance characteristics of each historical work order of the server supplier node to obtain the comprehensive performance characteristics of the server supplier node; determine the similarity between the comprehensive performance characteristics of the server supplier node and the performance characteristics of the work order; and determine the server supplier node with the highest similarity as the server supplier node for receiving work orders.

[0172] In some embodiments, the audit module 2553 is further configured to: execute a fault handling process corresponding to the fault type represented by the second fault analysis result.

[0173] In some embodiments, the second analysis module 2552 is further configured to: determine the urgency level of the work order corresponding to the first fault analysis result; query the sending method corresponding to the urgency level; and send the work order carrying fault data to the server supplier node according to the sending method.

[0174] In some embodiments, before executing the fault handling process based on the feedback result of the operation and maintenance node regarding the second fault analysis result, the audit module 2553 is further configured to: receive the feedback result sent by the operation and maintenance node regarding the second fault analysis result; wherein the feedback result is obtained by the operation and maintenance node performing any of the following processes: in response to the feedback operation regarding the second fault analysis result, obtaining the feedback result submitted by the feedback operation regarding the second fault analysis result; calling the second neural network model to perform feedback result prediction processing on the second fault analysis result to obtain the feedback result regarding the second fault analysis result; wherein the training samples of the second neural network model include fault samples, and the labeled data of the training samples includes the pre-labeled second fault analysis result of the fault samples.

[0175] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device (e.g., a computer device) reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the fault handling method for the business system described in this application.

[0176] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to execute a fault-handling method for a business system provided in this application. For example, ... Figures 4A-4D The fault handling method of the business system is shown.

[0177] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0178] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0179] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0180] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0181] In summary, the embodiments of this application enable collaborative processing of server faults between the operation and maintenance nodes and the server supplier nodes. Since the faults are obtained through multiple fault analyses, the reliability of fault handling can be improved. Because the faults are distributed among multiple nodes for processing, the workload of each node can be reduced, thereby improving the efficiency and reliability of fault handling.

[0182] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A fault handling method for a business system, characterized in that, The method is executed by an online diagnostic analysis system within a fault handling system. The fault handling system includes a business system, a monitoring system, the online diagnostic analysis system, business nodes, operation and maintenance nodes, and server provider nodes. The method includes: Monitor the servers that malfunction in the business system and bind the malfunction data with the corresponding work orders for the malfunction. When the fault report request originates from the business system, a work order carrying log links and information links from the fault data is sent to the operation and maintenance node. So that the operation and maintenance node performs a first fault analysis based on the fault logs corresponding to the log links and the basic information corresponding to the information links, and obtains a first fault analysis result; When the first fault analysis result indicates that the cause of the fault is unknown or / and the first fault analysis result indicates the cause of the fault, and the first fault analysis result meets the verification conditions, a work order carrying the fault data is sent to the server supplier node used to receive the work order, wherein... The server supplier node used to receive the work order has the minimum cumulative completion time of its assigned work orders, or the highest similarity to the performance characteristics of the work order, so that... The server supplier node performs a second fault analysis process based on the fault data to obtain a second fault analysis result. The second fault analysis result is sent to the operation and maintenance node, and the fault handling process is executed based on the feedback from the operation and maintenance node regarding the second fault analysis result.

2. The method according to claim 1, characterized in that, The method further includes: When the source of the fault report request is the monitoring system, a work order carrying the fault data is sent to the server supplier node so that the server supplier node can perform third fault analysis and obtain the third fault analysis result. The third fault analysis result is sent to the operation and maintenance node, and the fault handling process is executed based on the feedback result from the operation and maintenance node regarding the third fault analysis result.

3. The method according to claim 1, characterized in that, Sending a work order carrying log links and information links from the fault data to the maintenance node includes: Obtain fault data, including fault logs and basic information; The fault data is bound to the corresponding work order for the fault, and a work order carrying the log link and information link in the fault data is sent to the operation and maintenance node.

4. The method according to claim 3, characterized in that, The operation and maintenance node performs a first fault analysis based on the fault logs corresponding to the log links and the basic information corresponding to the information links, and obtains a first fault analysis result, including: In response to the data acquisition operation for the log link and the information link, the fault log corresponding to the log link and the basic information corresponding to the information link are acquired to present the fault log and the basic information.

5. The method according to claim 1, characterized in that, The servers in the business system are provided by different suppliers; Sending the work order carrying the fault data to the server supplier node used to receive the work order includes: Obtain the supplier identifier of the server that experienced the failure; Based on the supplier identifier, the supplier that provided the faulty server is identified, and the work order carrying the fault data is sent to the server supplier node of the supplier.

6. The method according to claim 5, characterized in that, Before sending the work order carrying the fault data to the supplier's server supplier node, the method further includes: When the number of server provider nodes of the provider is multiple, perform any one of the following operations: Based on the work orders of each server supplier node of the supplier, determine the server supplier node to receive the work orders from among the server supplier nodes of the plurality of suppliers. Based on the historical work orders of each server supplier node of the supplier, a server supplier node for receiving the work orders is determined from among the server supplier nodes of the multiple suppliers.

7. The method according to claim 6, characterized in that, The process of determining the server supplier node to receive the work order from among multiple server supplier nodes of the suppliers, based on the work order for each server supplier node of the suppliers, includes: Perform the following processing for each server provider node of the aforementioned provider: Obtain the assigned work orders of the server supplier node; For each of the assigned work orders of the server supplier node, a completion time prediction process is performed to obtain the predicted completion time of each of the assigned work orders. The predicted completion times of multiple assigned work orders are summed to obtain the cumulative completion time of the server supplier node. The server supplier node with the minimum cumulative completion time is selected as the server supplier node to receive the work order.

8. The method according to claim 6, characterized in that, The process of determining the server supplier node for receiving the work order from among multiple server supplier nodes based on historical work orders of each server supplier node of the supplier includes: Perform the following processing for each server provider node of the aforementioned provider: Obtain the historical work orders and corresponding performance characteristics of the server supplier node; The performance characteristics of each historical work order of the server supplier node are fused to obtain the comprehensive performance characteristics of the server supplier node. Determine the similarity between the overall performance characteristics of the server supplier node and the performance characteristics of the work order. The server supplier node with the highest similarity is identified as the server supplier node to receive the work order.

9. The method according to claim 1, characterized in that, The step of executing a fault handling process based on the feedback from the maintenance node regarding the second fault analysis result includes: Based on the fault type characterized by the second fault analysis result, execute the fault handling process corresponding to the fault type.

10. The method according to claim 1, characterized in that, Sending the work order carrying the fault data to the server supplier node used to receive the work order includes: Determine the urgency level of the work order corresponding to the first fault analysis result; Query the sending method corresponding to the stated urgency level; The work order carrying the fault data is sent to the server supplier node according to the sending method described.

11. The method according to claim 1, characterized in that, Before executing the fault handling process based on the feedback from the maintenance node regarding the second fault analysis result, the method further includes: Receive feedback results from the operation and maintenance node regarding the second fault analysis result; The feedback result is obtained by the operation and maintenance node performing any of the following processes: In response to a feedback operation on the second fault analysis result, obtain the feedback result submitted by the feedback operation on the second fault analysis result; The second neural network model is invoked to perform feedback result prediction processing on the second fault analysis result, and a feedback result is obtained for the second fault analysis result; The training samples of the second neural network model include fault samples, and the labeled data of the training samples includes the pre-labeled second fault analysis results of the fault samples.

12. A fault handling device for a business system, characterized in that, The device is applied to the online diagnostic analysis system within a fault handling system. The fault handling system includes a business system, a monitoring system, the online diagnostic analysis system, business nodes, operation and maintenance nodes, and server provider nodes. The device includes: The first analysis module is used to monitor servers that have failed in the business system, and bind the failure data with the corresponding work orders for the failure. When the failure report request originates from the business system, a work order carrying the log link and information link in the failure data is sent to the operation and maintenance node, so that the operation and maintenance node can perform the first failure analysis based on the failure log corresponding to the log link and the basic information corresponding to the information link, and obtain the first failure analysis result. The second analysis module is used to send a work order carrying the fault data to the server supplier node that receives the work order when the first fault analysis result indicates that the cause of the fault is unknown or / and the first fault analysis result indicates the cause of the fault, and the first fault analysis result meets the verification conditions. The work order is provided in the following case: the cumulative completion time of the assigned work orders of the server supplier node that receives the work order is the smallest, or the similarity with the performance characteristics of the work order is the largest, so that the server supplier node performs a second fault analysis process based on the fault data to obtain a second fault analysis result. The review module is used to send the second fault analysis result to the operation and maintenance node, and execute the fault handling process based on the feedback result from the operation and maintenance node regarding the second fault analysis result.

13. The apparatus according to claim 12, characterized in that, The first analysis module is further configured to receive a fault reporting request sent by the business system, the fault reporting request being generated when the business system performs business perception processing and determines the server that has failed; or, receive a fault reporting request sent by the monitoring system, the fault reporting request being generated when the monitoring system performs monitoring processing on the business system and determines the server that has failed; and in response to the fault reporting request, obtain the reporting information carried in the fault reporting request.

14. The apparatus according to claim 13, characterized in that, The device further includes: The third analysis module is used to send a work order carrying the fault data to the server supplier node when the fault reporting request originates from the monitoring system, so that the server supplier node can perform third fault analysis processing to obtain the third fault analysis result; send the third fault analysis result to the operation and maintenance node; and execute the fault handling process based on the feedback result from the operation and maintenance node regarding the third fault analysis result.

15. An electronic device, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the fault handling method of the business system according to any one of claims 1 to 11.

16. A computer-readable storage medium, characterized in that, It stores executable instructions for use by a processor to implement the fault handling method of the business system according to any one of claims 1 to 11.

17. A computer program product comprising: A computer-executable instruction, characterized in that, when executed by a processor, the computer-executable instruction implements the method described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Monitoring method and system for data center

    CN103684817A

  • Solid state disk monitoring method, device and equipment

    CN109741786A