Interface fault root cause analysis method and device, equipment and storage medium
By calculating the correlation score between the abnormal alarm information and the associated alarm information in the interface operation and maintenance information, the root cause of interface failure is determined, which solves the problem of low interface troubleshooting efficiency and improves the maintenance effect.
Patent Information
- Application Number
- CN202510303039.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
In the existing technology, the interface troubleshooting efficiency is low, and the lack of effective root cause analysis methods can cause repeated failures and poor maintenance results.
By obtaining the operation and maintenance information of the interface, including abnormal alarm information and associated alarm information, the correlation score between the abnormal alarm information and associated alarm information is calculated, and the root cause of the interface is determined.
It improves the convenience and accuracy of the analysis of root causes of interface failures, reduces the possibility of repeated failures, and improves the maintenance effect of interfaces.
Smart Images

Figure CN120216241A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of interface maintenance, and in particular, to a method, device, equipment, and storage medium for analyzing the root cause of interface failures. Background Art
[0002] In the related art, in the case of detecting abnormal alarm information of a system, it is generally necessary to check multiple interfaces of the system one by one to determine the faulty interface, which is likely to result in poor efficiency in troubleshooting interface failures. Moreover, due to the lack of analysis of the root cause of interface failures, it is easy to have a situation where the root cause of the interface failure is not found, leading to repeated occurrence of failures, and thus resulting in poor maintenance effect of the interface.
[0003] For example, for a digital financial system that processes financial business and a digital medical system that processes medical business, both rely on interface calls to maintain the operation of the digital financial system and the digital medical system. For example, the digital financial system needs to call an interface to obtain financial data related to financial business, such as insurance claim information, insurance renewal information, etc. The digital medical system needs to call an interface to obtain medical data related to medical business, such as medical examination reports, pharmacies, etc. If the maintenance effect of the interfaces in the digital financial system and the digital medical system is poor, it is likely to lead to poor operation effects of the digital financial system and the digital medical system.
[0004] Based on this, it is urgent to improve the convenience of analyzing the root cause of interface failures to enhance the maintenance effect of interfaces. Summary of the Invention
[0005] The main purpose of the present application is to provide a method, device, equipment, and storage medium for analyzing the root cause of interface failures, aiming to improve the convenience of analyzing the root cause of interface failures to enhance the maintenance effect of interfaces.
[0006] In the first aspect, the present application provides a method for analyzing the root cause of interface failures, including the following steps:
[0007] Obtain the operation and maintenance information corresponding to the interface;
[0008] When the operation and maintenance information includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, determine the correlation degree score between the abnormal alarm information and the associated alarm information according to the operation and maintenance information;
[0009] Determine the root cause of the interface failure according to the correlation degree score and the associated alarm information corresponding to the correlation degree score.
[0010] In the second aspect, the present application further provides an apparatus for analyzing the root cause of interface failures, and the apparatus for analyzing the root cause of failures includes:
[0011] An operation and maintenance information acquisition module, configured to acquire operation and maintenance information corresponding to an interface;
[0012] A scoring module, configured to determine a correlation degree score between the exception alarm information and the associated alarm information according to the operation and maintenance information when the operation and maintenance information includes the exception alarm information of the interface and the associated alarm information of the operation support environment associated with the interface;
[0013] A fault root cause determination module, configured to determine the fault root cause of the interface according to the correlation degree score and the associated alarm information corresponding to the correlation degree score.
[0014] In a third aspect, the present application further provides a computer device, where the computer device includes a memory and a processor;
[0015] The memory is configured to store a computer program;
[0016] The processor is configured to execute the computer program and implement the steps of the method for analyzing the fault root cause of the interface as described above when executing the computer program.
[0017] In a fourth aspect, the present application further provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for analyzing the fault root cause of the interface as described above are implemented.
[0018] The present application provides a method, apparatus, device, and storage medium for analyzing the fault root cause of an interface. The method for analyzing the fault root cause of an interface includes: acquiring operation and maintenance information corresponding to the interface; when the operation and maintenance information includes exception alarm information of the interface and associated alarm information of the operation support environment associated with the interface, determining a correlation degree score between the exception alarm information and the associated alarm information according to the operation and maintenance information; and determining the fault root cause of the interface according to the correlation degree score and the associated alarm information corresponding to the correlation degree score, so as to improve the maintenance effect of the interface.
[0019] When the operation and maintenance information corresponding to the interface is acquired, if the operation and maintenance information includes the exception alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, the correlation degree score between the exception alarm information and the associated alarm information can be determined according to the operation and maintenance information. The correlation degree score can be used to judge the possibility that the associated alarm information corresponding to the correlation degree score is the fault root cause of the interface, and then determine the fault root cause of the interface, which is beneficial to improving the convenience of analyzing the fault root cause of the interface. The fault root cause of the interface can be used as the basis for maintaining the interface, so as to trace the fault of the interface for maintenance, and then reduce the possibility of repeated occurrence of the fault of the interface, which is beneficial to improving the maintenance effect of the interface.
[0020] For example, there may be a digital financial system for processing financial services and a digital medical system for processing medical services. When processing financial services such as insurance claims and insurance premium renewals through the digital financial system, the digital financial system can obtain relevant information of the insured through the call interface to determine the corresponding insurance claim information, insurance premium renewal information, etc. When a fault occurs in the corresponding interface of the digital financial system, the root cause of the interface fault is determined through the root cause analysis method of the interface, and the root cause of the interface fault can be used as the basis for maintaining the interface, so that the data financial system can call the maintained interface to improve the operation effect of the digital financial system. When processing medical services such as obtaining medical examination reports and transmitting prescriptions through the digital medical system, the digital medical system can obtain the medical examination reports of relevant personnel through the call interface and convey the prescribed prescriptions to the pharmacy through the call interface. Correspondingly, when a fault occurs in the corresponding interface of the digital medical system, the root cause of the interface fault is determined through the root cause analysis method of the interface, and the root cause of the interface fault can be used as the basis for maintaining the interface, so that the data medical system can call the maintained interface to improve the operation effect of the digital medical system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a schematic diagram of an application environment of the root cause analysis method of the interface in an embodiment of the present application;
[0023] Figure 2 It is a schematic flowchart of the root cause analysis method of the interface in an embodiment of the present application;
[0024] Figure 3 It is a schematic diagram of the operation and maintenance information corresponding to the interface involved in an embodiment of the present application;
[0025] Figure 4 is Figure 2 a schematic flowchart of a specific implementation manner of step S20 in
[0026] Figure 5 It is another schematic flowchart of the root cause analysis method of the interface in an embodiment of the present application;
[0027] Figure 6 It is a schematic diagram of the topology diagram corresponding to the interface involved in an embodiment of the present application;
[0028] Figure 7It is another schematic flowchart of the method for analyzing the root cause of faults of an interface in an embodiment of the present application;
[0029] Figure 8 It is a schematic diagram of the recommended list of root causes of faults of an interface involved in an embodiment of the present application;
[0030] Figure 9 It is a schematic structural diagram of a device for analyzing the root cause of faults of an interface in an embodiment of the present application;
[0031] Figure 10 It is a schematic structural diagram of a computer device in an embodiment of the present application;
[0032] Figure 11 It is another schematic structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.
[0034] The method for analyzing the root cause of faults of an interface provided in the embodiments of the present application can be applied, for example, in Figure 1In the application environment, the client communicates with the server through the network. The server can obtain the operation and maintenance information corresponding to the interface through the client. When the operation and maintenance information includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, the server determines the correlation score between the abnormal alarm information and the associated alarm information according to the operation and maintenance information; according to the correlation score and the associated alarm information corresponding to the correlation score, the server determines the root cause of the interface failure, and feeds back the root cause of the interface failure to the client. In this application, it can be used in a digital financial system for processing financial services and a digital medical system for processing medical services. When processing financial services such as insurance claim settlement services and insurance renewal services through the digital financial system, the digital financial system can obtain relevant information of the insured through the interface to determine the corresponding insurance claim settlement information, insurance renewal information, etc. When a failure occurs in the corresponding interface of the digital financial system, the root cause of the interface failure is determined through the root cause analysis method of the interface, and the root cause of the interface failure can be used as a basis for maintaining the interface, so that the data financial system can call the maintained interface to improve the operation effect of the digital financial system. When processing medical services such as obtaining medical examination reports and transmitting prescriptions through the digital medical system, the digital medical system can obtain medical examination reports of relevant personnel through the interface, and convey the prescribed prescriptions to the pharmacy through the interface. Correspondingly, when a failure occurs in the corresponding interface of the digital medical system, the root cause of the interface failure is determined through the root cause analysis method of the interface, and the root cause of the interface failure can be used as a basis for maintaining the interface, so that the data medical system can call the maintained interface to improve the operation effect of the digital medical system. Among them, the client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The following describes this application in detail through specific embodiments.
[0035] Please refer to Figure 2 as shown in Figure 2 FIG. [X] is a schematic flowchart of a method for analyzing the root cause of an interface failure provided by an embodiment of the present application, including the following steps:
[0036] S10: Obtain the operation and maintenance information corresponding to the interface.
[0037] The method for analyzing the root cause of faults of the interface provided by this application can be applied to data processing systems in various application scenarios. The data processing system can include a digital financial system, a digital medical system, etc., without limitation here. The data processing system can achieve the interaction between users and computer devices by calling the interface. In some embodiments, the data processing system can include different types of components such as services, databases, container hosts, interfaces, etc. The number of each of the different types of components included in the data processing system can include one or more, without limitation here. Among them, during the process of calling the interface, since it depends on services, databases, container hosts, etc., the services, databases, and container hosts can be determined as the operating support environment associated with the interface.
[0038] For example, for a digital financial system, when a user processes corresponding financial services through the digital financial system, the digital financial system can obtain the information required to process the corresponding financial services by calling the interface to assist the user in processing financial services.
[0039] For example, for a digital medical system, when a user processes corresponding medical services through the digital medical system, the digital medical system can obtain the information required to process the corresponding medical services by calling the interface to assist the user in processing medical services.
[0040] During the process of calling the interface, operation and maintenance information corresponding to the interface can be formed. The operation and maintenance information can be used to monitor whether the interface fails and to determine the root cause of the interface fault.
[0041] In some embodiments, the service side can be used to monitor the faults of the interface. For example, the service side can obtain the operation and maintenance information corresponding to the interface in real time or at intervals, and based on the operation and maintenance information corresponding to the interface, determine whether the interface fails. Correspondingly, in the case where it is monitored that the interface fails, the root cause of the interface fault can be determined based on the operation and maintenance information corresponding to the interface. Of course, this is not limited to this. The data processing system itself can also monitor the faults of the interface, and the client where the data processing system is located can also obtain the operation and maintenance information corresponding to the interface in real time or at intervals, without limitation here.
[0042] Exemplarily, the operation and maintenance information corresponding to the interface may include the interface's exception warning information, the associated warning information of the operation support environment associated with the interface, the interface trace cache, the data related to the interface provided by the Application Performance Management (APM) tool, the cache stack data corresponding to the interface, etc., which is not limited herein. The exception warning information can be used to indicate the notification or log information automatically generated and sent by the data processing system when an error or exception occurs in the interface during the operation of the data processing system, so as to monitor whether the interface fails. The associated warning information can be used to indicate the notification or log information automatically generated and sent by the data processing system when an error or exception occurs in the operation support environment associated with the interface during the operation of the data processing system, so as to monitor whether the operation support environment fails. Both the exception warning information and the associated warning information can be used to monitor the health of the data processing system, diagnose problems, and respond to faults in a timely manner. The interface trace cache can be used to store the trace information of interface calls, so as to quickly retrieve and analyze the interface call situation when needed, helping to troubleshoot problems or optimize performance. The data related to the interface provided by the APM tool can be used to indicate the performance metrics of the interface, such as the response time, throughput, error count, error rate, etc. of the interface, so as to monitor and optimize the performance of the interface. The cache stack data can be used to indicate the usage of the interface cache, such as the hit rate, invalidation policy, etc., to assist in optimizing the interface cache mechanism. In an exemplary embodiment, the operation and maintenance information corresponding to the interface is as Figure 3 shown, the exception warning information of the interface can be reflected by bolding the operation and maintenance information corresponding to the interface. Of course, it is not limited to this. For example, it can also be highlighted in red, indicated by an arrow, etc., which is not limited herein. The operation and maintenance information corresponding to the interface may include the operation and maintenance analysis information corresponding to the interface, such as the interface address, the name of the high-time-consuming method, the parameters of the high-time-consuming method, the interface time-consuming, the proportion of the method time-consuming, etc. Of course, the exception warning information of the interface and the operation and maintenance information corresponding to the interface are not limited to this, which is not limited herein.
[0043] When the operation and maintenance information corresponding to the interface is obtained, the operation and maintenance information corresponding to the interface can be used to monitor whether the interface fails, so as to, when it is monitored that the interface fails, serve as a basis for analyzing the root cause of the interface failure, and then determine the root cause of the interface failure.
[0044] S20: When the operation and maintenance information includes the exception warning information of the interface and the associated warning information of the operation support environment associated with the interface, determine the correlation score between the exception warning information and the associated warning information according to the operation and maintenance information.
[0045] For example, when the operation and maintenance information includes the abnormal alarm information of an interface, it can be determined that the interface has a fault, and then the root cause analysis of the fault of the interface needs to be carried out. When the operation and maintenance information includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, since the operation of the interface depends on the operation support environment of the interface, it can be speculated that there may be an association between the associated alarm information and the root cause of the interface fault. Correspondingly, since the operation support environment of the interface can include services associated with the interface, databases associated with the interface, container hosts associated with the interface, service release information associated with the interface, etc., the associated alarm information can also include the associated alarm information of different types of operation support environments respectively. The root cause of the interface fault may be caused by at least one piece of associated alarm information, and then it is necessary to select the associated alarm information associated with the root cause of the interface fault from all the associated alarm information. Based on this, in the process of carrying out the root cause analysis of the interface fault, when the obtained operation and maintenance information includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, the correlation score between the abnormal alarm information and the associated alarm information can be determined according to the operation and maintenance information. The correlation score can be used to evaluate the possibility of the abnormal alarm information of the interface being triggered by the associated alarm information, so as to determine the root cause of the interface fault. In the case where the correlation score between the abnormal alarm information and the associated alarm information is higher, the possibility that the fault of the interface indicated by the abnormal alarm information is caused by the fault of the operation support environment indicated by the associated alarm information is higher, and the correlation score can be used to determine the root cause of the interface fault subsequently.
[0046] Taking the data processing system including the digital financial system as an example. When processing the insurance claim settlement business through the digital financial system, the digital financial system can obtain the relevant information of the insured person by calling the interface, such as the personal information of the insured person, policy information, claim settlement rules, etc. When the operation and maintenance information corresponding to the interface includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, the abnormal alarm information is used to indicate, for example, that the interface call is abnormal, resulting in the failure to obtain all the relevant information of the insured person, or the failure to obtain some of the information of the insured person. The associated alarm information is used to indicate, for example, the container host alarm associated with the interface, the database alarm associated with the interface, etc. Based on this, it can be speculated that the abnormal interface call is caused by the container host alarm associated with the interface and the database alarm associated with the interface, and then the correlation score between the abnormal interface call and the container host alarm and the database alarm can be determined according to the operation and maintenance information corresponding to the interface. For example, the abnormal interface call may be caused by at least one of the container host alarm and the database alarm, and the correlation score can be used as the basis for determining the root cause of the interface fault.
[0047] Of course, it is not limited to this. Taking a data processing system including a digital medical system as an example. When processing medical services through the digital medical system, the digital medical system can obtain relevant medical information of the patient through the call interface, such as the personal information of the patient, the medical examination report of the patient, the prescription opened by the medical staff for the patient, etc. When the operation and maintenance information corresponding to the interface includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, the abnormal alarm information is used to indicate, for example, that the interface call is abnormal, resulting in the failure to obtain all the relevant information of the patient, or the failure to obtain some information of the insured person. The associated alarm information is used to indicate, for example, the service alarm associated with the interface, the service release information alarm associated with the interface, etc. Based on this, it can be inferred that the interface call abnormality is caused by the service alarm and the service release information alarm associated with the interface. Furthermore, according to the operation and maintenance information corresponding to the interface, the correlation degree scores between the interface call abnormality and the service alarm and the service release information alarm can be determined. The correlation degree score can be used as a basis for determining the root cause of the interface failure. For example, if the interface call abnormality is caused by at least one of the service alarm and the service release information alarm, the correlation degree score can be used as a basis for determining the root cause of the interface failure. There is no limitation here.
[0048] In some embodiments, as Figure 4 shown, in step S20, that is, according to the operation and maintenance information, determining the correlation degree score between the abnormal alarm information and the associated alarm information includes the following steps:
[0049] S21: Extract features from the operation and maintenance information to obtain the correlation features between the abnormal alarm information and the associated alarm information; the correlation features include at least one of the alarm time correlation degree, the alarm type correlation degree, and the alarm level correlation degree.
[0050] S22: Based on the preset scoring model, according to the correlation features between the abnormal alarm information and the associated alarm information, determine the correlation degree score between the abnormal alarm information and the associated alarm information; the preset scoring model is trained according to the historical abnormal alarm information, the historical associated alarm information, and the preset correlation degree scores between the historical abnormal alarm information and the historical associated alarm information corresponding to each of multiple interfaces.
[0051] For step S21, when the operation and maintenance information corresponding to the interface is obtained, and the operation and maintenance information includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, the operation and maintenance information can be subjected to feature extraction to obtain the association features between the abnormal alarm information and the associated alarm information. The association features include at least one of the alarm time correlation degree, the alarm type correlation degree, and the alarm level correlation degree. For example, feature extraction can be performed on the interface trace cache included in the operation and maintenance information, the data related to the interface provided by the APM tool, and the cache stack data to obtain the association features between the abnormal alarm information and the associated alarm information, which is not limited here.
[0052] The alarm time correlation degree can be used to indicate the correlation degree between the alarm time of the abnormal alarm information and the alarm time of the associated alarm information. When the time difference between the alarm time of the abnormal alarm information and the alarm time of the associated alarm information is smaller, the alarm time correlation degree between the abnormal alarm information and the associated alarm information is higher. Taking the associated alarm information including the service alarm associated with the interface and the database alarm associated with the interface as an example. When the alarm time of the abnormal alarm information is time 1, the alarm time of the service alarm is time 2, and the alarm time of the database alarm is time 3, if the time difference between time 1 and time 2 is less than the time difference between time 1 and time 3, it can be determined that the alarm time correlation degree between the abnormal alarm information and the service alarm is higher than the alarm time correlation degree between the abnormal alarm information and the database alarm. Of course, the abnormal alarm information, the associated alarm information, and the alarm times of the abnormal alarm information and the associated alarm information are not limited to this, which is not limited here.
[0053] The correlation degree of alarm types can be used to indicate the correlation degree between the alarm type of abnormal alarm information and the alarm type of associated alarm information. When the similarity between the alarm type of abnormal alarm information and the alarm type of associated alarm information is higher, the correlation degree of alarm types between the abnormal alarm information and the associated alarm information is higher. For example, the alarm type can include any one of timeout and error. Taking the abnormal alarm information including the container host alarm associated with the interface and the service release information alarm associated with the interface as an example. When the abnormal alarm information indicates that the interface call times out, it can be determined that the alarm type of the abnormal alarm information is timeout. When the container host alarm indicates that the container host response times out, it can be determined that the alarm type of the container host alarm is timeout. When the service release information alarm indicates that the service release information fails to be released or a rollback occurs, it can be determined that the alarm type of the service release information alarm is error. Thus, it can be determined that the similarity between the alarm type of the abnormal alarm information and the alarm type of the container host alarm is higher than the similarity between the alarm type of the abnormal alarm information and the alarm type of the service release information alarm, and then it can be determined that the correlation degree of alarm types between the abnormal alarm information and the container service alarm is higher than the correlation degree of alarm types between the abnormal alarm information and the service release information alarm. Of course, the abnormal alarm information, the associated alarm information, and the alarm types of the abnormal alarm information and the associated alarm information are not limited to this, and no restrictions are imposed here.
[0054] The correlation degree of alarm levels can be used to indicate the correlation degree between the alarm level of abnormal alarm information and the alarm level of associated alarm information. When the level difference between the alarm level of abnormal alarm information and the alarm level of associated alarm information is smaller, the correlation degree of alarm levels between the abnormal alarm information and the associated alarm information is higher. Taking the associated alarm information including the database alarm associated with the interface and the container host alarm associated with the interface as an example. When the alarm level of the abnormal alarm information is the first level, the alarm level of the database alarm is the fifth level, and the alarm level of the container host alarm is the second level, it can be determined that the level difference between the alarm level of the abnormal alarm information and the alarm level of the database alarm is greater than the alarm level of the abnormal alarm information and the alarm level of the container host alarm, and then it can be determined that the correlation degree of alarm levels between the abnormal alarm information and the database alarm is lower than the correlation degree of alarm levels between the abnormal alarm information and the container host alarm. Of course, the abnormal alarm information, the associated alarm information, and the alarm levels of the abnormal alarm information and the associated alarm information are not limited to this, and no restrictions are imposed here.
[0055] In an exemplary embodiment, operation and maintenance information can be input into a preset feature extraction model, and the preset feature extraction model can have the function of extracting features from the operation and maintenance information. Through the preset feature extraction model, features of the operation and maintenance information can be extracted to obtain the correlation features between the abnormal alarm information and the associated alarm information. Of course, this is not limited thereto, and no limitation is made herein.
[0056] For step S22, in the case of obtaining the correlation features between the abnormal alarm information and the associated alarm information, the correlation features can be input into a preset scoring model, and the preset scoring model can determine the correlation score between the abnormal alarm information and the associated alarm information according to the correlation features.
[0057] Exemplarily, the preset scoring model is trained according to the historical abnormal alarm information, historical associated alarm information corresponding to multiple interfaces, and the preset correlation score between the historical abnormal alarm information and the historical associated alarm information.
[0058] For example, the preset correlation score can be determined according to the historical maintenance information indicated by the log information corresponding to the data processing system and associated with at least one of the historical abnormal alarm information and the historical associated alarm information, and the preset scoring rule. For example, in the case where after determining that the historical associated alarm information is successfully maintained according to the historical maintenance information, the historical abnormal alarm information is also successfully maintained accordingly, the preset correlation score between the historical abnormal alarm information and the historical associated alarm information can be determined as the first preset correlation score according to the preset scoring rule. Another example is that in the case where after determining that the historical associated alarm information is successfully maintained according to the historical maintenance information, the historical abnormal alarm information is not successfully maintained, the preset correlation score between the historical abnormal alarm information and the historical associated alarm information can be determined as the second preset correlation score according to the preset scoring rule. Among them, the first preset correlation score is greater than the second preset correlation score.
[0059] Of course, it is not limited to this. For example, the preset correlation score can be determined by manually scoring based on the correlation score between the historical abnormal alarm information and the historical associated alarm information by the operation and maintenance personnel. For example, the operation and maintenance personnel can score the correlation between the historical abnormal alarm information and the historical associated alarm information to obtain the corresponding preset correlation score. In some exemplary embodiments, in order to reduce the difference in manual scoring caused by the differences in personal experience of different operation and maintenance personnel, there may be multiple initial correlation scores between the same historical abnormal alarm information and the same historical associated alarm information. The multiple initial correlation scores can be obtained by manually scoring according to the correlation scores of different operation and maintenance personnel for the correlation between the same historical abnormal alarm information and the same historical associated alarm information. Accordingly, according to the preset correlation score selection rule, the corresponding preset correlation score can be determined based on the multiple initial correlation scores. For example, the corresponding preset correlation score can be determined according to any one of the average value, mode, weighted average value, median, and trimmed mean value corresponding to the multiple initial correlation scores.
[0060] When the historical abnormal alarm information, the historical associated alarm information, and the preset correlation score between the historical abnormal alarm information and the historical associated alarm information corresponding to each of the multiple interfaces are obtained, the historical abnormal alarm information, the historical associated alarm information, and the corresponding preset correlation score can be input into the initial scoring model to train the initial scoring model to obtain the preset scoring model. In an exemplary embodiment, during the process of training the initial scoring model, the correlation feature between the historical abnormal alarm information and the historical associated alarm information can be input into the initial scoring model. The initial scoring model can adopt the Gradient Boosting Decision Tree (GBDT) algorithm and combine the correlation feature between the historical abnormal alarm information and the historical associated alarm information to determine the predicted correlation score between the historical abnormal alarm information and the historical associated alarm information. The correlation feature between the historical abnormal alarm information and the historical associated alarm information is determined based on the historical operation and maintenance information jointly corresponding to the historical abnormal alarm information and the historical associated alarm information. Accordingly, the model parameters of the initial scoring model can be adjusted by combining the predicted correlation score and the preset correlation score, and then the preset scoring model can be trained.
[0061] When the preset scoring model is trained, the correlation feature between the abnormal alarm information and the associated alarm information can be input into the preset scoring model so that the preset scoring model outputs the correlation score between the abnormal alarm information and the associated alarm information. The correlation score between the abnormal alarm information and the associated alarm information can be used to determine the root cause of the interface failure.
[0062] When the operation and maintenance information includes the abnormal alarm information of the interface and the associated alarm information, the correlation score between the abnormal alarm information and the associated alarm information can be determined based on the operation and maintenance information, which is conducive to improving the convenience of determining the correlation score. Based on the correlation score, the root cause of the interface failure can be determined, which is conducive to improving the convenience of analyzing the root cause of the interface failure and the convenience of determining the root cause of the interface failure, and further conducive to improving the subsequent maintenance effect of the interface.
[0063] S30: Determine the root cause of the interface failure according to the correlation score and the associated alarm information corresponding to the correlation score.
[0064] For example, when the correlation score and the associated alarm information corresponding to the correlation score are obtained, the data processing system can determine the root cause of the interface failure from the associated alarm information corresponding to the correlation score according to the correlation score.
[0065] In an exemplary embodiment, the root cause of the interface failure can be determined according to the associated alarm information with the highest correlation score. Taking the associated alarm information of the operation support environment associated with the interface including database alarm associated with the interface, container host alarm associated with the interface, and service alarm associated with the interface as an example. When the correlation score corresponding to the database alarm is greater than the correlation score corresponding to the container host alarm, and the correlation score corresponding to the container host alarm is greater than the service alarm score, it can be determined that the root cause of the interface failure is due to the database alarm.
[0066] In another exemplary embodiment, the root cause of the interface failure can be determined according to all the associated alarm information whose correlation score is greater than or equal to the preset correlation score threshold. Taking the associated alarm information of the operation support environment associated with the interface including database alarm associated with the interface, container host alarm associated with the interface, and service alarm associated with the interface as an example. When the correlation score corresponding to the database alarm and the correlation score corresponding to the container host alarm are both greater than or equal to the preset correlation score threshold, and the correlation score corresponding to the service alarm is less than the preset correlation score threshold, it can be determined that the root cause of the interface failure is due to the database alarm and the container host alarm.
[0067] In yet another exemplary embodiment, when the interface has multiple root causes of failure, the data processing system can also determine the correlation between the multiple root causes of failure to determine the final root cause of the interface failure. For example, when the root causes of the interface failure include database alarm and container host alarm, it can be determined that the final root cause of the interface failure is abnormal application call. Another example is that when the root causes of the interface failure include service alarm and service release information alarm, it can be determined that the final root cause of the interface failure is upstream and downstream changes. Of course, it is not limited to this, and no restrictions are made here.
[0068] When determining the relevance score, by determining the root cause of the interface failure based on the relevance score and the associated alarm information corresponding to the relevance score, it is beneficial to improve the convenience of determining the root cause of the interface failure. Correspondingly, different methods can be adopted to determine the root cause of the interface failure according to the relevance score and the associated alarm information corresponding to the relevance score, which is beneficial to improving the flexibility of determining the root cause of the interface failure. Moreover, when performing root cause analysis on the interface to determine the root cause of the interface failure, the operation and maintenance personnel do not need to troubleshoot the interface failures one by one, which is beneficial to improving the maintenance effect of the interface.
[0069] An embodiment of the present application discloses another method for analyzing the root cause of an interface failure. As Figure 5 shown, based on the above Figure 2 corresponding embodiment, before step S10, it further includes:
[0070] S40, obtaining the operation support environment associated with the interface through a preset application performance management tool; the operation support environment includes at least one of the service associated with the interface, the database associated with the interface, the container host associated with the interface, and the service release information associated with the interface.
[0071] S50, based on the association relationship between the interface and the operation support environment, generating a topology diagram corresponding to the interface according to the abnormal alarm information of the interface, the associated alarm information of the operation support environment associated with the interface, and the operation and maintenance information corresponding to the interface other than the abnormal alarm information and the associated alarm information.
[0072] In this embodiment, step S10 includes:
[0073] S11, obtaining the operation and maintenance information corresponding to the interface through the topology diagram.
[0074] For step S40, the data processing system can interface with a preset application performance management tool, that is, an APM tool. The data processing system can obtain the operation support environment associated with the interface through the APM tool. The operation support environment includes at least one of the service associated with the interface, the database associated with the interface, the container host associated with the interface, and the service release information associated with the interface. Among them, at least one of the service, database, container host, and service release information associated with the interface included in the operation support environment can also be referred to as the upstream and downstream of the interface, which is not limited here. Of course, the operation support environment associated with the interface is not limited to this, which is not limited here.
[0075] For step S50, when the operating support environment associated with the interface is obtained, the association relationship between the interface and the operating support environment can be determined. Correspondingly, when the data processing system detects the operation and maintenance information corresponding to the interface, such as the abnormal alarm information of the interface, the association alarm information of the operating support environment associated with the interface, and the operation and maintenance information corresponding to the interface other than the abnormal alarm information and the association alarm information, it can generate the topology diagram corresponding to the interface in combination with the association relationship between the interface and the operating support environment. As Figure 6 shown, in the topology diagram corresponding to the interface, the interface and the operating support environment corresponding to the interface can have corresponding identification information. The identification information can include icons, names, etc., which are not limited here. The icons corresponding to the interface and the operating support environment can be different or the same, which are not limited here. For the interface and the operating support environment with an association relationship, they can be connected by a connection line to indicate the association relationship between the interface and the operating support environment. Of course, it is not limited to this. In an exemplary embodiment, when there is an association relationship between different interfaces, the different interfaces with an association relationship can also be connected by a connection line, which is not limited here. In an exemplary embodiment, the topology diagram corresponding to the interface may further include the operation and maintenance information corresponding to the interface other than the abnormal alarm information and the association alarm information, such as the call times of the interface, the average response time of the interface, the total time of the interface, etc., which are not limited here.
[0076] For step S11, in the process of obtaining the operation and maintenance information corresponding to the interface, since it is necessary to determine whether to determine the root cause of the interface failure, the abnormal alarm information of the interface can be obtained through the topology diagram, and the abnormal alarm information of the operating support environment associated with the interface in the topology diagram can be determined as the association alarm information of the operating support environment associated with the interface. Correspondingly, the data processing system can determine that it is necessary to determine the root cause of the interface failure based on the abnormal alarm information and the association alarm information obtained from the topology diagram, and then continue to obtain the operation and maintenance information other than the abnormal alarm information and the association alarm information.
[0077] When obtaining the operating support environment associated with the interface through a preset application performance management tool, and generating a topology map corresponding to the interface based on the association relationship between the interface and the operating support environment, according to the exception warning information of the interface and the exception warning information of the operating support environment associated with the interface, the topology map can be used to detect whether the operation and maintenance information corresponding to the interface includes the exception warning information of the interface and the associated warning information corresponding to the interface, which is beneficial to improving the detection intuitiveness and detection convenience of the operation and maintenance information corresponding to the interface. Accordingly, when it is determined according to the topology map that the operation and maintenance information does not include the exception warning information of the interface and the associated warning information corresponding to the interface, there is no need to continue obtaining the operation and maintenance information corresponding to the interface. When it is determined according to the topology map that the operation and maintenance information includes the exception warning information of the interface and the associated warning information corresponding to the interface, the operation and maintenance information corresponding to the interface can be continuously obtained, and then based on the operation and maintenance information corresponding to the interface, the root cause of the interface failure can be determined, which is beneficial to saving the process of analyzing the root cause of the interface failure to improve the efficiency of analyzing the root cause of the interface failure, and further beneficial to improving the maintenance effect of the interface.
[0078] An embodiment of the present application discloses another method for analyzing the root cause of interface failure. As Figure 7 shown, on the basis of the above Figure 2 corresponding embodiment, between step S10 and step S20, it further includes:
[0079] S60. When detecting the exception warning information of multiple interfaces, determine the target interface from the multiple interfaces according to the operation and maintenance information corresponding to each of the multiple interfaces.
[0080] In this embodiment, step S20 includes:
[0081] S23. When the operation and maintenance information includes the exception warning information of the target interface and the associated warning information of the operating support environment associated with the target interface, determine the correlation score between the exception warning information of the target interface and the associated warning information corresponding to the target interface according to the operation and maintenance information corresponding to the target interface.
[0082] For step S60, when the data processing system detects the exception warning information of multiple interfaces, since the computing resources of the data processing system are limited, the data processing system needs to determine the target interface from the multiple interfaces to preferentially determine the root cause of the target interface failure, and then preferentially maintain the target interface. In the process of determining the target interface from the multiple interfaces, the data processing system can obtain the operation and maintenance information corresponding to each of the multiple interfaces, and determine the target interface based on the operation and maintenance information corresponding to each of the multiple interfaces.
[0083] Taking multiple interfaces including Interface 1, Interface 2, and Interface 3 as an example. When the data processing system detects the respective exception warning information of Interface 1 to Interface 3, it can evaluate the urgency of performing a root cause analysis of the faults on Interface 1 to Interface 3 based on the respective corresponding operation and maintenance information of Interface 1 to Interface 3. For example, when the data processing system determines that the urgency of performing a root cause analysis of the fault on Interface 1 is higher than that of performing a root cause analysis of the faults on Interface 2 and Interface 3, the data processing system can determine that the target interface is Interface 1.
[0084] For step S23, when the target interface is determined, the root cause analysis of the fault can be preferentially performed on the target interface to determine the root cause of the fault of the target interface. The data processing system can determine that it is necessary to evaluate the possibility that the exception warning information of the target interface is caused by the associated warning information when the operation and maintenance information corresponding to the target interface includes the exception warning information of the target interface and the associated warning information of the operation support environment associated with the target interface. Based on this, the data processing system can determine the correlation score between the exception warning information of the target interface and the associated warning information corresponding to the target interface according to the operation and maintenance information corresponding to the target interface. The correlation score corresponding to the target interface can be used to determine the root cause of the fault of the target interface. By analogy, when the correlation score corresponding to the target interface and / or the root cause of the fault of the target interface is determined, the data processing system can continue to select a new target interface from multiple interfaces to obtain the correlation score between the exception warning information of the new target interface and the associated warning information corresponding to the new target interface for subsequent determination of the root cause of the fault of the new target interface.
[0085] Based on this, when the exception warning information of multiple interfaces is detected, by determining the target interface from multiple interfaces according to the respective corresponding operation and maintenance information of multiple interfaces to determine the correlation score between the exception warning information of the target interface and the associated warning information corresponding to the target interface, the computing resources of the data processing system can be preferentially utilized to determine the correlation score corresponding to the target interface for subsequent determination of the root cause of the fault of the target interface, thereby facilitating the root cause analysis of the fault of the target interface and improving the efficiency of the root cause analysis of the fault of the target interface.
[0086] Exemplarily, the operation and maintenance information further includes at least one of the error count of the interface and the response time of the interface.
[0087] In some embodiments, when there is an interface among multiple interfaces whose corresponding operation and maintenance information meets the preset screening conditions, the interface that meets the preset screening conditions is determined as the target interface; the preset screening conditions include at least one of the error count of the interface being greater than or equal to the preset count threshold and the response time of the interface being greater than or equal to the preset time threshold.
[0088] For example, in the case of determining a target interface from multiple interfaces, it is possible to determine whether the operation and maintenance information of each of the multiple interfaces meets a preset filtering condition. In the case where there is an interface among the multiple interfaces whose corresponding operation and maintenance information meets the preset filtering condition, the interface that meets the preset filtering condition can be determined as the target interface. Correspondingly, in the case where there is an interface among the multiple interfaces whose corresponding operation and maintenance information does not meet the preset filtering condition, the interface that does not meet the preset filtering condition can be determined as a non-target interface. The preset filtering condition can be set in advance, or can be set by relevant personnel, such as operation and maintenance personnel, without limitation here.
[0089] The operation and maintenance information can include at least one of the error count of the interface and the response time of the interface.
[0090] When there is an interface whose error count is greater than or equal to a preset count threshold, it can be determined that the interface frequently has errors, and the urgency of performing a root cause analysis of the interface is relatively high. Correspondingly, when there is an interface whose error count is less than the preset count threshold, it can be determined that the interface does not frequently have errors, and the urgency of performing a root cause analysis of the interface is relatively low. The preset count threshold can be set in advance, or can be set by relevant personnel, such as operation and maintenance personnel, without limitation here. For example, the preset count threshold can be at least one of the maximum value, median, average value, and mode of the error counts of each of the multiple interfaces, without limitation here.
[0091] When there is an interface whose response time is greater than or equal to a preset time threshold, it can be determined that the interface has a response timeout, and the urgency of performing a root cause analysis of the interface is relatively high. Correspondingly, when there is an interface whose response time is less than the preset time threshold, it can be determined that the interface does not have a response timeout, and the urgency of performing a root cause analysis of the interface is relatively low. The preset time threshold can be set in advance, or can be set by relevant personnel, such as operation and maintenance personnel, without limitation here. For example, the preset time threshold can be at least one of the maximum value, median, average value, and mode of the response times of each of the multiple interfaces, without limitation here.
[0092] In the case where setting the preset filtering condition can include at least one of the error count of the interface being greater than or equal to the preset count threshold and the response time of the interface being greater than or equal to the preset time threshold, the data processing system can, according to the preset filtering condition, preferentially determine the interface with a relatively high urgency of root cause analysis of the fault as the target interface for subsequent preferentially determining the root cause of the fault of the target interface.
[0093] Based on the setting of preset screening conditions, the operation and maintenance information corresponding to multiple interfaces can be combined to determine the target interface from multiple interfaces, which is conducive to improving the convenience of determining the target interface. In the case of determining the target interface, subsequent fault root cause analysis can be preferentially performed on the target interface, which is conducive to improving the flexibility of fault root cause analysis of the interface, thereby being conducive to improving the maintenance effect of the interface.
[0094] In some embodiments, when multiple response times of an interface are obtained, according to the interquartile range algorithm, the time-consuming outliers are determined from the multiple response times; when there is a time-consuming outlier greater than or equal to the preset time threshold, it is determined that the response time of the interface is greater than or equal to the preset time threshold.
[0095] For example, in the case of obtaining the operation and maintenance information corresponding to an interface, the operation and maintenance information corresponding to the interface may include the response time of the interface. Correspondingly, the same interface can be called multiple times, and the interface can have multiple response times. In the case where the interface has multiple response times, if it is determined whether the interface is the target interface based on whether the multiple response times of the interface are greater than or equal to the preset time threshold, it is easy to have some response times greater than or equal to the preset time threshold and some response times less than the preset time threshold among the multiple response times, resulting in a situation where it is impossible to determine whether the interface is the target interface. Based on this, in order to improve the flexibility and convenience of determining the target interface, when multiple response times of an interface are obtained, the time-consuming outliers can be determined from the multiple response times according to the interquartile range (IQR) algorithm. In the process of determining the time-consuming outliers from the multiple response times of the interface according to the interquartile range algorithm, they can be sorted in ascending order of time consumption to obtain the sorted multiple response times. According to the sorted multiple response times, the first quartile and the third quartile are determined. According to the first quartile and the third quartile, the preset time-consuming range corresponding to the multiple response times is determined. The response times that are not within the preset time-consuming range are determined as time-consuming outliers. The time-consuming outliers existing in the multiple response times can be used as the basis for judging whether the response time of the interface is greater than or equal to the preset time threshold, and further as the basis for judging whether the interface is the target interface. For example, in the case where there is a time-consuming outlier greater than or equal to the preset time threshold, it can be determined that the response time of the interface is greater than or equal to the preset time threshold. In the case where it is determined that the response time of the interface is greater than or equal to the preset time threshold, it can be determined that the interface meets the preset screening conditions, that is, it can be determined that the interface is the target interface for subsequent preferential fault root cause analysis of the target interface.
[0096] When obtaining the response time-consuming of multiple interfaces, by combining the interquartile range algorithm to determine the time-consuming outliers, and when the time-consuming outliers are greater than or equal to the preset time threshold, it is determined that the response time-consuming of the interface is greater than or equal to the preset time threshold, which is beneficial to improving the convenience of determining whether the response time-consuming of the interface is greater than or equal to the preset time threshold, so as to improve the convenience of determining the target interface subsequently.
[0097] In some embodiments, according to the root cause of the interface failure, operation and maintenance prompt information is output, and the operation and maintenance prompt information is used to prompt the maintenance of the interface.
[0098] When determining the root cause of the interface failure, the data processing system can output operation and maintenance prompt information to prompt relevant personnel, such as operation and maintenance personnel, to refer to the root cause of the interface failure indicated by the operation and maintenance prompt information to maintain the interface. The root cause of the interface failure may include at least one associated alarm information, and the operation and maintenance prompt information may be used to indicate at least one associated alarm information and the respective associated degree scores corresponding to at least one associated alarm information.
[0099] For example, when determining the root cause of the interface failure, since the root cause of the failure is determined according to the associated degree score and the associated alarm information corresponding to the associated degree score, the associated alarm information corresponding to the associated degree score can be sorted according to the associated degree score to obtain a recommended list of the root cause of the interface failure; among them, the higher the associated degree score of the associated alarm information, the higher the ranking in the recommended list of the root cause of the interface failure; according to the recommended list of the root cause of the interface failure, operation and maintenance prompt information is output.
[0100] Such as Figure 8As shown, take the fault root cause of the interface including associated alarm information 1 to associated alarm information 3 as an example. When the correlation score corresponding to associated alarm information 1 is greater than the correlation score corresponding to associated alarm information 3, and the correlation score corresponding to associated alarm information 3 is greater than the correlation score corresponding to associated alarm information 2, the associated alarm information 1 to associated alarm information 3 can be sorted to obtain a recommended list of the fault root cause of the interface. Among them, in the recommended list of the fault root cause, associated alarm information 1 is before associated alarm information 3, and associated alarm information 3 is before associated alarm information 2. Correspondingly, when the data processing system outputs operation and maintenance prompt information according to the fault root cause of the interface, it can display the recommended list of the fault root cause of the interface. When the operation and maintenance personnel view the recommended list of the fault root cause, they can determine that it is necessary to prioritize the investigation of associated alarm information 1. After investigating associated alarm information 1, if abnormal alarm information of the interface is still detected, the investigation of associated alarm information 3 can be continued. Correspondingly, after investigating associated alarm information 3, if abnormal alarm information of the interface is still detected, the investigation of associated alarm information 2 can be continued. And so on, until abnormal alarm information of the interface is not detected, it can be determined that the maintenance of the interface has been completed.
[0101] Of course, it is not limited to this. For example, the data processing system can detect abnormal alarm information of the interface in real time or at preset time intervals. Correspondingly, when the data processing system detects abnormal alarm information of the interface at preset time intervals, the data processing system can obtain the abnormal alarm information of all interfaces within the preset time interval and determine the correlation scores between the abnormal alarm information corresponding to each interface and the associated alarm information. Moreover, the data processing system can select the associated alarm information with a correlation score greater than or equal to the preset score threshold from the correlation scores corresponding to each interface and determine it as a potential problem affecting the stability or performance of the data processing system. For example, a potential problem affecting the stability or performance of the data processing system can also be called a system smoking point. The data processing system can output system maintenance information according to the system smoking point to prompt the maintenance of the system smoking point. There is no restriction here.
[0102] Based on this, after determining the fault root cause of the interface, by outputting operation and maintenance prompt information according to the fault root cause of the interface, it is beneficial to improve the convenience of maintaining the interface and thus improve the maintenance effect of the interface.
[0103] It can be seen that in the above embodiments, for the interfaces involved in data processing systems such as digital financial systems and digital medical systems, the operation and maintenance information corresponding to the interfaces can be obtained. When the operation and maintenance information includes the abnormal alarm information of the interfaces and the associated alarm information of the operation support environment associated with the interfaces, according to the operation and maintenance alarm information, the correlation score between the abnormal alarm information and the associated alarm information can be determined. Furthermore, the possibility that the associated alarm information corresponding to the correlation score is the root cause of the interface failure can be judged based on the correlation score to determine the root cause of the interface failure, which is beneficial to improving the convenience of analyzing the root cause of the interface failure. The root cause of the interface failure can be used as the basis for maintaining the interface to trace the source and maintain the interface failure, thereby reducing the possibility of repeated occurrence of the interface failure, which is beneficial to improving the maintenance effect of the interface. Correspondingly, since the operation of data processing systems such as digital financial systems and digital medical systems depends on the invocation of interfaces, when the maintenance effect of the interfaces is improved, it is beneficial to promote the normal operation of data processing systems such as digital financial systems and digital medical systems.
[0104] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0105] In one embodiment, a device for analyzing the root cause of interface failure is provided. The device for analyzing the root cause of interface failure corresponds one-to-one with the method for analyzing the root cause of interface failure in the above embodiments. As Figure 9 shown, the device for analyzing the root cause of failure includes an operation and maintenance information acquisition module 110, a scoring module 120, and a root cause of failure determination module 130. The detailed description of each functional module is as follows:
[0106] The operation and maintenance information acquisition module 110 is used to acquire the operation and maintenance information corresponding to the interface;
[0107] The scoring module 120 is used to determine the correlation score between the abnormal alarm information and the associated alarm information according to the operation and maintenance information when the operation and maintenance information includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface;
[0108] The root cause of failure determination module 130 is used to determine the root cause of the interface failure according to the correlation score and the associated alarm information corresponding to the correlation score.
[0109] In one embodiment, the scoring module 120 is used for:
[0110] Extract features from the operation and maintenance information to obtain the correlation features between the abnormal alarm information and the associated alarm information; the correlation features include at least one of the alarm time correlation degree, the alarm type correlation degree, and the alarm level correlation degree;
[0111] Based on a preset scoring model, determine the correlation degree score between the abnormal alarm information and the associated alarm information according to the correlation features between the abnormal alarm information and the associated alarm information; the preset scoring model is trained according to the historical abnormal alarm information, the historical associated alarm information corresponding to each of multiple interfaces, and the preset correlation degree score between the historical abnormal alarm information and the historical associated alarm information.
[0112] In one embodiment, the fault root cause analysis device is further configured to:
[0113] Obtain the operation support environment associated with the interface through a preset application performance management tool; the operation support environment includes at least one of the service associated with the interface, the database associated with the interface, the container host associated with the interface, and the service release information associated with the interface;
[0114] Based on the association relationship between the interface and the operation support environment, generate a topology diagram corresponding to the interface according to the abnormal alarm information of the interface, the associated alarm information of the operation support environment associated with the interface, and the operation and maintenance information corresponding to the interface other than the abnormal alarm information and the associated alarm information;
[0115] The operation and maintenance information acquisition module 110 is configured to:
[0116] Obtain the operation and maintenance information corresponding to the interface through the topology diagram.
[0117] In one embodiment, the fault root cause analysis device is further configured to:
[0118] When the abnormal alarm information of multiple interfaces is detected, determine a target interface from the multiple interfaces according to the operation and maintenance information corresponding to each of the multiple interfaces;
[0119] The scoring module 120 is further configured to:
[0120] When the operation and maintenance information includes the abnormal alarm information of the target interface and the associated alarm information of the operation support environment associated with the target interface, determine the correlation degree score between the abnormal alarm information of the target interface and the associated alarm information corresponding to the target interface according to the operation and maintenance information corresponding to the target interface.
[0121] In one embodiment, the operation and maintenance information further includes at least one of the error count of the interface and the response time of the interface;
[0122] The fault root cause analysis device is further configured to:
[0123] When the operation and maintenance information corresponding to each interface among the multiple interfaces meets the preset screening conditions, determine the interface that meets the preset screening conditions as the target interface;
[0124] The preset screening conditions include at least one of the error count of the interface being greater than or equal to a preset count threshold and the response time of the interface being greater than or equal to a preset time threshold.
[0125] In one embodiment, the fault root cause analysis device is further configured to:
[0126] When obtaining multiple response times of the interface, determine the time-consuming outlier from the multiple response times according to the interquartile range algorithm;
[0127] When there is a time-consuming outlier greater than or equal to the preset time threshold, determine that the response time of the interface is greater than or equal to the preset time threshold.
[0128] In one embodiment, the fault root cause analysis device is further configured to:
[0129] Output operation and maintenance prompt information according to the fault root cause of the interface, where the operation and maintenance prompt information is used to prompt the maintenance of the interface.
[0130] This application provides a fault root cause analysis device for an interface. When obtaining the operation and maintenance information corresponding to the interface, if the operation and maintenance information includes the abnormal alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, the correlation degree score between the abnormal alarm information and the associated alarm information can be determined based on the operation and maintenance information. The correlation degree score can be used to judge the possibility that the associated alarm information corresponding to the correlation degree score is the fault root cause of the interface, and then determine the fault root cause of the interface, which is beneficial to improving the convenience of fault root cause analysis of the interface. The fault root cause of the interface can be used as the basis for maintaining the interface to trace the source and maintain the fault of the interface, thereby reducing the possibility of repeated occurrence of the interface fault, which is beneficial to improving the maintenance effect of the interface. Correspondingly, since the operation of data processing systems, such as digital financial systems and digital medical systems, depends on the invocation of interfaces, when the maintenance effect of the interface is improved, it is beneficial to promote the normal operation of data processing systems, such as digital financial systems and digital medical systems.
[0131] For the specific limitations of the interface fault root cause analysis device, reference can be made to the limitations of the interface fault root cause analysis method in the above text, which will not be elaborated here. Each module in the above interface fault root cause analysis device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0132] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an interface fault root cause analysis method.
[0133] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 11 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an interface fault root cause analysis method
[0134] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0135] Obtain the operation and maintenance information corresponding to the interface;
[0136] When the operation and maintenance information includes the exception alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, determine the correlation score between the exception alarm information and the associated alarm information according to the operation and maintenance information;
[0137] Determine the root cause of the fault of the interface according to the correlation score and the associated alarm information corresponding to the correlation score.
[0138] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0139] Obtain the operation and maintenance information corresponding to the interface;
[0140] When the operation and maintenance information includes the exception alarm information of the interface and the associated alarm information of the operation support environment associated with the interface, determine the correlation score between the exception alarm information and the associated alarm information according to the operation and maintenance information;
[0141] Determine the root cause of the fault of the interface according to the correlation score and the associated alarm information corresponding to the correlation score.
[0142] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can implement, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0143] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0144] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0145] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for analyzing the root cause of an interface failure, characterized in that: include: Get the operation and maintenance information corresponding to the interface; When the operation and maintenance information includes abnormal alarm information of the interface and associated alarm information of an operation support environment associated with the interface, determining a correlation score between the abnormal alarm information and the associated alarm information according to the operation and maintenance information; The root cause of the failure of the interface is determined according to the correlation score and the associated alarm information corresponding to the correlation score.
2. The fault root cause analysis method according to claim 1, characterized in that: The step of determining, according to the operation and maintenance information, a correlation score between the abnormal alarm information and the associated alarm information comprises: Extracting features from the operation and maintenance information to obtain correlation features between the abnormal alarm information and the associated alarm information; the correlation features include at least one of an alarm time correlation, an alarm type correlation, and an alarm level correlation; Based on a preset scoring model, a correlation score between the abnormal alarm information and the associated alarm information is determined according to the correlation characteristics between the abnormal alarm information and the associated alarm information; the preset scoring model is trained based on historical abnormal alarm information, historical associated alarm information, and preset correlation scores between the historical abnormal alarm information and the historical associated alarm information corresponding to multiple interfaces.
3. The fault root cause analysis method according to claim 1 or 2, characterized in that: Also includes: Obtaining the operation support environment associated with the interface through a preset application performance management tool; The operation support environment includes at least one of a service associated with the interface, a database associated with the interface, a container host associated with the interface, and service publishing information associated with the interface; Based on the association relationship between the interface and the operation support environment, generating a topology map corresponding to the interface according to abnormal alarm information of the interface, associated alarm information of the operation support environment associated with the interface, and operation and maintenance information corresponding to the interface other than the abnormal alarm information and the associated alarm information; The operation and maintenance information corresponding to the acquisition interface includes: The operation and maintenance information corresponding to the interface is obtained through the topology map.
4. The fault root cause analysis method according to claim 1 or 2, characterized in that: After obtaining the operation and maintenance information corresponding to the interface, the method further includes: When abnormal alarm information of multiple interfaces is detected, a target interface is determined from the multiple interfaces according to operation and maintenance information corresponding to each of the multiple interfaces; When the operation and maintenance information includes abnormal alarm information of the interface and associated alarm information of the operation support environment associated with the interface, determining, according to the operation and maintenance information, a correlation score between the abnormal alarm information and the associated alarm information, includes: When the operation and maintenance information includes abnormal alarm information of the target interface and associated alarm information of the operation support environment associated with the target interface, the correlation score between the abnormal alarm information of the target interface and the associated alarm information corresponding to the target interface is determined according to the operation and maintenance information corresponding to the target interface.
5. The fault root cause analysis method according to claim 4, characterized in that: The operation and maintenance information also includes at least one of the number of interface errors and the response time of the interface; The determining the target interface from the multiple interfaces according to the operation and maintenance information corresponding to each of the multiple interfaces includes: When operation and maintenance information corresponding to each of the multiple interfaces meets the preset screening condition, determining the interface meeting the preset screening condition as the target interface; The preset screening condition includes at least one of the following: the number of interface errors is greater than or equal to a preset number threshold, and the response time of the interface is greater than or equal to a preset time threshold.
6. The fault root cause analysis method according to claim 5, characterized in that: The response time of the interface is greater than or equal to the preset time threshold, including: When multiple response times of the interface are obtained, determining a time consumption outlier from the multiple response times according to an interquartile range algorithm; When there is a time consumption abnormal value greater than or equal to the preset time threshold, it is determined that the response time consumption of the interface is greater than or equal to the preset time threshold.
7. The fault root cause analysis method according to claim 1 or 2, characterized in that: Also includes: Output operation and maintenance prompt information according to the root cause of the failure of the interface, where the operation and maintenance prompt information is used to prompt maintenance of the interface.
8. A device for analyzing root causes of interface failures, characterized in that: The fault root cause analysis device comprises: The operation and maintenance information acquisition module is used to obtain the operation and maintenance information corresponding to the interface; A scoring module, configured to determine, according to the operation and maintenance information, a correlation score between the abnormal alarm information and the associated alarm information when the operation and maintenance information includes abnormal alarm information of the interface and associated alarm information of an operation support environment associated with the interface; The fault root cause determination module is used to determine the fault root cause of the interface according to the correlation score and the associated alarm information corresponding to the correlation score.
9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the method for analyzing the root cause of a fault of an interface according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for analyzing the root cause of a fault of an interface according to any one of claims 1 to 7 are implemented.