An alarm analysis method and device

By introducing the concept of shock response into the transaction monitoring system and utilizing indicator data transformation and cluster analysis, the problem of difficulty in judging the importance and relevance of alarms in existing technologies has been solved, enabling rapid and accurate alarm analysis and fault diagnosis.

CN114297031BActive Publication Date: 2025-11-28CHINA UNIONPAY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111647311.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-11-28
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Existing transaction monitoring systems struggle to accurately determine the importance and relevance of alarms when faced with a large number of alerts, resulting in inefficient fault analysis.

Method used

By introducing the concept of impact response, the impact response is calculated by transforming and differentiating the indicator data of the monitored object. Cluster analysis and absolute value are used to determine the importance and relevance of alarms, and important alarms are quickly classified.

Benefits of technology

It enables rapid and accurate analysis of a large number of alarms from the monitoring system, improving troubleshooting efficiency, reducing computational load, and enhancing the accuracy of alarm analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114297031B_ABST
    Figure CN114297031B_ABST
Patent Text Reader

Abstract

The application relates to an alarm analysis method and device, which are used for quickly determining important alarms and related alarms from a large number of alarms. The method comprises the following steps: determining at least one index type for alarm analysis for each alarm generated in a preset period; acquiring each index data of each monitoring object in the preset period for any index type; obtaining an impact response of the monitoring object under the index type according to the index data of the monitoring object; and determining the importance of the alarm corresponding to each monitoring object according to the impact responses of the monitoring objects under the index types.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data analysis, and particularly relates to an alarm analysis method and device. BACKGROUND

[0002] In the daily monitoring process of a transaction monitoring system, a large number of alarms are generated, and the importance and relevance of the alarms need to be identified to further locate the alarm causes. If the alarms are analyzed and judged manually, the workload is large, and the accuracy is low.

[0003] In the prior art, a fault tree can be used to display the monitoring objects that may generate alarms. However, on the one hand, the existing transaction monitoring system divides the monitoring objects into a fixed hierarchical structure, and there is a deviation between the hierarchical division of the monitoring objects and the transaction dimensions of the monitoring objects that actually generate alarms. On the other hand, the generation of actual alarms is random, and the fault tree lacks alarm degree information when being displayed. That is, the fault tree displays all the levels of the monitoring objects that may be fault causes. Therefore, when there are multiple alarms caused by faults in the transaction monitoring system, it is difficult to determine which alarm caused by which fault is more important.

[0004] Therefore, there is an urgent need for a solution to quickly determine the importance of different alarms from a large number of alarms. SUMMARY

[0005] The present application provides an alarm analysis method and device to quickly determine important alarms from a large number of alarms.

[0006] In a first aspect, the present application provides an alarm analysis method, which includes: determining at least one index type for alarm analysis for each alarm generated in a preset period; obtaining, for any index type, each index data of a monitoring object in the preset period for the index type; obtaining an impact response of the monitoring object under the index type according to the index data of the monitoring object; and determining the importance of the alarm corresponding to the monitoring object according to the impact responses of the monitoring objects under the index types.

[0007] In the above technical solution, the concept of impact response is introduced into the monitoring system, and the impact response of the index data of each monitoring object is used as a determination basis to automatically analyze and classify a large number of alarms of the monitoring system, and quickly and accurately determine the importance of the alarms issued by each monitoring object.

[0008] Optionally, the impact response of the monitoring object under the index type is obtained according to the index data of the monitoring object, comprising: converting the index data according to a conversion processing mode of the index type to obtain index data meeting linear superposition requirements; and deriving the index data meeting linear superposition requirements according to time points corresponding to the index data to obtain the impact response of the monitoring object under the index type.

[0009] In the technical solution, the index data is converted to index data meeting linear superposition requirements to measure the influence and correlation between alarms of the monitoring objects, thereby reducing the calculation amount of alarm analysis of the monitoring system and improving the accuracy of alarm analysis.

[0010] Optionally, the index type comprises a TPS index, a success rate index and a time delay index; and the conversion processing mode of the index type comprises: when the index type is the success rate index, converting the index data into successful TPS and failed TPS for any index data; and when the index type is the time delay index, converting the index data into total TPS time delay for any index data.

[0011] In the technical solution, the index data is converted to index data meeting linear superposition requirements to measure the influence and correlation between alarms of the monitoring objects, thereby reducing the calculation amount of alarm analysis of the monitoring system and improving the accuracy of alarm analysis.

[0012] Optionally, after the impact response of the monitoring object under the index type is obtained, the method further comprises: clustering the impact responses of the monitoring objects under the index types to obtain at least one cluster group; wherein alarms corresponding to the monitoring objects in the same cluster group are caused by the same fault reason.

[0013] In the technical solution, the impact responses of the monitoring objects under the index types are clustered to determine that alarms in the same cluster group are caused by the same fault reason, thereby automatically analyzing and classifying a large number of alarms of the monitoring system and quickly and accurately determining the correlation between alarms of the monitoring objects.

[0014] Optionally, the importance of the alarms corresponding to the monitoring objects is determined according to the impact responses of the monitoring objects under the index types, comprising: determining the absolute value of the impact response of each monitoring object under any index type; wherein the greater the absolute value, the greater the importance of the alarm corresponding to the monitoring object.

[0015] In the technical solution, the importance of the alarm corresponding to each monitoring object is determined according to the absolute value of the impact response of each monitoring object under the index type, so that important alarms can be quickly and accurately found from a large number of alarms in the monitoring system, and the staff can first investigate the failure cause of the monitoring object causing the important alarm.

[0016] Optionally, after the at least one clustering group is obtained, the method further includes: if the impact response of the first index data of the first monitoring object and the impact response of the second index data of the second monitoring object exist in the same clustering group, considering that the alarm generated by the first monitoring object and the alarm generated by the second monitoring object are associated alarms; the first monitoring object and the second monitoring object are any two different monitoring objects in the monitoring objects; and the first index data and the second index data are any two different index data in the index data.

[0017] In the technical solution, the impact responses of the monitoring objects under the multiple index types are clustered and analyzed, so that the correlation between the alarms generated by the monitoring objects due to the data fluctuations of different index types can be determined.

[0018] Optionally, after the at least one clustering group is obtained, the method further includes: performing alarm root cause analysis on the alarm information of each monitoring object in the at least one clustering group.

[0019] In the design scheme, the impact responses of the monitoring objects under the multiple index types are clustered and analyzed, and each alarm group caused by the same failure cause is output, which can be used as a leading step for determining the alarm root cause, so as to reduce the search range for determining the alarm root cause and reduce the calculation amount.

[0020] In a second aspect, an embodiment of the present application provides an alarm analysis device, including:

[0021] A collection module is configured to determine N monitoring objects generating alarms according to transaction alarm information in a preset time period, N being a positive integer.

[0022] The collection module is further configured to collect index data of the monitoring objects at multiple time points in the preset time period for each monitoring object.

[0023] A processing module is configured to calculate impact responses of the index data of the monitoring objects at the multiple time points in the preset time period, and determine the most possible cause of generating the alarms according to the sizes of the impact responses of the index data of the N monitoring objects at the multiple time points in the preset time period.

[0024] Optionally, the processing module is further configured to perform conversion processing on the index data according to a conversion processing mode of the index type, to obtain index data meeting linear superposition requirements; and perform derivation on the index data meeting linear superposition requirements according to time points corresponding to the index data, to obtain an impact response of the monitoring object under the index type.

[0025] Optionally, the index type includes a TPS index, a success rate index, and a time delay index; the processing module is further configured to, when the index type is the success rate index, convert the index data into successful TPS and failed TPS for any index data; and when the index type is the time delay index, convert the index data into total TPS time delay for any index data.

[0026] Optionally, the processing module is further configured to cluster the impact responses of the monitoring objects under the index types, to obtain at least one cluster group; wherein alarms corresponding to monitoring objects in a same cluster group are caused by a same fault reason.

[0027] Optionally, the processing module is further configured to determine an absolute value of the impact response of each monitoring object under any index type; wherein a larger absolute value represents a greater importance of an alarm corresponding to the monitoring object.

[0028] Optionally, the processing module is further configured to, if, in a same cluster group, there are an impact response of first index data of a first monitoring object and an impact response of second index data of a second monitoring object, consider that an alarm generated by the first monitoring object and an alarm generated by the second monitoring object are associated alarms; the first monitoring object and the second monitoring object are any two different monitoring objects in the monitoring objects; and the first index data and the second index data are any two different index data in the index data.

[0029] Optionally, the processing module is further configured to perform alarm root cause analysis on alarm information of the monitoring objects in the at least one cluster group.

[0030] In a third aspect, an embodiment of the present application further provides a computer device, comprising:

[0031] a memory configured to store program instructions;

[0032] a processor configured to invoke the program instructions stored in the memory, and perform the method described in various possible designs of the first aspect according to the obtained program instructions.

[0033] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores computer readable instructions. When the computer readable instructions are read and executed by a computer, the method in the first aspect or any possible design of the first aspect is implemented. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative labor based on these drawings are within the protection scope of the present application.

[0035] Figure 1 A schematic diagram of a system architecture of a monitoring system provided by the embodiments of the present application;

[0036] Figure 2 A flowchart of an alarm analysis method provided by the embodiments of the present application;

[0037] Figure 3 A curve diagram of a monitoring object in a monitoring system provided by the embodiments of the present application;

[0038] Figure 4 A schematic diagram of an alarm list in a monitoring system provided by the embodiments of the present application;

[0039] Figure 5 A schematic diagram of a specific alarm analysis flow provided by the embodiments of the present application;

[0040] Figure 6 A schematic diagram of another specific alarm analysis flow provided by the embodiments of the present application;

[0041] Figure 7 A schematic diagram of an alarm analysis device provided by the embodiments of the present application;

[0042] Figure 8 A schematic diagram of a computer device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the protection scope of the present application.

[0044] In the embodiments of this application, "multiple" refers to two or more. Terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.

[0045] Figure 1 An exemplary schematic diagram of a monitoring system provided in an embodiment of this application is shown, such as... Figure 1 As shown, this monitoring system can be used to monitor transactions. Multiple hosts used for transaction processing can be connected to the system. The system samples data from the transactions processed on each host at regular intervals (e.g., every 10 seconds or 20 seconds). Then, based on the sampled data, indicator data for each monitored object under different indicator types is obtained according to the defined monitoring objects. Specifically, the monitoring objects are set according to different dimensions of the transactions, such as transaction type, trading institution, and transaction response code. Correspondingly, the indicator types can be the transaction per second (TPS) of the trading system (the number of transactions processed per second), as well as transaction processing success rate indicators, latency indicators, and other types of indicators. Based on the multiple sampled data, the monitoring system displays monitoring curves corresponding to different indicator types for each monitored object on the monitoring system's display page. The monitoring system has pre-set alarm thresholds for each indicator data. If the indicator data collected for a monitored object at a certain time exceeds the alarm threshold, the monitoring system will generate an alarm message for that monitored object, recording the alarm time and the corresponding indicator data for that monitored object at that time. In the daily operation of a monitoring system, a large number of alarm messages are generated. This application analyzes the large number of alarm messages collected by the monitoring system and quickly identifies important alarms and related alarms from the large number of alarms.

[0046] To better understand the embodiments of this application, the theoretical basis of this application is introduced below. This application treats the monitoring system as a linear time-invariant system, wherein the index T of the monitored object m is... m The input and output relationship satisfies the linear definition, as shown in Formula 1, x n Let x be the smallest granularity dimension that can possibly generate an alarm. n It is unknown. Referring to indicators such as the number of transactions, the coefficient 'a' of the indicator data is used to superimpose the coefficients. mn It is restricted to only {0, 1}.

[0047] T m (t)=a m1 x1(t)+a m2 x2(t)+…+a mn x n(t) Formula One

[0048] Derivation of Formula One, i.e. the impact input of each minimum granularity dimension on T m (t) is T' m , as shown in Formula Two.

[0049] T' m (t) = a m1 x'1(t) + a m2 x'2(t) + … + a mn x' n (t) Formula Two

[0050] Suppose that only one minimum granularity dimension x n=p (t) changes at t0, then the derivation result at t0 is T' m (t0) = x' n=p (t0), where a m,n=p = 1. That is, at t0, the rest of the minimum granularity dimensions have no fault change, the derivative value is 0, x' n≠p (t0) = 0, and the derivative value of the monitored object is finally calculated as the derivative value corresponding to the fault of the p dimension.

[0051] Then, when the alarms of multiple monitored objects are caused by the same minimum granularity dimension fault, the impact responses calculated by the multiple monitored objects are the same, i.e. the impact responses of all T m,n=p nodes that satisfy a m = 1 on x p (t) are the same. That is, when the impact responses of multiple monitored objects are calculated to be the same, it can be considered that the alarms issued by the multiple monitored objects are caused by the same fault reason.

[0052] In one possible assumption, when multiple minimum granularity dimensions x n (t) change at t0, the impact response of the monitored object is the linear superposition of the impact responses of each x n (t), and when the impact responses of multiple monitored objects are calculated to be the same, it can also be considered that the alarms issued by the multiple monitored objects are caused by multiple same fault reasons.

[0053] Figure 2 An alarm analysis method provided by an embodiment of the present application is exemplarily shown, and is applied to the above-mentioned monitoring system. As shown in Figure 2 , the method comprises the following steps:

[0054] Step 201, for each alarm generated in a preset period, at least one index type for alarm analysis of each alarm is determined.

[0055] In the embodiments of the present application, the preset time period is a specified time range for analyzing the alarms in this period. In the monitoring system, there are preset alarm thresholds of the index data corresponding to each index type. If the collected index data at a certain time exceeds the alarm threshold, the monitoring system will generate an alarm information. In a period of time, the monitoring system will record a large number of alarms. For the alarms caused by the changes of different index data, the corresponding index type can be selected for subsequent analysis, for example, for the reason determination of transaction volume fluctuation, the TPS index can be selected for analysis; for the reason determination of the increase of transaction failure volume, the success rate index can be selected for analysis, and for the reason determination of the increase of transaction delay, the time delay index can be selected for analysis.

[0056] In step 202, for any index type, the index data of each monitoring object in the preset time period for the index type is obtained, and the impact response of the monitoring object under the index type is obtained according to the index data of the monitoring object.

[0057] In the embodiments of the present application, the monitoring object refers to a where condition or an equivalent description in a relational database, specifically, each monitoring dimension for collecting index data in the monitoring system, for example, the transaction initiating institution can be a group of monitoring objects, then the different initiating institutions of the transaction initiating institution A, the transaction initiating institution B and the transaction initiating institution C can be individually used as a monitoring object, and the monitoring system collects the index data of each type of transaction initiating institution at a fixed time interval. For example, the transaction response code can be a group of monitoring objects, then the different types of transaction response codes of the transaction response code A, the transaction response code B and the transaction response code C can be individually used as a monitoring object, and the monitoring system collects the index data of each type of transaction response code at a fixed time interval. In addition, the transaction type and the internal system can also be used as a group of monitoring objects.

[0058] It should be noted that the monitoring objects can be flexibly set according to different monitoring dimensions, for example, for the transaction initiating institution, the transaction initiating institution of each province can be used as a monitoring object, each type of transaction initiating institution can be used as a monitoring object, and each type of transaction initiating institution of each province can be used as a monitoring object. Taking the transaction initiating institution as an example of each bank, all banks in Shanghai, all banks in Zhejiang province, the Industrial and Commercial Bank of China in all provinces, the Construction Bank of China in all provinces, the Industrial and Commercial Bank of China Shanghai Branch, the Construction Bank of China Shanghai Branch, the Industrial and Commercial Bank of China Zhejiang Branch, the Construction Bank of China Zhejiang Branch, etc. can be used as a monitoring object, and the monitoring system collects the index data of these monitoring objects at a fixed time interval.

[0059] Since the monitoring system is regarded as a linear time-invariant system in the present application, each index data needs to be converted according to the conversion processing mode of the index type before the impact response of the monitoring object under the index type is calculated, so as to obtain each index data meeting the linear superposition requirement. Specifically, when the index type is TPS index, the TPS index data can be directly used for subsequent processing; when the index type is success rate index, for any index data, the index data is converted into successful TPS and failed TPS, i.e. the transaction volume processed successfully per second in the transaction system and the transaction volume processed unsuccessfully per second in the transaction system; when the index type is time delay index, for any index data, the index data is converted into total TPS delay, where the total TPS delay is the product of the average delay and the total TPS.

[0060] For the index data of each monitoring object collected by the monitoring system within the preset time period, the derivative of each index data meeting the linear superposition requirement is calculated according to the time point corresponding to each index data, so as to obtain the impact response of each monitoring object under the corresponding index type, which serves as the basis for subsequent alarm analysis.

[0061] For example, it is assumed that the monitoring system collects the index data of 100 monitoring objects every 10s, the preset time period is 10 minutes, and the collected index data is TPS. Then, 60 TPSs are collected for each monitoring object within 10 minutes, and 6000 TPSs are collected for 100 monitoring objects in total. The derivative of the 6000 TPSs is calculated to obtain the impact response of the TPS of 100 monitoring objects, which serves as the basis for subsequent alarm analysis.

[0062] Step 203, determining the importance of the alarm corresponding to each monitoring object according to the impact response of each monitoring object under each index type.

[0063] In the embodiments of the present application, for any index type, the importance of the alarm corresponding to each monitoring object can be determined by monitoring the absolute value of the impact response of the monitoring object under the index type, wherein the greater the absolute value, the greater the importance of the alarm corresponding to the monitoring object. Specifically, after determining the absolute value of the impact response of each monitoring object under the index type, the absolute values of the impact responses of each monitoring object under the index type are sorted in descending order, and the alarm corresponding to the monitoring object with the largest absolute value of the impact response is output as an important alarm, and the fault reason causing the alarm is first investigated. When the absolute values of the impact responses of multiple monitoring objects under the index type are the largest, it is considered that the alarms corresponding to these monitoring objects are important alarms. If the impact response values of these monitoring objects under the index type are the same or similar, it is considered that these alarms are caused by the same fault reason. It should be noted that when the absolute values of the multiple impact responses are the largest, the absolute values of the multiple impact responses are not necessarily completely equal, and the absolute values of the impact responses can be approximately equal. Similarly, the impact response values can also be approximately equal.

[0064] Further, the alarms in a preset time period can be caused by multiple fault reasons. In order to distinguish the alarms caused by different fault reasons, after obtaining the impact responses of the monitoring objects under each index type, the impact responses of each monitoring object under each index type are clustered to obtain at least one cluster group. The clustering analysis can be performed using a clustering algorithm based on, for example, the Euclidean distance, which is not limited in the present application. It is considered that the alarms corresponding to each monitoring object in the same cluster group are caused by the same fault reason. The threshold of the impact response value of each index data can also be set in the monitoring system, and multiple cluster groups with impact response values greater than the threshold are output, and it is considered that the alarms in these cluster groups are relatively important.

[0065] In one example, the derivation of 6000 TPS of 100 monitoring objects is collected in total, and the impact response of 6000 TPS of 100 monitoring objects is obtained. The absolute values of the impact response of 6000 TPS are sorted in descending order, and it is assumed that the absolute values of three impact responses are the largest, which correspond to monitoring object A, monitoring object B and monitoring object C respectively. Then, the alarms of the three monitoring objects are output as important alarms, and the three alarms are caused by the same fault reason. Further, the 6000 TPS impact response values are clustered, and it is assumed that three impact response value larger clustering groups are output, which are group 1 {monitoring object A, monitoring object B, monitoring object C}, group 2 {monitoring object E, monitoring object F, monitoring object G} and group 3 {monitoring object H, monitoring object I}. Then, it can be considered that the alarms in the past 10 minutes are caused by three main fault reasons, and the alarms generated by the monitoring objects in each group are caused by the same fault reason.

[0066] In another example, if group 1 {monitoring object A, monitoring object B, monitoring object C} of the above example corresponds to all banks in Shanghai, all provincial banks of the Bank of Communications and the Bank of Communications Shanghai Branch respectively, it can be considered that the alarms generated by all banks in Shanghai, all provincial banks of the Bank of Communications and the Bank of Communications Shanghai Branch are caused by the same fault reason. Since the granularity dimension of the monitoring object of the Bank of Communications Shanghai Branch is smaller, it can be judged that the alarm is caused by a fault in the Bank of Communications Shanghai Branch, and more specific fault troubleshooting needs to be performed on the Bank of Communications Shanghai Branch.

[0067] In another example, if transaction response code A fails, it is assumed that there are 32 transaction initiation institutions using transaction response code A. Then, all 32 transaction initiation institutions using transaction response code A and the failure TPS indicator of transaction response code A will issue alarms. If it is calculated that the failure TPS impact response values of the 32 transaction initiation institutions are small, and the failure TPS impact response value of transaction response code A is large, it can be considered that these alarms are unrelated to the transaction initiation institutions, and the alarms are caused by the fault of transaction response code A.

[0068] In a possible implementation, after clustering the impact responses of each monitoring object under each index type, a plurality of clustering groups are obtained, each clustering group can be used as a judgment threshold for alarm root cause determination, the format of the judgment threshold is {time range: alarm list}, wherein the time range is the time range involved in the alarms clustered into the same group, and the alarm list is each alarm of each monitoring object and the alarm issuing monitoring object. The alarm information of each monitoring object in at least one clustering group is analyzed for alarm root cause analysis. Specifically, the correlation or decision tree algorithm can be used, and the node with the largest correlation with the fault is taken as the root cause of the domain analysis, and the root cause determination result is output.

[0069] In a possible implementation, the application can also determine the correlation between alarms caused by different index types of different monitoring objects according to the impact responses of a plurality of index types of different monitoring objects. Specifically, after obtaining at least one clustering group, if there are impact responses of first index data of a first monitoring object and impact responses of second index data of a second monitoring object in the same clustering group, the alarm generated by the first monitoring object and the alarm generated by the second monitoring object are considered as related alarms. Wherein the first monitoring object and the second monitoring object are any two different monitoring objects in the monitoring objects; the first index data and the second index data are any two different index data in the index data. For example, the reason for the decrease of TPS of the downstream monitoring object caused by the increase of the failure TPS of the upstream monitoring object can be determined by clustering the failure TPS index and the impact response of the TPS index of each monitoring object. Considering that the failure TPS and the TPS fluctuation may be positively correlated or negatively correlated, the absolute values of the impact responses of the TPS and the failure TPS of each monitoring object are used in clustering. After clustering, if there are absolute values of the impact responses of the TPS of monitoring object A and absolute values of the impact responses of the failure TPS of monitoring object B in the same clustering group, it is considered that the alarm generated by monitoring object A and the alarm generated by monitoring object B are related alarms, and it is considered that there is a causal relationship between the increase of the failure TPS of monitoring object B and the decrease of the TPS of downstream monitoring object A.

[0070] The monitoring system also generates a curve graph of each monitoring object when collecting the index data of each monitoring object, as shown in Figure 3 , which is Figure 3 the actual TPS curve and the success rate curve of four different monitoring objects. Figure 4 is the alarm list of different monitoring objects in the same time period. Among them, Figure 3 the impact responses of the TPS and the failure TPS of the four monitoring objects correspond to Figure 4the first four alarm data in the table. It can be seen that the absolute values of the impact responses of the TPS and the failed TPS of the four monitoring objects are relatively large and approximately equal, and the four actual TPS curves corresponding to the four monitoring objects are also similar. Therefore, it is considered that the alarms of the four monitoring objects are important alarms, and the alarms are caused by the same fault reason. Figure 3

[0071] In order to better explain the embodiments of the present application, Figure 5 a specific alarm analysis process provided by the embodiments of the present application is exemplarily shown, and the alarm analysis process is applied to analysis of alarm reasons of a single index type, such as analysis of alarm reasons of scenarios such as transaction volume fluctuation, transaction failure volume increase, and transaction delay increase in a transaction monitoring system. As shown in Figure 5 the alarm analysis process includes the following steps:

[0072] Step 501, according to alarms generated by a transaction monitoring system in a preset time period, determining an index type, a monitoring object, and a time range for which alarm analysis is performed.

[0073] Step 502, converting each index data into index data meeting linear superposition requirements.

[0074] Converting each index data of a plurality of monitoring objects of the selected index type in the preset time period into index data meeting linear superposition requirements, so as to meet the {0, 1} superposition condition.

[0075] Step 503, calculating a derivative value of each monitoring object at each time point under the corresponding index type, that is, an impact response.

[0076] Step 504, sorting the absolute values of the impact responses in descending order, and outputting an alarm corresponding to a monitoring object with the largest absolute value of the impact response.

[0077] Step 505, using a clustering algorithm to cluster according to dimensions such as impact response value, time, and topological distance, and outputting an alarm grouping after clustering.

[0078] Step 506, generating a judgment domain for root cause research according to the alarm grouping, and the general format is {time range: alarm list}.

[0079] Step 507, determining a root cause, for each judgment domain, using a correlation and a decision tree algorithm to take a node with the largest fault correlation as a root cause of the judgment domain, and outputting a root cause determination result.

[0080] In order to better explain the embodiments of the present application, Figure 6 ​Exemplarily, another specific alarm analysis process provided by the embodiment of the application is shown, which is applied to analysis of alarm causes of multiple index types. Taking the cause determination of the decrease of TPS of a downstream monitoring object caused by the increase of failure TPS of an upstream monitoring object as an example, the alarm analysis process includes the following steps. Figure 6

[0081] Step 601: According to alarms generated by the transaction monitoring system in a preset time period, determine the index types (TPS and failure TPS), monitoring objects and time range for alarm analysis.

[0082] Step 602: Calculate the derivative values of TPS and failure TPS of each monitoring object corresponding to each time point, i.e. impact response.

[0083] Step 603: Sort the absolute values of impact responses of TPS in descending order, and output the alarm corresponding to the monitoring object with the largest absolute value of TPS impact response.

[0084] Step 604: Sort the absolute values of impact responses of failure TPS in descending order, and output the alarm corresponding to the monitoring object with the largest absolute value of failure TPS impact response.

[0085] Step 605: Use a clustering algorithm to cluster according to the absolute values of impact responses of TPS and failure TPS, time, topological distance and other dimensions, and output the clustered alarm groups.

[0086] Considering that the fluctuations of failure TPS and TPS may be positively correlated or negatively correlated, the absolute values of impact responses of TPS and failure TPS of each monitoring object are used in clustering.

[0087] Step 606: Mark the alarms generated by TPS of monitoring object A and failure TPS of monitoring object B in the same clustering group as associated alarms.

[0088] For the same clustering group, if the absolute value of impact response of TPS of monitoring object A and the absolute value of impact response of failure TPS of monitoring object B exist, it is considered that the alarm generated by monitoring object A and the alarm generated by monitoring object B are associated alarms, and it is considered that the increase of failure TPS of monitoring object B and the decrease of TPS of downstream monitoring object A have a causal relationship.

[0089] The embodiment of the application provides an alarm analysis method, which introduces the concept of impact response in a monitoring system, takes the impact response of index data of each monitoring object as a determination basis, can automatically analyze and classify a large number of alarms of the monitoring system, and quickly and accurately determines the importance of alarms issued by each monitoring object and the correlation between alarms.​

[0090] based on the same technical concept, Figure 7 An alarm analysis device provided by the embodiment of the application is exemplarily shown, and the device is used to implement the alarm analysis method in the above embodiment. As shown in the figure, Figure 7 The device 700 includes:

[0091] The collection module 701 is configured to determine N monitoring objects that generate alarms according to transaction alarm information in a preset time period, N being a positive integer.

[0092] The collection module 701 is further configured to collect index data of the monitoring objects at a plurality of time points in the preset time period for each monitoring object.

[0093] The processing module 702 is configured to calculate impact responses of the index data of the monitoring objects at the plurality of time points in the preset time period, and determine the most possible cause of the alarms according to the sizes of the impact responses of the index data of the N monitoring objects at the plurality of time points in the preset time period.

[0094] Optionally, the processing module 702 is further configured to perform conversion processing on the index data according to a conversion processing mode of the index type, to obtain index data that meets linear superposition requirements; and perform derivation on the index data that meets linear superposition requirements according to the time points corresponding to the index data, to obtain the impact responses of the monitoring objects under the index type.

[0095] Optionally, the index type includes a TPS index, a success rate index, and a time delay index; and the processing module 702 is further configured to, when the index type is the success rate index, convert the index data into successful TPS and failed TPS for any index data; and when the index type is the time delay index, convert the index data into total TPS time delay for any index data.

[0096] Optionally, the processing module 702 is further configured to cluster the impact responses of the monitoring objects under the index type, to obtain at least one clustering group; and the alarms corresponding to the monitoring objects in the same clustering group are caused by the same fault cause.

[0097] Optionally, the processing module 702 is further configured to determine the absolute values of the impact responses of each monitoring object under any index type; and the greater the absolute value, the greater the importance of the alarm corresponding to the monitoring object.

[0098] Optionally, the processing module 702 is further configured to consider the alarm generated by the first monitoring object and the alarm generated by the second monitoring object as related alarms if, in the same cluster group, there is an impact response of the first indicator data of the first monitoring object and an impact response of the second indicator data of the second monitoring object; the first monitoring object and the second monitoring object are any two different monitoring objects among all monitoring objects; the first indicator data and the second indicator data are any two different indicator data among all indicator data.

[0099] Optionally, the processing module 702 is further configured to perform alarm root cause analysis on the alarm information of each monitored object within at least one cluster group.

[0100] Based on the same technical concept, embodiments of this application provide a computer device, such as... Figure 8 As shown, it includes at least one processor 801 and a memory 802 connected to at least one processor. In this embodiment, the specific connection medium between the processor 801 and the memory 802 is not limited. Figure 8 Taking the connection between the processor 801 and the memory 802 via a bus as an example, the bus can be divided into address bus, data bus, control bus, etc.

[0101] In this embodiment of the application, the memory 802 stores instructions that can be executed by at least one processor 801. By executing the instructions stored in the memory 802, the at least one processor 801 can implement the steps of the alarm analysis method described above.

[0102] The processor 801 is the control center of the computer device, capable of connecting various parts of the computer device via various interfaces and lines. It performs resource configuration by running or executing instructions stored in the memory 802 and accessing data stored in the memory 802. Optionally, the processor 801 may include one or more processing units. The processor 801 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and applications, while the modem processor primarily handles wireless communication. It is understood that the modem processor may not be integrated into the processor 801. In some embodiments, the processor 801 and the memory 802 may be implemented on the same chip; in other embodiments, they may be implemented on separate chips.

[0103] The processor 801 can be a general processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, and can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly embodied as completed by a hardware processor, or completed by a combination of hardware and software modules in the processor.

[0104] The memory 802 is a non-volatile computer readable storage medium, which can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 802 can include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card type memory, random access memory (RAM), static random access memory (SRAM), programmable read only memory (PROM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), magnetic storage, optical disc, etc. The memory 802 is any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but is not limited to this. The memory 802 in the embodiments of the present application can also be a circuit or any other device capable of realizing a storage function, used to store program instructions and / or data.

[0105] Based on the same technical concept, the embodiments of the present application also provide a computer readable storage medium, which stores computer readable instructions, when the computer reads and executes the computer readable instructions, the alarm analysis method listed in any of the above manners is realized.

[0106] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, and the like) embodying computer readable program code.

[0107] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0108] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0109] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0110] While preferred embodiments of the application have been described, modifications and variations can be apparent to those skilled in the art once aware of the general underlying concepts. Accordingly, the appended claims are intended to embrace all such modifications and variations as fall within the scope of the application.

[0111] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. An alarm analysis method, characterized in that, The method includes: For each alarm generated within a preset time period, determine at least one indicator type for alarm analysis of each alarm; For any given indicator type, acquire the indicator data for each monitored object within the preset time period for that indicator type; and obtain the impact response of the monitored object under that indicator type based on the indicator data of the monitored object. The importance of the alarms corresponding to each monitored object is determined based on the impact response of each monitored object under each indicator type. The step of obtaining the impact response of the monitored object under the specified indicator type based on the indicator data of the monitored object includes: According to the conversion processing method of the aforementioned indicator type, the data of each indicator are converted to obtain the data of each indicator that meets the requirements of linear superposition. Based on the time points corresponding to each indicator data, the derivative of each indicator data that meets the linear superposition requirement is calculated to obtain the impact response of the monitored object under the indicator type. The indicator type is: TPS indicator, success rate indicator, or latency indicator; The conversion processing method according to the indicator type, which converts the data of each indicator, includes: When the indicator type is a success rate indicator, for any indicator data, the indicator data is converted into success TPS and failure TPS; When the indicator type is a latency indicator, for any indicator data, the indicator data is converted into total TPS latency.

2. The method according to claim 1, characterized in that, After obtaining the impact response of the monitored object under the specified indicator type, the method further includes: Cluster the impact responses of each monitored object under each indicator type to obtain at least one cluster group; among them, the alarms corresponding to each monitored object belonging to the same cluster group are caused by the same fault.

3. The method according to any one of claims 1 to 2, characterized in that, The determination of the importance of alarms for each monitored object based on the impact response of each monitored object under each indicator type includes: For any given indicator type, determine the absolute value of the impact response of each monitored object under that indicator type; where a larger absolute value indicates a greater degree of importance of the alarm corresponding to the monitored object.

4. The method according to claim 3, characterized in that, After obtaining at least one cluster group, the process further includes: If, within the same cluster group, there exists an impact response of the first indicator data of the first monitored object and an impact response of the second indicator data of the second monitored object, then the alarm generated by the first monitored object and the alarm generated by the second monitored object are considered to be related alarms; the first monitored object and the second monitored object are any two different monitored objects among all monitored objects; the first indicator data and the second indicator data are any two different indicator data among all indicator data.

5. The method according to claim 3, characterized in that, After obtaining at least one cluster group, the process further includes: Perform root cause analysis on alarm information for each monitored object within at least one cluster group.

6. An alarm analysis device, characterized in that, include: The collection module is used to determine the N monitored objects that generated the alarms based on the transaction alarm information within a preset time period, where N is a positive integer; The collection module is also used to collect indicator data of each monitored object at multiple time points within the preset time period. The processing module is used to calculate the impact response of the indicator data of the monitored object at multiple time points within a preset time period; and to determine the most likely cause of the alarm based on the magnitude of the impact response of the indicator data of N monitored objects at multiple time points within the preset time period. The processing module is also used to convert the data of each indicator according to the conversion processing method of the indicator type to obtain the indicator data that meets the linear superposition requirements; and to calculate the derivative of the indicator data that meets the linear superposition requirements according to the time point corresponding to each indicator data to obtain the impact response of the monitored object under the indicator type; the indicator type is: TPS indicator, success rate indicator, or latency indicator. The processing module is further configured to, when the indicator type is a success rate indicator, convert the indicator data into successful TPS and failed TPS for any indicator data; and when the indicator type is a latency indicator, convert the indicator data into total TPS latency for any indicator data.

7. A computer device, characterized in that, include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the method as described in any one of claims 1 to 5 according to the obtained program instructions.

8. A computer-readable storage medium, characterized in that, Includes computer-readable instructions that, when read and executed by a computer, cause the method as described in any one of claims 1 to 5 to be implemented.

Citation Information

Patent Citations

  • Alarm convergence method and device based on clustering algorithm and time sequence association rule

    CN113052225A

  • Organizing network performance metrics into historical anomaly dependency data

    US20150033084A1