Cluster Alarm Control Method, Device, Electronic Device and Storage Medium

By judging the correlation between the management side and the user side alarm items in the cluster alarm system, the problem of high false alarm rate in cluster alarm processing is solved, the accuracy of alarm information is improved, and resource waste is reduced.

CN114461506BActive Publication Date: 2025-06-13BEIJING IQIYI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210055230.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-18
Publication Date
2025-06-13
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

The existing cluster alarm processing method can easily lead to false alarms, which will consume a lot of unnecessary manpower and material costs when viewing and handling cluster problems.

Method used

When the first alarm item on the management side is detected, it is determined whether there is a second alarm item on the user side associated with the alarm item, and whether an alarm information is issued based on the alarm data relationship between the two.

Benefits of technology

By combining the alarm items on the management side and the user side, the accuracy of alarm information is improved, the false alarm rate is reduced, and unnecessary waste of resources is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114461506B_ABST
    Figure CN114461506B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a cluster alarm control method, which includes the following steps: during the operation of the cluster, when a first alarm item on the management side is detected, it is determined whether there is a second alarm item on the user side that has an associated relationship with the first alarm item; if there is a second alarm item currently, it is determined whether to issue an alarm message for the first alarm item according to the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item. Applying the technical solution provided by the embodiment of the present invention, combining the alarm items on the management side and the alarm items on the user side to determine whether to issue an alarm message can improve the accuracy of the alarm message, reduce the false alarm rate of the alarm message, and effectively avoid wasting more unnecessary human and material costs. An embodiment of the present invention also provides a cluster alarm control device, an electronic device, and a computer-readable storage medium, which have corresponding technical effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technologies, and particularly to a method, device, electronic device, and storage medium for cluster alarm control. Background Art

[0002] With the rapid development of computer technologies, cluster technologies have gradually developed. Through cluster technologies, relatively high benefits in terms of performance, reliability, flexibility, etc. can be obtained at a relatively low cost. In enterprises and institutions, to improve business processing capabilities, the number of deployed clusters is increasing.

[0003] To ensure the stability of the cluster, an alarm mechanism is usually set on the management side to diagnose the health status of the cluster. Once a problem occurs in the cluster, an alarm message will be sent. The administrator views the cluster based on the alarm message and solves the problem.

[0004] Currently, this alarm processing method can solve the problems that occur in the cluster in a timely manner to a certain extent. However, because the cluster has a certain self-healing ability and can sometimes automatically recover within a short time after a problem occurs, if an alarm message is sent as soon as a problem occurs, it will lead to false alarms of the alarm message. If the false alarm rate is relatively high, then the process of the administrator viewing the cluster based on the alarm message will consume a lot of unnecessary human and material costs. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a method, device, electronic device, and storage medium for cluster alarm control to reduce the false alarm rate of alarm messages and avoid consuming a lot of unnecessary human and material costs. The specific technical solutions are as follows:

[0006] In the first aspect of the present invention, a method for cluster alarm control is first provided, including:

[0007] During the operation of the cluster, when a first alarm item on the management side is detected, it is determined whether there is a second alarm item on the user side that is related to the first alarm item;

[0008] If the second alarm item currently exists, it is determined whether to send an alarm message for the first alarm item according to the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item.

[0009] In a specific embodiment of the present invention, the determining whether to send an alarm message for the first alarm item according to the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item includes:

[0010] Determine the magnitude relationship between the alarm level of the first alarm item and the alarm level of the second alarm item according to the alarm data of the first alarm item and the alarm data of the second alarm item;

[0011] If the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, determine to send an alarm message for the first alarm item.

[0012] In a specific embodiment of the present invention, when the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, before determining to send an alarm message for the first alarm item, it further includes:

[0013] Determine the inclusion relationship between the set of management-side associated alarm items of the second alarm item and the first alarm item according to the alarm data of the first alarm item and the alarm data of the second alarm item, and the set of management-side associated alarm items of the second alarm item includes alarm items on the management side that are preset to have an associated relationship with the second alarm item;

[0014] If the set of management-side associated alarm items of the second alarm item includes the first alarm item, execute the step of determining to send an alarm message for the first alarm item.

[0015] In a specific embodiment of the present invention, when the second alarm item does not exist currently, it further includes:

[0016] Judge whether there is an unprocessed alarm item with the same alarm event as the first alarm item within the judgment waiting duration before the moment when the first alarm item is detected;

[0017] If there is no unprocessed alarm item with the same alarm event as the first alarm item, ignore the first alarm item.

[0018] In a specific embodiment of the present invention, when there is an unprocessed alarm item with the same alarm event as the first alarm item, it further includes:

[0019] Judge whether to send an alarm message for the first alarm item according to the number of unprocessed alarm items with the same alarm event as the first alarm item.

[0020] In a specific embodiment of the present invention, the judging whether to send an alarm message for the first alarm item according to the number of unprocessed alarm items with the same alarm event as the first alarm item includes:

[0021] If the number of unprocessed alarm items corresponding to the same alarm event as the first alarm item is within a set first quantity range, the first alarm item is ignored;

[0022] If the number of unprocessed alarm items corresponding to the same alarm event as the first alarm item is within a set second quantity range, an alarm message for the first alarm item is determined to be sent;

[0023] The upper limit value of the first quantity range is less than the lower limit value of the second quantity range.

[0024] In a specific implementation manner of the present invention, when detecting a first alarm item on the management side, before determining whether there is a second alarm item on the user side that has an associated relationship with the first alarm item, it further includes:

[0025] Determine the alarm type of the first alarm item;

[0026] If the alarm type of the first alarm item is a set low-sensitivity type, determine whether the alarm event corresponding to the first alarm item is successfully recovered when the self-healing allowable waiting duration of the first alarm item is reached;

[0027] If the alarm event corresponding to the first alarm item is not successfully recovered, perform the step of determining whether there is a second alarm item on the user side that has an associated relationship with the first alarm item.

[0028] In a specific implementation manner of the present invention, when the alarm type of the first alarm item is a set low-sensitivity type, it further includes:

[0029] Mark the alarm status of the first alarm item as a pre-recovery state;

[0030] If the alarm event corresponding to the first alarm item is successfully recovered within the self-healing allowable waiting duration of the first alarm item, update the alarm status of the first alarm item from the pre-recovery state to a recovery state;

[0031] If the alarm event corresponding to the first alarm item is not successfully recovered within the self-healing allowable waiting duration of the first alarm item, update the alarm status of the first alarm item from the pre-recovery state to a fault state, and when it is detected that the alarm event corresponding to the first alarm item is successfully processed after the alarm message for the first alarm item is sent, update the alarm status of the first alarm item from the fault state to a processed complete state.

[0032] In the second aspect of the implementation of the present invention, a cluster alarm control device is further provided, including:

[0033] An associated alarm item existence judgment module, configured to, during the operation of the cluster, when detecting a first alarm item on the management side, determine whether there is currently a second alarm item on the user side that has an associated relationship with the first alarm item;

[0034] An alarm information sending judgment module, configured to, when the second alarm item exists currently, determine whether to send alarm information for the first alarm item according to the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item.

[0035] In another aspect of the implementation of the present invention, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0036] The memory is used to store a computer program;

[0037] The processor is configured to, when executing the program stored on the memory, implement the steps of the above-mentioned cluster alarm control method.

[0038] In another aspect of the implementation of the present invention, there is also provided a computer-readable storage medium, on which a computer program is stored. The program is characterized in that when being executed by a processor, it implements the above-mentioned cluster alarm control method.

[0039] In another aspect of the implementation of the present invention, there is also provided a computer program product including instructions. When it runs on a computer, it enables the computer to execute the above-mentioned cluster alarm control method.

[0040] The technical solution provided by the embodiments of the present invention, when detecting a first alarm item on the management side, does not directly send alarm information for the first alarm item, but first determines whether there is currently a second alarm item on the user side that has an associated relationship with the first alarm item. If it exists, it is considered that the alarm event corresponding to the first alarm item on the management side may have affected the use on the user side. Then, according to the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item, it is determined whether to send alarm information for the first alarm item. That is to say, by combining the alarm items on the management side and the alarm items on the user side, it is determined whether to send alarm information, which can improve the accuracy of the alarm information, reduce the false alarm rate of the alarm information, and effectively avoid wasting more unnecessary human and material resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0042] Figure 1It is a schematic diagram of the composition architecture of the monitoring and alarm platform in the embodiment of the present invention;

[0043] Figure 2 It is a flowchart of the implementation of a cluster alarm control method in the embodiment of the present invention;

[0044] Figure 3 It is a schematic diagram of the specific process of cluster alarm control in the embodiment of the present invention;

[0045] Figure 4 It is a schematic diagram of the structure of a cluster alarm control device in the embodiment of the present invention;

[0046] Figure 5 It is a schematic diagram of the structure of an electronic device in the embodiment of the present invention. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0048] The core of the present invention is to provide a cluster alarm control method, which can be applied to a monitoring and alarm platform. For the convenience of understanding, first, the composition architecture of the monitoring and alarm platform applicable to the technical solution of the present invention will be introduced.

[0049] See Figure 1 , which shows the composition architecture of the monitoring and alarm platform. In the monitoring and alarm platform, the cluster control center can monitor the health status of one or more clusters, such as monitoring the health status of cluster A, cluster B, cluster C,..., cluster n - 1, cluster n, cluster n + 1, etc. When a problem occurs in a cluster, an alarm item on the management side can be generated. At the same time, when the problem that occurs in the cluster affects the user's use, the user application management center can generate an alarm item on the user side. The alarm control center can obtain the alarm item on the management side and the alarm item on the user side, and combine the alarm item on the management side and the alarm item on the user side to determine whether to send an alarm message. This can improve the accuracy of the alarm message, reduce the false alarm rate of the alarm message, and effectively avoid wasting more unnecessary human and material costs.

[0050] Above, the technical solution of the embodiment of the present invention has been described as a whole from the perspective of the monitoring and alarm platform applied in the embodiment of the present invention. Next, the specific implementation of the embodiment of the present invention will be described in detail.

[0051] See Figure 2 shown, which is a flowchart of the implementation of a cluster alarm control method provided by the embodiment of the present invention. This method can be applied to the alarm control center in the monitoring and alarm platform and can include the following steps:

[0052] S210: During the operation of the cluster, when a first warning item on the management side is detected, it is determined whether there is a second warning item on the user side that has an associated relationship with the first warning item.

[0053] It can be understood that during the operation of the cluster, some problems will inevitably occur. For example, problems such as the abnormal process of the application programming interface server (api-server), the abnormal process of the controller manager, and the remaining available disk space being less than the set space threshold. When problems occur in the cluster, warning items on the management side will be generated. The warning items on the management side can be regarded as warning items from the perspective of the administrator, and they are generated based on the warning mechanism set from the administrator's perspective.

[0054] When the problems that occur in the cluster affect the use on the user side, warning items on the user side will be generated. The warning items on the user side can be regarded as warning items from the perspective of the user, and they are generated based on the warning mechanism set from the user's perspective.

[0055] Specifically, the warning items on the management side can be generated by the cluster control center in the monitoring and warning platform when there are problems in the cluster, and the warning items on the user side can be generated by the user application management center in the monitoring and warning platform when it detects that the problems in the cluster affect the use. That is to say, the warning items on the management side can be obtained through the cluster control center, and the warning items on the user side can be obtained through the user application management center.

[0056] However, it can be understood that not all problems that occur in the cluster will affect the use on the user side. For example, when there is a problem in the cluster that the remaining available disk space is less than the set space threshold, a warning item on the management side will be generated for this problem, but this problem will not affect the use on the user side, so no warning item on the user side will be generated.

[0057] In the embodiment of the present invention, an associated relationship between the warning items on the management side and the warning items on the user side can be established in advance. For example, the warning items on the management side can be denoted as The set of warning items on the user side that has an associated relationship with it, that is, the set of user-side associated warning items, can be denoted as The warning items on the user side can be denoted as The set of warning items on the management side that has an associated relationship with it, that is, the set of management-side associated warning items, can be denoted as At the same time, the association between each warning item and each warning item in its corresponding associated warning item set can be set, such as strong association or weak association, etc. For a warning item, it can also be represented by multiple warning item sets to indicate warning items with different associated relationships with this warning item.

[0058] During the operation of the cluster, if a first warning item on the management side is detected, it is possible to first determine whether there is a second warning item on the user side that is associated with the first warning item. Specifically, it is possible to determine whether there is a second warning item on the user side that is associated with the first warning item according to the pre-established association relationship between the warning items on the management side and the warning items on the user side. The association relationship between the warning items on the management side and the warning items on the user side can be recorded in the database. When the warning control center detects a first warning item on the management side, it is possible to determine whether there is a second warning item on the user side that is associated with the first warning item by querying the association relationship recorded in the database.

[0059] If there is currently a second warning item on the user side that is associated with the first warning item, it can be considered that the warning event corresponding to the first warning item has affected the use on the user side, and the operations of the subsequent steps can be continued. If there is currently no second warning item on the user side that is associated with the first warning item, it can be considered that the warning event corresponding to the first warning item has not affected the use on the user side, and the first warning item can be left unprocessed, or it can be processed based on other processing rules.

[0060] S220: If there is currently a second warning item, determine whether to issue a warning message for the first warning item according to the relationship between the warning data of the first warning item and the warning data of the second warning item.

[0061] In the embodiments of the present invention, a first data format of the warning data of the warning items on the management side and a second data format of the warning data of the warning items on the user side can be preset in advance. After generating a warning item on the management side, the warning data of the warning item can be formatted according to the set first data format and converted into warning data with the first data format. After generating a warning item on the user side, the warning data of the warning item can be formatted according to the set second data format and converted into warning data with the second data format. This is to facilitate subsequent determination of the relationship based on the warning data.

[0062] For example, the warning data of the warning item on the management side with the first data format is: The warning data of the warning item on the user side with the second data format is:

[0063] Among them, Alert is the specific warning content, that is, the warning event; is the set of user-side associated warning items; is the set of management-side associated warning items; It is the alarm status of the alarm item on the management side; P is the alarm level; Htime is the self-healing allowable waiting duration. If the alarm type of the alarm item is a highly sensitive type, then Htime is 0, that is, the alarm item of the highly sensitive type has no self-healing waiting duration; Wtime is the judgment waiting duration, and the unit can be seconds.

[0064] It should be noted that the above first data format and second data format are only specific examples, and other different formats can be set during the specific implementation process, and the data items included can also be increased or decreased.

[0065] When the first alarm item on the management side is detected, if it is determined that there is currently a second alarm item on the user side that has an associated relationship with the first alarm item, it can be considered that the alarm event corresponding to the first alarm item has affected the use on the user side, and the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item can be further determined. Then, based on the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item, it is judged whether to send an alarm message for the first alarm item.

[0066] Applying the method provided by the embodiment of the present invention, when the first alarm item on the management side is detected, instead of directly sending an alarm message for the first alarm item, it is first judged whether there is currently a second alarm item on the user side that has an associated relationship with the first alarm item. If there is, it is considered that the alarm event corresponding to the first alarm item on the management side may have affected the use on the user side, and then, based on the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item, it is judged whether to send an alarm message for the first alarm item. That is to say, by combining the alarm item on the management side and the alarm item on the user side, it is judged whether to send an alarm message, which can improve the accuracy of the alarm message, reduce the false alarm rate of the alarm message, and effectively avoid wasting more unnecessary human and material costs.

[0067] In an embodiment of the present invention, step S220 may include the following steps:

[0068] Step 1: Determine the magnitude relationship between the alarm level of the first alarm item and the alarm level of the second alarm item according to the alarm data of the first alarm item and the alarm data of the second alarm item;

[0069] Step 2: If the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, determine to send an alarm message for the first alarm item.

[0070] For the convenience of description, the above two steps are combined for description.

[0071] In an embodiment of the present invention, when a first alarm item on the management side is detected, if it is determined that there is currently a second alarm item on the user side that has an associated relationship with the first alarm item, it can be considered that the alarm event corresponding to the first alarm item has affected the use on the user side, and the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item can be further determined.

[0072] Specifically, based on the alarm data of the first alarm item and the alarm data of the second alarm item, the magnitude relationship between the alarm level of the first alarm item and the alarm level of the second alarm item can be obtained by comparison. After determining the magnitude relationship between the alarm level of the first alarm item and the alarm level of the second alarm item, according to this magnitude relationship, it is determined whether to send an alarm message for the first alarm item.

[0073] The alarm items on the management side and the alarm items on the user side can correspond to alarm levels, and can be divided into multiple alarm levels according to the impact degree of the alarm event. For example, it is divided into 4 alarm levels: P1, P2, P3, P4. Among them, the alarm event with the alarm level of P1 is an important alarm event and needs to be processed immediately once it occurs; the alarm event with the alarm level of P2 is a relatively important alarm event and allows recovery within a few minutes; the alarm event with the alarm level of P3 is an unimportant alarm event and allows recovery within an hour; the alarm event with the alarm level of P4 is a notification event and generally does not need to be concerned. Of course, these 4 alarm levels are only specific examples, and more or fewer alarm levels can be set in the specific implementation process.

[0074] If the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, it can be considered that the alarm event corresponding to the first alarm item has a greater impact on the use on the user side. In this case, it can be determined to send an alarm message for the first alarm item. This can improve the accuracy of the alarm message and reduce false alarms.

[0075] In a specific implementation manner of the present invention, when the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, before determining to send an alarm message for the first alarm item, the inclusion relationship between the set of management-side associated alarm items of the second alarm item and the first alarm item can also be determined according to the alarm data of the first alarm item and the alarm data of the second alarm item. If the set of management-side associated alarm items of the second alarm item includes the first alarm item, the step of determining to send an alarm message for the first alarm item can be executed. The set of management-side associated alarm items of the second alarm item includes the alarm items on the management side that are preset to have an associated relationship with the second alarm item.

[0076] It can be understood that if the alarm data of the first alarm item is abnormal, it will cause the correlation between the first alarm item determined by the alarm data of the first alarm item and the second alarm item to be incorrect. In this case, false alarms of the alarm information for the first alarm item may still occur. Therefore, when the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, it can be further determined whether the set of management-side associated alarm items of the second alarm item contains the first alarm item. If the first alarm item is included, it is considered that there is indeed an association between the first alarm item and the second alarm item, and it can be determined to send the alarm information for the first alarm item. That is, when the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, and the set of management-side associated alarm items of the second alarm item contains the first alarm item, it is determined to send the alarm information for the first alarm item, further improving the accuracy of the alarm information and reducing false alarms.

[0077] If the alarm level of the second alarm item is less than the alarm level of the first alarm item, or the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, but the set of management-side associated alarm items of the second alarm item does not contain the first alarm item, the first alarm item can be ignored first, recorded in a file such as a log, and the alarm information for the first alarm item is not sent. This is to reduce false alarms and avoid wasting a lot of unnecessary human and material costs.

[0078] In the specific implementation process, when the first alarm item on the management side is detected and it is determined that there is a second alarm item on the user side that is associated with the first alarm item, the correlation between the second alarm item and the first alarm item can be determined first, such as strong correlation or weak correlation, etc.

[0079] If the second alarm item has a strong correlation with the first alarm item, it can be determined to send the alarm information for the first alarm item when the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, and the set of associated alarm items of the second alarm item contains the first alarm item.

[0080] If the second alarm item has a weak correlation with the first alarm item, it can be determined to send the alarm information for the first alarm item when the number of second alarm items is greater than the set number threshold.

[0081] In an embodiment of the present invention, when there is no second alarm item currently, the method may further include the following steps:

[0082] Step 1: Determine whether there is an unprocessed alarm item with the same alarm event as the first alarm item during the judgment waiting duration before the moment when the first alarm item is detected;

[0083] Step 2: If there is no unprocessed alarm item with the same alarm event as the first alarm item, then ignore the first alarm item.

[0084] For ease of description, the above two steps are combined for explanation.

[0085] As described above, the alarm data of the alarm items on the management side can have a first data format, which includes data items such as alarm events, the corresponding set of associated alarm items on the user side, alarm levels, alarm statuses, self-healing allowable waiting durations, judgment waiting durations, etc. The judgment waiting durations of different alarm items can be set based on the alarm levels and can be the same or different.

[0086] In the case of detecting the first alarm item on the management side, if there is no second alarm item on the user side that has an association relationship with the first alarm item at present, it can be considered that the alarm event corresponding to the first alarm item does not affect the use on the user side, and it can be further determined whether there is an unprocessed alarm item with the same alarm event as the first alarm item within the judgment waiting duration before the moment when the first alarm item is detected. The judgment waiting durations of the alarm items with the same alarm event are the same.

[0087] If there is no unprocessed alarm item with the same alarm event as the first alarm item within the judgment waiting duration before the moment when the first alarm item is detected, it can be considered that the first alarm item is the first alarm item with this alarm event detected within the judgment waiting duration, and the first alarm item can be ignored without sending an alarm message for the first alarm item. This can effectively avoid false alarms of alarm messages. At the same time, the alarm status of the first alarm item can be marked as the pre-restored state so that when an alarm item with the same alarm event as the first alarm item is detected again, it can be determined according to this alarm status that the first alarm item is an unprocessed alarm item. If the first alarm item is successfully restored within the judgment waiting duration after the moment when the first alarm item is detected, the alarm status of the first alarm item can be marked as the restored state so that when an alarm item with the same alarm event as the first alarm item is detected again, it can be determined according to this alarm status that the first alarm item is not an unprocessed alarm item.

[0088] In an embodiment of the present invention, in the case where there is an unprocessed alarm item with the same alarm event as the first alarm item, the method may further include the following steps:

[0089] Judge whether to send an alarm message for the first alarm item according to the number of unprocessed alarm items with the same alarm event as the first alarm item.

[0090] In the case where the first warning item on the management side is detected, if there is currently no second warning item on the user side that has an associated relationship with the first warning item, it can be considered that the warning event corresponding to the first warning item has not affected the user side usage. It is possible to further determine whether there are unprocessed warning items that are the same as the warning event corresponding to the first warning item within the judgment waiting duration before the moment when the first warning item is detected.

[0091] If there are unprocessed warning items that are the same as the warning event corresponding to the first warning item within the judgment waiting duration before the moment when the first warning item is detected, it can be considered that there were warning items with the same warning event before the first warning item was generated and they were not processed. In this case, the number of unprocessed warning items that are the same as the warning event corresponding to the first warning item can be determined. Based on this number, it can be judged whether to send a warning message for the first warning item.

[0092] Specifically, if the number of unprocessed warning items that are the same as the warning event corresponding to the first warning item is within the set first quantity range, the first warning item is ignored; if the number of unprocessed warning items that are the same as the warning event corresponding to the first warning item is within the set second quantity range, it is determined to send a warning message for the first warning item; the upper limit value of the first quantity range is less than the lower limit value of the second quantity range.

[0093] In the embodiments of the present invention, the first quantity range and the second quantity range can be preset, and the upper limit value of the first quantity range is less than the lower limit value of the second quantity range. For example, the first quantity range is set to (0, 1], and the second quantity range is set to [2, 5].

[0094] If the number of unprocessed warning items that are the same as the warning event corresponding to the first warning item is within the set first quantity range, it can be considered that the warning event corresponding to the first warning item may be self-recovering, and the first warning item can be ignored without sending a warning message for the first warning item. This can effectively avoid false alarms of warning messages.

[0095] If the number of unprocessed warning items that are the same as the warning event corresponding to the first warning item is within the set second quantity range, it can be considered that the warning event corresponding to the first warning item is relatively urgent and may not be self-recovering. It can be determined to send a warning message for the first warning item to promptly conduct problem troubleshooting based on this warning message. At the same time, the warning status of the first warning item can be marked as the fault status, and after processing is completed, the warning status of the first warning item is marked as the processed status.

[0096] In an embodiment of the present invention, when a first warning item on the management side is detected, before determining whether there is a second warning item on the user side that has an associated relationship with the first warning item, the method may further include the following steps:

[0097] The first step: Determine the warning type of the first warning item;

[0098] The second step: If the warning type of the first warning item is a set low-sensitivity type, then determine whether the warning event corresponding to the first warning item has been successfully recovered when the self-healing allowable waiting duration of the first warning item is reached;

[0099] The third step: If the warning event corresponding to the first warning item has not been successfully recovered, then execute the step of determining whether there is a second warning item on the user side that has an associated relationship with the first warning item.

[0100] For ease of description, the above three steps are combined for description.

[0101] There can be various warning types for warning items on the management side, such as high-sensitivity types, low-sensitivity types, etc. Warning items of the high-sensitivity type need to send warning information in a timely manner, such as warning items corresponding to abnormal application program interface server processes, abnormal controller manager processes, DNS server / forwarder unable to start, etc. Warning items of the low-sensitivity type can have a certain period of waiting redundancy, such as warning items corresponding to the remaining available disk space being less than the set space threshold, a certain machine restarting, excessive container set backlog, etc.

[0102] The self-healing allowable waiting duration for warning items of the high-sensitivity type can be set to 0, and the self-healing allowable waiting duration for warning items of the low-sensitivity type can be set according to the warning level.

[0103] In an embodiment of the present invention, when a first warning item on the management side is detected, the warning type of the first warning item can be determined first.

[0104] If the warning type of the first warning item is a set low-sensitivity type, then it can be determined whether the warning event corresponding to the first warning item has been successfully recovered when the self-healing allowable waiting duration of the first warning item is reached. If the warning event corresponding to the first warning item has not been successfully recovered within the self-healing allowable waiting duration, it can be considered that the warning event corresponding to the first warning item cannot self-heal, and the steps of determining whether there is a second warning item on the user side that has an associated relationship with the first warning item and below can be executed, that is, in combination with the warning items on the user side, determine whether to send a warning message for the first warning item to reduce the false alarm rate.

[0105] If the alarm type of the first alarm item is the set high-sensitivity type, the steps of directly determining whether there is a second alarm item on the user side related to the first alarm item and below can be directly executed, that is, combining the alarm items on the user side to determine whether to issue an alarm message for the first alarm item, so as to make a timely judgment on whether to issue an alarm message, which can improve the alarm efficiency while reducing the false alarm rate.

[0106] In an embodiment of the present invention, when the alarm type of the first alarm item is the set low-sensitivity type, the method may further include the following steps:

[0107] The first step: Mark the alarm status of the first alarm item as the pre-recovery state;

[0108] The second step: If within the self-healing allowable waiting duration of the first alarm item, the alarm event corresponding to the first alarm item is successfully recovered, update the alarm status of the first alarm item from the pre-recovery state to the recovery state;

[0109] The third step: If within the self-healing allowable waiting duration of the first alarm item, the alarm event corresponding to the first alarm item is not successfully recovered, update the alarm status of the first alarm item from the pre-recovery state to the fault state, and when it is detected that the alarm event corresponding to the first alarm item is successfully processed after issuing the alarm message for the first alarm item, update the alarm status of the first alarm item from the fault state to the processed state.

[0110] For the convenience of description, the above three steps are combined for description.

[0111] In an embodiment of the present invention, when the first alarm item on the management side is detected, if it is determined that the alarm type of the first alarm item is the low-sensitivity type, the alarm status of the first alarm item can be marked as the pre-recovery state, and then start timing, and poll and query at a set time interval to determine whether the alarm event of the first alarm item is successfully recovered within the self-healing allowable waiting duration of the first alarm item.

[0112] If within the self-healing allowable waiting duration of the first alarm item, the alarm event corresponding to the first alarm item is successfully recovered, the alarm status of the first alarm item can be updated from the pre-recovery state to the recovery state. To determine whether to perform further operations according to the alarm status of the alarm item, such as combining the alarm items on the user side to determine whether to issue an alarm message.

[0113] If the alarm event corresponding to the first alarm item is not successfully recovered within the self-healing allowable waiting duration of the first alarm item, it can be considered that the alarm event corresponding to the first alarm item cannot be self-healed. The alarm status of the first alarm item can be updated from the pre-recovery status to the fault status, and it is combined with the alarm items on the user side to determine whether to send an alarm message. After sending the alarm message for the first alarm item, if it is detected that the alarm event corresponding to the first alarm item is successfully processed, the alarm status of the first alarm item can be updated from the fault status to the processed status.

[0114] That is to say, in the embodiments of the present invention, the alarm status has diversity and can include a pre-recovery status, a recovery status, a fault status, and a processed status. Such as (pre_recovery, recovery, problem, OK), where no alarm message is sent for the alarm items in the pre-recovery status pre_recovery and the recovery status recovery, and they can be recorded in the log. For the alarm items in the fault status problem, alarm messages can be sent after further determination. After processing is completed, the alarm status of the alarm item is OK.

[0115] That is, the update process of the alarm status of the highly sensitive type of alarm item can be: problem -> after successful processing -> OK;

[0116] The update process of the alarm status of the low-sensitive type of alarm item can be one of the following two: pre_recovery -> successfully recovered within the self-healing allowable waiting duration -> recovery; pre_recovery -> not successfully recovered within the self-healing allowable waiting duration -> problem -> after successful processing -> OK.

[0117] Through the multi-faceted processing of the alarm status of the alarm item, it helps to accurately determine whether to send an alarm message for the alarm item.

[0118] For ease of understanding, take Figure 3 the specific implementation process shown as an example to illustrate the embodiments of the present invention again.

[0119] When there is an alarm item on the management side, first determine the alarm type of the alarm item.

[0120] If it is a highly sensitive type, directly format the alarm data and transfer it to the alarm control processing process.

[0121] If it is of the low-sensitivity type, pre-recovery processing is performed, the alarm status is marked as the pre-recovery status, and it is judged whether it is successfully recovered. If it is successfully recovered, the alarm status can be marked as the recovered status and the process ends. If it is not successfully recovered, it can be judged whether it times out, that is, the self-healing allowed waiting duration is reached. If it does not time out, it can continue to judge whether it is successfully recovered. If it times out, the alarm data can be formatted and transferred to the alarm control processing process.

[0122] When there is an alarm item on the user side, the alarm data is directly formatted and transferred to the alarm control processing process.

[0123] During the alarm control processing process, it is judged whether there is currently an alarm item on the user side that is related to the alarm item on the management side. If it exists, it is judged whether to send an alarm message according to the relationship between the alarm data of the alarm item on the management side and the alarm data of the alarm item on the user side, and when it is determined to send an alarm message, the alarm message is sent by means of voice, email, text, etc.

[0124] In the embodiment of the present invention, by combining the alarm item on the management side with the alarm item on the user side, it is judged whether to send an alarm message, which can filter out the self-healable alarm items, reduce the false alarm rate of the alarm message, and improve the stability of the cluster and its services.

[0125] Corresponding to the above method embodiment, the embodiment of the present invention also provides a cluster alarm control device, and the cluster alarm control device described below can be mutually corresponding and referred to the cluster alarm control method described above.

[0126] See Figure 4 As shown, the device may include the following modules:

[0127] The associated alarm item existence judgment module 410 is used to judge whether there is currently a second alarm item on the user side that is related to the first alarm item on the management side when the first alarm item on the management side is detected during the operation of the cluster;

[0128] The alarm message sending judgment module 420 is used to judge whether to send an alarm message for the first alarm item according to the relationship between the alarm data of the first alarm item and the alarm data of the second alarm item when the second alarm item exists currently.

[0129] When the device provided by the embodiment of the present invention is applied, when the first warning item on the management side is detected, instead of directly sending out the warning information for the first warning item, it first judges whether there is a second warning item on the user side that has an associated relationship with the first warning item. If so, it is considered that the warning event corresponding to the first warning item on the management side may have affected the use on the user side. Then, according to the relationship between the warning data of the first warning item and the warning data of the second warning item, it judges whether to send out the warning information for the first warning item. That is to say, by combining the warning items on the management side and the warning items on the user side to judge whether to send out the warning information, the accuracy of the warning information can be improved, the false alarm rate of the warning information can be reduced, and the waste of a large amount of unnecessary human and material resources can be effectively avoided.

[0130] In a specific embodiment of the present invention, the warning information sending judgment module 420 is used for:

[0131] According to the warning data of the first warning item and the warning data of the second warning item, determine the magnitude relationship between the warning level of the first warning item and the warning level of the second warning item;

[0132] If the warning level of the second warning item is greater than or equal to the warning level of the first warning item, determine to send out the warning information for the first warning item.

[0133] In a specific embodiment of the present invention, the warning information sending judgment module 420 is further used for:

[0134] When the warning level of the second warning item is greater than or equal to the warning level of the first warning item, before determining to send out the warning information for the first warning item, according to the warning data of the first warning item and the warning data of the second warning item, determine the inclusion relationship between the set of management-side associated warning items of the second warning item and the first warning item. The set of management-side associated warning items of the second warning item includes the warning items on the management side that are preset to have an associated relationship with the second warning item;

[0135] If the set of management-side associated warning items of the second warning item includes the first warning item, execute the step of determining to send out the warning information for the first warning item.

[0136] In a specific embodiment of the present invention, it further includes:

[0137] The warning item identical judgment module is used to judge whether there is an unprocessed warning item with the same warning event as the first warning item during the judgment waiting duration before the moment when the first warning item is detected when there is no second warning item currently;

[0138] An alarm item processing module, which is used to ignore the first alarm item when there is no unprocessed alarm item that is the same as the alarm event corresponding to the first alarm item.

[0139] In a specific embodiment of the present invention, the alarm information sending judgment module 420 is further used for:

[0140] When there is an unprocessed alarm item that is the same as the alarm event corresponding to the first alarm item, judge whether to send the alarm information for the first alarm item according to the number of unprocessed alarm items that are the same as the alarm event corresponding to the first alarm item.

[0141] In a specific embodiment of the present invention, the alarm information sending judgment module 420 is used for:

[0142] When the number of unprocessed alarm items that are the same as the alarm event corresponding to the first alarm item is within a set first quantity range, ignore the first alarm item;

[0143] When the number of unprocessed alarm items that are the same as the alarm event corresponding to the first alarm item is within a set second quantity range, determine to send the alarm information for the first alarm item;

[0144] The upper limit value of the first quantity range is less than the lower limit value of the second quantity range.

[0145] In a specific embodiment of the present invention, it further includes:

[0146] An alarm type determination module, which is used to determine the alarm type of the first alarm item on the management side when the first alarm item on the management side is detected and before judging whether there is a second alarm item on the user side that has an associated relationship with the first alarm item;

[0147] A recovery judgment module, which is used to judge whether the alarm event corresponding to the first alarm item is successfully recovered when the self-healing allowable waiting duration of the first alarm item is reached when the alarm type of the first alarm item is a set low-sensitivity type; if the alarm event corresponding to the first alarm item is not successfully recovered, trigger the associated alarm item existence judgment module 410 to execute the step of judging whether there is a second alarm item on the user side that has an associated relationship with the first alarm item.

[0148] In a specific embodiment of the present invention, it further includes an alarm status marking module, which is used for:

[0149] When the alarm type of the first alarm item is a set low-sensitivity type, mark the alarm status of the first alarm item as a pre-recovery state;

[0150] If, within the self-healing allowed waiting duration of the first warning item, the warning event corresponding to the first warning item is successfully recovered, then update the warning status of the first warning item from the pre-recovery state to the recovery state;

[0151] If, within the self-healing allowed waiting duration of the first warning item, the warning event corresponding to the first warning item is not successfully recovered, then update the warning status of the first warning item from the pre-recovery state to the fault state, and in the case where it is detected that the warning event corresponding to the first warning item is successfully processed after the warning information for the first warning item is sent, update the warning status of the first warning item from the fault state to the processed completed state.

[0152] An embodiment of the present invention further provides an electronic device, as Figure 5 shown, including a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 complete mutual communication through the communication bus 504. Among them,

[0153] The memory 503 is used to store a computer program;

[0154] The processor 501, when executing the program stored on the memory 503, implements the following steps:

[0155] During the operation of the cluster, when a first warning item on the management side is detected, determine whether there is currently a second warning item on the user side that has an associated relationship with the first warning item;

[0156] If there is currently a second warning item, then determine whether to send a warning information for the first warning item according to the relationship between the warning data of the first warning item and the warning data of the second warning item.

[0157] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of easy representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0158] The communication interface is used for communication between the above terminal and other devices.

[0159] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0160] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0161] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium storing instructions, which, when running on a computer, cause the computer to execute the cluster alarm control method in any one of the above embodiments.

[0162] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which, when running on a computer, cause the computer to execute the cluster alarm control method in any one of the above embodiments.

[0163] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0164] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0165] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.

[0166] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A cluster alarm control method, characterized in that, it includes: During the operation of the cluster, when a first alarm item on the management side is detected, it is judged whether there is a second alarm item on the user side that has an associated relationship with the first alarm item; If the second alarm item currently exists, according to the alarm data of the first alarm item and the alarm data of the second alarm item, determine the magnitude relationship between the alarm level of the first alarm item and the alarm level of the second alarm item; If the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, according to the alarm data of the first alarm item and the alarm data of the second alarm item, determine the inclusion relationship between the set of management-side associated alarm items of the second alarm item and the first alarm item, and the set of management-side associated alarm items of the second alarm item includes the alarm items on the management side that are preset to have an associated relationship with the second alarm item; If the set of management-side associated alarm items of the second alarm item includes the first alarm item, determine to issue an alarm message for the first alarm item.

2. The cluster alarm control method according to claim 1, characterized in that, When the second alarm item does not currently exist, it further includes: Judge whether there is an unprocessed alarm item with the same alarm event as the first alarm item during the judgment waiting duration before the moment when the first alarm item is detected; If there is no unprocessed alarm item with the same alarm event as the first alarm item, ignore the first alarm item.

3. The cluster alarm control method according to claim 2, characterized in that, When there is an unprocessed alarm item with the same alarm event as the first alarm item, it further includes: According to the number of unprocessed alarm items with the same alarm event as the first alarm item, judge whether to issue an alarm message for the first alarm item.

4. The cluster alarm control method according to claim 3, characterized in that, The judging whether to issue an alarm message for the first alarm item according to the number of unprocessed alarm items with the same alarm event as the first alarm item includes: If the number of unprocessed alarm items with the same alarm event as the first alarm item is within a set first quantity range, ignore the first alarm item; If the number of unprocessed alarm items with the same alarm event as the first alarm item is within a set second quantity range, determine to issue an alarm message for the first alarm item; The upper limit value of the first quantity range is less than the lower limit value of the second quantity range.

5. The cluster alarm control method according to any one of claims 1 to 4, characterized in that, When a first alarm item on the management side is detected, before judging whether there is a second alarm item on the user side that has an associated relationship with the first alarm item, it further includes: Determine the alarm type of the first alarm item; If the alarm type of the first alarm item is a set low-sensitivity type, it is determined whether the alarm event corresponding to the first alarm item is successfully recovered when the self-healing allowable waiting duration of the first alarm item is reached; If the alarm event corresponding to the first alarm item is not successfully recovered, the step of determining whether there is a second alarm item on the user side that has an associated relationship with the first alarm item is executed.

6. The cluster alarm control method according to claim 5, characterized in that when the alarm type of the first alarm item is a set low-sensitivity type, it further includes: marking the alarm status of the first alarm item as a pre-recovery status; if the alarm event corresponding to the first alarm item is successfully recovered within the self-healing allowable waiting duration of the first alarm item, updating the alarm status of the first alarm item from the pre-recovery status to a recovery status; if the alarm event corresponding to the first alarm item is not successfully recovered within the self-healing allowable waiting duration of the first alarm item, updating the alarm status of the first alarm item from the pre-recovery status to a fault status, and when it is detected that the alarm event corresponding to the first alarm item is successfully processed after the alarm information for the first alarm item is sent, updating the alarm status of the first alarm item from the fault status to a processed status.

7. A cluster alarm control device, characterized in that it includes: an associated alarm item existence judgment module, configured to, during the operation of the cluster, when detecting a first alarm item on the management side, judge whether there is a second alarm item on the user side that has an associated relationship with the first alarm item; an alarm information sending judgment module, configured to, when the second alarm item exists, determine the magnitude relationship between the alarm level of the first alarm item and the alarm level of the second alarm item according to the alarm data of the first alarm item and the alarm data of the second alarm item; if the alarm level of the second alarm item is greater than or equal to the alarm level of the first alarm item, determine the inclusion relationship between the management-side associated alarm item set of the second alarm item and the first alarm item according to the alarm data of the first alarm item and the alarm data of the second alarm item, and the management-side associated alarm item set of the second alarm item includes the alarm items on the management side that are preset to have an associated relationship with the second alarm item; if the management-side associated alarm item set of the second alarm item includes the first alarm item, it is determined to send the alarm information for the first alarm item.

8. An electronic device, characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used for storing a computer program; the processor is configured to, when executing the program stored on the memory, implement the steps of the cluster alarm control method according to any one of claims 1-6.

9. A computer-readable storage medium, on which a computer program is stored, characterized in that When executed by a processor, the program implements the cluster alarm control method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Alarm processing method, device and system

    CN102820995A

  • Power distribution network alarm method and system, electronic equipment and computer readable storage medium

    CN109829833A

  • Alarm processing method and device and electronic equipment

    CN110650036A