Processing method of alarm self-healing system and alarm self-healing system
By using an alarm self-healing system, combined with Zabbix and ITIL, the system automates the processing of operation and maintenance alarms, solving the problem of low IT operation and maintenance efficiency and achieving efficient operation and maintenance management and fault self-healing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU ROBAM APPLIANCES CO LTD
- Filing Date
- 2024-10-18
- Publication Date
- 2026-04-21
AI Technical Summary
IT operations and maintenance management systems are large and complex. Traditional operations and maintenance personnel are inefficient, prone to human error and distraction, making it difficult to handle complex tasks. In addition, the proliferation of alarms reduces the value of the monitoring system.
An alarm self-healing system is adopted, which uses the Zabbix monitoring module and the ITIL work order module to automatically process operation and maintenance alarm events, generate processing work orders or alarm work orders, and reduce manual intervention.
Improve operational efficiency, reduce manual intervention, free up operations and maintenance personnel from tedious tasks, and allow them to focus on system optimization and innovative projects, thereby enhancing system stability.
Smart Images

Figure CN121901003A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a processing method and a self-healing alarm system. Background Technology
[0002] Currently, IT (Information Technology) operations and maintenance management systems are becoming increasingly large and complex. However, traditional operations and maintenance personnel spend most of their time and energy dealing with simple, repetitive tasks, which is not only inefficient but may also lead to human error and distraction, making it difficult to concentrate on more complex and creative tasks.
[0003] At the same time, as the number of hosts involved in the business system increases and the types of services provided by the applications also increase, the scale of monitoring indicators also increases, and the number of alarms also increases. If the operation and maintenance personnel do not reduce the number of alarms, the alarms may become too numerous and exceed the recipient's capacity, causing the recipient to be annoyed by the alarms or question the alarms, ultimately reducing the utilization value of the monitoring system. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a processing method and an alarm self-healing system to change the previous passive operation and maintenance mode, and to establish an alarm self-healing system through automated alarms, automated event processing and other means, thereby reducing manual intervention and improving operation and maintenance efficiency.
[0005] In a first aspect, embodiments of the present invention provide a processing method for an alarm self-healing system, applied to an alarm self-healing system. The method includes: collecting monitoring data of the monitored device; determining maintenance alarm events based on the monitoring data; determining a self-healing processing flow for the maintenance alarm events based on pre-set matching rules; processing the maintenance alarm events based on the self-healing processing flow; and generating a work order for the maintenance alarm events. The work order includes a processing work order or an alarm work order. The processing work order is used to characterize the result of the self-healing processing flow for the maintenance alarm events; the alarm work order is used to prompt maintenance personnel to manually process the maintenance alarm events.
[0006] In an optional embodiment of this application, the alarm self-healing system includes: a monitoring module, an event handling module, and a work order module, which are connected sequentially. The step of collecting monitoring data from the monitored device includes: the monitoring module collecting monitoring data from the monitored device and sending the monitoring data to the event handling module; the step of determining an operation and maintenance alarm event based on the monitoring data, determining a self-healing process for the operation and maintenance alarm event based on pre-set matching rules, and processing the operation and maintenance alarm event based on the self-healing process includes: the event handling module determining the operation and maintenance alarm event based on the monitoring data, determining a self-healing process for the operation and maintenance alarm event based on pre-set matching rules, and processing the operation and maintenance alarm event based on the self-healing process; and the step of generating a work order for the operation and maintenance alarm event includes: the work order module generating a work order for the operation and maintenance alarm event.
[0007] In an optional embodiment of this application, the monitoring module is a Zabbix monitoring module and the work order module is an ITIL work order module.
[0008] In optional embodiments of this application, the alarm self-healing system further includes a display interface; the method further includes displaying monitoring data on the display interface.
[0009] In an optional embodiment of this application, the step of the monitoring module collecting monitoring data of the monitored device includes: if the monitored device is a hardware device, the monitoring module collects monitoring data through a specified protocol; if the monitored device is an operating system, the monitoring module collects monitoring data through a monitoring agent.
[0010] In optional embodiments of this application, the event handling module determines a self-healing process for maintenance alarm events based on pre-set matching rules, and the step of handling maintenance alarm events based on the self-healing process includes: the event handling module determining whether the maintenance alarm event has reached a threshold based on the matching rules; if the maintenance alarm event has not reached the threshold, the event handling module determines a self-healing process for the maintenance alarm event and handles the maintenance alarm event based on the self-healing process; if the maintenance alarm event has reached the threshold, the event handling module determines that there is no self-healing process for the maintenance alarm event; the step of the work order module generating a work order for the maintenance alarm event includes: if the maintenance alarm event has not reached the threshold, the work order module generates a processing work order for the maintenance alarm event; if the maintenance alarm event has reached the threshold, the work order module generates an alarm work order for the maintenance alarm event.
[0011] In an optional embodiment of this application, the step of determining the operation and maintenance alarm event based on monitoring data includes: determining the monitoring data within a time window; and determining whether an operation and maintenance alarm event exists within the time window based on the monitoring data within the time window.
[0012] In optional embodiments of this application, the monitoring data includes: hardware monitoring indicators, system performance monitoring indicators, and service monitoring indicators.
[0013] In an optional embodiment of this application, the monitored device is equipped with a monitoring agent, and the alarm self-healing system is equipped with a monitoring service; the monitoring agent sends monitoring data to the monitoring service based on a specified frequency.
[0014] In an optional embodiment of this application, the step of determining whether there is an operation and maintenance alarm event within a time window based on monitoring data within the time window includes: if there is no monitoring data within the time window, determining that there is an operation and maintenance alarm event within the time window.
[0015] In an optional embodiment of this application, the step of determining whether there is an operation and maintenance alarm event within a time window based on monitoring data within the time window includes: determining the fluctuation value of the indicator of the monitoring data within the time window; if the fluctuation value of the indicator is greater than a preset fluctuation threshold, determining that there is an operation and maintenance alarm event within the time window.
[0016] In an optional embodiment of this application, the step of determining the fluctuation value of the monitoring data indicator within the time window includes: calculating the fluctuation value of the monitoring data indicator within the time window using the following formula: in, Let be the fluctuation value of the i-th indicator in the monitoring data within the k-th time window, and S be the number of times the indicator is collected within the k-th time window. Let i be the value of the i-th metric of the monitoring data within the k-th time window. It represents the average value of the i-th indicator of the monitoring data within the k-th time window.
[0017] Secondly, embodiments of the present invention provide an alarm self-healing system for performing the above-described alarm self-healing system processing method.
[0018] The embodiments of the present invention bring the following beneficial effects:
[0019] This invention provides a method and system for processing alarm self-healing systems. The method involves collecting monitoring data from monitored devices; identifying maintenance alarm events based on the monitoring data; determining a self-healing process for the maintenance alarm events based on pre-set matching rules; processing the maintenance alarm events based on the self-healing process; and generating work orders for the maintenance alarm events. Each work order includes either a processing work order or an alarm work order. The processing work order characterizes the result of the self-healing process for the maintenance alarm events, while the alarm work order prompts maintenance personnel to manually handle the maintenance alarm events. This approach can change the previous passive maintenance model by establishing an alarm self-healing system through automated alarms and automated event processing, thereby reducing manual intervention and improving maintenance efficiency.
[0020] Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.
[0021] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the structure of an alarm self-healing system provided in an embodiment of the present invention;
[0024] Figure 2 A flowchart illustrating a processing method for an alarm self-healing system provided in an embodiment of the present invention;
[0025] Figure 3 This is a schematic diagram of the overall design framework of an alarm self-healing system provided in an embodiment of the present invention;
[0026] Figure 4 A flowchart of an alarm convergence method provided in an embodiment of the present invention;
[0027] Figure 5 This is a schematic diagram of an alarm convergence method provided in an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Currently, IT operations and maintenance management systems are becoming increasingly large and complex. However, traditional operations and maintenance personnel spend most of their time and energy dealing with simple, repetitive tasks, which is not only inefficient but can also lead to human error and distraction, making it difficult to concentrate on more complex and creative tasks.
[0030] At the same time, as the number of hosts involved in the business system increases and the types of services provided by the applications also increase, the scale of monitoring indicators also increases, and the number of alarms also increases. If the operation and maintenance personnel do not reduce the number of alarms, the alarms may become too numerous and exceed the recipient's capacity, causing the recipient to be annoyed by the alarms or question the alarms, ultimately reducing the utilization value of the monitoring system.
[0031] Based on this, the present invention provides a method for processing alarm self-healing systems and an alarm self-healing system, specifically providing a method for alarm convergence and automated operation and maintenance fault handling based on Zabbix. This method can change the previous passive operation and maintenance mode, and establish an alarm self-healing system through automated alarms, automated event processing and other means, thereby reducing manual intervention and improving operation and maintenance efficiency.
[0032] Zabbix is an enterprise-grade open-source solution that provides distributed system monitoring and network monitoring functions based on a web (World Wide Web) interface.
[0033] Existing ITIL (Information Technology Infrastructure Library) systems have significant advantages in operation and maintenance services, but lack automated monitoring and prevention of IT equipment; while the Zabbix monitoring system provides platform support for monitoring equipment, it lacks unified management and service processes for equipment operation and maintenance. Therefore, effectively combining the advantages of both and complementing each other can greatly enhance the quality and efficiency of IT operation and maintenance.
[0034] This embodiment, combining ITIL and Zabbix, can change the previous passive operation and maintenance model. By establishing a self-healing system through automated alarms and automated event handling, it reduces manual intervention, improves operation and maintenance efficiency, and helps operation and maintenance personnel to be freed from tedious and repetitive tasks, allowing them to focus on system optimization, performance tuning, and innovative projects, thereby improving overall operation and maintenance efficiency and system stability.
[0035] The alarm convergence algorithm provided in this embodiment can converge alarm data, solve the problem of alarm redundancy, improve the efficiency of operation and maintenance personnel in obtaining alarm information, and help them quickly find the root cause of the fault from a large number of alarm messages. At the same time, it reduces the cost of alarm information storage, SMS or email notifications, etc.
[0036] To facilitate understanding of this embodiment, a detailed description of an alarm self-healing system disclosed in this embodiment of the invention will be provided first.
[0037] Example 1:
[0038] This invention provides a processing method for an alarm self-healing system, applicable to the alarm self-healing system, see [link to documentation]. Figure 1 The diagram shows the structure of an alarm self-healing system, which includes a monitoring module, an event handling module, and a work order module, which are connected in sequence.
[0039] Based on the above description, see Figure 2 The flowchart shown illustrates a processing method for an alarm self-healing system, which includes the following steps:
[0040] Step S202: Collect monitoring data from the monitored equipment.
[0041] like Figure 1 As shown, the monitoring module in this embodiment can collect monitoring data from the monitored device and send the monitoring data to the event handling module.
[0042] See also Figure 3 The diagram shows the overall design framework of an alarm self-healing system. The monitoring module can be a Zabbix monitoring module. The Zabbix monitoring module is responsible for collecting monitoring data from the monitored devices and can transmit the monitoring data to the event handling module through a custom alarm medium.
[0043] In some embodiments, the alarm self-healing system further includes: a display interface; the display interface displays monitoring data.
[0044] In this embodiment, the Zabbix monitoring module can display the collected relevant data information on the display interface.
[0045] In some embodiments, if the monitored device is a hardware device, the monitoring module collects monitoring data through a specified protocol; if the monitored device is an operating system, the monitoring module collects monitoring data through a monitoring agent.
[0046] like Figure 3 As shown, for hardware device monitoring, this embodiment can use protocols such as SNMP (Simple Network Management Protocol) and ICMP (Internet Control Message Protocol) to collect monitoring data; for operating systems such as Windows and Linux, this embodiment can use the monitoring agent zabbix-agent to collect data.
[0047] Step S204: Determine the operation and maintenance alarm events based on monitoring data, determine the self-healing process of the operation and maintenance alarm events based on pre-set matching rules, and process the operation and maintenance alarm events based on the self-healing process.
[0048] like Figure 1 As shown, the event handling module in this embodiment determines the operation and maintenance alarm event based on monitoring data, determines the self-healing process of the operation and maintenance alarm event based on the pre-set matching rules, and processes the operation and maintenance alarm event based on the self-healing process.
[0049] like Figure 3 As shown, the event handling module is used for alarm event convergence. Specifically, the event handling module receives monitoring data sent by the Zabbix monitoring module, performs multi-metric aggregation detection, and identifies operational alarm events. The event handling module can be pre-configured with matching rules, matches operational alarm events based on these rules, determines the self-healing processing flow for operational alarm events, and executes the self-healing processing flow to automatically handle operational alarm events.
[0050] In some embodiments, the event handling module determines whether the operation and maintenance alarm event has reached a threshold based on matching rules; if the operation and maintenance alarm event has not reached the threshold, the event handling module determines the self-healing process of the operation and maintenance alarm event and processes the operation and maintenance alarm event based on the self-healing process; if the operation and maintenance alarm event has reached the threshold, the event handling module determines that there is no self-healing process for the operation and maintenance alarm event; if the operation and maintenance alarm event has not reached the threshold, the work order module generates a processing work order for the operation and maintenance alarm event; if the operation and maintenance alarm event has reached the threshold, the work order module generates an alarm work order for the operation and maintenance alarm event.
[0051] In this embodiment, the event handling module can determine whether an operational alarm event has reached a threshold based on matching rules. If the operational alarm event has not reached the threshold, it is automatically processed through a self-healing process. In this case, the work order module needs to generate a processing work order representing the result of the self-healing process for the operational alarm event. If the operational alarm event has reached the threshold, it requires manual processing by operations personnel. In this case, the work order module needs to generate a processing work order prompting operations personnel to manually process the operational alarm event.
[0052] Step S206: Generate a work order for the operation and maintenance alarm event; wherein, the work order includes: a processing work order or an alarm work order; wherein, the processing work order is used to characterize the result of the self-healing process of the operation and maintenance alarm event; the alarm work order is used to prompt the operation and maintenance personnel to manually handle the operation and maintenance alarm event.
[0053] like Figure 1As shown, the work order module generates work orders for operation and maintenance alarm events; among them, work orders include: processing work orders or alarm work orders; wherein, processing work orders are used to characterize the result of the self-healing process of operation and maintenance alarm events; alarm work orders are used to prompt operation and maintenance personnel to manually handle operation and maintenance alarm events.
[0054] like Figure 3 As shown, the work order module can be used with the ITIL work order module. Figure 3 As shown, the work order module can generate processing work orders or alarm work orders. After the operation and maintenance alarm event is resolved, the work order module can send a processing work order, and the service will return to normal.
[0055] This invention provides a method for processing alarm self-healing systems. The method involves collecting monitoring data from the monitored equipment; determining maintenance alarm events based on the monitoring data; determining a self-healing processing flow for the maintenance alarm events based on pre-set matching rules; processing the maintenance alarm events based on the self-healing processing flow; and generating work orders for the maintenance alarm events. Each work order includes either a processing work order or an alarm work order. The processing work order characterizes the result of the self-healing processing flow for the maintenance alarm events, while the alarm work order prompts maintenance personnel to manually handle the maintenance alarm events. This method can change the previous passive maintenance model by establishing an alarm self-healing system through automated alarms and automated event processing, thereby reducing manual intervention and improving maintenance efficiency.
[0056] Example 2:
[0057] This embodiment provides an alarm convergence method, which is implemented based on the above embodiments. This embodiment focuses on describing an alarm convergence method that can be applied to the event handling module of the alarm self-healing system provided in the aforementioned embodiments.
[0058] See Figure 4 The flowchart shown represents an alarm convergence method, which includes the following steps:
[0059] Step S402, based on monitoring data within a defined time window.
[0060] In some embodiments, the monitored device is equipped with a monitoring agent, and the alarm self-healing system is equipped with a monitoring service; the monitoring agent sends monitoring data to the monitoring service at a specified frequency.
[0061] In this embodiment, a monitoring service called Zabbix Server can be installed on the main control server of the alarm self-healing system, and a monitoring agent called Zabbix Agent can be installed on the controlled servers of each monitored device. The controlled servers push data to the main control server at a certain frequency through the Zabbix Agent, and the data is stored in the database corresponding to the Zabbix Server.
[0062] In some embodiments, the monitoring data includes: hardware monitoring metrics, system performance monitoring metrics, and business monitoring metrics.
[0063] For example, hardware monitoring metrics for a controlled server may include disk bad sectors and disk damage; system performance monitoring metrics for a controlled server may include CPU load and memory utilization; and business monitoring metrics for a controlled server may include the number of concurrent requests and response latency.
[0064] While the Zabbix server continuously acquires monitoring metrics from each controlled server, it performs anomaly detection on each metric. In this embodiment, anomaly detection can be performed using time windows to determine if any operational alarm events exist.
[0065] Step S404: Determine whether there are any operation and maintenance alarm events within the time window based on the monitoring data within the time window.
[0066] In this embodiment, W can be set as the number of time windows. Let W be the value of the indicator collected in the i-th time window (i.e., the value of the i-th indicator), where 1 ≤ k ≤ W, and S is the number of times the indicator is collected within that time window (1 ≤ i ≤ S). This embodiment can continuously slide the time window to detect whether each indicator is normal. Abnormal indicators in this embodiment include abnormal values and abnormal fluctuations.
[0067] In some embodiments, if no monitoring data exists within a time window, it is determined that an operation and maintenance alarm event exists within the time window.
[0068] If monitoring data from the controlled server cannot be obtained within a certain time window, the controlled server is defined as having an abnormal value retrieval.
[0069] In some embodiments, the fluctuation value of the monitoring data metric within the time window is determined; if the fluctuation value of the metric is greater than a preset fluctuation threshold, it is determined that an operation and maintenance alarm event exists within the time window.
[0070] This embodiment can determine whether there is any abnormal fluctuation by comparing the fluctuation value of the indicator with a preset fluctuation threshold.
[0071] Specifically, standard deviation can be used to measure the magnitude of volatility. In some embodiments, the volatility of an indicator for monitoring data within a time window can be calculated using the following formula: in, Let be the fluctuation value of the i-th indicator in the monitoring data within the k-th time window, and S be the number of times the indicator is collected within the k-th time window. Let i be the value of the i-th metric of the monitoring data within the k-th time window. It represents the average value of the i-th indicator of the monitoring data within the k-th time window.
[0072] The following formula can be used to calculate it.
[0073] For example, the second indicator (i=2) of the monitoring data in the first time window (i.e., k=1) is a hardware monitoring indicator, collected 3 times (i.e., S=3), and the values of the hardware monitoring indicator are 10, 20, and 30 respectively (i.e., ... Therefore, we can calculate first. Recalculate The fluctuation value of the second indicator (i.e., hardware monitoring indicator) of the monitoring data in the first time window is 10.
[0074] In the calculation After that, if This indicates an abnormal fluctuation, and an abnormal fluctuation alarm can be issued. Wherein, δ i This is the preset fluctuation threshold for the i-th indicator.
[0075] The alarm self-healing method provided in this embodiment can detect alarms in real time, perform pre-diagnostic analysis through an alarm convergence algorithm, automatically recover from faults, and connect with surrounding systems to achieve rapid fault recovery. By improving the availability of business systems and reducing manpower investment in troubleshooting, it achieves a shift from "manual handling" to "unattended operation."
[0076] To address the need for fault self-healing capabilities in operations and maintenance, a visual operations and maintenance configuration module is provided. By giving users the ability to customize and edit fault self-healing processes, it eliminates the need for manual handling of alarms. Users only need to pre-program the alarm handling process, and the event handling module will automatically trigger it according to the scenario, thereby achieving fault self-healing.
[0077] Let's take automatic disk cleanup as an example:
[0078] Performance requirement: When the server disk usage exceeds 90%, trigger the automatic cleanup policy to free up disk space.
[0079] Step 1: Include the servers that need to be managed in the platform for Zabbix monitoring, and set the monitor to issue a critical alert when disk usage exceeds 90%.
[0080] Step 2: Enter the matching rules and self-healing process orchestration of the event handling module, and create an automatic disk full cleanup strategy. Based on the actual troubleshooting process, plan the self-healing process by dragging and dropping strategy nodes.
[0081] Step 3: Configure the triggering method. The triggering strategy is activated when an alarm message is received from the Zabbix alarm self-healing system.
[0082] See also Figure 5 The diagram illustrates an alarm convergence method. This embodiment can start from obtaining accurate alarms, proceed to pre-diagnosis analysis, determine the alarm type and level, trigger a self-healing process for general alarms, and automatically recover the platform. For serious and complex alarms, the system notifies the operation and maintenance personnel through alarm notifications, alarm work orders, etc., for manual handling, thereby achieving rapid fault recovery.
[0083] The alarm self-healing provided in this embodiment of the invention can automatically complete the processing of alarm events according to matching rules and processing event flow after receiving an alarm, reducing manual intervention; it can also automatically handle alarms without the need for personnel to be on duty 24 / 7; it can also obtain monitoring alarms to detect anomalies, perform data analysis, and detect and handle faults in advance.
[0084] The alarm self-healing method provided in this invention can detect alarms in real time, perform pre-diagnostic analysis through an alarm convergence algorithm, automatically recover from faults, and connect with surrounding systems to achieve rapid fault recovery. By improving the availability of business systems and reducing manpower investment in troubleshooting, it achieves a shift from "manual handling" to "unattended operation."
[0085] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the alarm self-healing method described above can be referred to the corresponding process in the aforementioned embodiments of the alarm self-healing system, and will not be repeated here.
[0086] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0087] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0088] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0089] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for processing alarm self-healing systems, characterized in that, The method, applied to an alarm self-healing system, includes: Collect monitoring data from the monitored equipment; Based on the monitoring data, an operation and maintenance alarm event is determined; based on the pre-set matching rules, a self-healing process for the operation and maintenance alarm event is determined; and the operation and maintenance alarm event is processed based on the self-healing process. Generate a work order for the operation and maintenance alarm event; wherein the work order includes: a processing work order or an alarm work order; wherein the processing work order is used to characterize the result of the self-healing process of the operation and maintenance alarm event; and the alarm work order is used to prompt operation and maintenance personnel to manually handle the operation and maintenance alarm event.
2. The method according to claim 1, characterized in that, The alarm self-healing system includes a monitoring module, an event handling module, and a work order module, which are connected sequentially. The steps for collecting monitoring data from the monitored device include: the monitoring module collecting monitoring data from the monitored device and sending the monitoring data to the event handling module; The steps of determining an operation and maintenance alarm event based on the monitoring data, determining a self-healing process for the operation and maintenance alarm event based on pre-set matching rules, and processing the operation and maintenance alarm event based on the self-healing process include: the event handling module determining an operation and maintenance alarm event based on the monitoring data, determining a self-healing process for the operation and maintenance alarm event based on pre-set matching rules, and processing the operation and maintenance alarm event based on the self-healing process. The step of generating a work order for the maintenance alarm event includes: the work order module generating a work order for the maintenance alarm event.
3. The method according to claim 2, characterized in that, The monitoring module is a Zabbix monitoring module, and the work order module is an ITIL work order module. The alarm self-healing system further includes a display interface; the method further includes: the display interface displays the monitoring data.
4. The method according to claim 2, characterized in that, The steps for the monitoring module to collect monitoring data from the monitored device include: If the monitored device is a hardware device, the monitoring module collects the monitoring data through a specified protocol; If the monitored device is an operating system, the monitoring module collects the monitoring data through a monitoring agent.
5. The method according to claim 2, characterized in that, The event handling module determines the self-healing process for the operation and maintenance alarm event based on pre-set matching rules, and the steps for handling the operation and maintenance alarm event based on the self-healing process include: The event handling module determines whether the operation and maintenance alarm event has reached the threshold based on the matching rules; If the maintenance alarm event does not reach the threshold, the event handling module determines the self-healing process of the maintenance alarm event and processes the maintenance alarm event based on the self-healing process. If the operation and maintenance alarm event reaches the threshold, the event handling module is used to determine that the operation and maintenance alarm event does not have a self-healing process; The steps of generating a work order for the maintenance alarm event by the work order module include: If the maintenance alarm event does not reach the threshold, the work order module generates a processing work order for the maintenance alarm event; If the maintenance alarm event reaches the threshold, the work order module generates an alarm work order for the maintenance alarm event.
6. The method according to any one of claims 1-5, characterized in that, The steps for determining operation and maintenance alarm events based on the monitoring data include: Based on monitoring data within a defined time window; Based on the monitoring data within the time window, determine whether there are any operation and maintenance alarm events within the time window.
7. The method according to claim 6, characterized in that, The monitoring data includes: hardware monitoring indicators, system performance monitoring indicators, and business monitoring indicators; the monitored device is equipped with a monitoring agent, and the alarm self-healing system is equipped with a monitoring service; the monitoring agent sends the monitoring data to the monitoring service at a specified frequency.
8. The method according to claim 6, characterized in that, The steps for determining whether there are maintenance alarm events within the time window based on the monitoring data within the time window include: If the monitoring data is not present within the time window, it is determined that an operation and maintenance alarm event exists within the time window.
9. The method according to claim 6, characterized in that, The steps for determining whether there are maintenance alarm events within the time window based on the monitoring data within the time window include: Determine the fluctuation value of the monitoring data indicators within the time window; If the fluctuation value of the indicator is greater than the preset fluctuation threshold, it is determined that there is an operation and maintenance alarm event within the time window; The fluctuation value of the monitoring data indicator within the time window is calculated using the following formula: in, Let S be the fluctuation value of the i-th indicator of the monitored data within the k-th time window, and let S be the number of times the indicator is collected within the k-th time window. The value of the i-th indicator of the monitoring data within the k-th time window. It is the average value of the i-th indicator of the monitoring data within the k-th time window.
10. An alarm self-healing system, characterized in that, The alarm self-healing system is used to execute the processing method of the alarm self-healing system according to any one of claims 1-9.