Alarm information convergence processing method and device

By using the configuration management database CMDB in the alarm analysis server of the data center to determine the associated services of the faulty equipment and perform multiple rounds of alarm information convergence processing, the problem of excessive alarm information and inability to effectively troubleshoot problems in the existing technology is solved, exponential convergence and intuitive description of alarm information is achieved, SMS storm and sending costs are reduced, and strong technical support is provided for fault alarms and operation and maintenance of the data center.

CN120196499APending Publication Date: 2025-06-24NETSUNION CLEARING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311789686.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing data center alarm processing methods cannot be linked in conjunction with context, resulting in too large alarm information and difficult to effectively troubleshoot problems, and lack of advanced alarm rules.

Method used

The alarm analysis server obtains multiple equipment failure information, uses the configuration management database CMDB to determine the associated service of the faulty device, performs a first convergence process to obtain the first alarm information related to the equipment and services, and performs a second convergence process through service layer feedback to obtain the second alarm information related to the equipment and services, and finally push it to the relevant personnel.

Benefits of technology

The exponential convergence of alarm information is achieved, making the description of alarm information more intuitive, effectively reducing SMS storms, reducing SMS sending costs, and providing strong technical support for fault alarms and operation and maintenance of data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196499A_ABST
    Figure CN120196499A_ABST
Patent Text Reader

Abstract

The invention provides an alarm information convergence processing method and device, and relates to the technical field of computer data processing.The method is executed by an alarm analysis server and comprises the steps that multiple pieces of monitored equipment fault information are obtained, and each piece of equipment fault information corresponds to fault equipment; according to the equipment information stored in the configuration management database CMDB, determining associated services of the fault equipment; carrying out first convergence processing on the multiple pieces of equipment fault information according to the associated service of each piece of fault equipment to obtain first alarm information related to the equipment and the service; acquiring service alarm information fed back by the service layer, and performing second convergence processing according to the first alarm information and the service alarm information to obtain second alarm information related to the equipment and the service; and pushing the second alarm information to related personnel. According to the invention, exponential convergence of alarm information can be realized, the short message sending cost is reduced, and powerful technical support is provided for fault alarm and operation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data processing, and in particular to a method and device for processing convergence of alarm information. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the invention recited in the claims. No admission is made that the description herein is prior art by inclusion in this section.

[0003] When a data center alarm occurs, it usually contains a large amount of information. For example, a router failure may affect multiple connected devices and be accompanied by a large amount of business alarm information. At this time, the mobile phone will receive a large number of alarm messages, which is like being attacked by SMS floods, and it is impossible to understand the root cause of the failure in the first place. At the same time, all business teams collectively report the impact of the failure, which will cause more noise in troubleshooting.

[0004] Common data center alarm devices are generally one layer of logic, and monitor according to the target threshold. When a failure occurs, the threshold rule judgment is performed and the alarm measures are executed; however, the existing alarm processing methods have the following main defects: 1. It cannot be linked with the context. Because the alarm rules are only one layer of logic, it is impossible to form a context linkage with other monitoring targets. For example, when a router device fails, it is impossible to combine the device connection status (generally connected to multiple physical servers) for linkage processing. 2. The alarm information is too large, which is not conducive to problem troubleshooting. The settings may cause the mobile phone to not work properly. 3. Lack of advanced alarm rules. The current alarm rules are relatively simple and cannot support more flexible rule settings.

[0005] In summary, there is an urgent need for a technical solution that can overcome the above-mentioned defects and improve the way of processing alarm information. Summary of the invention

[0006] In order to solve the problems existing in the prior art, the present invention proposes a method and device for processing alarm information convergence.

[0007] In a first aspect of an embodiment of the present invention, a method for processing alarm information convergence is proposed. The method is executed by an alarm analysis server and includes:

[0008] Acquire multiple pieces of monitored equipment failure information, each piece of equipment failure information corresponds to a failed equipment respectively;

[0009] Determine the associated business of each of the faulty devices according to the device information stored in the configuration management database CMDB;

[0010] Performing first convergence processing on the plurality of pieces of equipment fault information according to the associated services of each of the faulty equipment to obtain first alarm information related to both the equipment and the service;

[0011] Obtain the service alarm information fed back by the service layer, and perform a second convergence process according to the first alarm information and the service alarm information to obtain second alarm information related to both the device and the service;

[0012] Push the second alarm information to relevant personnel.

[0013] In the second aspect of the embodiments of the present invention, an alarm information convergence processing device is provided, which is set in an alarm analysis server and includes:

[0014] A fault information acquisition module, configured to acquire multiple pieces of device fault information monitored, and each of the device fault information corresponds to a faulty device respectively;

[0015] A fault information analysis module, configured to determine the associated services of each of the faulty devices according to the device information stored in the Configuration Management Database (CMDB);

[0016] A first convergence processing module, configured to perform a first convergence process on the multiple pieces of device fault information according to the associated services of each of the faulty devices to obtain first alarm information related to both the device and the service;

[0017] A second convergence processing module, configured to obtain the service alarm information fed back by the service layer, and perform a second convergence process according to the first alarm information and the service alarm information to obtain second alarm information related to both the device and the service;

[0018] An information push module, configured to push the second alarm information to relevant personnel.

[0019] In the third aspect of the embodiments of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, an alarm information convergence processing method is implemented.

[0020] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, an alarm information convergence processing method is implemented.

[0021] The alarm information convergence processing method and device proposed by the present invention can obtain multiple device fault messages monitored, and each of the device fault messages corresponds to a faulty device respectively; determine the associated services of each faulty device according to the device information stored in the Configuration Management Database (CMDB); perform a first convergence process on the multiple device fault messages according to the associated services of each faulty device to obtain first alarm information related to both devices and services; obtain the service alarm information fed back by the service layer, and perform a second convergence process according to the first alarm information and the service alarm information to obtain second alarm information related to both devices and services; push the second alarm information to relevant personnel; the overall solution realizes exponential convergence of alarm information, makes the description of alarm information more intuitive, effectively reduces the SMS storm, reduces the SMS sending cost, and provides strong technical support for fault alarm and operation and maintenance of the data center. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a schematic flowchart of the alarm information convergence processing method according to an embodiment of the present invention.

[0024] Figure 2 It is a schematic architecture diagram of an exemplary scenario of the present invention.

[0025] Figure 3 It is a schematic diagram of the alarm information convergence processing relationship according to an embodiment of the present invention.

[0026] Figure 4 It is a schematic flowchart of determining the associated services of each faulty device according to the device information stored in the Configuration Management Database (CMDB) according to an embodiment of the present invention.

[0027] Figure 5 It is a detailed schematic flowchart of determining the associated services of each faulty device for faulty devices of different device types according to an embodiment of the present invention.

[0028] Figure 6 It is a schematic flowchart of performing a first convergence process on the multiple device fault messages according to the associated services of each faulty device to obtain first alarm information related to both devices and services according to an embodiment of the present invention.

[0029] Figure 7It is a schematic flowchart of the second convergence process according to the first alarm information and the service alarm information in an embodiment of the present invention to obtain the second alarm information related to both the device and the service.

[0030] Figure 8 It is a schematic structural diagram of an alarm information convergence processing device in an embodiment of the present invention.

[0031] Figure 9 It is a schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0032] Next, the principles and spirit of the present invention will be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, rather than limiting the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to convey the scope of the present disclosure fully to those skilled in the art.

[0033] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, a device, an equipment, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0034] According to the embodiments of the present invention, an alarm information convergence processing method and device are proposed, which relate to the technical field of computer data processing. The present invention can implement the convergence processing of alarm information in the data center, make the description of alarm information more intuitive, effectively reduce alarm noise, and prevent the generation of alarm SMS storms.

[0035] In the embodiments of the present invention, the terms to be explained are:

[0036] Data center alarm information: Specifically refers to a series of SMS or email alarms caused by power failure or network failure of data center devices.

[0037] Service alarm information: Alarm information customized by the business team for the business health status.

[0038] Next, with reference to several representative embodiments of the present invention, the principles and spirit of the present invention will be explained in detail.

[0039] Figure 1 It is a schematic flowchart of an alarm information convergence processing method in an embodiment of the present invention. As Figure 1 shown, this method is executed by an alarm analysis server and includes:

[0040] S101. Obtain multiple pieces of monitored device fault information, where each piece of the device fault information corresponds to a faulty device respectively;

[0041] S102. Determine the associated services of each of the faulty devices according to the device information stored in the Configuration Management Database (CMDB);

[0042] S103. Perform a first convergence process on the multiple pieces of device fault information according to the associated services of each of the faulty devices to obtain first alarm information related to both the devices and the services;

[0043] S104. Obtain the service alarm information fed back by the service layer, and perform a second convergence process according to the first alarm information and the service alarm information to obtain second alarm information related to both the devices and the services;

[0044] S105. Push the second alarm information to relevant personnel.

[0045] In an actual application scenario, the alarm information convergence processing method proposed by the present invention can be executed by an alarm analysis server. The working principle is as follows: Obtain multiple pieces of monitored device fault information, where each piece of the device fault information corresponds to a faulty device respectively; Determine the associated services of each of the faulty devices according to the device information stored in the Configuration Management Database (CMDB); Perform a first convergence process on the multiple pieces of device fault information according to the associated services of each of the faulty devices to obtain first alarm information related to both the devices and the services; Obtain the service alarm information fed back by the service layer, and perform a second convergence process according to the first alarm information and the service alarm information to obtain second alarm information related to both the devices and the services; Push the second alarm information to relevant personnel. The present invention can achieve exponential convergence of alarm information, make the description of alarm information more intuitive, effectively reduce the SMS storm, reduce the SMS sending cost, and provide strong technical support for the fault alarm and operation and maintenance of the data center.

[0046] For a clearer explanation of the above data center alarm information convergence processing method, the following will be combined with Figure 2 the exemplary scenario of Figure 3 and the relationship schematic diagram shown in

[0047] In one embodiment, referring to Figure 2 the exemplary scenario of

[0048] Specifically, in combination withFigure 3 The schematic diagram of the relationship shown is used to elaborate in detail on the process of converging and processing the alarm information of the alarm analysis server.

[0049] In one embodiment, (S101) Obtain multiple pieces of device fault information monitored, and each of the device fault information corresponds to a faulty device respectively.

[0050] In an actual application scenario, when a fault occurs in the data center, attempt to push the monitored fault information to the alarm analysis service for in-depth analysis and processing. If an anomaly is analyzed, maintain the original method and directly send a text message alarm so that personnel can receive the fallback alarm text message.

[0051] In one embodiment, (S102) Determine the associated services of each of the faulty devices according to the device information stored in the Configuration Management Database (CMDB).

[0052] Refer to Figure 4 , and the specific process is as follows:

[0053] S401, Through the IP, host name, or device name in the device fault information, combined with the device information stored in the CMDB, query the corresponding device type; wherein, the device type includes basic devices and key devices;

[0054] S402, For the faulty devices with the device type of basic devices, execute the step of determining the associated services of each of the faulty devices.

[0055] Specifically, Figure 5 is the detailed process schematic diagram for determining the associated services of faulty devices of different device types in an embodiment of the present invention.

[0056] Refer to Figure 5 As shown, for the faulty devices with the device type of basic devices, the detailed process of determining the associated services of each of the faulty devices is as follows:

[0057] S501, For the faulty devices with the device type of basic devices, combine the device information stored in the CMDB to query the directly associated and indirectly associated devices as the first target devices, and further query the services associated with the first target devices as the associated services of the faulty devices.

[0058] In this embodiment, the faulty devices with the device type of basic devices can be power supplies, cabinets, etc. For the faulty devices with the device type of basic devices, query the devices associated with the faulty device, and further query the services affected by the device; Converge all the fault information associated with the basic device and the services affected by the device within the monitoring period into one piece of fault information.

[0059] The impact caused by basic equipment failures is relatively large, so it is only necessary to inform the affected service information. Specifically, the converged fault information includes: the device name of the faulty device and the service names affected; the format of the fault information is: Device A fails, affecting services B and C.

[0060] For faulty devices of the key device type, refer to Figure 5 As shown, the detailed process for converging fault information is as follows:

[0061] S502, for faulty devices of the key device type, query the devices directly associated with it as the second target devices;

[0062] S503, converge all the fault information associated with the faulty device and the second target devices within the monitoring period into one fault information; among them, the fault information includes the identification of the faulty device and the identification of the second target device.

[0063] In this embodiment, the faulty devices of the key device type can be network devices, firewalls, etc.

[0064] The impact scope caused by key device failures is smaller than that of basic equipment failures, and it affects the directly connected devices, such as servers, etc. The identification of the faulty device in the fault information can be the name of the faulty device, and the identification of the second target device can be the name of the server affected by the faulty device; the format of the fault information is: Device D fails, affecting servers E and F.

[0065] In the actual application scenario, it may also occur that the faulty device information or the associated devices cannot be queried. Refer to Figure 5 As shown, the specific processing method is:

[0066] S504, if the faulty device information or the associated devices cannot be queried, generate an analysis failure message and push it to the relevant personnel.

[0067] In one embodiment, the device type at least includes: power supply, cabinet, network device, security device, firewall, physical machine server, virtual machine, and bare metal server.

[0068] In the actual application scenario, if the alarm analysis service has an analysis exception, directly send an alarm message to the relevant personnel via text message.

[0069] CMDB (Configuration Management Database) is used to record and manage the device configuration information of the data center, which can help the relevant personnel better manage the devices and services, and quickly respond to and handle device failures. The main functions include:

[0070] Data entry: Enter the basic information of the device into the CMDB, including device name, IP address, device type, affiliated business, location of the computer room, manufacturer, etc.

[0071] Data update: As the device changes, it is necessary to update the device information stored in the CMDB in a timely manner, including the IP address of the device, affiliated business, location of the computer room, etc.

[0072] Device fault analysis: When a device fails, the faulty device can be quickly located through the CMDB, and the relevant information of the device, such as affiliated business, location of the computer room, manufacturer, etc., can be viewed.

[0073] Device impact analysis: Through the device relationships stored in the CMDB, devices and services related to the faulty device can be quickly found, and then the impact of the faulty device on the business can be analyzed.

[0074] Fault handling: Through the device information and device relationships stored in the CMDB, the faulty device can be quickly located, and the corresponding maintenance personnel can be notified for handling.

[0075] Fault record: Record the fault information in the CMDB, including information such as faulty device, fault time, fault cause, handling personnel, handling time, etc., for future query and analysis.

[0076] In summary, using the CMDB can better manage devices and services, and quickly respond to and handle device failures. Through the CMDB, the faulty device can be quickly located, and the impact of the faulty device on the business can be analyzed, thereby improving the overall operation and maintenance efficiency.

[0077] In one embodiment, (S103) perform a first convergence process on the multiple pieces of device fault information according to the associated business of each of the faulty devices to obtain first warning information related to both the device and the business.

[0078] Reference Figure 6 , the specific process is as follows:

[0079] S601, in the case where the number of faulty devices is multiple, perform deduplication and convergence on the device fault information according to the multiple faulty devices to obtain the converged device fault information; the device fault information includes a faulty device identifier;

[0080] Specifically, the faulty device identifier can be the device name of the faulty device.

[0081] S602, according to the converged device fault information, mark the affected associated businesses in the associated business of the faulty device to obtain first warning information related to both the device and the business.

[0082] Specifically, for the device fault information of multiple faulty devices, based on the information such as the business to which the device belongs, the computer room where it is located, and the manufacturer in the device fault information, the device fault information with an information coincidence degree reaching the set threshold can be selected for convergence processing, which can avoid the interference of a large number of duplicate fault information to operation and maintenance personnel and business personnel, and improve the efficiency and accuracy of fault handling.

[0083] In actual application scenarios, continuously monitor and collect various device fault information generated by the system, and classify and count it for subsequent processing. Based on the device fault information, analyze the collected device fault information to find the same or similar information, which can facilitate the subsequent induction of the cause of the fault, the scope of influence, and the solution.

[0084] Information coincidence degree calculation: Calculate the coincidence degree of information such as the business to which the device belongs, the computer room where it is located, and the manufacturer in the device fault information, and select the device fault information with a coincidence degree reaching the set threshold for convergence processing. Converge the selected device fault information and merge it into one fault information, which can avoid information overload for operation and maintenance personnel and business personnel, and can also avoid information duplication and chaos. Through the convergence processing of device fault information, a large number of duplicate fault information can be avoided from interfering with operation and maintenance personnel and business personnel, improving the efficiency and accuracy of fault handling. At the same time, it also helps to deeply analyze and optimize the problems of the system or application.

[0085] In one embodiment, (S104) Obtain the business alarm information fed back by the business layer, and perform a second convergence processing according to the first alarm information and the business alarm information to obtain the second alarm information related to both the device and the business.

[0086] Reference Figure 7 , the specific process is as follows:

[0087] S701, Analyze the error code information in the business alarm information, and perform a third convergence processing on the business alarm information to obtain the third alarm information;

[0088] S702, Perform another convergence processing according to the consistent key fields in the third alarm information and the first alarm information to obtain the second alarm information related to both the device and the business.

[0089] Specifically, in S701, analyze the error code information in the business alarm information, and perform a third convergence processing on the business alarm information to obtain the third alarm information. The detailed process is as follows:

[0090] Analyze the error code information in the business alarm information, find the error codes with the same or similarity reaching the set threshold, and merge them into one error code information.

[0091] Specifically, in S702, a convergence process is performed again based on the keyword fields that are the same in the third alarm information and the first alarm information to obtain second alarm information related to both the device and the service, including:

[0092] Compare the error code information with the keyword fields in the first alarm information, and merge the error code alarm information corresponding to the same device to obtain second alarm information related to both the device and the service; the format of the second alarm information is the name of the faulty device, the name of the affected server, the name of the service, and the error code alarm information that needs to be ignored.

[0093] When performing the convergence process on the error code alarm information, analyze the error code alarm information, find error codes that are the same or have a similarity reaching the set threshold, and merge them into one error code alarm information; if the IP or application name in the error code alarm information is consistent with the scope of the analysis result based on CMDB in (S102), then perform a convergence process on the error code alarm information (i.e., the convergence process in S702), and merge the error code alarm information of the same device to obtain the converged alarm information. The information format of the converged alarm information is the faulty device, the IP of the affected server, the name of the service, and the error code alarm information that needs to be ignored; through the above processing process, a large number of duplicate and invalid fault information can be reduced, the number of subsequent alarms and the number of alarm message pushes can be reduced, the interference to business personnel can be reduced, and the fault handling efficiency can be improved.

[0094] Specifically, when a device fails, there are also a large number of alarm information at the service level, and the general form is: System X fails, error code Y, and the IP is Z. Therefore, it is necessary to perform a convergence on the error code alarm information at the service level as well.

[0095] If the IP or application name in the error code alarm information is consistent with the scope in the previous CMDB analysis, then perform a convergence action on the error code alarm information again. Here, it is considered that the alarm information at the service level this time is also caused by a device failure, and it can be converged again based on the faulty device. The alarm information is converged to: Device L fails, the affected range includes the IPs of multiple servers, Service M, and the relevant error code N alarm can be ignored.

[0096] In the actual application scenario, through the error code convergence process, a large number of identical or similar error codes generated by the system or application can be converged and processed to avoid information overload on operation and maintenance personnel and business personnel and improve the fault handling efficiency.

[0097] Specifically, the error code processing process is as follows: monitoring and collection. The system or application generates various error codes. The monitoring system needs to continuously collect these error codes, classify and count them for subsequent processing. Analyze the collected error codes to find the same or similar error codes. Here, the cause of their generation, the scope of influence, and the solution can be understood through the error codes. Converge the same or similar error codes, merge them into one error message, and avoid information overload for operation and maintenance personnel and business personnel. Push the converged error message to relevant operation and maintenance personnel and business personnel to ensure that they only see one error message and avoid information duplication and chaos. Based on the converged error message, operation and maintenance personnel can quickly locate and handle problems, improving the efficiency of fault handling. Further, the converged error message can also be recorded and analyzed to understand the generation rules and trends of error codes for subsequent optimization of the system or application.

[0098] Through the convergence processing of error codes, a large number of repeated error messages can be avoided from interfering with operation and maintenance personnel and business personnel, improving the efficiency and accuracy of fault handling. At the same time, it also helps to deeply analyze and optimize the problems of the system or application.

[0099] In the actual application scenario, this method also includes:

[0100] If any convergence processing is abnormal, the corresponding fault information will be alarmed in a preset manner.

[0101] Specifically, when the convergence processing process in steps S103, S104, S701, S702, etc. is abnormal, the alarm can be triggered in a preset manner according to the fault information.

[0102] When the convergence processing is abnormal, it may cause the fault information not to be pushed to relevant personnel in time, thus affecting the efficiency of fault handling. To avoid this situation, a preset alarm mechanism for fault information can be set to notify relevant personnel in time in case of an abnormality. In the actual application scenario, preset alarm rules can be set, and corresponding preset alarm rules, including alarm levels, alarm methods, alarm objects, etc., can be set according to the importance and scope of influence of the fault information.

[0103] At the same time, monitor the status of error code convergence processing. If abnormal situations are found, such as processing timeouts, processing failures, etc., immediately trigger the preset alarm mechanism. When sending the alarm message, the alarm message can be sent to relevant personnel according to the preset alarm rules, including the content of the fault information, the processing status, the estimated processing time, etc.

[0104] After receiving the alarm information, relevant personnel need to process the fault information as soon as possible to ensure that the fault is resolved in a timely manner. In addition, the fault information and the processing process can be recorded and analyzed, recording the triggering situation and processing results of the preset alarm mechanism, and analyzing the processing efficiency and accuracy of the fault information, so as to optimize the system or application subsequently. Through the preset alarm mechanism, it can be ensured that the fault information is processed and notified in a timely manner under abnormal conditions, avoiding the omission and delay of the fault information, and improving the efficiency and accuracy of fault handling.

[0105] In one embodiment, (S105) the second alarm information is pushed to relevant personnel.

[0106] Specifically, the converged alarm information can be pushed to relevant personnel through alarm text messages. Relevant personnel (such as operation and maintenance personnel, business personnel) will only see one converged alarm text message, which will not cause a mobile phone text message explosion.

[0107] In an actual application scenario, if the alarm information convergence processing method proposed by the present invention fails to converge the alarm information, the original information text message can be directly sent to ensure that relevant personnel can still receive the alarm when the device is unavailable. Based on the CMDB data, the context combination mode is adopted to compress the alarm information and focus on the fault. It can also be linked with the upper-layer service monitoring to eliminate the noise related to business alarms.

[0108] The present invention can perform fault convergence processing in an organized and systematic manner, and use CMDB and alarm analysis services to quickly locate and process device faults.

[0109] When the fault information is pushed to the alarm analysis service, it can be ensured that the fault information is deeply analyzed and processed. At the same time, an alternative text message alarm method is also provided to ensure that personnel can receive the fault information in a timely manner. When analyzing in combination with CMDB data, querying the device type and associated devices through CMDB helps to quickly understand the impact range of the faulty device and its impact on the business. By converging a large number of fault information into one, the number of alarm information received by operation and maintenance personnel and business personnel can be reduced, improving the processing efficiency. Pushing the converged alarm text message can ensure that operation and maintenance personnel and business personnel only receive one converged alarm text message, avoiding a mobile phone text message explosion and improving the processing efficiency.

[0110] Overall, through multiple convergence processes, the present invention obtains the final alarm information related to both the device and the business, and pushes it to relevant personnel so that relevant personnel can process the fault information in a timely manner, helping relevant personnel to quickly and efficiently respond to and process device faults, and improving the overall operation and maintenance efficiency.

[0111] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0112] After introducing the method of the exemplary embodiment of the present invention, next, with reference to Figure 8 the alarm information convergence processing device of the exemplary embodiment of the present invention will be introduced.

[0113] The implementation of the alarm information convergence processing device can refer to the implementation of the above method, and the repeated parts will not be elaborated. The term "module" or "unit" used hereinafter can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0114] Based on the same inventive concept, the present invention also proposes an alarm information convergence processing device, as Figure 8 shown, this device is disposed in the alarm analysis server and includes:

[0115] A fault information acquisition module 810, configured to acquire multiple pieces of device fault information monitored, and each of the device fault information corresponds to a faulty device respectively;

[0116] A fault information analysis module 820, configured to determine the associated services of each of the faulty devices according to the device information stored in the configuration management database CMDB;

[0117] A first convergence processing module 830, configured to perform a first convergence processing on the multiple pieces of device fault information according to the associated services of each of the faulty devices, and obtain first alarm information related to both the device and the service;

[0118] A second convergence processing module 840, configured to acquire service alarm information fed back by the service layer, and perform a second convergence processing according to the first alarm information and the service alarm information, and obtain second alarm information related to both the device and the service;

[0119] An information push module 850, configured to push the second alarm information to relevant personnel.

[0120] In one embodiment, the first convergence processing module 830 performs a first convergence processing on the multiple pieces of device fault information according to the associated services of each of the faulty devices, and obtains first alarm information related to both the device and the service, including:

[0121] When the number of the faulty devices is multiple, deduplicate and converge the device fault information according to the multiple faulty devices to obtain the converged device fault information; the device fault information includes a faulty device identifier.

[0122] According to the converged device fault information, mark the affected associated services in the associated services of the faulty devices to obtain the first alarm information related to both the devices and the services.

[0123] In one embodiment, the fault information analysis module 820 determines the associated services of each faulty device according to the device information stored in the configuration management database CMDB, including:

[0124] Through the IP, host name or device name in the device fault information, query the corresponding device type in combination with the device information stored in CMDB; wherein, the device type includes basic devices and key devices.

[0125] For the faulty devices with the device type of basic devices, perform the process of determining the associated services of each faulty device.

[0126] The faulty device identifier can be the device name of the faulty device.

[0127] Specifically, the process of the fault information analysis module 820 determining the associated services of each faulty device includes:

[0128] For the faulty devices with the device type of basic devices, query the directly associated and indirectly associated devices in combination with the device information stored in CMDB as the first target devices, and further query the services associated with the first target devices as the associated services of the faulty devices.

[0129] Specifically, the fault information analysis module 820 is further configured to:

[0130] For the faulty devices with the device type of key devices, query the directly associated devices as the second target devices;

[0131] Converge all the fault information associated with the faulty device and the second target devices within the monitoring period into one piece of fault information; wherein, the fault information includes the faulty device identifier and the identifier of the second target device.

[0132] The faulty device identifier in the fault information can be the faulty device name, and the identifier of the second target device can be the server name affected by the faulty device.

[0133] Furthermore, the fault information analysis module 820 is further configured to:

[0134] If the fault device information or associated devices cannot be queried, generate an analysis failure message and push it to the relevant personnel.

[0135] In one embodiment, the second convergence processing module 840 obtains the service alarm information fed back by the service layer, performs second convergence processing based on the first alarm information and the service alarm information, and obtains second alarm information related to both the device and the service, including:

[0136] Analyze the error code information in the service alarm information, perform third convergence processing on the service alarm information, and obtain third alarm information;

[0137] Perform another convergence processing based on the consistent key fields in the third alarm information and the first alarm information, and obtain second alarm information related to both the device and the service.

[0138] Specifically, the second convergence processing module 840 analyzes the error code information in the service alarm information, performs third convergence processing on the service alarm information, and obtains third alarm information, including:

[0139] Analyze the error code information in the service alarm information, find the error codes with the same or similarity reaching the set threshold, and merge them into one error code information.

[0140] Specifically, the second convergence processing module 840 performs another convergence processing based on the consistent key fields in the third alarm information and the first alarm information, and obtains second alarm information related to both the device and the service, including:

[0141] Compare the error code information with the key fields in the first alarm information, merge the error code alarm information corresponding to the same device, and obtain second alarm information related to both the device and the service; the format of the second alarm information is the name of the faulty device, the name of the affected server, the service name, and the error code alarm information that needs to be ignored.

[0142] In one embodiment, for the first convergence processing module 830 and the second convergence processing module 840, they are also used for:

[0143] If any convergence processing is abnormal, alarm the corresponding fault information in a preset manner.

[0144] It should be noted that although several modules of the alarm information convergence processing device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.

[0145] Based on the foregoing inventive concept, asFigure 9 As shown in the figure, the present invention also provides a computer device 900, including a memory 910, a processor 920, and a computer program 930 stored on the memory 910 and executable on the processor 920. When the processor 920 executes the computer program 930, the foregoing alarm information convergence processing method is implemented.

[0146] Based on the foregoing inventive concept, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the foregoing alarm information convergence processing method is implemented.

[0147] Based on the foregoing inventive concept, the present invention provides a computer program product including a computer program, and when the computer program is executed by a processor, an alarm information convergence processing method is implemented.

[0148] The alarm information convergence processing method and device provided by the present invention can obtain multiple monitored device fault information, and each of the device fault information corresponds to a faulty device respectively; determine the associated services of each faulty device according to the device information stored in the configuration management database CMDB; perform a first convergence processing on the multiple device fault information according to the associated services of each faulty device to obtain first alarm information related to both the device and the service; obtain service alarm information fed back by the service layer, and perform a second convergence processing on the first alarm information and the service alarm information to obtain second alarm information related to both the device and the service; push the second alarm information to relevant personnel; the overall solution realizes exponential convergence of alarm information, makes the description of alarm information more intuitive, effectively reduces the SMS storm, reduces the SMS sending cost, and provides strong technical support for fault alarm and operation and maintenance of the data center.

[0149] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of laws and regulations.

[0150] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0151] The present invention is described with reference to the flowcharts and / or block diagrams of methods and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in one block or multiple blocks.

[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in one block or multiple blocks.

[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in one block or multiple blocks.

[0154] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, and are not intended to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for converging and processing alarm information, characterized in that This method is executed by an alarm analysis server and includes: Obtaining multiple pieces of device fault information that has been monitored, with each piece of device fault information corresponding to a faulty device respectively; Determining the associated services of each faulty device according to the device information stored in the Configuration Management Database (CMDB); Performing a first convergence process on the multiple pieces of device fault information according to the associated services of each faulty device to obtain first alarm information that is related to both devices and services; Obtaining service alarm information fed back by the service layer, and performing a second convergence process according to the first alarm information and the service alarm information to obtain second alarm information that is related to both devices and services; Pushing the second alarm information to relevant personnel.

2. The alarm information convergence processing method according to claim 1, wherein Performing a first convergence process on the multiple pieces of device fault information according to the associated services of each faulty device to obtain first alarm information that is related to both devices and services, including: In the case where the number of faulty devices is multiple, performing duplicate removal and convergence on the device fault information according to the multiple faulty devices to obtain converged device fault information; the device fault information includes a faulty device identifier; According to the converged device fault information, marking the affected associated services in the associated services of the faulty devices to obtain first alarm information that is related to both devices and services.

3. The alarm information convergence processing method according to claim 1, characterized in that Determining the associated services of each faulty device according to the device information stored in the Configuration Management Database (CMDB), including: Querying the corresponding device type through the IP, host name, or device name in the device fault information, in combination with the device information stored in the CMDB; wherein, the device type includes basic devices and key devices; For faulty devices with a device type of basic devices, execute the step of determining the associated services of each faulty device.

4. The alarm information convergence processing method according to claim 3, wherein, Determining the associated services of each faulty device, including: For faulty devices with a device type of basic devices, querying the devices directly and indirectly associated with them in combination with the device information stored in the CMDB as first target devices, and further querying the services associated with the first target devices as the associated services of the faulty devices.

5. The alarm information convergence processing method according to claim 3, wherein This method further includes: For faulty devices with a device type of key devices, querying the devices directly associated with them as second target devices; Converging all the fault information associated with the faulty device and the second target devices within the monitoring period into one piece of fault information; wherein, the fault information includes the faulty device identifier and the identifier of the second target device.

6. The alarm information convergence processing method according to claim 1, wherein Obtaining service alarm information fed back by the service layer, and performing a second convergence process according to the first alarm information and the service alarm information to obtain second alarm information that is related to both devices and services, including: Analyzing the error code information in the service alarm information and performing a third convergence process on the service alarm information to obtain third alarm information; Performing another convergence process according to the keyword fields that are consistent in the third alarm information and the first alarm information to obtain second alarm information that is related to both devices and services.

7. The alarm information convergence processing method according to claim 6, wherein Performing another convergence process according to the keyword fields that are consistent in the third alarm information and the first alarm information to obtain second alarm information that is related to both devices and services, including: Compare the error code information with the keyword fields in the first warning information, merge the error code warning information that is consistent for the same device, and obtain the second warning information that is related to both the device and the service; the format of the second warning information is the name of the faulty device, the name of the affected server, the name of the service, and the error code warning information that needs to be ignored.

8. An alarm information convergence processing device, characterized in that, It is set in the warning analysis server and includes: A fault information acquisition module, which is used to acquire multiple pieces of device fault information monitored, and each of the device fault information corresponds to a faulty device respectively; A fault information analysis module, which is used to determine the associated services of each of the faulty devices according to the device information stored in the Configuration Management Database (CMDB); A first convergence processing module, which is used to perform a first convergence processing on the multiple pieces of device fault information according to the associated services of each of the faulty devices, and obtain the first warning information that is related to both the device and the service; A second convergence processing module, which is used to acquire the service warning information fed back by the service layer, and perform a second convergence processing according to the first warning information and the service warning information, and obtain the second warning information that is related to both the device and the service; An information push module, which is used to push the second warning information to relevant personnel.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 7.