Alarm root cause output methods, devices, equipment, media and program products
By determining the alarm rules and equipment topology structure in the data center computer room infrastructure, identifying the alarm trigger status of the current monitoring data item, and using topological association relationships to quickly locate the root cause equipment, solving the problems of slow positioning speed and low accuracy of the alarm cause in the existing technology, improving operation and maintenance efficiency and security.
Patent Information
- Application Number
- CN202210110134.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-01-28
AI Technical Summary
During the operation and maintenance of data center computer room infrastructure, it is difficult for the existing technology to quickly and accurately locate the root causes of alarms, resulting in low operation and maintenance efficiency and insufficient security.
By determining the alarm rules for upstream monitoring data, identify the alarm trigger status of the current monitoring data item relative to the historical monitoring data item, and traverse the alarm root cause in the device topology structure, use the device topology association relationship to quickly locate the root cause device, and output the alarm root cause.
It realizes the rapid and accurate positioning of the alarm cause without relying on expert experience and big data analysis, improves operation and maintenance efficiency and security, and avoids the problems of huge workload and poor flexibility in traditional methods.
Smart Images

Figure CN114443437B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data technology, and in particular to an alarm root cause output method, apparatus, device, medium, and program product. Background Art
[0002] Currently, data center system operations, such as those for computer room infrastructure, require the collection and analysis of current operational status of various infrastructure components using sensors and IoT technology. The number of monitored sensor points ranges from hundreds of thousands to millions, depending on the size of the data center. When a core node fails, downstream devices simultaneously trigger numerous alarms. For example, computer room infrastructure alarms are generated based on IoT data collection technology and, through the integration of tree-structured retrieval algorithms, an alarm aggregation scenario is constructed to inform operations personnel of the root cause of the alarm. Summary of the Invention
[0003] In view of at least one of the technical problems existing in the process of determining the root cause of alarms in the above-mentioned data center computer room infrastructure, the present disclosure provides an alarm root cause output method, device, equipment, medium and program product that improve the alarm root cause location processing speed.
[0004] According to a first aspect of the present disclosure, a method for outputting an alarm root cause is provided, comprising: determining an alarm rule for upstream monitoring data; identifying an alarm triggering status of a current monitoring data item in the upstream monitoring data relative to a historical monitoring data item based on the alarm rule; and traversing the root cause device of the alarm root cause in a device topology structure according to the alarm triggering status of the current monitoring data item, to output the alarm root cause of the root cause device.
[0005] According to an embodiment of the present disclosure, before determining the alarm rule of the upstream monitoring data, the method further includes: determining device standard information corresponding to the upstream monitoring data; and configuring a device topology structure according to the device standard information.
[0006] According to an embodiment of the present disclosure, in determining the equipment standard information corresponding to the upstream monitoring data, it includes: standardizing the equipment type information corresponding to the upstream monitoring data; based on the standardized equipment type information, standardizing the monitoring point information corresponding to the equipment type information to generate equipment standard information; wherein the equipment type information includes the equipment name and the corresponding equipment number; the monitoring point information includes the monitoring point name corresponding to the equipment name and the corresponding monitoring point number.
[0007] According to an embodiment of the present disclosure, configuring a device topology structure according to device standard information includes: configuring a device topology association corresponding to upstream monitoring data according to the device standard information through device topology editing to complete the configuration of the device topology structure.
[0008] According to an embodiment of the present disclosure, in determining the alarm rules of upstream monitoring data, it includes: creating rule items of new alarm rules based on the equipment standard information corresponding to the upstream monitoring data; generating alarm conditions corresponding to the rule items to determine the alarm rules; wherein the alarm conditions include alarm level, alarm range, alarm equipment type, alarm monitoring point, alarm trigger condition, alarm recovery condition and alarm topology relationship.
[0009] According to an embodiment of the present disclosure, before identifying the alarm trigger status of the current monitoring data item in the upstream monitoring data relative to the historical monitoring data item based on the alarm rule, it also includes: cleaning and standardizing the received upstream monitoring data to generate upstream standardized data; identifying the data change status of the current standardized data item in the upstream standardized data relative to the historical standardized data item; and determining the stored monitoring data in the upstream standardized data based on the data change status of the current standardized data item.
[0010] According to an embodiment of the present disclosure, in identifying the alarm trigger status of the current monitoring data item in the upstream monitoring data relative to the historical monitoring data item based on the alarm rule, it includes: querying the alarm rule that matches the current monitoring data item based on the equipment standard information corresponding to the current monitoring data item in the incoming monitoring data; and performing a trigger check on the current monitoring data item in response to the queried alarm rule to determine the alarm trigger status of the current monitoring data item relative to the historical monitoring data item.
[0011] According to an embodiment of the present disclosure, in response to the queried alarm rules, a trigger check is performed on the current monitoring data item to determine the alarm trigger status of the current monitoring data item relative to the historical monitoring data item, including: determining through a trigger check that the alarm trigger status of the equipment monitoring point corresponding to the historical monitoring data item is not triggered; in response to the not triggered alarm trigger status, when the monitoring information of the equipment monitoring point corresponding to the current monitoring data item meets the alarm trigger condition of the alarm rule, the alarm trigger status of the current monitoring data is added.
[0012] According to an embodiment of the present disclosure, in response to the queried alarm rules, a trigger check is performed on the current monitoring data item to determine the alarm trigger status of the current monitoring data item relative to the historical monitoring data item, and it also includes: determining, through the trigger check, that the alarm trigger status of the equipment monitoring point corresponding to the historical monitoring data item is triggered; in response to the triggered alarm trigger status, when the monitoring information of the equipment monitoring point corresponding to the current monitoring data item meets the alarm trigger condition of the alarm rule, the alarm trigger status of the current monitoring data is updated.
[0013] According to an embodiment of the present disclosure, in response to the queried alarm rules, a trigger check is performed on the current monitoring data item to determine the alarm trigger status of the current monitoring data item relative to the historical monitoring data item, and also includes: based on the alarm trigger status, when there is no alarm change corresponding to the current monitoring data item, updating the alarm information corresponding to the current monitoring data item; or based on the alarm trigger status, when there is an alarm change corresponding to the current monitoring data item, adding the alarm information corresponding to the current monitoring data item.
[0014] According to an embodiment of the present disclosure, in traversing the root cause device of the alarm root cause in the device topology structure according to the alarm trigger status of the current monitoring data item, and outputting the alarm root cause of the root cause device, it includes: determining the position status information of the current monitoring device corresponding to the current monitoring data item in the device topology structure according to the alarm trigger status of the current monitoring data item; traversing its upstream devices in the device topology structure based on the position status information of the current monitoring device, and outputting the root cause device.
[0015] According to an embodiment of the present disclosure, in a device topology structure, based on the location status information of the current monitoring device, a traversal is performed on its upstream devices, and the root cause device is output, including: determining the last final triggering device in an alarm triggering state during the execution of the upstream device traversal; based on the alarm aggregation information of the final triggering device, outputting the final triggering device as the root cause device.
[0016] According to an embodiment of the present disclosure, in traversing the root cause device of the alarm root cause in the device topology structure according to the alarm trigger status of the current monitoring data item, outputting the alarm root cause of the root cause device includes: updating the aggregated alarm information of the root cause device; and outputting the updated aggregated alarm information as the alarm root cause.
[0017] A second aspect of the present disclosure provides an alarm root cause output device, comprising a rule determination module, a state identification module, and a device traversal module. The rule determination module is configured to determine an alarm rule for upstream monitoring data; the state identification module is configured to identify, based on the alarm rule, the alarm trigger state of a current monitoring data item in the upstream monitoring data relative to a historical monitoring data item; and the device traversal module is configured to traverse a device in a device topology structure for a root cause of the alarm root cause based on the alarm trigger state of the current monitoring data item, and output the alarm root cause of the root cause device.
[0018] The third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned alarm root cause output method.
[0019] A fourth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned alarm root cause output method.
[0020] A fifth aspect of the present disclosure further provides a computer program product, including a computer program, which implements the above-mentioned alarm root cause output method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0022] Figure 1 Schematically illustrates an application scenario diagram of the alarm root cause output method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0023] Figure 2 Schematically shows a flow chart of a method for outputting an alarm root cause according to an embodiment of the present disclosure;
[0024] Figure 3 A diagram schematically illustrates a composition diagram of a device topology structure according to an embodiment of the present disclosure;
[0025] Figure 4 The following schematically shows a composition diagram of an alarm rule according to an embodiment of the present disclosure;
[0026] Figure 5 Schematically shows a data processing flow chart of upstream monitoring data-incoming monitoring data according to an embodiment of the present disclosure;
[0027] Figure 6 The following schematically illustrates a flow chart of a process for identifying an alarm triggering state of a current monitoring data item according to an embodiment of the present disclosure;
[0028] Figure 7 Schematically shows an output flow chart of a root cause device according to an embodiment of the present disclosure;
[0029] Figure 8 Schematically shows a structural block diagram of an alarm root cause output device according to an embodiment of the present disclosure; and
[0030] Figure 9 A block diagram of an electronic device suitable for implementing the alarm root cause output method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0031] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0032] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0033] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0034] When expressions such as "at least one of A, B and C, etc." are used, they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0035] In order to realize the output of alarm root cause location, there are mainly two traditional methods in the existing technology.
[0036] First, extract the characteristic values or key information of each alarm. Using pre-set rules, group and categorize key information, and aggregate alarms of the same category into alerts. This approach primarily utilizes technologies such as alarm information standardization and regular expression matching analysis. However, this approach primarily analyzes and categorizes alarm information, relying on the pre-established organization of business rules and the ongoing operational refinement of alarm grouping information and characteristic values. This approach requires the accumulation of expert experience and cannot automatically or semi-automatically resolve the alarm aggregation problem through technical means.
[0037] Second, alarm aggregation is accomplished through big data analysis and conditional judgment. After an upstream system generates an alarm, the system receives it and identifies it through big data analysis. This analysis compares it with historical data, determines the pre-set conditions, and determines the root cause of the alarm through big data analysis and scenario elimination. However, this technical solution relies on the analytical and computational capabilities of big data. If the amount of data analyzed is large, the processing time is long, and it cannot meet the requirements for rapid identification and location of alarm information, resulting in a certain delay. Furthermore, the judgment conditions must be exhaustively set, and parameters must be continuously modified based on the analysis results, preventing rapid operational use.
[0038] In view of at least one of the technical problems existing in the process of determining the root cause of alarms in the above-mentioned data center computer room infrastructure, the present disclosure provides an alarm root cause output method, device, equipment, medium and program product that improve the alarm root cause location processing speed.
[0039] It should be noted that the above-mentioned alarm root cause output method and device disclosed in the present invention can be used in the fields of big data technology and artificial intelligence technology, and can also be used in the financial field and any field outside the financial field. The application field of the alarm root cause output method and device disclosed in the present invention is not limited.
[0040] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information and other data are in compliance with relevant laws and regulations, necessary confidentiality measures are taken, and they do not violate public order and good morals. Furthermore, the user's authorization or consent is obtained before obtaining or collecting user personal information.
[0041] An embodiment of the present disclosure provides a method for outputting an alarm root cause, comprising: determining an alarm rule for upstream monitoring data; identifying an alarm triggering status of a current monitoring data item in the upstream monitoring data relative to a historical monitoring data item based on the alarm rule; and traversing the root cause device of the alarm root cause in a device topology structure according to the alarm triggering status of the current monitoring data item, to output the alarm root cause of the root cause device.
[0042] Figure 1 The application scenario diagram of the alarm root cause output method, apparatus, device, medium and program product according to an embodiment of the present disclosure is schematically shown.
[0043] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0044] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0045] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0046] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0047] It should be noted that the alarm root cause output method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the alarm root cause output device provided in the embodiments of the present disclosure can generally be set in the server 105. The alarm root cause output method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the alarm root cause output device provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0048] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0049] The following will be based on Figure 1 The scene described by Figures 2 to 6 The alarm root cause output method of the disclosed embodiment is described in detail.
[0050] Figure 2 The flowchart of the alarm root cause output method according to an embodiment of the present disclosure is schematically shown.
[0051] like Figure 2 As shown, the alarm root cause output method of this embodiment includes operations S201 to S203.
[0052] In operation S201, an alarm rule for upstream monitoring data is determined;
[0053] In operation S202 , an alarm triggering state of a current monitoring data item in upstream monitoring data relative to a historical monitoring data item is identified based on an alarm rule;
[0054] In operation S203, root cause devices of the alarm root cause are traversed in the device topology structure according to the alarm triggering state of the current monitoring data item, so as to output the alarm root cause of the root cause device.
[0055] The upstream monitoring data is the monitoring data of each device in the operation and maintenance system in the embodiment of the present disclosure. Specifically, for each device in the operation and maintenance system, it has a corresponding supporting monitoring module or monitoring device, which monitors the operation of the device regularly, periodically or in real time, and forms monitoring data. Among them, the monitoring data set formed by all the devices in the operation and maintenance system can be used as the above-mentioned upstream monitoring data. The alarm rule is an alarm setting condition that matches each monitoring data item in the upstream monitoring data. When the alarm setting condition is met, if the monitoring data corresponding to the monitoring data item is abnormal, a monitoring alarm for the monitoring data item can be realized. Therefore, through the alarm rule of the upstream monitoring data, alarm matching can be performed on each data item in the upstream monitoring data.
[0056] For each monitoring data item in the upstream monitoring data, since timed, periodic, and even real-time status monitoring is required, alarm trigger monitoring needs to be implemented at different times for this monitoring data item. The current monitoring data item is the current data item in the upstream monitoring data that is currently undergoing alarm trigger monitoring. The historical monitoring data item is a historical data item that is consistent with the data item corresponding to the current monitoring data item, and the historical data item is a historical data item in the upstream monitoring data that has completed alarm trigger monitoring. In other words, the data item name and the device to which the current monitoring data item and the historical monitoring data item correspond are consistent, but the corresponding data item contents are different. The data item content of the current monitoring data item is the data item content at the current monitoring trigger moment, while the data item content of the historical monitoring data item is the data item content at the historical moment when the monitoring trigger has been completed.
[0057] The alarm trigger state of the current monitoring data item is identified according to the alarm rule. The current data content of the current monitoring data item can be compared with the alarm trigger data content of the data item matched in the alarm rule to identify whether the state of the current monitoring data item is an alarm trigger. For example, when the current monitoring data item is a voltage data item of a battery pack device, and the monitored value of the current voltage data item is 220V (i.e., data content), the alarm trigger data content of the voltage data item matched in the corresponding alarm rule is 220V. This indicates that the two data contents are consistent, and the alarm is identified as no abnormality. The alarm trigger state of the current voltage data item is marked as no trigger, i.e., no alarm processing is performed. Correspondingly, when the two data contents are inconsistent, it can be identified as an alarm abnormality, and the alarm trigger state of the current voltage data item is marked as triggered, i.e., an alarm processing is performed. Therefore, the alarm trigger state is the data abnormality feedback state of each monitoring data item in the corresponding upstream monitoring data matched according to the alarm rule. When the data is abnormal, the data item is marked as an alarm trigger, otherwise the data item is marked as no trigger.
[0058] For example, the current monitoring data item at 11:00 am on December 11, 2021 is the voltage data item of the battery pack equipment. The monitoring value of the current voltage data item is 220V, there is no abnormality, and the alarm trigger status is not triggered; correspondingly, the historical monitoring data item at 10:00 am on December 11, 2021 is also the voltage data item of the battery pack equipment. The monitoring value of the historical voltage data item is 210V, which is abnormal, and the alarm trigger status is triggered.
[0059] Among them, the alarm trigger state of the current monitoring data item can be judged according to the alarm trigger state of the historical monitoring data item. For example, according to the alarm rule, when the alarm trigger state of the historical monitoring data item is triggered, if the monitoring data content of the current monitoring data item is the same as it, then the alarm trigger state of the current monitoring data item can also be identified as triggered. On the contrary, when the data of the two are not the same, the monitoring data content of the current monitoring data item can be further identified by the alarm rule. In this way, the connection between the historical monitoring data item and the current historical monitoring data can be established, and the current historical monitoring data can be judged by means of the alarm state of the historical monitoring data item, thereby helping to realize parallel operation of data and speeding up data processing.
[0060] The device topology structure is the topology structure of all upstream devices in operation in the operation and maintenance system disclosed herein. Through this device topology structure, a topological association relationship can be established for all upstream devices. Through this topological association relationship, the alarm trigger of the current monitoring data item can be traversed to query the root cause device of the alarm root cause. For each current monitoring data item in the upstream monitoring data, once its corresponding alarm trigger state is triggered, i.e., the content of the current data item is abnormal, then monitoring anomalies may also exist on other upstream devices with which it has a topological association relationship. That is, an abnormality exists in the upstream device of the current device corresponding to the current monitoring data item, which may cause the corresponding current monitoring data item of the current device to have an abnormal state of alarm triggering. By traversing the upstream devices with a topological association relationship with the current device, the device where the fundamental anomaly that causes the abnormality of the current monitoring data to occur can be found, i.e., the root cause device. The abnormal data item and its content that appear in the root cause device and cause the abnormal state of the current monitoring data of the current device to be alarm triggered can serve as the above-mentioned root cause of the alarm.
[0061] Therefore, compared with the traditional method of aggregating alarms by configuring the relationship of alarm rules one by one and aggregating alarms by using big data analysis in the prior art, the above-mentioned alarm root cause output method of the embodiment of the present invention utilizes the topological association relationship between each device in the operation and maintenance system to establish upstream and downstream device connections, standardizes the equipment and alarm information that may cause an alarm storm, and realizes rapid positioning of the root cause device and accurate output of the alarm root cause without relying on expert technical experience and big data analysis, thereby avoiding the huge workload and poor flexibility caused by the traditional method of managing and configuring alarm rules one by one, as well as the traditional disadvantages of being unable to quickly update associations, greatly improving the output efficiency of the alarm root cause, ensuring the accuracy and output speed of the alarm root cause, and improving the operation and maintenance safety and efficiency.
[0062] like Figure 2 As shown, according to an embodiment of the present disclosure, before determining the alarm rule of the upstream monitoring data in operation S201, the method further includes:
[0063] Determine the equipment standard information corresponding to the upstream monitoring data;
[0064] Configure the device topology according to the device standard information.
[0065] Standardized device information is the standardized basic information of all devices in the operation and maintenance system corresponding to upstream monitoring data. This information primarily includes the operation and maintenance monitoring equipment in data center infrastructure, such as computer rooms. This basic information includes the device type, basic device information, device space information, and monitoring point information. Furthermore, standardization can include data initialization management to facilitate the configuration of the basic data for upstream monitoring data.
[0066] After initializing the upstream monitoring data, the basic information of the corresponding devices can be standardized to form device standard information, such as the unification of standard fields for each data item, so as to facilitate subsequent processing of each data item according to the unified standard fields.
[0067] Furthermore, providing this device with standard information can establish topological relationships between various devices, thereby forming a corresponding device topology structure based on this topological relationship. This facilitates the subsequent root cause tracing of the alarms of various abnormal data items, thereby enabling the search and location of upstream and downstream topological relationships.
[0068] like Figure 2 As shown, according to an embodiment of the present disclosure, determining the device standard information corresponding to the upstream monitoring data includes:
[0069] Standardize the device type information corresponding to upstream monitoring data;
[0070] Based on the standardized equipment type information, the monitoring point information corresponding to the equipment type information is standardized to generate equipment standard information;
[0071] The device type information includes the device name and the corresponding device number; the monitoring point information includes the monitoring point name corresponding to the device name and the corresponding monitoring point number.
[0072] Initialize the basic equipment information corresponding to the upstream monitoring data, including the standardization of equipment type information, monitoring point information and equipment naming, so as to complete the standardization of upstream and downstream equipment information.
[0073] First, in the process of standardizing the equipment type information and the monitoring point information of the equipment, considering that the infrastructure equipment of the computer room center has multiple equipment types, such as battery packs, high-voltage cabinets, low-voltage cabinets, precision air conditioners, UPS, etc., the equipment type information can be used as the type or category information of the operation and maintenance equipment, including the equipment name. Among them, each type of equipment can have multiple monitoring points, and each monitoring point can realize feedback on different monitoring data items. For example, the monitoring point of the equipment is a battery pack, including monitoring data items such as voltage, internal resistance, temperature, and discharge status, wherein the monitoring values and monitoring contents of these specific data items can be used as the data content of the data items. Therefore, the monitoring point information can be the information of the monitoring data items of each monitoring point of the corresponding equipment. Therefore, after completing the standardization of the monitoring point information, the monitoring alarm rules can be set for the monitoring points under the corresponding equipment type in the subsequent alarm rule configuration process. For example, the monitoring value of the monitoring voltage data item of the equipment battery pack can be set in the alarm rule to issue an alarm when it is less than 0.
[0074] Based on device type, the equipment in the computer room center can be divided into multiple categories, such as transformers, high-voltage cabinets, low-voltage cabinets, UPS devices, battery packs, precision air conditioners, temperature and humidity sensors, and gas detection equipment. To standardize the type information for each of these devices, each category can be named and identified. For example, devices can be uniformly named according to a four-digit alphabetical and numerical numbering scheme, such as BA01 for battery packs. This can be adjusted based on actual circumstances. This allows device type information to also include the device number corresponding to the device name, enabling standardized type information for each device.
[0075] After completing the standardization of the above-mentioned equipment type information, the monitoring points corresponding to the equipment type number can be divided into monitoring point monitoring parameters such as monitoring voltage, monitoring internal resistance, monitoring temperature rise, and monitoring discharge status of battery pack equipment. Therefore, the monitoring point information includes the monitoring point name corresponding to the equipment name. Furthermore, the monitoring location corresponding to the equipment can also be uniformly named according to the numbering rules of 4-digit letters or numbers. For example, the monitoring point of the battery pack equipment is battery voltage, and the number of the battery voltage is 1001. The specific adjustment can be made according to the actual situation. In this way, the monitoring point information corresponding to each equipment type can be standardized by defining the monitoring point number for the monitoring point name corresponding to the equipment name.
[0076] Furthermore, with the help of the standardization of the above-mentioned equipment type information and monitoring point information, unified naming and numbering of equipment types and monitoring points is achieved. In this way, unified naming of the equipment itself can be achieved. Specifically, the unified naming of the equipment can also reflect the professional information and spatial information of the equipment. For example, the equipment can be uniformly named according to the numbering naming rules of building (1 digit) - floor (3 digits) - room (5 digits) - professional (2 digits) - equipment type (4 digits) - equipment number code (8 digits). For example, the second battery of the 5th group of the power storage battery group in the battery room No. 1 on the negative second floor of Building A can be named A-B02-BAR01-01-BA01-00050002. Correspondingly, unified naming of the corresponding monitoring points can also be achieved.
[0077] Therefore, based on the standardization of equipment type information and equipment monitoring point information, equipment standard information can be formed to ensure the unique readability of the equipment.
[0078] Figure 3 The diagram schematically shows a composition diagram of a device topology structure according to an embodiment of the present disclosure.
[0079] like Figure 2 and Figure 3 As shown, according to an embodiment of the present disclosure, configuring a device topology structure according to device standard information includes:
[0080] By editing the device topology, you can configure the device topology association corresponding to the upstream monitoring data according to the device standard information to complete the configuration of the device topology structure.
[0081] By using a device topology editing program or tool to uniformly edit the unified device naming and numbering within the standard device information for each device in the computer room data center, a device topology structure with device topology associations can be constructed. Specifically, in the device topology editing process, visual interaction can be used to perform topological editing of the unified device naming and numbering for maintained devices, completing the association between upper-level and lower-level devices, forming a device upper- and lower-level topology association. Once configured, a tree-like topology structure of the computer room infrastructure devices is formed. Multiple topologies can be defined based on specific disciplines and device associations, allowing a single device to be in multiple topologies.
[0082] like Figure 3As shown, in the device topology 300, the high-voltage distribution cabinet 310 can be used as an upstream device to be associated with the downstream transformers 320 and 330; wherein, the transformer 320 can be used as an upstream device to be associated with the downstream low-voltage distribution cabinets 340, 350, and 360; wherein, the low-voltage distribution cabinet 340 can be used as an upstream device to be associated with the downstream UPS devices 370, 380, and 390. In this way, a device topology 300 with a topological association relationship can be formed. Therefore, when the data item of the corresponding monitoring point of any device is abnormal, the device corresponding to the monitoring point and all upstream and downstream devices that may cause the abnormal data item of the device can be searched and located to determine its impact range and the abnormal device at the most upstream.
[0083] With the help of the above-mentioned device topology structure, the root cause of the alarm of each abnormal data item can be traced in the later stage, thereby realizing the search and positioning of the upstream and downstream topological relationships, as well as the configuration of alarm rules for the corresponding device monitoring data items.
[0084] Figure 4 The figure schematically shows the composition of the alarm rule according to an embodiment of the present disclosure.
[0085] like Figure 2-Figure 4 As shown, according to an embodiment of the present disclosure, in operation S201, determining the alarm rule of upstream monitoring data includes:
[0086] Create new alarm rule items based on the device standard information corresponding to the upstream monitoring data;
[0087] Generate alarm conditions corresponding to rule items to determine alarm rules;
[0088] Among them, the alarm conditions include alarm level, alarm scope, alarm device type, alarm monitoring point, alarm triggering condition, alarm recovery condition and alarm topology relationship.
[0089] After completing the basic configuration of the above-mentioned equipment standard information and equipment topology structure, alarm rules can be set for the data items of the monitoring points corresponding to the equipment standard information. The alarm rules can perform batch input matching corresponding to different data items, and use the matched alarm rules to make alarm judgments on abnormal data items of the above-mentioned upstream monitoring data received.
[0090] Each data item's corresponding alarm rule has a corresponding unified name, which can be defined based on the alarm condition and the corresponding monitoring point, device type, and topological relationship between the topological devices. Specifically, a new rule item can be created for a data item of upstream monitoring data based on the device standard information corresponding to the data item. This rule item can uniquely define the alarm condition for the data item, forming a data item alarm rule.
[0091] like Figure 4 As shown, the alarm conditions include alarm device type 410 (which can be reflected by the device name, etc.), alarm level 420, alarm effective area 430 (i.e., alarm range), alarm association topology 440 (i.e., alarm topology relationship), alarm recovery condition 450, alarm trigger condition 460, and alarm monitoring point (i.e., monitoring indicator 470). Therefore, a systematic definition of alarm rules can be achieved.
[0092] Alarm device type 410 setting: configure the device type defined by the data item corresponding to the upstream monitoring data, such as a battery pack. Alarm level 420 setting: define the level of abnormal alarm for abnormal data items as early warning, general alarm, severe alarm, etc., which can be set specifically according to operation and maintenance needs. Effective area 430 setting: refers to the spatial scope of the monitoring equipment of the alarm rule of this data item. You can select the entire computer room, the entire building or the entire floor according to the situation. The default is the entire computer room. If a battery pack is selected above, it refers to all battery packs in the entire computer room. If a building is selected, it refers to the battery pack in the building. Similarly, the monitoring point settings are deduced to define the monitoring points under the equipment type involved in the alarm, such as the voltage under the battery pack. Associated topology 440 setting: Select as above Figure 3 The device topology structure 300 of the defined device upstream and downstream relationship, if the above-mentioned trigger condition 460 is triggered, will complete the aggregation calculation trigger association according to the device topology structure. The recovery condition 450 is set: when the above-mentioned trigger condition is reached, the alarm exits, and the definition content can be consistent with the trigger condition, such as the voltage of the battery pack is greater than 0, indicating that the fault is recovered. Among them, the trigger condition 460 is set: if the threshold of the monitoring point position of the data item corresponding to the corresponding monitoring indicator is set to a threshold reference relationship such as greater than, less than, greater than or equal to, and less than or equal to, when the value of the data content of the corresponding data line in the upstream monitoring data or the defined corresponding value is compared with the threshold size, it can be defined whether the data content is abnormal, whether the data item is an abnormal data item, whether the corresponding device is an abnormal device, and whether other upstream and downstream devices with topological associations are also abnormal devices.
[0093] It can be seen that through the above-mentioned alarm rule settings corresponding to each data item, the data items can be directly matched with batch alarm rules, and the judgment criteria for abnormality and normality of each data item are established, which speeds up the matching speed of subsequent alarm rules, improves the efficiency of alarm aggregation processing, and ensures the judgment accuracy of data item alarms.
[0094] Figure 5 The data processing flow chart of upstream monitoring data-incoming monitoring data according to an embodiment of the present disclosure is schematically shown.
[0095] like Figure 2-Figure 5As shown, according to an embodiment of the present disclosure, before operation S202 identifies the alarm triggering status of the current monitoring data item in the upstream monitoring data relative to the historical monitoring data item based on the alarm rule, the method further includes:
[0096] Clean and standardize the received upstream monitoring data to generate upstream standardized data;
[0097] Identify the data change status of the current standardized data item relative to the historical standardized data item in the upstream standardized data;
[0098] Determine the incoming monitoring data in the upstream standardized data based on the data change status of the current standardized data item.
[0099] like Figure 4 As shown, the upstream monitoring data provided by the upstream monitoring system can be received through the Internet of Things technology. These upstream monitoring data may include monitoring data streams of monitoring equipment such as power monitoring systems (such as high-voltage cabinets, low-voltage cabinets UPS, battery packs, etc.), operation monitoring systems, and dynamic environment monitoring systems (such as precision air conditioners, etc.), such as operation S501.
[0100] Specifically, the upstream monitoring data received above includes not only each monitoring data item, but also the data content corresponding to each data item, such as the value of a specific monitoring point, such as the battery numbered A-B02-BAR01-01-BA01-00050002, whose monitoring point number is 1001 and the voltage reading is 320V.
[0101] After receiving the real-time upstream monitoring data collected and transmitted by the Internet of Things, it is possible to monitor whether the equipment involved in the upstream monitoring data belongs to the equipment corresponding to the above-mentioned standardized equipment standard information, that is, whether the equipment corresponding to the upstream monitoring data has completed the above-mentioned standardized equipment information maintenance operation, such as operation S502. Among them, if the query confirms that the equipment corresponding to the upstream monitoring data has completed the standardization operation and generated the corresponding equipment standard information, then proceed to the next step, otherwise the data is directly discarded, such as operation S521. For example, the upstream monitoring data of a device uniformly named C-B01-BAR11-01-BA01-00000001 is received, but the corresponding equipment basic information of the device has not undergone the above-mentioned information standardization processing, and there is no corresponding equipment standard maintenance information, then it is discarded. Therefore, the preliminary screening of the upstream monitoring data to be processed can be completed, the amount of useless data can be preliminarily reduced, redundant data can be screened out, and the data processing process can be accelerated.
[0102] On the basis of the preliminary screening results of the above-mentioned equipment, the received upstream monitoring data is further cleaned and standardized. The corresponding monitoring data items in the upstream monitoring data generally include the corresponding data item generation time, the corresponding equipment unique naming identification information, the equipment monitoring point information and the corresponding monitoring value. Therefore, after determining that the corresponding equipment is a maintained equipment, it is necessary to supplement the remaining attribute information of the equipment according to the equipment standardization information, such as supplementing the equipment type, equipment space information and equipment monitoring point conversion data that the upstream monitoring data does not have, and complete the cleaning and standardization of the upstream monitoring data, thereby generating the corresponding upstream standardized data, which is the corresponding upstream monitoring data that has undergone data supplementation and standardization processing, such as operation S503. In this way, the precise processing of the upstream monitoring data can be further achieved, so that the upstream standardized data generated by the upstream monitoring data can be more perfectly matched with the subsequent alarm processing process, and is conducive to achieving accurate alarm aggregation.
[0103] After completing the above data cleaning and standardization, enter the data warehousing stage and store the above upstream standardized data to facilitate real-time calls, such as transmitting these data calls to alarm rule matching and accessing alarm logic processing, such as operation S505.
[0104] After data cleaning, the warehousing logic of the corresponding database first queries the current values of the corresponding equipment and monitoring points of each current standardized data item in the upstream standardized data, and combines the historical values of the corresponding equipment and monitoring points of the historical standardized data items relative to the current standardized data item. When the current value has not changed relative to the historical value, the data change status of the current standardized data item is determined to be unchanged, and the monitoring data generation time of the historical standardized data item is directly updated to the data generation time of the corresponding equipment and monitoring points of the current standardized data item, and the process of data warehousing to form the stored monitoring data is completed, such as operation S504.
[0105] On the contrary, when the current value changes relative to the historical value, the data change status of the current standardized data item is determined to be changed, and a new data item record is added to the data table in the database, and the current standardized data item and its corresponding equipment and monitoring point data, as well as the data generation time are filled into the newly added data item record at the same time, thereby completing the recording of the changed current standardized data item into the warehouse and forming the stored monitoring data.
[0106] Therefore, by cleaning and standardizing the above-mentioned upstream monitoring data, and performing warehousing operations on the corresponding data, and filtering out the data content of data items that have not changed during the warehousing operation, updating the generation time of the data content, or adding the data content and data content generation time of data items that have changed, the amount of upstream monitoring data can be greatly reduced, and the amount of data for subsequent processing is greatly reduced. At the same time, the records of changed data can be kept, and the log time of unchanged data is updated, ensuring the accuracy of data processing, maintaining data integrity, preventing data omission and loss, and facilitating data input in the subsequent alarm aggregation process.
[0107] Figure 6 The following schematically shows a flow chart of the process of identifying the alarm triggering status of the current monitoring data item according to an embodiment of the present disclosure.
[0108] like Figure 2-Figure 6 As shown, according to an embodiment of the present disclosure, in operation S202, identifying the alarm triggering status of the current monitoring data item in the upstream monitoring data relative to the historical monitoring data item based on the alarm rule includes:
[0109] According to the equipment standard information corresponding to the current monitoring data item in the incoming monitoring data, query the alarm rules that match the current monitoring data item;
[0110] In response to the queried alarm rule, a trigger check is performed on the current monitoring data item to determine the alarm trigger status of the current monitoring data item relative to the historical monitoring data item.
[0111] For the incoming monitoring data that has been cleaned and put into storage, each of its data items needs to be matched with the corresponding alarm rules. For the current monitoring data item in the incoming monitoring data, it is necessary to query and match the alarm rules based on the unified naming information of the equipment, the equipment type information, and the corresponding monitoring point information in the corresponding equipment standard information. Query the rule items that can be matched with it in the alarm rules. Each rule item can define the alarm conditions of the current monitoring data item, which is conducive to completing the following alarm trigger check. Among them, the query of the alarm rules is based on the preset alarm rule matching conditions, and the rule items are matched with the corresponding data items through the matching relationship between the equipment type information and monitoring point information and the matching rule items.
[0112] After completing the query matching of the alarm rule, a trigger check can be performed based on the alarm rule that matched the query to determine whether the data content of the current monitoring data item meets the trigger condition of the alarm rule. If the trigger condition is met, the trigger check is successful; otherwise, the trigger check fails. The alarm trigger status of the current monitoring data item for which the trigger check succeeds can be "alarm triggered," while the alarm trigger status of the current monitoring data item for which the trigger check fails can be "alarm not triggered."
[0113] Therefore, the alarm triggering status of the current monitoring data item can be judged thereby, and the processing process of alarm aggregation can be further facilitated thereby, ensuring the processing accuracy of alarm aggregation.
[0114] Furthermore, with the help of trigger checks, alarm judgment can be performed on the incoming monitoring data after the above data cleaning, streaming data can be calculated and judged, and alarm rules can be matched in real time, and an alarm can be triggered after the alarm conditions are met.
[0115] like Figure 6 As shown, after receiving the equipment inbound monitoring data, the alarm rules are queried through the equipment standard information corresponding to the current monitoring data item in the inbound monitoring data to check whether the equipment and monitoring points corresponding to the current monitoring data item are configured with the rule item information of the alarm rule. If so, the alarm trigger check is entered; if not, the process ends and exits, such as operations S601-S602. Furthermore, if the equipment monitoring data is configured with an alarm rule, a trigger check is performed to check the trigger status corresponding to the alarm rule. If the alarm rule has been triggered, that is, an alarm has been generated under the alarm rule, the existing alarm processing link is entered. If it has not been triggered before, the new alarm processing link is entered, such as operations S603-S607.
[0116] like Figure 2-Figure 6 As shown, according to an embodiment of the present disclosure, in response to the queried alarm rule, a trigger check is performed on the current monitoring data item to determine the alarm trigger status of the current monitoring data item relative to the historical monitoring data item, including:
[0117] Through trigger check, it is determined that the alarm trigger status of the equipment monitoring point corresponding to the historical monitoring data item is not triggered;
[0118] In response to the untriggered alarm trigger state, when the monitoring information of the equipment monitoring point corresponding to the current monitoring data item meets the alarm trigger condition of the alarm rule, the alarm trigger state of the current monitoring data is added.
[0119] If the equipment entry monitoring data is configured with an alarm rule, a trigger check is performed to check the trigger status of the current monitoring data item corresponding to the alarm rule. If the alarm rule is not triggered by the historical monitoring data item, that is, no alarm status is generated under the alarm rule, then enter the new alarm processing procedure, such as operations S603-S606.
[0120] A trigger check is performed on the historical monitoring data item corresponding to the current monitoring data item. If the monitoring data content corresponding to the historical monitoring data item does not meet the alarm trigger conditions of the matching alarm rules, the alarm trigger status of the device monitoring point corresponding to the monitoring data content of the historical monitoring data item may be not triggered.
[0121] A trigger check is performed on the monitoring data content of the current monitoring data item. If the monitoring data content corresponding to the current monitoring data item meets the alarm trigger condition of the matching alarm rule, it means that the alarm trigger status of the device monitoring point of the monitoring data content corresponding to the current monitoring data item is triggered.
[0122] Therefore, the alarm trigger status of the current monitoring data item has changed compared with that of the historical monitoring data item. Therefore, the current monitoring data item and its corresponding data content are newly processed in the corresponding data table in the database, and the current monitoring data item is added to the data table.
[0123] For untriggered alarm rules corresponding to historical monitoring data items in the incoming monitoring data, the trigger check process is initiated to determine whether the current monitoring data item meets the trigger conditions. If so, an alarm processing process is added under the alarm rule, the alarm status is continuously monitored, and the alarm information, first alarm time, etc. are recorded. The alarm rule status is then changed to triggered. If the trigger conditions are not met, the process ends, as in operations S606-S607.
[0124] Therefore, it is possible to add new items to the currently triggered monitoring data, thereby ensuring the integrity of the triggered monitoring data and avoiding the omission or loss of necessary abnormal data, thereby ensuring the accuracy of abnormal triggering alarms.
[0125] like Figure 2-Figure 6 As shown, according to an embodiment of the present disclosure, in response to the queried alarm rule, a trigger check is performed on the current monitoring data item to determine the alarm trigger status of the current monitoring data item relative to the historical monitoring data item, further comprising:
[0126] Through trigger checking, it is determined that the alarm trigger status of the equipment monitoring point corresponding to the historical monitoring data item is triggered;
[0127] In response to the triggered alarm trigger state, when the monitoring information of the equipment monitoring point corresponding to the current monitoring data item meets the alarm trigger condition of the alarm rule, the alarm trigger state of the current monitoring data is updated.
[0128] If the equipment entry monitoring data is configured with an alarm rule, a trigger check is performed to check the trigger status of the current monitoring data item corresponding to the alarm rule. If the alarm rule has been triggered by a historical monitoring data item, that is, an existing alarm status has been generated under the alarm rule, then the existing alarm processing procedure is entered, such as operations S603-S605.
[0129] A trigger check is performed on the historical monitoring data item corresponding to the current monitoring data item. If the monitoring data content corresponding to the historical monitoring data item meets the alarm trigger condition of the matching alarm rule, the alarm trigger status of the device monitoring point corresponding to the monitoring data content of the historical monitoring data item can be triggered.
[0130] A trigger check is performed on the monitoring data content of the current monitoring data item. If the monitoring data content corresponding to the current monitoring data item meets the alarm trigger condition of the matching alarm rule, it means that the alarm trigger status of the device monitoring point of the monitoring data content corresponding to the current monitoring data item is triggered.
[0131] Therefore, the alarm triggering status of the current monitoring data item may have changed compared to the alarm triggering status of the historical monitoring data item. Therefore, when the data content of the alarm triggering of the current monitoring data item is consistent with the data content of the alarm triggering of the historical monitoring data item, the current monitoring data item and its corresponding data content generation time are updated in the corresponding data table in the database, and the data generation time of the current monitoring data item is added to the data table.
[0132] Among them, if the historical monitoring data item in the equipment inventory monitoring data has triggered an alarm, it is determined whether the latest current monitoring data item still meets the triggering conditions. If the triggering conditions are met and an alarm is issued, the latest alarm time is recorded and the alarm time of the historical monitoring data item is replaced, such as operation S605.
[0133] Therefore, it is possible to filter the monitoring data items that have triggered the alarm rules, and update the alarm time and / or data generation time corresponding to the monitoring data items, thereby completing further filtering of the monitoring data items.
[0134] like Figure 2-Figure 6 As shown, according to an embodiment of the present disclosure, in response to the queried alarm rule, a trigger check is performed on the current monitoring data item to determine the alarm trigger status of the current monitoring data item relative to the historical monitoring data item, further comprising:
[0135] Based on the alarm trigger status, when there is no alarm change corresponding to the current monitoring data item, update the alarm information corresponding to the current monitoring data item; or
[0136] Based on the alarm trigger status, when there is an alarm change corresponding to the current monitoring data item, the alarm information corresponding to the current monitoring data item is added.
[0137] If the latest current monitoring data item does not meet the trigger conditions and there is no corresponding trigger alarm change, the data content corresponding to the current monitoring data item and the matching alarm information (such as alarm time, alarm rules, etc.) are updated to replace the data content and alarm information of the historical monitoring data item.
[0138] If the latest current monitoring data item meets the trigger conditions and there is a corresponding trigger alarm change, the data content corresponding to the current monitoring data item and the matching alarm information (such as alarm time, alarm rules, etc.) will be added and set in the data table alongside the data content and alarm information of the historical monitoring data items.
[0139] like Figure 6 As shown, in operation S608, when updating the alarm record, if there is an alarm processing for the historical monitoring data item, and if the alarm of the current monitoring data item has not changed, the latest alarm occurrence time and trigger count, etc. are updated; if the alarm of the current monitoring data item is restored, the alarm restoration time and alarm status, etc. are updated. On the contrary, if the alarm of the current monitoring data item has changed, a new alarm processing alarm record is added, such as alarm information (including the first alarm occurrence time, alarm level status, etc.). After the record is updated, the alarm aggregation program module is entered, such as operation S609.
[0140] If the latest current monitoring data item does not meet the trigger condition, it is determined whether it meets the recovery condition. If it does, the latest alarm status and alarm time are recorded and updated. The alarm rule daemon is also canceled, and the alarm rule status is changed to untriggered. If the latest current monitoring data item meets neither the trigger condition nor the recovery condition, the process ends.
[0141] Therefore, it is possible to achieve real-time parallel judgment of each data item in the data stream of the incoming monitoring data, and to filter and update the data volume of the data stream, so that the data processing volume is greatly reduced, resource consumption and data processing delay are reduced, data processing speed is improved, the integrity of alarm data is guaranteed, the omission and loss of abnormal data are avoided, and the accuracy of data alarm aggregation is improved.
[0142] Figure 7 The output flow chart of the root cause device according to an embodiment of the present disclosure is schematically shown.
[0143] like Figure 2-Figure 7 As shown, according to an embodiment of the present disclosure, in operation S203, traversing the root cause device of the alarm root cause in the device topology structure according to the alarm triggering state of the current monitoring data item, and outputting the alarm root cause of the root cause device includes:
[0144] Determine the position status information of the current monitoring device corresponding to the current monitoring data item in the device topology structure according to the alarm trigger status of the current monitoring data item;
[0145] In the device topology structure, based on the location status information of the current monitoring device, traverse its upstream devices and output the root cause device.
[0146] In order to meet the alarm aggregation processing requirements of the aggregate alarm scenario and complete the continuous monitoring and updating of the alarm status, the device corresponding to the current monitoring data item can be located according to the device topology association relationship defined by the device topology structure. Figure 3 As shown, if the current monitoring data item is in a triggered alarm state, the corresponding UPS device 370 can be located and traced back to the upstream low-voltage distribution cabinet 340 associated with the UPS device 370. Therefore, with the help of the device topology association relationship, the position of the device in the device topology structure can be defined. For example, the high-voltage distribution cabinet 310 is the original upstream device, and the defined position is (0, 0); the transformers 320 and 330, which are downstream devices of the high-voltage distribution cabinet 310, are in the first-level association relationship and can be located at (1, 0) and (1, 1) respectively; accordingly, the low-voltage distribution cabinets 340, 350, and 360, which are downstream devices of the transformer 320, are in the second-level association relationship and can be located at (2, 0), (2, 1), and (2, 2) respectively; the UPS devices 370, 380, and 390, which are downstream devices of the low-voltage distribution cabinet 340, are in the third-level association relationship and can be located at (3, 0), (3, 1), and (3, 2) respectively.
[0147] Therefore, the location of the current monitoring device in the device topology structure for the current monitoring data item, i.e., the location status information, can be understood as a set of associated coordinates defined by the device management level. The location status information of each device allows for the definition of the associated location of each device, thereby facilitating the location or traceability of each device. By inputting the corresponding alarm information for the current monitoring data item and performing an aggregated alarm check, it is possible to further determine the associated topology of the alarm rule, as in operations S701-S703.
[0148] When the current monitoring data item is judged to be in the triggered alarm state, the abnormal impact range in the entire device topology structure can be determined by tracing back to its upstream device. Figure 3As shown, when UPS device 370 is the current monitoring device corresponding to the current monitoring data item, and the data item is in the alarm triggering state, then according to the location status information associated with it in the device topology structure, the device traversal of the current monitoring device UPS device 370 is performed, as shown in operations S704-S706. If there is no corresponding device topology structure set, the operation and maintenance data processing is directly performed.
[0149] Therefore, when the alarm trigger status of the current monitoring data item is triggered, the position of the current monitoring device corresponding to the current monitoring data item in the device topology structure can be obtained, and the position can be marked as an alarm trigger, so that the corresponding root cause device can be traversed and traced with the help of the position of the device topology structure.
[0150] like Figure 2-Figure 7 As shown, according to an embodiment of the present disclosure, in the device topology structure, based on the location status information of the current monitoring device, traversing its upstream devices and outputting the root cause device include:
[0151] Determine the last final triggering device in the alarm triggering state during the upstream device traversal execution process;
[0152] Based on the alarm aggregation information of the final triggering device, the final triggering device is output as the root cause device.
[0153] In the embodiments of the present disclosure, alarm aggregation is actually to collect and converge the alarms of various upstream systems or upstream devices through a series of algorithms, and find the data processing means of the root cause node that triggers the alarm to reduce the alarm storm.
[0154] Specifically, traverse the upper-level device of the current monitoring device and determine whether the upper-level device is also in the triggered alarm trigger state. If so, continue to traverse to the upper-level device until traversing to a device in the untriggered alarm trigger state. Output the last device in the triggered alarm trigger state in the traversal process to the upper level as the final triggering device, and determine the alarm information of the device at the same time, that is, the final triggering device is the root cause device, such as operations S704-S706.
[0155] The alarm information of the final triggering device can be used to form the aforementioned aggregated alarm information. This aggregated alarm information can be a collection of alarm information records for all triggered devices determined through the traversal process in the device topology. Based on this aggregated alarm information, the output of the final triggering device can be realized, and the output is the root cause device.
[0156] Therefore, compared with the traditional method of managing and configuring alarm rules one by one, which causes a huge workload and poor flexibility, the topological association relationship defined by the device topology structure can greatly reduce the data processing workload, make data processing more flexible, and meet the needs of rapid updating of related information, which is more conducive to related application.
[0157] Therefore, with the help of matching alarm rules, alarm triggering check is realized, and the corresponding alarm device is identified based on the alarm triggering check results of each data item (i.e., alarm triggering status), and the traceability output of the root cause device is realized according to the topological association status of the alarm device. This can basically avoid the large-scale generation of various alarm storms, ensure the operation and maintenance stability during the alarm aggregation process, significantly improve the operation and maintenance safety, and improve the operation and maintenance efficiency.
[0158] like Figure 2-Figure 7 As shown, according to an embodiment of the present disclosure, in operation S203, traversing the root cause device of the alarm root cause in the device topology structure according to the alarm triggering state of the current monitoring data item, and outputting the alarm root cause of the root cause device further includes:
[0159] Update the aggregated alarm information of the root cause device;
[0160] The updated aggregate alarm information is output as the alarm root cause.
[0161] If it is determined that there is no root cause device corresponding to the current monitoring device in the existing alarm aggregation record, the dimension of the current root cause can be added and new alarm aggregation information can be added according to the topological association relationship of the device topology structure of the current detection device, and aggregated alarm information such as the aggregated topology structure information, the root cause device type, the matching alarm rule name, the alarm occurrence time, and the number of alarm aggregations can be recorded, as in operation S707. Finally, the alarm aggregation record is output and forwarded for processing as the above-mentioned alarm root cause, as in operation S708. Among them, when the alarm information of the final alarm root cause device is restored, the alarm aggregation status can be modified to restored.
[0162] Therefore, the above-mentioned method of the embodiment of the present disclosure can utilize the upstream and downstream correlation relationship of the computer room infrastructure to standardize the equipment and alarm information that may cause an alarm storm. Without relying on expert judgment and data analysis, through the device topology association setting, the alarm rules are matched to the corresponding data items. Combined with alarm aggregation, the root cause device of the alarm can be quickly found. It is suitable for operation and maintenance monitoring scenarios with relatively close upstream and downstream relationships, and can significantly improve operation and maintenance security and improve operation and maintenance efficiency.
[0163] like Figure 3 As shown, in order to further reflect the beneficial effects of the above technical solution, specific implementation cases are further provided as follows to provide those skilled in the art with a more comprehensive understanding of the above method of the embodiment of the present disclosure.
[0164] First, maintain the association relationship information between the equipment standard information and the equipment topology structure of the current monitoring equipment. Specifically, the high-voltage distribution cabinet 310 of Building A is used as the root node. Its downstream includes transformers 320 and 330, and the transformer 320 includes low-voltage distribution cabinets 340, 350, and 360. Further down are UPS devices 370, 380, and 390. The equipment naming and topology settings are standardized according to the above information standardization scheme.
[0165] After that, configure the alarm rule information of each device, and when the switch status of the current monitoring device is 0 (such as setting the alarm rule to "0-abnormal, 1-normal"), an alarm will be triggered, and the alarm rule will be matched to the following Figure 3 the corresponding device as shown.
[0166] Receive the current monitoring data collected by the Internet of Things technology, clean the data and store it in the database to form the stored monitoring data. Figure 3 The topological position defined by the transformer 320 shown is (1, 0), the topological position defined by the low-voltage distribution cabinet 340 is (2, 0), the topological position defined by the low-voltage distribution cabinet 350 is Device (2, 1), the topological position defined by the low-voltage distribution cabinet 360 is (2, 2), and the topological positions defined by the UPS devices 370-390 are (3, 0), (3, 1), and (3, 2) respectively. The status of the device is adjusted to 0, which triggers an alarm.
[0167] The system determines that the device alarm is triggered and generates an alarm.
[0168] Among them, when no alarm device is found above the transformer 320 (1, 0), the transformer 320 (1, 0) is the final aggregated alarm device. When the low-voltage distribution cabinet 340 (2, 0) and other 6 devices traverse the topology to the upper-level device, they are all located at the transformer 320 (1, 0). The system determines the alarm generated by the transformer 320 (1, 0) as the root cause alarm, and the alarms generated by the other 6 devices as aggregated alarms, and outputs the root cause alarm.
[0169] Therefore, in the process of operating and maintaining the infrastructure of the computer room of the data center, if the upstream equipment of the power system, such as transformers, high-voltage distribution cabinets, low-voltage distribution cabinets, UPS and other facilities and equipment, fails or operates abnormally, the downstream equipment will be affected. When the upstream main node equipment triggers an alarm, the associated downstream equipment will generate an alarm storm, which is not conducive to the rapid location and analysis of the alarm. Through the output method of the above-mentioned alarm root cause of the embodiment of the present disclosure, based on the equipment topology structure of the computer room infrastructure, through the collection and topological analysis of equipment alarm information, when a large number of alarms and alarm storms occur, the root cause can be quickly located to achieve the effect of alarm aggregation and alarm convergence.
[0170] Based on the above alarm root cause output method, the present disclosure also provides an alarm root cause output device. Figure 8 The device is described in detail.
[0171] Figure 8 The structural block diagram of the alarm root cause output device according to an embodiment of the present disclosure is schematically shown.
[0172] like Figure 8 As shown, the alarm root cause output device 800 of this embodiment includes a rule determination module 810 , a state identification module 820 and a device traversal module 830 .
[0173] The rule determination module 810 is used to determine the alarm rules of the upstream monitoring data. In one embodiment, the rule determination module 810 can be used to perform the operation S201 described above, which will not be repeated here.
[0174] The state identification module is used to identify the alarm triggering state of the current monitoring data item in the upstream monitoring data relative to the historical monitoring data item based on the alarm rule. In one embodiment, the state identification module 820 can be used to perform the operation S202 described above, which will not be repeated here.
[0175] The device traversal module is used to traverse the root cause device of the alarm root cause in the device topology according to the alarm trigger state of the current monitoring data item, and output the alarm root cause of the root cause device. In one embodiment, the device traversal module 830 can be used to perform the operation S203 described above, which will not be repeated here.
[0176] According to an embodiment of the present disclosure, any multiple modules among the rule determination module 810, the state identification module 820, and the device traversal module 830 can be combined into a single module, or any one of them can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present disclosure, at least one of the rule determination module 810, the state identification module 820, and the device traversal module 830 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware by any other reasonable means of integrating or packaging circuits, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, at least one of the rule determination module 810, the state identification module 820, and the device traversal module 830 can be at least partially implemented as a computer program module, which can perform the corresponding function when the computer program module is executed.
[0177] Figure 9 A block diagram of an electronic device suitable for implementing the alarm root cause output method according to an embodiment of the present disclosure is schematically shown.
[0178] like Figure 9 As shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. The processor 901 may, for example, include a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include an onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0179] Various programs and data required for the operation of the electronic device 900 are stored in the RAM 903. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0180] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the I / O interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card or a modem. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 910 as needed, so that a computer program read therefrom can be installed into the storage portion 908 as needed.
[0181] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0182] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.
[0183] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiments of the present disclosure.
[0184] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the processor 901 executes the computer program. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0185] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0186] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the processor 901, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0187] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0188] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0189] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.
[0190] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A method for outputting the root cause of an alarm, wherein: include: Determining equipment standard information corresponding to upstream monitoring data specifically includes: standardizing equipment type information corresponding to the upstream monitoring data; based on the standardized equipment type information, standardizing monitoring point information corresponding to the equipment type information to generate the equipment standard information; wherein the equipment type information includes an equipment name and a corresponding equipment number; and the monitoring point information includes a monitoring point name corresponding to the equipment name and a corresponding monitoring point number; Configuring a device topology structure according to the device standard information; Determine the alarm rules for upstream monitoring data; Cleaning and standardizing the received upstream monitoring data to generate upstream standardized data; Identifying a data change status of a current standardized data item relative to a historical standardized data item in the upstream standardized data; Determining the incoming monitoring data in the upstream standardized data according to the data change status of the current standardized data item; Identifying, based on the alarm rule, an alarm triggering status of a current monitoring data item in the upstream monitoring data relative to a historical monitoring data item, specifically comprising: querying, based on the equipment standard information corresponding to the current monitoring data item in the incoming monitoring data, the alarm rule that matches the current monitoring data item; performing a trigger check on the current monitoring data item in response to the query alarm rule to determine the alarm triggering status of the current monitoring data item relative to the historical monitoring data item; and According to the alarm triggering state of the current monitoring data item, the root cause device of the alarm root cause is traversed in the device topology structure, and the alarm root cause of the root cause device is output.
2. The method according to claim 1, wherein Configuring the device topology structure according to the device standard information includes: By editing the device topology, the device topology association corresponding to the upstream monitoring data is configured according to the device standard information to complete the configuration of the device topology structure.
3. The method according to claim 1, wherein Determine the alarm rules for upstream monitoring data, including: Creating a rule item of the alarm rule according to the equipment standard information corresponding to the upstream monitoring data; Generating an alarm condition corresponding to the rule item to determine the alarm rule; The alarm conditions include alarm level, alarm scope, alarm device type, alarm monitoring point, alarm triggering condition, alarm recovery condition and alarm topology relationship.
4. The method according to claim 1, wherein In response to the queried alarm rule, performing a trigger check on the current monitoring data item to determine an alarm trigger status of the current monitoring data item relative to the historical monitoring data item, including: Determining, through the trigger check, that the alarm trigger status of the equipment monitoring point corresponding to the historical monitoring data item is not triggered; In response to the untriggered alarm trigger state, when the monitoring information of the equipment monitoring point corresponding to the current monitoring data item meets the alarm trigger condition of the alarm rule, the alarm trigger state of the current monitoring data is newly added.
5. The method according to claim 1, wherein In response to the queried alarm rule, performing a trigger check on the current monitoring data item to determine an alarm trigger status of the current monitoring data item relative to the historical monitoring data item, further comprising: Determining, through the trigger check, that the alarm trigger status of the equipment monitoring point corresponding to the historical monitoring data item is triggered; In response to the triggered alarm trigger state, when the monitoring information of the equipment monitoring point corresponding to the current monitoring data item meets the alarm trigger condition of the alarm rule, the alarm trigger state of the current monitoring data is updated.
6. The method according to claim 1, wherein In response to the queried alarm rule, performing a trigger check on the current monitoring data item to determine an alarm trigger status of the current monitoring data item relative to the historical monitoring data item, further comprising: Based on the alarm triggering state, when there is no alarm change corresponding to the current monitoring data item, updating the alarm information corresponding to the current monitoring data item; or Based on the alarm triggering state, when there is an alarm change corresponding to the current monitoring data item, alarm information corresponding to the current monitoring data item is added.
7. The method according to claim 1, wherein Traversing the root cause device of the alarm root cause in the device topology structure according to the alarm triggering state of the current monitoring data item, and outputting the alarm root cause of the root cause device, including: Determining, according to the alarm triggering state of the current monitoring data item, position status information of the current monitoring device corresponding to the current monitoring data item in the device topology structure; In the device topology structure, according to the location status information of the current monitoring device, its upstream devices are traversed to output the root cause device.
8. The method according to claim 7, wherein: In the device topology structure, based on the location status information of the current monitoring device, traverse its upstream devices and output the root cause device, including: Determining the last final triggering device in the alarm triggering state during the traversal of the upstream devices; Based on the alarm aggregation information of the final triggering device, the final triggering device is output as the root cause device.
9. The method according to claim 7, wherein: Traversing the root cause device of the alarm root cause in the device topology structure according to the alarm triggering state of the current monitoring data item, and outputting the alarm root cause of the root cause device, including: Updating the aggregated alarm information of the root cause device; The updated aggregated alarm information is output as the alarm root cause.
10. An alarm root cause output device, wherein: include: A rule determination module is configured to determine an alarm rule for upstream monitoring data, wherein, before determining the alarm rule for the upstream monitoring data, the module further comprises: determining device standard information corresponding to the upstream monitoring data; and configuring a device topology structure according to the device standard information; a state identification module, configured to identify, based on the alarm rule, the alarm trigger state of the current monitoring data item in the upstream monitoring data relative to the historical monitoring data item, specifically comprising: querying, based on the equipment standard information corresponding to the current monitoring data item in the incoming monitoring data, an alarm rule matching the current monitoring data item; and performing a trigger check on the current monitoring data item in response to the query alarm rule to determine the alarm trigger state of the current monitoring data item relative to the historical monitoring data item; and A device traversal module is used to traverse the root cause device of the alarm root cause in the device topology according to the alarm trigger status of the current monitoring data item, and output the alarm root cause of the root cause device, Determining the equipment standard information corresponding to the upstream monitoring data includes: Standardizing device type information corresponding to the upstream monitoring data; Based on the standardized device type information, the monitoring point information corresponding to the device type information is standardized to generate the device standard information, wherein the device type information includes a device name and a corresponding device number; the monitoring point information includes a monitoring point name corresponding to the device name and a corresponding monitoring point number, Before identifying the alarm triggering status of the current monitoring data item in the upstream monitoring data relative to the historical monitoring data item based on the alarm rule, the method further includes: Cleaning and standardizing the received upstream monitoring data to generate upstream standardized data; Identifying a data change status of a current standardized data item relative to a historical standardized data item in the upstream standardized data; The incoming monitoring data in the upstream standardized data is determined according to the data change status of the current standardized data item.
11. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to perform the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method and device for identifying index exception reasons
CN110262937A
Fault root cause positioning method based on network topology and real-time alarm
CN112181758A