Server power failure monitoring method and device and electronic equipment
By standardizing the power sensor information of the server cluster and matching it with power alarm policies, the problem of automated monitoring of power failures in a large number of servers is solved, achieving efficient unified management and avoidance of false alarms.
Patent Information
- Application Number
- CN202110396596.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-13
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2041-09-13
AI Technical Summary
Existing technologies cannot achieve automated monitoring of power failures in large numbers of servers, and the significant differences between server models from different manufacturers make unified management impossible, which can easily lead to false alarms at the baseboard management controller level.
By acquiring power sensor information from the target server cluster, standardizing it, and matching it with power alarm policy configuration information, automated monitoring of power failures can be achieved, avoiding false alarms.
It enables automated monitoring and unified management of power failures for a large number of servers of different models, reduces false alarms at the baseboard management controller level, and improves operation and maintenance efficiency and monitoring accuracy.
Smart Images

Figure CN113704049B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a server power failure monitoring method and device and electronic equipment. BACKGROUND
[0002] The power failure of a server will have a huge impact on the business on the server, so it is necessary to monitor the power failure of the server.
[0003] In the related art, the power failure of a server can be preliminarily judged through the out-of-band server event log, or the current power-related alarm of the server can be viewed through the out-of-band Web page logged into the specific server. Among them, the out-of-band server event log is defined by each server manufacturer, and the content defined by different server manufacturers of different models is very different. Not only can it not achieve unified management, but also it is easy to produce false alarms at the Baseboard Management Controller (BMC) level. In addition, the above monitoring methods of the related art are only applicable to the power failure monitoring scene of a small batch of servers, and cannot realize the automatic monitoring of the power failure of a large batch (such as one million) of servers. SUMMARY
[0004] In order to solve the problems of the prior art, the embodiments of the present application provide a server power failure monitoring method, device and electronic equipment. The technical solution is as follows:
[0005] On the one hand, a server power failure monitoring method is provided, and the method comprises:
[0006] Obtaining power sensor information of a first server in a target server cluster; the power sensor information comprises server attribute information and power collection item information;
[0007] Standardizing the power collection item information to obtain standard power collection item information;
[0008] Determining at least one power alarm strategy in the power alarm strategy configuration information that matches the server attribute information; each power alarm strategy in the at least one power alarm strategy comprises the server attribute information and alarm power collection item information;
[0009] Matching the standard power collection item information with the alarm power collection item information corresponding to each power alarm strategy in the at least one power alarm strategy;
[0010] If there is matching target alarm power collection item information, determining that the monitoring result of the first server is a power failure alarm.
[0011] In another aspect, a server power failure monitoring apparatus is provided, the apparatus comprising:
[0012] a sensor information obtaining module configured to obtain power sensor information of a first server in a target server cluster; the power sensor information comprising server attribute information and power collection item information;
[0013] a standardization module configured to perform standardization processing on the power collection item information to obtain standard power collection item information;
[0014] a first matching module configured to determine at least one power alarm strategy in the power alarm strategy configuration information that matches the server attribute information; each power alarm strategy in the at least one power alarm strategy comprising the server attribute information and alarm power collection item information;
[0015] a second matching module configured to match the standard power collection item information with the alarm power collection item information corresponding to each power alarm strategy in the at least one power alarm strategy;
[0016] a monitoring result determining module configured to determine that the monitoring result of the first server is a power failure alarm if there is matching target alarm power collection item information.
[0017] In one possible implementation, the apparatus further comprises:
[0018] a first obtaining module configured to obtain the monitoring result of the first server in a preset monitoring period;
[0019] a first work order creating module configured to create a power failure alarm work order corresponding to the first server if the monitoring result in the preset monitoring period is a power failure alarm;
[0020] a first sending module configured to send the power failure alarm work order corresponding to the first server to a maintenance node and a spare part node, so that the spare part node determines a target spare part according to the power failure alarm work order, and the maintenance node performs failure maintenance processing on a power component of the first server according to the target spare part.
[0021] In one possible implementation, the apparatus further comprises:
[0022] a deployment unit determining unit configured to determine a deployment unit corresponding to the first server;
[0023] a first number determining module configured to determine the number of power failure alarm work orders corresponding to a target server in the deployment unit; the target server being a server in the deployment unit that has the same server attribute information as the first server;
[0024] a second work order creation module, configured to perform the step of creating the power failure alarm work order corresponding to the first server when the number does not exceed the preset number threshold.
[0025] In one possible implementation, the apparatus further includes:
[0026] a first generation module, configured to generate power abnormality information corresponding to the first server according to the server attribute information of the first server and the standard power collection item information when the number exceeds the preset number threshold;
[0027] a second sending module, configured to send the power abnormality information corresponding to the first server to a power abnormality confirmation node, and perform confirmation processing on the power abnormality information by the power abnormality confirmation node.
[0028] In one possible implementation, the apparatus further includes:
[0029] a judgment module, configured to judge whether the standard power collection item information matches abnormal power collection item information when there is no matching target alarm power collection item information;
[0030] a power abnormality confirmation module, configured to determine that the monitoring result of the first server is power abnormality when the standard power collection item information matches the abnormal power collection item information.
[0031] a second generation module, configured to generate power abnormality information corresponding to the first server according to the server attribute information of the first server and the standard collection item information if the monitoring result of the first server in the preset monitoring period is power abnormality;
[0032] a third sending module, configured to send the power abnormality information corresponding to the first server to a power abnormality confirmation node, and perform confirmation processing on the power abnormality information by the power abnormality confirmation system.
[0033] In one possible implementation, the apparatus further includes:
[0034] a second acquisition module, configured to acquire a confirmation result of the power abnormality information of the first server from the power abnormality confirmation node;
[0035] a third work order creation module, configured to create a power failure alarm work order corresponding to the first server when the confirmation result is power failure alarm.
[0036] In one possible implementation, the apparatus further includes:
[0037] The third generating module is configured to generate a power alarm strategy to be configured according to the power abnormality information of each server in the target server cluster.
[0038] The second quantity determining module is configured to determine a first quantity of the power abnormality information corresponding to the power alarm strategy to be configured.
[0039] The third obtaining module is configured to obtain, from the power abnormality confirmation node, a second quantity of confirmation results of power failure alarms for the power abnormality information corresponding to the power alarm strategy to be configured.
[0040] The updating module is configured to update the power alarm strategy configuration information according to the power alarm strategy to be configured when a ratio of the second quantity to the first quantity exceeds a preset ratio.
[0041] In one possible implementation, the power collection item information includes an original power collection item and an original collection value; and the standardization module includes:
[0042] The first determining module is configured to determine a standard power collection item matched with the original power collection item in preset power collection item mapping information, and the preset power collection item mapping information represents a corresponding relationship between the original power collection item and the standard power collection item.
[0043] The second determining module is configured to determine a standard collection value matched with the standard power collection item and the original collection value in preset collection value mapping information, and the preset collection value mapping information represents a corresponding relationship between the standard power collection item, the original collection value, and the standard collection value.
[0044] The fourth generating module is configured to obtain the standard power collection item information according to the standard power collection item and the standard collection value.
[0045] In another aspect, an electronic device is provided, including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the above-mentioned server power failure monitoring method.
[0046] In another aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by a processor to implement the above-mentioned server power failure monitoring method.
[0047] In another aspect, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the server power failure monitoring method provided in the various optional implementations described above.
[0048] The embodiment of the present application can realize automatic monitoring and alarming of power failures of a large number of servers of different models by obtaining power sensor information of a first server in a target server cluster, performing standardization processing on power collection item information in the power sensor information to obtain standard power collection item information, determining at least one power alarm strategy in power alarm strategy configuration information that matches server attribute information, then matching the standard power collection item information with power alarm collection items corresponding to each power alarm strategy in the at least one power alarm strategy, and determining that a monitoring result of the first server is a power failure alarm when there is matching target alarm power collection item information, so as to realize unified management while avoiding false alarms at the BMC level. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on these drawings.
[0050] Figure 1 is an architecture schematic diagram of a server power failure monitoring method provided by the embodiment of the present application;
[0051] Figure 2 is a flow schematic diagram of a server power failure monitoring method provided by the embodiment of the present application;
[0052] Figure 3 is an example of power collection item information provided by the embodiment of the present application;
[0053] Figure 4a is an example of preset power collection item mapping information provided by the embodiment of the present application;
[0054] Figure 4b is an example of preset collection value mapping information provided by the embodiment of the present application;
[0055] Figure 5 is an example of power alarm strategy configuration information provided by the embodiment of the present application;
[0056] Figure 6 is a flowchart of another server power failure monitoring method provided by an embodiment of the present application;
[0057] Figure 7 is an example of a power failure alarm work order provided by an embodiment of the present application;
[0058] Figure 8 is a flowchart of another server power failure monitoring method provided by an embodiment of the present application;
[0059] Figure 9 is a flowchart of another server power failure monitoring method provided by an embodiment of the present application;
[0060] Figure 10 is a structural block diagram of a server power failure monitoring device provided by an embodiment of the present application;
[0061] Figure 11 is a hardware structural block diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION
[0062] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0063] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0064] Please refer to Figure 1 which is an architecture schematic diagram of a server power failure monitoring method provided by an embodiment of the present application, including a server cluster 110, a power failure monitoring node 120, a power abnormality confirmation node 130, an operation and maintenance node 140, and a spare parts node 150, wherein:
[0065] The number of servers in the server cluster 110 can be deployed according to actual needs, and can be a small batch of servers or a large batch (such as millions) of servers. The server models in the server cluster 110 can also be deployed as different server models according to actual needs. In addition, the server cluster 110 can be located in the same deployment unit, such as in the same machine room.
[0066] The power failure monitoring node 120 can monitor the power failure of each server in the server cluster 110. Specifically, the power failure monitoring node 120 can determine whether a power failure alarm occurs in the server based on the power sensor information of the server and the power alarm strategy configuration information. If a power failure alarm occurs, a corresponding power failure alarm work order can be generated, and the power failure alarm work order is sent to the operation and maintenance node 140 and the spare parts node 150 respectively. The spare parts node 150 determines the target spare parts according to the power failure alarm work order, and the operation and maintenance node 140 obtains the target spare parts from the spare parts node 150 according to the power failure alarm work order, and then performs fault maintenance processing on the power component of the server.
[0067] In one possible implementation, in order to avoid false alarms, the server that appears the power failure alarm can also be confirmed by the power abnormality confirmation node 130 in combination with Figure 1 When it is confirmed that the power component of the server appears the power failure alarm, the power failure monitoring node 120 creates a power failure alarm work order and sends the power failure alarm work order to the operation and maintenance node 140 and the spare parts node 150 respectively.
[0068] The server power failure monitoring method of the embodiment of the application can be implemented based on cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network in a wide area network or a local area network to realize data calculation, storage, processing, and sharing.
[0069] Among them, cloud computing is a computing mode which distributes computing tasks on a large number of computing resources, so that various application systems can obtain computing power, storage space and information services according to needs. The network providing resources is called "cloud". The resources in the "cloud" are infinitely expandable to users and can be obtained at any time, used on demand, expanded at any time and paid according to use. As a basic capability provider of cloud computing, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established, and a plurality of types of virtual resources are deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines containing operating systems), storage devices, network devices. According to logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer is deployed on the PaaS layer. SaaS is a variety of business software such as web portal websites and SMS mass senders. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0070] It should be noted that the server in the embodiment of the application can be a stand-alone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms.
[0071] In one possible implementation, the server and the node involved in the embodiment of the application can be node devices in a blockchain system, capable of sharing the acquired and generated information to other node devices in the blockchain system, realizing information sharing between multiple node devices. The plurality of node devices in the blockchain system can be configured with the same blockchain, which is composed of a plurality of blocks, and the adjacent blocks have an association relationship, so that when the data in any block is tampered with, it can be detected through the next block, so as to avoid the data in the blockchain being tampered with, and ensure the security and reliability of the data in the blockchain.
[0072] Please refer to Figure 2As shown in the flow diagram of a server power failure monitoring method provided by the embodiment of the application, the method can be applied to Figure 1 the power failure monitoring node 120 in the server cluster 100. It should be noted that the present specification provides method operation steps as described in the embodiments or flow diagrams, but more or fewer operation steps can be included based on conventional or non-creative labor. The order of steps listed in the embodiments is only one of the many step execution orders, and does not represent the only execution order. In actual system or product execution, the method order can be executed in sequence or in parallel (such as parallel processor or multi-threaded processing environment) as shown in the embodiments or drawings. Specifically as shown in Figure 2 the method can include:
[0073] S201, obtaining power sensor information of a first server in a target server cluster; the power sensor information includes server attribute information and power collection item information.
[0074] The target server cluster refers to a monitored server cluster, which can be servers located in the same deployment unit, which can be but is not limited to a machine room unit. In addition, the size of the target server cluster can not be limited, and can include a small batch of servers, or a large batch of servers such as one million servers, and the models of the servers in the target server cluster are not limited, and can include hundreds or more different models.
[0075] The first server can be any server in the target server cluster. Specifically, sensor information collection instructions can be sent to each server in the target server cluster at a preset time interval, and the servers in the target server cluster report sensor information (Sensor Data Record, SDR) after receiving the sensor information collection instructions, and the sensor information includes power sensor information. The preset time interval can be set according to actual needs, for example, it can be set to 10-30 minutes, so that minute-level power failure monitoring of the target server cluster can be achieved. In a specific implementation, the servers in the target server cluster can be collected for minute-level sensor information through an intelligent platform management interface (Intelligent Platform Management Interface, IPMI).
[0076] The power supply sensor information includes server attribute information and power supply collection item information. The server attribute information can include information of dimensions such as server version number, version identifier, model, device type, machine room, operation and maintenance part, and business module. The power supply collection item information can include original power supply collection items and original collection values. The original power supply collection items are composed of power supply component identifiers (such as names) and collection items, and the original collection values are specific description information corresponding to the collection items. For example, the collection items can include input power, output power, health status, input voltage, output voltage, redundancy status, fan status, and the like.
[0077] Figure 3 is an example of power supply collection item information provided by the embodiment of the application, wherein "PSU0_STATUS", "PSU1_STATUS", "PSU0_PIN" and the like are original power supply collection items, PSU0 and PSU1 are names of power supply components, STATUS and PIN are collection items, and "Presence detected", "Presence detected, AC lost or out-of-range", and "232Watts" are corresponding original collection values, that is, specific descriptions of STATUS and PIN.
[0078] S203, standardizing the power supply collection item information to obtain standard power supply collection item information.
[0079] The standardization processing can unify the power supply collection item information, so as to facilitate subsequent processing. Specifically, the original power supply collection items and original collection values in the power supply collection item information can be standardized respectively, so as to obtain standard power supply collection items and standard collection values, which constitute the standard power supply collection item information. Based on this, in a possible implementation, the standardization processing of the power supply collection item information to obtain the standard power supply collection item information can include the following steps:
[0080] (1) determining a standard power supply collection item matched with the original power supply collection item in preset power supply collection item mapping information. The preset power supply collection item mapping information represents a corresponding relationship between the original power supply collection item and the standard power supply collection item.
[0081] In actual application, in order to reduce the data amount of the preset power supply collection item mapping information and improve the matching speed, when configuring the preset power supply collection item mapping information, for the original power supply collection item in the preset power supply collection item mapping information, the characters (such as 0, 1, 2, etc.) used to distinguish different power supply components in the power supply component identifier can be fuzzily processed, such as being replaced by a general character "*". That is, for "PSU0_STATUS" and "PSU1_STATUS", in the preset power supply collection item mapping information, it can be represented as "PSU*_STATUS", so that in the above matching, "PSU0_STATUS" and "PSU1_STATUS" can be both corresponded to "PSU*_STATUS", and then the standard power supply collection item having a corresponding relationship with "PSU*_STATUS" is determined from the preset power supply collection item mapping information, and then "PSU0_STATUS" and "PSU1_STATUS" are uniformly mapped to the standard power supply collection item.
[0082] Figure 4a Fig. 1 shows a preset power supply collection item mapping information provided by an embodiment of the present application, wherein the collection item ID is a standard power supply collection item, and the parameter is an original power supply collection item, and the number used to distinguish different power supply components is represented by "*" in the original power supply collection item. For example, as shown in Fig. 1, the original power supply collection item "PSU*_STATUS" has a corresponding relationship with the standard power supply collection item "power_status", so that "PSU1_STATUS" can be mapped to the standard power supply collection item "power_status". Figure 4a
[0083] In the embodiment of the present application, the standard power supply collection item can include power input power (power_input_power), power output power (power_ouptpower), power health status (power_status), power input voltage (power_input_voltage), power output voltage (power_output_voltage), power redundancy status (power_redundancy_status) and power fan status (power_fan_status).
[0084] (2) determining a standard collection value matching the standard power supply collection item and the original collection value in preset collection value mapping information. The preset collection value mapping information represents the corresponding relationship between the standard power supply collection item, the original collection value and the standard collection value.
[0085] Specifically, the current standard power collection item determined based on the step (1) can determine a candidate corresponding relationship matching the current standard power collection item from the preset collection value mapping information, the candidate corresponding relationship can include a plurality of corresponding relationships, each of the plurality of corresponding relationships includes the current determined standard power collection item, then a target corresponding relationship corresponding to the current original collection value is determined from the plurality of corresponding relationships, the target corresponding relationship includes the current original collection value, and then the standard collection value in the target corresponding relationship can be determined as the standard collection value corresponding to the current original collection value.
[0086] Figure 4b An example of the preset collection value mapping information provided by the embodiment of the application is shown, which Figure 4b The standard power collection item in the Figure 4a corresponds to the Figure 4b The corresponding relationship between the collection item ID, the original value and the standard value is shown, wherein the original value is the original collection value, and the standard value is the standard collection value.
[0087] In actual application, in order to improve the monitoring efficiency of the server power failure, when the preset collection value mapping information is configured, the standard collection value in the preset collection value mapping information can be corresponded to a numerical value, and the numerical value is used to uniquely identify a standard collection value. For example, in order to distinguish from the original collection value, the numerical value used to uniquely identify the standard collection value can be a negative number, as shown in Figure 4b The numerical value -9 is used to uniquely identify the standard collection value RedundantLost. It can be understood that the corresponding relationship in the preset collection value mapping information can be set and adjusted according to actual needs.
[0088] (3) obtaining the standard power collection item information according to the standard power collection item and the standard collection value.
[0089] The embodiment of the application can flexibly standardize the power sensor information of a plurality of different servers into standard power collection item information through the preset power collection item mapping information and the preset collection value mapping information.
[0090] S205, determining at least one power alarm strategy matching the server attribute information in the power alarm strategy configuration information.
[0091] Each of the at least one power alarm strategy includes the server attribute information and the alarm power collection item information.
[0092] The power supply alarm strategy configuration information includes server dimension configuration and power supply parameter dimension configuration corresponding to the server dimension configuration. The server dimension configuration can be configured according to server attribute information, for example, can include one or more combinations of server version number, version ID, model, device type, computer room, operation and maintenance department, and business module dimensions. In a specific implementation, the specific configuration of the server dimension can be determined according to the sensitivity of the business to the power supply failure, so that different alarm configurations can be realized according to different business needs. The power supply parameter dimension configuration can be configured by alarm power supply collection item information. The alarm power supply collection item information can include any combination of multiple standard power supply collection items (such as power supply input power, power supply output power, power supply health status, power supply input voltage, power supply output voltage, power supply redundancy status, and power supply fan status) and alarm standard collection values (such as AC lost, Failure detected, Redundancy Lost) corresponding to the standard power supply collection items in the combination.
[0093] Figure 5 An example of power supply alarm strategy configuration information provided by an embodiment of the application is shown as follows: Figure 5 As shown, each row represents a power supply alarm strategy, which includes server attribute information (such as model) and corresponding alarm power supply collection item information. In addition, in order to improve the flexibility of power supply failure alarm monitoring, whether automatic alarm is configured for each power supply alarm strategy in the power supply alarm strategy configuration information.
[0094] S207, matching the standard power supply collection item information with the alarm power supply collection item information corresponding to each power supply alarm strategy in the at least one power supply alarm strategy.
[0095] Specifically, if there is matching target alarm power supply collection item information, step S209 can be performed. The matching means that the standard power supply collection item and the standard collection value in the standard power supply collection item information are matched with the alarm power supply collection item and the alarm collection value in the target alarm power supply collection item information.
[0096] S209, determining that the monitoring result of the first server is a power supply failure alarm.
[0097] The embodiment of the application can realize automatic monitoring and alarm of power supply failure of a large number of servers of different models, and is very suitable for power supply failure monitoring of servers in a million-level computer room, which realizes unified management while avoiding false alarms at the BMC level.
[0098] In order to improve the operation and maintenance efficiency and avoid power supply false alarms caused by abnormal power-off of the computer room or rack, in a possible implementation manner, as shown in Figure 6Another flowchart of the method for monitoring server power failure is provided. After determining that the monitoring result of the first server is a power failure alarm in step S209, the method can further include:
[0099] S601, obtaining a monitoring result of the first server in a preset monitoring period.
[0100] The preset monitoring period can include N (N≥1) monitoring time points, and the time interval between adjacent two monitoring time points can be the same. The steps S201 to S209 are performed at each monitoring time point, so that N monitoring results of the first server at N monitoring time points can be obtained. Specifically, the specific value of N can be set according to actual needs. For example, N can be set to 30.
[0101] S603, determining whether the monitoring results in the preset monitoring period are all power failure alarms. If yes, step S605 can be performed.
[0102] Specifically, if the N continuous monitoring results are all power failure alarms, step S605 can be performed.
[0103] S605, creating a power failure alarm work order corresponding to the first server.
[0104] Specifically, the power failure alarm work order can include the deployment unit (such as the management unit of the machine room to which the first server belongs) of the first server, server attribute information, standard power acquisition item information, and maintenance processing suggestions and other information. For example, Figure 7 is an example of the power failure alarm work order provided by the embodiment of the present application.
[0105] S607, sending the power failure alarm work order corresponding to the first server to the operation and maintenance node and the spare part node.
[0106] After the embodiment of the present application monitors the power failure alarm of the server, the power failure alarm work order is created and sent to the operation and maintenance node and the spare part node, so that the spare part node can determine the target spare part according to the power failure alarm work order, and the operation and maintenance node can perform fault maintenance processing on the power component of the first server according to the target spare part, thereby reducing the human input, improving the operation and maintenance efficiency, and reducing the influence of power failure on business.
[0107] In addition, the embodiment of the present application creates the power failure alarm work order only when the monitoring results in the preset monitoring period are all power failure alarms, thereby avoiding the power false alarm caused by abnormal power failure of the machine room or rack, and having high anti-shake performance.
[0108] In order to further reduce the generation of false alarms, in one possible implementation, as shown inFigure 8 Another flowchart of the server power failure monitoring method is provided. Before step S605 of creating the power failure alarm work order corresponding to the first server, the method can further include:
[0109] S801, determining a deployment unit corresponding to the first server.
[0110] Specifically, the level of the deployment unit can be set according to actual application needs, and can be a computer room, a computer room management unit, or a rack.
[0111] S803, determining a number of power failure alarm work orders corresponding to a target server in the deployment unit; the target server is a server in the deployment unit having the same server attribute information as the first server.
[0112] Taking the server attribute information as a model and the deployment unit as a computer room as an example, the number of power failure alarm work orders generated by the servers having the same model as the first server and located in the same computer room can be counted.
[0113] S805, determining whether the number exceeds a preset number threshold.
[0114] Specifically, if the number does not exceed the preset number threshold, step S605 can be executed; if the number exceeds the preset number threshold, the alarm does not create a work order this time, and steps S807 to S809 can be executed.
[0115] The preset number threshold can be set according to experience in actual application, for example, can be 5.
[0116] S807, generating power abnormal information corresponding to the first server according to the server attribute information of the first server and standard power collection item information.
[0117] S809, sending the power abnormal information corresponding to the first server to a power abnormality confirmation node, and performing confirmation processing on the power abnormal information by the power abnormality confirmation node.
[0118] Specifically, the power abnormality confirmation node can arrange personnel to perform on-site inspection and confirmation processing on the power components of the first server according to the power abnormal information. The on-site inspection and confirmation processing mode can be confirmation of the on-site power alarm lamp.
[0119] In one possible implementation, continuing to refer to Figure 8 After step S809, the method can further include:
[0120] S811, obtaining a confirmation result of the power abnormal information of the first server from the power abnormality confirmation node.
[0121] S813, when the confirmation result is the power failure alarm, creating a power failure alarm work order corresponding to the first server.
[0122] Specifically, after sending the power abnormality information of the first server to the power abnormality confirmation node, the confirmation result of the power abnormality information of the first server can be obtained from the power abnormality confirmation node, and when the confirmation result is the power failure alarm, the step S605 is executed to create the power failure alarm work order corresponding to the first server.
[0123] Before creating the power failure alarm work order of the server, the power abnormality of the server is determined, and when the power abnormality is determined, the abnormality is confirmed by the power abnormality confirmation node, and the corresponding power failure alarm work order is created only when the confirmation result is the power failure alarm. Therefore, the server power failure alarm can be buffered, and the false alarm can be reduced.
[0124] In one possible implementation, as Figure 9 Another flowchart of a server power failure monitoring method is provided. After the step S207, the method can further include:
[0125] S901, if there is no matching target alarm power collection item information, determining whether the standard power collection item information matches the abnormal power collection item information.
[0126] Specifically, the standard power collection item information is matched with the alarm power collection item information corresponding to each power alarm strategy in the at least one power alarm strategy. If there is no matching target alarm power collection item information, it indicates that the first server does not currently meet the power failure alarm condition, and no power failure alarm is performed. However, in order to improve the accuracy of the server power failure monitoring, the embodiment of the present application further determines whether the standard power collection item information matches the abnormal power collection item information when it is determined that the server does not meet the power failure alarm condition. If they match, step S903 can be performed; otherwise, if they do not match, it indicates that the power of the first server is currently normal.
[0127] The abnormal power collection item information is used to determine whether the power component of the server is abnormal, and the abnormal power collection item information includes an abnormal power collection item and an abnormal value corresponding to the abnormal power collection item. The abnormal power collection item can be determined according to the standard power collection item in the power alarm strategy configuration information. For example, the abnormal power collection item can include power input power, power output power, power health status, and power redundancy status. Specifically, refer to the abnormal power collection item information table shown in Table 1 below.
[0128] Table 1
[0129] Abnormal power collection item Outlier Power input power 0 / No Reading Power output power 0 / No Reading Power health status AC lost; Failure detected; Predictive failure Power redundancy status Redundancy Lost
[0130] S903, determining that the monitoring result of the first server is power abnormality.
[0131] S905, judging whether the monitoring result of the first server in a preset monitoring period is power abnormality, if yes, step S907 can be executed.
[0132] The preset monitoring period can include N (N≥1) monitoring time points, the time interval between adjacent two monitoring time points can be same, and the above steps S201 to S209 are executed at each monitoring time point, so that N monitoring results corresponding to N monitoring time points of the first server can be obtained. Specifically, the specific value of N can be set according to actual needs, for example, N can be set to 30. If the above N continuous monitoring results are all power abnormality, step S907 can be executed.
[0133] S907, generating power abnormality information corresponding to the first server according to the server attribute information and the standard collection item information of the first server.
[0134] S909, sending the power abnormality information corresponding to the first server to a power abnormality confirmation node, and confirming the power abnormality information by the power abnormality confirmation system.
[0135] Specifically, the power abnormality confirmation node can arrange personnel to perform on-site inspection and confirmation on the power component of the first server according to the power abnormality information.
[0136] In one possible implementation, continuing to refer to Figure 9 After the above step S909, the method can further include:
[0137] S911, obtaining the confirmation result of the power abnormality information of the first server from the power abnormality confirmation node.
[0138] S913, when the confirmation result is power failure alarm, creating a power failure alarm work order corresponding to the first server.
[0139] The embodiment of the application delivers the server which does not meet the power failure alarm but meets the power abnormality to the power abnormality confirmation node for confirmation processing, and then creates a corresponding power failure alarm work order when the confirmation result is power failure alarm, so that the accuracy of server power failure monitoring is improved.
[0140] In view of the different degrees of association between different power supply parameters (such as model) of different server attribute information and power supply failure, in order to further improve the accuracy of server power supply failure monitoring, the power supply alarm strategy configuration information can also be updated according to the power supply abnormal information generated when the power supply is abnormal. Based on this, in a possible implementation, the method can further include the following steps:
[0141] (1) generating a power supply alarm strategy to be configured according to the power supply abnormal information of each server in the target server cluster;
[0142] (2) determining a first number of the power supply abnormal information corresponding to the power supply alarm strategy to be configured;
[0143] (3) obtaining a second number of confirmation results that are power supply failure alarms from the power supply abnormal confirmation node for the power supply abnormal information corresponding to the power supply alarm strategy to be configured;
[0144] Specifically, the power supply abnormal confirmation node can determine the power supply failure alarm in the power supply abnormal information through on-site inspection and manufacturer failure analysis, etc. confirmation processing mode, wherein the on-site inspection is to judge whether it belongs to power supply failure according to the situation of the power supply lamp and whether the power supply line is plugged in, and the manufacturer failure analysis is to analyze the reason for power failure by analyzing the bad parts replaced by failure.
[0145] (4) when the ratio of the second number to the first number exceeds a preset ratio, updating the power supply alarm strategy configuration information according to the power supply alarm strategy to be configured.
[0146] The preset ratio can be set according to actual application, for example, the preset ratio can be set to 95%. The specific updating method can be to add the power supply alarm strategy to be configured to the power supply alarm strategy configuration information. Table 2 is a confirmation table of the power supply alarm strategy to be configured provided by the embodiment of the present application. If the preset ratio is 95% for example, 3 power supply alarm strategies to be configured in Table 2 can be updated to the power supply alarm strategy configuration information.
[0147] Table 2
[0148]
[0149] Corresponding to the server power supply failure monitoring method provided by the above several embodiments, the embodiment of the present application also provides a server power supply failure monitoring device. Since the server power supply failure monitoring device provided by the embodiment of the present application corresponds to the server power supply failure monitoring method provided by the above several embodiments, the implementation modes of the foregoing server power supply failure monitoring method are also applicable to the server power supply failure monitoring device provided by the present embodiment, which will not be described in detail in the present embodiment.
[0150] Referring to Figure 10 , it is a structural schematic diagram of a server power failure monitoring device provided by an embodiment of the application. The device 1000 has the function of implementing the server power failure monitoring method in the above method embodiment. The function can be implemented by hardware, or corresponding software can be executed by hardware. As Figure 10 shown, the device 1000 can include:
[0151] A sensor information acquisition module 1010 is configured to acquire power sensor information of a first server in a target server cluster. The power sensor information includes server attribute information and power collection item information.
[0152] A standardization module 1020 is configured to perform standardization processing on the power collection item information to obtain standard power collection item information.
[0153] A first matching module 1030 is configured to determine at least one power alarm strategy in power alarm strategy configuration information that matches the server attribute information. Each power alarm strategy in the at least one power alarm strategy includes the server attribute information and alarm power collection item information.
[0154] A second matching module 1040 is configured to match the standard power collection item information with alarm power collection item information corresponding to each power alarm strategy in the at least one power alarm strategy.
[0155] A monitoring result determination module 1050 is configured to determine that a monitoring result of the first server is a power failure alarm if there is target alarm power collection item information that matches.
[0156] In one possible implementation, the device can further include:
[0157] A first acquisition module is configured to acquire a monitoring result of the first server in a preset monitoring period.
[0158] A first work order creation module is configured to create a power failure alarm work order corresponding to the first server if the monitoring result in the preset monitoring period is a power failure alarm.
[0159] A first sending module is configured to send the power failure alarm work order corresponding to the first server to a maintenance node and a spare part node, so that the spare part node determines a target spare part according to the power failure alarm work order, and the maintenance node performs failure maintenance processing on a power component of the first server according to the target spare part.
[0160] In one possible implementation, the device can further include:
[0161] The deployment unit determination unit is configured to determine a deployment unit corresponding to the first server;
[0162] The first quantity determination module is configured to determine a quantity of power failure alarm work orders corresponding to a target server in the deployment unit; the target server is a server in the deployment unit that has the same server attribute information as the first server;
[0163] The second work order creation module is configured to, when the quantity does not exceed the preset quantity threshold, execute the step of creating the power failure alarm work order corresponding to the first server.
[0164] In one possible implementation, the apparatus can further include:
[0165] The first generation module is configured to, when the quantity exceeds the preset quantity threshold, generate power abnormality information corresponding to the first server according to the server attribute information of the first server and standard power collection item information;
[0166] The second sending module is configured to send the power abnormality information corresponding to the first server to a power abnormality confirmation node, and perform confirmation processing on the power abnormality information by the power abnormality confirmation node.
[0167] In one possible implementation, the apparatus can further include:
[0168] The judgment module is configured to, when there is no matching target alarm power collection item information, judge whether the standard power collection item information matches abnormal power collection item information;
[0169] The power abnormality confirmation module is configured to, when the matching exists, determine that a monitoring result of the first server is power abnormality;
[0170] The second generation module is configured to, if the monitoring result of the first server in the preset monitoring period is power abnormality, generate power abnormality information corresponding to the first server according to the server attribute information of the first server and standard collection item information;
[0171] The third sending module is configured to send the power abnormality information corresponding to the first server to a power abnormality confirmation node, and perform confirmation processing on the power abnormality information by the power abnormality confirmation system.
[0172] In one possible implementation, the apparatus can further include:
[0173] The second acquisition module is configured to acquire, from the power abnormality confirmation node, a confirmation result of the power abnormality information of the first server;
[0174] The third work order creation module is configured to create a power failure alarm work order corresponding to the first server when the confirmation result is a power failure alarm.
[0175] In one possible implementation, the apparatus can further include:
[0176] The third generation module is configured to generate a power alarm policy to be configured according to the power abnormality information of each server in the target server cluster.
[0177] The second quantity determination module is configured to determine a first quantity of the power abnormality information corresponding to the power alarm policy to be configured.
[0178] The third acquisition module is configured to acquire, from the power abnormality confirmation node, a second quantity of confirmation results that are power failure alarms, with respect to the power abnormality information corresponding to the power alarm policy to be configured.
[0179] The updating module is configured to update the power alarm policy configuration information according to the power alarm policy to be configured, when a ratio of the second quantity to the first quantity exceeds a preset ratio.
[0180] In one possible implementation, the power collection item information includes an original power collection item and an original collection value; and the standardization module includes:
[0181] The first determination module is configured to determine a standard power collection item that matches the original power collection item in preset power collection item mapping information, wherein the preset power collection item mapping information represents a correspondence between the original power collection item and the standard power collection item.
[0182] The second determination module is configured to determine a standard collection value that matches the standard power collection item and the original collection value in preset collection value mapping information, wherein the preset collection value mapping information represents a correspondence between the standard power collection item, the original collection value, and the standard collection value.
[0183] The fourth generation module is configured to obtain the standard power collection item information according to the standard power collection item and the standard collection value.
[0184] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, only uses the division of the above functional modules as an example for illustration. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be repeated here.
[0185] The electronic device provided by the embodiment of the present application comprises a processor and a memory, the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program is loaded and executed by the processor to implement the server power failure monitoring method provided by the above method embodiment.
[0186] The memory can be used to store software programs and modules, and the processor executes various functions and the server power failure monitoring by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory can also include a memory controller to provide the processor with access to the memory.
[0187] The method provided by the embodiment of the present application can be executed in a computer terminal, a server or a similar computing device. Taking the case of running on a server, Figure 11 is a hardware structure block diagram of a server provided by the embodiment of the present application for running a server power failure monitoring method, as Figure 11 shown, the server 1100 can have a large difference due to different configurations or performances, and can include one or more central processing units (CPU) 1110 (the processor 1110 can include but not limited to a microprocessor MCU or a programmable logic device FPGA processing device), a memory 1130 for storing data, one or more storage media 1120 (such as one or more mass storage devices) for storing application programs 1123 or data 1122. Among them, the memory 1130 and the storage medium 1120 can be temporary storage or persistent storage. The program stored in the storage medium 1120 can include one or more modules, each module can include a series of instruction operations in the server. Further, the central processing unit 1110 can be configured to communicate with the storage medium 1120 and execute a series of instruction operations in the storage medium 1120 on the server 1100. The server 1100 can also include one or more power supplies 1160, one or more wired or wireless network interfaces 1150, one or more input and output interfaces 1140, and / or one or more operating systems 1121, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0188] The input / output interface 1140 can be configured to receive or transmit data via a network. Examples of the network can include a wireless network provided by a communication provider of the server 1100. In an example, the input / output interface 1140 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In an example, the input / output interface 1140 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.
[0189] Those skilled in the art can understand that, Figure 11 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 1100 can further include more or fewer components than those shown, or have a different configuration of components than those shown. Figure 11 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 1100 can further include more or fewer components than those shown, or have a different configuration of components than those shown. Figure 11 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 1100 can further include more or fewer components than those shown, or have a different configuration of components than those shown.
[0190] The embodiments of the present application also provide a computer readable storage medium, which can be arranged in an electronic device to store at least one instruction or at least one program for implementing a server power failure monitoring method. The at least one instruction or the at least one program is loaded and executed by the processor to implement the server power failure monitoring method provided by the above-mentioned method embodiments.
[0191] Optionally, in the present embodiment, the computer readable storage medium can include, but is not limited to: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0192] The embodiments of the present application also provide a computer program product or computer program, which includes computer instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the electronic device execute the server power failure monitoring method provided in the above-mentioned various optional implementation manners.
[0193] It should be noted that the above-mentioned embodiments of the present application are merely intended to describe the present application and are not intended to limit the present application. The above-mentioned embodiments of the present application are described in a progressive manner, and the same or similar parts among the embodiments can be mutually referred to. Each embodiment focuses on the difference from other embodiments. In particular, the device embodiments are described simply because they are basically similar to the method embodiments, and the relevant parts can be referred to the description of the method embodiments.
[0194] The embodiments in the present specification are described in a progressive manner, and the same or similar parts among the embodiments can be mutually referred to. Each embodiment focuses on the difference from other embodiments. In particular, the device embodiments are described simply because they are basically similar to the method embodiments, and the relevant parts can be referred to the description of the method embodiments.
[0195] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or can be completed by a program instructing the relevant hardware, and the program can be stored in a computer readable storage medium, such as a read-only memory, a magnetic disk or an optical disk.
[0196] The above-mentioned embodiments are merely preferred embodiments of the present application, and are not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for monitoring power failure of a server, the method comprising: The method comprises: obtaining power supply sensor information of a first server in a target server cluster; the power supply sensor information comprises server attribute information and power supply collection item information, the server attribute information comprises a model, and the power supply collection item information comprises an original power supply collection item and an original collection value, the original power supply collection item being composed of a power supply component identifier and a collection item; determining a standard power supply collection item matched with the original power supply collection item in preset power supply collection item mapping information; the preset power supply collection item mapping information represents a corresponding relationship between the original power supply collection item and the standard power supply collection item, and characters in the power supply component identifier of the original power supply collection item used to distinguish different power supply components are subjected to fuzzy processing; determining a standard collection value matched with the standard power supply collection item and the original collection value in preset collection value mapping information; the preset collection value mapping information represents a corresponding relationship among the standard power supply collection item, the original collection value and the standard collection value; obtaining standard power supply collection item information according to the standard power supply collection item and the standard collection value; determining at least one power supply alarm strategy matched with the server attribute information in power supply alarm strategy configuration information; each power supply alarm strategy in the at least one power supply alarm strategy comprises the server attribute information and alarm power supply collection item information; the power supply alarm strategy configuration information comprises server dimension configuration and power supply parameter dimension configuration corresponding to the server dimension configuration; wherein the server dimension configuration is configured according to the server attribute information, and the power supply parameter dimension configuration is configured for alarm through the alarm power supply collection item information; matching the standard power supply collection item information with alarm power supply collection item information corresponding to each power supply alarm strategy in the at least one power supply alarm strategy; if there is matched target alarm power supply collection item information, determining that a monitoring result of the first server is a power supply fault alarm of a power supply component.
2. The method of claim 1, wherein, The method further comprises: obtaining a monitoring result of the first server in a preset monitoring period; if the monitoring result in the preset monitoring period is all power supply fault alarms, creating a power supply fault alarm work order corresponding to the first server; sending the power supply fault alarm work order corresponding to the first server to an operation and maintenance node and a spare part node, so that the spare part node determines a target spare part according to the power supply fault alarm work order, and the operation and maintenance node performs fault maintenance processing on the power supply component of the first server according to the target spare part.
3. The method of claim 2, wherein the server power failure monitoring method is characterized by, Before creating the power supply fault alarm work order corresponding to the first server, the method further comprises: determining a deployment unit corresponding to the first server; determining a number of power supply fault alarm work orders corresponding to a target server in the deployment unit; the target server is a server in the deployment unit having the same server attribute information as the first server; if the number does not exceed a preset number threshold, performing the step of creating the power supply fault alarm work order corresponding to the first server.
4. The method of claim 3, wherein the server power failure monitoring method is characterized by, The method further comprises: If the number exceeds a preset number threshold, power abnormal information corresponding to the first server is generated according to server attribute information of the first server and standard power collection item information; The power abnormal information corresponding to the first server is sent to a power abnormality confirmation node, and the power abnormal information is processed by the power abnormality confirmation node.
5. The method of claim 2, wherein the server power failure monitoring method is characterized by, The method further comprises: If there is no matching target alarm power collection item information, it is determined whether the standard power collection item information matches abnormal power collection item information; If they match, it is determined that the monitoring result of the first server is a power abnormality; If the monitoring result of the first server in the preset monitoring period is a power abnormality, power abnormal information corresponding to the first server is generated according to server attribute information of the first server and standard collection item information; The power abnormal information corresponding to the first server is sent to a power abnormality confirmation node, and the power abnormal information is processed by the power abnormality confirmation system.
6. The method of claim 4 or 5, wherein, The method further comprises: From the power abnormality confirmation node, the confirmation result of the power abnormal information of the first server is obtained; When the confirmation result is a power failure alarm, a power failure alarm work order corresponding to the first server is created.
7. The method of claim 4 or 5, wherein, The method further comprises: According to the power abnormal information of each server in the target server cluster, a power alarm strategy to be configured is generated; A first number of the power abnormal information corresponding to the power alarm strategy to be configured is determined; For the power abnormal information corresponding to the power alarm strategy to be configured, a second number of confirmation results that are power failure alarms is obtained from the power abnormality confirmation node; When the ratio of the second number to the first number exceeds a preset ratio, the power alarm strategy configuration information is updated according to the power alarm strategy to be configured.
8. A server power failure monitoring apparatus, characterized by comprising: The device comprises: A sensor information acquisition module for acquiring power sensor information of a first server in a target server cluster; the power sensor information includes server attribute information and power collection item information, the server attribute information includes a model, and the power collection item information includes an original power collection item and an original collection value, the original power collection item being composed of a power component identifier and a collection item; A standardization module for determining a standard power collection item that matches the original power collection item in preset power collection item mapping information; the preset power collection item mapping information represents the correspondence between the original power collection item and the standard power collection item, and characters in the power component identifier of the original power collection item that are used to distinguish different power components are blurred; determining a standard collection value that matches the standard power collection item and the original collection value in preset collection value mapping information, the preset collection value mapping information representing the correspondence between the standard power collection item, the original collection value, and the standard collection value; obtaining standard power collection item information according to the standard power collection item and the standard collection value; The first matching module is configured to determine at least one power alarm strategy matched with the server attribute information in the power alarm strategy configuration information; each power alarm strategy in the at least one power alarm strategy comprises the server attribute information and alarm power collection item information; the power alarm strategy configuration information comprises a server dimension configuration and a power parameter dimension configuration corresponding to the server dimension configuration; wherein the server dimension configuration is configured according to the server attribute information, and the power parameter dimension configuration is configured by the alarm power collection item information; The second matching module is configured to match the standard power collection item information with the alarm power collection item information corresponding to each power alarm strategy in the at least one power alarm strategy; The monitoring result determination module is configured to determine that the monitoring result of the first server is a power failure alarm of a power component when there is the matched target alarm power collection item information.
9. The server power failure monitoring apparatus of claim 8, wherein, The device further comprises: The first obtaining module is configured to obtain a monitoring result of the first server in a preset monitoring period; The first work order creation module is configured to create a power failure alarm work order corresponding to the first server if the monitoring result in the preset monitoring period is a power failure alarm; The first sending module is configured to send the power failure alarm work order corresponding to the first server to a maintenance node and a spare part node, so that the spare part node determines a target spare part according to the power failure alarm work order, and the maintenance node performs fault maintenance processing on the power component of the first server according to the target spare part.
10. The server power failure monitoring apparatus of claim 9, wherein, The device further comprises: The deployment unit determination unit is configured to determine a deployment unit corresponding to the first server; The first number determination module is configured to determine a number of power failure alarm work orders corresponding to a target server in the deployment unit; the target server is a server in the deployment unit having the same server attribute information as the first server; The second work order creation module is configured to perform the step of creating the power failure alarm work order corresponding to the first server if the number does not exceed a preset number threshold.
11. The server power failure monitoring apparatus of claim 10, wherein, The device further comprises: The first generation module is configured to generate power anomaly information corresponding to the first server according to the server attribute information of the first server and standard power collection item information if the number exceeds the preset number threshold; The second sending module is configured to send the power anomaly information corresponding to the first server to a power anomaly confirmation node, and perform confirmation processing on the power anomaly information by the power anomaly confirmation node.
12. The server power failure monitoring apparatus of claim 9, wherein, The device further comprises: The judgment module is configured to judge whether the standard power collection item information matches abnormal power collection item information if there is no matched target alarm power collection item information; The power anomaly confirmation module is configured to determine that the monitoring result of the first server is a power anomaly if the standard power collection item information matches the abnormal power collection item information. The second generation module is configured to generate power supply abnormality information corresponding to the first server according to server attribute information and standard collection item information of the first server if the monitoring results of the first server in the preset monitoring period are all power supply abnormality. The third sending module is configured to send the power supply abnormality information corresponding to the first server to a power supply abnormality confirmation node, and perform confirmation processing on the power supply abnormality information by the power supply abnormality confirmation system.
13. The server power failure monitoring apparatus of claim 11 or 12, wherein, The apparatus further includes: The second obtaining module is configured to obtain a confirmation result of the power supply abnormality information of the first server from the power supply abnormality confirmation node. The third work order creation module is configured to create a power supply failure alarm work order corresponding to the first server when the confirmation result is a power supply failure alarm.
14. The server power failure monitoring apparatus of claim 11 or 12, wherein, The apparatus further includes: The third generation module is configured to generate a power supply alarm strategy to be configured according to power supply abnormality information of each server in the target server cluster. The second quantity determination module is configured to determine a first quantity of the power supply abnormality information corresponding to the power supply alarm strategy to be configured. The third obtaining module is configured to obtain, from the power supply abnormality confirmation node, a second quantity of confirmation results that are power supply failure alarms for the power supply abnormality information corresponding to the power supply alarm strategy to be configured. The updating module is configured to update the power supply alarm strategy configuration information according to the power supply alarm strategy to be configured when a ratio of the second quantity to the first quantity exceeds a preset ratio.
15. An electronic device, comprising: The computer readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the server power failure monitoring method in any one of claims 1-7.
16. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the server power failure monitoring method in any one of claims 1-7.
Citation Information
Patent Citations
Example monitoring method, computer readable storage medium, and terminal device
CN109240876A
UPS monitoring method, equipment, storage medium and device based on cloud platform
CN111443791A
Cluster alarm method and related device
CN112636979A