A server power supply reliability early warning method and system
By conducting multi-faceted monitoring of server power supplies and verifying them with CPLD, combined with fault diagnosis and repair robot positioning and repair, the problem of lack of early warning in power supply monitoring in existing technologies has been solved, thereby improving the stability and reliability of server power supplies.
Patent Information
- Application Number
- CN202210902093.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-07-28
AI Technical Summary
Existing server power monitoring strategies lack accurate fault warnings, affecting server power stability and operational continuity.
By monitoring the power supply status information, component characteristic parameter information, and operating index data, verifying BMC information using CPLD, and combining fault repair robots for location and repair, early warning prompts are achieved.
It improves the accuracy of server power supply warnings, avoids manual intervention, saves labor costs, and ensures the stability and reliability of server power supplies.
Smart Images

Figure CN115237719B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of server power management, and in particular to a server power reliability early warning method and system. BACKGROUND
[0002] With the rapid popularization and development of the Internet, cloud computing technology has made great progress, and cloud data centers have been established in various places. Cloud service products have gradually entered people's daily life. People's daily life relies more on network communication, and servers, which serve as the network hub, have become increasingly important. With the extensive use of servers, data center machine rooms have been established, and the number and scale of server rooms have been increasing. To ensure the data security of data center server rooms, the stability of data center server room power supply is particularly important. Server power is the most important power supply module of the server.
[0003] Currently, the server rack in the data center machine room is generally stored in a cabinet. Each cabinet contains multiple servers, and the rack density is high. Each server is independently powered and operated by at least two redundant power supplies. To meet the long-term and uninterrupted operation requirements of servers and complex front-end data processing conditions, the server power supply needs to have high reliability. If the server power supply is powered off due to internal hardware or software failure or external complex working conditions, it may cause the server to shut down due to the disappearance of redundancy and power supply interruption, which poses a security risk to customer data.
[0004] In actual work, the server room server is mainly monitored by the server BMC (Baseboard Management Controller) to monitor the power supply working state. For example, if the server power supply working alarm occurs, the BMC will record the alarm content and transmit it to the front-end monitoring interface through the BMC communication port via network communication. The maintenance personnel of the server room determines the fault cause by analyzing the BMC feedback log. However, this monitoring method is only suitable for replacing and maintaining the power supply after failure, and cannot avoid the risk of server power supply power failure by predicting the power failure in advance, which greatly affects the continuity and safety of the server operation. SUMMARY
[0005] The present application provides a server power reliability early warning method and system to solve the problem of lack of accurate fault early warning in the existing server power monitoring strategy, which affects the stability of the server power supply.
[0006] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0007] The present application provides a server power reliability early warning method, which comprises the following steps:
[0008] The state information of the power supply, the characteristic parameter information of each component of the power supply, and the index data of the operation of the power supply are monitored respectively.
[0009] The abnormal information monitored is compared with the preset abnormal value to obtain a corresponding risk level.
[0010] Based on the preset response strategy corresponding to the risk level, the response strategy includes on-site inspection, and the on-site inspection result and the risk level are combined to perform early warning prompt of the server power supply.
[0011] Further, the method further comprises the steps of:
[0012] For the server power supply that issues the early warning prompt, the fault point is located and repaired by the fault repair robot.
[0013] Further, the state information includes temperature information of the power supply, overcurrent signal and overvoltage signal of the power supply output.
[0014] Further, the monitoring of the state information is specifically:
[0015] In response to the state information abnormal alarm reported by the baseboard management controller, the complex programmable logic device is called to obtain the state information of the alarm item corresponding to the current abnormal alarm, and the state information obtained by the baseboard management controller is compared, and if the comparison result is consistent, the abnormal information is formed.
[0016] Further, the power supply components for monitoring the characteristic parameter information include power factor correction feedback circuit, diode circuit, communication optocoupler, each drive chip and standby control circuit.
[0017] Further, the monitoring of the characteristic parameter information is specifically:
[0018] The complex programmable logic device polls the characteristic parameter information in the server power supply register, and the characteristic parameter information is collected in real time by a sensor;
[0019] The characteristic parameter information is compared with the preset value, and the number of times of occurrence of the abnormal information is recorded, the number of times is marked according to a preset rule, and the marked value is taken as the abnormal information.
[0020] Further, the index data includes output power consumption, output current and voltage value of output signal of the power supply.
[0021] Further, the monitoring of the index data is specifically:
[0022] The complex programmable logic device polls the index data in the server power supply register, and the index data is obtained and / or calculated in real time by a power supply chip;
[0023] The index data is compared with preset values, and the number of occurrences of abnormal information is recorded, the number is marked according to a preset rule, and the marked value is taken as the abnormal information.
[0024] The second aspect of the application provides a server power reliability early warning system, the system comprises:
[0025] A power online monitoring module is configured to monitor state information of the power, component characteristic parameter information of the power, and index data of the power operation, respectively.
[0026] A reliability early warning module is configured to compare the monitored abnormal information with preset abnormal values to obtain a corresponding risk level.
[0027] A data center machine room control module is configured to correspond to a preset response strategy based on the risk level, the response strategy comprises on-site inspection, and the server power is prompted for early warning in combination with an on-site inspection result and the risk level, and the on-site inspection is configured to obtain server external environment information.
[0028] Further, the system further comprises a server power maintenance module, the server power maintenance module is configured to locate a fault point through a fault maintenance robot and maintain according to a preset strategy for the server power that issues the early warning prompt.
[0029] The server power reliability early warning system of the second aspect of the application can implement the method in the first aspect and the implementation manners of the first aspect, and achieve the same effects.
[0030] The effects provided in the summary are only the effects of the embodiments, not all the effects of the application, and one of the technical solutions has the following advantages or beneficial effects:
[0031] The application sets up multi-directional monitoring of the server power, including state information, index data and characteristic parameters, performs polling monitoring, and obtains accurate power state information through CPLD verification for the state information monitored by the BMC, avoids the false alarm condition of the existing single BMC monitoring mode, and ensures the accuracy of the early warning. The early warning power is positioned and maintained by the machine room robot, personnel are prevented from entering the machine room, the influence of the machine room environment is avoided, and the labor cost is saved. BRIEF DESCRIPTION OF DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0033] Figure 1 is a flowchart of the method embodiment of the present application;
[0034] Figure 2 is a structural schematic diagram of the system embodiment of the present application. DETAILED DESCRIPTION
[0035] In order to clearly illustrate the technical features of the present application, the present application will be described in detail below with reference to the specific embodiments and the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing the various structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. In addition, the present application can repeatedly refer to the same numbers and / or letters in different examples. Such repetition is for the purpose of simplification and clarity, and does not indicate the relationship between the various embodiments and / or settings discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. The present application omits the description of well-known components and processing techniques and processes to avoid unnecessary limitation of the present application.
[0036] The embodiment of the present application provides a server power supply reliability early warning method, and the method comprises the following steps:
[0037] S1, respectively, the state information of the power supply, the characteristic parameter information of each component of the power supply and the index data of the power supply operation are monitored;
[0038] S2, compare the abnormal information monitored with the preset abnormal value to obtain the corresponding risk level;
[0039] S3, based on the risk level corresponding to the preset coping strategy, the coping strategy includes on-site inspection, combined with the on-site inspection result and the risk level, the early warning prompt of the server power supply is carried out, and the on-site inspection is used to obtain the external environment information of the server.
[0040] In one implementation mode of the embodiment of the present application, the method further comprises the following steps:
[0041] For the server power supply that sends the early warning prompt, the fault point is located and repaired by the fault repair robot.
[0042] In step S1, the state information includes the temperature information of the power supply, the power supply output overcurrent signal and the overvoltage signal.
[0043] The monitoring of the state information is specifically: the server BMC polls the server power state information and compares with the specification value, if the power state information does not meet the requirement of the specification value, the BMC displays the power alarm state, based on the power alarm state, the parameter information of the server alarm power is read through the CPLD and compared with the alarm information fed back by the BMC, if the comparison result is different, the BMC is commanded to read and feed back again until the comparison result is the same, if the fault alarm information fed back by the BMC is matched, it is determined as abnormal information.
[0044] The server CPLD polls the characteristic parameters (real-time voltage values and working states of PFC feedback circuit, PFC OVP detection loop, key circuit diode, communication optocoupler, isolation drive IC, standby control chip, working temperature and state of standby integrated chip) of each component of the power supply collected by the power supply sensor in the register of the server power in real time, the CPLD records the parameter data of each key component of the power supply, analyzes the key parameter data of each component of the power supply and compares with the specification interval, if the polling data is within the standard specification interval, the power state output is 00, if the polling data exceeds the standard specification interval, the power state output of each power abnormal component is 01, if the polling data of the same component or the same characteristic parameter exceeds the standard specification interval for three times continuously, the power state output of each power abnormal component is 10.
[0045] The server CPLD monitors the output power consumption, output current, and each output signal voltage (12V, Vingood, Alert, PG) of the server power supply in real time. The index data is obtained through the power supply register, and the power chip obtains and / or calculates the index data in real time and stores it in the power supply register. The CPLD polls and records each parameter data and compares it with the specification value. To ensure the power supply redundancy of the whole machine, the output power consumption of the server power supply in the machine room should be less than 50% of the rated power (standard specification range); to ensure the stable power supply of the whole machine, the output current of the two power supplies of the server power supply in the machine room should meet the current sharing requirement (standard specification range: below 20% load, current unbalance less than 10%; 20% load and above, current unbalance less than 5%); to ensure the reliability of server power supply and communication, the output signal quality of 12V (standard specification range: 12.0V-12.8V), Vingood (standard specification range: 2.4V-3.46V), Alert (standard specification range: 2.4V-3.46V), and PG (standard specification range: 2.4V-3.46V) of the server power supply should be within the specification range. If the polling data is within the standard specification range, the power supply state output is 00; if the polling data exceeds the standard specification range, the abnormal state output of each power supply is 01 each time; if the polling data exceeds the standard specification range for three consecutive times, the power supply state output is 10. The server power supply state online monitoring module summarizes the above power supply alarm information and power supply state output value, and transmits the summarized information to the server power supply reliability warning module.
[0046] In step S2, the alarm information and power supply state output value of each power supply of the server are received and summarized, and a server power supply reliability state table is generated. Based on the power supply reliability state table information, the server power supply is divided into a low-risk area (power supply state output value is 00), a medium-risk area (power supply state output value is less than 10), and a high-risk area (power supply state output value is greater than or equal to 10).
[0047] For the low-risk area power supply, a low-risk power supply list is generated; for the medium-risk area power supply, a risk power supply list and corresponding alarm information are generated to the data center machine room control module for power supply reliability identification analysis. For the high-risk area power supply, the server power supply reliability warning module transmits the high-risk power supply list and corresponding alarm information to the data center machine room control module for power supply maintenance process analysis.
[0048] In step S3, for the medium-risk server power supply, the data center management personnel need to determine whether the risk warning information of each power supply is an effective alarm that needs to be controlled, and finally assess the warning level of the server power supply. If the final warning level of the server power supply is medium-risk and there is a need for on-site inspection, the accurate location of the faulty server component is located, the server power supply index collection command is issued to the data center automation maintenance robot through the Internet of Things wireless transmission technology, the data center automation maintenance robot moves to the fault server location according to the fault location, and the working state of the faulty power supply is photographed, videoed, and odor and other index collection is performed. The collected data is transmitted to the data center machine room control module, and the data center management personnel view and process the feedback information through the visual interface of the control module.
[0049] For the high-risk server power supply, the data center management personnel determine whether the risk warning information of each power supply is an effective alarm that needs to be controlled, and finally assess the warning level of the server power supply. If the final warning level of the server power supply is high-risk and there is a need for on-site power supply replacement, the data center management personnel issue a power supply replacement on-site confirmation instruction to the server power supply automation maintenance module, the server power supply automation maintenance module receives the requirement and locates the accurate position of the faulty server component, issues a server power supply index collection command to the data center automation maintenance robot, and the data center automation maintenance robot moves to the fault server location according to the fault location. The working state of the faulty power supply is photographed, videoed, and odor and other index collection is performed. The collected data is transmitted to the data center machine room control module. The data center management personnel view the feedback information and finally confirm the maintenance requirement, issue a formal maintenance instruction, and the data center automation maintenance robot moves to the fault power supply location through the mechanical arm to complete the automatic replacement of the faulty power supply and the re-powering operation.
[0050] As shown in FIG. 1, Figure 2 The embodiment of the present application also provides a server power supply reliability warning system, which comprises a power supply online monitoring module 1, a reliability warning module 2, a data center machine room control module 3, and a server power supply maintenance module 4.
[0051] The power supply online monitoring module 1 is used for monitoring the state information of the power supply, the characteristic parameter information of each component of the power supply, and the index data of the power supply operation, respectively; the reliability warning module 2 is used for comparing the monitored abnormal information with the preset abnormal value to obtain the corresponding risk level; the data center machine room control module 3 corresponds to a preset response strategy based on the risk level, the response strategy includes on-site inspection, and the on-site inspection is used for obtaining the external environment information of the server. The server power supply maintenance module 4 locates the fault point through the fault maintenance robot for the server power supply that issues the warning prompt and maintains according to the preset strategy.
[0052] The power supply online monitoring module: on the one hand, the server BMC polls the server power state information and compares with the specification value, if the power state information does not meet the specification value requirement, the BMC displays the power alarm state and transmits 10 to the power supply online monitoring module. The power supply online monitoring module receives the abnormal alarm information fed back by the BMC, reacts immediately and reads the parameter information of the server alarm power through the CPLD, and compares with the alarm information fed back by the BMC. If the comparison result is different, the BMC is instructed to read and feed back again until the comparison result is the same. If the fault alarm information fed back by the BMC matches, the power supply online monitoring module determines that the fault alarm information is correct, and transmits the fault alarm information and the power state output value to the reliability early warning module.
[0053] On the one hand, the server CPLD polls the characteristic parameters (real-time voltage values and working states of PFC feedback circuit, PFC OVP detection loop, key circuit diode, communication optocoupler, isolation drive IC, standby control chip, standby integrated chip, working temperature and state) of each component of the power supply collected by the power supply sensor in the register of the server power supply in real time, the CPLD records the parameter data of each key component of the power supply and transmits to the power supply online monitoring module, the server power supply state online monitoring module analyzes the key parameter data of each component of the power supply and compares with the specification interval. If the polling data is within the standard specification interval, the power state output is 00; if the polling data exceeds the standard specification interval, the power state output of each abnormal power supply component is 01; if the polling data of the same component or the same characteristic parameter exceeds the standard specification interval for three times in succession, the power state output of each abnormal power supply component is 10. The server power supply state online monitoring module summarizes the above power alarm information and power state output value, and transmits the summarized information to the reliability early warning module.
[0054] In one aspect, the server CPLD monitors the output power consumption, output current, and each output signal voltage (12V, Vingood, Alert, PG) of the server power supply in real time. The CPLD polls and records each parameter data and compares it with the specification value. To ensure the power supply redundancy of the whole machine, the output power consumption of the server power supply in the machine room should be less than 50% of the rated power (standard specification range); to ensure the power supply stability of the whole machine, the output current of the two power supplies of the server power supply in the machine room should meet the current sharing requirements (standard specification range: below 20% load, current imbalance less than 10%; 20% load and above, current imbalance less than 5%); to ensure the power supply and communication reliability of the server power supply, the 12V (standard specification range: 12.0V-12.8V), Vingood (standard specification range: 2.4V-3.46V), Alert (standard specification range: 2.4V-3.46V), and PG (standard specification range: 2.4V-3.46V) output signal quality of the server power supply should be within the specification range. If the polling data is within the standard specification range, the power supply state output is 00; if the polling data exceeds the standard specification range, the abnormal state output of each power supply is 01; if the polling data exceeds the standard specification range for three consecutive times, the power supply state output is 10. The server power supply state online monitoring module summarizes the above power supply alarm information and power supply state output value, and transmits the summarized information to the reliability warning module.
[0055] The reliability warning module receives and summarizes the alarm information and power supply state output value of each server power supply transmitted by the power supply online monitoring module, generates a server power supply reliability state table, and divides the server power supply into a low-risk area (power supply state output value is 00), a medium-risk area (power supply state output value is less than 10), and a high-risk area (power supply state output value is greater than or equal to 10) based on the power supply reliability state table information.
[0056] For the low-risk area power supply, the reliability warning module transmits the low-risk power supply list to the data center machine room control module for display. For the medium-risk area power supply, the reliability warning module transmits the medium-risk power supply list and corresponding alarm information to the data center machine room control module for power supply reliability identification analysis. For the high-risk area power supply, the reliability warning module transmits the high-risk power supply list and corresponding alarm information to the data center machine room control module for power supply repair process analysis.
[0057] The data center machine room control module receives the server power supply risk list and corresponding alarm information fed back by the reliability warning module, and the machine room management personnel views the server power supply risk list and corresponding alarm information through the visual interface of the data center machine room control module.
[0058] For the medium-risk server power supply, the room management personnel needs to determine whether the risk warning information of each power supply is an effective alarm that needs to be controlled, and finally assess the warning level of the server power supply. If the final warning level of the server power supply is medium-risk and there is a need for on-site inspection, the room management personnel issues an inspection instruction to the repair module, the repair module receives this requirement and locates the accurate position of the faulty server component, issues a server power supply index collection command to the data center automation repair robot through the Internet of Things wireless transmission technology, and the data center automation repair robot moves to the fault server position according to the fault positioning to collect indicators such as photos, videos, and odors of the working state of the faulty power supply, and transmits the collected data to the data center room control module. The room management personnel views and processes this feedback information through the visual interface of the control module.
[0059] For the medium-risk server power supply, the room management personnel needs to determine whether the risk warning information of each power supply is an effective alarm that needs to be controlled, and finally assess the warning level of the server power supply. If the final warning level of the server power supply is medium-risk and there is a need for on-site inspection, the room management personnel issues an inspection instruction to the repair module, the repair module receives this requirement and locates the accurate position of the faulty server component, issues a server power supply index collection command to the data center automation repair robot through the Internet of Things wireless transmission technology, and the data center automation repair robot moves to the fault server position according to the fault positioning to collect indicators such as photos, videos, and odors of the working state of the faulty power supply, and transmits the collected data to the data center room control module. The room management personnel views and processes this feedback information through the visual interface of the control module.
[0060] The above scheme can also realize the automatic monitoring, early warning, and repair of the reliability of the server power supply in the computer room.
[0061] Although the specific embodiments of the present application have been described in detail with reference to the accompanying drawings, this is not a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications or variations made on the basis of the technical solutions of the present application without creative labor are still within the scope of protection of the present application.
Claims
1. A method for early warning of server power supply reliability, characterized in that, The method comprises the following steps: Respectively monitoring state information of the power supply, characteristic parameter information of each component of the power supply and index data of the power supply operation; The state information comprises temperature information of the power supply, power supply output overcurrent signal and overvoltage signal; The monitoring of the state information is specifically as follows: In response to an abnormal alarm of the state information reported by the baseboard management controller, the complex programmable logic device is called to acquire state information of an alarm item corresponding to the current abnormal alarm, and the acquired state information is compared with the state information acquired by the baseboard management controller, if the comparison result is consistent, abnormal information is formed; The power supply components monitored for the characteristic parameter information comprise a power factor correction feedback circuit, a diode circuit, a communication optocoupler, each drive chip and a standby control circuit; The monitoring of the characteristic parameter information is specifically as follows: The complex programmable logic device polls the characteristic parameter information in the server power supply register, and the characteristic parameter information is collected in real time by a sensor; The characteristic parameter information is compared with a preset value, the number of times of occurrence of abnormal information is recorded, the number of times is marked according to a preset rule, and the marked value is taken as an abnormal value; The server CPLD polls the characteristic parameter index data of each component of the power supply collected in real time by the power supply sensor in the register of the server power supply, the CPLD records the parameter data of each key component of the power supply, analyzes the key parameter data of each component of the power supply and compares with a specification interval, if the polling data is within the standard specification interval range, the power supply state output is 00, if the polling data exceeds the standard specification interval range, the power supply state output of each power supply abnormal component is 01, if the polling data of the same component or the same characteristic parameter continuously exceeds the standard specification interval range for three times, the power supply state output of each power supply abnormal component is 10; The index data comprises output power consumption, output current and voltage value of an output signal of the power supply; The monitoring of the index data is specifically as follows: The complex programmable logic device polls the index data in the register of the server power supply, and the index data is acquired and / or calculated in real time by a power supply chip; The index data is compared with a preset value, the number of times of occurrence of abnormal information is recorded, the number of times is marked according to a preset rule, and the marked value is taken as an abnormal value; The server CPLD monitors the index data of the server power supply in real time. The index data is obtained through the power supply register. The power supply chip obtains and / or calculates the index data in real time and stores it in the power supply register. The CPLD polls and records each parameter data and compares it with the specification value. To ensure the power supply redundancy of the whole machine, the output power consumption of the server power supply in the machine room should be less than 50% of the rated power of the power supply. To ensure the stable power supply of the whole machine, the output current of the two power supplies of the server power supply in the machine room should meet the current sharing requirements. The standard specification range is as follows: below 20% load, the current sharing degree is less than 10%; 20% load and above, the current sharing degree is less than 5%. To ensure the power supply and communication reliability of the server power supply, the quality of the 12V, Vingood, Alert, and PG output signals of the server power supply should be within the specification range. If the polling data is within the standard specification range, the power supply state output is 00. If the polling data exceeds the standard specification range, the abnormal state output of each power supply is 01 each time. If the polling data exceeds the standard specification range for three consecutive times, the power supply state output is 10. The server power supply state online monitoring module summarizes the above power supply alarm information and power supply state output value, and transmits the summarized information to the server power supply reliability warning module. The abnormal value monitored is compared with the preset abnormal value to obtain the corresponding risk level. Specifically: The server power supply state online monitoring module transmits the alarm information and power supply state output value of each power supply of the server to the server power supply reliability state summary table generation module. Based on the power supply reliability state summary table information, the server power supply is divided into a low-risk area, i.e. the power supply state output value is 00; a medium-risk area, i.e. the power supply state output value is less than 10; and a high-risk area, i.e. the power supply state output value is greater than or equal to 10. For the power supply in the low-risk area, a low-risk power supply list is generated. For the power supply in the medium-risk area, a risk power supply list and corresponding alarm information are generated to the data center machine room control module for power supply reliability identification analysis. For the power supply in the high-risk area, the server power supply reliability warning module transmits the high-risk power supply list and corresponding alarm information to the data center machine room control module for power supply repair process analysis. Based on the risk level, a preset response strategy is determined. The response strategy includes on-site inspection. Based on the on-site inspection results and the risk level, a warning prompt for the server power supply is given. The on-site inspection is used to obtain the external environment information of the server.
2. The method of claim 1, wherein the server power reliability early warning method is characterized by, The method further includes the following steps: For the server power supply that gives the warning prompt, a fault locating robot is used to locate the fault point and repair it.
3. A server power supply reliability early warning system, characterized by, The system is used to implement the method of claim 1, and the system includes: A power supply online monitoring module for monitoring the state information of the power supply, the characteristic parameter information of each component of the power supply, and the index data of the power supply operation, respectively; A reliability warning module for comparing the monitored abnormal information with the preset abnormal value to obtain the corresponding risk level; The data center machine room control module is based on the risk level corresponding to a preset response strategy, the response strategy includes on-site inspection, combined with the on-site inspection result and the risk level, a pre-warning prompt of a server power supply is performed, and the on-site inspection is used to acquire server external environment information.
4. The server power reliability early warning system of claim 3, wherein, The system further comprises a server power supply maintenance module, which positions a fault point through a fault maintenance robot and maintains according to a preset strategy for a server power supply that issues the pre-warning prompt.
Citation Information
Patent Citations
Server power supply fault monitoring method and device and electronic equipment
CN113704049A