A server power failure simulation device, method, terminal and storage medium

By simulating the server power failure mode and recording the power indicator status and reading value, the problem of low power failure analysis efficiency in the existing technology is solved, and rapid fault positioning and maintenance efficiency are improved.

CN115995173BActive Publication Date: 2025-08-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211720361.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-08-01
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

In the prior art, when the server power supply fails, fault analysis usually requires on-site inspection, resulting in inefficient maintenance.

Method used

Provide a server power failure simulation device and method, which simulates multiple power failure modes through hardware and software, records the power indicator status and reading value, and realizes rapid positioning of faults.

Benefits of technology

During the test phase, the power failure mode is simulated and the power information is recorded. The server can quickly locate the cause of failure when it is actually running, improve maintenance efficiency, and avoid the server being offline after bursts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115995173B_ABST
    Figure CN115995173B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of server power failure, and specifically discloses a server power failure simulation device, method, terminal and storage medium. A number of power failure modes are predefined, and an electronic switch is set on the circuit corresponding to each power failure mode. The simulation controller is electrically connected to each electronic switch respectively, triggers each electronic switch to close in sequence, simulates the corresponding power failure mode, and obtains and records the power indicator status and power reading under each power failure mode. The present invention simulates multiple power failure modes, obtains the power indicator status and power reading under each power failure mode. When a power failure occurs during the actual operation of the server, the operation and maintenance personnel can correspond to the corresponding power failure mode according to the power indicator status and power reading, locate the cause of the failure, greatly improve the maintenance efficiency, and save time and effort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of server power failure, and in particular, to a server power failure simulation device, method, terminal, and storage medium. Background Art

[0002] The focus of attention in the operation and maintenance of server systems is first on reliability and reducing the offline rate. Among them, the server power supply (Server PSU) is a very important operating component. If the Server PSU fails, it may cause the server system to shut down or go offline. Therefore, for any Server PSU alarm, server operation and maintenance personnel need to go to the site to troubleshoot.

[0003] Currently, the reliability requirements for Server PSUs are quite high, and current designs can all provide a stable and continuously operating server power supply. However, there is still a possibility of failure for Server PSUs. For the failure mode (DFMEA) of the server system during a failure, often only the root cause analysis (RCA) is carried out when it actually occurs on-site, which is time-consuming and laborious and affects the maintenance efficiency. Summary of the Invention

[0004] To solve the above problems, the present invention provides a server power failure simulation device, method, terminal, and storage medium, which simulate multiple power failure modes, obtain the power indicator status and power readings under each power failure mode. When a power failure occurs during the actual operation of the server, operation and maintenance personnel can correspond to the corresponding power failure mode according to the power indicator status and power readings, locate the cause of the failure, greatly improve the maintenance efficiency, and save time and effort.

[0005] In a first aspect, the technical solution of the present invention provides a server power failure simulation device, which is characterized by including: a simulation controller and a plurality of electronic switches;

[0006] Pre-define a plurality of power failure modes, and set an electronic switch on the circuit corresponding to each power failure mode;

[0007] The simulation controller is electrically connected to each electronic switch respectively, and sequentially triggers each electronic switch to close, simulating the corresponding power failure mode, and obtaining and recording the power indicator status and power readings under each power failure mode.

[0008] Further, the power failure modes include: fan power failure, fan jamming failure, SDA communication line fault, SCL communication line fault, main power overvoltage protection failure, standby power overvoltage protection failure, main power undervoltage protection failure, standby power undervoltage protection failure, main power overcurrent protection failure, standby power overcurrent protection failure, primary side overtemperature protection failure, secondary side overtemperature protection failure, and power inlet over temperature protection failure.

[0009] Further, an electronic switch is provided on the circuit corresponding to each power failure mode, specifically including:

[0010] 1) For fan power failure, a first electronic switch is provided on the fan power supply line;

[0011] 2) For fan jamming failure, a second electronic switch is provided on the fan reading value transmission line;

[0012] 3) For SDA communication line fault, a third electronic switch is provided on the SDA signal transmission line;

[0013] 4) For SCL communication line fault, a fourth electronic switch is provided on the SCL signal transmission line;

[0014] 5) For main power overvoltage protection failure, a fifth electronic switch is provided on the main power output feedback circuit. When the fifth electronic switch is closed, the feedback resistor is short-circuited;

[0015] 6) For standby power overvoltage protection failure, a sixth electronic switch is provided on the standby power output feedback circuit. When the sixth electronic switch is closed, the feedback resistor is short-circuited;

[0016] 7) For main power undervoltage protection failure, a seventh electronic switch is provided on the main power output feedback circuit. The seventh electronic switch is connected in parallel with a first test resistor. When the seventh electronic switch is closed, the feedback resistor is connected in parallel with the first test resistor;

[0017] 8) For standby power undervoltage protection failure, an eighth electronic switch is provided on the standby power output feedback circuit. The eighth electronic switch is connected in parallel with a second test resistor. When the eighth electronic switch is closed, the feedback resistor is connected in parallel with the second test resistor;

[0018] 9) For main power overcurrent protection failure, a ninth electronic switch is provided on the main power overcurrent protection feedback circuit. When the ninth electronic switch is closed, the reference power supply is input into the main power overcurrent protection feedback circuit;

[0019] 10) For standby power overcurrent protection failure, a tenth electronic switch is provided on the standby power overcurrent protection feedback circuit. When the tenth electronic switch is closed, the reference power supply is input into the standby power overcurrent protection feedback circuit;

[0020] 11) For the failure of the primary side overtemperature protection, an eleventh electronic switch is set in parallel with the primary side temperature sensor, and when the eleventh electronic switch is closed, the primary side temperature sensor is short-circuited;

[0021] 12) For the failure of the secondary side overtemperature protection, a twelfth electronic switch is set in the secondary side temperature detection circuit, and when the twelfth electronic switch is closed, the secondary side temperature sensor is short-circuited;

[0022] 13) For the failure of the power inlet air over-temperature protection, a thirteenth electronic switch is set in the power inlet air temperature detection circuit, and when the thirteenth electronic switch is closed, the inlet air temperature sensor is short-circuited.

[0023] Furthermore, the analog controller uses a secondary side MCU. When the secondary side MCU detects a power failure, it sends an alarm signal to the BMC.

[0024] In a second aspect, the technical solution of the present invention provides a method for simulating server power failure, including the following steps:

[0025] Pre-define the PSU PMBus1.2 instruction set, and each instruction in the instruction set corresponds to a power failure mode;

[0026] Execute each instruction of the PSU PMBus1.2 instruction set in sequence to simulate the corresponding power failure mode;

[0027] Obtain the power indicator status and power reading values under each power failure mode;

[0028] Save and record the power failure mode and its corresponding power indicator status and power reading values.

[0029] Furthermore, the power failure modes include: fan power failure, fan locked failure, main power overvoltage protection failure, standby power overvoltage protection failure, main power undervoltage protection failure, standby power undervoltage protection failure, main power overcurrent protection failure, standby power overcurrent protection failure, power inlet air over-temperature protection failure, and power over-temperature protection.

[0030] Furthermore, the method specifically includes the following steps:

[0031] Detect the status value of the PSU PMBus instruction address D1h Bit 0;

[0032] When the status value of D1h Bit 0 is 1, enter the power failure simulation program and execute each instruction of the PSU PMBus1.2 instruction set in sequence;

[0033] When the status value of D1h Bit 0 is 0, forcefully jump out of the power failure simulation program;

[0034] Detect the status value of bit 0 of the PSU PMBus instruction address D2h;

[0035] When the status value of bit 0 of D2h is 1, it indicates that the power failure simulation program has been successfully completed;

[0036] When the status value of bit 0 of D2h is 0, it indicates that the power failure simulation program has not been successfully completed.

[0037] Furthermore, the method further includes the following steps:

[0038] When a power failure is detected, send an alarm message to the BMC.

[0039] In a third aspect, the technical solution of the present invention provides a terminal, including:

[0040] A memory for storing a server power failure simulation program;

[0041] A processor for implementing the steps of the server power failure simulation method as described in any one of the above when executing the server power failure simulation program.

[0042] In a fourth aspect, the technical solution of the present invention provides a computer-readable storage medium, on which a server power failure simulation program is stored, and when the server power failure simulation program is executed by a processor, it implements the steps of the server power failure simulation method as described in any one of the above.

[0043] A server power failure simulation device, method, terminal and storage medium provided by the present invention have the following beneficial effects compared with the prior art: Simulate multiple power failure modes during the test stage, obtain the power indicator status and power reading values under each power failure mode. When a power failure occurs during the actual operation of the server, the operation and maintenance personnel can correspond to the corresponding power failure mode according to the power indicator status and power reading values, locate the cause of the failure, optimize the troubleshooting process, accelerate the server troubleshooting time, and enable the server to return to the line earlier or detect offline problems of large-scale system vulnerabilities earlier. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1It is a schematic structural diagram of a server power failure simulation device provided in Embodiment 1 of the present invention.

[0046] Figure 2 It is a schematic structural diagram of the electronic switch setting for fan power failure.

[0047] Figure 3 It is a schematic structural diagram of the electronic switch setting for fan jamming failure.

[0048] Figure 4 It is a schematic structural diagram of the electronic switch setting for SDA communication line failure / SCL communication line failure.

[0049] Figure 5 It is a schematic structural diagram of the electronic switch setting for main power overvoltage protection failure.

[0050] Figure 6 It is a schematic structural diagram of the electronic switch setting for standby power overvoltage protection failure.

[0051] Figure 7 It is a schematic structural diagram of the electronic switch setting for main power undervoltage protection failure.

[0052] Figure 8 It is a schematic structural diagram of the electronic switch setting for standby power undervoltage protection failure.

[0053] Figure 9 It is a schematic structural diagram of the electronic switch setting for main power overcurrent protection failure.

[0054] Figure 10 It is a schematic structural diagram of the electronic switch setting for standby power overcurrent protection failure.

[0055] Figure 11 It is a schematic structural diagram of the electronic switch setting for primary side overtemperature protection failure.

[0056] Figure 12 It is a schematic structural diagram of the electronic switch setting for secondary side overtemperature protection failure.

[0057] Figure 13 It is a schematic structural diagram of the electronic switch setting for power inlet air overtemperature protection failure.

[0058] Figure 14 It is a schematic diagram of the power indicator light status on the GUI interface.

[0059] Figure 15 It is a schematic diagram of the power reading structure on the GUI interface.

[0060] Figure 16 It is a schematic flowchart of a server power failure simulation method provided in Embodiment 2 of the present invention.

[0061] Figure 17It is a block diagram of the architecture for converting the analog signal feedback from the sensor to the digital signal by the MCU.

[0062] Figure 18 It is the PMBus 1.2 instruction set.

[0063] Figure 19 It is the Server PSU LED signal table.

[0064] Figure 20 It is a schematic structural diagram of a terminal provided in the third embodiment of the present invention. Detailed implementation manners

[0065] The following explains some terms related to the present invention.

[0066] Server PSU (Server Power Supply Unit): Server power supply, or simply server power.

[0067] IPMI (Intelligent Platform Management Interface): An industrial standard adopted by the peripheral devices of an Intel architecture enterprise system.

[0068] CRPS (Common Redundant Power Supply): A common power supply specification led by Intel.

[0069] I2C (Inter-Integrated Circuit): A serial communication bus with a complete communication protocol.

[0070] PMBus 1.2 (Power Management Bus 1.2): Power management bus version 1.2, a set of common power interface methods and languages based on the I2C communication protocol.

[0071] EEPROM (Electrically-Erasable Programmable Read-Only Memory): Electrically erasable programmable read-only memory, a semiconductor storage device that can be rewritten electronically multiple times.

[0072] BMC (Baseboard Management Controller): Baseboard management controller, a small dedicated processor for remote monitoring and management of the host system.

[0073] MCU (Microcontroller Unit): Microprocessor.

[0074] RCA (Root Cause Analysis): Root Cause Analysis

[0075] Failure Mode & Effect Analysis (FMEA): Failure Mode & Effect Analysis

[0076] Design & Quality verification Stage (DQ Stage): Design & Quality verification Stage

[0077] QA: Quality Analysis, Quality Analysis

[0078] OVP: Over Voltage Protection, Over Voltage Protection

[0079] UVP: Under Voltage Protection, Under Voltage Protection

[0080] OCP: Over Current Protection, Over Current Protection

[0081] OTP: Over Temperature Protection, Over Temperature Protection

[0082] GUI (Graphical User Interface): Graphical computer operation user interface

[0083] In order to enable those skilled in the art to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0084] Currently, the reliability requirements for Server PSUs are quite high, and current designs can all provide a stable and continuously operating server power supply. However, there is still a possibility of failure for Server PSUs. For the failure modes (DFMEA) of the server system during a failure, often the root cause analysis (RCA) is only carried out when it actually occurs on-site, which is time-consuming and laborious. It is not until the Server PSU has a function abnormality alarm or the server suddenly goes offline that the server maintenance personnel will be reminded to go to the site to troubleshoot the problem and the losses caused by the system offline. The core of the present invention is to address the deficiencies of the prior art and provide a server power failure simulation solution. During the test phase, the power failure modes are simulated through both hardware and software methods, and the corresponding power supply information is recorded under various power failure modes. When a power failure occurs during the actual operation of the server, the corresponding power failure mode can be determined based on the power supply information, enabling rapid location of the power failure. Also, the power supply information can be monitored in real time to promptly detect problems and prevent the server from suddenly going offline.

[0085] Embodiment 1

[0086] A server power failure simulation device provided in Embodiment 1 of the present invention realizes the simulation of power failure modes through hardware means.

[0087] Pre-define the server power failure modes, and set the corresponding hardware implementation circuits according to the power failure modes. The hardware implementation method in this embodiment mainly configures electronic switches on the corresponding circuits to cut off or short-circuit the corresponding circuits, etc., to realize the simulation of power failure modes.

[0088] Figure 1 is a schematic structural diagram of a server power failure simulation device provided in Embodiment 1 of the present invention, as Figure 1 shown. The device includes a simulation controller and several electronic switches. Pre-define several power failure modes, and set an electronic switch on the circuit corresponding to each power failure mode; the simulation controller is electrically connected to each electronic switch respectively, triggers each electronic switch to close in sequence, simulates the corresponding power failure mode, and obtains and records the power indicator status and power readings under each power failure mode.

[0089] During the test phase, after triggering the electronic switch to close, the corresponding power failure mode is simulated. At this time, there will be corresponding power indicator status and power readings. Record this power supply information, correspond it to the corresponding power failure mode, and save the correspondence between the power failure mode and the power supply information. When the server is running, if a power failure alarm occurs, the corresponding power failure mode can be determined based on the current power supply information, thereby locating the power failure; or when a certain power supply information corresponding to the corresponding power failure mode appears, promptly check the power supply to prevent the server from suddenly going offline.

[0090] The main power failure modes of the server are shown in Table 1 below, including: fan power failure, fan jamming failure, SDA communication line fault, SCL communication line fault, main power overvoltage protection failure, standby power overvoltage protection failure, main power undervoltage protection failure, standby power undervoltage protection failure, main power overcurrent protection failure, standby power overcurrent protection failure, primary side overtemperature protection failure, secondary side overtemperature protection failure, and power inlet over-temperature protection failure.

[0091] Table 1: Main Power Failure Modes of the Server

[0092] project Expiration Name Failure modes and characteristics of server power supplies 1 Fan Power Off The server power supply fan power is cut off, causing the fan to stop. At this time, the power supply MCU cannot read the fan speed. 2 Fan Tachline Failure The server power supply fan has stopped due to a stuck or disconnected fan control cable. The power supply MCU cannot read the fan speed. 3 I2C SDA communication line failure (PSU_I2C_SDA line fail) The system cannot read the server power information or the server power delivery information. 4 I2C SCL communication line failure (PSU_I2C_SCL line fail) The system cannot read the server power information or the server power delivery information. 5 Main power overvoltage protection (Main OutputOVP) The main output voltage of the server power supply is continuously shut down due to overvoltage protection, resulting in no output voltage. 6 Standby Output OVP The server auxiliary power supply output voltage is continuously shut down due to overvoltage protection, and the auxiliary power supply has no output voltage. 7 Main power undervoltage protection (Main OutputUVP) This fault causes the +12VDC to be shut down in latching mode. 8 UVP +12Vsb This fault causes +12V and +12Vsb to shut down in retry hiccup mode. 9 OCP +12V This fault turns off the +12V output. 10 OCP +12VSB This fault causes the power supply to shut down in retry mode. 11 Primary side OTP This fault shuts down the PSU output 12 Secondary side OTP This fault shuts down the PSU output 13 Power inlet OTP This fault shuts down the PSU output

[0093] The current Server PSU will have the MCU complete functions such as converter switch control, fan control, LED control, monitoring, protection, and communication in the power supply... The division of labor will be divided into the PRIMARY Side MCU and the SECONDARY Side MCU. In some specific embodiments, the secondary side MCU performs analog judgment and outputs alarm information to the system BMC. In terms of hardware, an electronic switch is set on the circuit corresponding to each power failure mode.

[0094] The following describes the setting of the electronic switch for various power failure modes.

[0095] 1) Fan power failure

[0096] Figure 2 It is a schematic diagram of the electronic switch setting for fan power failure. A first electronic switch is set on the fan power supply line. The first electronic switch simulates the interruption of the fan power supply and disconnects the fan power. It simulates the failure of the fan function caused by the breakage of the power cord of the fan power supply.

[0097] 2) Fan jamming failure

[0098] Figure 3 It is a schematic diagram of the electronic switch setting for fan jamming failure. A second electronic switch is set on the fan reading transmission line. The second electronic switch simulates no reading value of the fan and disconnects the fan reading RPM transmission line. It simulates that the fan speed reading value cannot be fed back to the secondary side MCU for some unknown reason, so the correct fan reading value cannot be read, resulting in no reading value or misjudgment.

[0099] 3) SDA communication line fault / SCL communication line fault

[0100] Figure 4It is a schematic diagram of the electronic switch setting structure for SDA communication line failure / SCL communication line failure. A third electronic switch is set on the SDA signal transmission line, and a fourth electronic switch is set on the SCL signal transmission line. The third electronic switch and the fourth electronic switch simulate the I2C communication transmission failure and disconnect the SDA and SCL transmission lines. Since the I2C communication cannot feed back to the secondary side MCU due to the disconnection of the SDA and SCL transmission lines, the correct read values of each PSU cannot be read, resulting in no read value or misjudgment.

[0101] 4) Failure of main power overvoltage protection

[0102] Figure 5 It is a schematic diagram of the electronic switch setting structure for the failure of main power overvoltage protection. A fifth electronic switch is set on the main power output feedback circuit. When the fifth electronic switch is closed, the feedback resistor is short-circuited. The fifth electronic switch simulates the overvoltage protection of the main power output (+12V) and short-circuits the output voltage feedback resistor. Simulating the main output feedback resistor, the resistance value changes due to long-term operation, resulting in too high an output voltage and causing overprotection of the power output voltage.

[0103] 5) Failure of standby power overvoltage protection

[0104] Figure 6 It is a schematic diagram of the electronic switch setting structure for the failure of standby power overvoltage protection. A sixth electronic switch is set on the standby power output feedback circuit. When the sixth electronic switch is closed, the feedback resistor is short-circuited. The sixth electronic switch simulates the overvoltage protection of the standby power output (+12Vsb) and short-circuits the standby power output voltage feedback resistor. Simulating the standby power feedback resistor, the resistance value changes due to long-term operation, resulting in too high an output voltage of the standby power and causing overprotection of the standby power output voltage.

[0105] 6) Failure of main power undervoltage protection

[0106] Figure 7 It is a schematic diagram of the electronic switch setting structure for the failure of main power undervoltage protection. A seventh electronic switch is set on the main power output feedback circuit. The seventh electronic switch is connected in parallel with a first test resistor. When the seventh electronic switch is closed, the feedback resistor is connected in parallel with the first test resistor.

[0107] The seventh electronic switch simulates that the main output (+12V) voltage is too low. A new resistance value is paralleled to the output voltage feedback resistor, simulating that the feedback resistor value changes due to long-term operation, resulting in too low a main output voltage and generating protection.

[0108] 7) Failure of standby power undervoltage protection

[0109] Figure 8It is a schematic diagram of the electronic switch setting structure for the failure of the standby power supply under-voltage protection. An eighth electronic switch is set on the standby power supply output feedback circuit. The eighth electronic switch is connected in parallel with a second test resistor. When the eighth electronic switch is closed, the feedback resistor is connected in parallel with the second test resistor.

[0110] The eighth electronic switch simulates the situation where the standby power supply (+12Vsb) voltage is too low. A new resistance value is paralleled to the standby power supply voltage feedback resistor to simulate the change in the feedback resistor value caused by long-term use, resulting in the main standby power supply voltage being too low and triggering protection.

[0111] 8) Failure of the main power supply over-current protection

[0112] Figure 9 It is a schematic diagram of the electronic switch setting structure for the failure of the main power supply over-current protection. A ninth electronic switch is set on the main power supply over-current protection feedback circuit. When the ninth electronic switch is closed, the reference power supply is input into the main power supply over-current protection feedback circuit.

[0113] The ninth electronic switch simulates the over-current protection of the main output (+12V). By inputting the reference power supply into the main output over-current protection feedback line, a feedback voltage value sufficient to achieve over-current protection is generated to achieve over-current protection.

[0114] 9) Failure of the standby power supply over-current protection

[0115] Figure 10 It is a schematic diagram of the electronic switch setting structure for the failure of the standby power supply over-current protection. A tenth electronic switch is set on the standby power supply over-current protection feedback circuit. When the tenth electronic switch is closed, the reference power supply is input into the standby power supply over-current protection feedback circuit.

[0116] The tenth electronic switch simulates the over-current protection of the standby power supply (+12Vsb). By inputting the reference power supply into the standby power supply over-current protection feedback line, a feedback voltage value sufficient to achieve over-current protection is generated to achieve over-current protection.

[0117] 10) Failure of the primary side over-temperature protection

[0118] Figure 11 It is a schematic diagram of the electronic switch setting structure for the failure of the primary side over-temperature protection. An eleventh electronic switch is set in parallel with the primary side temperature sensor. When the eleventh electronic switch is closed, the primary side temperature sensor is short-circuited.

[0119] The eleventh electronic switch simulates the primary side over-temperature protection. The primary side temperature sensor is short-circuited. It simulates the failure of the primary side temperature sensor due to long-term operation, thereby causing the primary side over-temperature protection.

[0120] 11) Failure of the secondary side over-temperature protection

[0121] Figure 12It is a schematic diagram of the electronic switch setting structure for the failure of the secondary-side over-temperature protection. An eleventh electronic switch is set on the secondary-side temperature detection circuit. When the eleventh electronic switch is closed, the secondary-side temperature sensor is short-circuited.

[0122] The twelfth electronic switch simulates the secondary-side over-temperature protection. It short-circuits the secondary-side temperature sensor, simulating the failure of the secondary-side temperature sensor due to long-term operation, thereby causing the secondary-side over-temperature protection to fail.

[0123] 12) Failure of the over-temperature protection at the power inlet

[0124] Figure 13 It is a schematic diagram of the electronic switch setting structure for the failure of the over-temperature protection at the power inlet. A thirteenth electronic switch is set on the temperature detection circuit at the power inlet. When the thirteenth electronic switch is closed, the temperature sensor at the inlet is short-circuited.

[0125] The thirteenth electronic switch simulates the over-temperature protection at the power inlet. It short-circuits the temperature sensor at the inlet, simulating the failure of the temperature sensor at the inlet due to long-term operation, thereby causing the over-temperature protection at the inlet to fail.

[0126] The above is the simulation of the power failure mode in a hardware manner. When the trigger electronic switch is closed, it will cause a power failure, and the corresponding simulation controller will obtain and record the status of the power indicator and the power reading under each power failure mode.

[0127] The following explains the status of the power indicator and the power reading under the power failure mode.

[0128] In the test stage, a GUI interface is used to display and observe the power phenomenon. The GUI is a graphical operation interface programmed through LabVIEW software. The status of the power indicator and the power reading are displayed on the GUI interface. Figure 14 It is a schematic diagram of the status of the power indicator on the GUI interface. Figure 15 It is a schematic diagram of the power reading structure on the GUI interface.

[0129] Under normal power conditions, all indicators are green, and the power reading also has corresponding normal values. When a power failure occurs, the corresponding indicator will turn red and / or the power reading will change.

[0130] For example, when the fan power supply fails, the red light "Fan Fail" is displayed. Both the 79h PSU status and the 81h PSU Fan status in the PMBus 1.2 specification are triggered. At this time, the power readings are all normal, but there is no rotational speed information for the fan. When the fan is stuck and fails, the red light "Fan Fail" is displayed. Both the 79h PSU status and the 81h PSU Fan status in the PMBus 1.2 specification are triggered. At this time, the power readings are all normal, but there is no rotational speed information for the fan. When the SDA communication line fails, the outputs are all normal. However, there are no readings in each status column. At this time, communication anomalies occur. When the SCL communication line fails, the outputs are all normal at this time. However, there are no readings in each status column. At this time, communication anomalies occur. When the main power overvoltage protection fails, the GUI immediately displays the red light "Vout Fail". Both the 79h PSU status and the 7Ah PSU Vout status in the PMBus 1.2 specification are triggered. At this time, the power output is shut down due to protection, and the readings of the output voltage, current, and power consumption are all 0. When the standby power overvoltage protection fails, since the auxiliary power supply also supports the PSU internal MCU power supply, it is also shut down. At this time, the PSU is in a down state, the GUI shows a disconnection, and there are no readings in each status column. When the main power undervoltage protection fails, the GUI immediately displays the red light "Vout Fail". Both the 79h PSU status and the 7Ah PSU Vout status in the PMBus 1.2 specification are triggered. At this time, the power output is shut down due to protection, and the readings of the output voltage, current, and power consumption are all 0. When the standby power undervoltage protection fails, since the auxiliary power supply also supports the PSU internal MCU power supply, it is also shut down. At this time, the PSU is in a down state, the GUI shows a disconnection, and there are no readings in each status column. When the main power overcurrent protection fails, the GUI immediately displays the red light "Iout Fail". Both the 79h PSU status and the 7Bh PSU Iout status in the PMBus 1.2 specification are triggered. At this time, the power output is shut down due to protection, and the readings of the output voltage, current, and power consumption are all 0. When the standby power overcurrent protection fails, since the auxiliary power supply also supports the PSU internal MCU power supply, it is also shut down. At this time, the PSU is in a down state, the GUI shows a disconnection, and there are no readings in each status column. When the primary side overtemperature protection fails, the GUI immediately displays the red light "TEMP Fail". Both the 79h PSU status and the 7Dh PSU TEMP status in the PMBus 1.2 specification are triggered. At this time, the power output is shut down due to protection, and the readings of the output voltage, current, and power consumption are all 0, and the 8EH reading is as high as 125 degrees. When the secondary side overtemperature protection fails, the GUI immediately displays the red light "TEMP Fail". Both the 79h PSU status and the 7Dh PSU TEMP status in the PMBus 1.2 specification are triggered. At this time, the power output is shut down due to protection, and the readings of the output voltage, current, and power consumption are all 0, and the 8FH reading is as high as 125 degrees. When the power inlet air overtemperature protection fails, the GUI immediately displays the red light "TEMP Fail".Both the 79h PSU status and 7Dh PSU TEMP status in the PMBus 1.2 specification are triggered. At this time, the power output is shut down due to protection, and the read values of output voltage, current, and power consumption are all 0, and the 8DH read value is as high as 90 degrees.

[0131] Embodiment 2

[0132] In the above Embodiment 1, the simulation of the power failure mode is realized by hardware. Embodiment 2 provides a method for simulating the power failure of a server power supply, which realizes the simulation of the power failure mode by software.

[0133] Figure 16 It is a schematic flowchart of a method for simulating the power failure of a server power supply provided in Embodiment 2 of the present invention. As Figure 16 shown, the method includes the following steps.

[0134] S1, Pre-define the PSU PMBus1.2 instruction set, and each instruction in the instruction set corresponds to a power failure mode.

[0135] S2, Sequentially execute each instruction in the PSU PMBus1.2 instruction set to simulate the corresponding power failure mode.

[0136] S3, Obtain the power indicator status and power read values in each power failure mode.

[0137] S4, Save and record the power failure mode and its corresponding power indicator status and power read values.

[0138] The method for simulating the power failure of a server power supply provided in this embodiment configures instructions for various power failure modes in the PSU PMBus1.2 instruction set, and executes the corresponding instructions to realize the simulation of the power failure mode.

[0139] To further understand the present invention, the following further details this embodiment.

[0140] Figure 17It is a block diagram of the architecture for converting the analog signal feedback from the sensor to the digital signal by the MCU. The power supply needs to read the feedback value content, which can be classified into voltage detection, current detection, temperature detection, power detection, fan detection, etc. The server uses BMC to access the server power supply MCU through the I2C Bus to obtain the readings of various sensors. The internal sensor hardware circuit of the server power supply obtains the analog voltage signal of the required detection source through a "sensor unit" composed of a shunt resistor, a differential amplifier, and a setting resistor, and returns it to the digital-to-analog converter (ADC) of the internal MCU of the Server PSU, and converts it into a digital signal for calculation and judgment inside the MCU. Generally, 5-10 groups of "sensor units" are required inside the power supply to detect and protect the signals required for the entire PSU power management.

[0141] Usually, there are high and low standards for the detection values set in the power supply firmware, and protection will be triggered if they are too high or too low. Therefore, a self-simulation fault mode is set for the analog server power failure mode module. After the power supply enters the analog server power failure mode, various types of failure modes can be simulated using instructions. Figure 18 It is the PMBus1.2 instruction set. First, in the PMBus1.2 specification, the self-learn function is defined. In the PMBus1.2 instruction set, the D1h–D3h instruction sets are reserved for function expansion. We can use the D1h address as the function definition and expansion of the software-simulated power failure mode.

[0142] In this embodiment, D1h is to enter the analog server power failure mode. When the PSU PMBus instruction address D1h Bit0 is written with 1, it starts to enter the analog server power failure mode. Writing 0 to the address D1h Bit 1 can forcefully jump out of the analog server power failure mode. When entering the analog server power failure mode, the LED signal behavior at this time will be the same as Figure 19 the Server PSU LED signal table. The output remains normal in this case, but software mode failure simulation is performed. And each failure mode instruction can be executed step by step through the PMBUS instruction as follows:

[0143] (1) 00000 000: Fan power failure.

[0144] (2) 00000 001: No reading of fan speed.

[0145] (3) 00000 010: Main output overvoltage protection (OVP).

[0146] (4) 00000 011: Auxiliary power overvoltage protection (OVP).

[0147] (5) 00000 100: Main output undervoltage protection (UVP).

[0148] (6) 00000 101: Main output overcurrent protection (OCP).

[0149] (7) 00000 110: Inlet air over-temperature protection (OTP).

[0150] (8) 00000 111: Power supply over-temperature protection (OTP).

[0151] It can be seen that the power failure modes that can be achieved in this embodiment include: fan power failure, fan jamming failure, main power overvoltage protection failure, standby power overvoltage protection failure, main power undervoltage protection failure, standby power undervoltage protection failure, main power overcurrent protection failure, standby power overcurrent protection failure, power supply inlet air over-temperature protection failure, and power supply over-temperature protection.

[0152] Define D2h as the status feedback of the analog server power failure mode. When the PSU PMBus instruction address D2h Bit 0 reports 1, it represents that the analog server power failure mode is completed. When the address D2h Bit 0 feedbacks 0, it represents that the analog server power failure mode is not successful. It can be used for the system to judge whether to perform the analog test again. When the PSU PMBus instruction address D2h Bit 1 feedbacks 1, it represents entering the analog server power failure mode. When the address D2h Bit 1 feedbacks 0, it represents jumping out of the analog server power failure mode.

[0153] Correspondingly, the server power failure simulation method of this embodiment specifically includes the following steps:

[0154] Step 1, detect the status value of the PSU PMBus instruction address D1h Bit 0;

[0155] Step 2, when the status value of D1h Bit 0 is 1, enter the power failure simulation program and execute each instruction of the PSU PMBus1.2 instruction set in sequence;

[0156] Step 3, when the status value of D1h Bit 0 is 0, forcefully jump out of the power failure simulation program;

[0157] Step 4, detect the status value of the PSU PMBus instruction address D2h Bit 0;

[0158] Step 5, when the status value of D2h Bit 0 is 1, indicate that the power failure simulation program is successfully completed;

[0159] Step 6, when the status value of D2h Bit 0 is 0, it indicates that the power failure simulation program has not been successfully completed.

[0160] In addition, in this embodiment, when a power failure is detected, an alarm message is also sent to the BMC.

[0161] This embodiment obtains the power indicator status and power reading values in each power failure mode, and saves and records the power failure mode and its corresponding power indicator status and power reading values.

[0162] Embodiment 3

[0163] Figure 20 The following is a schematic structural diagram of a terminal device 2000 provided by an embodiment of the present invention, including: a processor 2010, a memory 2020, and a communication unit 2030. When the processor 2010 is used to implement the server power failure simulation program saved in the memory 2020, the following steps are implemented:

[0164] S1, predefined PSU PMBus1.2 instruction set, and each instruction in the instruction set corresponds to a power failure mode;

[0165] S2, sequentially execute each instruction of the PSU PMBus1.2 instruction set to simulate the corresponding power failure mode;

[0166] S3, obtain the power indicator status and power reading values in each power failure mode;

[0167] S4, save and record the power failure mode and its corresponding power indicator status and power reading values.

[0168] The present invention simulates multiple power failure modes during the test phase, obtains the power indicator status and power reading values in each power failure mode. When a power failure occurs during the actual operation of the server, the operation and maintenance personnel can correspond to the corresponding power failure mode according to the power indicator status and power reading values, locate the cause of the failure, optimize the troubleshooting process, accelerate the server troubleshooting time, and let the server go back online earlier or discover the offline problem of large-scale system vulnerabilities earlier.

[0169] The terminal device 2000 includes a processor 2010, a memory 2020, and a communication unit 2030. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and can also include more or fewer components than shown in the figure, or combine some components, or different component arrangements.

[0170] Among them, the memory 2020 can be used to store the execution instructions of the processor 2010. The memory 2020 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. When the execution instructions in the memory 2020 are executed by the processor 2010, the terminal 2000 can execute some or all of the steps in the above method embodiments.

[0171] The processor 2010 is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 2020, and calling the data stored in the memory, it executes various functions of the electronic terminal and / or processes data. The processor can be composed of an integrated circuit (IC). For example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 2010 may only include a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single arithmetic core or can include multiple arithmetic cores.

[0172] The communication unit 2030 is used to establish a communication channel so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.

[0173] Embodiment 4

[0174] The present invention also provides a computer storage medium. The storage medium here can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), etc.

[0175] The computer storage medium stores a server power failure simulation program. When the server power failure simulation program is executed by the processor, the following steps are implemented:

[0176] S1, predefined PSU PMBus1.2 instruction set, and each instruction in the instruction set corresponds to a power failure mode;

[0177] S2, sequentially execute each instruction of the PSU PMBus1.2 instruction set to simulate the corresponding power failure mode;

[0178] S3. Obtain the power indicator status and power reading under each power failure mode;

[0179] S4. Save and record the power failure mode and its corresponding power indicator status and power reading.

[0180] The present invention realizes the simulation of multiple power failure modes during the test stage, obtains the power indicator status and power reading under each power failure mode. When a power failure occurs during the actual operation of the server, the operation and maintenance personnel can correspond to the corresponding power failure mode according to the power indicator status and power reading, locate the cause of the failure, optimize the troubleshooting process, accelerate the server troubleshooting time, and enable the server to return to the line earlier or detect the offline problem of large-scale system vulnerabilities earlier.

[0181] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disc, etc., which can store program codes, including several instructions to enable a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention.

[0182] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0183] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0184] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.

[0185] The above disclosure is only the preferred embodiment of the present invention, but the present invention is not limited thereto. Any non-creative changes that can be thought of by those skilled in the art, as well as several improvements and refinements made without departing from the principle of the present invention, should fall within the protection scope of the present invention.

Claims

1. A server power failure simulation method based on a server power failure simulation device, characterized in that, The device includes: An analog controller and several electronic switches; Pre - define several power failure modes, and set an electronic switch on the circuit corresponding to each power failure mode; The analog controller is electrically connected to each electronic switch respectively, triggers each electronic switch to close in sequence, simulates the corresponding power failure mode, and obtains and records the power indicator status and power reading under each power failure mode; The power failure modes include: fan power failure, fan stuck failure, SDA communication line fault, SCL communication line fault, main power over - voltage protection failure, standby power over - voltage protection failure, main power under - voltage protection failure, standby power under - voltage protection failure, main power over - current protection failure, standby power over - current protection failure, primary - side over - temperature protection failure, secondary - side over - temperature protection failure, and power inlet over - temperature protection failure; Setting an electronic switch on the circuit corresponding to each power failure mode specifically includes: For fan power failure, set a first electronic switch on the fan power supply line; For fan stuck failure, set a second electronic switch on the fan reading value transmission line; For SDA communication line fault, set a third electronic switch on the SDA signal transmission line; For SCL communication line fault, set a fourth electronic switch on the SCL signal transmission line; For main power over - voltage protection failure, set a fifth electronic switch on the main power output feedback circuit. When the fifth electronic switch closes, the feedback resistor is short - circuited; For standby power over - voltage protection failure, set a sixth electronic switch on the standby power output feedback circuit. When the sixth electronic switch closes, the feedback resistor is short - circuited; For main power under - voltage protection failure, set a seventh electronic switch on the main power output feedback circuit. The seventh electronic switch is in parallel with a first test resistor. When the seventh electronic switch closes, the feedback resistor is in parallel with the first test resistor; For standby power under - voltage protection failure, set an eighth electronic switch on the standby power output feedback circuit. The eighth electronic switch is in parallel with a second test resistor. When the eighth electronic switch closes, the feedback resistor is in parallel with the second test resistor; For main power over - current protection failure, set a ninth electronic switch on the main power over - current protection feedback circuit. When the ninth electronic switch closes, the reference power supply is input into the main power over - current protection feedback circuit; For standby power over - current protection failure, set a tenth electronic switch on the standby power over - current protection feedback circuit. When the tenth electronic switch closes, the reference power supply is input into the standby power over - current protection feedback circuit; For primary - side over - temperature protection failure, set an eleventh electronic switch in parallel with the primary - side temperature sensor. When the eleventh electronic switch closes, the primary - side temperature sensor is short - circuited; For secondary - side over - temperature protection failure, set a twelfth electronic switch on the secondary - side temperature detection circuit. When the twelfth electronic switch closes, the secondary - side temperature sensor is short - circuited; For power inlet over - temperature protection failure, set a thirteenth electronic switch on the power inlet temperature detection circuit. When the thirteenth electronic switch closes, the inlet temperature sensor is short - circuited; The analog controller uses a secondary - side MCU. When the secondary - side MCU detects a power failure, it sends an alarm signal to the BMC; The method includes the following steps: Pre - define the PSU PMBus1.2 instruction set, where each instruction in the instruction set corresponds to a power failure mode; Execute each instruction of the PSU PMBus1.2 instruction set in sequence to simulate the corresponding power failure mode; Obtain the power indicator status and power reading values under each power failure mode; Save and record the power failure mode together with its corresponding power indicator status and power reading values.

2. The server power failure simulation method based on the server power failure simulation device according to claim 1, wherein, Specifically, the method includes the following steps: Detect the status value of the PSU PMBus instruction address D1h Bit 0; When the status value of D1h Bit 0 is 1, enter the power failure simulation program and execute each instruction of the PSU PMBus1.2 instruction set in sequence; When the status value of D1h Bit 0 is 0, forcefully jump out of the power failure simulation program; Detect the status value of the PSU PMBus instruction address D2h Bit 0; When the status value of D2h Bit 0 is 1, indicate that the power failure simulation program has been successfully completed; When the status value of D2h Bit 0 is 0, indicate that the power failure simulation program has not been successfully completed.

3. The server power failure simulation method based on the server power failure simulation device according to claim 2, characterized in that, The method further includes the following steps: When a power failure is detected, send an alarm message to the BMC.

4. A terminal, characterized in that, It includes: A memory for storing the server power failure simulation program based on the server power failure simulation device; A processor for implementing the steps of the server power failure simulation method based on the server power failure simulation device as described in any one of claims 1 - 3 when executing the server power failure simulation program based on the server power failure simulation device.

5. A computer-readable storage medium, characterized in that, The server power failure simulation program based on the server power failure simulation device is stored on the readable storage medium, and when the server power failure simulation program based on the server power failure simulation device is executed by the processor, it implements the steps of the server power failure simulation method based on the server power failure simulation device as described in any one of claims 1 - 3.

Citation Information

Patent Citations

  • Experiment tester for spare power automatic switching apparatus

    CN101504449A

  • Fault detection device of complete machine internal power supply module

    CN211426733U