A server power module state detection apparatus and method
Patent Information
- Application Number
- CN202310330277.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-03-30
AI Technical Summary
然而,实际上在客户现场难免会出现一种情况:当市电故障时,引发了大面积宕机现象
[0015]本发明提供的一种服务器电源模块状态检测装置及方法,相对于现有技术,具有以下有益效果:实时采集电源模块的状态信号和服务器待机供电信号,将所采集信号记录到日志供维护人员查看,维护人员可根据电源模块状态信号和服务器待机供电信号分析故障原因,为确定故障原因和后续维护提供数据支持,降低后续宕机故障概率。
Smart Images

Figure CN116301276B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power status detection, and more specifically to a server power module status detection device and method. Background Technology
[0002] Currently, servers in user sites generally use a dual-power supply mode, with each power module providing independent power. Typically, one power source is AC mains power, and the other is an uninterruptible power supply (UPS). When the AC mains power fails, the system switches to the UPS to ensure server operation and prevent downtime. Figure 1 This is a schematic diagram of the current server power supply structure, such as... Figure 1 As shown, the server motherboard is connected to a main power module and a backup power module, respectively. The main power module is powered by AC power, and the backup power module is powered by an uninterruptible power supply.
[0003] While theoretically, the system should automatically switch to another power source when one power source fails (e.g., switching to an uninterruptible power supply during a mains power outage), in reality, a situation inevitably arises at customer sites where a mains power failure triggers a widespread outage. This outage disrupts customer usage and leads to complaints. Theoretically, with dual power supplies in a redundant configuration, a power outage in one source shouldn't affect machine operation. However, after an outage occurs, it's difficult to determine whether the problem lies with the server's internal power switching or the customer's on-site power supply environment. This is because outages are instantaneous, and customers can resolve the issue by simply restarting the server. Currently, there's no effective record of information to analyze and determine the cause of the outage, which is detrimental to subsequent maintenance of the server itself and its power supply environment. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides a server power module status detection device and method, which provides data support for determining the cause of a fault and subsequent maintenance, thereby reducing the probability of subsequent system downtime.
[0005] In a first aspect, the technical solution of the present invention provides a server power module status detection device, comprising: a power module status collector and a detection manager; Power module status collector: Connects to the server motherboard and is configured to collect status signals of the backup power module and server standby power supply signals from the server motherboard and send them to the detection manager; Detection Manager: Connects to the power module status collector, configured to receive the backup power module status signal and the server standby power supply signal sent by the power module status collector, and record the backup power module status signal and the server standby power supply signal to generate a detection log.
[0006] In an optional implementation, the detection manager is also connected to the server motherboard and configured to collect server status signals and power switching signals. It determines whether the server has crashed based on the server status signals. If the server crashes when the main power module switches to the backup power module, it analyzes the cause of the failure based on the status signals of the backup power module and the server standby power supply signals, and records the cause of the failure in the log. The fault analysis is based on the status signals of the backup power module and the standby power supply signal of the server. Specifically, if the status signals of the backup power module and the standby power supply signal of the server are abnormal, it is determined that the power supply environment of the backup power module is abnormal; if the status signals of the backup power module and the standby power supply signal of the server are normal, it is determined that the reliability of the server redundancy logic is abnormal.
[0007] In an optional implementation, the power module status collector is also configured to collect status signals of the main power module from the server motherboard and send them to the detection manager. Correspondingly, the detection manager is also configured to receive the main power module status signal sent by the power module status collector and record the main power module status signal to the detection log.
[0008] In an optional implementation, the detection manager is also configured to analyze the cause of the failure based on the status signal of the main power module and the standby power supply signal of the server in response to a server crash when the backup power module switches to the main power module, and record the cause of the failure in the log. Specifically, the fault cause is analyzed based on the status signals of the main power module and the standby power supply signals of the server. If the status signals of the main power module and the standby power supply signals of the server are abnormal, it is determined that the power supply environment of the main power module is abnormal; if the status signals of the main power module and the standby power supply signals of the server are normal, it is determined that the reliability of the server's redundancy logic is abnormal.
[0009] In one optional implementation, the status signal of the power module includes an output voltage normal indication signal and a power control enable signal; in response to the output voltage normal indication signal and the power control enable signal being high, the status signal of the power module is determined to be normal; in response to the output voltage normal indication signal and the power control enable signal being low, the status signal of the power module is determined to be abnormal. If the server standby power supply signal is high, the server standby power supply signal is considered normal; if the server standby power supply signal is low, the server standby power supply signal is considered abnormal.
[0010] In one alternative implementation, the power module status acquisition device employs a programmable logic device.
[0011] In an optional implementation, the device further includes a main connector and a secondary connector located on the server motherboard; the power module status acquisition unit is connected to the secondary connector via the main connector.
[0012] In an optional embodiment, the device further includes a battery module connected to the power module status collector and the detection manager, respectively, to supply power to the power module status collector and the detection manager.
[0013] Secondly, the technical solution of the present invention provides a server power module status detection method, comprising the following steps: Receives status signals from the main power module, the backup power module, and the server standby power supply signal; The status signals of the main power module, the backup power module, and the server standby power supply signal are recorded in the log.
[0014] In an optional implementation, the method further includes the following steps: Receive server status signals and power switching signals; In response to server downtime during power switching, the cause of the failure is analyzed based on the power module status signals and the server standby power supply signals, and the cause of the failure is recorded in the log. Specifically, this includes: In response to a server crash when the main power module switches to the backup power module, if the status signal of the backup power module and the server standby power supply signal are abnormal, it is determined that the power supply environment of the backup power module is abnormal; if the status signal of the backup power module and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal. In response to a server crash when the backup power module switches to the main power module, if the main power module status signal and the server standby power supply signal are abnormal, it is determined that the main power module power supply environment is abnormal; if the main power status signal and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal.
[0015] The present invention provides a server power module status detection device and method, which has the following advantages over the prior art: real-time acquisition of power module status signals and server standby power supply signals, recording the acquired signals in a log for maintenance personnel to view, and maintenance personnel can analyze the cause of the fault based on the power module status signals and server standby power supply signals, providing data support for determining the cause of the fault and subsequent maintenance, and reducing the probability of subsequent downtime faults. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the current server power supply structure.
[0018] Figure 2 This is a schematic diagram of an application scenario for the server power module status detection device provided in this embodiment of the invention. Figure 1 .
[0019] Figure 3 This is a schematic diagram of an application scenario for the server power module status detection device provided in this embodiment of the invention. Figure 2 .
[0020] Figure 4 This is a schematic diagram of a server power module status detection device provided in an embodiment of the present invention.
[0021] Figure 5 This is a schematic diagram of a server power module status detection device provided in an embodiment of the present invention.
[0022] Figure 6 This is a schematic diagram of a specific embodiment of a server power module status detection device provided in this invention.
[0023] Figure 7 This is a schematic flowchart of a server power module status detection method provided in an embodiment of the present invention.
[0024] Figure 8 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0027] The key terms used in this invention will be explained below.
[0028] CPLD: Complex Programmable Logic Device.
[0029] PSU: Power supply unit, power module.
[0030] P12V_STBY: The AC power input of the server power supply refers to the 220V AC power input to the server power input terminal. This 220V AC power is connected to the server power input terminal via the power cord (AC line). Internally, the server power supply converts the 220V AC power into two DC power supplies: P12V_main and P12V_STBY. Both are 12V, but P12V_main provides a larger current and requires a power enable signal to be converted into the system main voltage for other chips or components. P12V_STBY provides a smaller current and automatically outputs its voltage after the AC line is plugged in, without a power enable signal. The 12V STBY is generally converted into P3V3_STBY or other lower voltages to power chips on the board that are powered up early, such as CPLDs. P12V_main is converted into the main voltage of other systems after the server is powered on, and used by chips or components on the board that only work after the power-on, such as CPU, memory, hard drive, and fan lights.
[0031] Figure 2 This is a schematic diagram of an application scenario for the server power module status detection device provided in this embodiment of the invention. Figure 1 In this application scenario, each server is equipped with a separate server power module status detection device. The server power module status detection device can be configured as a board inside the server. The server power module status detection device collects signals such as the power module status of the server it is located in and sends the generated logs to the management terminal separately. Maintenance personnel perform fault analysis based on the logs of each server.
[0032] Figure 3 This is a schematic diagram of an application scenario for the server power module status detection device provided in this embodiment of the invention. Figure 2In this second application scenario, a server power status detection device can collect signals such as the power module status of multiple servers. This device can operate within a separate computer device or as a standalone terminal. During information collection, each server is identified, and the collected power module status signals for each server are mapped to that server identifier to generate logs. The server power status detection device then sends all the generated logs of the power module status of all servers to the management terminal.
[0033] Figure 4 This is a schematic diagram of a server power module status detection device provided in an embodiment of the present invention, as shown below. Figure 4 As shown, the device includes a power module status acquisition unit and a detection manager.
[0034] The server motherboard is connected to a main power module and a backup power module. The main power module is powered by AC mains, while the backup power module is powered by an uninterruptible power supply (UPS). Under normal circumstances, the main power module powers the server. In the event of a mains power failure, power is switched to the backup power module powered by the UPS. If a server crashes during this switchover, it indicates an abnormality in the UPS power supply environment or a problem with the server's redundancy logic reliability.
[0035] Power module status collector: Connects to the server motherboard and is configured to collect status signals of the backup power module and server standby power supply signals from the server motherboard and send them to the detection manager.
[0036] Detection Manager: Connects to the power module status collector, configured to receive the backup power module status signal and the server standby power supply signal sent by the power module status collector, and record the backup power module status signal and the server standby power supply signal to generate a detection log.
[0037] If the backup power module's power supply environment is normal, meaning the uninterruptible power supply (UPS) is functioning correctly, both the backup power module's status signals and the server's standby power supply signals will be normal, and no abnormalities will occur. If a system crash occurs under these conditions, it indicates an issue with the server's redundancy logic reliability. Conversely, if the backup power module's power supply environment is abnormal, meaning the UPS is malfunctioning, both the backup power module's status signals and the server's standby power supply signals will exhibit abnormal states. A system crash in this case can be attributed to an abnormality in the backup power module's power supply environment. Maintenance personnel can view the backup power module's status signals and the server's standby power supply signals in the logs, and by analyzing their status, the cause of the fault can be determined.
[0038] In an optional implementation, the status signal of the backup power module may be an output voltage normal indication signal and a power control enable signal.
[0039] Under normal uninterruptible power supply (UPS) conditions, the normal output voltage indication signal and power control enable signal of the backup power module, as well as the server standby power supply signal, are all at a high level; under abnormal UPS conditions, the normal output voltage indication signal and power control enable signal of the backup power module, as well as the server standby power supply signal, are all at a low level.
[0040] Maintenance personnel checked the high and low level signals in the logs to determine the cause of the fault. When the output voltage normal indication signal, power control enable signal, and server standby power supply signal were all at high levels, it was determined that the power module's status signals and the server's standby power supply signal were normal, thus determining that the uninterruptible power supply (UPS) environment was normal, and the fault was caused by an abnormality in the server's redundancy logic reliability. When the output voltage normal indication signal, power control enable signal, and server standby power supply signal were all at low levels, it was determined that the power module's status signals and the server's standby power supply signal were abnormal, thus determining that the UPS environment was abnormal, and the fault was caused by an abnormal UPS environment.
[0041] In one alternative implementation, the power module status acquisition device employs a programmable logic device.
[0042] The device is also equipped with a main connector and a secondary connector on the server motherboard; the power module status acquisition unit connects to the secondary connector via the main connector. The server motherboard sends the backup power module status signal and the server standby power supply signal to the power module status acquisition unit through the connector.
[0043] In an optional implementation, the detection manager is also connected to the server motherboard, which can be wirelessly connected. It is configured to collect server status signals and power switching signals, determine whether the server has crashed based on the server status signals, and respond to server crashes when the main power module switches to the backup power module. It analyzes the cause of the fault based on the status signals of the backup power module and the server standby power supply signals, and records the cause of the fault in the log.
[0044] In this embodiment, the device automatically performs logical judgments to determine the cause of the fault and provides data to maintenance personnel. Specifically, it can determine whether the server has crashed based on the server status signal and whether a power switching operation has occurred based on the power switching signal. If the server crashes when the main power module switches to the backup power module, and the backup power module status signal and the server standby power supply signal are abnormal, it is determined that the backup power module power supply environment is abnormal; if the backup power status signal and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal.
[0045] In an optional embodiment, the device further includes a battery module connected to the power module status collector and the detection manager, respectively, to supply power to the power module status collector and the detection manager.
[0046] In the aforementioned application scenario, the battery module supplies power to the device independently, ensuring continued monitoring of the power module's status signals even in the event of a server outage. The battery module can be a rechargeable battery, connected to the main connector and charged by the server motherboard, thus drawing power from the motherboard. Alternatively, a power switch can be configured, with one input connected to the main connector and the other to the battery module; its outputs connected to the power module status acquisition unit and the detection manager; and its controller connected to the detection manager. When the server is functioning normally, the server motherboard supplies power to the device; when the server malfunctions, power is switched to the battery module.
[0047] Figure 5 This is a schematic diagram of a server power module status detection device provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the device includes a power module status acquisition unit and a detection manager.
[0048] Power module status collector: Connects to the server motherboard and is configured to collect status signals from the main power module, the backup power module, and the server standby power supply signal from the server motherboard, and send them to the detection manager.
[0049] Detection Manager: Connects to the power module status collector, configured to receive the main power module status signal, backup power module status signal, and server standby power supply signal sent by the power module status collector, and record the main power module status signal, backup power module status signal, and server standby power supply signal to generate a detection log.
[0050] The device in this embodiment simultaneously detects the status signals of the main power module, the backup power module, and the server standby power supply signal. When the main power module switches to the backup power module, the device analyzes the cause of the fault based on the status signals of the backup power module and the server standby power supply signal, as detailed in the above embodiment, and will not be repeated here. When the backup power module switches to the main power module, the device analyzes the cause of the fault based on the status signals of the main power module and the server standby power supply signal.
[0051] If the mains power supply environment is normal, meaning the AC power is normal, both the mains power module status signal and the server standby power supply signal will be normal, and no abnormalities will occur. If a system crash occurs under these conditions, it indicates an issue with the server's redundancy logic reliability. Conversely, if the mains power supply environment is abnormal, both the mains power module status signal and the server standby power supply signal will show abnormalities. A system crash in this case can be attributed to an abnormal mains power module power supply environment. Maintenance personnel can view the mains power module status signal and the server standby power supply signal in the logs and determine the cause of the fault based on their status.
[0052] In an optional implementation, the main power module status signal may be an output voltage normal indication signal and a power control enable signal.
[0053] Under normal mains power conditions, the main power module's output voltage normal indication signal, power control enable signal, and server standby power supply signal are all at a high level; under abnormal mains power conditions, the main power module's output voltage normal indication signal, power control enable signal, and server standby power supply signal are all at a low level.
[0054] Maintenance personnel checked the high and low level signals in the logs to determine the cause of the fault. When the output voltage normal indication signal, power control enable signal, and server standby power supply signal were all at high levels, it was determined that the power module's status signals and the server standby power supply signal were normal, thus determining that the mains power supply environment was normal, and the fault was caused by an abnormality in the reliability of the server's redundant logic. When the output voltage normal indication signal, power control enable signal, and server standby power supply signal were all at low levels, it was determined that the power module's status signals and the server standby power supply signal were abnormal, thus determining that the mains power supply environment was abnormal, and the fault was caused by an abnormality in the mains power supply environment.
[0055] In an optional implementation, the detection manager is also connected to the server motherboard, which can be wirelessly connected. It is configured to collect server status signals and power switching signals, determine whether the server has crashed based on the server status signals, and respond to server crashes when the main power module switches to the backup power module. It analyzes the cause of the fault based on the status signals of the backup power module and the server standby power supply signals, and records the cause of the fault in the log.
[0056] In this embodiment, the device automatically performs logical judgments to determine the cause of the fault and provides data to maintenance personnel. Specifically, it can determine whether the server has crashed based on the server status signal and whether a power switching operation has occurred based on the power switching signal. If the server crashes when the backup power module switches to the main power module, and the main power module status signal and the server standby power supply signal are abnormal, it is determined that the main power module power supply environment is abnormal; if the main power status signal and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal.
[0057] In an optional implementation, the detection manager is also connected to the server motherboard, which can be wirelessly connected. It is configured to collect server status signals and power switching signals, determine whether the server has crashed based on the server status signals, and respond to server crashes when the backup power module switches to the main power module. It analyzes the cause of the fault based on the main power module status signal and the server standby power supply signal, and records the cause of the fault in the log.
[0058] In this embodiment, the device automatically performs logical judgments to determine the cause of the fault and provides data to maintenance personnel. Specifically, it can determine whether the server has crashed based on the server status signal and whether a power switching operation has occurred based on the power switching signal. If the server crashes when the backup power module switches to the main power module, and the main power module status signal and the server standby power supply signal are abnormal, it is determined that the main power module power supply environment is abnormal; if the main power status signal and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal.
[0059] In one optional implementation, the power module status acquisition device employs a programmable logic device. The device also includes a main connector and a secondary connector on the server motherboard; the power module status acquisition device connects to the secondary connector via the main connector. The server motherboard transmits standby power module status signals and server standby power signals to the power module status acquisition device via the connectors.
[0060] In an optional embodiment, the device further includes a battery module connected to the power module status collector and the detection manager, respectively, to supply power to the power module status collector and the detection manager.
[0061] To further understand the present invention, a specific embodiment is provided below to illustrate the invention in a more detailed manner. Figure 6 This is a schematic diagram of the structure of this specific embodiment.
[0062] like Figure 6As shown, the device in this specific embodiment includes a CPLD, a detection manager, a battery module, and a main connector. A secondary connector is provided on the server motherboard, which is connected to a first PSU and a second PSU. The first PSU is powered by AC mains, and the second PSU is powered by a UPS.
[0063] The main connector is connected to the secondary connector, the CPLD is connected to the detection manager and the main connector respectively, the battery module is connected to the CPLD and the detection manager respectively, and the server motherboard charges the battery module.
[0064] The CPLD monitors the status of the PSU1_POWER_GOOD, PSU1_EN, PSU2_POWER_GOOD, and PSU2_EN signals of the first and second PSUs, as well as the status of the P12V_STBY signal.
[0065] Working Principle: When a user is conducting a power-off test (or in a power supply anomaly scenario, such as a sudden failure of one power supply line), and one of the power supplies is shut down for testing (AC mains power is turned off), the CPLD can monitor the signal status of two PSUs. For example, after one power line is cut off, according to the power redundancy mode, the power supply will immediately switch to the other power line, and the server will not crash. However, in abnormal situations, after one power line fails to switch back to the other power line in time, causing a large-scale server crash. In this case, by monitoring the signal status of the two PSUs, it can be clearly determined whether the problem is on the server side or due to factors in the power supply environment of the data center. Under normal circumstances, the CPLD monitors the following signals as high-level: PSU1_POWER_GOOD, PSU1_EN, PSU2_POWER_GOOD, PSU2_EN, and P12V_STBY. If any signal is abnormal, the CPLD will notify the detection manager, which will generate a log and issue an alarm via a buzzer. If the PSU1_POWER_GOOD, PSU1_EN, PSU2_POWER_GOOD, PSU2_EN, and P12V_STBY signals remain in a normal state (e.g., after the AC mains power is cut off and the server crashes), and the PSU2_POWER_GOOD, PSU2_EN, and P12V_STBY signals are all pulled low, it indicates a problem with the customer's power supply environment. If the PSU2_POWER_GOOD, PSU2_EN, and P12V_STBY signals remain unchanged, but the server still crashes, it indicates a problem with the reliability of the server's redundancy logic.
[0066] The foregoing has described in detail an embodiment of a server power module status detection device. Based on the server power module status detection device described in the above embodiment, this invention also provides a server power module status detection method corresponding to the device.
[0067] Figure 7 This is a schematic flowchart of a server power module status detection method provided in an embodiment of the present invention, as shown below. Figure 7 As shown, the method includes the following steps.
[0068] S1 receives the status signals of the main power module, the backup power module, and the server standby power supply signal.
[0069] S2 records the status signals of the main power module, the backup power module, and the server standby power supply signal to the log.
[0070] Maintenance personnel can analyze the causes of faults based on the signals recorded in the logs. For example, when switching from the main power module to the backup power module, if the backup power module's power supply environment is normal (i.e., the uninterruptible power supply is normal), both the backup power module's status signals and the server's standby power supply signals will be normal, without any abnormalities. If a system crash occurs at this time, it indicates an abnormality in the server's redundancy logic reliability. If the backup power module's power supply environment is abnormal (i.e., the uninterruptible power supply is abnormal), both the backup power module's status signals and the server's standby power supply signals will show abnormal states, and a system crash at this time can be determined to be due to an abnormality in the backup power module's power supply environment. Similarly, when switching from the backup power module to the main power module, if the main power module's power supply environment is normal (i.e., the mains power is normal), both the main power module's status signals and the server's standby power supply signals will be normal, without any abnormalities. If a system crash occurs at this time, it indicates an abnormality in the server's redundancy logic reliability. If the main power module's power supply environment is abnormal (i.e., the mains power is abnormal), both the main power module's status signals and the server's standby power supply signals will show abnormal states, and a system crash at this time can be determined to be due to an abnormality in the main power module's power supply environment. Maintenance personnel can view the main power module status signal and the server standby power supply signal in the log. Based on the status of the main power module status signal and the server standby power supply signal, they can determine the cause of the fault.
[0071] In one optional implementation, the cause of the fault can be automatically determined, specifically including the following steps.
[0072] S1 receives the status signals of the main power module, the backup power module, and the server standby power supply signal.
[0073] S2 receives server status signals and power switching signals.
[0074] S3 records the status signals of the main power module, the backup power module, and the server standby power supply signal to the log.
[0075] S4 responds to server crashes during power switching by analyzing the cause of the fault based on the power module status signal and the server standby power supply signal, and records the cause of the fault in the log.
[0076] Specifically, step S4 includes: In response to a server crash when the main power module switches to the backup power module, if the status signal of the backup power module and the server standby power supply signal are abnormal, it is determined that the backup power module's power supply environment is abnormal; if the status signal of the backup power module and the server standby power supply signal are normal, it is determined that the server's redundancy logic reliability is abnormal. Similarly, in response to a server crash when the backup power module switches to the main power module, if the status signal of the main power module and the server standby power supply signal are abnormal, it is determined that the main power module's power supply environment is abnormal; if the status signal of the main power module and the server standby power supply signal are normal, it is determined that the server's redundancy logic reliability is abnormal.
[0077] The server power module status detection method of this embodiment is based on the aforementioned server power module status detection device. Therefore, the specific implementation method of this method can be found in the embodiment section of the server power module status detection device above. Thus, the specific implementation method can be referred to the description of the corresponding embodiments, and will not be elaborated here.
[0078] Furthermore, since the server power module status detection method in this embodiment is based on the aforementioned server power module status detection device, its function corresponds to that of the aforementioned device, and will not be described again here.
[0079] Figure 8 A schematic diagram of a terminal 800 provided in an embodiment of the present invention includes: a processor 810, a memory 820, and a communication unit 830. The processor 810 is used to implement the following steps when executing the server power module status detection program stored in the memory 820: Receives status signals from the main power module, the backup power module, and the server standby power supply signal; The status signals of the main power module, the backup power module, and the server standby power supply signal are recorded in the log.
[0080] This invention collects the status signals of the power module and the standby power supply signals of the server in real time, and records the collected signals in a log for maintenance personnel to view. Maintenance personnel can analyze the cause of the fault based on the status signals of the power module and the standby power supply signals of the server, providing data support for determining the cause of the fault and subsequent maintenance, and reducing the probability of subsequent downtime faults.
[0081] The terminal 800 includes a processor 810, a memory 820, and a communication unit 830. These components communicate via one or more buses. Those skilled in the art will understand that the server structure shown in the figure does not constitute a limitation of the present invention. It can be a bus topology or a star topology, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0082] The memory 820 can be used to store the execution instructions of the processor 810. The memory 820 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 820 are executed by the processor 810, the terminal 800 is able to perform some or all of the steps in the above method embodiments.
[0083] The processor 810 serves as the control center of the storage terminal, connecting various parts of the electronic terminal via various interfaces and lines. It executes software programs and / or modules stored in the memory 820, and calls data stored in the memory to perform various functions of the electronic terminal and / or process data. The processor can be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, the processor 810 may consist only of a central processing unit (CPU). In this embodiment of the invention, the CPU may have a single processing core or include multiple processing cores.
[0084] The communication unit 830 is used to establish a communication channel, enabling the storage terminal to communicate with other terminals. It can receive user data sent by other terminals or send user data to other terminals.
[0085] The present invention also provides a computer storage medium, wherein the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0086] The computer storage medium stores a server power module status detection program, which, when executed by the processor, performs the following steps: Receives status signals from the main power module, the backup power module, and the server standby power supply signal; The status signals of the main power module, the backup power module, and the server standby power supply signal are recorded in the log.
[0087] This invention collects the status signals of the power module and the standby power supply signals of the server in real time, and records the collected signals in a log for maintenance personnel to view. Maintenance personnel can analyze the cause of the fault based on the status signals of the power module and the standby power supply signals of the server, providing data support for determining the cause of the fault and subsequent maintenance, and reducing the probability of subsequent downtime faults.
[0088] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or other media capable of storing program code. It includes several instructions to cause a computer terminal (which may be a personal computer, server, or a second terminal, network terminal, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0089] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0090] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0091] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0092] The above-disclosed embodiments are merely preferred embodiments of the present invention, but the present invention is not limited thereto. Any non-creative variations that can be conceived by those skilled in the art, as well as any improvements and modifications made without departing from the principles of the present invention, should fall within the protection scope of the present invention.
Claims
1. A server power module status detection device, characterized in that, include: Power module status acquisition and detection manager; The power module status acquisition device is connected to the server motherboard and is configured to acquire status signals of the backup power module and server standby power supply signals from the server motherboard and send them to the detection manager. The detection manager is connected to the power module status collector and is configured to receive the backup power module status signal and the server standby power supply signal sent by the power module status collector, and record the backup power module status signal and the server standby power supply signal to generate a detection log. When a server crashes after switching from the main power module to the backup power module, the detection manager analyzes the cause of the failure based on the status signals of the backup power module and the server standby power supply signal, and records the cause of the failure in the log. Specifically, the analysis of the cause of the failure based on the status signals of the backup power module and the server standby power supply signal includes: if the status signals of the backup power module and the server standby power supply signal are abnormal, it is determined that the backup power module power supply environment is abnormal; if the status signals of the backup power module and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal. The server power module status detection device can collect the server standby power supply signals and backup power module status signals of multiple servers. When collecting information, each server is identified. The collected server standby power supply signals and backup power module status signals of each server are mapped to the server identifier to generate logs, and the generated logs are sent to the management terminal.
2. The server power module status detection device according to claim 1, characterized in that, The detection manager is also connected to the server motherboard and is configured to collect server status signals and power switching signals, and determine whether the server has crashed based on the server status signals.
3. The server power module status detection device according to claim 1 or 2, characterized in that, The power module status acquisition device is also configured to acquire status signals of the main power module from the server motherboard and send them to the detection manager; Correspondingly, the detection manager is also configured to receive the main power module status signal sent by the power module status collector and record the main power module status signal in the detection log.
4. The server power module status detection device according to claim 3, characterized in that, The detection manager is also configured to analyze the cause of the failure based on the status signal of the main power module and the standby power supply signal of the server in response to a server crash when the backup power module switches to the main power module, and record the cause of the failure in the log. Specifically, analyzing the cause of the fault based on the status signal of the main power module and the standby power supply signal of the server includes: if the status signal of the main power module and the standby power supply signal of the server are abnormal, it is determined that the power supply environment of the main power module is abnormal; if the status signal of the main power module and the standby power supply signal of the server are normal, it is determined that the reliability of the server redundancy logic is abnormal.
5. The server power module status detection device according to claim 4, characterized in that, The status signals of the power module include the normal output voltage indication signal and the power control enable signal; when the normal output voltage indication signal and the power control enable signal are at a high level, the status signal of the power module is determined to be normal; when the normal output voltage indication signal and the power control enable signal are at a low level, the status signal of the power module is determined to be abnormal. The server standby power supply signal is considered to be normal when it is high. The server is deemed to have an abnormal standby power supply signal when the standby power supply signal is low.
6. The server power module status detection device according to claim 5, characterized in that, The power module status acquisition device uses a programmable logic device.
7. The server power module status detection device according to claim 6, characterized in that, The device also includes a main connector, and a secondary connector is provided on the server motherboard; the power module status collector is connected to the secondary connector through the main connector.
8. The server power module status detection device according to claim 7, characterized in that, The device also includes a battery module, which is connected to the power module status collector and the detection manager respectively, and supplies power to the power module status collector and the detection manager.
9. A method for detecting the status of a server power module, characterized in that, Includes the following steps: Receives status signals from the main power module, the backup power module, and the server standby power supply signal; The status signals of the main power module, the backup power module, and the server standby power supply signal are recorded in the log. In response to a server crash when the main power module switches to the backup power module, the cause of the failure is analyzed based on the status signal of the backup power module and the server standby power supply signal, and the cause of the failure is recorded in the log. Specifically, the analysis of the cause of the failure based on the status signal of the backup power module and the server standby power supply signal includes: if the status signal of the backup power module and the server standby power supply signal are abnormal, it is determined that the backup power module power supply environment is abnormal; if the status signal of the backup power module and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal. The method also includes: collecting server standby power supply signals and backup power module status signals from multiple servers; identifying each server during information collection; mapping the collected server standby power supply signals and backup power module status signals of each server to the server identifier to generate logs; and sending the generated logs to the management terminal.
10. The server power module status detection method according to claim 9, characterized in that, The method also includes the following steps: Receive server status signals and power switching signals; In response to server downtime during power switching, the cause of the failure is analyzed based on the power module status signals and the server standby power supply signals, and the cause of the failure is recorded in the log. Specifically, this includes: In response to a server crash when the main power module switches to the backup power module, if the status signal of the backup power module and the server standby power supply signal are abnormal, it is determined that the backup power module power supply environment is abnormal; if the status signal of the backup power module and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal. In response to a server crash when the backup power module switches to the main power module, if the status signal of the main power module and the server standby power supply signal are abnormal, it is determined that the power supply environment of the main power module is abnormal; if the status signal of the main power module and the server standby power supply signal are normal, it is determined that the server redundancy logic reliability is abnormal.
Citation Information
Patent Citations
Method and device for monitoring power supplies on server mainboard, and equipment
CN108919935A