Fault detection method and device for server testing, storage medium and electronic device
By providing temporary power and collecting status information during the alternating power-on and power-off testing of the server, the problem of inaccurate determination of the cause of PSU power failure was solved, achieving efficient fault detection and reliability assurance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSPUR SUZHOU INTELLIGENT TECH CO LTD
- Filing Date
- 2023-09-19
- Publication Date
- 2026-07-24
AI Technical Summary
In existing technologies, the cause of PSU power failure cannot be accurately determined during alternating power-on and power-off testing of servers, resulting in a waste of testing resources and manpower, and may leave safety hazards, with low fault detection efficiency.
When an undervoltage condition is detected in the power supply component, a temporary power supply that meets the target power supply conditions is connected to the server, server status information and power distribution equipment power status information are collected, and the cause of the fault is located.
It improves the efficiency of fault detection in server testing, reduces the waste of testing resources and manpower, shortens the debugging cycle, and ensures the reliability of server operation.
Smart Images

Figure CN117331731B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to a fault detection method and apparatus for server testing, a storage medium, and an electronic device. Background Technology
[0002] In real-world use, servers typically face dynamic power consumption demands in complex scenarios. For example, the power supply system of servers in data center server rooms requires regular switching and maintenance, and the server's power modules also need regular maintenance and replacement. These real-world scenarios fall under the category of alternating power-on and power-off cycles for servers. The alternating power-on and power-off process involves single power supply and the disappearance of redundancy, resulting in a higher power supply risk. Therefore, to ensure the stability of server power supply under these conditions, up to 10,000 alternating power-on and power-off tests are conducted during the R&D phase. In the alternating power-on and power-off mode, the power supply typically operates on for 3 seconds and off for 3 seconds.
[0003] However, during 10,000 alternating 3S on and 3S off power cycles on and off of the server, if the server experiences an abnormal power outage, although the system restart log can be viewed in the system event log (SEL), it's impossible to determine whether the power outage was caused by a server-side malfunction or a power supply failure in the PDU (Power Distribution Unit). This poses a significant challenge for developers debugging. If the cause of the fault cannot be clarified, the alternating power outage test must be restarted from scratch to ensure the reliability of the delivered server, resulting in a significant waste of equipment, manpower, and laboratory resources, and potentially causing project delivery delays. Furthermore, on-site debugging of anomalies is necessary, leading to poor timeliness.
[0004] There is still no effective solution to the problem of low efficiency in fault detection during server testing in related technologies. Summary of the Invention
[0005] This application provides a method and apparatus for fault detection in server testing, as well as a storage medium and electronic device, to at least solve the problem of low efficiency in fault detection in server testing in related technologies.
[0006] According to one embodiment of the present application, a fault detection method for server testing is provided, comprising:
[0007] During the alternating power-on and power-off test of the server, if the power supply component is detected to be in an undervoltage state, a temporary power supply that meets the target power supply conditions is connected to the server. The target power supply conditions are used to indicate the operating environment that allows the server to transmit the information it generates.
[0008] Server status information is collected from the server, wherein the server status information is used to indicate the historical operating status of each component within the server;
[0009] The cause of the undervoltage state of the server is located based on the server status information and the power status information of the server's power distribution device. The power distribution device is used to distribute voltage to the power supply components of the server, and the power status information is used to indicate the historical operating status of the power distribution device.
[0010] Optionally, providing the server with a temporary power supply that meets the target power supply conditions includes:
[0011] A power supply command is sent to the backup power supply corresponding to the server, wherein the backup power supply is connected to the server and is in an on state during the server's alternating power-on and power-off test. When the backup power supply in the on state receives the power supply command, it is allowed to temporarily supply power to the server.
[0012] Optionally, before connecting the server to a temporary power supply that meets the target power supply conditions when the power supply component is detected to be in an undervoltage state, the method further includes:
[0013] The voltage of the power bus of the server is detected to obtain the bus voltage, wherein the power supply component supplies power to the server through the power bus;
[0014] If the bus voltage is less than or equal to a first voltage threshold, it is determined that the power supply component is in an undervoltage state.
[0015] Optionally, after detecting the voltage of the power bus of the server to obtain the bus voltage, the method further includes:
[0016] When the bus voltage recovers to the second voltage threshold, a shutdown command is sent to the backup power supply, wherein the second voltage threshold is greater than the first voltage threshold, the backup power supply is connected to the server, and the backup power supply is in the on state during the server's alternating power-on / off test. When the backup power supply in the on state receives the shutdown command, it is prohibited from providing temporary power to the server.
[0017] Optionally, collecting server status information from the server includes:
[0018] The operation logs of each associated component in the server are collected to obtain an operation log set, and the register information of each register in the power supply component is polled to obtain a register information set. The associated component is a component that is allowed to cause the power supply component to be in an undervoltage state when a fault occurs. The operation log set is used to indicate the historical operation status of each associated component, and the register information set is used to indicate the historical operation status of each power supply component.
[0019] The set of running logs and the set of register information are determined as the server status information.
[0020] Optionally, before locating the cause of the fault leading to the server being in the undervoltage state based on the server status information and the power status information of the server's power distribution device, the method further includes:
[0021] The duration of the power supply component being in an undervoltage state is detected; and the voltage waveform of the output voltage of the power distribution device during the period when the power supply component is in an undervoltage state is obtained;
[0022] The duration and the voltage waveform are determined as the power supply status information.
[0023] Optionally, locating the cause of the server being in the undervoltage state based on the server status information and the power status information of the server's power distribution device includes:
[0024] If the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, the duration is less than or equal to the preset duration, and the voltage waveform indicates that the output voltage is intermittent, then the cause of the fault is determined to be an intermittent fault in the power distribution equipment.
[0025] If the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, and the voltage waveform indicates that the output voltage has experienced a long-term failure, then the cause of the failure is determined to be a long-term failure of the power distribution equipment.
[0026] If the running log set indicates that the associated component has failed, and / or the register information set indicates that the register has failed, and / or the duration is longer than the preset duration, the cause of the failure is determined to be a failure of the server.
[0027] Optionally, after locating the cause of the fault leading to the server being in the undervoltage state based on the server status information and the power status information of the server's power distribution device, the method further includes:
[0028] If the cause of the failure is a momentary power outage in the power distribution equipment, or if the cause of the failure is a prolonged power outage in the power distribution equipment, the alternating power-on and power-off test shall continue.
[0029] If the cause of the failure is a server malfunction, the alternating power-on / off test shall be terminated.
[0030] According to another embodiment of the present application, a fault detection device for server testing is also provided, comprising:
[0031] The first detection module is used to connect a temporary power supply that meets the target power supply conditions to the server when the power supply component is detected to be in an undervoltage state during the alternating power-on and power-off test of the server. The target power supply conditions are used to indicate the operating environment that allows the server to transmit the information it generates.
[0032] The acquisition module is used to acquire server status information from the server, wherein the server status information is used to indicate the historical operating status of each component within the server;
[0033] The positioning module is used to locate the cause of the fault that causes the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device, wherein the power distribution device is used to distribute voltage to the power supply components of the server, and the power status information is used to indicate the historical operating status of the power distribution device.
[0034] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described fault detection method for server testing when running.
[0035] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described fault detection method for server testing through the computer program.
[0036] In this embodiment, during the alternating power-on / off test of the server, if an undervoltage condition is detected in the power supply component, a temporary power supply meeting the target power supply conditions is provided to the server. The target power supply conditions indicate the operating environment that allows the server to transmit generated information. Server status information is collected from the server, indicating the historical operating status of each component within the server. The cause of the undervoltage condition is located based on the server status information and the power status information of the server's power distribution equipment. The power distribution equipment distributes voltage to the server's power supply component, and the power status information indicates the historical operating status of the power distribution equipment. In other words, when an undervoltage condition is detected in the power supply component during the alternating power-on / off test, a temporary power supply meeting the target power supply conditions is immediately provided to the server to allow the server to transmit generated information. Then, server status information indicating the historical operating status of each component within the server is collected from the server. Finally, the cause of the undervoltage condition is located based on the server status information and the power status information of the server's power distribution equipment. This method allows for the immediate determination of the cause of the undervoltage condition when the power supply component is undervoltage. By adopting the above technical solution, the problem of low efficiency in fault detection during server testing is solved, and the technical effect of improving the efficiency of fault detection during server testing is achieved. Attached Figure Description
[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a schematic diagram of the hardware environment for a fault detection method for server testing according to an embodiment of this application.
[0040] Figure 2 This is a flowchart of a fault detection method for server testing according to an embodiment of this application;
[0041] Figure 3 This is a schematic diagram of the hardware connection for fault detection in a server test according to an embodiment of this application;
[0042] Figure 4This is a schematic diagram of a server testing fault detection system according to an embodiment of this application;
[0043] Figure 5 This is a schematic diagram illustrating the function of a server-side fault monitoring module according to an embodiment of this application;
[0044] Figure 6 This is a schematic diagram illustrating the function of a PDU-end intelligent monitoring module according to an embodiment of this application;
[0045] Figure 7 This is a structural block diagram of a fault detection device for server testing according to an embodiment of this application. Detailed Implementation
[0046] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0047] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0048] The methods and embodiments provided in this application can be executed on a computer terminal, device terminal, or similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a schematic diagram of the hardware environment for a fault detection method for server testing according to an embodiment of this application. Figure 1 As shown, a computer terminal may include one or more ( Figure 1Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. In one exemplary embodiment, the computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 Equivalent functions or ratios shown Figure 1 The functions shown have more different configurations.
[0049] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the server testing fault detection method in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0050] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0051] The terms used in the embodiments of this application are explained as follows:
[0052] PSU: Power supply unit;
[0053] PDU: Power Distribution Unit, also known as power distribution equipment;
[0054] CPU: Central Processing Unit;
[0055] GPU: Graphics Processing Unit;
[0056] BMC: Base-board management controller;
[0057] CPLD: Complex Programmable Logic Device.
[0058] Before introducing the fault detection method for server testing proposed in this application, let's first explain the relevant technologies: In these technologies, during 10,000 cycles of alternating 3S on and 3S off power-on / off cycles on the server's power supply components, if the server experiences an abnormal power outage and shutdown, the common approach is to check the PSU's black-box logs for a preliminary analysis of the cause of the PSU power outage, checking if the PDU is malfunctioning and unable to function properly. If it's a server or power supply component failure, the problem needs to be reproduced and the bug fixed. For any test interruptions caused by the above reasons, a second test must be performed.
[0059] Therefore, related technologies cannot accurately determine whether a PSU power failure is caused by a server-side malfunction or a PSU failure itself. While checking if a PDU is malfunctioning and unable to function properly can identify prolonged PDU malfunctions, it cannot monitor PSU power failures caused by intermittent PDU power supply interruptions. In reality, if the problem is a server or power supply component failure, it needs to be reproduced and resolved. Failure to reproduce the problem may lead to overlooking potential issues and jeopardizing server reliability. If the problem is caused by a PDU malfunction or intermittent interruption, secondary testing is unnecessary if the test can continue. The aforementioned related technologies suffer from several problems: the inability to accurately clarify the cause of PSU power failures necessitates secondary testing, resulting in wasted testing resources and manpower, and project delivery delays. Furthermore, addressing abnormal power failures caused by servers or power supplies leads to lengthy debugging cycles, the inability to reproduce fault phenomena, and potential safety hazards. There are also issues with accurately identifying abnormal server power failures caused by PDU malfunctions or intermittent interruptions, and the inability to continue "alternating power-on / off tests" in PDU malfunction or intermittent interruption states. Therefore, this application proposes a server testing fault detection method to address these problems in related technologies.
[0060] This embodiment provides a fault detection method for server testing, applied to the aforementioned computer terminal. Figure 2 This is a flowchart of a fault detection method for server testing according to an embodiment of this application, such as... Figure 2As shown, the process includes the following steps:
[0061] Step S202: During the alternating power-on and power-off test of the server, if the power supply component is detected to be in an undervoltage state, a temporary power supply that meets the target power supply conditions is connected to the server. The target power supply conditions are used to indicate the operating environment that allows the server to transmit the information it generates.
[0062] Step S204: Collect server status information from the server, wherein the server status information is used to indicate the historical operating status of each component within the server;
[0063] Step S206: Locate the cause of the fault that causes the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device, wherein the power distribution device is used to distribute voltage to the power supply components of the server, and the power status information is used to indicate the historical operating status of the power distribution device.
[0064] Through the above steps, during the alternating power-on / off test of the server, if an undervoltage condition is detected in the power supply components, temporary power is supplied to the server to meet the target power supply conditions. The target power supply conditions indicate the operating environment that allows the server to transmit generated information. Server status information is collected from the server, indicating the historical operating status of each component within the server. Based on the server status information and the power status information of the server's power distribution equipment, the cause of the server's undervoltage condition is located. The power distribution equipment distributes voltage to the server's power supply components, and the power status information indicates the historical operating status of the power distribution equipment. In other words, when an undervoltage condition is detected in the power supply components during the alternating power-on / off test, temporary power is immediately supplied to the server to meet the target power supply conditions, allowing the server to transmit generated information. Then, server status information indicating the historical operating status of each component within the server is collected. Finally, based on the server status information and the power status information of the server's power distribution equipment, the cause of the server's undervoltage condition is located. This method allows for the immediate identification of the cause of the server's undervoltage condition when the power supply components are undervoltage. By adopting the above technical solution, the problem of low efficiency in fault detection during server testing is solved, and the technical effect of improving the efficiency of fault detection during server testing is achieved.
[0065] In the technical solution provided in step S202 above, alternating power-on / off testing is a common test item for servers. The following is an introduction to "alternating power-on / off testing": Current servers usually adopt a redundant power supply design, the most common being 1+1 redundant power supply, that is, there are 2 PSUs (PSU0, PSU1) in the server to power the server. Since servers may face scenarios where the power supply system needs to be switched and maintained regularly, and the power modules also need to be maintained and replaced regularly, these actual scenarios all fall under the category of alternating power-on / off testing of server power supply. That is, the process of alternating power-on / off testing of server power supply is single power supply, the redundancy mode disappears, and the power supply risk is high. Therefore, to ensure the stability of server power supply at this time, it is necessary to conduct up to 10,000 alternating power-on / off tests of server power supply during the research and development stage. In the alternating power-on / off mode, the power supply is generally 3S on and 3S off.
[0066] Optionally, in this embodiment, the hardware connection method for alternating power-on and power-off testing is described as follows: Figure 3 This is a schematic diagram of a hardware connection for fault detection in server testing according to an embodiment of this application, as shown below. Figure 3 As shown, the server is equipped with PSUs (i.e., power supply units): PSU0 and PSU1. PSU0 and PSU1 are alternately powered on and off by an external PDU (i.e., power distribution unit) to complete the alternating power-on and power-off test of the server.
[0067] Optionally, in this embodiment, corresponding to the hardware connection for fault detection in the server test described above, this application also proposes a fault detection system for server testing. Figure 4 This is a schematic diagram of a server testing fault detection system according to an embodiment of this application, as shown below. Figure 4As shown, the server testing fault detection system includes: a server test status remote monitoring module, a server-side fault monitoring module, and a PDU-side intelligent monitoring module. The CPLD and BMC within the server work together to implement the server testing fault detection function. The server test status remote monitoring module is deployed on the client side (which can be a remote terminal device such as a laptop or mobile phone), the server-side fault monitoring module is deployed on the server, and the PDU-side intelligent monitoring module is deployed on the PDU. The server-side fault monitoring module and the PDU-side intelligent monitoring module operate independently but can be remotely interconnected via a local area network. The operating status and alarm information of the server-side fault monitoring module and the PDU-side intelligent monitoring module are promptly transmitted to the server test status remote monitoring module. The server test status remote monitoring module then promptly issues alarm information to notify the R&D personnel. Simultaneously, the R&D personnel can remotely download fault logs, polling logs, and PDU power supply waveforms stored in the cloud by the server-side fault monitoring module and the PDU-side intelligent monitoring module through the server test status remote monitoring module, thereby analyzing and locating the cause of the fault for timely and accurate on-site handling.
[0068] Optionally, in this embodiment, if an undervoltage condition is detected in the power supply component during the alternating power-on / off test of the server, it indicates that the power supply component is about to lose power. Taking a general-purpose server powered by a 1+1 redundant power supply as an example, after the server power supply fails, the 12V bus high level of the server system can only be maintained for about 60ms. During this 60ms period, the functions of various parts of the system can still work normally, but it is not possible to complete the collection of server status information, i.e., polling, recording, and saving of server-side BMC fault logs and power register fault power failure information. To solve the problem of the short duration of the 12V bus high level after the PSU power failure, a temporary power supply that meets the target power supply conditions is connected to the server. After the server connects to the temporary power supply, it is allowed to transmit the generated information, i.e., it supports the collection of server status information.
[0069] In one exemplary embodiment, the server may be connected to a temporary power supply that meets the target power supply conditions by means of, but not limited to, sending a power supply command to the backup power supply corresponding to the server, wherein the backup power supply is connected to the server, the backup power supply is in an on state during the server performing the alternating power-on and power-off test, and the backup power supply in the on state is allowed to temporarily supply power to the server upon receiving the power supply command.
[0070] Optionally, in this embodiment, the backup power supply may be, but is not limited to, a server backup battery module, which aims to solve the problem of the short duration of the 12V high level on the bus after the PSU loses power. The backup power supply can be connected to the server by reserving an external power supply port on the server power supply PCB board during the R&D stage, allowing the external server backup battery module to be connected, and equipping it with an enable switch. Figure 5 This is a schematic diagram illustrating the function of a server-side fault monitoring module according to an embodiment of this application, such as... Figure 5 As shown, at the start of the alternating power-on / off test of the server, the enable switch of the server's backup battery module is turned on (i.e., the backup power supply is in the on state during the server's alternating power-on / off test). When the bus voltage sensor of the server's backup battery module detects that the server power bus voltage has dropped to 11V or below, the server's CPLD detects the voltage sensor value and triggers an alarm signal, immediately issuing an access operation command (i.e., a power supply command) to the server's backup battery module to control the temporary power supply to the server. During the above process, the server power bus voltage monitoring and the server backup battery module start / stop command are monitored and issued in real time through the server CPLD.
[0071] When the server's CPLD detects that the server's power bus voltage has dropped to 11V (i.e., the first voltage threshold) or below, it indicates that the power supply component is in an undervoltage state. Simultaneously, it sends a one-click fault log collection command to the system-side BMC (i.e., collects the operating logs of various related components in the server to obtain an operating log set), and begins polling the PSU register information 20 times (i.e., polls the register information of each register within the power supply component to obtain a register information set). The fault log and polling log are stored in the server system's local folder and can be downloaded remotely via the local area network. After the system BMC completes the fault log collection and polling and saving of the PSU register information, it sends an action completion command to the system-side CPLD. When the server's backup battery module's bus voltage sensor detects that the server's power bus voltage is greater than 11.6V (i.e., the second voltage threshold) and transmits this information to the CPLD, the system-side CPLD sends a shutdown command to the server's backup battery module (i.e., a shutdown command), and the server's backup battery pack stops working and completes automatic charging. By collecting and saving fault logs and PSU register fault information with one click, the R&D risks caused by the inability to reproduce low-probability bugs (faults) can be effectively resolved, the debugging cycle can be shortened, and the reliability of server operation can be improved.
[0072] In one exemplary embodiment, before connecting the server to a temporary power supply that meets the target power supply conditions when the power supply component is detected to be in an undervoltage state, the following methods may be included, but are not limited to: detecting the voltage of the power bus of the server to obtain the bus voltage, wherein the power supply component supplies power to the server through the power bus; and determining that the power supply component is in an undervoltage state when the bus voltage is less than or equal to a first voltage threshold.
[0073] Optionally, in this embodiment, the voltage of the server's power bus can be detected, but is not limited to, by a bus voltage sensor. The voltage of the power bus can be monitored by a CPLD. The first voltage threshold can be the aforementioned 11V, or it can be set according to actual needs.
[0074] In one exemplary embodiment, after detecting the voltage of the power bus of the server and obtaining the bus voltage, the method may include, but is not limited to, the following: when the bus voltage recovers to a second voltage threshold, sending a shutdown command to the backup power supply, wherein the second voltage threshold is greater than the first voltage threshold, the backup power supply is connected to the server, the backup power supply is in an on state during the server performing the alternating power-on / off test, and the backup power supply in the on state is prohibited from temporarily supplying power to the server upon receiving the shutdown command.
[0075] Optionally, in this embodiment, the second voltage threshold can be the aforementioned 11.6V, or it can be set according to actual needs. In the initial stage, if the bus voltage is less than or equal to the first voltage threshold, the power supply component is in an undervoltage state, and temporary server power supply is activated. Once the bus voltage recovers to the second voltage threshold, indicating that the undervoltage state of the power supply component has been repaired, the temporary server power supply is shut down.
[0076] In the technical solution provided in step S204 above, in step S202, in order to solve the problem of the short holding time of the 12V high level on the bus after the PSU power failure, a temporary power supply that meets the target power supply conditions is connected to the server. After the server is connected to the temporary power supply, the generated information is allowed to be transmitted, which prepares the preconditions for collecting server status information from the server.
[0077] In one exemplary embodiment, server status information can be collected from the server in the following ways, but not limited to: collecting the operation logs of each associated component in the server to obtain an operation log set, and polling the register information of each register in the power supply component to obtain a register information set, wherein the associated component is a component that is allowed to cause the power supply component to be in an undervoltage state when a fault occurs, the operation log set is used to indicate the historical operation status of each associated component, and the register information set is used to indicate the historical operation status of each power supply component; the operation log set and the register information set are determined as the server status information.
[0078] Optionally, in this embodiment, collecting the operation logs of each associated component in the server to obtain the operation log set can be, but is not limited to, the CPLD issuing a one-click fault log collection command to the BMC. After receiving the one-click collection command, the BMC immediately collects the fault logs in the server to obtain the operation log set. The specific collection object in the server can be an associated component, which can be one or more. When an associated component fails, it may cause the power supply component to be in an undervoltage state. According to the operation log set, the historical operation status of each associated component can be known. The operation failures or errors of the associated components can be known in the operation log set.
[0079] Optionally, in this embodiment, polling the register information of each register in the power supply component can be, but is not limited to, polling the PSU register information 20 times to obtain server status information. Faults or errors in each PSU register can be found in the register information.
[0080] In an exemplary embodiment, before locating the cause of the fault that leads the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device, the following methods may be included, but are not limited to: detecting the duration of the undervoltage state of the power supply component; acquiring the voltage waveform of the output voltage of the power distribution device during the period when the power supply component is in the undervoltage state; and determining the duration and the voltage waveform as the power status information.
[0081] Optionally, in this embodiment, the aforementioned power status information can be obtained through, but is not limited to, a PDU-side intelligent monitoring module. The PDU-side intelligent monitoring module includes: a PDU power supply status monitoring module and a PDU power supply compensation module. Figure 6 This is a schematic diagram illustrating the function of a PDU-end intelligent monitoring module according to an embodiment of this application, as shown below. Figure 6As shown, taking a general-purpose server with a 1+1 redundant power supply as an example, during a test of 10,000 alternating power-on and power-off cycles, the server is supplied with power to server PSU0 and PSU1 by the main PDU, alternating between 3S on and 3S off. It must be ensured that at any given time, one of PSU0 and PSU1 is in the 3S on state. Each output port of the PDU is equipped with an external voltage monitoring port, which can be connected to a portable oscilloscope differential carbon rod clamp to monitor the real-time voltage waveform of the PDU power supply port (i.e., to obtain the voltage waveform of the output voltage of the power distribution device during the period when the power supply component is in an undervoltage state). If the server CPLD detects that the server power bus voltage has dropped to 11V or below, it immediately sends an alarm command to the PDU's built-in PDU power supply status monitoring module. After receiving the alarm command from the server CPLD, the PDU power supply status monitoring module triggers an oscilloscope screenshot after a 30-second delay (the specific duration can be adjusted according to actual needs) and automatically saves it to a local folder, which can be downloaded remotely via the local area network. The above methods effectively solve the problem of not being able to pinpoint the exact location of abnormal power outage faults in PDUs.
[0082] The PDU power supply compensation module includes a primary PDU and a backup PDU. The primary PDU and the backup PDU are interconnected via a network cable and are uniformly allocated and used by the PDU power supply status monitoring module. Only one primary or backup PDU can be in output state at the same time.
[0083] Upon receiving an alarm command from the server CPLD, the PDU power supply status monitoring module immediately activates the backup PDU of the PDU power supply compensation module to simultaneously power the server's PSU0 and PSU1 while the server's backup battery module is providing temporary power to the server. If the server CPLD detects that the server power bus voltage has recovered to above 11.6V, it sends a command to the PDU power supply status monitoring module, instructing the PDU power supply compensation module to use the backup PDU to continue powering the server and to call the main PDU's alternating power-on / off program to continue the remaining alternating power-on / off tests. The main PDU's alternating power-on / off program records the number of times the main PDU has completed the current 10,000 alternating power-on / off tests. For example, if the alternating power-on / off program shows that 4,000 tests have been completed, the backup PDU will take over from the main PDU to complete the remaining 6,000 tests.
[0084] If the server power bus voltage received by the PDU power supply status monitoring module from the CPLD remains at or below 11V, it indicates a server or PSU malfunction. In this case, the PDU power supply status monitoring module issues a command to stop the primary and backup PDUs in the PDU power supply compensation module from supplying power to the server, and the server remains powered off, awaiting on-site handling by R&D personnel. This method effectively solves the problem of server power outages caused by fluctuations in the primary PDU power supply DIP (Digital Interference Pulse) (i.e., intermittent power distribution equipment failure) and test interruptions caused by PDU power supply abnormalities (i.e., long-term power distribution equipment failure), requiring secondary testing. This effectively improves testing efficiency and ensures project testing timeliness.
[0085] In the technical solution provided in step S206 above, the causes of the failure come from two main directions: one is a server failure, and the other is a power distribution device failure. Further, a server failure may be due to a failure in the server's associated components or a failure in the PSU's registers; a power distribution device failure may be a power distribution device intermittent failure or a long-term failure. In this case, the alternating power-on / off test needs to be terminated when the server fails, and the R&D personnel need to analyze the specific problem and start the test from scratch after debugging. However, when the power distribution device fails, since it is an external cause to the server and not a cause of the server itself, it is not necessary to start the test from scratch, and the remaining alternating power-on / off tests can continue.
[0086] In one exemplary embodiment, the cause of the server being in the undervoltage state can be located, but is not limited to, based on the server status information and the power status information of the server's power distribution device, in the following ways: If the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, the duration is less than or equal to a preset duration, and the voltage waveform indicates that the output voltage is intermittent, the cause of the fault is determined to be an intermittent fault in the power distribution device; if the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, and the voltage waveform indicates that the output voltage is experiencing a prolonged fault, the cause of the fault is determined to be a prolonged fault in the power distribution device; if the operation log set indicates that the associated component is faulty, and / or the register information set indicates that the register is faulty, and / or the duration is greater than the preset duration, the cause of the fault is determined to be a server fault.
[0087] Optionally, in this embodiment, the preset duration may be, but is not limited to, 60 seconds, depending on... Figure 6 The control process of the PDU-side intelligent monitoring module shown demonstrates that when a power supply component is detected to be in an undervoltage state, the backup PDU in the PDU power supply compensation module will be activated to power PSU0 and PSU1. If the undervoltage state of the power supply component is restored, it indicates that the server and PSU are normal and the reason for the undervoltage state is that the external primary PDU has failed. In this case, the backup PDU will be used to take over from the primary PDU to continue the remaining alternating power-on and power-off tests. However, if the power supply component is still in an undervoltage state for a long time (e.g., more than 60 seconds), it is determined that the server and / or PSU has failed, and the alternating power-on and power-off tests will end.
[0088] In one exemplary embodiment, after locating the cause of the fault that leads the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device, the method may include, but is not limited to, the following: if the fault cause is a momentary power failure of the power distribution device, or if the fault cause is a long-term power failure of the power distribution device, continue the alternating power-on / off test; if the fault cause is a server failure, end the alternating power-on / off test.
[0089] Optionally, in this embodiment, if the cause of the failure is a momentary failure of the power distribution device, or if the cause of the failure is a long-term failure of the power distribution device, it indicates that the reason for the power supply component being in an undervoltage state is a failure of the external main PDU, and the alternating power-on and power-off test can continue.
[0090] Optionally, in this embodiment, if the cause of the failure is a server malfunction, indicating that the power supply component is in an undervoltage state is due to a server (associated component, and / or, PSU) malfunction, the alternating power-on / off test is terminated.
[0091] Furthermore, the fault detection method for server testing proposed in this application can not only be applied to the 3S on and 3S off alternating power-on and power-off test of a 1+1 redundant power supply for a server, but also to any redundant power supply, such as the 3S on and 3S off alternating power-on and power-off test of a 1+2 redundant power supply. This application aims to accurately locate the cause of the power supply component being in an undervoltage state during the alternating power-on and power-off test, and to control the alternating power-on and power-off test based on the located cause.
[0092] The above methods effectively solve the problems of server power failure during 10,000 cycles of alternating 3S on and 3S off power-on and power-off cycles with 1+1 redundant power supplies. These problems include the inability to accurately identify the cause of the failure, leading to secondary testing and verification, wasted testing resources and manpower, and project delivery delays. They also effectively solve the problems of long debugging cycles and potential safety hazards caused by abnormal power failures due to server or power supply issues. Furthermore, they effectively solve the problems of inaccurate identification of abnormal server power failures caused by PDU malfunctions or intermittent power outages, and the inability to continue system testing in PDU malfunction or intermittent power outage states.
[0093] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0094] Figure 7 This is a structural block diagram of a fault detection device for server testing according to an embodiment of this application; as shown... Figure 7 As shown, it includes:
[0095] The first detection module 702 is used to connect a temporary power supply that meets the target power supply conditions to the server when the power supply component is detected to be in an undervoltage state during the alternating power-on and power-off test of the server. The target power supply conditions are used to indicate the operating environment that allows the server to transmit the information it generates.
[0096] The acquisition module 704 is used to acquire server status information from the server, wherein the server status information is used to indicate the historical operating status of each component within the server;
[0097] The positioning module 705 is used to locate the cause of the fault that causes the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device, wherein the power distribution device is used to distribute voltage to the power supply components of the server, and the power status information is used to indicate the historical operating status of the power distribution device.
[0098] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0099] Through the above embodiments, during the alternating power-on / off test of the server, if an undervoltage condition is detected in the power supply component, a temporary power supply meeting the target power supply conditions is connected to the server. The target power supply conditions indicate the operating environment that allows the server to transmit the generated information. Server status information is collected from the server, indicating the historical operating status of each component within the server. Based on the server status information and the power status information of the server's power distribution equipment, the cause of the server's undervoltage condition is located. The power distribution equipment distributes voltage to the server's power supply component, and the power status information indicates the historical operating status of the power distribution equipment. In other words, when an undervoltage condition is detected in the power supply component during the alternating power-on / off test, a temporary power supply meeting the target power supply conditions is immediately connected to the server to allow the server to transmit the generated information. Then, server status information indicating the historical operating status of each component within the server is collected from the server. Finally, the cause of the server's undervoltage condition is located based on the server status information and the power status information of the server's power distribution equipment. This method allows for the immediate determination of the cause of the server's undervoltage condition when the power supply component is undervoltage. By adopting the above technical solution, the problem of low efficiency in fault detection during server testing is solved, and the technical effect of improving the efficiency of fault detection during server testing is achieved.
[0100] In an exemplary embodiment, the first detection module includes:
[0101] The first sending unit is used to send a power supply command to the backup power supply corresponding to the server, wherein the backup power supply is connected to the server, and the backup power supply is in an on state during the server's alternating power-on and power-off test. When the backup power supply in the on state receives the power supply command, it is allowed to temporarily supply power to the server.
[0102] In one exemplary embodiment, the apparatus further includes:
[0103] The second detection module is used to detect the voltage of the power bus of the server before connecting the server to a temporary power supply that meets the target power supply conditions when the power supply component is detected to be in an undervoltage state, and to obtain the bus voltage, wherein the power supply component supplies power to the server through the power bus;
[0104] The first determining module is used to determine that the power supply component is in an undervoltage state when the bus voltage is less than or equal to a first voltage threshold.
[0105] In one exemplary embodiment, the apparatus further includes:
[0106] The second sending unit is configured to, after detecting the voltage of the power bus of the server and obtaining the bus voltage, send a shutdown command to the backup power supply when the bus voltage recovers to a second voltage threshold, wherein the second voltage threshold is greater than the first voltage threshold, the backup power supply is connected to the server, the backup power supply is in an on state during the server performing the alternating power-on / off test, and the backup power supply in the on state is prohibited from temporarily supplying power to the server upon receiving the shutdown command.
[0107] In one exemplary embodiment, the acquisition module includes:
[0108] The acquisition unit is used to acquire the operation logs of each associated component in the server to obtain an operation log set, and to poll the register information of each register in the power supply component to obtain a register information set. The associated component is a component that is allowed to cause the power supply component to be in an undervoltage state when a fault occurs. The operation log set is used to indicate the historical operation status of each associated component, and the register information set is used to indicate the historical operation status of each power supply component.
[0109] The first determining unit is used to determine the set of running logs and the set of register information as the server status information.
[0110] In one exemplary embodiment, the apparatus further includes:
[0111] The third detection module is used to detect the duration of the power supply component being in the undervoltage state before locating the fault cause that causes the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device; and to obtain the voltage waveform of the output voltage of the power distribution device during the time period when the power supply component is in the undervoltage state.
[0112] The second determining module is used to determine the duration and the voltage waveform as the power supply state information.
[0113] In one exemplary embodiment, the third detection module includes:
[0114] The second determining unit is configured to determine that the cause of the fault is a power distribution device experiencing a flashover fault when the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, the duration is less than or equal to a preset duration, and the voltage waveform indicates that the output voltage is interrupted.
[0115] The third determining unit is used to determine that the cause of the fault is a long-term fault in the power distribution equipment when the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, and the voltage waveform indicates that the output voltage has a long-term fault.
[0116] The fourth determining unit is configured to determine that the cause of the failure is a server failure when the running log set indicates that the associated component has failed, and / or the register information set indicates that the register has failed, and / or the duration is longer than the preset duration.
[0117] In one exemplary embodiment, the apparatus further includes:
[0118] The first test module is used to continue the alternating power-on / off test after locating the cause of the fault that causes the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device, in the case that the fault cause is a momentary power failure of the power distribution device or a long-term power failure of the power distribution device.
[0119] The second test module is used to terminate the alternating power-on / off test if the cause of the failure is a server malfunction.
[0120] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0121] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0122] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0123] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0124] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0125] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0126] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A fault detection method for server testing, characterized in that, include: During the alternating power-on and power-off test of the server, if the power supply component is detected to be in an undervoltage state, a temporary power supply that meets the target power supply conditions is connected to the server. The target power supply conditions are used to indicate the operating environment that allows the server to transmit the information it generates. Server status information is collected from the server, wherein the server status information is used to indicate the historical operating status of each component within the server; The cause of the server being in the undervoltage state is located based on the server status information and the power status information of the server's power distribution device. The power distribution device is used to distribute voltage to the power supply components of the server, and the power status information is used to indicate the historical operating status of the power distribution device. The step of collecting server status information from the server includes: The operation logs of each associated component in the server are collected to obtain an operation log set, and the register information of each register in the power supply component is polled to obtain a register information set. The associated component is a component that is allowed to cause the power supply component to be in an undervoltage state when a fault occurs. The operation log set is used to indicate the historical operation status of each associated component, and the register information set is used to indicate the historical operation status of each power supply component. The set of running logs and the set of register information are determined as the server status information; Before locating the cause of the undervoltage state based on the server status information and the power status information of the server's power distribution device, the method further includes: The duration of the power supply component being in an undervoltage state is detected; and the voltage waveform of the output voltage of the power distribution device during the period when the power supply component is in an undervoltage state is obtained; The duration and the voltage waveform are determined as the power supply status information; The step of locating the cause of the server being in the undervoltage state based on the server status information and the power status information of the server's power distribution device includes: If the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, the duration is less than or equal to the preset duration, and the voltage waveform indicates that the output voltage is intermittent, then the cause of the fault is determined to be an intermittent fault in the power distribution equipment. If the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, and the voltage waveform indicates that the output voltage has experienced a long-term failure, then the cause of the failure is determined to be a long-term failure of the power distribution equipment. If the running log set indicates that the associated component has failed, and / or the register information set indicates that the register has failed, and / or the duration is longer than the preset duration, the cause of the failure is determined to be a failure of the server.
2. The method according to claim 1, characterized in that, The provision of temporary power supply to the server that meets the target power supply conditions includes: A power supply command is sent to the backup power supply corresponding to the server, wherein the backup power supply is connected to the server and is in an on state during the server's alternating power-on and power-off test. When the backup power supply in the on state receives the power supply command, it is allowed to temporarily supply power to the server.
3. The method according to claim 2, characterized in that, Before connecting the server to a temporary power supply that meets the target power supply conditions when the power supply component is detected to be in an undervoltage state, the method further includes: The voltage of the power bus of the server is detected to obtain the bus voltage, wherein the power supply component supplies power to the server through the power bus; If the bus voltage is less than or equal to a first voltage threshold, it is determined that the power supply component is in an undervoltage state.
4. The method according to claim 3, characterized in that, After detecting the voltage of the power bus of the server and obtaining the bus voltage, the method further includes: When the bus voltage recovers to the second voltage threshold, a shutdown command is sent to the backup power supply, wherein the second voltage threshold is greater than the first voltage threshold, the backup power supply is connected to the server, and the backup power supply is in the on state during the server's alternating power-on / off test. When the backup power supply in the on state receives the shutdown command, it is prohibited from providing temporary power to the server.
5. The method according to claim 1, characterized in that, After locating the cause of the fault leading to the server being in the undervoltage state based on the server status information and the power status information of the server's power distribution device, the method further includes: If the cause of the failure is a momentary power outage in the power distribution equipment, or if the cause of the failure is a prolonged power outage in the power distribution equipment, the alternating power-on and power-off test shall continue. If the cause of the failure is a server malfunction, the alternating power-on / off test shall be terminated.
6. A fault detection device for server testing, characterized in that, include: The first detection module is used to connect a temporary power supply that meets the target power supply conditions to the server when the power supply component is detected to be in an undervoltage state during the alternating power-on and power-off test of the server. The target power supply conditions are used to indicate the operating environment that allows the server to transmit the information it generates. The acquisition module is used to acquire server status information from the server, wherein the server status information is used to indicate the historical operating status of each component within the server; The positioning module is used to locate the cause of the fault that causes the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device, wherein the power distribution device is used to distribute voltage to the power supply components of the server, and the power status information is used to indicate the historical operating status of the power distribution device. The acquisition module includes: The acquisition unit is used to acquire the operation logs of each associated component in the server to obtain an operation log set, and to poll the register information of each register in the power supply component to obtain a register information set. The associated component is a component that is allowed to cause the power supply component to be in an undervoltage state when a fault occurs. The operation log set is used to indicate the historical operation status of each associated component, and the register information set is used to indicate the historical operation status of each power supply component. The first determining unit is used to determine the set of running logs and the set of register information as the server status information; The device further includes: The third detection module is used to detect the duration of the power supply component being in the undervoltage state before locating the fault cause that causes the server to be in the undervoltage state based on the server status information and the power status information of the server's power distribution device; and to obtain the voltage waveform of the output voltage of the power distribution device during the time period when the power supply component is in the undervoltage state. The second determining module is used to determine the duration and the voltage waveform as the power state information; The third detection module includes: The second determining unit is configured to determine that the cause of the fault is a power distribution device experiencing a flashover fault when the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, the duration is less than or equal to a preset duration, and the voltage waveform indicates that the output voltage is interrupted. The third determining unit is used to determine that the cause of the fault is a long-term fault in the power distribution equipment when the operation log set indicates that the associated component is operating normally, the register information set indicates that the register is operating normally, and the voltage waveform indicates that the output voltage has a long-term fault. The fourth determining unit is configured to determine that the cause of the failure is a server failure when the running log set indicates that the associated component has failed, and / or the register information set indicates that the register has failed, and / or the duration is longer than the preset duration.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 5.
8. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 5 through the computer program.
Citation Information
Patent Citations
CN109101358A
CN116301276A