Server failure processing method, storage medium, and electronic device

CN119902918BActive Publication Date: 2026-09-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202412000082.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2026-09-08
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

[0003]本申请实施例提供了一种服务器的故障处理方法、存储介质、电子设备,以至少解决相关技术中针对集成PMIC芯片的内存暂无故障检测及相应故障处理措施方法的问题

Benefits of technology

[0014] By obtaining first-time information and second-time information, and controlling the server to perform preset fault handling operations (such as alarms) based on the first-time information and second-time information, the technical effect of handling faults according to different types of faults and different times of occurrence is achieved, making fault handling more timely and intelligent, and minimizing the impact on the server. In related technologies, there is currently no fault detection and corresponding fault handling measures for memory with integrated PMIC chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902918B_ABST
    Figure CN119902918B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a server fault processing method, a storage medium and an electronic device, wherein the method comprises: obtaining first time information, the first time information comprising at least one of a first time and a second time, the first time being a time when a first fault occurs, and the second time being a time when a second fault occurs; obtaining second time information, the second time information being a third time when a power supply management integrated chip (PMIC) voltage regulator is powered on; and generating a target control instruction set based on the first time information and the second time information, the target control instruction set being used to control the server to perform a preset fault processing operation. Through the present application, the problem of no memory fault detection and corresponding fault processing measure method for the integrated PMIC chip in the related art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of servers, and more specifically, to a server fault handling method, storage medium, and electronic device. Background Technology

[0002] In related technologies, the power supply module (PMIC chip) is integrated into the memory. The PMIC chip on the memory refers to a power management integrated circuit used to manage and control the power supply to the memory. Specifically, the memory PMIC chip is mainly responsible for providing a stable power supply, managing and monitoring the power supply to the memory, and protecting the memory from power fluctuations and overloads. However, existing technologies have not yet proposed an effective solution for how the aforementioned PMIC chip can perform real-time fault reporting and fault handling. Summary of the Invention

[0003] This application provides a server fault handling method, storage medium, and electronic device to at least solve the problem in the related art of the lack of fault detection and corresponding fault handling measures for memory with integrated PMIC chips.

[0004] According to one embodiment of this application, a server fault handling method is provided, comprising: acquiring first time information, the first time information including at least one of a first moment and a second moment, the first moment being the moment when a first fault occurs, the second moment being the moment when a second fault occurs, the first fault being a fault detected by a processor through fault detection of the registers of a power management integrated chip, the processor being used to execute BIOS firmware, and the second fault being a fault obtained by a baseboard controller through fault detection of the registers of the power management integrated chip; acquiring second time information, the second time information being a third moment when the voltage regulator of the power management integrated chip is powered on; and generating a target control instruction set based on the first time information and the second time information, the target control instruction set being used to control the server to perform preset fault handling operations.

[0005] In an exemplary embodiment, generating a target control instruction set based on first time information and second time information includes: in response to the first time being earlier than the third time, generating a first control instruction in the target control instruction set, the first control instruction being used to determine fault information, generate a fault log, and shield the communication channel where the faulty memory is located, the fault information including at least the fault occurrence address, and the fault log being generated by the baseboard controller based on the fault information.

[0006] In one exemplary embodiment, generating a target control instruction set based on first time information and second time information includes: in response to the first time being later than the third time, generating a second control instruction in the target control instruction set, the second control instruction being used to perform a power-on reprogramming on the server.

[0007] In an exemplary embodiment, generating a target control instruction set based on first time information and second time information includes: in response to the second time being later than the third time, generating a third control instruction in the target control instruction set, the third control instruction being used to perform a shutdown of the server, a power-on of the server, and during the power-on process of the server, detecting whether a first fault has occurred and locating the slot where the first fault occurred, the third control instruction being used to shield the communication channel where the faulty memory is located when the first fault occurs.

[0008] In one exemplary embodiment, determining fault information includes: polling the registers of the power management integrated chip on the memory module using the processor to obtain a polling result; and, in response to the polling result including a register fault, locating the address where the fault occurred based on the polling result to obtain fault information.

[0009] In one exemplary embodiment, the method for detecting a second fault includes: acquiring memory information, the memory information including a voltage signal and a temperature signal, wherein the voltage signal is generated by a voltage regulator when a register fault occurs, the voltage signal is acquired by a complex programmable logic device inside the server, and the temperature signal is acquired by a temperature sensor on the memory module, the temperature signal being used to determine the temperature of the memory; determining whether the voltage signal is the same as a set voltage signal, wherein the set voltage signal is the level signal generated by the voltage regulator corresponding to the memory when the memory is normal; if the voltage signal is different from the set voltage signal, determining that the memory corresponding to the voltage signal has experienced a second fault; determining whether the temperature of the memory is greater than a set temperature value, wherein the set temperature value is the highest temperature allowed for normal operation of the memory; if the temperature of the memory is greater than the set temperature value, determining that the memory corresponding to the temperature signal has experienced a second fault.

[0010] In one exemplary embodiment, performing a power-on reprogramming on the server includes: controlling the baseboard controller to restart, and sending a power-on command to other target components within the server through the restarted baseboard controller.

[0011] According to yet another embodiment of this application, a computer-readable storage medium is also provided, in which a computer program is stored, wherein the computer program is configured to perform the steps in any of the above method embodiments when it is run.

[0012] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0013] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0014] By obtaining first-time information and second-time information, and controlling the server to perform preset fault handling operations (such as alarms) based on the first-time information and second-time information, the technical effect of handling faults according to different types of faults and different times of occurrence is achieved, making fault handling more timely and intelligent, and minimizing the impact on the server. In related technologies, there is currently no fault detection and corresponding fault handling measures for memory with integrated PMIC chips. Attached Figure Description

[0015] Figure 1 This is a hardware structure block diagram of a server device for a server fault handling method according to an embodiment of this application;

[0016] Figure 2 This is a flowchart of a server fault handling method according to an embodiment of this application;

[0017] Figure 3 This is a PMIC fault reporting system architecture diagram according to an embodiment of this application;

[0018] Figure 4 This is a hardware diagram of BMC detecting memory power signals according to an embodiment of this application;

[0019] Figure 5 This is a flowchart of a server fault handling method according to an embodiment of this application;

[0020] Figure 6 This is a flowchart of a server fault handling method according to an embodiment of this application. Detailed Implementation

[0021] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.

[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0023] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a server device for a server fault handling method according to an embodiment of this application. For example... Figure 1 As shown, the server device may include one or more ( Figure 1Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0024] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the server fault handling method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the server device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0025] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0026] Specifically, the existing terms appearing in this application are explained as follows: NGN: Next Generation Network. DDR5: Double Data Rate 5 (SDRAM) technology. PMIC: Power Management Integrated Circuit. CPLD: Complex Programmable Logic Device. TSOD: Temperature sensor on DIMM. RCD: Registering Clock Drive. BMC: Baseboard Management Controller. VR: Voltage Regulator. BIOS: Basic Input and Output System. PSU: Power Supply Unit. MUX: Multiplexer. POST: Power On Self-Test. PECI: Platform Environment Control Interface.

[0027] A PMIC (Power Management Integrated Circuit) chip is an integrated circuit that integrates multiple power management functions. It is commonly used in electronic devices to manage and control the power supply, providing a stable power source to various components and subsystems. PMIC chips can perform functions such as power switching, voltage regulation, current control, battery charge / discharge management, and temperature monitoring to ensure the normal operation and safety of the device. The design and selection of PMIC chips have a significant impact on the power consumption, efficiency, and performance of electronic devices.

[0028] Prior to DDR5 (Double Data Rate 5 SDRAM technology), memory power was supplied by a separate memory power supply circuit on the motherboard. Starting with DDR5, the power supply module (PMIC chip) is integrated into the memory module. The PMIC chip on DDR5 memory refers to the power management integrated circuit used to manage and control the power supply to the DDR5 memory. Specifically, the DDR5 memory PMIC chip is mainly responsible for providing a stable power supply, managing and monitoring the power of the DDR5 memory, and protecting the DDR5 memory from power fluctuations and overloads. Different DDR5 memory products may use different PMIC chips; common PMIC chip manufacturers include Infineon, Texas Instruments, and Analog Devices.

[0029] I3C is a new type of serial bus interface, short for Improved Inter-Integrated Circuit. Developed by the MIPI (Mobile Industry Processor Interface) alliance, it aims to replace I2C (Inter-Integrated Circuit) and SMBus (System Management Bus) interfaces. While maintaining backward compatibility with I2C devices, the I3C interface offers higher data transfer rates, lower power consumption, and more functionality. It can be used to connect various sensors, memories, and other peripherals, and is widely used in mobile devices, the Internet of Things (IoT), automotive electronics, and other fields. The speed of the I3C interface varies depending on the specific implementation and device support. Generally, the maximum speed of the I3C interface can reach 12.5 Mbps, several times faster than the traditional I2C interface. Furthermore, I3C supports multi-master parallel transmission and high-speed data modes, which can further improve data transfer rates.

[0030] This embodiment provides a server fault handling method. Figure 2 This is a flowchart of a server fault handling method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0031] Step S102: Obtain first time information, which includes at least one of a first moment and a second moment. The first moment is the time when the first fault occurs, and the second moment is the time when the second fault occurs. The first fault is a fault detected by the processor through fault detection of the registers of the power management integrated chip. The processor is used to execute BIOS firmware. The second fault is a fault obtained by the baseboard controller through fault detection of the registers of the power management integrated chip. Fault detection includes various types of detection, such as monitoring the voltage, current, and temperature values ​​of each module of the PMIC. When at least one of the above parameters exceeds a set threshold, an alarm is triggered and a fault is detected.

[0032] Step S104: Obtain the second time information, which is the third moment when the voltage regulator of the power management integrated chip is powered on.

[0033] Step S106: Generate a target control instruction set based on the first time information and the second time information. The target control instruction set is used to control the server to perform preset fault handling operations.

[0034] By obtaining first and second time information and controlling the server to perform preset fault handling operations based on the first and second time information, this application achieves the technical effect of handling faults according to different types of faults and different times of occurrence, making fault handling more timely and intelligent, and minimizing the impact on the server. In related technologies, there is currently no fault detection and corresponding fault handling measures for memory with integrated PMIC chips.

[0035] The entities that perform the above steps can be servers, terminals, etc., but are not limited to these.

[0036] In related technologies, when a PMIC fails, the server will shut down directly. Since DDR5 integrates the PMIC chip into the memory for the first time, there is currently no intuitive method for fault alarm and fault handling. Based on this, this application proposes an error reporting method and fault handling method when a PMIC fails, which allows users to easily locate the specific memory and handle the fault.

[0037] The execution order of steps S102 and S104 can be interchanged; that is, step S104 can be executed first, and then step S102 can be executed.

[0038] like Figure 3The diagram illustrates the PMIC fault reporting system architecture. All devices on the DDR5 DIMM (Power Management IC [PMIC], Thermal Sensor [TS] / Temperature Sensor [TSOD], Register Clock Driver [RCD]) are connected to the SPD HUB via a local I3C* bus. The SPD HUB is connected to the CPU via the I3C* bus. All local DIMM devices (including the SPD HUB) are visible as terminal devices on the SPD I3C* bus. The SPD controller in the CPU is the I3C* master, allowing the CPU to access the PMIC on each DIMM via I3C.

[0039] like Figure 4 The diagram shows the hardware diagram for the BMC to detect memory power signals. In current Intel platform hardware designs, the POWER GOOD signals of every four DIMMs are connected together to form a single POWER GOOD signal, which is then connected to the CPLD. The CPLD is in turn connected to the BMC, allowing the BMC to detect the memory's POWER GOOD signal. When the memory is functioning normally, the memory's POWER GOOD signal is high. When the PMIC malfunctions, the POWER GOOD signal is pulled low. The BMC can detect this low signal, but it cannot determine which specific memory module is receiving the signal.

[0040] The PMIC specification defines 11 events that trigger internal VR disable, meaning a PMIC failure will trigger memory VR disable. When VR disable is triggered, the memory's POWER GOOD signal is pulled low, and the CPLD detects the abnormal memory voltage, at which point the CPLD shuts down the server. When the server is running normally, if a memory DIMM's PMIC fails, the server will immediately shut down. If a command is sent or the power button is pressed manually to power on at this point, the server will fail to boot because the memory's POWER GOOD fail state is not cleared, violating normal power-on logic. Therefore, the server will shut down again at some point during the power-on process. In the current system design, because the power good signals of the four memory DIMMs are connected together, the BMC cannot detect and locate which DIMM has experienced a PMIC failure. However, by detecting the first fault, the CPU can perform precise detection of the PMIC fault address.

[0041] In specific embodiments of this application, PMIC failures fall into two categories. The first is during the initial power-on phase, before the BIOS enables memory VR (BIOS firmware enables memory voltage regulators). In this case, the BIOS isolates and reports faulty memory information. The second is after the BIOS enables memory VR, i.e., later in the power-on process and while running under the OS. In this case, the system first shuts down and then performs an AC reboot, isolating and reporting faulty memory information during the power-on process. Enabling memory VR means that the memory is powered on normally.

[0042] In an exemplary embodiment, generating a target control instruction set based on first time information and second time information includes: in response to the first time being earlier than the third time, generating a first control instruction in the target control instruction set, the first control instruction being used to determine fault information, generate a fault log, and shield the communication channel where the faulty memory is located, the fault information including at least the fault occurrence address, and the fault log being generated by the baseboard controller based on the fault information.

[0043] In one exemplary embodiment, generating a target control instruction set based on first time information and second time information includes: in response to the first time being later than the third time, generating a second control instruction in the target control instruction set, the second control instruction being used to perform a power-on reprogramming on the server.

[0044] like Figure 6 The diagram shows a flowchart of an example of a failure occurring before memory VR enable during the boot process. Step 1: AC power on, server boots up, PMIC enters configuration mode, which is non-write-protected mode.

[0045] Step 2: The BIOS polls the PMIC registers of all DIMMs via I3C to locate the DIMM with the faulty PMIC. Because error information is recorded in non-volatile registers after a PMIC failure, checking these registers can help determine if a PMIC is faulty.

[0046] Step 3: The BIOS sends the specific information about the memory module that experienced the PMIC failure, along with the PMIC register information, to the BMC. Upon receiving this information, the BMC generates a log and issues an alarm. This allows the user to pinpoint the specific memory slot and replace the memory module.

[0047] Step 4: Configure the faulty PMIC register in the BIOS to clear the power good exception status and disable the channel where the faulty memory is located, thus isolating the faulty memory.

[0048] Step 5: Enable VR on all normal memory modules to power on the memory normally.

[0049] Step 6: Initialize all normal memory. After initialization, set the PMIC of all memory to write-protected mode.

[0050] Step 7: The BIOS polls the PMIC registers of all memory modules again to check for any remaining memory PMIC faults. If a fault is found, information about the faulty memory is sent to the BMC, and the BMC performs an AC reboot, starting from step one again. If there are no faults, the system continues booting.

[0051] In an exemplary embodiment, generating a target control instruction set based on first time information and second time information includes: in response to the second time being later than the third time, generating a third control instruction in the target control instruction set, the third control instruction being used to perform a shutdown of the server, a power-on of the server, and during the power-on process of the server, detecting whether a first fault has occurred and locating the slot where the first fault occurred, the third control instruction being used to shield the communication channel where the faulty memory is located when the first fault occurs.

[0052] In DDR5 memory, multiple memory modules (DIMMs) typically share one or more data channels. If the PMIC of a particular DIMM fails, for system stability and performance reasons, it's necessary to disable (stop using) the channel (communication channel) containing the failed DIMM, rather than the entire memory system. This is because: unstable power supply to the failed memory can lead to data read / write errors; disabling the channel prevents data corruption or system instability. Although disabling a channel reduces available memory bandwidth, compared to a complete shutdown or restart, this method can maintain system operation to some extent, avoiding greater performance loss. Disabling the failed memory channel reduces unnecessary power load, contributing to system thermal management and energy efficiency. Through these operations, the system can safely isolate and handle power supply failures in DDR5 memory while minimizing the impact on system performance, ensuring the continuous and stable operation of the server or computer.

[0053] DIMM is an abbreviation for "Dual In-line Memory Module," a standard for memory modules used in computers and servers. In DDR5 technology, DIMM refers to a memory module that conforms to the DDR5 specification. It not only contains DRAM chips (Dynamic Random Access Memory) for storing data but also integrates other important components and circuitry to achieve higher performance, lower power consumption, and better signal integrity.

[0054] like Figure 5The diagram illustrates the process of PMIC failure after memory VR is enabled. Step 1: After a PMIC failure, the PMIC triggers internal VR disable, and the POWER GOOD signal is pulled low.

[0055] Step 2: The CPLD detects an abnormal memory POWER GOOD signal, the server shuts down, and the BMC reports a memory POWER GOOD fail log. Normally, the POWER GOOD signal is high; a low POWER GOOD signal indicates a power failure.

[0056] Step 3: After the BMC detects that a certain memory module has a POWER GOOD error, it actively sends a command to the power supply to execute ACreboot, which means powering down the entire server and then powering it back on.

[0057] The reason for AC rebooting is as follows: During normal operation, the BIOS sends a VR Enable command to the DIMMs during the boot process, powering on the memory and then performing some initialization work on the PMIC. After the initialization work is completed, the BIOS sets all DIMM PMICs to a "write-protected" state. When a PMIC fails, if the system is powered off and then powered on again without interruption, the POWER GOOD signal of the memory will remain low. To remove this state, an AC reboot is required to remove the "write-protected" state from the PMIC, allowing the BIOS to configure the PMIC registers to isolate the faulty memory.

[0058] Step 4: After the BMC restarts, send the power on command to power on the server.

[0059] When AC reboot is executed, all power to the motherboard on the server will be restored, and the BMC will also restart. After the BMC restarts, if it detects any logs indicating that abnormal memory voltage caused the server to shut down before the restart, it will send a power-on command to the server.

[0060] Step 5: During the boot process, after the BIOS detects a PMIC fault, it isolates the corresponding memory and reports the specific memory and PMIC register information to the BMC. The BMC records this in its log and issues an alert. The user can then obtain the specific error information from the BMC and perform actions such as replacing the memory.

[0061] In one exemplary embodiment, determining fault information includes: polling the registers of the power management integrated chip on the memory module using the processor to obtain a polling result; and, in response to the polling result including a register fault, locating the address where the fault occurred based on the polling result to obtain fault information.

[0062] Specifically, polling is a computer communication and data acquisition method in which the central processing unit (CPU) or other master device periodically or continuously checks the status or data of slave devices (such as sensors, memory, peripherals, etc.) to determine if there is any new information or event that needs to be processed. In polling mode, the master device actively initiates queries instead of waiting for slave devices to send interrupt signals to notify of status changes.

[0063] During operation, the CPU periodically accesses the temperature sensor (TSOD) on the memory to read the current temperature information. The TSOD is a temperature sensor embedded in the DDR memory module, used to monitor the memory temperature and prevent overheating that could lead to performance degradation or hardware damage.

[0064] To optimize resource utilization and avoid bus conflicts between the CPU and BMC, a MUX (multiplexer) is used to switch control of the I3C bus at different stages, allowing the CPU to poll during the POST stage and the BMC to perform real-time monitoring and polling during the OS runtime stage.

[0065] In this way, both the CPU and BMC can effectively monitor memory temperature, but the BMC's monitoring is more real-time and detailed, and can respond to temperature anomalies in a timely manner without interrupting normal system operation, thereby improving system stability and security.

[0066] In one exemplary embodiment, the method for detecting a second fault includes: acquiring memory information, the memory information including a voltage signal and a temperature signal, wherein the voltage signal is generated by a voltage regulator when a register fault occurs, the voltage signal is acquired by a complex programmable logic device inside the server, and the temperature signal is acquired by a temperature sensor on the memory module, the temperature signal being used to determine the temperature of the memory; determining whether the voltage signal is the same as a set voltage signal, wherein the set voltage signal is the level signal generated by the voltage regulator corresponding to the memory when the memory is normal; if the voltage signal is different from the set voltage signal, determining that the memory corresponding to the voltage signal has experienced a second fault; determining whether the temperature of the memory is greater than a set temperature value, wherein the set temperature value is the highest temperature allowed for normal operation of the memory; if the temperature of the memory is greater than the set temperature value, determining that the memory corresponding to the temperature signal has experienced a second fault.

[0067] In one exemplary embodiment, powering back the server includes: restarting the control board controller and sending power-on commands to other target components within the server via the restarted control board controller. In related technologies, servers can only print PMIC fault information in the BIOS serial port, without alarming the BMC. Furthermore, due to hardware design limitations, it cannot accurately pinpoint which memory voltage is abnormal. This application, without modifying the existing hardware design, can accurately report PMIC fault information through the cooperation of the BIOS and BMC, automatically shutting down when a PMIC fault occurs. If a command is sent at this time or the power button is manually pressed to turn on the server, it will fail to boot normally. To restore normal operation, the AC must be manually disconnected. This application automatically performs an AC reboot after detecting an abnormal memory POWER GOOD signal. In a data center, this eliminates the need for maintenance personnel to manually unplug the power cord for an AC reboot, simplifying maintenance.

[0068] Using the technical solution of this application, after the BMC detects a memory POWER GOOD error causing a shutdown, it sends a command to the PSU to perform an AC reboot. After the AC is powered on, the BIOS locates the PMIC faulty memory information and sends detailed information about the faulty memory to the BMC. Upon receiving the PMIC faulty memory information from the BIOS, the BMC logs the information, issues an alarm, and notifies the user to replace the memory.

[0069] The detection methods for the first and second faults are different. The first fault is detected by the processor, while the second fault is detected by the BMC. In an optional embodiment, the BIOS (Basic Input / Output System): During the power-on phase, the BIOS is mainly responsible for hardware initialization and basic function testing. It reads the PMIC registers via the I3C bus to check for fault records. Once a fault is detected, the BIOS takes action, such as isolating the faulty memory channel to prevent instability during system power-on. The BIOS also sets the PMIC registers to write-protected mode to prevent unauthorized access. POST (Power-On Self-Test) is a series of hardware checks executed by the BIOS after the power-on phase to verify the basic functionality of the server hardware, such as the CPU, memory, and hard drive. During the POST phase, the BIOS firmware is executed by the processor to access the PMIC registers via the I3C bus, check the PMIC status, identify and handle potential hardware faults, such as PMIC failures or memory overheating. The BIOS, short for Basic Input / Output System, is a program stored in a ROM (Read-Only Memory) chip on the server motherboard. It contains a series of basic hardware initialization codes and self-test programs, and is the first software to run when a PC system starts up. The main responsibilities of the BIOS include: identifying and initializing hardware resources in the system, including the CPU, memory, disk, and display adapter. The BIOS performs the POST (Power-On Self-Test) process to check the system hardware and verify that it is in normal working order. If a hardware fault is detected, the BIOS will notify the user through a specific error code or audible signal. If the POST process passes, the BIOS will read the boot program (usually the operating system's boot loader) from the preset boot device (such as the hard drive) and transfer control to the boot program, thereby starting the operating system.

[0070] The CPU is the "brain" of a computer, responsible for executing instructions and performing calculations. During the POST (Power-On Self-Test) phase, the CPU's main role is as follows: The CPU reads the BIOS program from the motherboard ROM and executes its initialization and self-test code. The CPU accesses and tests hardware resources according to BIOS instructions, such as reading the PMIC registers via the I3C bus to check the memory's power management status. If the BIOS detects a hardware fault during POST, the CPU will follow the BIOS's instructions, potentially halting the boot process or skipping the initialization of the faulty hardware, allowing the system to continue booting to a usable state.

[0071] BMC (Baseboard Management Controller): After POST (Power-On Post-Processing) completes and enters the OS phase, the BMC takes over the server's hardware monitoring responsibilities. This includes real-time monitoring of the PMIC's status registers via the I3C bus and monitoring the temperature information from the TSOD (Temperature Sensor). The BMC's monitoring focuses on real-time and continuous operation. It continuously checks the PMIC's voltage, current, and temperature, and immediately records and alarms any exceeding of thresholds. If necessary, it can also notify the CPU to implement temperature-related protection measures via the PECI interface.

[0072] Optionally, the BMC can switch control of the PMIC on the DIMM via a multiplexer and CPU.

[0073] The MUX (Multiplexer) acts as a switcher for communication channels, allowing the CPU or BMC to control the I3C bus at different times to access devices on the DIMM. Specifically:

[0074] One of the ports of the MUX connects to the CPU on the motherboard. During the POST (Power On Self-Test) phase, the CPU uses this port to access devices such as the PMIC and TSOD on the DIMM for initialization, configuration, and fault detection.

[0075] Another port of the MUX connects to the BMC (Baseboard Management Controller). During OS (Operating System) operation, control switches from the CPU to the BMC. The BMC uses this port to monitor the PMIC status on the DIMM in real time, including parameters such as voltage, current, and temperature, to ensure the stability and security of the server hardware during operation.

[0076] This design allows the CPU and BMC to use the I3C bus at different times, avoiding communication conflicts that may occur when they simultaneously access devices on the DIMM. It also ensures that devices on the DIMM can be efficiently and in real-time monitored and managed. Controlling the switching of the MUX via a CPLD (Complex Programmable Logic Device) further enhances the system's flexibility and reliability in managing DIMM devices.

[0077] The CPU detects PMIC faults via the I3C component, while the BMC can either obtain the PMIC's VR power-good signal through the CPLD or communicate with the SPD HUB via the I3C component. The SPD HUB, or SerialPresence Detect Hub, is a key component on the DDR5 DIMM used to manage communication between various devices connected to the memory module via the I3C bus. The SPD HUB functions somewhat like a small communication management center; it receives data from various devices on the DIMM (such as PMIC, TSOD, RCD, etc.), integrates this data, and sends it to the CPU or BMC on the motherboard via the I3C bus for processing.

[0078] On DDR5 DIMMs, the SPD HUB connects to these devices via a local I3C bus. "Local" here means that these connections and communications are limited to within the DIMM module. The I3C bus is an improved version of the I2C bus, offering higher data transfer rates and richer functionality, making it ideal for high-speed communication between devices within the module.

[0079] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0080] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.

[0081] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0082] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0083] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0084] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0085] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0086] The embodiments described herein also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.

[0087] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0088] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0089] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A method for handling server faults, characterized in that, include: Obtain first-time information, which includes at least one of the following: a first moment, which is the moment when a first fault occurs, the first fault including a fault detected by the processor through fault detection of the registers of the power management integrated chip, the processor being used to execute BIOS firmware; and a second moment, which is the moment when a second fault occurs, the second fault including a fault obtained by the baseboard controller through fault detection of the registers of the power management integrated chip. Acquire the second time information, which is the third moment when the voltage regulator of the power management integrated chip is powered on; A target control instruction set is generated based on the first time information and the second time information. The target control instruction set is used to control the server to perform preset fault handling operations. The generation of the target control instruction set based on the first time information and the second time information includes at least one of the following: If the first time information includes the first moment, in response to the first moment being earlier than the third moment, a first control instruction in the target control instruction set is generated. The first control instruction is used to determine fault information, generate a fault log, and shield the communication channel where the faulty memory is located. The fault information includes at least the fault occurrence address, and the fault log is generated by the baseboard controller based on the fault information. If the first time information includes the second time moment, in response to the second time moment being later than the third time moment, a third control instruction in the target control instruction set is generated. The third control instruction is used to perform a shutdown of the server, a power-on of the server, and to detect whether the first fault occurs during the power-on process and locate the slot where the first fault occurs. The third control instruction is also used to shield the communication channel where the faulty memory is located when the first fault occurs during the power-on process.

2. The method according to claim 1, characterized in that, The method further includes: If the first time information includes the first moment, in response to the first moment being later than the third moment, a second control instruction is generated in the target control instruction set, the second control instruction being used to perform a power-on reprogramming on the server.

3. The method according to claim 1, characterized in that, Determine fault information, including: The processor polls the registers of the power management integrated chip on the memory module to obtain the query results; In response to the query result including a register fault, the address where the fault occurred is located based on the query result to obtain the fault information.

4. The method according to claim 1, characterized in that, The detection method for the second fault includes: The system acquires memory information, which includes voltage and temperature signals. The voltage signal is generated by a voltage regulator when a register malfunctions. The voltage signal is acquired by a complex programmable logic device inside the server. The temperature signal is acquired by a temperature sensor on the memory module. The temperature signal is used to determine the temperature of the memory. Determine whether the voltage signal is the same as the set voltage signal, wherein the set voltage signal is the level signal generated by the voltage regulator corresponding to the memory when the memory is normal; If the voltage signal is different from the set voltage signal, it is determined that the memory corresponding to the voltage signal has experienced a second fault; Determine whether the temperature of the memory is greater than a set temperature value, wherein the set temperature value is the highest temperature allowed for normal operation of the memory; If the memory temperature exceeds the set temperature value, a second memory fault is determined to have occurred corresponding to the temperature signal.

5. The method according to claim 1 or 2, characterized in that, Performing a power-on process on the server includes: controlling the baseboard controller to restart, and sending a power-on command to other target components within the server through the restarted baseboard controller.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 5.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 5.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method, system and device for monitoring power management chip and medium

    CN117707884A

  • PMIC power supply fault processing method and device, computer equipment and storage medium

    CN118034983A