Fault handling method and electronic device
Patent Information
- Application Number
- CN202610958119.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-30
AI Technical Summary
[0003]本申请提供了一种故障处理方法及电子设备,以解决相关技术中难以精准识别第二控制器的启动失败故障,并实现自动修复的技术问题
[0007] The fault handling method provided in this application, after confirming that the board containing the second controller is in place, sends a communication command to the board identification storage unit on the second controller; this avoids the first controller performing invalid I2C communication operations on empty slots or boards that are not inserted. After the first controller obtains the response data, it sends a read command to the upgrade address of the second controller to ensure that the communication link of the board identification storage unit is normal. In response to the first controller obtaining the upgrade address, the first controller sends a read command to the communication address of the second controller; in response to the first controller not obtaining the communication address, it is determined that the second controller has a startup failure fault. The first controller writes a repair command to the upgrade address of the second controller. The repair command is used to trigger the second controller to re-execute the firmware loading process. By detecting the reachability of the upgrade address and the communication address, the communication address loss fault caused by the abnormal termination of the startup process of the second controller is accurately distinguished from other types of faults, avoiding misjudgment and misoperation. During fault repair, the first controller writes a repair command to the upgrade address to trigger the second controller to re-execute the firmware loading process, achieving service-free fault repair without cutting off the chip power supply or affecting the operation of peripherals.
Smart Images

Figure CN122470428B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fault handling technology, and in particular to a fault handling method, electronic device, storage medium and program product. Background Technology
[0002] When a Complex Programmable Logic Device (CPLD) terminates its Self-Loading Mode (SDM) process due to an abnormal I2C level, a CPLD communication address loss fault occurs, leading to the hard drive not powering on and the system failing to boot. Related technologies typically attribute this type of fault simply to a complete CPLD failure or an I2C bus malfunction, addressing it by resetting the entire device or replacing the board. However, this approach results in low fault response efficiency and a high risk of service interruption. Currently, there is no precise way to identify this type of fault, and automatic repair methods are lacking. Summary of the Invention
[0003] This application provides a fault handling method and electronic device to solve the technical problem in the related art of accurately identifying the startup failure fault of the second controller and achieving automatic repair.
[0004] This application provides a fault handling method applied to a first controller, the fault handling method comprising: After the first controller starts up, check whether the board containing the second controller is in place; In response to the presence of the board containing the second controller, a communication command is sent to the board identifier storage unit on the second controller; In response to the received communication command response data, a read command is sent to the upgrade address of the second controller; In response to obtaining the upgrade address, a read command is sent to the communication address of the second controller; In response to the failure to obtain a communication address, it is determined that the second controller has a startup failure fault. A repair instruction is written to the upgrade address of the second controller. The repair instruction is used to trigger the second controller to re-execute the firmware loading process.
[0005] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the fault handling method described in the following embodiments.
[0006] After the first controller starts up, check whether the board containing the second controller is in place; In response to the presence of the board containing the second controller, a communication command is sent to the board identifier storage unit on the second controller; In response to the received communication command response data, a read command is sent to the upgrade address of the second controller; In response to obtaining the upgrade address, a read command is sent to the communication address of the second controller; In response to the failure to obtain a communication address, it is determined that the second controller has a startup failure fault. A repair instruction is written to the upgrade address of the second controller. The repair instruction is used to trigger the second controller to re-execute the firmware loading process.
[0007] The fault handling method provided in this application, after confirming that the board containing the second controller is in place, sends a communication command to the board identification storage unit on the second controller; this avoids the first controller performing invalid I2C communication operations on empty slots or boards that are not inserted. After the first controller obtains the response data, it sends a read command to the upgrade address of the second controller to ensure that the communication link of the board identification storage unit is normal. In response to the first controller obtaining the upgrade address, the first controller sends a read command to the communication address of the second controller; in response to the first controller not obtaining the communication address, it is determined that the second controller has a startup failure fault. The first controller writes a repair command to the upgrade address of the second controller. The repair command is used to trigger the second controller to re-execute the firmware loading process. By detecting the reachability of the upgrade address and the communication address, the communication address loss fault caused by the abnormal termination of the startup process of the second controller is accurately distinguished from other types of faults, avoiding misjudgment and misoperation. During fault repair, the first controller writes a repair command to the upgrade address to trigger the second controller to re-execute the firmware loading process, achieving service-free fault repair without cutting off the chip power supply or affecting the operation of peripherals. Attached Figure Description
[0008] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A flowchart illustrating a fault handling method provided in an embodiment of this application; Figure 2 This is a schematic diagram of hardware storage isolation for a second controller provided in an embodiment of this application; Figure 3 A flowchart illustrating the power-on steps of a server as provided in an embodiment of this application; Figure 4 A schematic diagram of the communication hardware between the first controller and the second controller provided in an embodiment of this application; Figure 5 A schematic diagram of a fault log provided in an embodiment of this application; Figure 6 A schematic diagram of a fault log provided for another embodiment of this application; Figure 7A flowchart illustrating a fault handling method provided in another embodiment of this application; Figure 8 This is a schematic diagram of the structure of a fault handling device provided in an embodiment of this application; Figure 9 This is an internal structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0011] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0012] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0013] In computer hardware systems such as servers and storage devices, complex programmable logic devices (CPLDs) are core components for board communication, power supply control, and hardware status monitoring. Key boards such as backplanes, switch boards (SW boards), and storage backplanes (SDB boards) are equipped with CPLDs and communicate with the baseboard management controller (BMC) via the I2C bus (a serial communication bus commonly used for inter-chip communication). During AC power-on (i.e., power-off and power-on of the entire system), the CPLD needs to complete the SDM (self-download mode) process to load firmware. Subsequently, it opens preset communication addresses (such as 0x10 / 0x11 / 0x18) to interact with the BMC. Simultaneously, the CPLD reserves a dedicated UFM (user flash memory) upgrade address (such as 0x40) for firmware upgrade operations.
[0014] In practical applications, servers generally adopt a distributed independent low-dropout linear regulator (LDO) 3.3V standby power supply architecture. There is no unified power-on synchronization enable signal between the various boards, and the power-on time difference between the backplane and the switching board is stable between 2 and 8 milliseconds. At the moment of power-on, the I2C bus level floats in the middle range of 0.3V to 1.2V. Since the CPLD has no analog level filtering hardware inside, the CPLD will misinterpret the middle level for more than 200 nanoseconds as an I2C start signal, thus forcibly interrupting the SDM firmware self-loading process.
[0015] The CPLD's internal storage is divided into a hardware-isolated Configuration Flash (CFM) running partition and a User Flash (UFM) dual-sector partition. An SDM interrupt only causes the CFM running partition to fail to initialize and the regular I2C communication addresses (such as 0x10 / 0x18) to become unreachable, but the UFM hardware bus is unaffected, and the upgrade address (0x40) remains permanently readable and writable.
[0016] A single peripheral board contains two independent I2C devices: a Board Identifier Memory Unit (FRU) chip and a CPLD. Their power supply is completely isolated from the I2C bus. The FRU chip is used solely to store the board's static identification information (such as serial number and part number), and is unaffected by power-on timing interference. It is implemented using a standard electrically erasable programmable read-only memory (EEPROM) chip. However, related technologies do not utilize the dual-presence identification mechanism of FRU combined with General Purpose Input / Output (GPIO), nor do they utilize a dual-channel hierarchical approach of FRU and UFM for fault differentiation.
[0017] To address such faults, relevant technologies typically employ methods such as hardware watchdog reset, BMC polling reset, or firmware upgrade reload. For example, a separate hardware watchdog timer chip can be deployed externally to the CPLD. During normal CPLD operation, the watchdog is periodically fed; if the CPLD stops working due to a fault, a hardware reset signal is triggered to force the CPLD reset pin low to reload the firmware. Alternatively, the BMC periodically polls the CPLD's internal register status, and a reset pulse is sent via GPIO when consecutive read failures occur. Another approach is to send a reload command during the CPLD firmware upgrade process to activate the new firmware. However, all of the above methods have significant drawbacks: the hardware watchdog cannot distinguish between a CPLD logic failure and a failure where the AC power-on simply fails to load. In the AC power-on failure scenario, the CPLD fails to function properly from start to finish and cannot perform the watchdog feeding operation, meaning the watchdog may never be activated in the first place. The BMC polling monitoring logic is too simple and will trigger a reset in any I2C communication failure, including BMC restarts, I2C bus transient interference, or CPLD performing a normal firmware upgrade. This can all lead to false triggers, and the reset operation usually causes the CPLD to completely restart, resulting in power-down and power-on failures of the peripherals it controls, which is unacceptable for running services. The firmware upgrade reloading technology is only applied to firmware upgrade scenarios, aiming to achieve seamless switching between old and new firmware or ensure immediate effectiveness after the upgrade. It does not address fault detection and repair triggering in the specific failure scenario of AC power-on failure. Furthermore, most of the above solutions are single-step logic, lacking hierarchical processing and self-verification mechanisms. If a reset fails once, more complex upper-level logic intervention or direct error reporting is often required.
[0018] To address the aforementioned technical issues, such as Figure 1 As shown, one embodiment of this application provides a fault handling method applied to a first controller, the method specifically including the following steps: S1: After the first controller starts up, check whether the board containing the second controller is in place.
[0019] S2: In response to the presence of the board containing the second controller, a communication command is sent to the board identifier storage unit on the second controller.
[0020] The first controller obtains the hardware version information of the board where the second controller is located; in response to the hardware version being the first version, the first controller obtains the level data of the presence signal pin of the board where the second controller is located, and determines whether the board where the second controller is located is in place by using the level data; in response to the hardware version being the second version, the first controller determines whether the board where the second controller is located is in place by sending a read command to the board identifier storage unit of the board where the second controller is located.
[0021] In response to a low-level signal, the first controller determines that the board containing the second controller is in place and sends a communication command to the board identifier storage unit on the second controller; or, in response to the first controller receiving response data from the board identifier storage unit that sent a read command, the first controller determines that the board containing the second controller is in place and sends a communication command to the board identifier storage unit on the second controller.
[0022] In response to a high-level data level; or, if the first controller does not receive response data from the board identifier storage unit sending a read command, it is determined that the board containing the second controller is not in place, and the repair process for the second controller is terminated.
[0023] In this application, the first controller may be a baseboard management controller (BMC) and the second controller may be a complex programmable logic controller (CPLD).
[0024] Hardware version information indicates the hardware design iteration status of the board. Different versions of the board may use different hardware designs, such as whether or not a dedicated GPIO in-situ signal pin is configured. In this application, the main difference between the first version and the second version is whether or not a GPIO in-situ signal pin is configured. The first version has a GPIO in-situ signal pin configured, while the second version does not.
[0025] For boards equipped with a dedicated presence signal pin, the first controller determines board presence by reading the pin's level data. Specifically, in the hardware design, the backplane side connects the board presence signal line to the power supply via a pull-up resistor, while the board side connects the signal line to ground via a pin. When the board is not inserted, the signal line is maintained at a high level through the pull-up resistor; after the board is inserted, the signal line is pulled low. The first controller reads the GPIO pin level; if it is low, the board is considered present; if it is high, the board is considered absent. For hot-swapping scenarios, the first controller delays for 50 to 100 milliseconds after detecting a level change before reading again to confirm signal stability and avoid misjudgment due to contact jitter during insertion and removal.
[0026] For older hardware or low-cost design boards that lack a dedicated presence signal pin, the first controller determines board presence by sending a read command to the board identification storage unit, such as the FRU EEPROM chip. Specifically, the first controller sends a read command to the FRU chip via the I2C bus to attempt to read the board identification information stored in the FRU chip. If the first controller successfully receives a response from the FRU chip, it proves that the I2C bus is active and the FRU chip is powered, thus determining that the board is present. If the first controller does not receive a response (e.g., read timeout or NACK received), it determines that the board is not present or that the board power supply / I2C bus is disconnected.
[0027] When the first controller determines that the board containing the second controller is in place, the first controller sends a communication command to the board identification storage unit. Specifically, the first controller sends the communication command to the board identification storage unit via the I2C bus. This communication command can be a read command, used to read the board identification information stored in the FRU chip, such as the board serial number, part number, model, and asset tag. If the first controller successfully obtains the response data returned by the FRU chip, it proves that the FRU chip and the I2C link are healthy. By verifying whether the I2C communication link of the board identification storage unit is normal, it distinguishes between two states: the board is in place but there is a power supply or bus fault, and the board is completely normal. If the communication of the board identification storage unit is abnormal, it indicates that even if the board is physically in place, there is still a problem with its power supply or I2C bus. CPLD diagnosis should not continue; instead, the fault should be marked and the process terminated. By completing the FRU communication verification before CPLD diagnosis, this application effectively avoids misjudging power supply or bus faults as CPLD faults.
[0028] In response to the first controller's failure to receive response data from the board identifier storage unit's communication command, a fault is determined in the board identifier storage unit, and the repair process for the second controller is terminated. Specifically, the following situations constitute failure to receive response data: the board identifier storage unit does not return an acknowledgment signal after the first controller issues a communication command; the first controller does not receive any return data within a preset time window, and the communication bus is in a suspended or blocked state; the data returned by the board identifier storage unit does not conform to the expected format or fails verification, and cannot be identified as valid response data. When any of the above situations occur, the first controller determines that the board identifier storage unit has a communication fault. Possible causes include: abnormal board power supply, such as unstable or insufficient standby power supply; physical circuit break in the I2C bus, such as poor connector contact or damaged lines; damage to the board identifier storage unit chip itself, such as a damaged or expired EEPROM chip. In this case, since the FRU communication link is unavailable, it cannot be confirmed whether the subsequent CPLD diagnostics are based on a valid hardware path, and the first controller terminates the repair process for the second controller without performing any subsequent operations. Thus, this application utilizes a hybrid in-situ identification mechanism combining general purpose input / output pins (GPIO) and board identifier storage units (FRU) to differentiate between faults such as empty slots, full board power supply or bus failures, and CPLD chip hardware damage. In multi-board scenarios, after terminating the repair process for the current board, the first controller continues processing the next board.
[0029] S3: In response to receiving the response data of the communication command, send a read command to the upgrade address of the second controller.
[0030] The first controller sends a probe read command to the upgrade address of the second controller via the communication bus; the probe read command is used to read the status information corresponding to the upgrade address of the second controller.
[0031] The upgrade address refers to a dedicated User Flash Memory (UFM) access address reserved by the CPLD manufacturer during chip design to support online firmware upgrades. The I2C response logic for this upgrade address is embedded in the CPLD hardware and does not depend on the user firmware loading status. Even if the CPLD's SDM (Self-Loading Mode) process is abnormally interrupted, causing the communication address to fail to load successfully, the upgrade address can still respond to I2C bus access. The communication address is used for regular communication with the second controller; the upgrade address is independent of the communication address and remains reachable even if the second controller's firmware loading process is abnormally interrupted.
[0032] The first controller sends a probe read command to the upgrade address of the second controller via the communication bus. This probe read command is used to read the status information corresponding to the upgrade address of the second controller. This command differs from the version number read command sent to the communication address; its purpose is to probe the response status of the upgrade address, rather than to obtain specific firmware version information. The status information refers to the response data returned by the user flash memory area corresponding to the upgrade address. The first controller determines whether the upgrade address is reachable by whether it can successfully obtain this response data.
[0033] If the first controller does not receive a response from the second controller to the probe / read command, the first controller resends the probe / read command to the upgrade address of the second controller via the communication bus after a preset time interval. If the number of times the probe / read command is sent to the upgrade address of the second controller via the communication bus reaches a preset threshold and the first controller receives a response from the second controller to the probe / read command, it is determined that the first controller has obtained status information. If the first controller does not obtain status information, it is determined that the second controller has a hardware failure, and the repair process for the second controller is terminated.
[0034] To improve detection accuracy, this application further includes a retry mechanism. If the first controller fails to acquire status information on its first attempt, it can wait for a preset time interval and then send a probe read command to the upgrade address of the second controller via the communication bus again. The preset time interval can be 1 second, 2 seconds, etc., and its specific value can be set according to actual needs. This application does not limit the specific value of the preset time interval. If the read fails again, the above retry operation is repeated until the read is successful or the number of retries reaches a preset threshold. The preset threshold can be 3 times, 4 times, etc., and its specific value can be set according to actual needs. This application does not limit the specific value of the preset threshold. If the status data is still not successfully acquired after reaching the preset threshold, the upgrade address of the second controller is determined to be unreachable. This retry mechanism can effectively avoid misjudging a single read failure as a CPLD fault due to momentary interference on the I2C bus, thus improving the accuracy of fault detection. If the number of retries is within the preset threshold range, and the first controller receives a response corresponding to the probe read command on any one of these retries, it indicates that the upgrade address is reachable.
[0035] If the first controller fails to obtain status information after a preset number of retries, it determines that the upgrade address of the second controller is unreachable. In this case, since the FRU chip can communicate normally but the UFM upgrade address is inaccessible, it indicates that the CPLD chip's UFM storage hardware or I2C path is physically damaged, rather than the firmware loading process being interrupted. The first controller determines that the second controller has a hardware failure, terminates the repair process for the second controller, and does not execute subsequent communication address detection and repair instruction writing. Through the above processing, this application effectively distinguishes between hardware damage and firmware loading failure, avoiding invalid firmware reloading operations when the chip is physically damaged.
[0036] S4: In response to obtaining the upgrade address, send a read command to the communication address of the second controller.
[0037] In response to the first controller obtaining status information, it is determined that the first controller has obtained the upgrade address; the first controller sends a version number read instruction to the communication address of the second controller through the communication bus, and the version number read instruction is used to read the version number data of the second controller.
[0038] In I2C bus communication, each slave device has a unique device address, which the master uses to communicate with a specific slave device. When a CPLD acts as an I2C slave device, it is configured with a fixed communication address, such as 0x10, 0x11, or 0x18, for routine control commands and data exchange with the BMC, such as reading the version number and obtaining status information.
[0039] Once the first controller obtains the status information from the upgrade address via a probe read command, it determines that the first controller has acquired the upgrade address, meaning the upgrade address is reachable. At this point, the first controller sends a version number read command to the second controller's communication address via the communication bus. This command reads the version number data of the second controller. The version number data refers to the firmware version identifier information stored in the second controller's internal register. It uniquely identifies the version of the firmware currently running on the CPLD and is typically represented by numbers or a combination of numbers and letters, such as "V1.0", "2.03", or "0x01". This data is written into the firmware image by the developers during firmware compilation and is stored in the CPLD's version number register after the firmware is loaded.
[0040] If the first controller successfully obtains the version number data, it indicates that the communication address is reachable, the CPLD's SDM process has been completed, the firmware has been loaded normally, and the second controller is in normal working condition.
[0041] S5: In response to the failure to obtain the communication address, it is determined that the second controller has a startup failure fault. A repair instruction is written to the upgrade address of the second controller. The repair instruction is used to trigger the second controller to re-execute the firmware loading process.
[0042] In response to the first controller failing to obtain the version number data of the second controller, it is determined that the first controller has not obtained the communication address, and the startup failure of the second controller is a communication address loss fault caused by an abnormal interruption of the self-loading mode process of the second controller.
[0043] To improve detection accuracy, this application further includes a retry mechanism. After the first controller fails to acquire the version number data from the second controller on its first attempt, it can wait for a preset time interval and then send a version number read command to the communication address of the second controller via the communication bus again. The preset time interval can be 1 second, 2 seconds, etc., and its specific value can be set according to actual needs; this application does not limit the specific value of the preset time interval. If the read fails again, the above retry operation is repeated until the read is successful or the number of retries reaches a preset threshold. The preset threshold can be 3 times, 4 times, etc., and its specific value can be set according to actual needs; this application does not limit the specific value of the preset threshold. If the version number data is still not successfully acquired after reaching the preset threshold, it is determined that the first controller has not acquired the communication address. This retry mechanism can effectively avoid misjudging a single read failure as a CPLD fault due to momentary interference on the I2C bus, thus improving the accuracy of fault detection.
[0044] If the number of times the version number read command is sent to the communication address of the second controller via the communication bus reaches a preset threshold and the first controller fails to obtain the version number data of the second controller, it is determined that the first controller has failed to obtain the communication address.
[0045] If the first controller fails to obtain the communication address after the aforementioned retry mechanism, it can be determined that the second controller has a startup failure because the communication address is unreachable while the upgrade address is reachable. This indicates that the SDM process was abnormally interrupted, preventing the user logic (communication address) from loading, while the upgrade address, being permanently stored in the hardware, continues to respond normally. If the first controller fails to obtain status information from the upgrade address, it indicates that the upgrade address is also unreachable. In this case, the second controller may be in a complete deadlock state or have an I2C bus fault, which is not due to an abnormal interruption of the SDM process. The first controller logs and terminates the test.
[0046] This application provides a schematic diagram of hardware isolation between configuration flash (CFM) and user flash (UFM) partitions in one embodiment. Figure 2 The hardware isolation structure of the internal storage partitions of the CPLD is shown.
[0047] like Figure 2 As shown, the internal storage of the CPLD is divided into two hardware-isolated partitions: a configuration flash memory partition and a user flash memory partition.
[0048] A flash memory partition is configured to store the CPLD's operating firmware. After the SDM process is completed, the firmware in this partition is loaded into the operating logic area, exposing regular I2C communication addresses (such as 0x10 / 0x11 / 0x18) for routine control command and data interaction with the BMC. If the SDM process is interrupted due to an abnormal I2C level, the CFM partition initialization fails, and the regular communication addresses become unavailable.
[0049] The user flash partition is an independent partition bound to an upgrade address. It stores the firmware image and can receive reload or upgrade commands. The I2C response logic of this partition is embedded in the CPLD hardware and does not depend on the completion status of the SDM firmware loading process. It is always accessible after power-on. Even if the SDM process is abnormally interrupted, causing the CFM partition initialization to fail, the hardware bus of the UFM partition is unaffected, and the upgrade address remains readable and writable.
[0050] The CFM partition and the UFM partition are completely isolated at the hardware level, and their I2C paths are independent of each other. SDM interrupts only affect the initialization of the CFM partition and the opening of regular addresses, and have no impact on the accessibility of the hardware paths and upgrade addresses of the UFM partition.
[0051] Based on the aforementioned hardware characteristic differences, the second controller exhibits different combinations of communication address and upgrade address acquisition states under different fault states, as detailed below: Under normal circumstances, after the entire device is connected to AC power, the standby power supply outputs, and the CPLD powers on first. After the CPLD power-on reset, the hardware logic forcibly switches to SDM initialization mode, suspends all I2C external responses, and begins firmware loading. The CPLD's internal logic accesses the local user flash memory, reads the pre-stored complete firmware program and logic configuration file, and performs integrity verification on the firmware. After successful verification, the firmware is loaded into the CPLD's running logic area, completing internal logic initialization. After successful firmware loading, the CPLD exits SDM self-loading mode and activates the board's preset communication address. At this time, the communication address responds normally to the BMC's version reading, status query, and control commands. The upgrade address remains resident and accessible for subsequent firmware upgrade and reload command reception. Subsequently, the BMC completes initialization, the motherboard main power supply powers on, the hard drive backplane power supply is enabled, and the server enters standby or power-on process. The normal process result is that both I2C addresses are accessible, and the device starts normally.
[0052] An abnormal interruption occurred during the SDM process: After the standby power-on, the CPLD started normally and entered SDM mode, beginning to read the UFM firmware. However, due to the power-on timing of the peripheral boards being ahead of schedule, the I2C bus generated an intermediate abnormal level that was neither high nor low. During SDM operation, the CPLD monitors the I2C bus in real time and mistakenly interprets the abnormal level as an I2C start signal or bus activity. The CPLD hardware logic triggers a protection mechanism, forcibly terminating the current SDM firmware loading process, exiting self-loading mode, and switching to a state waiting for external I2C commands. The critical result of this anomaly is that the firmware is not fully loaded, the running logic area is not initialized, and the communication address is permanently closed and unresponsive; while the UFM partition hardware path is normal, and the upgrade address can still be read and written normally. Ultimately, this manifests as the BMC not detecting the CPLD communication address, resulting in hard drive power failure, board disconnection, and other malfunctions.
[0053] However, if the CPLD completely fails or the I2C bus malfunctions, the entire CPLD chip fails. This can be due to power abnormalities, chip damage, or I2C bus faults such as the bus being pulled low or the clock line being short-circuited. In this case, both the communication address and the upgrade address become unreachable. This presents a combination of unreachable communication and upgrade addresses, significantly different from the specific faults mentioned above. Therefore, by setting the first controller to detect the communication address and upgrade address of the second controller, this application can accurately identify whether the communication address loss is caused by an abnormal interruption of the SDM process.
[0054] As described above, when a CPLD experiences a communication address loss, possible causes include abnormal SDM process interruption, complete CPLD deadlock, I2C bus failure, firmware corruption, and power supply abnormalities. Related technologies do not differentiate between these fault types, uniformly employing reset or restart as the solution. If the fault is caused by firmware corruption or a complete chip deadlock, a simple reset operation will not resolve the issue. If the fault is caused by I2C bus interference or the CPLD is performing a firmware upgrade, a reset operation would be an unnecessary false trigger. For example, when the BMC itself is restarting, there is transient interference on the I2C bus, or the CPLD is performing a firmware upgrade, the BMC polling scheme may trigger a reset due to read failure. More seriously, this type of reset is often a system-wide hardware reset, cutting off the CPLD power supply or pulling the reset pin low, causing controlled peripherals, such as hard drives, to lose power and then be powered on again. For servers running services, unexpected hard drive power loss may trigger RAID rebuilding, cached data loss, causing service interruption or even data corruption.
[0055] Furthermore, the lack of differentiation between fault types means a lack of effective diagnostic information. Maintenance personnel are faced with a vague alarm like "CPLD communication failed," unable to determine the specific cause of the fault. They are forced to resort to trial and error methods such as restarting or replacing boards, severely impacting fault response efficiency and system availability in large-scale data centers. Different fault causes require different repair strategies: abnormal SDM process interruption requires firmware reloading, corrupted firmware requires re-flashing, and complete CPLD deadlock requires checking the power supply or replacing hardware. Because related technologies cannot differentiate between fault types, they can only use a uniform reset method, failing to implement the most effective repair measures for different faults, resulting in low repair success rates and recurring faults.
[0056] In other words, the reload command in related technologies is applied to CPLD firmware upgrade scenarios. After the upgrade is completed, it is triggered manually or by management software to make the newly burned firmware effective, and it is usually executed under the premise that the communication address is reachable. The reset operation of some solutions is a hardware reset, which is achieved by powering off or pulling the reset pin low, which will cause the peripheral device to lose power. The repair command in this application is applied to the CPLD power-on startup failure scenario. It is automatically triggered after the BMC accurately identifies the abnormal termination of the SDM process and the loss of the communication address through dual address detection. It does not involve writing new firmware data, but only triggers the CPLD to re-execute the loading process of the existing firmware. The command is sent through the independent channel of the upgrade address, and it can still be executed in the fault state of the CPLD communication address being unreachable. It triggers a soft reset, does not cut off the CPLD power supply, and does not affect the power supply status of the controlled peripheral device, thus achieving non-destructive repair.
[0057] To ensure the reliability of writing repair commands to the upgrade address, the first controller can employ a tiered retry mechanism to send the recovery trigger command. Specifically, after a write failure, the first controller waits for a preset time interval, such as 1 second, and then writes the recovery trigger command to the upgrade address again. If the write fails again, the retry operation is repeated until the write is successful or the number of retries reaches a preset threshold, such as 3 times. If the write still fails after reaching the preset threshold, a fault log is recorded and the repair process ends; if any write is successful, the subsequent recovery verification process begins. By setting a tiered retry mechanism, write failures caused by momentary interference on the I2C bus can be effectively avoided, improving the reliability of the repair process.
[0058] The first controller writes a preset recovery trigger instruction to the upgrade address of the second controller. The recovery trigger instruction is a preset byte sequence. In response to the second controller detecting that the upgrade address has been written to the preset byte sequence, the second controller triggers a soft reset, reads firmware data from the preset firmware storage area, and loads the firmware data into the second controller's runtime storage area for execution.
[0059] The preset byte sequence can be "0x79 0 0", used to trigger the second controller to re-execute the firmware loading process. This instruction may have different names in different CPLD manufacturers (such as Refresh instruction, Reload instruction, or LSC_REFRESH instruction), but its function is to trigger the CPLD to re-execute the SDM (Self-Download Mode) firmware loading process, causing the CPLD to reread firmware data from the firmware storage area and load it into the runtime storage area. In this application, this instruction is used to force the CPLD to re-execute the loading process to recover the communication address when an abnormal interruption of the SDM process is detected, resulting in the loss of the communication address.
[0060] The second controller can be configured with hardware-embedded loading or boot logic to monitor in real time whether the upgrade address is written to a preset byte sequence. When the upgrade address is detected to be written to the preset byte sequence, the second controller triggers a soft reset. A soft reset is an internal reset operation triggered by the second controller after receiving a recovery trigger command. This operation does not cut off the power supply to the second controller, but only resets the internal state machine of the second controller, causing it to re-enter the firmware loading process, reread the firmware data from the preset firmware storage area (e.g., user flash memory UFM), and load the firmware data into the second controller's runtime storage area for execution.
[0061] The soft reset in this application is triggered by an I2C instruction, resetting only the internal state machine of the CPLD without cutting off the CPLD power supply. Therefore, the output pin state controlled by the CPLD remains locked during the soft reset process, and the controlled peripherals (such as hard drives) do not lose power. In contrast, a hardware reset (such as pulling the reset pin low via an external watchdog chip) will cut off the CPLD power supply or cause the output pin to enter a high-impedance state, resulting in the controlled peripherals losing power. This application uses a soft reset method, which does not affect the output pin state already latched by the second controller. Therefore, peripherals controlled by the second controller (such as hard drives, fans, etc.) will not lose power during the soft reset process. That is, the power supply state of the controlled peripherals is not affected during the repair process, achieving non-destructive repair and avoiding secondary faults caused by the repair operation.
[0062] After loading the firmware data into the runtime storage area of the second controller and executing it, the process includes: the first controller sending a probe read command to the communication address of the second controller again via the communication bus after a preset time; in response to the first controller obtaining the version number data of the second controller, determining that the repair is successful; in response to the first controller not obtaining the version number data of the second controller, determining that the repair has failed. After determining that the repair has failed, the process includes: the first controller reading the firmware integrity verification flag bit in the status register of the second controller; in response to the first controller not obtaining the firmware integrity verification flag bit, determining that the second controller has a firmware corruption fault, and the first controller triggering the firmware recovery process; in response to the first controller obtaining the firmware integrity verification flag bit, the first controller writing a preset recovery trigger command to the upgrade address of the second controller again.
[0063] After the second controller loads the firmware data into the runtime storage area and completes the execution, the first controller sends a probe read command to the communication address of the second controller again via the communication bus after a preset duration (e.g., 1 second). The preset duration is configured to be greater than or equal to the maximum time required for the second controller to complete firmware loading and open the communication address, to ensure that the second controller has sufficient time to complete the firmware loading process.
[0064] Upon receiving the version number data of the second controller via a probe read command, the first controller determines that the repair was successful. At this point, the communication address has returned to normal, and the second controller can respond normally to the first controller's regular control commands. The first controller records a successful repair log and ends the repair process.
[0065] In response to the first controller's failure to obtain the version number data of the second controller, the repair is deemed to have failed. At this point, the communication address remains unreachable, but the upgrade address remains reachable. This state combination has ruled out possibilities such as I2C bus failure, complete CPLD deadlock, and complete power supply failure. The remaining possible causes mainly fall into two categories: firmware corruption (the firmware data itself is corrupted, causing loading failure) and the SDM process being interrupted again (the firmware is intact, but the loading process is abnormally interrupted again during reloading). To effectively distinguish between these two types of failures and take targeted measures, this application further performs a firmware integrity check after the recovery verification fails: if the check passes, the firmware is deemed intact, but the loading process has failed again, and the first controller writes the recovery trigger command again; if the check fails, the firmware is deemed corrupted, and the first controller triggers the firmware recovery process.
[0066] Specifically, when the recovery verification fails, the first controller reads the firmware integrity verification flag from the status register of the second controller. The firmware integrity verification flag is a flag information pre-written into the status register by the CRC check engine or firmware verification logic inside the second controller, used to indicate whether the current firmware of the second controller is complete and valid.
[0067] After each firmware loading process, the second controller performs an integrity check on the loaded firmware data (e.g., CRC check, checksum comparison), and writes the check result to the status register as a flag. If the firmware data is complete and valid, the check flag is set to the first value, such as 1 or passed; if the firmware data is incomplete or corrupted, the check flag is set to the second value, such as 0 or failed. The first controller can determine the firmware integrity check result by reading this flag.
[0068] If the firmware integrity verification flag indicates that the firmware is incomplete or corrupted—for example, if the first controller fails to read the expected flag value or reads a failed flag value—then the first controller determines that the second controller has a firmware corruption fault. In this case, since the firmware data itself is corrupted, a simple reload operation cannot restore the communication address. The first controller triggers a firmware recovery process, such as re-flashing the firmware to the second controller's UFM area from the firmware backup area stored in the BMC or a firmware image obtained via the network, and then triggers a reload again. After the firmware recovery process is completed, the first controller performs a recovery verification again. If it still fails, it records the final fault log and reports it.
[0069] If the firmware integrity verification flag indicates that the firmware is intact and valid, for example, if the first controller successfully reads the flag value as passing, then the possibility of firmware corruption is ruled out. At this point, the firmware of the second controller itself is intact, but the loading process fails again. Possible causes include I2C bus transient interference causing the second loading to be interrupted, unstable power supply, etc. In response, the first controller writes the preset recovery trigger instruction (i.e., the Refresh instruction) to the upgrade address of the second controller again and performs recovery verification again. If the communication address is restored after the second repair, the repair is considered successful; if the communication address still fails to be restored after multiple attempts, the final fault log is recorded and reported for further investigation of hardware problems by maintenance personnel.
[0070] This application employs a multi-level repair mechanism. First, a recovery trigger instruction is written to the upgrade address to trigger the CPLD to re-execute the SDM process. This level of repair addresses soft faults where the firmware is intact but the loading process is interrupted. If the repair fails and the verification flag indicates firmware corruption, a firmware recovery process is triggered to re-flash the firmware. This level of repair addresses hard faults where the firmware data itself is corrupted. Through this multi-level repair mechanism, this application can handle various types of boot failures. For soft faults (SDM interrupted), the Refresh instruction can be used to reload and recover. For hard faults (firmware corruption), the firmware recovery process can be used to re-flash the firmware. For cases where the firmware is intact but multiple reloads still fail, logs are recorded and reported for manual troubleshooting of hardware problems, effectively improving the coverage and automation of fault handling.
[0071] In one embodiment, there are multiple second controllers. After the first controller starts up, the first controller detects whether the board containing the second controller is in place, including: after the first controller starts up, the first controller identifies the operating condition type of the board containing the second controller, which includes AC power-on, DC cold restart, and board hot-swapping; in response to the operating condition type being AC power-on, the first controller determines the processing order of the multiple second controllers, and the first controller checks whether the board containing the second controller is in place according to the processing order of the second controllers; in response to the operating condition type being DC cold restart, the first controller determines the processing order of the multiple second controllers, and after the host main power is restored, the first controller delays for a preset time and then checks whether the board containing the second controller is in place according to the processing order of the second controllers, the preset time being used to wait for the power to stabilize; in response to the operating condition type being board hot-swapping, the first controller identifies the trigger event by capturing the interrupt signal of the board in place signal pin, and only checks whether the board that generated the interrupt signal is in place.
[0072] In a typical server system, multiple boards are usually configured, each equipped with a secondary controller (CPLD). For example, backplanes, SW boards, SDB boards, power distribution boards, and fan boards all have CPLDs. The primary controller needs to execute differentiated detection strategies based on different operating conditions. In this application, the primary controller first identifies the current operating condition type, which includes AC power-on, DC cold restart, and board hot-swapping. For the different hardware states under different operating conditions, the primary controller employs corresponding processing logic to adapt to the fault detection needs of the server system in different operational scenarios.
[0073] AC power-on: This refers to the startup process after the entire machine has been completely disconnected from AC power and then reconnected to AC power. Under this condition, all components of the machine are powered on again from a power-off state. Both the BMC and CPLD undergo a complete startup process. The CPLD is initialized in advance at the millisecond level. After the BMC starts up, it executes the GPIO / FRU presence detection and hierarchical fault detection process.
[0074] When the operating condition is AC power-on, the first controller determines the processing order of multiple second controllers, and then checks the presence of each second controller's board according to this order. The processing order can be determined by a fixed board priority order or by a calculated comprehensive weight value. Under AC power-on conditions, all boards are powered on again from a power-off state, and the CPLD has completed initialization. Therefore, the first controller can directly perform a complete fault detection and repair process on all boards.
[0075] DC cold restart: This refers to the restart process where the main power supply to the host computer fails but the standby power supply continues to provide power. Under this condition, the BMC continues to run in standby mode, while only the main power supply to the host computer experiences a power outage and subsequent power-on. Because the standby power supply is not interrupted, the BMC remains operational and does not require system reloading; however, it must wait for the main power supply to stabilize before performing fault detection.
[0076] When the operating condition is a DC cold restart, the first controller determines the processing order of multiple second controllers and, after the main power supply is restored, delays for a preset time before sequentially checking whether the boards containing each second controller are in place. This preset time is used to wait for power stability, avoiding misjudgments caused by voltage fluctuations immediately after the main power supply is restored. In one implementation, the preset time is 300 milliseconds. Compared to the AC power-on condition, the BMC runs continuously under the DC cold restart condition without restarting, but requires additional waiting for the main power supply to stabilize before performing the checks.
[0077] Hot-swapping of boards: refers to the insertion and removal of peripheral boards (such as backplanes, switching boards, etc.) with CPLDs during equipment operation. Under this condition, the BMC runs continuously, sensing the insertion or removal of boards by capturing GPIO in-place interrupt signals, and only performs detection on the currently inserted or removed single board, without needing to perform batch traversal of the entire machine.
[0078] When the operating condition is hot-swapping of a board, the first controller identifies the trigger event by capturing the interrupt signal of the board's presence signal pin and only checks whether the board that generated the interrupt signal is present. Specifically, when a board is inserted, the level of the presence signal pin changes, for example, from a high level to a low level, triggering a GPIO interrupt. The first controller responds to this interrupt signal and performs presence detection and subsequent fault detection and repair processes only on the currently inserted board. Unlike AC power-on and DC cold restart operating conditions, in the hot-swapping operating condition, the first controller does not perform a batch traversal of all boards in the entire machine, but only processes the currently inserted single board to improve response efficiency and avoid interference with other normally operating boards. Compared to traditional repair logic that only supports AC power-on scenarios, this application supports the verification process in AC power-on, DC cold restart, and hot-swapping scenarios, and can be compatible with multiple operation and maintenance scenarios.
[0079] Please see Figure 3 , Figure 3 This is a schematic diagram of the power-on timing of the whole machine under multiple operating conditions provided in an embodiment of this application.
[0080] In AC power-on mode, after the entire machine is connected to AC power, the standby power is established, the CPLD starts up in milliseconds and completes SDM initialization, followed by the BMC startup. Since the BMC needs to load the operating system and initialize various services, its startup time is significantly slower than the CPLD's. After the BMC starts up, it executes the complete GPIO / FRU presence detection and hierarchical fault handling process. The CPLD starts earlier than the BMC, with a significant time difference between them. This timing relationship ensures that the CPLD has already completed initialization when the BMC begins fault detection; there is no situation where the CPLD is still initializing.
[0081] In a DC cold restart scenario, after the main power supply fails, the standby power supply continues to provide power, and the BMC remains operational without requiring a system reload. After the main power supply is restored, the first controller delays for a preset period, waiting for the power to stabilize before performing fault detection. This delay period is marked as a power stabilization waiting window to avoid misjudgments caused by voltage fluctuations immediately after the main power supply is restored.
[0082] In the case of hot-swapping of circuit boards, during equipment operation, the level of the presence signal pin changes when a board is inserted, for example, transitioning from a high level to a low level, triggering a GPIO interrupt signal. The first controller responds to this interrupt signal, performing presence detection and subsequent fault detection and repair processes only on the currently inserted / removed board. The detection and triggering of hot-swapping events are independent of BMC startup and main power supply status. The first controller responds immediately after capturing the interrupt signal, without waiting for changes in the status of other components, achieving automated plug-and-play detection.
[0083] This application is designed to adapt to three maintenance scenarios: AC power-on, DC cold restart, and hot-swapping of boards. This eliminates the need to develop separate firmware versions for different operating conditions, reducing development and maintenance costs. In the DC cold restart scenario, increasing the power stabilization delay effectively avoids misjudgments caused by voltage fluctuations when the main power supply is restored, improving detection accuracy. In the hot-swapping scenario, triggering board detection via GPIO interrupt signals enables automated plug-and-play testing with fast response times and no impact on the normal operation of other boards. This enhanced multi-condition coverage allows this application to adapt to various maintenance scenarios in data centers, increasing its application value.
[0084] Because different CPLDs on different boards vary in terms of fault impact range, power-on time, and repair urgency, treating them indiscriminately may result in failure to repair critical board faults in a timely manner, affecting the overall machine startup.
[0085] The first controller determines the processing order of multiple second controllers by: obtaining the score values of each second controller's board in multiple dimensions, including fault impact range, power-on time, repair urgency, business dependency and hardware redundancy; the score values of each dimension are preset according to the hardware characteristics of the second controller's board and stored in the first controller's configuration file.
[0086] The scope of impact of a fault refers to the degree to which the function of the entire machine is affected when the board containing the second controller fails. For example, a backplane CPLD failure will cause the hard drive to be unrecognizable, resulting in a large scope of impact and a high score; a fan board CPLD failure only affects heat dissipation, which can be degraded by directly controlling the fan through the BMC, resulting in a small scope of impact and a low score.
[0087] Power readiness time refers to the time required for the second controller to become accessible to the first controller after being powered on. A second controller powered by standby power completes initialization immediately upon AC power-on, resulting in a short power readiness time and a high score. Conversely, a second controller powered by mains power requires the host computer to be powered on before receiving power, resulting in a long power readiness time and a low score.
[0088] Repair urgency refers to the maximum allowable repair time after the second controller fails. A CPLD failure on the power distribution board affects the power supply of the entire machine and needs to be repaired in a very short time, resulting in high urgency and a high score. A CPLD failure on the fan board can be handled by the BMC directly controlling the fan strategy for degradation, resulting in a long allowable repair time, low urgency, and a low score.
[0089] Business dependency refers to the degree to which other business modules depend on the second controller. The backplane CPLD is widely relied upon by the BMC and BIOS, with high dependency and a high score; the fan board CPLD can be directly controlled by the BMC for fan control, with low dependency and a low score.
[0090] Hardware redundancy refers to whether the board containing the second controller has redundant backups. A board without redundant backups cannot degrade to a lower level after a failure, resulting in a high score; a board with redundant backups can switch to the backup after a failure, resulting in a low score.
[0091] The first controller determines the comprehensive weight value of each second controller based on its score across various dimensions using the entropy weight method. The processing order of the second controllers is then determined according to their comprehensive weight values, from highest to lowest. Fault detection and repair are then performed on each second controller sequentially according to this processing order. A higher comprehensive weight value indicates a higher processing priority for that second controller, meaning it should be detected and repaired first. Specifically: Construct a data matrix by treating each second controller as a row and each dimension as a column, using the score of each second controller in each dimension as an element. The ratio of each element in each column to the sum of all elements in that column is taken as the probability value of each second controller in each dimension. Specifically, for each dimension, the score of each second controller in that dimension is divided by the sum of the scores of all second controllers in that dimension, and the quotient is taken as the probability value of each second controller in that dimension. The calculation formula is as follows: ; Where, p ij Let x represent the probability value of the i-th second controller in the j-th dimension, m represent the total number of second controllers, and x represent the probability value of the i-th second controller in the j-th dimension. ij This represents the score of the i-th second controller in the j-th dimension. Through the above normalization process, the influence of dimensions and orders of magnitude between dimensions can be eliminated, making the data of different dimensions comparable.
[0092] The entropy weights for each dimension are determined based on the dispersion of the probability values of each second controller across all dimensions. Specifically, the information entropy value for each dimension is calculated using the probability values of each second controller across all dimensions and the formula for calculating information entropy. The information entropy value characterizes the dispersion of the probability values of each second controller across all dimensions. The lower the information entropy value, the more dispersed the probability values of each second controller are within that dimension, and the stronger the ability of that dimension to distinguish between the second controllers. Conversely, the higher the information entropy value, the more concentrated the probability values are, and the weaker the ability to distinguish between them. The entropy weights for each dimension are calculated using the information entropy values of each dimension and the entropy weight calculation formula. The formula for calculating the information entropy value is as follows: ; e j p represents the information entropy value of the j-th dimension. ij Let p represent the probability value of the i-th second controller in the j-th dimension, and m represent the total number of second controllers; in particular, if p ij =0, then define =0.
[0093] The entropy weight of each dimension is calculated based on its information entropy value. Entropy weight reflects the importance of each dimension in the overall evaluation; the lower the information entropy value, the greater its entropy weight and the greater its impact on the overall weight value; conversely, the higher the information entropy value, the smaller its entropy weight and the smaller its impact on the overall weight value. The formula for calculating entropy weight is shown below: ; w j Let represent the entropy weight of the j-th dimension, and n represent the total number of dimensions. This indicates the preset magnification factor. Greater than 1, It represents the smallest positive number that is preset.
[0094] This application introduces a preset magnification factor. And the preset minimum positive number In the CPLD board priority evaluation scenario, even if the score values of each board are exactly the same under a certain dimension (for example, the hardware redundancy of all boards is 0), the basic weight of that dimension is retained, thus avoiding the complete loss of useful information.
[0095] And when e j When it approaches 1, 1-e j Approaching zero, the general entropy weight formula will cause the entropy weight of this dimension to approach zero, but small fluctuations in the entropy value can lead to drastic changes in the weight. This application introduces an amplification factor... ( >1), when e j When it approaches 1, Approaching 0 further allows for a smoother handling of weight allocation when information entropy approaches 1, resulting in a more reasonable weight allocation.
[0096] The scores of the second controller in each dimension are weighted and summed according to the entropy weight of each dimension to obtain the comprehensive weight value of each second controller. The calculation formula is as follows: ; Among them, W i w represents the overall weight value of the i-th second controller. j Let x represent the entropy weight of the j-th dimension. ij This represents the score of the i-th second controller in the j-th dimension.
[0097] The first controller determines the processing order of each second controller based on their overall weight values, ranked from highest to lowest. The higher the overall weight value, the higher the processing priority of that second controller, and the more likely it is to be detected and repaired.
[0098] The first controller performs the aforementioned fault detection and repair process on each of the second controllers in the above-mentioned processing order, that is, it checks whether the board on which the second controller is located is in place.
[0099] Through the above-mentioned multi-dimensional evaluation and entropy weight method, this application realizes differentiated processing of multiple second controllers, ensuring that the CPLD on the key board is detected and repaired first, thereby improving the success rate and reliability of the whole machine startup.
[0100] After determining the processing order of each second controller, the first controller checks in sequence whether the board containing the second controller is in place.
[0101] The first controller first obtains a configuration file, which contains the board type, communication channel, and communication address mapping relationship of each second controller's board. The first controller then searches the configuration file for the communication channel and address corresponding to the second controller currently being processed, according to the processing order. Next, the first controller writes the channel value corresponding to the communication channel to the multiplexer via the communication bus and sends a stop condition to officially activate the channel switch. It should be noted that in the I2C bus protocol, the channel switch of the multiplexer only truly takes effect after the stop condition is met. If the first controller performs read / write operations on the second controller directly without sending a stop condition after writing the channel value, the read / write operation will still be performed through the original channel, leading to erroneous access. Therefore, this application must send a stop condition after writing the channel value to ensure that the channel switch is complete.
[0102] To ensure successful channel switching, the first controller further reads the control register of the multiplexer to obtain the current register value (i.e., the value of the currently active channel) and compares it with the previously written target channel value. If the current register value matches the target channel value, the channel switching is successful. The first controller then sends a read command to the communication address of the second controller through the switched channel to execute the fault detection and repair process. If the current register value does not match the target channel value, the channel switching fails. The first controller rewrites the target channel value to the multiplexer and verifies it again. In one embodiment, the first controller sets a preset number of retries, which can be 3, 4, etc. The specific value of the preset number of retries can be set according to actual needs. If the channel switching still fails after the preset threshold is reached, the first controller records a channel switching failure log and skips the current second controller's detection, continuing to process the next second controller.
[0103] This effectively avoids address conflicts and misaccess problems caused by multiple CPLDs using the same communication address; verification by writing and reading the control register ensures that the channel switch is successful before executing subsequent I2C operations, avoiding access to the wrong device or misjudgment of faults due to channel switch failure, thus improving the system's fault tolerance and maintainability; by managing the channel and address information of each board through configuration files and combining processing priority scheduling, automated processing of multi-board CPLDs is realized.
[0104] Please see Figure 4 In this application, the BMC and CPLD adopt a dual I2C address communication structure. The BMC is connected to an I2C multiplexer (address 0x74) via an I2C bus, and communicates with the CPLDs (Complex Programmable Logic Devices) on various boards such as the backplane, SW board (switch board), and SDB board (storage backplane) through channel switching of the multiplexer. Each CPLD is configured with a communication address (possibly 0x10 / 0x11 / 0x18) and a UFM upgrade address (0x40). The two addresses are independent of each other. The communication address is used for routine control (such as obtaining the version number, status query, etc.), and the upgrade address is used for firmware upgrades and receiving REFRESH_CMD reload commands.
[0105] When accessing CPLDs on different boards, the BMC first switches to the corresponding channel via a multiplexer, and then communicates with the target CPLD via I2C through that channel. Each CPLD may have the same communication address (e.g., all are 0x10), but because they are connected to different I2C channels, the BMC can uniquely identify each CPLD by combining the channel with the address, thereby enabling separate access and batch processing of CPLDs on multiple boards.
[0106] This application also sets up multi-level fault log recording. Specifically, in response to the first controller failing to obtain the communication address of the second controller, the identification information of the second controller, the number of communication address reading failures and the failure timestamp are recorded as the first level log. In response to the failure of the first controller to write a repair instruction to the upgrade address of the second controller, the identification information of the second controller, the number of repair instruction writing failures and the failure timestamp are recorded as the second-level log, and an alarm is triggered on the management interface. If the communication address is still not obtained after the first controller writes the repair instruction to the upgrade address of the second controller, the identification information of the second controller, the number of communication address verification failures after repair and the failure timestamp are recorded as the third level log, and the front panel fault code display and remote operation and maintenance alarm are triggered. In response to the first controller obtaining the communication address after writing the repair instruction, the second controller's identification information and repair success timestamp are recorded as a repair success log.
[0107] Please see Figure 5 This application also includes a multi-level fault log recording system to record the execution results of key nodes during fault detection and repair, so as to facilitate subsequent fault tracing and operation and maintenance analysis.
[0108] Specifically, the multi-level fault logging in this application includes the following three levels: First Fault Log: In response to the first controller failing to obtain the communication address of the second controller (i.e., communication address read failure), the first controller records the identification information of the second controller (e.g., board number, board type, or device address), the number of communication address read failures (e.g., the number of consecutive failures), and the failure timestamp. The triggering of the first fault log indicates that the communication address of the second controller is unreachable, and the first controller has determined that it has a startup failure fault and is about to trigger the repair process. By recording the first fault log, maintenance personnel can identify which boards' CPLDs experienced communication address loss faults during the power-on initialization phase, and the specific time when the fault occurred.
[0109] Second Fault Log: In response to the failure of the first controller to write a repair command to the upgrade address of the second controller, the first controller records the identification information of the second controller, the number of repair command writing failures, and the failure timestamp. The triggering of the second fault log indicates that the first controller has attempted to send a repair command to the upgrade address, but failed to write it successfully for some reason. At this time, the first controller will terminate the repair operation on the second controller. By recording the second fault log, maintenance personnel can learn the specific circumstances of the repair command writing failure, facilitating troubleshooting whether the write failure was caused by I2C bus interference, an abnormal upgrade address, or other reasons.
[0110] The third fault log: In response to the first controller failing to obtain a communication address after writing a repair command to the upgrade address of the second controller (i.e., recovery verification failed), the first controller records the identification information of the second controller, the number of communication address verification failures after repair, and the failure timestamp. The triggering of the third fault log indicates that the first controller has successfully written the repair command, but the communication address of the second controller has still not been restored after firmware reloading. At this point, the first controller will terminate the repair operation on the second controller. By recording the third fault log, maintenance personnel can know that the repair operation failed to successfully restore the communication address, requiring further investigation into whether it is firmware corruption, Flash physical damage, or other hardware failure.
[0111] In addition, this application also records a successful repair log: in response to the first controller obtaining the communication address (including initial successful detection or successful verification after repair), the first controller records the identification information of the second controller and the recovery success timestamp, indicating that the second controller has returned to normal working status.
[0112] In one implementation, before writing the repair instruction to the upgrade address of the second controller, the first controller performs a pre-verification operation to ensure that the second controller is in a state where it can normally receive and execute instructions when the repair instruction is written. Specifically, the pre-verification operation includes at least one of power supply status detection, clock signal detection, and communication bus detection. Power supply status detection includes: the first controller determines whether the current power supply to the second controller is normal by reading the power status register of the second controller or detecting the power supply normal signal (PGOOD) of the second controller. If the power supply is abnormal (e.g., the voltage is lower than the operating threshold or there are voltage fluctuations), the second controller will not be able to perform a soft reset and firmware reload normally even if it receives a repair command. If the power supply status is abnormal, the first controller records a power supply abnormality log and delays or abandons writing the repair command, and retryes it after the power supply is restored.
[0113] Clock signal detection includes: the first controller determines whether the clock signal of the second controller is stable by reading the clock status register of the second controller or detecting the clock output signal of the second controller. If the clock signal is abnormal (e.g., frequency deviation exceeds the range or clock jitter exists), the internal state machine of the second controller may not function properly, and soft reset and firmware reload operations may fail. If the clock signal is abnormal, the first controller records a clock abnormality log and delays or abandons writing repair instructions.
[0114] The communication bus detection includes: the first controller sending a probe command (e.g., a read command to the UFM status register) to the upgrade address of the second controller and checking for timeouts or error responses to determine if there is an anomaly in the communication bus (e.g., the bus is pulled low, the clock line is short-circuited, or there is electrical noise). If there is an anomaly in the communication bus, the repair command may not be correctly transmitted to the second controller. If the communication bus is abnormal, the first controller records a bus anomaly log and delays or abandons writing the repair command.
[0115] If any pre-verification operation fails, the first controller records the corresponding fault log and terminates the current repair process, preventing invalid writing of repair instructions in the event of an abnormal state of the second controller. If all the above pre-verification operations pass, the first controller continues to execute the step of writing repair instructions to the upgrade address of the second controller. In this way, this application ensures that the second controller supports soft reset and firmware reload before writing repair instructions, avoiding invalid execution of repair instructions due to power supply abnormalities, clock abnormalities, or bus abnormalities, thereby improving the success rate and reliability of the repair process.
[0116] Please see Figure 6 The hierarchical fault log in this application includes three levels: low-level log, medium-level log, and high-level log.
[0117] Low level: The trigger condition is that the board identification storage unit communicates normally, but the communication address fails to be read multiple times, indicating that the second controller has a startup failure. Logs at this level are only stored in the local log system for subsequent maintenance personnel to review and do not trigger active alarms.
[0118] Medium level: The trigger condition is a failure to write the repair command, indicating that the first controller has attempted to send the repair command to the upgrade address but failed to write it successfully for some reason. In addition to local storage, this level of log also triggers a pop-up alarm in the web management interface of the baseboard management controller, prompting maintenance personnel to intervene in a timely manner.
[0119] High-level logs: The trigger condition is that after a repair command is successfully written, the regular address remains unreachable during the recovery verification phase, indicating a hidden hardware fault in the second controller. In addition to local storage, this level of log also triggers front panel fault code display and Simple Network Management Protocol (SNMP) remote maintenance alarm reporting, achieving dual-channel emergency alarms both locally and remotely.
[0120] Successful repair log: The trigger condition is that the normal address is restored to reachability after the repair command is written. It records information such as board identification, storage unit slot identification, operating status identification, total repair time and power-on timing difference, which are used for subsequent operation and maintenance statistics and hardware optimization analysis.
[0121] Monthly statistics module: Automatically summarizes the failure frequency of each board and outputs power supply timing optimization statistics reports, providing data support for hardware iteration optimization of printed circuit board power supply timing.
[0122] In one specific implementation, such as Figure 7 As shown, the fault handling method of this application can be as follows: a multi-level verification mechanism is adopted to accurately distinguish four types of faults: empty slot, hardware open circuit, CPLD flash memory damage, and SDM timing interruption, so as to avoid meaningless whole machine restart.
[0123] Level 0 is for board presence detection. Specifically, the first controller reads the level of the general purpose input / output pins or the communication status of the board identification memory unit. If the general purpose input / output pins are high or the board identification memory unit does not respond, it is directly determined that the slot is empty or the board power supply or integrated circuit bus hardware is disconnected, terminating the current board repair process and not initiating any subsequent operations.
[0124] Level 1 is for board identification storage unit access detection. Specifically, assuming the board is in place, the first controller reads the data from the board identification storage unit. If the board identification storage unit cannot be read or written normally, the entire board hardware is considered to have failed, a level 1 fault alarm is marked, and the controller is no longer accessed.
[0125] Level 2 is for upgrade address detection. Specifically, after the board identifies the storage unit and communication is normal, the first controller accesses the controller's fixed upgrade address. If there is no response at the upgrade address, it is determined that the controller or the user's flash memory is physically damaged, a severe hardware alarm is marked, and the repair process is terminated.
[0126] Level 3 is for regular communication address verification. Specifically, when both the board identification storage unit and the upgrade address can communicate normally, the first controller reads the controller's regular communication address multiple times consecutively. If multiple reads fail, it is determined that interference at the intermediate level of the integrated circuit caused the self-loading mode process to be interrupted, and the user flash memory reload self-healing process is entered; if the regular communication address is read normally, the normal operation log is recorded directly.
[0127] In this application, the basic hardware of the board is first verified through the board identification storage unit channel. If the board identification storage unit does not respond, it is determined that the board power supply or integrated circuit bus is open, and the relevant operations of the controller are skipped directly, thereby avoiding invalid firmware reload for power supply or bus failures. After the board identification storage unit communicates normally, it accesses the controller's dedicated upgrade address. If the upgrade address does not respond, it means that the controller's user flash memory is physically damaged. If the upgrade address is normal, it means that the controller's underlying hardware is intact, thus effectively distinguishing between two types of faults: hardware damage and firmware loading failure, avoiding invalid repairs to physically damaged chips. Under the premise that both the board identification storage unit and the upgrade address are communicating normally, the controller's regular communication address is verified. Only when the regular communication address fails to read can it be determined that the self-healing soft fault caused by the intermediate level interference of the integrated circuit is the self-loading mode interruption. A reload command can be issued to complete the repair, thereby accurately locking the self-loading mode interruption fault after eliminating interference factors such as empty slots, power supply failures, and hardware damage, avoiding meaningless whole machine restarts and service interruptions. Thus, this application enables precise differentiation and differentiated handling of four types of faults: empty slots, power supply interruption, hardware damage, and self-loading mode interruption, effectively improving the accuracy of fault diagnosis and repair efficiency.
[0128] This application features GPIO hardware level recognition and board identification storage unit read / write recognition, compatible with new server boards equipped with in-situ GPIO pins and older boards without dedicated in-situ pins. It can proactively identify scenarios such as empty slots and board power supply interruptions, eliminating the need for I2C communication on invalid hardware slots, reducing invalid bus accesses, and avoiding the generation of numerous meaningless fault logs. Employing a multi-level verification mechanism, it can accurately distinguish between four types of differentiated faults: empty slot power supply failure, overall board hardware failure, CPLD flash memory physical damage, and SDM loading interruption caused by I2C level interference. Corresponding handling procedures are matched for each fault, eliminating the need for a unified system reboot and significantly reducing the probability of peripheral service interruptions. The firmware of this application incorporates multi-condition branch judgment logic, which can be adapted to AC power supply interruptions. This solution covers three typical data center operation and maintenance scenarios: power-on, DC standby cold restart, and hot-swapping of equipment operating boards. A single program covers the entire lifecycle of server power-on and operation and maintenance, making the solution more versatile. It also sets up a hierarchical fault recording mechanism, which simultaneously supports alarm output from multiple channels, including local storage, BMC management page pop-ups, equipment front panel fault codes, and SNMP remote operation and maintenance. It also periodically summarizes the frequency of various faults and generates hardware optimization statistical reports, providing comprehensive fault tracing capabilities and quantitative data support for iterative optimization of server PCB power supply and I2C hardware circuits. All detection and log alarm logic in this application can be implemented by upgrading the BMC firmware, without changing the server PCB hardware routing or component layout. The hardware modification cost is low, and existing servers can be quickly adapted and deployed.
[0129] An embodiment of this application provides a fault handling device applied to a first controller. The fault handling device is specifically as follows: Figure 8 As shown, the fault handling device includes a sending module 20 and a processing module 21.
[0130] The sending module 20 is used to detect whether the board where the second controller is located is in place after the first controller starts up; in response to the board where the second controller is located being in place, send a communication command to the board identifier storage unit on the second controller; in response to obtaining the response data of the communication command, send a read command to the upgrade address of the second controller; and in response to obtaining the upgrade address, send a read command to the communication address of the second controller.
[0131] The processing module 21 is used to respond to the failure to obtain the communication address, determine that the second controller has a startup failure fault, write a repair instruction to the upgrade address of the second controller, and trigger the second controller to re-execute the firmware loading process.
[0132] For specific limitations regarding the fault handling device, please refer to the limitations of the fault handling method above, which will not be repeated here. Each module in the aforementioned fault handling device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independently of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to each module.
[0133] like Figure 9 As shown, embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described fault handling method embodiments.
[0134] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described fault handling method embodiments when it is run.
[0135] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0136] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described fault handling method embodiments.
[0137] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described fault handling method embodiments.
[0138] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "electronic device," or "apparatus" used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.
[0139] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0140] The above provides a detailed description of a fault handling method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A fault handling method, characterized in that, Applied to a first controller, the method includes: After the first controller starts up, check whether the board containing the second controller is in place; In response to the presence of the board containing the second controller, a communication command is sent to the board identifier storage unit on the second controller; In response to receiving the response data of the communication instruction, a read instruction is sent to the upgrade address of the second controller; In response to obtaining the upgrade address, a read command is sent to the communication address of the second controller; In response to the failure to obtain the communication address, it is determined that the second controller has a startup failure fault, and a repair instruction is written to the upgrade address of the second controller. The repair instruction is used to trigger the second controller to re-execute the firmware loading process.
2. The fault handling method according to claim 1, characterized in that, The step of detecting whether the board containing the second controller is in place includes: Obtain the hardware version information of the board where the second controller is located; In response to the hardware version being the first version, the level data of the presence signal pin of the board where the second controller is located is obtained, and the presence of the board where the second controller is located is determined by the level data; In response to the hardware version being the second version, a read command is sent to the board identifier storage unit of the board containing the second controller to determine whether the board containing the second controller is in place.
3. The fault handling method according to claim 2, characterized in that, The step of sending a communication instruction to the board identifier storage unit on the second controller in response to the presence of the board containing the second controller includes: In response to the low-level data, it is determined that the board containing the second controller is in place, and a communication command is sent to the board identifier storage unit on the second controller; or, In response to receiving the response data of the board identifier storage unit sending a read command, it is determined that the board where the second controller is located is in place, and a communication command is sent to the board identifier storage unit on the second controller.
4. The fault handling method according to claim 2, characterized in that, The step of detecting whether the board containing the second controller is in place also includes: In response to the level data being high; or, In response to the failure to receive a read command from the board identifier storage unit, it is determined that the board containing the second controller is not in place, and the repair process for the second controller is terminated.
5. The fault handling method according to claim 2, characterized in that, The step of sending communication instructions to the board identifier storage unit on the second controller further includes: In response to the failure to receive response data from the board identifier storage unit to send a communication command, the board identifier storage unit is determined to be faulty, and the repair process for the second controller is terminated.
6. The fault handling method according to claim 1, characterized in that, Sending the read command to the upgrade address of the second controller includes: A probe read command is sent to the upgrade address of the second controller via the communication bus; the probe read command is used to read the status information corresponding to the upgrade address of the second controller.
7. The fault handling method according to claim 6, characterized in that, Sending the read instruction to the upgrade address of the second controller further includes: In response to the failure to obtain the status information, it is determined that the second controller has a hardware damage fault, and the repair process for the second controller is terminated.
8. The fault handling method according to claim 7, characterized in that, The step of sending a read command to the communication address of the second controller in response to obtaining the upgrade address includes: In response to obtaining the status information, it is determined that the upgrade address has been obtained; A version number read instruction is sent to the communication address of the second controller via the communication bus. The version number read instruction is used to read the version number data of the second controller.
9. The fault handling method according to claim 8, characterized in that, The step of determining whether the upgrade address has been obtained in response to obtaining the status information includes: In response to not receiving a response from the second controller to the probe read command, the probe read command is sent again to the upgrade address of the second controller via the communication bus after a preset time interval; In response to the first controller obtaining the status information, it is determined that the first controller has acquired the status information when the number of times the probe read command is sent to the upgrade address of the second controller through the communication bus reaches a preset threshold and the first controller receives a response to the probe read command from the second controller.
10. The fault handling method according to claim 1, characterized in that, The response to not obtaining the communication address, determining that the second controller has a startup failure fault includes: In response to the failure to obtain the version number data of the second controller, it is determined that the communication address has not been obtained, and the startup failure of the second controller is a communication address loss fault caused by abnormal interruption of the second controller self-loading mode process.
11. The fault handling method according to claim 10, characterized in that, The determination that the communication address has not been obtained in response to the failure to obtain the version number data of the second controller includes: In response to the failure to obtain the version number data of the second controller, a version number read command is sent again to the communication address of the second controller via the communication bus after a preset time interval; If the number of times the version number read command is sent to the communication address of the second controller via the communication bus reaches a preset threshold and the first controller fails to obtain the version number data of the second controller, it is determined that the communication address has not been obtained.
12. The fault handling method according to claim 1, characterized in that, The step of writing a repair instruction to the upgrade address of the second controller, wherein the repair instruction is used to trigger the second controller to re-execute the firmware loading process, includes: Write a preset recovery trigger instruction to the upgrade address of the second controller, wherein the recovery trigger instruction is a preset byte sequence; In response to the second controller detecting that the upgrade address has been written to a preset byte sequence, the second controller triggers a soft reset, reads firmware data from a preset firmware storage area, and loads the firmware data into the second controller's runtime storage area for execution.
13. The fault handling method according to claim 12, characterized in that, After loading the firmware data into the runtime storage area of the second controller and executing it, the process includes: After a preset time, a probe and read command is sent again to the communication address of the second controller via the communication bus; Upon receiving the version number data of the second controller, the repair is deemed successful. The repair failed because the version number data of the second controller was not obtained.
14. The fault handling method according to claim 13, characterized in that, The determination of repair failure includes: Read the firmware integrity check flag bit from the status register of the second controller; In response to the failure to obtain the firmware integrity verification flag, it is determined that the second controller has a firmware corruption fault, and the firmware recovery process is triggered; In response to obtaining the firmware integrity verification flag, a preset recovery trigger instruction is written again to the upgrade address of the second controller.
15. The fault handling method according to claim 1, characterized in that, There are multiple second controllers. The step of detecting whether the board containing the second controller is in place after the first controller has started includes: After the first controller starts up, the operating condition type of the board where the second controller is located is identified. The operating condition type includes AC power-on, DC cold restart and board hot-swap. In response to the operating condition being AC power-on, the processing order of multiple second controllers is determined, and the board containing the second controller is checked sequentially according to the processing order of the second controllers to see if it is in place; In response to the operating condition being a DC cold restart, the processing order of multiple second controllers is determined. After the main power of the host is restored, a preset time is delayed, and the board containing the second controller is checked in turn according to the processing order of the second controllers. The preset time is used to wait for the power supply to stabilize. In response to the operating condition being hot-plugging of the board, the trigger event is identified by capturing the interrupt signal of the board in-place signal pin, and only the board that generated the interrupt signal is detected to be in-place.
16. The fault handling method according to claim 15, characterized in that, Determining the processing order of multiple second controllers includes: Obtain the score values of each board where the second controller is located in multiple dimensions, including fault impact range, power supply readiness time, repair urgency, business dependency and hardware redundancy; A data matrix is constructed based on the rating values. The rating values of each dimension in the data matrix are normalized, and the probability values of each second controller in each dimension are calculated. The entropy weights for each dimension are determined based on the degree of dispersion of the probability values of each second controller in each dimension. The scores of the second controller in each dimension are weighted and summed according to the entropy weight of each dimension to obtain the comprehensive weight value of each second controller. The processing order of each of the second controllers is determined according to the order of the comprehensive weight values from high to low.
17. The fault handling method according to claim 16, characterized in that, The step of constructing a data matrix based on the rating values, normalizing the rating values for each dimension in the data matrix, and calculating the probability value of each second controller in each dimension includes: The data matrix is constructed by treating each second controller as a row, each dimension as a column, and the score of each second controller in each dimension as an element. The ratio of each element in each column of the data matrix to the sum of all elements in that column is used as the probability value of each second controller in each dimension.
18. The fault handling method according to claim 17, characterized in that, The determination of the entropy weights for each dimension based on the dispersion of the probability values of each second controller in each dimension includes: The information entropy value of each dimension is calculated based on the probability value of each second controller in each dimension and the information entropy value calculation formula; the information entropy value is used to characterize the degree of dispersion of the probability value of each second controller in each dimension; Calculate the entropy weight of each dimension based on the information entropy value of each dimension and the entropy weight calculation formula; The formula for calculating the information entropy value is as follows: ; e j p represents the information entropy value of the j-th dimension. ij Let represent the probability value of the i-th second controller in the j-th dimension, and m represent the total number of second controllers; The formula for calculating entropy weight is as follows: ; w j Let represent the entropy weight of the j-th dimension, and n represent the total number of dimensions. This indicates the preset magnification factor. Greater than 1, It represents the smallest positive number that is preset.
19. The fault handling method according to claim 15, characterized in that, According to the processing order of the second controller, the following checks are performed sequentially to determine whether the board containing the second controller is in place: Obtain the configuration file, which contains the mapping relationship between the board type, communication channel and communication address of the second controller; The communication channels corresponding to each second controller are searched sequentially from the configuration file according to the processing order described above. The channel value corresponding to the communication channel is written to the multiplexer via the communication bus, and a stop condition is sent to make the channel switching effective. Read the control register of the multiplexer to obtain the current register value; In response to the current register value being consistent with the channel value, an in-situ detection is performed on the board containing the second controller through the switched channel.
20. The fault handling method according to claim 1, characterized in that, The method further includes: In response to the failure to obtain the communication address of the second controller, the identification information of the second controller, the number of communication address reading failures and the failure timestamp are recorded as the first level log; In response to the failure to write the repair instruction to the upgrade address of the second controller, the identification information of the second controller, the number of repair instruction writing failures and the failure timestamp are recorded as the second level log, and an alarm is triggered on the management interface. If the communication address is still not obtained after writing the repair instruction to the upgrade address of the second controller, the identification information of the second controller, the number of communication address verification failures after repair and the failure timestamp are recorded as the third level log, and the front panel fault code display and remote operation and maintenance alarm are triggered. In response to obtaining the communication address after writing the repair instruction, the identification information of the second controller and the repair success timestamp are recorded as a repair success log.
21. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the fault handling method as described in any one of claims 1 to 20.
Citation Information
Patent Citations
Server startup fault maintenance method and system, storage medium and computer program product
CN120973576A
Hard disk power-on and power-off time sequence control method and electronic equipment
CN121680592A