A system and method for quickly locating and repairing hard disk hardware failures

Through the BMC control voltage conversion chip power outage and re-powering, combined with the PCH identification process, the hardware failure of the M.2 SATA SSD disk is automatically repaired, solving the MCU inoperable problem caused by VR failure, and reducing labor and equipment costs.

CN115794512BActive Publication Date: 2025-08-29INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211403282.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-10
Publication Date
2025-08-29
Estimated Expiration
2042-11-10

AI Technical Summary

Technical Problem

In the prior art, the VR fails to power when the M.2 SATA SSD disk is disturbed by external interference, resulting in the MCU not working and the server cannot be recognized. It requires manual restart or replacement of the hard disk, and the hardware failure cannot be directly repaired in the system operation state.

Method used

The BMC controls the power outage and re-powering of the voltage conversion chip, combined with the PCH identification process, automatically repairs the hard disk hardware failures to ensure that other parts of the server are operating normally.

Benefits of technology

It realizes automatic positioning and repairing hardware failures of the M.2 SATA SSD disk without affecting the operation of other parts of the server, reducing labor and equipment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115794512B_ABST
    Figure CN115794512B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of hard disk fault detection, and specifically provides a system and method for quickly locating and repairing hard disk hardware faults. The method includes the following steps: when the server is powered on, the BMC reads the hard disk hardware presence signal to determine whether the hard disk is in place; after the server is powered on, the BMC checks whether the hard disk is properly recognized by reading the hard disk information in the asset information file transmitted by the PCH; if the hard disk is not recognized, the BMC reads the hard disk presence information in the PCH and, at the same time, reads the level status of the indicator light control signal output by the hard disk to determine whether the microprocessor is functioning properly; if the microprocessor on the in-place hard disk is determined to be abnormal, the BMC controls the voltage conversion chip to power off, and after a set time, controls the voltage conversion chip to power on again, so that the PCH re-identifies the hard disk and establishes a connection to complete the fault repair. This hardware fault is automatically repaired during the boot process without affecting the operation of other parts of the server system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hard disk fault detection, and in particular to a system and method for quickly locating and repairing hard disk hardware faults. Background Art

[0002] SATA SSDs are storage devices used in servers and are typically connected to the server's PCH SATA interface. SATA SSDs can be structurally divided into 29-pin SATA hard drive interfaces and M.2 interfaces. M.2 SATA SSDs are widely used due to their high-speed transmission performance and compact size. M.2 SATA SSDs are M.2-structured solid-state drives that comply with the SATA transmission protocol. Their compact size and high speed make them suitable for use as server system drives. Due to their compact size, the VR power supply design on M.2 drives is generally small, resulting in poor interference resistance. When the server is exposed to external low-frequency interference, the VR voltage converter chip on the M.2 drive may fail to power up, causing the server's PCH to fail to recognize the M.2 drive.

[0003] Because the current M.2 SATA protocol definition does not define a pin to indicate the power ok of the M.2 SATA disk, when an M.2 SATA disk VR hardware failure occurs, it is impossible to directly determine whether the M.2 disk is not recognized due to a VR failure or a SATA signal transmission problem. The server's M.2 SATA SSD disk has no VR output, the MCU does not work, the M.2 SATA disk cannot be used normally, the BMC log will have a disk loss record, and there will be no log record on the M.2 disk. Operation and maintenance personnel are required to intervene and check the status of the LED light controlled by the DAS signal of the server's M.2 SATA disk and measure the abnormal VR level on the M.2 SATA disk to locate the cause. Operation and maintenance personnel need to restart the entire server or replace the M.2 SATA disk device. The fault cannot be directly repaired while the system is running. Summary of the Invention

[0004] When the VR output on the server's M.2 SATA SSD disk is missing, the MCU is inoperative, and the M.2 SATA disk cannot be used normally. The BMC log will record a lost disk, and there will be no log record on the M.2 disk. This requires the intervention of maintenance personnel. They must check the LED status of the server's M.2 SATA disk's DAS signal control and measure the abnormal VR level on the M.2 SATA disk to locate the cause. Maintenance personnel must restart the entire server or replace the M.2 SATA disk device, and cannot directly repair the fault while the system is running. The present invention provides a system and method for quickly locating and repairing hard disk hardware faults. After detecting an M.2 SATA disk hardware fault in the server, the system can locate the voltage converter chip VR problem that causes the M.2 SATA disk to be unrecognized. The system then controls the M.2 SATA disk to power off and then power on again, allowing the PCH to re-recognize the SATA disk. This hardware fault is automatically repaired during the boot process without affecting the operation of other parts of the server system.

[0005] In a first aspect, the technical solution of the present invention provides a system for quickly locating and repairing hard disk hardware failures, including a BMC, a PCH, a voltage conversion chip, and a hard disk equipped with a microprocessor;

[0006] The PCH, BMC, and voltage conversion chip are connected to the hard disk respectively;

[0007] The PCH and voltage conversion chip are connected to the BMC respectively;

[0008] The BMC reads the hard drive hardware presence signal to determine whether the hard drive is in place. After the hard drive is in place and the server is powered on, the BMC reads the hard drive information in the asset information file transmitted by the PCH to determine whether the hard drive is recognized normally. If the hard drive is not recognized, the BMC reads the hard drive presence information in the PCH to determine whether the hard drive's microprocessor is in place. At the same time, it reads the level status of the indicator light control signal output by the hard drive to determine whether the hard drive's microprocessor is normal. If the microprocessor of the hard drive in place is abnormal, there is a hard drive hardware failure. The BMC controls the voltage conversion chip to power off. After a set time, the control voltage conversion chip is powered on again. The PCH re-identifies the hard drive and establishes a connection to complete the fault repair.

[0009] The present invention proposes a system for quickly locating and repairing hard disk hardware failures. After detecting a hard disk hardware failure in a server, this solution can locate that the problem is a voltage conversion chip problem that causes the hard disk to be unrecognized. The hard disk is controlled to be powered off and then powered on again, and the PCH re-recognizes the hard disk. The hardware failure is automatically repaired during the startup process without affecting the operation of other parts of the server system.

[0010] Preferably, when the PCH is enabled, it periodically polls the hard disk slot to check whether a hard disk exists, and opens a link with the hard disk when a hard disk exists in the hard disk slot.

[0011] Preferably, the system further comprises a BIOS, which completes identification of all ports of the server CPU and PCH during the POST self-test process. After the port identification is completed, the PCH transmits the asset information file to the BMC via PCIe.

[0012] Preferably, a PCS register for storing enable information and presence information is provided in the PCH;

[0013] The BMC is connected to the hard disk through an analog-to-digital conversion chip;

[0014] The BMC reads the hard disk information in the asset information file and checks whether the hard disk is recognized normally. If the hard disk information does not exist, that is, the hard disk is not recognized, the BMC reads the in-place information in the PCS register to determine whether the hard disk's microprocessor is in place. At the same time, the BMC reads the level status of the indicator light control signal output by the hard disk through the analog-to-digital conversion chip to determine whether the microprocessor of the hard disk in place is working normally.

[0015] Preferably, when the BMC detects through the analog-to-digital conversion chip that the level state of the indicator light control signal output by the hard disk is low or the voltage changes between high and low levels, it indicates that the microprocessor is working normally; when it detects that the level state of the indicator light control signal output by the hard disk is always high or the level of the first set threshold is detected, it indicates that the microprocessor is not working.

[0016] Preferably, when the BMC cannot identify the hard disk presence information through the asset information file, but the microprocessor is operating normally, it is determined that there is no hardware failure of the power conversion chip, and the BMC outputs an alarm message.

[0017] Preferably, the system further comprises a power supply and an indicator light, wherein the power supply is connected to the voltage conversion chip, and the power supply is further connected to the analog-to-digital conversion chip via a first resistor;

[0018] The indicator light control signal output by the hard disk is connected to one end of the indicator light, and the other end of the indicator light is connected to the power supply through the second resistor.

[0019] The BMC reads the hard drive hardware presence signal through GPIO. The BIOS sends the asset information detection results to the BMC, which then checks whether the hard drive is recognized normally. The BMC reads the presence signal in the PCS register of the PCH to determine whether the hard drive has responded. The BMC reads the level status of the indicator light control signal output by the hard drive through the analog-to-digital conversion chip to determine whether the hard drive microprocessor is working properly. The BMC determines that the microprocessor is not working and determines that it is a hard drive hardware failure. The BMC controls the power supply of the voltage conversion chip to turn off and then turn it on again after a set time to power on the hard drive. The PCH periodically polls the back-end hard drive to re-identify the hard drive and establish a connection.

[0020] Due to the small size of the VR power supply and poor anti-interference ability of M.2, when the server is subjected to external low-frequency interference, the M.2 may not power on. This patent detects the PCH register and the M.2 DAS signal to determine that it is an M.2 SATA disk hardware failure. The BMC controls the M.2 SATA disk to restart, allowing the PCH to re-identify the M.2 SATA SSD and repair this hardware failure.

[0021] In a second aspect, the technical solution of the present invention further provides a method for quickly locating and repairing hard disk hardware failures based on the system of the first aspect, comprising the following steps:

[0022] When the server is powered on, the BMC reads the hard drive hardware presence signal to determine whether the hard drive is in place. After the hard drive is in place and the server is powered on, the BMC checks whether the hard drive is recognized normally by reading the hard drive information in the asset information file transmitted by the PCH.

[0023] If the hard drive is not recognized, the BMC reads the hard drive presence information in the PCH to determine whether the microprocessor on the hard drive is in place. At the same time, the BMC reads the level status of the indicator light control signal output by the hard drive to determine whether the microprocessor is working properly.

[0024] If the microprocessor on the hard disk is determined to be abnormal, indicating a hard disk hardware failure, the BMC controls the voltage conversion chip to power off. After a set time, the voltage conversion chip is controlled to power on again. The PCH re-identifies the hard disk and establishes a connection to complete the fault repair.

[0025] Preferably, the method further comprises:

[0026] If the BMC cannot identify the hard disk's presence information through the asset information file, but the hard disk's microprocessor is working properly, it determines that there is no power conversion chip hardware failure and outputs an alarm message.

[0027] Preferably, the step of the BMC reading the level state of the indicator light control signal output by the hard disk and judging whether the microprocessor is working normally includes:

[0028] The BMC reads the level status of the indicator light control signal output by the hard disk through the analog-to-digital conversion chip;

[0029] When it is detected that the level of the indicator light control signal output by the hard disk is low or the voltage level changes, it means that the microprocessor is working normally;

[0030] When it is detected that the level of the indicator light control signal output by the hard disk is always at a high level or the level of the first set threshold is detected, it indicates that the microprocessor is not working.

[0031] The BMC checks whether the M.2 SATA drive hardware is present and recognized by the PCH. If the hardware is present but not recognized, the BMC reads the Portx_present bit in the PCH's PCS register to determine whether the M.2 SATA drive's MCU is responding. The BMC also uses the ADC to read the DAS_LED_N level to determine whether the MCU is functioning properly. If the M.2 SATA drive's MCU is abnormal, it indicates a hardware failure, possibly caused by a VR (Virtual Reality) issue. The BMC then powers the M.2 drive back on, and the PCH then polls the backend device to re-identify the M.2 SATA SSD and establish a connection, completing the fault repair. This resolves the issue of the server PCH failing to recognize the M.2 drive due to the VR on the M.2 SATA SSD failing to power, resulting in a non-functional MCU. This requires maintenance personnel to manually restart the server or replace the M.2 SATA SSD, incurring both labor and equipment costs.

[0032] As can be seen from the above technical solution, the present invention has the following advantages: by detecting the level status of the PCH register and the control indication signal output by the hard drive, it determines that there is a hard drive hardware failure. The BMC controls the hard drive to restart, allowing the PCH to re-identify the hard drive, thus repairing the hardware failure. This design solves the problem of the hard drive's voltage conversion chip not being able to power up, causing the hard drive microprocessor to not work, resulting in the server PCH being unable to recognize the hard drive. This then requires maintenance personnel to manually restart the server or replace the hard drive, resulting in labor and equipment costs.

[0033] In addition, the present invention has a reliable design principle, a simple structure and a very broad application prospect.

[0034] It can be seen that compared with the prior art, the present invention has outstanding substantial features and significant progress, and the beneficial effects of its implementation are also obvious. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0036] Figure 1 FIG. 4 is a schematic block diagram of a system according to an embodiment of the present invention.

[0037] Figure 2 is a schematic block diagram of a system according to another embodiment of the present invention.

[0038] Figure 3 is a schematic flow chart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0039] When the VR output on a server's M.2 SATA SSD is missing, the MCU is inoperative, and the M.2 SATA drive cannot be used normally. The BMC log will show a lost disk record, and no log records are kept for the M.2 drive. This requires the intervention of maintenance personnel. They must check the LED status of the server's M.2 SATA drive's DAS signal control and measure the abnormal VR level on the M.2 SATA drive to locate the cause. Maintenance personnel must restart the entire server or replace the M.2 SATA drive device, and cannot directly fix the fault while the system is running. The present invention provides a system and method for quickly locating and repairing hard drive hardware faults. After detecting a hardware fault in a server's M.2 SATA drive, the system can locate a voltage converter chip VR issue that causes the M.2 SATA drive to be unrecognized. The system then powers down the M.2 SATA drive and restarts it, allowing the PCH to re-recognize the SATA drive. This hardware fault is automatically fixed during the boot process without affecting the operation of other parts of the server system. To help those skilled in the art better understand the technical solutions of the present invention, the following will provide a clear and complete description of the technical solutions in the embodiments of the present invention, combined with the accompanying drawings. Obviously, the described embodiments represent only some, not all, of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work should fall within the scope of protection of the present invention.

[0040] like Figure 1 As shown, an embodiment of the present invention provides a system for quickly locating and repairing hard disk hardware failures, including a BMC, a PCH, a voltage conversion chip, and a hard disk equipped with a microprocessor;

[0041] The PCH, BMC, and voltage conversion chip are connected to the hard disk respectively;

[0042] The PCH and voltage conversion chip are connected to the BMC respectively;

[0043] The BMC reads the hard drive hardware presence signal to determine whether the hard drive is in place. After the hard drive is in place and the server is powered on, the BMC reads the hard drive information in the asset information file transmitted by the PCH to determine whether the hard drive is recognized normally. If the hard drive is not recognized, the BMC reads the hard drive presence information in the PCH to determine whether the hard drive's microprocessor is in place. At the same time, it reads the level status of the indicator light control signal output by the hard drive to determine whether the hard drive's microprocessor is normal. If the microprocessor of the hard drive in place is abnormal, there is a hard drive hardware failure. The BMC controls the voltage conversion chip to power off. After a set time, the control voltage conversion chip is powered on again. The PCH re-identifies the hard drive and establishes a connection to complete the fault repair.

[0044] like Figure 2As shown, an embodiment of the present invention provides a system for quickly locating and repairing hard disk hardware failures, including a BMC, a PCH, a voltage conversion chip, and a hard disk equipped with a microprocessor;

[0045] The PCH, BMC, and voltage conversion chip are connected to the hard disk respectively;

[0046] The PCH and voltage conversion chip are connected to the BMC respectively;

[0047] The PCH is equipped with PCS registers for storing enable information and in-position information; the BMC is connected to the hard disk through an analog-to-digital conversion chip;

[0048] After the server is powered on, the PCH is enabled and periodically polls the hard disk slot to check whether a hard disk is present. If a hard disk is present in the hard disk slot, the connection to the hard disk is opened. The BIOS completes all port identification of the server CPU and PCH during the POST self-test process. After the port identification is completed, the PCH transmits the asset information file to the BMC via PCIe. The BMC reads the hard disk information in the asset information file to determine whether the hard disk is normally identified under the PCH. If the hard disk information described in the asset information file does not exist, that is, the hard disk is not identified, the BMC reads the hard disk presence information in the PCS register in the PCH to determine whether the hard disk's microprocessor is in place, and reads the indicator light output by the hard disk at the same time. The level status of the control signal determines whether the hard disk microprocessor is normal. When the level status of the indicator light control signal output by the hard disk is detected to be low or the voltage changes between high and low levels, it indicates that the microprocessor is working normally. When the level status of the indicator light control signal output by the hard disk is detected to be always high or the level of the first set threshold is detected, it indicates that the microprocessor is not working. If the microprocessor is not working, there is a hard disk hardware failure. The BMC controls the voltage conversion chip to power off. After a set time, the voltage conversion chip is controlled to power on again. The PCH re-identifies the hard disk and establishes a connection to complete the fault repair. If the microprocessor is working normally, it is determined that there is no power conversion chip hardware failure, and the BMC outputs an alarm message.

[0049] It should be noted that the system also includes a power supply and an indicator light. The power supply is connected to the voltage conversion chip, and the power supply is also connected to the analog-to-digital conversion chip through a first resistor R1; the indicator light control signal output by the hard disk is connected to one end of the indicator light, and the other end of the indicator light is connected to the power supply through a second resistor R2.

[0050] Due to the compact size of M.2 SATA SSDs, the VR power supply (or power conversion chip in this case) on the M.2 drive is typically small, resulting in poor noise immunity. When a server is exposed to external low-frequency interference, the VR power supply on the M.2 drive may fail to power on, causing the microprocessor (MCU) to malfunction and the server's PCH to fail to recognize the M.2 drive. The SATA M.2 specification does not include a pin to indicate the SATA M.2 drive's power supply (PWRGD). When an M.2 SATA drive VR hardware failure occurs, it's difficult to determine whether the VR failure or SATA signal transmission issues are causing the M.2 drive to fail properly. If the M.2 drive VR hardware fails, the M.2 MCU controller will not function, and the cause of the drive removal will not be recorded. The SATA M.2 specification does include a DAS (Drive Activity Signal) pin, which is generally used to control an external LED to indicate whether the M.2 SATA SSD is connected. An off LED indicates no SATA SSD is connected, while a lit or blinking LED indicates a connected SATA SSD. The BMC monitors the level of the M.2 SATA drive indicator control signal DAS_LED_N through the analog-to-digital converter chip to determine whether the M.2 SATA drive MCU output signal is normal, and thus whether the MCU is functioning properly. If the BMC detects a low level or voltage fluctuations through the analog-to-digital converter chip, it indicates that the MCU is functioning properly. If the BMC detects a consistently high level or an abnormally low level, it indicates that the M.2 SATA drive MCU is not functioning properly.

[0051] The PCS register of the PCH's SATA port contains the enable information (Portx enable) and the presence information (Portx present). This controls the SATA port's enablement and detects the presence of a SATA drive. When in the Portx enable state, the PCH periodically polls for the presence of a device and then initiates the link process. The BMC can read the M.2 SATA drive's presence signal (M2_PRSNT_N) to verify that the device is inserted into the slot. When the server boots up, the BIOS completes POST (Post Processing) to identify all ports on the CPU and PCH and sends the asset information file (asset.json) to the BMC for web display. The BMC checks whether the M.2 SATA SSD drive has established a connection. If the asset information does not contain any M.2 SATA drive information, the BMC reads the PCH's Port Control and Status (PCS) register and the corresponding Portx_present bit to determine whether the M.2 SATA drive has responded. If the PCS register does not detect presence information, the M.2 SATA SSD's MCU is not functioning, indicating a hardware failure with the M.2 SATA drive VR. If the register contains in-position information, the MCU on the M.2 is working, and the M.2 SATA disk failure is not caused by a VR hardware problem. The BMC only generates an alarm message.

[0052] When the BMC determines a hardware fault on the M.2 SATA drive by checking the DAS pin level and reading the PCH register, it controls the I2C interface to shut down the M.2's power supply, the P3V3 VR, for a period of time and then restart it, powering the MCU on the M.2 SATA drive. The PCH then polls the backend device to re-identify the M.2 SATA SSD and establish a connection.

[0053] like Figure 3 As shown, an embodiment of the present invention also provides a method for quickly locating and repairing a hard disk hardware failure, comprising the following steps:

[0054] Step 1: The server is powered on, and the BMC reads the hard disk hardware presence signal to determine whether the hard disk is in place.

[0055] Step 2: After the hard drive is in place and the server is powered on, the PCH transmits the asset information file to the BMC.

[0056] Step 3: The BMC checks whether the hard disk is recognized normally by reading the hard disk information in the asset information file;

[0057] If the hard disk is not recognized, proceed to step 4. If the hard disk is recognized normally, end.

[0058] Step 4: The BMC reads the hard drive presence information in the PCH to determine whether the microprocessor on the hard drive is in place. At the same time, the BMC reads the level status of the indicator light control signal output by the hard drive to determine whether the microprocessor is working properly.

[0059] If it is determined that the microprocessor on the hard disk is abnormal, go to step 5; if the microprocessor on the hard disk is working properly, go to step 7;

[0060] Step 5: If there is a hard disk hardware failure, the BMC controls the voltage conversion chip to power off and then powers it back on after a set time.

[0061] Step 6: PCH re-identifies the hard drive and establishes a connection to complete the fault repair.

[0062] Step 7: It is determined that there is no hardware failure of the power conversion chip, and the BMC outputs an alarm message.

[0063] It should be noted that the steps for the BMC to read the level status of the indicator light control signal output by the hard disk and determine whether the microprocessor is working properly include:

[0064] The BMC reads the level status of the indicator light control signal output by the hard disk through the analog-to-digital conversion chip;

[0065] When it is detected that the level of the indicator light control signal output by the hard disk is low or the voltage level changes, it means that the microprocessor is working normally;

[0066] When it is detected that the level of the indicator light control signal output by the hard disk is always at a high level or the level of the first set threshold is detected, it indicates that the microprocessor is not working.

[0067] When the server boots up, the BMC reads the M.2 SATA drive hardware presence signal via GPIO to determine if it is present. After booting, the PCH transmits the asset information file, asset.json, to the BMC via PCIe. The BMC then checks whether the M.2 SATA drive is recognized properly. If not, the BMC reads the Portx_present bit in the PCH's PCS register to determine whether the M.2 SATA drive's MCU is responding. Simultaneously, the BMC reads the DAS_LED_N level via the analog-to-digital converter chip to determine if the MCU is functioning properly. If the M.2 SATA drive's MCU is abnormal, it indicates a hardware failure, likely caused by a VR. The BMC controls the P3V3 VR via I2C to disable output. After a period of time, the BMC controls the P3V3 VR to power on again, enabling the M.2 SATA drive to function properly. The PCH then polls the backend device to re-identify the M.2 SATA SSD and establish a connection, completing the repair.

[0068] Although the present invention has been described in detail with reference to the accompanying drawings and in conjunction with preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and such modifications or substitutions shall be within the scope of the present invention. Any person skilled in the art who can easily conceive of changes or substitutions within the technical scope disclosed in the present invention shall be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope of protection of the claims.

Claims

1. A system for quickly locating and repairing hard disk hardware failures, characterized in that: Includes BMC, PCH, voltage conversion chip and hard disk with microprocessor; The PCH, BMC, and voltage conversion chip are connected to the hard disk respectively; The PCH and voltage conversion chip are connected to the BMC respectively; The BMC reads the hard drive hardware presence signal to determine whether the hard drive is in place. After the hard drive is in place and the server is powered on, the BMC reads the hard drive information in the asset information file transmitted by the PCH to determine whether the hard drive is recognized normally. If the hard drive is not recognized, the BMC reads the hard drive presence information in the PCH to determine whether the hard drive's microprocessor is in place. At the same time, it reads the level status of the indicator light control signal output by the hard drive to determine whether the hard drive's microprocessor is normal. If the microprocessor of the hard drive in place is abnormal, there is a hard drive hardware failure. The BMC controls the voltage conversion chip to power off. After a set time, the control voltage conversion chip is powered on again. The PCH re-identifies the hard drive and establishes a connection to complete the fault repair. The system also includes a BIOS, which completes all port identification of the server CPU and PCH during the POST self-test process. After the port identification is completed, the PCH transmits the asset information file to the BMC via PCIe; The PCH is provided with a PCS register for storing enable information and in-position information; The BMC is connected to the hard disk through an analog-to-digital conversion chip; The BMC reads the hard disk information in the asset information file and checks whether the hard disk is recognized normally. If the hard disk information does not exist, that is, the hard disk is not recognized, the BMC reads the in-place information in the PCS register to determine whether the hard disk's microprocessor is in place. At the same time, the BMC reads the level status of the indicator light control signal output by the hard disk through the analog-to-digital conversion chip to determine whether the microprocessor of the hard disk in place is working normally.

2. The system for quickly locating and repairing hard disk hardware failures according to claim 1, characterized in that: When PCH is enabled, it periodically polls the hard disk slot to check whether there is a hard disk. When a hard disk is present in the hard disk slot, it opens the connection with the hard disk.

3. The system for quickly locating and repairing hard disk hardware failures according to claim 2, characterized in that: When the BMC detects through the analog-to-digital conversion chip that the level of the indicator light control signal output by the hard disk is low or the voltage changes between high and low levels, it indicates that the microprocessor is working normally; when it detects that the level of the indicator light control signal output by the hard disk is always high or the level of the first set threshold is detected, it indicates that the microprocessor is not working.

4. The system for quickly locating and repairing hard disk hardware failures according to claim 1, characterized in that: If the BMC cannot identify the hard disk's presence information through the asset information file, but the microprocessor is working normally, it determines that there is no power conversion chip hardware failure and outputs an alarm message.

5. The system for quickly locating and repairing hard disk hardware failures according to claim 1, characterized in that: The system further includes a power supply and an indicator light, wherein the power supply is connected to the voltage conversion chip, and the power supply is also connected to the analog-to-digital conversion chip via a first resistor; The indicator light control signal output by the hard disk is connected to one end of the indicator light, and the other end of the indicator light is connected to the power supply through the second resistor.

6. A method for quickly locating and repairing hard disk hardware failures, applied to the system according to any one of claims 1 to 5, characterized in that: The steps include: When the server is powered on, the BMC reads the hard drive hardware presence signal to determine whether the hard drive is in place. After the hard drive is in place and the server is powered on, the BMC checks whether the hard drive is recognized normally by reading the hard drive information in the asset information file transmitted by the PCH. If the hard drive is not recognized, the BMC reads the hard drive presence information in the PCH to determine whether the microprocessor on the hard drive is in place. At the same time, the BMC reads the level status of the indicator light control signal output by the hard drive to determine whether the microprocessor is working properly. If the microprocessor on the hard disk is determined to be abnormal, indicating a hard disk hardware failure, the BMC controls the voltage conversion chip to power off. After a set time, the voltage conversion chip is controlled to power on again. The PCH re-identifies the hard disk and establishes a connection to complete the fault repair.

7. The method for quickly locating and repairing hard disk hardware failure according to claim 6, characterized in that: The method further includes: If the BMC cannot identify the hard disk's presence information through the asset information file, but the hard disk's microprocessor is working properly, it determines that there is no power conversion chip hardware failure and outputs an alarm message.

8. The method for quickly locating and repairing hard disk hardware failure according to claim 7, characterized in that: The BMC reads the level of the indicator light control signal output by the hard disk and determines whether the microprocessor is working properly. The steps include: The BMC reads the level status of the indicator light control signal output by the hard disk through the analog-to-digital conversion chip; When it is detected that the level of the indicator light control signal output by the hard disk is low or the voltage level changes, it means that the microprocessor is working normally; When it is detected that the level of the indicator light control signal output by the hard disk is always at a high level or the level of the first set threshold is detected, it indicates that the microprocessor is not working.

Citation Information

Patent Citations

  • Hard disk fault detection method and related device

    CN111048138A

  • Hard disk fault early warning method and system, terminal and storage medium

    CN115221015A