A hard disk fault processing system, method, electronic device and storage medium

By working together with the slot management module, protocol management module, and array driver module, the system automatically detects and recovers faulty hard drives, solving the problem of low efficiency in RAID hard drive failure handling and enabling rapid recovery of RAID performance and prevention of data risks.

CN120803805BActive Publication Date: 2025-11-11INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511244114.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-11
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

The existing RAID hard drive failure handling process is cumbersome, resulting in low processing efficiency and affecting RAID performance.

Method used

The system employs a slot management module, a protocol management module, and an array driver module working together to automatically detect faulty hard drives and restore their status. This includes reading physical information through current and signal quality sensors, clearing alarms after confirming successful slot access, and performing health checks and status restoration.

Benefits of technology

It improves the automation and efficiency of RAID hard drive failure handling, ensuring that failed hard drives can be quickly restored to normal use, guaranteeing RAID performance, and preventing secondary data risks caused by human error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803805B_ABST
    Figure CN120803805B_ABST
Patent Text Reader

Abstract

This application discloses a hard disk failure handling system, method, electronic device, and storage medium, relating to the field of computer technology. By utilizing a slot management module and a protocol management module, a series of tests are performed on the faulty hard disk after it has been plugged and unplugged from its original slot. After a series of tests determine that the faulty hard disk can continue to be used normally after plugging and unplugging, the array driver module performs state recovery on the faulty hard disk. The entire process is automatically triggered by the user plugging and unplugging the faulty hard disk from its original slot, resulting in a high degree of automation. Furthermore, once it is determined that the faulty hard disk can continue to be used normally after plugging and unplugging, the array driver module only needs to perform a simple state recovery. The configuration process for state recovery is simple, improving the efficiency of RAID hard disk failure handling and enabling RAID to quickly return to normal business status, thereby ensuring RAID performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a hard disk fault handling system, method, electronic device, and storage medium. Background Technology

[0002] Redundant Array of Independent Disks (RAID) is a technology that combines multiple independent physical hard drives in different ways to form a hard drive group. By distributing data across multiple hard drives, it achieves data redundancy and backup.

[0003] In related technologies, when any hard drive in a RAID array fails, it is typically replaced with a new one. This new hard drive is then manually configured through a series of tests and setup steps to enable it to participate in data redundancy backup as a member of the RAID array. However, this manual configuration process is cumbersome, reducing the efficiency of handling RAID hard drive failures and compromising RAID performance. Summary of the Invention

[0004] This application provides a hard disk failure handling system, method, electronic device, and storage medium to at least solve the problem in the related art that reduces the handling efficiency of RAID hard disk failures and is detrimental to ensuring RAID performance.

[0005] This application provides a hard disk fault handling system, including: a slot management module, a protocol management module, and an array driver module; the slot management module includes a current sensor and a signal quality sensor;

[0006] The slot management module is used to open the original slot after the user has plugged and unplugged the faulty hard drive. It reads the physical information of the faulty hard drive through current sensor and signal quality sensor. The physical information includes at least the interface status signal, current data and signal error rate of the original slot. Based on the interface status signal, current data and signal error rate of the original slot, it determines whether the original slot has been opened successfully. If the original slot has been opened successfully, it clears the alarm information of the original slot and sends a hard drive online notification to the protocol management module.

[0007] The protocol management module is used to perform a health check on the faulty hard drive when it receives a hard drive online notification, and send a hard drive recovery notification to the array driver module if the faulty hard drive passes the health check.

[0008] The array driver module is used to reuse the original logical resources of the failed hard disk in the array member list when a hard disk recovery notification is received, so as to restore the state of the failed hard disk and restore it to a member disk.

[0009] The array member list includes the allocation relationship between each member disk and logical resources in the independent redundant disk array. The logical resources include at least the unique identifier of the hard disk, the unique identifier of the logical slot, and the address mapping information.

[0010] This application also provides a hard disk failure handling method, including:

[0011] After the user performs the process of plugging and unplugging the faulty hard drive in the original slot, the original slot is opened and the physical information of the faulty hard drive is read. The physical information includes at least the interface status signal, current data and signal error rate of the original slot.

[0012] Based on the interface status signal, current data, and signal error rate of the original slot, determine whether the original slot has been successfully opened.

[0013] If the original slot is successfully opened, clear the alarm information of the original slot;

[0014] Perform a health check on the faulty hard drive. If the faulty hard drive passes the health check, reuse the original logical resources of the faulty hard drive in the array member list to restore the state of the faulty hard drive and restore it to a member disk.

[0015] The array member list includes the allocation relationship between each member disk and logical resources in the independent redundant disk array. The logical resources include at least the unique identifier of the hard disk, the unique identifier of the logical slot, and the address mapping information.

[0016] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described hard disk failure handling methods when executing the computer program.

[0017] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described hard disk fault handling methods.

[0018] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described hard disk failure handling methods.

[0019] This application utilizes a slot management module and a protocol management module to perform a series of tests on a faulty hard drive after it has been plugged and unplugged from its original slot. Once these tests determine that the faulty hard drive can continue to be used normally after being plugged and unplugged, the array driver module performs state recovery on the faulty hard drive. The entire process is automatically triggered by the user plugging and unplugging the faulty hard drive from its original slot, resulting in a high degree of automation. Furthermore, once it is determined that the faulty hard drive can continue to be used normally after being plugged and unplugged, the array driver module only needs to perform a simple state recovery. The configuration process for state recovery is simple, which improves the efficiency of handling RAID hard drive failures and allows the RAID to recover to a normal business state as quickly as possible, thereby ensuring RAID performance. Attached Figure Description

[0020] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the interaction process of the hard disk fault handling system provided in the embodiments of this application;

[0022] Figure 2 This is a schematic diagram illustrating a scenario of inserting and removing a faulty hard drive in the original slot, as provided in an embodiment of this application.

[0023] Figure 3 This is a schematic diagram illustrating a target slot failure hard drive insertion / removal scenario provided in an embodiment of this application.

[0024] Figure 4 A schematic diagram illustrating a scenario of switching a new hard drive to the original slot, as provided in an embodiment of this application.

[0025] Figure 5 This is a schematic diagram of the hard disk state switching logic provided in an embodiment of this application;

[0026] Figure 6 A schematic diagram of the interaction flow of an exemplary hard disk fault handling system provided in this application embodiment;

[0027] Figure 7 A flowchart illustrating the hard disk failure handling method provided in this application embodiment;

[0028] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0030] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0031] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] In related technologies, the troubleshooting process for replacing a RAID disk is complex. It generally involves releasing resources occupied by the failed disk, clearing related alarms, changing the new disk's status to managed, and manually adding it to the target RAID array that previously had the failed disk. The RAID will then automatically recover. This process involves multiple functional modules of the storage system, making it prone to various problems. After a hard drive failure, replacing the disk is a complex process with a low success rate. Furthermore, the new hard drive requires a series of manual tests and configurations to ensure it participates in data redundancy backup as a RAID member. However, this manual configuration process is cumbersome, reducing the efficiency of handling RAID hard drive failures and negatively impacting RAID performance.

[0033] To address the aforementioned technical problems, this application provides a hard disk failure handling system, method, electronic device, and storage medium. The system includes a slot management module, a protocol management module, and an array driver module. By utilizing the slot management module and protocol management module to perform a series of tests on the faulty hard disk after it has been plugged and unplugged from its original slot, and after a series of tests confirming that the faulty hard disk can continue to be used normally after plugging and unplugging, the array driver module performs state recovery on the faulty hard disk. The entire process is automatically triggered by the user plugging and unplugging the faulty hard disk from its original slot, resulting in a high degree of automation. Furthermore, once it is determined that the faulty hard disk can continue to be used normally after plugging and unplugging, the array driver module only needs to perform a simple state recovery. The configuration process for state recovery is simple, improving the efficiency of RAID hard disk failure handling and enabling RAID to quickly return to normal business status, thereby ensuring RAID performance.

[0034] This application provides a hard disk failure handling system for handling hard disk failures after any hard disk in a RAID array fails.

[0035] like Figure 1 The diagram shown is an interactive flow diagram of the hard disk fault handling system provided in this application embodiment. The hard disk fault handling system includes: a slot management module, a protocol management module, and an array driver module. The slot management module includes a current sensor and a signal quality sensor.

[0036] The slot management module is used to open the original slot after the user has unplugged and plugged in the faulty hard drive. It reads the physical information of the faulty hard drive using current and signal quality sensors. The physical information includes at least the interface status signal, current data, and signal error rate of the original slot. Based on the interface status signal, current data, and signal error rate of the original slot, it determines whether the original slot has been successfully opened. If the original slot has been successfully opened, it clears the alarm information of the original slot and sends a hard drive online notification to the protocol management module. When the protocol management module receives the hard drive online notification, it performs a health check on the faulty hard drive. If the faulty hard drive passes the health check, it sends a hard drive recovery notification to the array driver module. When the array driver module receives the hard drive recovery notification, it reuses the original logical resources of the faulty hard drive in the array member list to restore the status of the faulty hard drive and restore it to a member disk.

[0037] The array member list includes the allocation relationship between each member disk and logical resources in the independent redundant disk array. The logical resources include at least the unique identifier of the hard disk, the unique identifier of the logical slot, and the address mapping information.

[0038] It should be noted that the slot management module can be implemented based on the device management module (enclosure) in the storage system; the slot management module is abbreviated as EN module. The protocol management module can be implemented based on the hard disk protocol management module (VL - Volume Layer) in the storage system; the protocol management module is abbreviated as VL module. The array driver module can be abbreviated as RAID module.

[0039] Specifically, when a user plugs or unplugs a faulty hard drive in its original slot, the slot management module detects the change in the slot's physical status via a hardware interface and then initiates the slot's activation process, such as activating the slot's power supply and data transmission channels. Simultaneously, the EN module reads the faulty hard drive's physical information, including its physical ID, interface connection status signal, current data, and signal error rate, to determine if the slot has been successfully activated. If the physical information shows stable current, adequate signal strength, and no contact errors (low signal error rate), and the interface connection status signal indicates a normal connection, then the slot is considered successfully activated, thus ruling out a fault in the original slot and confirming that the hard drive's failure is not due to a fault in the original slot. After successful activation, the EN module first clears the original slot's alarm information to prevent historical alarms from interfering with subsequent processes. It then sends a hard drive online notification to the VL module, informing it that the hard drive is physically connected and ready for the next functional test. The stability of the current can be determined based on the current fluctuations represented by the current data. When the current fluctuation is below a preset fluctuation threshold, the current is considered stable. Correspondingly, if the signal error rate is below a preset error rate threshold, no contact failure error is detected. By accurately analyzing the original slot from aspects such as interface connection status signals, current data, and signal error rate, the accuracy of the fault judgment results for the original slot is improved, laying the foundation for improving the reliability of subsequent hard drive fault handling procedures.

[0040] Furthermore, after receiving the hard drive online notification from the EN module, the VL module initiates a health check process for the faulty hard drive. If the faulty hard drive passes the health check, the VL module determines that the faulty hard drive can continue to be used normally. The previous fault was likely caused by a loose connection between the original slot and the faulty hard drive, which is resolved by reseating the drive. Upon confirming that the faulty hard drive has passed the health check, the VL module sends a hard drive recovery notification to the RAID module, informing it that the hard drive is functioning normally and can continue to be used.

[0041] Furthermore, after receiving the hard drive recovery notification from the VL module, the RAID module executes a status recovery process for the failed hard drive to restore it as a member drive. Moreover, the RAID module can reuse the original array configuration (original logical resources) of the failed hard drive in the queue member list without reconfiguration, further improving the efficiency of RAID fault recovery.

[0042] For example, such as Figure 2The diagram illustrates a scenario of plugging and unplugging a faulty hard drive in the original slot, as provided in this embodiment of the application. First, the faulty hard drive is currently in a faulty state, the original slot is closed, and the physical state change indicator for the slot is "No." After the user plugs and unplugs the faulty hard drive in the original slot, the physical state change indicator for the slot switches to "Yes." The EN module first opens the original slot and obtains the physical information of the faulty hard drive through the Discovery operation. If the physical information indicates that the original slot has been successfully opened, the EN VRT flag is modified to trigger the EN module to set drive_replacement (hard drive replacement flag) to the IN_PROGRESS state. Simultaneously, the swap (slot physical state change flag) and bypass (original slot closed) flags are cleared, and any unrepaired alarms are cleared. Then, drive_replacement is changed to the EN_COMPLETE state to trigger a hard drive online notification to be sent to the protocol management module (VL module). The VL module performs a health check on the faulty hard drive by running a disk test program. If the health check passes, it clears any unrepaired alarms and sets the drive_replacement to VL_COMPLETE (VL module complete) state, triggering a hard drive recovery notification to the array driver module. The array driver module first changes the hard drive's usage status to Spare, then restores it as a member drive and sets the drive_replacement to RD_COMPLETE (RAID module complete) state, triggering the EN module to change the drive_replacement to IDLE, signifying the end of the faulty hard drive processing (END). The entire process involves collaboration among modules, performing checks and status modifications on the re-inserted faulty hard drive.

[0043] Based on the above embodiments, as an implementable approach, in one embodiment, the slot management module is further configured to:

[0044] If the original slot fails to open, a slot switching notification is sent to the user so that the user can respond to the slot switching notification, select the target slot from the available slots, and perform the unplug and plug-in of the faulty hard drive in the target slot.

[0045] The target slot can be randomly selected from the available slots. If the physical information indicates that the original slot has failed to open, it can be determined that the original slot has a slot fault. Therefore, it is necessary to switch the slot, that is, to select the target slot from the available slots and then perform the faulty hard drive insertion and removal process in the target slot.

[0046] Accordingly, in one embodiment, the slot management module is further configured to enable the target slot after the user performs plug-and-play processing on the faulty hard drive in the target slot, and read the physical information of the faulty hard drive through a current sensor and a signal quality sensor; the physical information includes at least the interface status signal, current data and signal error rate of the target slot; determine whether the target slot has been successfully enabled based on the interface status signal, current data and signal error rate of the target slot; and send a hard drive online notification to the protocol management module if the target slot has been successfully enabled.

[0047] The process for opening a target slot in the slot management module is the same as that for opening the original slot, and will not be described again here.

[0048] For example, such as Figure 3 The diagram illustrates a scenario of plugging and unplugging a faulty hard drive in a target slot, as provided in this embodiment of the application. First, the faulty hard drive is currently in a faulty state, the original slot is closed, and the physical state change indicator for the slot is "no." After the user plugs and unplugs the faulty hard drive in the target slot, the physical state change indicator for the slot switches to "yes." The EN module first opens the target slot and obtains the physical information of the faulty hard drive through the Discovery operation. If the physical information indicates that the target slot has been successfully opened, the EN VRT flag is modified to trigger the EN module to set drive_replacement (hard drive replacement flag) to the IN_PROGRESS state. Simultaneously, the swap (slot physical state change flag) and bypass (target slot closed) flags are cleared, and any unrepaired alarms are cleared. Then, drive_replacement is changed to the EN_COMPLETE state to trigger a hard drive online notification to be sent to the protocol management module (VL module). The VL module performs a health check on the faulty hard drive by running a disk test program. If the health check passes, it clears any unrepaired alarms and sets the drive_replacement to the VL_COMPLETE (VL module complete) state, triggering a hard drive recovery notification to the array driver module. The array driver module first changes the hard drive's usage status to Spare, then restores it as a member disk and sets the drive_replacement to the RD_COMPLETE (RAID module complete) state, triggering the EN module to change the drive_replacement to the IDLE (idle) state, indicating the end of the faulty hard drive's processing flow.

[0049] Based on the above embodiments, as an implementable approach, in one embodiment, the protocol management module is further configured to:

[0050] If the faulty hard drive fails the health check, a hard drive switch notification will be sent to the user so that the user can respond to the hard drive switch notification and switch to a new hard drive in the original slot.

[0051] Specifically, if the faulty hard drive fails the health check, it can be determined that the hard drive itself is faulty and should be replaced with a new hard drive, that is, the user should be notified of the hard drive replacement.

[0052] Accordingly, in one embodiment, the slot management module is further configured to restart the original slot after the user performs the new hard drive switching process in the original slot, and read the physical information of the faulty hard drive through a current sensor and a signal quality sensor; the physical information includes at least the interface status signal, current data and signal error rate of the original slot; determine whether the original slot has been successfully restarted based on the interface status signal, current data and signal error rate of the original slot; if the original slot has been successfully restarted, clear the alarm information of the original slot and send a new hard drive switching notification to the protocol management module.

[0053] The restart process when inserting a new hard drive into the original slot is the same as the startup process when inserting a faulty hard drive into the original slot, and will not be described again here.

[0054] Furthermore, in one embodiment, the protocol management module is also used to perform legitimate authentication of the new hard drive when a new hard drive switching notification is received; and to send a hard drive configuration notification to the array driver module if the new hard drive passes the legitimate authentication.

[0055] It should be noted that legitimate authentication refers to verifying the compatibility and functionality between the new hard drive and the storage system to ensure that the new hard drive is suitable for the storage system. This avoids the installation of an incompatible hard drive due to user error, thereby improving the reliability of the storage system.

[0056] For example, such as Figure 4The diagram illustrates a scenario of switching a new hard drive to the original slot as provided in this application embodiment. First, the faulty hard drive is currently in a faulty state, the original slot is in a closed state, and the physical state change indicator of the slot is "no". After the user plugs or unplugs the new hard drive in the original slot, the physical state change indicator of the slot changes to "yes". The EN module first opens the original slot and obtains the physical information of the new hard drive through the Discovery operation. If the physical information indicates that the original slot has been successfully opened, the EN VRT flag is modified to trigger the EN module to set drive_replacement (hard drive replacement indicator) to IN_PROGRESS (in progress) state, and at the same time clear the swap (slot physical state change indicator) and bypass (original slot closed) flags, clearing any unrepaired alarms. Then, drive_replacement is changed to EN_COMPLETE (EN module completed) state to trigger the sending of a new hard drive switching notification to the protocol management module (VL module). The VL module performs a disk test program to authenticate the new hard drive. If authentication is successful, it clears any unrepaired alarms and sets `drive_replacement` to the VL_COMPLETE (VL module complete) state, triggering a hard drive configuration notification to the array driver module. The array driver module first changes the previously faulty hard drive to the unused state and the new hard drive to the spare state. Then, it configures the new hard drive as a member drive and sets `drive_replacement` to the RD_COMPLETE (RAID module complete) state, triggering the EN module to change `drive_replacement` to the IDLE state, signifying the end of the new hard drive's processing flow.

[0057] Among them, such as Figure 5 The diagram shown illustrates the hard disk state switching logic provided in this application embodiment. The hard disk states include at least the unused state, candidate state, member state, and failed state. The unused state represents the initial state after the hard disk is inserted, before any state transition is performed. The candidate state indicates that the current hard disk meets the system requirements (such as disk format requirements) for joining a RAID array. The member state refers to the hard disk being added to or creating a RAID array, performing I / O functions as a member hard disk. The failed state occurs when the member disk no longer has the ability to respond to I / O commands due to various reasons such as a faulty disk, reaching the end of its lifespan, or poor contact between the hard disk and its slot due to foreign objects, and is therefore removed from the system, with its usage state set to the failed state.

[0058] Furthermore, in one embodiment, the array driver module is also configured to obtain the attribute information of the new hard disk when a hard disk configuration notification is received; and configure the new hard disk as a member disk according to the attribute information of the new hard disk.

[0059] The new hard drive's attribute information includes at least the hard drive capacity, hard drive type (NVMe, SAS SSD, or SASHDD), and rotational speed (HDD).

[0060] Specifically, in one embodiment, for a new hard drive, the protocol management module can assign a unique hard drive identifier to the new hard drive; based on the unique hard drive identifier, a target protocol initialization command is sent to the new hard drive to determine whether the new hard drive has the function of responding to the target protocol initialization command; if it is determined that the new hard drive has the function of responding to the target protocol initialization command, it is determined that the new hard drive has passed the legitimate authentication.

[0061] The unique identifier for a hard drive is its hard drive ID.

[0062] Specifically, the VL module assigns a hard drive ID to the new hard drive. Based on the assigned hard drive ID, the VL module sends a protocol initialization command compatible with the system to the new hard drive, such as IDENTIFYDEVICE for the SATA protocol and OPEN ADDRESSFRAME for the SAS protocol. This protocol initialization command requests the new hard drive to return its own hardware parameters. Based on these parameters, the VL module verifies whether the hard drive supports the system's dependent protocol specifications, such as data transmission formats and error handling mechanisms. If the new hard drive returns a response conforming to the protocol specifications within a specified time, such as complete and correctly formatted parameters, it is determined that it has responsiveness, meaning the communication link between the new hard drive and the system is working and the protocol is compatible, and the basic hardware functions are not disabled. At this point, the VL module confirms that the new hard drive has passed legitimate authentication and allows it to enter the subsequent array configuration process by sending a ready_to_spare status request to the RAID module.

[0063] The protocol management module can filter out hard drives that are incompatible with the system protocol in advance by verifying the response to protocol commands, thus preventing storage errors caused by communication format mismatch and ensuring the stability of the storage system.

[0064] Accordingly, in one embodiment, for the reuse of a faulty hard drive, the protocol management module can send a logic unit function detection command to the faulty hard drive to obtain the command response information fed back by the faulty hard drive in response to the logic unit function detection command; if the command response information indicates that the logic unit function of the faulty hard drive is normal, various characteristic function checks are performed on the faulty hard drive; if the faulty hard drive passes the various characteristic function checks, it is determined that the faulty hard drive has passed the health check.

[0065] Specifically, the VL module sends logical unit (LU) function test commands to the faulty hard drive, such as TESTUNIT READY for the SATA protocol and LOGICAL UNIT RESET for the SAS protocol. The logical unit is the core functional carrier of the hard drive's storage service provision. The test includes: determining whether the hard drive can respond to the command to verify whether the communication link has been restored; determining whether the logical unit status is Ready to rule out Busy or Error lock states due to residual faults; and determining whether the basic read / write channels are unobstructed, such as performing a small amount of data read / write test on free sectors. The command response information must clearly show that the logical unit has no abnormal response and is in a ready state before proceeding to the next step of testing. Multiple feature function checks are performed after confirming that the logical unit functions normally. The VL module further checks the hard drive's core feature functions, including whether its WRITE CACHE function is enabled and whether the power management mode supports system policies. When the logical unit function test passes and all feature function checks are normal, the VL module determines that the faulty hard drive has passed the health check, confirming that its fault has been repaired and it is ready to be reconnected to the array.

[0066] Specifically, in one embodiment, for a new hard disk, the array driver module can determine whether the new hard disk meets the preset configuration conditions based on the attribute information of the new hard disk; if the new hard disk meets the preset configuration conditions, the faulty hard disk is removed from the array member list to release the logical resources occupied by the faulty hard disk; the new hard disk is added to the array member list and logical resources are allocated to the new hard disk to make the new hard disk a member disk.

[0067] Specifically, the RAID module compares the new hard drive's attribute information with the system's preset configuration rules for capacity, type, and performance. The requirements are: new hard drive capacity ≥ faulty hard drive capacity; new hard drive type (e.g., NVMe, SAS SSD) consistent with the faulty hard drive; and if it's a mechanical hard drive (HDD), new drive speed ≥ original drive speed. When the new hard drive meets these conditions, it is determined that it meets the preset configuration requirements. After releasing the logical resources occupied by the faulty hard drive, these released resources are allocated to the new hard drive, allowing the new hard drive to take over subsequent I / O operations.

[0068] Among them, the verification of preset configuration conditions avoids incompatibility issues such as replacing large-capacity hard drives with small-capacity hard drives or mixing SSDs and HDDs, and prevents problems such as array degradation or data read / write errors caused by hardware incompatibility, ensuring that the array can operate stably after a new hard drive is connected.

[0069] Specifically, in one embodiment, the slot management module is also used to monitor the processing time of the protocol management module and the array driver module, and to suspend the current hard disk failure processing service when the protocol management module or the array driver module experiences a timeout.

[0070] It should be noted that by monitoring the processing time of the protocol management module and the array driver module, it is possible to prevent hard drive slots and logical resources from being locked due to the VL module or RAID module freezing, thereby ensuring the reliability of the fault hard drive handling process.

[0071] Specifically, after a new hard drive is inserted, the EN module detects that the swap bit of the insertion / removal flag is set to 1 (indicating a change in the physical state of the slot), and the hard drive status is failed. The EN actions include checking the slot insertion, copying the drive_id to the old_drive_id, and invalidating the drive_id. That is, copying the new hard drive ID (drive_id) to the location of the old hard drive ID (old_drive_id) and marking the new hard drive ID as invalid, so that if the new hard drive fails subsequent checks, it can resume using the old hard drive ID to restore the fault situation. The EN module opens the hard drive slot. Previously, after a hard drive failed due to a fault, the EN module would actively close the slot, clear the swap bit, and start a 90s timer to time the subsequent physical information reading and alarm information clearing process of the EN module. If it is completed within 90s, the VL module is notified to start, and an automatic managed timer is started to monitor the business processing time of the VL module. After the VL module is pulled back into the slot, a Validation request (health check or valid authentication) is executed: if there are alarms for this module on the original disk, the alarms are cleared; if the disk is being replaced, a new hard drive ID is assigned, and the hard drive management module performs a test process on the new disk (format check, logical unit ready status, etc.) according to the Validation request; if it is the original disk, a re-initialization process is executed (logical unit ready status, feature support check, etc.). If the above test or re-initialization process is completed, a RAID ready_to_spare status request is triggered. After receiving the ready_to_spare request, the RAID module compares the bad disk serial number and the new disk serial number to confirm whether it is a disk replacement process. If it is a disk replacement process, it needs to compare the disk capacity, type (NVMe, SAS SSD, or SAS HDD), and rotation speed (HDD). The disk usage status is changed from spare to member, and existing alarms are cleared (if the disk is replaced, the old disk alarms are cleared; if the disk is reinserted, the previous alarms are repaired). RAID non-redundancy alarms are also cleared. The RAID module automatic management completion time is sent to the EN module. The EN module clears the timeout timer, the automatic managed process ends, and each module performs a corresponding status reset.

[0072] Specifically, during operation, the system provided in this application records event logs to document the automatic hard drive management process. This mainly includes generating logs (success / failure logs), recording the automatic hard drive management process, and recording failure scenarios in the entire disk replacement process. These failure scenarios include timeouts in the protocol management module or array driver module, failures in the validation process on the new hard drive (e.g., failure of the new drive's initialization process, authentication failure), and discrepancies between the new and old hard drives' required information during the replacement process (the new hard drive does not meet preset configuration conditions).

[0073] For example, such as Figure 6 The diagram illustrates the interactive flow of an exemplary hard drive failure handling system provided in this application embodiment. When the EN module detects a failure in any hard drive in the RAID array, it issues a slot or hard drive failure alarm. The VL module issues a drive alarm. When the EN module detects that the failed hard drive is offline, the RAID module issues a RAID non-redundancy alarm, initiates RAID reconstruction, and completes the process. After the user switches a new hard drive to the original slot, the SWAP flag of the slot is updated to 1. The EN module opens the original slot and clears the alarms related to the original slot. The EN module confirms the online status of the new hard drive with the VL module (reports LOGIN). The VL module marks the hard drive status as unused, indicating that the hard drive has been initially connected but has not yet undergone verification or other operations. The EN module sends a new hard drive switch notification to the VL module so that the VL module can perform valid authentication of the new hard drive. If the new hard drive passes the authentication, it sends a hard drive configuration notification to the RAID module. When the RAID module confirms that a new hard drive has been switched, it removes the faulty hard drive from the array member list, configures the new hard drive as a member drive, clears the non-redundancy alarm, and sends a message to the EN module that automatic management has been successful. The EN module confirms that the hard drive replacement is complete, and the VL module sets a completion flag to indicate that the entire faulty hard drive handling process is over.

[0074] Specifically, in one embodiment, an intelligent model can be introduced into the protocol management module. This model establishes a fault diagnosis and repair strategy library based on historical hard drive failure data, such as fault type, hard drive parameters at the time of failure, and repair process time. When performing health checks on a faulty hard drive, the intelligent model not only performs routine logic unit function tests and feature function checks, but also analyzes the hard drive's historical fault records and real-time data generated during the current test, such as command response latency and specific failures in feature function checks. Based on the analysis results, the intelligent model dynamically adjusts the repair strategy. For example, if it is determined that the fault is due to temporary communication interference causing abnormal logic unit response, multiple restarts of the logic unit can be attempted as a repair operation; if a specific hardware feature is found to have a potential problem, the test parameters and repair methods for that feature can be adjusted accordingly to ensure that the hard drive can be quickly restored to use.

[0075] The hard drive failure handling system provided in this application includes a slot management module, a protocol management module, and an array driver module. By utilizing the slot management module and protocol management module, a series of tests are performed on the faulty hard drive after it has been plugged and unplugged from its original slot. After a series of tests determine that the faulty hard drive can continue to be used normally after plugging and unplugging, the array driver module performs state recovery on the faulty hard drive. The entire process is automatically triggered by the user plugging and unplugging the faulty hard drive from its original slot, resulting in a high degree of automation. Furthermore, once it is determined that the faulty hard drive can continue to be used normally after plugging and unplugging, the array driver module only needs to perform a simple state recovery. The configuration process for state recovery is simple, improving the efficiency of RAID hard drive failure handling and enabling the RAID to quickly return to normal business status, thereby ensuring RAID performance. Moreover, through an automated verification process, it ensures that only qualified hard drives can be added to the RAID, preventing secondary data risks caused by manual mis-plugging of faulty drives. If the re-plugged hard drive passes verification, it can automatically recover to the system, avoiding resource waste caused by missed operations in traditional maintenance. After the new hard drive passes verification, the system intelligently determines whether to immediately add it to the RAID to meet the optimal array member target, improving the redundancy and performance consistency of the storage pool. It automatically cleans up old hard drive configuration information to prevent invalid hard drives from consuming system resources. The entire process is also recorded in the system log for easy auditing and root cause analysis.

[0076] Through the above description of the embodiments, those skilled in the art can clearly understand that the system according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0077] The embodiments of this application also provide a hard disk failure handling method, which is applied to the hard disk failure handling system provided in the above embodiments.

[0078] like Figure 7 The diagram shown is a flowchart illustrating a hard disk failure handling method provided in an embodiment of this application. The hard disk failure handling method includes:

[0079] Step 701: After the user performs the plug-and-play process for the faulty hard drive in the original slot, the original slot is opened and the physical information of the faulty hard drive is read; the physical information includes at least the interface status signal, current data and signal error rate of the original slot.

[0080] Step 702: Determine whether the original slot has been successfully opened based on the interface status signal, current data, and signal error rate of the original slot.

[0081] Step 703: If the original slot is successfully opened, clear the alarm information of the original slot;

[0082] Step 704: Perform a health check on the faulty hard drive. If the faulty hard drive passes the health check, reuse the original logical resources of the faulty hard drive in the array member list to restore the state of the faulty hard drive and restore it to a member disk.

[0083] The array member list includes the allocation relationship between each member disk and logical resources in the independent redundant disk array. The logical resources include at least the unique identifier of the hard disk, the unique identifier of the logical slot, and the address mapping information.

[0084] For a description of the features in the embodiments corresponding to the hard disk failure handling method, please refer to the relevant descriptions in the embodiments corresponding to the hard disk failure handling system, which will not be repeated here.

[0085] Embodiments of this application also provide an electronic device, such as... Figure 8 The diagram shown is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, including a processor 10 and a memory 20. The memory 20 stores a computer program, and the processor 10 is configured to run the computer program to execute the steps in any of the above embodiments of the hard disk failure handling method.

[0086] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the hard disk failure handling method when it is run.

[0087] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0088] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the hard disk fault handling method.

[0089] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the hard disk failure handling method.

[0090] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0091] The above provides a detailed description of a hard disk fault handling system, method, electronic device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A hard disk failure handling system, characterized in that, include: Slot management module, protocol management module, and array driver module; The slot management module includes a current sensor and a signal quality sensor; The slot management module is used to open the original slot after the user has plugged and unplugged the faulty hard drive in the original slot. It reads the physical information of the faulty hard drive through the current sensor and the signal quality sensor. The physical information includes at least the interface status signal, current data and signal error rate of the original slot. Based on the interface status signal, current data and signal error rate of the original slot, it determines whether the original slot has been successfully opened. If the original slot has been successfully opened, it clears the alarm information of the original slot and sends a hard drive online notification to the protocol management module. The protocol management module is used to perform a health check on the faulty hard drive when it receives the hard drive online notification, and send a hard drive recovery notification to the array driver module if the faulty hard drive passes the health check. The array driver module is used to reuse the original logical resources of the faulty hard disk in the array member list when it receives the hard disk recovery notification, so as to restore the state of the faulty hard disk and restore the faulty hard disk as a member disk. The array member list includes the allocation relationship between each member disk and logical resources in the independent redundant disk array, and the logical resources include at least a unique hard disk identifier, a unique logical slot identifier, and address mapping information.

2. The hard disk fault handling system according to claim 1, characterized in that, The slot management module is also used for: If the original slot fails to open, a slot switching notification is sent to the user so that the user can respond to the slot switching notification, select a target slot from the available slots, and perform hard drive plugging / unplugging in the target slot.

3. The hard disk fault handling system according to claim 2, characterized in that, The slot management module is also used for: After the user performs the plug-and-unplug procedure on the faulty hard drive in the target slot, the target slot is opened, and the physical information of the faulty hard drive is read through the current sensor and the signal quality sensor; the physical information includes at least the interface status signal, current data, and signal error rate of the target slot. Based on the interface status signal, current data, and signal error rate of the target slot, determine whether the target slot has been successfully opened; If the target slot is successfully opened, a hard drive online notification is sent to the protocol management module.

4. The hard disk fault handling system according to claim 1, characterized in that, The protocol management module is also used for: If the faulty hard drive fails the health check, a hard drive switching notification will be sent to the user so that the user can respond to the notification and switch to a new hard drive in the original slot.

5. The hard disk fault handling system according to claim 4, characterized in that, The slot management module is also used for: After the user performs the process of switching a new hard drive to the original slot, the original slot is restarted, and the physical information of the faulty hard drive is read through the current sensor and the signal quality sensor; the physical information includes at least the interface status signal, current data and signal error rate of the original slot; Based on the interface status signal, current data, and signal error rate of the original slot, determine whether the original slot has been successfully restarted; If the original slot is successfully restarted, the alarm information for the original slot is cleared, and a new hard drive switching notification is sent to the protocol management module.

6. The hard disk fault handling system according to claim 5, characterized in that, The protocol management module is also used for: When the notification of the new hard drive switch is received, the new hard drive is authenticated. Once the new hard drive has been successfully authenticated, a hard drive configuration notification is sent to the array driver module.

7. The hard disk fault handling system according to claim 6, characterized in that, The array driving module is also used for: When the hard drive configuration notification is received, the attribute information of the new hard drive is obtained; Based on the attribute information of the new hard drive, the new hard drive is configured to be used as a member drive.

8. The hard disk fault handling system according to claim 6, characterized in that, The protocol management module is specifically used for: Assign a unique hard drive identifier to the new hard drive; Based on the unique identifier of the hard disk, a target protocol initialization command is sent to the new hard disk to determine whether the new hard disk has the function of responding to the target protocol initialization command; If it is determined that the new hard disk has the function of responding to the target protocol initialization command, then it is determined that the new hard disk has passed legitimate authentication.

9. The hard disk failure handling system according to claim 1, characterized in that, The protocol management module is specifically used for: Send a logic unit function test command to the faulty hard disk to obtain the command response information fed back by the faulty hard disk in response to the logic unit function test command; If the logical unit function of the faulty hard disk is normal as indicated by the instruction response information, various characteristic function checks are performed on the faulty hard disk. If the faulty hard drive passes the various feature function checks, it is determined that the faulty hard drive has passed the health check.

10. The hard disk fault handling system according to claim 7, characterized in that, The array driving module is specifically used for: Based on the attribute information of the new hard drive, determine whether the new hard drive meets the preset configuration conditions; If the new hard drive meets the preset configuration conditions, the faulty hard drive is removed from the array member list to release the logical resources occupied by the faulty hard drive; Add the new hard drive to the array member list and allocate logical resources to the new hard drive to make it a member disk.

11. The hard disk fault handling system according to claim 1, characterized in that, The slot management module is also used for: The processing time of the protocol management module and the array driver module is monitored. When the protocol management module or the array driver module times out, the current hard disk failure processing service is terminated.

12. A hard disk failure handling method, characterized in that, include: After the user performs the process of plugging and unplugging the faulty hard drive in the original slot, the original slot is opened and the physical information of the faulty hard drive is read. The physical information includes at least the interface status signal, current data, and signal error rate of the original slot. Based on the interface status signal, current data, and signal error rate of the original slot, determine whether the original slot has been successfully opened; If the original slot is successfully opened, clear the alarm information for the original slot; A health check is performed on the faulty hard drive. If the faulty hard drive passes the health check, the original logical resources of the faulty hard drive are reused in the array member list to restore the state of the faulty hard drive and restore it to a member disk. The array member list includes the allocation relationship between each member disk and logical resources in the independent redundant disk array, and the logical resources include at least a unique hard disk identifier, a unique logical slot identifier, and address mapping information.

13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the hard disk failure handling method as described in claim 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the hard disk failure handling method as described in claim 12.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the hard disk failure handling method as described in claim 12.

Citation Information

Patent Citations

  • Bad hard disk disk repairing method

    CN111966541A

  • Method, device and equipment for automatically repairing hard disk fault of storage system and medium

    CN117290141A