Fault reconstruction method and device for mechanical hard disk, equipment and storage medium

By reconstructing the data in the mechanical hard disk failure area in the target reconstruction area, the problems of low reconstruction efficiency and reduced host IO performance caused by mechanical hard disk failure in the RAID group are solved, and the effect of improving the reconstruction efficiency of the RAID group and the utilization rate of the faulty hard disk is achieved.

CN120233949APending Publication Date: 2025-07-01SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510363075.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Mechanical hard disks are prone to damage or bad channels in the RAID group, resulting in low reconstruction efficiency of RAID group, reduced host IO performance, and low utilization rate of faulty disks.

Method used

By obtaining the fault area information of the currently failed hard disk in the disk array group, the target reconstruction area is determined, and the data of the fault area in the faulty hard disk is reconstructed in this area, and the corresponding relationship between the target reconstruction area and the fault area is recorded.

Benefits of technology

It improves the reconstruction efficiency of the RAID group, improves the utilization rate of the faulty mechanical hard disk, reduces costs, and thus improves the IO performance of the host.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233949A_ABST
    Figure CN120233949A_ABST
Patent Text Reader

Abstract

The invention discloses a fault reconstruction method, device and equipment for a mechanical hard disk and a storage medium, and relates to the technical field of storage, and the method comprises the following steps: obtaining fault area information of a current fault hard disk in a disk array group; wherein the fault area information is information of a strip of a disk array group where a fault occurring in the current fault hard disk is located; the fault area information comprises a fault hard disk identifier and a fault strip identifier; determining a target reconstruction region corresponding to the fault region information; wherein the target reconstruction area is an area in the current mechanical hard disk; reconstructing data of a fault area in the current fault hard disk in the target reconstruction area, and recording a corresponding relation between the target reconstruction area and the fault area; reconstruction is still carried out on the faulty mechanical hard disk, a new disk does not need to be inserted, the reconstruction efficiency of the RAID group and the utilization rate of the faulty mechanical hard disk are improved, the cost is greatly reduced, and therefore the IO performance of a host is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technologies, and particularly to a method, device, equipment and storage medium for fault reconstruction of a mechanical hard disk. Background Art

[0002] In a data center server system, in order to improve the read / write performance, scalability and security of hard disk storage, server manufacturers will integrate a RAID (Redundant Array of Independent Disks) board based on a PCIE (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) interface in the server to manage hard disks using hardware RAID. The RAID board is generally used to support HDDs (Hard Disk Drives, mechanical hard disks) that use the SATA (Serial Advanced Technology Attachment, a computer bus interface standard) protocol. However, with the rapid development of computer technology and the sharp increase in data processing requirements, the performance of mechanical hard disks is increasingly difficult to meet the needs.

[0003] Since the magnetic head of a mechanical hard disk floats above a rapidly rotating disk during operation, and the distance between the magnetic head and the disk is small, any vibration may cause the magnetic head to collide with the disk, resulting in the mechanical hard disk being prone to damage or bad sectors. This phenomenon causes the system to frequently layout (Layout) and insert new mechanical hard disks (such as Figure 1 the hot spare disk in to perform data reconstruction for the entire RAID group. However, usually only a few stripes in the faulty area are affected, resulting in low efficiency of RAID group reconstruction. Moreover, the data reconstruction of the entire disk causes the system to be unable to process host I / O in a timely manner, reducing the host I / O (Input Output) performance and the utilization rate of the faulty disk.

[0004] Therefore, how to improve the reconstruction efficiency of the RAID group, improve the utilization rate of faulty mechanical hard disks, and thus enhance the host I / O performance is an urgent problem to be solved today. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, device, equipment and computer-readable storage medium for fault reconstruction of a mechanical hard disk to improve the reconstruction efficiency of the RAID group, improve the utilization rate of faulty mechanical hard disks, and thus enhance the host I / O performance.

[0006] To solve the above technical problems, the present invention provides a method for fault reconstruction of a mechanical hard disk, including:

[0007] Obtain the fault area information of the currently faulty hard disk in the disk array group; wherein, the currently faulty hard disk is any mechanical hard disk in the disk array group, and the fault area information is the information of a stripe in the disk array group where the fault occurs in the currently faulty hard disk; the fault area information includes a faulty hard disk identifier and a faulty stripe identifier;

[0008] Determine the target reconstruction area corresponding to the fault area information; wherein, the target reconstruction area is an area in the current mechanical hard disk;

[0009] Reconstruct the data in the fault area of the currently faulty hard disk in the target reconstruction area, and record the corresponding relationship between the target reconstruction area and the fault area; wherein, the fault area is the stripe area corresponding to the fault area information.

[0010] On the other hand, the target reconstruction area is the stripe area of any backup stripe in the pre - device backup area of the disk array group.

[0011] On the other hand, the determining the target reconstruction area corresponding to the fault area information includes:

[0012] Determine the stripe identifier of the target reconstruction area.

[0013] On the other hand, the faulty hard disk identifier is the encoding of the currently faulty hard disk, and the faulty stripe identifier is the encoding of the stripe corresponding to the fault area information; the determining the stripe identifier of the target reconstruction area includes:

[0014] According to the fault area information, determine the current number of faults; wherein, the current number of faults is the number of faults of the currently faulty hard disk or the number of stripe faults of the disk array group;

[0015] Judge whether the current number of faults reaches the maximum number of faults;

[0016] If not, determine the stripe identifier of the target reconstruction area as the reconstruction - used stripe encoding; wherein, the initial value of the reconstruction - used stripe encoding is 0;

[0017] Add 1 to the reconstruction - used stripe encoding and update the reconstruction - used stripe encoding;

[0018] The recording the corresponding relationship between the target reconstruction area and the fault area includes:

[0019] Record the corresponding relationship between the stripe identifier of the target reconstruction area and the encoding of the stripe corresponding to the fault area information.

[0020] On the other hand, the recording the corresponding relationship between the target reconstruction area and the fault area includes:

[0021] The current entry in the fault area address mapping table records the correspondence between the target reconstruction area and the fault area; wherein, the current entry is further used to record the reconstruction status identifier of the target reconstruction area, and the reconstruction status identifier is the in-reconstruction status or the reconstruction completed status.

[0022] On the other hand, the method further includes:

[0023] According to the received host input / output command, obtain the information of the area to be accessed corresponding to the host input / output command; wherein, the information of the area to be accessed is the address information of each area to be accessed by the host input / output command, and the area to be accessed is a stripe area of a mechanical hard disk in a disk array group to be accessed.

[0024] According to the information of the area to be accessed, determine whether there is a recorded fault area in each area to be accessed.

[0025] If so, replace the address information of the target access area in the information of the area to be accessed with the address information of the corresponding target reconstruction area; wherein, the target access area is the area to be accessed that is the same as the fault area.

[0026] On the other hand, the information of the area to be accessed includes: the access disk array group code, the access stripe code, and the starting mechanical hard disk code; the access stripe code is io_slba / (Chunk_size*Data_disk_num), and the starting mechanical hard disk code is (io_slba / Chunk_size)%Data_disk_num; wherein, io_slba is the device-side starting logical block address of the host input / output command, Chunk_size is the preset area size, and Data_disk_num is the preset number of data disks.

[0027] The determining whether there is a recorded fault area in each area to be accessed according to the information of the area to be accessed includes:

[0028] Determine the starting mechanical hard disk code as the current access hard disk code.

[0029] Judge whether the current access hard disk code is less than or equal to the number of access disks; wherein, the number of access disks is [(io_slba+io_nlb) / Chunk_size] %Data_disk_num, and io_nlb is the number of device-side logical blocks of the host input / output command.

[0030] If it is less than or equal to the number of accessed disks, then, using the current accessed hard disk code and the disk array group code as indexes, check whether there is a target entry in the entry corresponding to the current accessed hard disk code in the fault area address mapping table; wherein, the fault stripe identifier in the target entry is the accessed stripe code.

[0031] If there is the target entry, then determine the area to be accessed corresponding to the current accessed hard disk code as the recorded fault area.

[0032] The present invention also provides a fault reconstruction device for a mechanical hard disk, including:

[0033] An acquisition module, configured to acquire fault area information of a current faulty hard disk in a disk array group; wherein, the current faulty hard disk is any mechanical hard disk in the disk array group, and the fault area information is information of a stripe in the disk array group where a fault occurs in the current faulty hard disk; the fault area information includes a faulty hard disk identifier and a fault stripe identifier.

[0034] A determination module, configured to determine a target reconstruction area corresponding to the fault area information; wherein, the target reconstruction area is an area in the current mechanical hard disk.

[0035] A reconstruction module, configured to reconstruct data in the fault area of the current faulty hard disk in the target reconstruction area, and record the corresponding relationship between the target reconstruction area and the fault area; wherein, the fault area is the stripe area corresponding to the fault area information.

[0036] The present invention also provides a fault reconstruction device for a mechanical hard disk, including:

[0037] A memory, configured to store a computer program.

[0038] A processor, configured to implement the steps of the fault reconstruction method for the mechanical hard disk as described above when executing the computer program.

[0039] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of the fault reconstruction method for the mechanical hard disk as described above when being executed by a processor.

[0040] A method for reconstructing a failed mechanical hard disk provided by the present invention includes: obtaining failure area information of a currently failed hard disk in a disk array group; wherein, the currently failed hard disk is any mechanical hard disk in the disk array group, and the failure area information is information of a stripe in the disk array group where the failure occurs in the currently failed hard disk; the failure area information includes a failed hard disk identifier and a failed stripe identifier; determining a target reconstruction area corresponding to the failure area information; wherein, the target reconstruction area is an area in the current mechanical hard disk; reconstructing data in the failure area of the currently failed hard disk in the target reconstruction area, and recording the corresponding relationship between the target reconstruction area and the failure area; wherein, the failure area is the stripe area corresponding to the failure area information.

[0041] It can be seen that by reconstructing the data in the failure area of the currently failed hard disk in the target reconstruction area and recording the corresponding relationship between the target reconstruction area and the failure area, the present invention enables the reconstruction to still be performed on the failed mechanical hard disk without inserting a new disk, improving the reconstruction efficiency of the RAID group and the utilization rate of the failed mechanical hard disk, greatly reducing the cost, and thus improving the IO performance of the host. In addition, the present invention also provides a device, equipment, and computer-readable storage medium for reconstructing a failed mechanical hard disk, which also have the above beneficial effects. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0043] Figure 1 A schematic diagram of a reconstruction solution for a mechanical hard disk in a failure scenario in the related art;

[0044] Figure 2 A flowchart of a method for reconstructing a failed mechanical hard disk provided by an embodiment of the present invention;

[0045] Figure 3 A schematic diagram of a reconstruction solution for a mechanical hard disk in a failure scenario provided by an embodiment of the present invention;

[0046] Figure 4 A flowchart of another method for reconstructing a failed mechanical hard disk provided by an embodiment of the present invention;

[0047] Figure 5 A flowchart of the IO processing process of a method for reconstructing a failed mechanical hard disk provided by an embodiment of the present invention;

[0048] Figure 6A schematic flowchart of another IO processing procedure provided by an embodiment of the present invention;

[0049] Figure 7 A structural block diagram of a failure reconstruction device for a mechanical hard disk provided by an embodiment of the present invention;

[0050] Figure 8 A structural schematic diagram of a failure reconstruction device for a mechanical hard disk provided by an embodiment of the present invention. Detailed implementation manners

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] Please refer to Figure 2 , Figure 2 A flowchart of a failure reconstruction method for a mechanical hard disk provided by an embodiment of the present invention. The method may include:

[0053] Step 101: Obtain failure area information of the current failed hard disk in the disk array group; wherein, the current failed hard disk is any mechanical hard disk in the disk array group, and the failure area information is information of a stripe in the disk array group where the failure occurs in the current failed hard disk; the failure area information includes a failed hard disk identifier and a failed stripe identifier.

[0054] It can be understood that the disk array (RAID) group in this embodiment may be a disk group composed of multiple independent disks (such as HDDs or HDDs and other disks) to utilize the additive effect generated by individual disks to improve the performance of the entire disk system. This embodiment takes the failure reconstruction of a certain mechanical hard disk (i.e., the current failed hard disk) in a disk array group in a stripe area (i.e., the area of the stripe on the current failed hard disk) as an example for display. For the failure reconstruction of other mechanical hard disks in the disk array group and the failure reconstruction of mechanical hard disks in other disk array groups, the same or similar methods as those provided in this embodiment may be used for implementation, and this embodiment does not make any restrictions thereon.

[0055] Correspondingly, the current faulty hard disk in this embodiment can be any mechanical hard disk in the disk array group where a strip in the disk array group fails. The faulty area information provided in this embodiment can be the information of a strip area where the fault occurs in the current faulty hard disk. The faulty area information in this embodiment can include a faulty hard disk identifier for representing the faulty hard disk and a faulty strip identifier for representing the strip where the fault is located, so as to determine a strip area where the current faulty hard disk fails, that is, the faulty area to be reconstructed, through the faulty hard disk identifier and the faulty strip identifier; the faulty area information can also include a faulty array group identifier to distinguish different disk array groups.

[0056] Step 102: Determine the target reconstruction area corresponding to the faulty area information; wherein, the target reconstruction area is an area in the current mechanical hard disk.

[0057] Among them, the target reconstruction area in this embodiment can be an area in the current mechanical hard disk used to reconstruct the faulty area represented by the faulty area information. For the specific location of the target reconstruction area in this embodiment, it can be set by the designer according to the practical scenario and user requirements. For example, the target reconstruction area is the strip area of any backup strip in the preset backup area of the disk array group; that is to say, in this embodiment, the disk backup area (preset backup area) of the preset mechanical hard disk can be used as the area for reconstructing data, as Figure 3 shown, when Disk 1 (i.e., the current faulty hard disk) of Strip 0 (Strip_id = 0) in the RAID group fails, at this time, it is necessary to reconstruct and replace through the corresponding disk in the backup area (i.e., the preset backup area) to complete the reconstruction task of Strip 0. The reconstructed area changes from a hot spare disk in the related art to a single strip, and at the same time, there is no need to add a new mechanical hard disk. The target reconstruction area can also be other areas in the current faulty hard disk. As long as the current faulty hard disk can use its own area to reconstruct the faulty area, this embodiment does not make any restrictions on this.

[0058] Correspondingly, for the specific method of determining the target reconstruction area corresponding to the faulty area information in this step, it can be set by the designer. For example, the strip identifier (i.e., the strip identifier of the backup strip, such as the strip code) of the target reconstruction area corresponding to the faulty area information can be determined to find the location of the target reconstruction area using this strip identifier. As long as the location of the target reconstruction area used to reconstruct the faulty area of the faulty area information can be determined, this embodiment does not make any restrictions on this.

[0059] Correspondingly, for the specific process of determining the target reconstruction area corresponding to the fault area information in this step, it can be set by the designer himself. For example, when the faulty hard disk is identified as the encoding of the current faulty hard disk and the faulty stripe is identified as the encoding of the stripe corresponding to the fault area information, the process of determining the stripe identifier of the target reconstruction area corresponding to the fault area information may include: determining the current number of faults according to the fault area information; determining whether the current number of faults reaches the maximum number of faults; if not, determining the stripe identifier of the target reconstruction area as the stripe encoding for reconstruction use; where the initial value of the stripe encoding for reconstruction use is 0; adding 1 to the stripe encoding for reconstruction use and updating the stripe encoding for reconstruction use. Correspondingly, the recorded correspondence between the target reconstruction area and the fault area may be the correspondence between the stripe identifier (i.e., stripe encoding) of the target reconstruction area and the encoding of the stripe corresponding to the fault area information.

[0060] It can be understood that the above current number of faults may be the number of faults of the current faulty hard disk or the number of stripe faults of the disk array group. For example, when the current number of faults is the number of stripe faults of the disk array group, the number of faulty stripes in the disk array group can be restricted to be less than the maximum number of faults; when the current number of faults is the number of faults of the current faulty hard disk, the number of stripes where each mechanical hard disk in the disk array group fails can be restricted to be less than the maximum number of faults.

[0061] Correspondingly, for the case where the above current number of faults reaches the maximum number of faults, this process can be directly ended and the reconstruction process provided in this embodiment is no longer performed; or corresponding fault prompt information can be output to prompt the information of excessive number of faults.

[0062] For example, in this embodiment, the correspondence between the target reconstruction area and the fault area can be recorded by using a fault area address mapping table. As shown in Table 1 below, an entry in the fault area address mapping table can record the hard disk encoding (Disk_id) corresponding to the target reconstruction area, the stripe encoding (Bad_strip_id) corresponding to the target reconstruction area, and the stripe encoding (new_strip_id) corresponding to the target reconstruction area.

[0063] Table 1 Fault Area Address Mapping Table

[0064]

[0065] As shown in Table 1, Raid_id can be the encoding of the disk array group. Max_bad_num can be the maximum number of faulty stripes (i.e., the above-mentioned maximum number of faults) to limit the maximum number of stripes that can fail. Cur_bad_num can be the above-mentioned current number of faults, such as the number of stripe faults of the disk array group counted or the number of faults of the currently faulty hard disk. Used_strip_id (i.e., the above-mentioned reconstruction-used stripe encoding) can record the starting stripe ID of the backup area (pre-equipment backup area) of all disks in the RAID group, and can default to the next stripe ID after the RAID group is created; for example, Figure 4 as shown, at the first fault, Used_strip_id can be Raid_size / Data_disk_num. Data_disk_num can be the number of data disks. disk_num can include data disks and parity disks. Chunk_size can represent the size of each stripe area (such as the faulty area or the target reconstruction area) in the mechanical hard disk, such as the number of blocks. Valid can indicate whether the entry is valid. Bad_strip_id can represent the stripe ID of the faulty area; new_strip_id can represent the stripe ID of the area where the data is stored after the reconstruction of the faulty area (i.e., the target reconstruction area), which is used for subsequent IO processing. Disk_id can represent the ID of the faulty mechanical hard disk (such as the currently faulty hard disk), and can be specified as consecutive integers starting from 0; Rebuild_status is the above-mentioned reconstruction status flag, which is used to indicate whether the target reconstruction area has been reconstructed; for example, the reconstruction status flag can be the reconstruction completed status (such as 1, indicating reconstruction completed) or the reconstructing status (such as 0, indicating that it is being or about to be reconstructed).

[0066] Among them, the faulty area address mapping table can be stored and indexed according to the granularity of RAID. At the same time, entries can be searched through the disk logical ID (Disk_id) inside each RAID group. The number of entries for each disk depends on the maximum number of stripes that can fail (Max_bad_num) and can be configured according to requirements.

[0067] Correspondingly, the configuration of the fault area address mapping table can be indexed by Raid_id and Disk_id, and the calculation formula for the storage space address is [(Disk_id * Max_bad_num * 128 bits) + 80] * raid_id; when the first bad disk appears in the RAID group, Used_strip_id needs to be configured according to the size of the RAID group, and its size is determined by the value of Raid_size. At the same time, the configuration information of the RAID is obtained, and information such as Data_disk_num, disk_num, Chunk_size, and Raid_size is recorded in the address space corresponding to the current RAID group.

[0068] Correspondingly, considering the size and performance of the mechanical hard disk, when a partial failure occurs in a single mechanical hard disk in a stripe, the method provided in this embodiment can remap and reconstruct the entire area of the mechanical hard disk under this stripe. When multiple consecutive stripe failures occur, stripe splitting is required. Each stripe can occupy an entry in the fault area address mapping table, and a larger fault area is represented by multiple entries; at the same time, the maximum number of faulty stripes (Max_bad_num) can be set. After exceeding Max_bad_num, the reconstruction process will not be entered.

[0069] Step 103: Reconstruct the data in the fault area of the current faulty hard disk in the target reconstruction area, and record the corresponding relationship between the target reconstruction area and the fault area; where the fault area is the stripe area corresponding to the fault area information.

[0070] It can be understood that in this embodiment, the data in the fault area of the current faulty hard disk can be reconstructed into the target reconstruction area, and by recording the corresponding relationship between the target reconstruction area and the fault area, the data in the fault area can be read from the target reconstruction area during subsequent IO processing.

[0071] Correspondingly, since IO processing cannot be performed during the reconstruction process, it is necessary to wait for the reconstruction to complete before processing. In this embodiment, during the process of recording the corresponding relationship between the target reconstruction area and the fault area in the current entry of the fault area address mapping table, the reconstruction status flag of the target reconstruction area can also be recorded in the current entry. The reconstruction status flag is the in-reconstruction status or the reconstruction-completed status, so as to facilitate determining the reconstruction status during subsequent IO processing.

[0072] Correspondingly, the current entry in the fault area address mapping table can also record a validity flag (such as Valid) to indicate whether the current entry is valid.

[0073] Such as Figure 4As shown, when the current number of failures (Cur_bad_num) has not reached the maximum number of failures (Max_bad_num), the current number of failures can be incremented by 1 to facilitate the determination of the current number of failures during the next failure; by determining the stripe identifier (new_strip_id) of the target reconstruction area as the reconstruction-used stripe code (Used_strip_id), it enables the subsequent writing of the data in the failed area (i.e., the Disk_id area of the Bad_strip_id stripe) of the current failed hard disk into the target reconstruction area corresponding to new_strip_id; by updating Used_strip_id to the value after incrementing it by 1, it facilitates the use of Used_strip_id during the next failure; before specific data reconstruction, the validity identifier can be configured as valid (such as 1), the reconstruction status identifier as the ongoing reconstruction status (such as 1), and after the data reconstruction is completed, the reconstruction status identifier is configured as the reconstruction-completed status (such as 0) to facilitate subsequent IO processing.

[0074] That is to say, for a single mechanical hard disk with multiple failure points, it is necessary to count the number of currently failed areas, and configure new_strip_id as Used_strip_id, and Used_strip_id increases as the number of failed areas increases. Before starting the reconstruction, it is necessary to set the entry corresponding to the current disk as valid, the reconstruction status identifier as 1, and at the same time, map the stripe ID of the failed area to the new area stripe ID. After the reconstruction task is completed, reset the reconstruction status identifier to 0, indicating that the reconstruction has been completed.

[0075] In this embodiment, the embodiment of the present invention reconstructs the data in the failed area of the current failed hard disk in the target reconstruction area and records the corresponding relationship between the target reconstruction area and the failed area, so that the reconstruction still occurs on the failed mechanical hard disk without the need to insert a new disk, improving the reconstruction efficiency of the RAID group and the utilization rate of the failed mechanical hard disk, greatly reducing the cost, and thus enhancing the IO performance of the host.

[0076] Based on the above embodiment, the method provided by the embodiment of the present invention may further include the host IO processing process corresponding to failure reconstruction to utilize the reconstructed data to implement the processing of host IO commands. Specifically, please refer to Figure 5 , Figure 5 which is the flowchart of the IO processing process of a failure reconstruction method for a mechanical hard disk provided by the embodiment of the present invention. The method may include:

[0077] Step 201: According to the received host input / output command, obtain the information of the area to be accessed corresponding to the host input / output command.

[0078] Among them, the information of the area to be accessed is the address information of each area to be accessed in the host input / output command, and the area to be accessed is a strip area of a mechanical hard disk in a disk array group to be accessed.

[0079] It can be understood that the information of the area to be accessed in this embodiment can be the address information of each area to be accessed obtained by parsing the host IO (input / output) command. For example, when the host IO command in this embodiment is a cross-strip IO command, it can be split into corresponding single-strip IO commands; obtain the parameters related to the single-strip IO command, such as Figure 6 the raid_id (raid ID) of the disk array group to be accessed, the device-side starting logical block address (io_slba) of the host IO command (or single-strip IO command), and the number of device-side logical blocks of the host IO command (or single-strip IO command); through io_slba / (Chunk_size*Data_disk_num), determine the access strip ID (such as Figure 6 the strip_id in it); through (io_slba / Chunk_size)%Data_disk_num, determine the starting mechanical hard disk ID (such as Figure 6 the disk_id in it); thus, by comparing (io_slba / Chunk_size)%Data_disk_num with the continuously increasing value after the starting mechanical hard disk ID, determine each mechanical hard disk to be accessed by the single-strip IO command.

[0080] Step 202: According to the information of the area to be accessed, determine whether there are recorded faulty areas in each area to be accessed; if so, enter Step 203.

[0081] Among them, in this step, by comparing the information of the area to be accessed with the relevant information of the recorded faulty areas (such as the correspondence between the above-mentioned target reconstruction area and the faulty area), determine whether there are reconstructed faulty areas in the area to be accessed. Thus, when there are, in Step 203, replace the address information of the faulty area to be accessed with the address information of the corresponding target reconstruction area, so as to complete the IO processing using the corresponding target reconstruction area; when there are none, it means that there is no involvement of faulty areas, and normal IO processing can be performed.

[0082] Correspondingly, when the area information to be accessed includes the access disk array group code, the access stripe code of a single-strip IO command, and the starting mechanical hard disk code, this step may include determining the starting mechanical hard disk code as the current access hard disk code; judging whether the current access hard disk code is less than or equal to the number of accessed disks; where the number of accessed disks is [(io_slba + io_nlb) / Chunk_size] % Data_disk_num, and io_nlb is the number of device-side logical blocks of the host input / output command; if it is less than or equal to the number of accessed disks, then using the current access hard disk code and the disk array group code as indexes, check whether there is a target entry in the entry corresponding to the current access hard disk code in the fault area address mapping table; where the fault stripe identifier in the target entry is the access stripe code; if there is a target entry, then determine the area to be accessed corresponding to the current access hard disk code as the recorded fault area.

[0083] Correspondingly, after the entry corresponding to the current access hard disk code is found, the current access hard disk code can be incremented by 1 to update the current access hard disk code, and the step of judging whether the current access hard disk code is less than or equal to the number of accessed disks is executed to implement the search for the entries corresponding to all the mechanical hard disks to be accessed by a single-strip IO command.

[0084] Step 203: Replace the address information of the target access area in the area information to be accessed with the address information of the corresponding target reconstruction area; where the target access area is the area to be accessed that is the same as the fault area.

[0085] It can be understood that in this step, by replacing the address information of the target access area in the area information to be accessed with the address information of the corresponding target reconstruction area, when the data to be accessed during the IO processing is in the fault area, data access can be performed in the corresponding target reconstruction area to utilize the reconstructed target reconstruction area for IO processing.

[0086] Correspondingly, as Figure 6 shown, before replacing the address information of the target access area in the area information to be accessed with the address information of the corresponding target reconstruction area in this step, it is also possible to detect whether the reconstruction status flag (Rebuild_status) corresponding to the current target access area is in the reconstruction completed state (such as 1); if so, replace the stripe id (i.e., Bad_strip_id) of the current target access area in the area information to be accessed with the new_strip_id in the corresponding entry; if not, it is possible to delay and wait for the reconstruction status flag to become the reconstruction completed state.

[0087] Correspondingly, before replacing the address information of the target access area in the to-be-accessed area information with the address information of the corresponding target reconstruction area, it is also possible to detect whether the validity flag (such as Valid) corresponding to the current target access area is in a valid state (such as 1), so as to perform subsequent address replacement and IO processing only when the corresponding entry is valid; when the corresponding entry is invalid, this process can be directly ended or an alarm process can be performed.

[0088] For example, as Figure 3 shown, taking raid5 with 4 mechanical hard disks (disks) as an example, when the first failure occurs on disk 1 with stripe code 0, after entering the reconstruction process, set raid_id = 0, Max_bad_num to 100, Cur_bad_num to 0, the starting space address of the available disk is the starting address strip_n of the backup area, the number of data disks is 3, the number of disks is 4, and chunk_size = 64KB; since it is disk 1 with disk code 0 corresponding to stripe code 0 that has failed, add an entry at the position corresponding to disk code 0, set valid = 1, the failed stripe code is 0, the newly allocated stripe code is strip_n, and at the same time set the reconstruction status flag to 0 (reconstructing status); during the reconstruction process, write the reconstructed data to the area of disk 1 with disk code 0 corresponding to stripe strip_n; after the reconstruction is completed, adjust the reconstruction status flag to 1, indicating that the reconstruction has been completed.

[0089] Correspondingly, when receiving the host IO command sent by the host and parsing to obtain io_slba = 0 and io_size = 64KB, and its related raid_id = 0, it is obvious that this IO only involves stripe 0 and the disk code is 0; traverse the stripe codes (i.e., the failed stripe codes) in all the entries of the failed area mapping table under the first disk (disk 1). If it matches the stripe code and rebuild_status is also 1, it means that the IO can be processed at this time; modify io_slba = 0 in the IO information to io_slba = strip_n * 3 * chunk_size, and after the IO processing is completed, return a message indicating successful execution to the host.

[0090] In this embodiment, the present invention replaces the failed area to be accessed by the host IO command with the corresponding target access area, and can use the reconstructed target access area to replace the failed area for IO processing, improving the host's IO performance.

[0091] Corresponding to the above method embodiments, an embodiment of the present invention further provides a fault reconstruction device for a mechanical hard disk. A fault reconstruction device for a mechanical hard disk described below can be correspondingly referred to the method for reconstructing a mechanical hard disk described above.

[0092] Please refer to Figure 7 , Figure 7 which is a structural block diagram of a fault reconstruction device for a mechanical hard disk provided by an embodiment of the present invention. The device may include:

[0093] An obtaining module 10, configured to obtain fault area information of a current faulty hard disk in a disk array group; wherein, the current faulty hard disk is any mechanical hard disk in the disk array group, and the fault area information is information of a stripe in the disk array group where the fault occurs in the current faulty hard disk; the fault area information includes a faulty hard disk identifier and a faulty stripe identifier;

[0094] A determining module 20, configured to determine a target reconstruction area corresponding to the fault area information; wherein, the target reconstruction area is an area in the current mechanical hard disk;

[0095] A reconstruction module 30, configured to reconstruct data in the fault area of the current faulty hard disk in the target reconstruction area, and record the correspondence between the target reconstruction area and the fault area; wherein, the fault area is the stripe area corresponding to the fault area information.

[0096] In some embodiments, the target reconstruction area is the stripe area of any backup stripe in the pre-device backup area of the disk array group.

[0097] In some embodiments, the determining module 20 may be specifically configured to determine the stripe identifier of the target reconstruction area.

[0098] In some embodiments, the faulty hard disk identifier is the encoding of the current faulty hard disk, and the faulty stripe identifier is the encoding of the stripe corresponding to the fault area information; the determining module 20 may include:

[0099] A times determining sub-module, configured to determine the current fault times according to the fault area information; wherein, the current fault times are the fault times of the current faulty hard disk or the stripe fault times of the disk array group;

[0100] A times judging sub-module, configured to judge whether the current fault times reach the maximum fault times;

[0101] A reconstruction determining sub-module, configured to, if the maximum fault times are not reached, determine the stripe identifier of the target reconstruction area as the reconstruction used stripe encoding; wherein, the initial value of the reconstruction used stripe encoding is 0;

[0102] An encoding updating sub-module, configured to add 1 to the reconstruction used stripe encoding and update the reconstruction used stripe encoding;

[0103] The reconstruction module 30 can be specifically used to record the correspondence between the stripe identifier of the target reconstruction area and the encoding of the stripe corresponding to the fault area information.

[0104] In some embodiments, the reconstruction module 30 may include:

[0105] A mapping table recording sub-module, configured to record the correspondence between the target reconstruction area and the fault area in the current entry of the fault area address mapping table; wherein, the current entry is further configured to record the reconstruction status identifier of the target reconstruction area, and the reconstruction status identifier is the in-reconstruction status or the reconstruction completed status.

[0106] In some embodiments, the device may further include:

[0107] A command receiving module, configured to obtain the information of the area to be accessed corresponding to the host input / output command according to the received host input / output command; wherein, the information of the area to be accessed is the address information of each area to be accessed by the host input / output command, and the area to be accessed is a stripe area of a mechanical hard disk in a disk array group to be accessed;

[0108] A fault query module, configured to determine whether there is a recorded fault area in each area to be accessed according to the information of the area to be accessed;

[0109] An address replacement module, configured to replace the address information of the target access area in the information of the area to be accessed with the address information of the corresponding target reconstruction area if there is a recorded fault area.

[0110] In some embodiments, the information of the area to be accessed includes: the access disk array group encoding, the access stripe encoding, and the starting mechanical hard disk encoding; the access stripe encoding is io_slba / (Chunk_size*Data_disk_num), and the starting mechanical hard disk encoding is (io_slba / Chunk_size)%Data_disk_num; wherein, io_slba is the device-side starting logical block address of the host input / output command, Chunk_size is the preset area size, and Data_disk_num is the preset number of data disks;

[0111] The fault query module may include:

[0112] An encoding determination sub-module, configured to determine the starting mechanical hard disk encoding as the current access hard disk encoding;

[0113] An access judgment sub-module, which is used to judge whether the current accessed hard disk code is less than or equal to the number of accessed disks; where the number of accessed disks is [(io_slba + io_nlb) / Chunk_size] % Data_disk_num, and io_nlb is the number of device-side logical blocks of the host input / output command;

[0114] A query judgment sub-module, which is used to, if it is less than or equal to the number of accessed disks, use the current accessed hard disk code and the disk array group code as indexes to find whether there is a target entry in the entry corresponding to the current accessed hard disk code in the fault area address mapping table; where the fault stripe identifier in the target entry is the accessed stripe code;

[0115] A query judgment sub-module, which is used to, if there is a target entry, determine the area to be accessed corresponding to the current accessed hard disk code as the recorded fault area; where the target access area is the area to be accessed that is the same as the fault area.

[0116] In this embodiment, the embodiment of the present invention reconstructs the data in the fault area of the current faulty hard disk in the target reconstruction area through the reconstruction module 30 and records the corresponding relationship between the target reconstruction area and the fault area, so that the reconstruction still proceeds on the faulty mechanical hard disk without the need to insert a new disk, improving the reconstruction efficiency of the RAID group and the utilization rate of the faulty mechanical hard disk, greatly reducing the cost, and thus improving the IO performance of the host.

[0117] Corresponding to the above method embodiment, the embodiment of the present invention also provides a fault reconstruction device for a mechanical hard disk. The following described fault reconstruction device for a mechanical hard disk can be mutually corresponding and referred to with the above described fault reconstruction method for a mechanical hard disk.

[0118] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a fault reconstruction device for a mechanical hard disk provided by the embodiment of the present invention. The device may include:

[0119] A memory D1, which is used to store a computer program;

[0120] A processor D2, which is used to implement the steps of the fault reconstruction method for a mechanical hard disk provided by the above method embodiment when executing the computer program.

[0121] Among them, the fault reconstruction device for a mechanical hard disk provided by this embodiment may specifically be a RAID card; or it may specifically be a mechanical hard disk.

[0122] Corresponding to the above method embodiment, the embodiment of the present invention also provides a computer program product. The following described computer program product can be mutually corresponding and referred to with the above described fault reconstruction method for a mechanical hard disk.

[0123] A computer program product includes a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method for reconstructing a failure of a mechanical hard disk provided in the above method embodiment are implemented.

[0124] Corresponding to the above method embodiment, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium described below can be correspondingly referred to with the method for reconstructing a failure of a mechanical hard disk described above.

[0125] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the method for reconstructing a failure of a mechanical hard disk in the above method embodiment are implemented.

[0126] Specifically, the computer-readable storage medium can be various readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0127] The various embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices, equipment, computer program products, and computer-readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0128] The method, device, equipment, and computer-readable storage medium for reconstructing a failure of a mechanical hard disk provided by the present invention have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A fault reconstruction method for a mechanical hard disk, characterized in that: include: Obtaining fault area information of a currently faulty hard disk in a disk array group; wherein the currently faulty hard disk is any mechanical hard disk in the disk array group, and the fault area information is information of a stripe of the disk array group where the fault occurs in the currently faulty hard disk; the fault area information includes a faulty hard disk identifier and a faulty stripe identifier; Determine a target reconstruction area corresponding to the fault area information; wherein the target reconstruction area is an area in the current mechanical hard disk; The data of the faulty area in the current faulty hard disk is reconstructed in the target reconstruction area, and the corresponding relationship between the target reconstruction area and the faulty area is recorded; wherein the faulty area is a stripe area corresponding to the faulty area information.

2. The fault reconstruction method of a mechanical hard disk according to claim 1, characterized in that: The target reconstruction area is a stripe area of ​​any backup stripe in a preset backup area of ​​the disk array group.

3. The fault reconstruction method of a mechanical hard disk according to claim 2, characterized in that: The determining of the target reconstruction area corresponding to the fault area information includes: Determine a stripe identifier of the target reconstruction area.

4. The fault reconstruction method of a mechanical hard disk according to claim 3, characterized in that: The faulty hard disk identifier is the code of the current faulty hard disk, and the faulty stripe identifier is the code of the stripe corresponding to the faulty area information; The determining of the stripe identifier of the target reconstruction area includes: Determine the current number of failures according to the failure area information; wherein the current number of failures is the number of failures of the currently failed hard disk or the number of stripe failures of the disk array group; Determine whether the current number of faults has reached the maximum number of faults; If not, the stripe identifier of the target reconstruction area is determined as the reconstruction-used stripe code; wherein the initial value of the reconstruction-used stripe code is 0; The reconstruction-used strip code is incremented by 1, and the reconstruction-used strip code is updated; The recording of the correspondence between the target reconstruction area and the fault area includes: The corresponding relationship between the stripe identifier of the target reconstruction area and the encoding of the stripe corresponding to the fault area information is recorded.

5. The fault reconstruction method of a mechanical hard disk according to claim 1, characterized in that: The recording of the correspondence between the target reconstruction area and the fault area includes: The current entry in the fault area address mapping table records the correspondence between the target reconstruction area and the fault area; wherein the current entry is also used to record the reconstruction state identifier of the target reconstruction area, and the reconstruction state identifier is a reconstruction state or a reconstruction completed state.

6. The fault reconstruction method of a mechanical hard disk according to any one of claims 1 to 5, characterized in that: Also includes: According to the received host input / output command, the to-be-accessed area information corresponding to the host input / output command is obtained; wherein the to-be-accessed area information is the address information of each to-be-accessed area to be accessed by the host input / output command, and the to-be-accessed area is a stripe area of ​​a mechanical hard disk in a disk array group to be accessed; Determine whether there is a recorded fault area in each area to be accessed according to the information of the area to be accessed; If so, the address information of the target access area in the to-be-accessed area information is replaced with the address information of the corresponding target reconstruction area; wherein the target access area is the to-be-accessed area that is the same as the fault area.

7. The fault reconstruction method of a mechanical hard disk according to claim 6, characterized in that: The area information to be accessed includes: an access disk array group code, an access stripe code and a starting mechanical hard disk code; the access stripe code is io_slba / (Chunk_size*Data_disk_num), and the starting mechanical hard disk code is (io_slba / Chunk_size)%Data_disk_num; wherein io_slba is the device-side starting logical block address of the host input and output command, Chunk_size is the preset area size, and Data_disk_num is the preset number of data disks; The determining, according to the information of the area to be accessed, whether there is a recorded fault area in each area to be accessed includes: The starting mechanical hard disk code is determined as the currently accessed hard disk code; Determine whether the currently accessed hard disk code is less than or equal to the number of accessed disks; wherein the number of accessed disks is [(io_slba+io_nlb) / Chunk_size] %Data_disk_num, and io_nlb is the number of device-side logic blocks of the host input and output commands; If it is less than or equal to the number of accessed disks, then using the current accessed hard disk code and the disk array group code as indexes, searching whether there is a target entry in the entry corresponding to the current accessed hard disk code in the fault area address mapping table; wherein the fault stripe identifier in the target entry is the access stripe code; If the target entry exists, the to-be-accessed area corresponding to the currently-accessed hard disk code is determined as the recorded fault area.

8. A fault reconstruction device for a mechanical hard disk, characterized in that: include: An acquisition module is used to acquire the fault area information of the current faulty hard disk in the disk array group; wherein the current faulty hard disk is any mechanical hard disk in the disk array group, and the fault area information is information of a stripe of the disk array group where the fault occurs in the current faulty hard disk; the fault area information includes a faulty hard disk identifier and a faulty stripe identifier; A determination module, used to determine a target reconstruction area corresponding to the fault area information; wherein the target reconstruction area is an area in the current mechanical hard disk; The reconstruction module is used to reconstruct the data of the faulty area in the current faulty hard disk in the target reconstruction area, and record the corresponding relationship between the target reconstruction area and the faulty area; wherein the faulty area is the stripe area corresponding to the faulty area information.

9. A mechanical hard disk fault reconstruction device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the fault reconstruction method for a mechanical hard disk as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the fault reconstruction method of the mechanical hard disk as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Data processing method and device, electronic equipment and readable storage medium

    CN120872254A

  • Redundant array of independent disks (RAID) reconstruction method and electronic equipment

    CN120892264A