Raid data processing method and device, chip, electronic equipment, storage medium and computer program product

By dividing the striped area of ​​RAID50 into two area groups and generating second parity information, the problem of traditional RAID50 being unable to decode under multiple failed disks is solved, achieving efficient data recovery and improved fault tolerance.

CN120704957BActive Publication Date: 2026-01-27SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511212563.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-01-27
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Traditional distributed RAID50 cannot decode multiple failed disks, especially when two failed disks are located in the same RAID, resulting in system unavailability, low fault tolerance, and poor recovery performance.

Method used

The RAID50 stripe area is divided into two area groups, each containing at least two data areas, one parity area, and one first hot spare area. Second parity information is generated by encoding and stored in the hot spare area. Decoding and recovery are performed using multiple formulas, and data and parity information of up to four faulty areas can be recovered simultaneously.

Benefits of technology

It significantly improves the fault tolerance and recovery performance of RAID50, enabling efficient data recovery in the event of multiple failed disks, avoiding the overhead of reading global data, and improving the system's response performance and processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704957B_ABST
    Figure CN120704957B_ABST
Patent Text Reader

Abstract

The application provides a RAID data processing method and device, a chip, an electronic device, a storage medium and a computer program product, relates to the field of data processing, and is applied to RAID. The method comprises the following steps: detecting that at most four storage devices are faulty, determining at most four to-be-recovered areas corresponding to the strip in the strip, the to-be-recovered area being a data area, a check area and / or a first hot backup area; determining a target area based on the at most four to-be-recovered areas, the target area being a data area, a check area and / or a first hot backup area except the to-be-recovered area; reading target information in the target area, and recovering data information and / or first check information of the to-be-recovered area based on the target information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data processing, and in particular to a RAID data processing method and device, a chip, an electronic device, a storage medium and a computer program product. BACKGROUND

[0002] In the field of RAID technology, the traditional distributed RAID50 system includes two traditional distributed RAID5. For the traditional distributed RAID5, when a disk fails, all member disks participate in the reconstruction process in the data reconstruction process, eliminating the single disk write bottleneck. RAID5 can only tolerate one failed disk, while the traditional distributed RAID50 can tolerate at most two disk failures, and the two failed disks can only be decoded when they are distributed in two RAIDs. When the two failed disks are distributed in one RAID, they cannot be decoded, at which time the system is unavailable, the fault tolerance is low, and the recovery performance is poor. SUMMARY

[0003] The present application provides a RAID data processing method, device, chip, electronic device, storage medium and computer program product.

[0004] The present application provides a RAID data processing method, device, chip, electronic device, storage medium and computer program product.

[0005] Detecting that at most four storage devices have failed, determining at most four to-be-recovered regions in the strip corresponding to the at most four storage devices, the to-be-recovered region being a data region, a check region and / or a first hot standby region;

[0006] Determining a target region based on the at most four to-be-recovered regions, the target region being a data region, a check region and / or a first hot standby region other than the to-be-recovered region;

[0007] Reading target information in the target region, and recovering data information and / or first check information of the to-be-recovered region based on the target information.

[0008] The method further includes:

[0009] detecting that one storage device fails, and the corresponding to-be-recovered region in the stripe is a data region or a check region, determining the target region as other data regions and / or check regions in a region group to which the to-be-recovered region belongs.

[0010] The determining the target region based on the at most four to-be-recovered regions comprises:

[0011] detecting that two storage devices fail, and the corresponding two to-be-recovered regions in the stripe are data regions and / or check regions and belong to a same region group, determining the target region as all other data regions and / or check regions and the first hot backup region.

[0012] The determining the target region based on the at most four to-be-recovered regions comprises:

[0013] detecting that two storage devices fail, and the corresponding two to-be-recovered regions in the stripe are data regions and / or check regions and belong to different region groups, determining the target region as all other data regions and / or check regions.

[0014] The determining the target region based on the at most four to-be-recovered regions comprises:

[0015] detecting that two storage devices fail, and the corresponding one to-be-recovered region in the stripe is a data region or a check region, and the other to-be-recovered region is a first hot backup region, determining the target region as other data regions and / or check regions in a region group to which the data region or the check region to be recovered belongs.

[0016] The determining the target region based on the at most four to-be-recovered regions comprises:

[0017] detecting that three storage devices fail, and the corresponding three to-be-recovered regions in the stripe are data regions and / or check regions and belong to a same region group, determining the target region as all other data regions and / or check regions and the first hot backup region.

[0018] The determining the target region based on the at most four to-be-recovered regions comprises:

[0019] detecting that three storage devices fail, and the corresponding three to-be-recovered regions in the stripe are data regions and / or check regions, two of which belong to a same region group and the other belongs to another region group, determining the target region as all other data regions and / or check regions and the first hot backup region.

[0020] The determining the target region based on the at most four to-be-recovered regions comprises:

[0021] In a case where three storage devices are detected to be faulty, two corresponding to-be-restored regions in the stripe are data regions and / or check regions, another to-be-restored region is a first hot spare region, and the two to-be-restored data regions and / or check regions belong to a same region group, the target region is determined to be all other data regions and / or check regions and the first hot spare region.

[0022] The determining the target region based on the at most four to-be-restored regions comprises:

[0023] In a case where three storage devices are detected to be faulty, two corresponding to-be-restored regions in the stripe are data regions and / or check regions, another to-be-restored region is a first hot spare region, and the two to-be-restored data regions and / or check regions belong to a same region group, the target region is determined to be all other data regions and / or check regions and the first hot spare region.

[0024] The determining the target region based on the at most four to-be-restored regions comprises:

[0025] In a case where three storage devices are detected to be faulty, two corresponding to-be-restored regions in the stripe are data regions and / or check regions, another to-be-restored region is a first hot spare region, and the two to-be-restored data regions and / or check regions belong to a same region group, the target region is determined to be all other data regions and / or check regions and the first hot spare region.

[0026] The determining the target region based on the at most four to-be-restored regions comprises:

[0027] In a case where four storage devices are detected to be faulty, four corresponding to-be-restored regions in the stripe are data regions and / or check regions, three of which belong to a same region group and another belongs to another region group, the target region is determined to be all other data regions and / or check regions and the first hot spare region.

[0028] The determining the target region based on the at most four to-be-restored regions comprises:

[0029] In a case where four storage devices are detected to be faulty, two corresponding to-be-restored regions in the stripe are data regions and / or check regions, another two to-be-restored regions are first hot spare regions, and the two to-be-restored data regions and / or check regions belong to different region groups, the target region is determined to be all other data regions and / or check regions.

[0030] Another aspect of the embodiments of the present application provides a RAID data processing device, which comprises:

[0031] The processing module is configured to detect that at most four storage devices fail, determine at most four corresponding to-be-recovered areas in a stripe, the to-be-recovered areas being data areas, check areas and / or first hot backup areas;

[0032] The computing module is configured to determine a target area based on the at most four to-be-recovered areas, the target area being a data area, a check area and / or a first hot backup area other than the to-be-recovered areas.

[0033] The computing module is further configured to read target information in the target area, and recover data information and / or first check information of the to-be-recovered areas based on the target information.

[0034] Another aspect of the embodiments of the present application provides a chip, the chip comprising a processor, the processor being capable of executing the RAID data processing method.

[0035] Another aspect of the embodiments of the present application provides an electronic device, the electronic device comprising a chip, the chip comprising a processor, the processor being capable of executing the RAID data processing method.

[0036] Another aspect of the embodiments of the present application provides a computer readable storage medium, the storage medium storing a computer program, the computer program being used to execute the RAID data processing method.

[0037] Another aspect of the embodiments of the present application provides a computer program product, comprising a computer program or instructions, used to cause a processor to execute the RAID data processing method provided by the embodiments of the present application.

[0038] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0039] The above and other objects, features and advantages of the present exemplary embodiments will be readily understood through reading the following detailed description in conjunction with the accompanying drawings, in which:

[0040] In the drawings, identical or corresponding reference signs refer to identical or corresponding parts.

[0041] Figure 1 A flow chart of a RAID data processing method according to one embodiment of the present application is shown;

[0042] Figure 2Fig. 1 shows a structural schematic diagram of a RAID data processing apparatus according to an embodiment of the present application;

[0043] Figure 3 Fig. 2 shows a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to make the objectives, features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0045] The scheme of the present application is used in distributed RAID50 (a RAID level).

[0046] Each stripe in the distributed RAID50 divides all the regions into two region groups, each of which stores data information through at least two data regions and stores one copy of first check information through one check region, and the first check information is obtained by encoding based on the data information stored in the region group to which the first check information belongs. When disk errors occur, at most two disk errors can be recovered at the same time.

[0047] For example, k+2 storage devices are used to form a distributed RAID50, wherein each stripe includes two check regions (regions storing first check information), k-2 data regions (regions storing data information) and two first hot standby regions (used for temporarily storing recovered data information and / or first check information), and these regions belong to different storage devices. All the regions of the stripe are divided into two region groups, each of which includes at least two data regions, one check region and one first hot standby region. The two copies of first check information are p and q. The two copies of first check information p and q can be determined through the data information stored in the respective region groups and the following formulas (1) and (2).

[0048]

[0049] wherein, is an exclusive OR operation, is the data information in the i th data region in the stripe, is the data information in the i th data region in the stripe, is the data information in a group of region groups in the stripe, is the data information in another group of region groups in the stripe, , and is the first check information in the stripe.

[0050] When a disk error occurs, if there is only one disk error, the data information and / or the first check information in the area group to which the disk error belongs and the above formula (1) or formula (2) can be used to decode and recover the single disk error. If there are two disk errors, and the two disk errors belong to different area groups, the data information and / or the first check information in the non-disk error area, the above formula (1) and formula (2) can be used to decode and recover the data in the two disk errors. However, in the case of two disk errors belonging to the same area group, or three disk errors, four disk errors, recovery cannot be performed.

[0051] For example, as shown in Table 1, Table 1 shows a distributed RAID 50 including 10 storage devices, divided into 2 stripe groups, each stripe group including 4 stripes and written in a left-rotated misaligned manner. Each stripe has 6 data areas storing data information, 2 check areas storing first check information, and 2 first hot spare areas. The first hot spare area is used to temporarily store the recovered data in the disk error area when a disk error occurs in the RAID 50. When storage device 2 and storage device 7 in the RAID 50 fail, the data information and / or the first check information in storage devices 2 and 7 needs to be recovered by reading the data information and / or the first check information in all areas except the disk error area and the first hot spare area as shown in the shaded part of Table 1.

[0052] Table 1

[0053]

[0054] For example, as shown in Table 2, Table 2 also shows a distributed RAID 50 including 10 storage devices, divided into 2 stripe groups, each stripe group including 4 stripes and written in a left-rotated misaligned manner. Each stripe has 6 data areas storing data information, 2 check areas storing first check information, and 2 first hot spare areas. The first hot spare area is used to temporarily store the recovered data in the disk error area when a disk error occurs in the RAID 50.

[0055] When storage device 2 and storage device 3 in the RAID 50 fail, since two data information and / or first check information are lost in each stripe, and the two data information and / or first check information belong to the same area group, the RAID 50 cannot recover the disk error.

[0056] When storage device 2, storage device 3 and storage device 6 in the RAID 50 fail, since three data information and / or first check information are lost in each stripe, the RAID 50 cannot recover the disk error.

[0057] Table 2

[0058]

[0059] To overcome the performance bottleneck of data recovery, improve the speed and efficiency of data recovery, enhance the recovery performance of RAID50, and enable the simultaneous recovery of up to four faulty disks in RAID50, this application provides a RAID data processing method applied to RAID50. The RAID50 includes at least one stripe group, which includes multiple stripes. The area within each stripe is divided into two area groups. Each area group includes at least two data areas, a parity area, and a first hot spare area. The data areas store data information, the parity area stores first parity information, and the first hot spare area stores second parity information. The first parity information is determined based on all data information in its respective area group, and the second parity information is determined based on all data information and all first parity information.

[0060] For example, as shown in Table 3, based on distributed RAID50, all data information and the first parity information are encoded to obtain two second parity information messages, which are then stored in the first hot spare areas of two area groups respectively. Stripe 1 in stripe group 1 will... , , , and The region is divided into group 1 of this strip. , , , and The region is divided into two groups, and then encoded based on all data information and the first check information using the following formulas (3) and (4) to obtain the two second check information for the strip. and The two second verification messages are then stored in the first hot standby area of ​​the two area groups of this stripe. Subsequent stripes follow the same method and will not be described again here.

[0061]

[0062] in, For XOR operation, For the first in the strip Data information in each data area This refers to the location information of the storage device belonging to the corresponding region. and This is the first verification information. and a second check information.

[0063] Table 3

[0064]

[0065] As Figure 1 shown, the method comprises:

[0066] Step 101, detecting that at most four storage devices fail, determining corresponding at most four to-be-recovered areas in the stripe, the to-be-recovered areas being data areas, check areas and / or first hot backup areas.

[0067] On the basis of the conventional RAID50, after the two second check information is obtained through the above-mentioned encoding, the information in the corresponding to-be-recovered areas can be recovered when at most four storage devices fail.

[0068] Step 102, determining a target area based on the at most four to-be-recovered areas, the target area being a data area, a check area and / or a first hot backup area other than the to-be-recovered areas.

[0069] According to the number and type of the to-be-recovered areas, the corresponding target area needs to be determined, and then the target information in the target area is read to recover the information in the to-be-recovered areas. The target area can be a data area, a check area and / or a first hot backup area other than the to-be-recovered areas.

[0070] Step 103, reading the target information in the target area, and recovering the data information and / or the first check information of the to-be-recovered areas based on the target information.

[0071] When only one storage device fails at the same time, the corresponding to-be-recovered area in the stripe can be a data area, a check area or a first hot backup area.

[0072] If the to-be-recovered area is a data area or a check area, the target area can be determined as all other data areas and / or check areas in the area group where the to-be-recovered area is located, the target data in the target area is read, and the information in the to-be-recovered area is recovered through the above-mentioned formula (1) or formula (2).

[0073] If the to-be-recovered area is a first hot backup area, it is not necessary to recover, and after the failure of the storage device ends or is repaired, all data information, first check information and second check information are read, the second check information is re-determined through the above-mentioned formula (3) and formula (4) and stored in the recovered first hot backup area.

[0074] When only two storage devices fail at the same time, there are four cases:

[0075] The first kind, the two corresponding to-be-recovered regions in the stripe are data regions and / or check regions, and the two to-be-recovered regions belong to the same region group. At this time, since there are two lost data information and / or first check information in a single region group, the recovery cannot be performed through formula (1) or formula (2). Therefore, the target region is determined as all other data regions and / or check regions and the first hot spare region, the target information in the target region is read, and the data information and / or first check information in the two to-be-recovered regions are recovered through formula (3) and formula (4).

[0076] The second kind, the two corresponding to-be-recovered regions in the stripe are data regions and / or check regions, and the two to-be-recovered regions belong to different region groups. At this time, there is only one lost data information or first check information in each region group in the stripe, therefore, the target region is determined as all other data regions and / or check regions. The other data information and / or first check information in each region group is read, and the data information and / or first check information in the two to-be-recovered regions are recovered through formula (1) and formula (2) respectively.

[0077] The third kind, one corresponding to-be-recovered region in the stripe is a data region or a check region, and the other to-be-recovered region is a first hot spare region. At this time, there is only one region group losing information in the stripe, and only one data information or first check information is lost. Therefore, the target region is determined as other data regions and / or check regions in the region group to which the to-be-recovered data region or check region belongs. The other data information and / or first check information in the region group is read, and the data information or first check information in the to-be-recovered region is recovered through formula (1) or formula (2). The second check information in the first hot spare region can be read after the failure of the storage device is repaired, all data information, first check information and / or second check information are read, the second check information is re-determined through formula (3) and formula (4) and stored in the recovered first hot spare region.

[0078] The fourth kind, the two corresponding to-be-recovered regions in the stripe are first hot spare regions. At this time, since no data information is lost, the second check information does not need to be recovered immediately. After the failure of the storage device is repaired, all data information and first check information are read, the two second check information are re-determined through formula (3) and formula (4) and stored in the two recovered first hot spare regions respectively.

[0079] When only three storage devices fail at the same time, there are four cases:

[0080] The first kind, the corresponding three to be recovered areas in the strip are data areas and / or check areas, and the three to be recovered areas belong to the same area group. At this time, since there are three lost data information and / or first check information in a single area group, the recovery cannot be performed through formula (1) or formula (2). Therefore, the target area is determined as all other data areas and / or check areas and the first hot backup area. The target information in the target area is read, formula (5) is obtained through formula (1) and formula (2), and the data information and / or first check information in the two to be recovered areas are recovered through joint decoding of formula (3), formula (4) and formula (5).

[0081]

[0082] The second kind, the corresponding three to be recovered areas in the strip are data areas and / or check areas, two of the to be recovered areas belong to the same area group, and the other to be recovered area belongs to another area group. At this time, there are two lost data information and / or first check information in a region group, and the two to be recovered areas in the region group cannot be recovered through formula (1) or formula (2). Therefore, the target area is determined as all other data areas and / or check areas and the first hot backup area. The other data information and / or first check information in the region group which loses only one data information or first check information is read first, and the data information or first check information is recovered through formula (1) or formula (2). Then, the other data information and / or first check information and the second check information in the strip are read, and the data information or first check information in the remaining two to be recovered areas is recovered through joint decoding of formula (3) and formula (4).

[0083] The third kind, the corresponding two to be recovered areas in the strip are data areas and / or check areas, the other to be recovered area is the first hot backup area, and the two to be recovered data areas and / or check areas belong to different area groups. At this time, the two area groups in the strip each have only one lost data information or first check information, which can be recovered through formula (1) and formula (2) respectively. Therefore, the target area is determined as all other data areas and / or check areas. The other data information and / or first check information in each area group is read, and the data information and / or first check information in the two to be recovered areas is recovered through decoding of formula (1) and formula (2) respectively. The second check information in the first hot backup area can be read after the failure of the storage device is repaired, all data information, first check information and / or second check information are read, the second check information is re-determined through formula (3) and formula (4) above and stored in the recovered first hot backup area.

[0084] Fourthly, the corresponding two to-be-recovered regions in the stripe are the first hot backup regions, and the other to-be-recovered region is the data region or the check region. At this time, only one data information or first check information in one region group in the stripe is lost, and the data information or first check information can be recovered through the formula (1) or formula (2). Therefore, the target region is determined as the other data region and / or check region in the region group to which the to-be-recovered data region or check region belongs. The other data region and / or check region in the region group is read, and the data information or first check information is recovered through the formula (1) or formula (2). The two second check information does not need to be recovered immediately, and after the failure of the storage device is repaired, all data information and first check information are read, the two second check information is re-determined through the above formula (3) and formula (4) and stored in the two first hot backup regions after recovery respectively.

[0085] When only four storage devices fail at the same time, there are two cases:

[0086] Firstly, the corresponding four to-be-recovered regions in the stripe are data regions and / or check regions, three to-be-recovered regions belong to the same region group, and the other to-be-recovered region belongs to another region group. At this time, since four data information and / or first check information are lost, it is not possible to directly recover through the formula (3) and formula (4), but the region group in which only one data information or first check information is lost can recover the lost data information or first check information in the region group through the formula (1) or formula (2). Then, based on the recovered data information or first check information, all other data information and / or first check information which is not lost, and the two second check information, the three data information and / or first check information in the other region group are recovered through the formula (3), formula (4) and formula (5) joint decoding. Therefore, the target region is determined as all other data regions and / or check regions and the first hot backup region.

[0087] The second, the corresponding two to-be-recovered areas in the stripe are data areas and / or check areas, the other two to-be-recovered areas are both first hot backup areas, and the two to-be-recovered data areas and / or check areas belong to different area groups. At this time, the two area groups in the stripe each have only one lost data information or first check information, which can be recovered through formula (1) and formula (2) respectively. Therefore, the target area is determined as all other data areas and / or check areas. The other data information and / or first check information in each area group is read, and the data information and / or first check information in the two to-be-recovered areas is recovered through formula (1) and formula (2) respectively. The second check information in the first hot backup area can be read after the failure of the storage device is repaired, all data information, first check information and / or second check information are read, the second check information is re-determined through formula (3) and formula (4) and stored in the recovered first hot backup area.

[0088] In the above scheme, by jointly encoding all data information and first check information in the traditional distributed RAID50, two global second check information is determined and stored in the first hot backup area, thereby significantly improving the fault tolerance and recovery performance of the system. When three faulty data areas and / or check areas appear at the same time and belong to the same area group, the data information and / or first check information of the two area groups and the second check information of the first hot backup area are read, formula (5) is determined using formula (1) and formula (2), and finally the three lost data information and / or first check information is recovered by jointly encoding formula (3), formula (4) and formula (5), thereby significantly improving the fault tolerance and recovery performance of the RAID system. When four faulty data areas and / or check areas appear at the same time, the lost data information and / or first check information in some cases can also be recovered, thereby further improving the fault tolerance and recovery performance of the RAID system.

[0089] In an example of the present application, a RAID data processing method is also provided, which determines the target area based on the at most four to-be-recovered areas, comprising:

[0090] When it is detected that one storage device fails and the corresponding to-be-recovered area in the stripe is a data area or a check area, the target area is determined as the other data areas and / or check areas in the area group where the to-be-recovered area is located.

[0091] For example, as shown in Table 4, in the distributed RAID50 shown in Table 4, storage device 3 fails, and the shaded part in Table 4 is the to-be-recovered area of each stripe. The to-be-recovered area of stripe 1 in stripe group 1 is the data area where it is located. The to-be-recovered area of stripe 2 in stripe group 1 is the data area where it is located. the data area where the to-be-recovered area of the stripe 1 in the stripe group 2 is located. the data area where the to-be-recovered area of the stripe 2 in the stripe group 2 is located. the data area where the to-be-recovered area of the stripe 2 in the stripe group 2 is located.

[0092] When only one storage device fails, and the corresponding to-be-recovered area in the stripe is a data area or a check area, the target area can be determined as all other data areas and / or check areas in the area group where the to-be-recovered area is located.

[0093] In the distributed RAID 50 shown in Table 4, in the stripe 1 of the stripe group 1, the to-be-recovered area is located in the data area, and the target area of the stripe 1 in the stripe group 1 is determined as the data area where the to-be-recovered area is located and the check area where the to-be-recovered area is located. In the stripe 2 of the stripe group 1, the to-be-recovered area is located in the check area, and the target area of the stripe 2 in the stripe group 1 is determined as the data area where the to-be-recovered area is located. In the stripe 1 of the stripe group 2, the to-be-recovered area is located in the data area, and the target area of the stripe 1 in the stripe group 2 is determined as the data area where the to-be-recovered area is located and the check area where the to-be-recovered area is located. In the stripe 2 of the stripe group 2, the to-be-recovered area is located in the data area, and the target area of the stripe 2 in the stripe group 2 is determined as the data area where the to-be-recovered area is located. After reading the target information of the target area in the stripe, the data information or the first check information in the to-be-recovered area is recovered through Formula (1) or Formula (2) (based on the area group where the to-be-recovered area is located).

[0094] It should be noted that if only one storage device fails, and the corresponding to-be-recovered area in the stripe is the first hot backup area, it is not necessary to recover immediately. After the failure of the storage device is repaired, all data information, first check information and / or second check information in the stripe can be read, and the second check information is re-determined through Formula (3) and Formula (4) and stored in the recovered first hot backup area.

[0095] Table 4

[0096]

[0097] ​​​​​​​​​​​​In the above scheme, by explicitly limiting the target region to other data regions and / or check regions within the region group where the to-be-recovered regions are located, the recovery process is more targeted and efficient. Unnecessary data access across the entire stripe or region group is avoided, significantly reducing the I / O overhead and computing resources required for the recovery operation, thereby speeding up the data recovery speed under single disk failure and improving the response performance and processing efficiency of the system.

[0098] In an example of the present application, a RAID data processing method is also provided, which determines the target region based on the at most four to-be-recovered regions, comprising:

[0099] When it is detected that two storage devices have failed, and the corresponding two to-be-recovered regions in the stripe are data regions or check regions and belong to the same region group, the target region is determined to be all other data regions and / or check regions and the first hot standby region.

[0100] For example, as shown in Table 5, in the distributed RAID50 shown in Table 5, storage device 3 and storage device 4 fail at the same time, then the shaded part in Table 5 is the to-be-recovered region of each stripe. The to-be-recovered region of stripe 1 in stripe group 1 is the data region and the check region where it is located. The to-be-recovered region of stripe 2 in stripe group 1 is the data region and the check region where it is located.

[0101] When two storage devices fail at the same time, and the corresponding two to-be-recovered regions in the stripe are data regions or check regions, and the two to-be-recovered regions belong to the same region group, the target region can be determined to be all other data regions and / or check regions and the first hot standby region.

[0102] In the distributed RAID50 shown in Table 5, in stripe 1 of stripe group 1, the to-be-recovered region is the data region and the check region where it is located, then the target region of stripe 1 in stripe group 1 is determined to be , , , , the data region where it is located, the check region where it is located, and , the first hot standby region where it is located. In stripe 2 of stripe group 1, the to-be-recovered region is the data region and the check region where it is located, then the target region of stripe 2 in stripe group 1 is determined to be , 、 、 、 the data region where the data region, the check region where the check region and 、 the first hot spare region where the first hot spare region. After reading the target information of the target region in the stripe, the data information and / or the first check information in the two to-be-recovered regions are recovered by joint decoding according to formula (3) and formula (4).

[0103] Table 5

[0104]

[0105] In the above scheme, when two storage devices fail at the same time, the corresponding two to-be-recovered regions are data regions and / or check regions, and the two to-be-recovered regions belong to the same region group, the target region is explicitly specified as all other data regions and / or check regions and the first hot spare region, and the second check information is used for joint decoding, effectively solving the problem that the double disk failure in the same region group cannot be recovered in the traditional method. Under the premise of maintaining data integrity, the precise reconstruction of the data information and the first check information in the failed region is realized, and the fault tolerance and recovery performance of the system to double disk failure are significantly improved.

[0106] In an example of the present application, a RAID data processing method is also provided, which determines the target region based on the at most four to-be-recovered regions, comprising:

[0107] When it is detected that two storage devices fail, the corresponding two to-be-recovered regions in the stripe are data regions or check regions and belong to different region groups, the target region is determined as all other data regions and / or check regions.

[0108] For example, as shown in Table 6, in the distributed RAID50 shown in Table 6, storage device 3 and storage device 7 fail at the same time, then the shaded part in Table 6 is the to-be-recovered region of each stripe. The to-be-recovered region of stripe 1 in stripe group 1 is 、 the data region where the data region. The to-be-recovered region of stripe 2 in stripe group 1 is the data region where the data region and the check region where the check region. The to-be-recovered region of stripe 1 in stripe group 2 is the data region where the data region and the check region where the check region. The to-be-recovered region of stripe 2 in stripe group 2 is 、 the data region where the data region.

[0109] When two storage devices fail at the same time, and the corresponding two to-be-recovered areas in the stripe are data areas or check areas, and the two to-be-recovered areas belong to different area groups, it can be determined that the target area is all other data areas and / or check areas.

[0110] In the distributed RAID 50 shown in Table 6, in the stripe 1 of the stripe group 1, the to-be-recovered area is the data area where , is located, it is determined that the target area of the stripe 1 in the stripe group 1 is the data area where , , , is located and the check area where , is located, based on , and , is recovered through formula (1), based on , and , is recovered through formula (2). In the stripe 2 of the stripe group 1, the to-be-recovered area is the data area where is located and the check area where is located, it is determined that the target area of the stripe 2 in the stripe group 1 is the data area where , , , , is located and the check area where is located, based on , , , is recovered through formula (1), based on , , , is recovered. In the stripe 1 of the stripe group 2, the to-be-recovered area is the data area where is located and the check area where is located, it is determined that the target area of the stripe 1 in the stripe group 2 is the data area where , , , , is located and the check area where is located, based on , , , is recovered through formula (1), based on , , recovery In the stripe 2 of the stripe group 2, the to-be-recovered area is the data area where , is located, and the target area of the stripe 2 in the stripe group 2 is determined as the data area where , , , is located and the check area where , is located, the , , is recovered through formula (1) , and the , , is recovered based on .

[0111] Table 6

[0112]

[0113] In the above scheme, when two storage devices fail at the same time, the corresponding two to-be-recovered areas are data areas and / or check areas, and the two to-be-recovered areas belong to different area groups, the target area is explicitly specified as all other data areas and / or check areas, and the lost data information and / or first check information in each group is independently and in parallel recovered by using the original check relationship in each area group. Avoid the overhead of reading global data or first hot standby area, significantly reduce the I / O load and computational complexity in the recovery process, so as to realize fast and accurate fault recovery, and effectively improve the efficiency and overall performance of the system in processing distributed faults across groups.

[0114] In an example of the present application, a RAID data processing method is also provided, and the target area is determined based on the at most four to-be-recovered areas, comprising:

[0115] When it is detected that two storage devices fail, and the corresponding one to-be-recovered area in the stripe is a data area or a check area, and the other to-be-recovered area is a first hot standby area, the target area is determined as other data areas and / or check areas in the area group to which the to-be-recovered data area or check area belongs.

[0116] For example, as shown in Table 5, in the distributed RAID50 shown in Table 5, storage device 3 and storage device 4 fail at the same time, and the to-be-recovered area of each stripe is the shaded part in Table 5. The to-be-recovered area of the stripe 1 in the stripe group 2 is the data area where is located and the first hot standby area where is located. The to-be-recovered area of the stripe 2 in the stripe group 2 is the data region and the first hot spare region.

[0117] When two storage devices fail simultaneously, and the corresponding one of the to-be-recovered regions in the stripe is a data region or a check region, and the other to-be-recovered region is a first hot spare region, the target region can be determined as the other data region and / or check region in the region group to which the to-be-recovered data region or check region belongs.

[0118] In the distributed RAID 50 shown in Table 5, in the stripe 1 of the stripe group 2, the to-be-recovered region is the data region and the first hot spare region, the target region of the stripe 1 in the stripe group 2 is determined as , the data region and the check region, based on , , recovered through formula (1) In the stripe 2 of the stripe group 2, the to-be-recovered region is the data region and the first hot spare region, the target region of the stripe 2 in the stripe group 2 is determined as , the data region and the check region, based on , , recovered through formula (1) It should be noted that if the to-be-recovered data region or check region is in another region group, only all other data information and / or first check information in the other region group are read, and the second check information is recovered through formula (2). The first hot spare region does not need to be recovered immediately. After the failure of the storage device is repaired, all data information, first check information and / or second check information in the stripe are read, the second check information is re-determined through formula (3) and formula (4), and stored in the recovered first hot spare region.

[0119] In the above solution, when two storage devices fail simultaneously, with one recovery area being a data area or verification area and the other a first hot standby area, the target area is explicitly designated as other data areas and / or verification areas within the same area group as the data area or verification area to be recovered. The lost data or verification information is quickly recovered using the verification relationships within the area group, eliminating the need to immediately address the failure of the first hot standby area. This significantly reduces the scope and computational burden of immediate recovery operations, prioritizing the accessibility of user data and the continuity of core system functions. Simultaneously, the recovery of the first hot standby area is delayed until after the failed storage device is repaired, thereby optimizing resource allocation and improving the recovery speed and efficiency of the system under mixed failure conditions.

[0120] This application also provides a RAID data processing method in one example, the method further comprising:

[0121] If two storage devices are detected to have failed, and the two corresponding recovery areas in the stripe are both first hot standby areas, then after the storage device failure is resolved, all data information and first verification information in the stripe are read, and the second verification information in the first hot standby area is re-determined based on all data information and the first verification information.

[0122] When two storage devices fail simultaneously, and the two corresponding areas to be recovered in the stripe are both first hot standby areas, immediate recovery is not required. After the storage devices are repaired, all data information and first verification information in the stripe can be read, and the two second verification information can be re-determined using formulas (3) and (4) and stored in the two recovered first hot standby areas respectively.

[0123] This application also provides a RAID data processing method in one example, wherein determining the target area based on the at most four areas to be recovered includes:

[0124] If three storage devices are detected to have malfunctioned, and the three areas to be recovered in the stripe are all data areas and / or verification areas and belong to the same area group, the target area is determined to be all other data areas and / or verification areas and the first hot standby area.

[0125] For example, as shown in Table 7, in the distributed RAID 50 array shown in Table 7, if storage devices 2, 3, and 4 fail simultaneously, the shaded areas in Table 7 represent the recoverable regions of each stripe. The recoverable region of stripe 1 in stripe group 1 is... , The data area and The verification area it is located in. The area to be recovered for stripe 2 in stripe group 1 is... , The data area and the data region in which the target region is located.

[0126] When three storage devices fail at the same time, and the corresponding three to-be-recovered regions in the stripe are all data regions and / or check regions, and the three to-be-recovered regions belong to the same region group, it can be determined that the target region is all other data regions and / or check regions, and the first hot backup region.

[0127] In the distributed RAID 50 shown in Table 7, in the stripe 1 of the stripe group 1, the to-be-recovered regions are the data region in which , the target region is located and the check region in which the target region is located, it is determined that the target region of the stripe 1 in the stripe group 1 is the data region in which , , , the target region is located, the check region in which the target region is located, and the first hot backup region in which , the target region is located. In the stripe 2 of the stripe group 1, the to-be-recovered regions are the data region in which , the target region is located and the check region in which the target region is located, it is determined that the target region of the stripe 2 in the stripe group 1 is the data region in which , , , the target region is located, the check region in which the target region is located, and the first hot backup region in which , the target region is located. After reading the target information of the target region in the stripe, the data information and / or the first check information in the three to-be-recovered regions are recovered by joint decoding according to formula (3), formula (4) and formula (5).

[0128] Table 7

[0129]

[0130] In the above solution, when three storage devices fail simultaneously, and the three corresponding areas to be recovered are all data areas and / or parity areas, the target area is explicitly designated as all other data areas and / or parity areas, as well as the first hot spare area. By comprehensively utilizing the first and second parity information within the stripe for joint decoding, the three lost data information and / or first parity information within the same area group can be successfully recovered. This overcomes the limitation of traditional RAID50 architectures in handling more than two concurrent failures, significantly enhancing the system's tolerance to severe local failures. It ensures data recoverability and system continuity under extreme failure conditions, thereby significantly improving the system's fault tolerance and recovery performance, and ultimately enhancing the overall reliability and resilience of the storage solution.

[0131] This application also provides a RAID data processing method in one example, wherein determining the target area based on the at most four areas to be recovered includes:

[0132] If three storage devices are detected to have failed, and the three areas to be recovered in the stripe are all data areas and / or verification areas, with two belonging to the same area group and the other to another area group, the target area is determined to be all other data areas and / or verification areas as well as the first hot standby area.

[0133] For example, as shown in Table 8, in the distributed RAID 50 array shown in Table 8, if storage devices 4, 6, and 7 fail simultaneously, the shaded areas in Table 8 represent the recoverable regions of each stripe. The recoverable region of stripe 1 in stripe group 1 is... , The data area and The verification area it is located in. The area to be recovered for stripe 2 in stripe group 1 is... , , The data area it is located in.

[0134] When three storage devices fail simultaneously, and the three recovery areas in the stripe are all data areas and / or parity areas, and two of the three recovery areas belong to the same area group while the other recovery area group belongs to another area group, then the target area can be determined as all other data areas and / or parity areas, as well as the first hot standby area.

[0135] In the distributed RAID50 shown in Table 8, in stripe 1 of stripe group 1, the area to be recovered is... , The data area and If the region is within the verification area, then the target region of stripe 1 in stripe group 1 is determined as... , , , The data area where it is located The verification area and , The first hot standby zone is based on , , Recover using formula (1) Based on , , , , , , , By jointly decoding using formulas (3) and (4), the results can be recovered. and In strip 2 of strip group 1, the region to be recovered is... , , If the data region is located within a certain range, then the target region for stripe 2 in stripe group 1 is determined as follows. , , The data area where it is located and The verification area and , The first hot standby zone is based on , , Recover using formula (1) Based on , , , , , , , By jointly decoding using formulas (3) and (4), the results can be recovered. and It should be noted that if only one data information or verification information area group is lost, then it is sufficient to first recover it based on formula (2).

[0136] Table 8

[0137]

[0138] In the above scheme, when three storage devices fail simultaneously, the three corresponding recovery areas are all data areas and / or parity areas. Two of these recovery areas belong to the same area group, while the other belongs to a different area group. The target area is explicitly designated as all other data areas and / or parity areas, as well as the first hot spare area. The system first recovers the information in the area group that lost a single piece of data or first parity information. Then, it uses the second parity information for joint decoding, successfully recovering the two lost pieces of data and / or first parity information in the other area group. This overcomes the limitation of traditional RAID50 architecture in handling more than two concurrent failures, significantly enhancing the system's tolerance to severe local failures. It ensures data recoverability and system continuity under extreme failure conditions, thereby significantly improving the system's fault tolerance and recovery performance, and ultimately enhancing the overall reliability and resilience of the storage solution.

[0139] This application also provides a RAID data processing method in one example, wherein determining the target area based on the at most four areas to be recovered includes:

[0140] If three storage devices are detected to have failed, and two of the corresponding areas to be recovered in the stripe are data areas and / or verification areas, and another area to be recovered is a first hot standby area, and the two data areas and / or verification areas to be recovered belong to the same area group, then the target area is determined to be all other data areas and / or verification areas and the first hot standby area.

[0141] For example, as shown in Table 9, in the distributed RAID 50 array shown in Table 9, if storage devices 4, 6, and 10 fail simultaneously, the shaded areas in Table 9 represent the recoverable regions of each stripe. The recoverable region of stripe 1 in stripe group 2 is... , The data area and The first hot standby area. The area to be restored for strip 2 in strip group 2 is... The data area where it is located The verification area and It is located in the first hot standby zone.

[0142] When three storage devices fail simultaneously, and two of the corresponding recovery areas in the stripe are data areas and / or verification areas, and the other recovery area is the first hot standby area, and the two data areas and / or verification areas to be recovered belong to the same area group, then the target area can be determined as all other data areas and / or verification areas, as well as the first hot standby area.

[0143] In the distributed RAID50 shown in Table 9, in stripe 1 of stripe group 2, the area to be recovered is... , The data area and If the location is in the first hot standby region, then the target region of stripe 1 in stripe group 2 is determined as... , , , The data area where it is located and The verification area and The first hot standby zone, based on , , , , , , , By jointly decoding using formulas (3), (4), and (5), the results are recovered. , and In strip 2 of strip group 2, the region to be recovered is... The data area where it is located The verification area and If the first hot standby region is located there, then the target region of stripe 2 in stripe group 2 is determined as follows. , , , , The data area where it is located The verification area and The first hot standby zone, based on , , , , , , By jointly decoding using formulas (3), (4), and (5), the results are recovered. , and The first hot standby area does not need to be restored immediately. After the storage device is repaired, all data information, first verification information and / or second verification information in the stripe can be read, and the second verification information can be re-determined and stored in the restored first hot standby area using formulas (3) and (4).

[0144] Table 9

[0145]

[0146] In the above scheme, when three storage devices fail simultaneously, the two corresponding areas to be recovered are the data area and / or parity area, and the other area to be recovered is the first hot spare area. Furthermore, the two data areas and / or parity areas to be recovered belong to the same area group. Therefore, the target area is explicitly designated as all other data areas and / or parity areas, as well as the first hot spare area. By comprehensively utilizing the first and second parity information within the stripe for joint decoding, the two lost data and / or parity information within the same area group can be successfully recovered. This overcomes the limitation of traditional RAID50 architecture in handling more than two concurrent failures, significantly enhancing the system's tolerance to severe local failures. It ensures data recoverability and system continuity under extreme failure conditions, thereby significantly improving the system's fault tolerance and recovery performance, and ultimately enhancing the overall reliability and resilience of the storage solution.

[0147] This application also provides a RAID data processing method in one example, wherein determining the target area based on the at most four areas to be recovered includes:

[0148] If three storage devices are detected to have failed, and the two areas to be recovered in the stripe are data areas and / or verification areas, and the other area to be recovered is the first hot standby area, and the two data areas and / or verification areas to be recovered belong to different area groups, then the target area is determined to be all other data areas and / or verification areas.

[0149] For example, as shown in Table 9, in the distributed RAID 50 array shown in Table 9, if storage devices 4, 6, and 10 fail simultaneously, the shaded areas in Table 9 represent the recoverable regions of each stripe. The recoverable region of stripe 1 in stripe group 1 is... The data area where it is located The verification area and The first hot standby area. The area to be restored for strip 2 in strip group 1 is... , The data area and It is located in the first hot standby zone.

[0150] When three storage devices fail simultaneously, and two of the corresponding recovery areas in the stripe are data areas and / or verification areas, and the other recovery area is the first hot standby area, and the two data areas and / or verification areas to be recovered belong to different area groups, then the target area can be determined to be all other data areas and / or verification areas.

[0151] In the distributed RAID50 shown in Table 9, in stripe 1 of stripe group 1, the area to be recovered is... The data area where it is located The verification area and If it is located in the first hot standby region, then the target region of stripe 1 in stripe group 1 is determined as follows. , , , , The data area where it is located The verification area and The first hot standby zone, based on , , Recover using formula (1) ,based on , , Recover using formula (2) In strip 2 of strip group 1, the region to be recovered is... , The data area and If the location is in the first hot standby region, then the target region of stripe 2 in stripe group 1 is determined as... , , , The data area where it is located and The verification area and The first hot standby zone, based on , , Recover using formula (1) ,based on , , Recover using formula (2) The first hot standby area does not need to be restored immediately. After the storage device is repaired, all data information, first verification information and / or second verification information in the stripe can be read, and the second verification information can be re-determined and stored in the restored first hot standby area using formulas (3) and (4).

[0152] In the above scheme, when three storage devices fail simultaneously, the two corresponding areas to be recovered are the data area and / or parity area, and the other area to be recovered is the first hot spare area. Furthermore, since the two data areas and / or parity areas to be recovered belong to different area groups, the target area is explicitly designated as all other data areas and / or parity areas. Only the first parity information within the stripe is needed to recover the two lost data and / or parity information within each area group. This overcomes the limitation of traditional RAID50 architectures in handling more than two concurrent failures, significantly enhancing the system's tolerance to severe local failures. It ensures data recoverability and system continuity under extreme failure conditions, thereby significantly improving the system's fault tolerance and recovery performance, and ultimately enhancing the overall reliability and resilience of the storage solution.

[0153] This application also provides a RAID data processing method in one example, wherein determining the target area based on the at most four areas to be recovered includes:

[0154] Three storage devices were detected to have failed, and the two areas to be recovered in the stripe were first hot standby areas, and the other area to be recovered was a data area or a verification area. The target area was determined to be all other data areas and / or verification areas in the area group to which the data area or verification area to be recovered belonged.

[0155] For example, as shown in Table 10, in the distributed RAID 50 shown in Table 10, if storage devices 5, 6, and 10 fail simultaneously, then the shaded areas in Table 10 represent the recoverable areas of each stripe. The recoverable area of ​​stripe 1 in stripe group 1 is... The data area and , The first hot standby area. The area to be restored for strip 2 in strip group 1 is... The data area and , It is located in the first hot standby zone.

[0156] When three storage devices fail simultaneously, and the two corresponding areas to be recovered in the stripe are the first hot standby areas, and the other area to be recovered is a data area or a verification area, then the target area can be determined to be all other data areas and / or verification areas in the area group to which the data area or verification area to be recovered belongs.

[0157] In the distributed RAID50 shown in Table 10, in stripe 1 of stripe group 1, the area to be recovered is... The data area and , If it is located in the first hot standby region, then the target region of stripe 1 in stripe group 1 is determined as follows. , The data area and The verification area is based on , , Recover using formula (2) In strip 2 of strip group 1, the region to be recovered is... The data area and , If the location is in the first hot standby region, then the target region of stripe 2 in stripe group 1 is determined as... , The data area and The verification area is based on , , Recover using formula (2) The first hot standby area does not need to be restored immediately. After the storage device is repaired, all data information, first verification information and / or second verification information in the stripe can be read, and the second verification information can be re-determined and stored in the restored first hot standby area using formulas (3) and (4).

[0158] Table 10

[0159]

[0160] In the above scheme, when three storage devices fail simultaneously, the two corresponding areas to be recovered are designated as the first hot standby areas, and the other area to be recovered is either a data area or a parity area. The target area is explicitly specified as all other data areas and / or parity areas within the same area group as the data area or parity area to be recovered. The lost data area or parity area can be recovered solely through the first parity information within the stripe. This overcomes the limitation of traditional RAID50 architectures in handling more than two concurrent failures, significantly enhancing the system's tolerance to severe local failures. It ensures data recoverability and system continuity under extreme failure conditions, thereby significantly improving the system's fault tolerance and recovery performance, and ultimately enhancing the overall reliability and resilience of the storage solution.

[0161] This application also provides a RAID data processing method in one example, wherein determining the target area based on the at most four areas to be recovered includes:

[0162] Four storage devices were detected to have malfunctioned; the four regions to be recovered in the stripe were all data regions and / or verification regions; three of them belonged to the same region group and the other to another region group; the target region was determined to be all other data regions and / or verification regions and the first hot standby region.

[0163] For example, as shown in Table 11, in the distributed RAID 50 array shown in Table 11, if storage devices 1, 2, 3, and 9 fail simultaneously, then the shaded areas in Table 11 represent the recoverable regions of each stripe. The recoverable region of stripe 1 in stripe group 1 is... , , The data area and The verification area it is located in. The area to be recovered for stripe 2 in stripe group 1 is... , , The data area and The verification area it is located in.

[0164] When four storage devices fail simultaneously, and the four areas to be recovered in the stripe are all data areas and / or parity areas, and three of the areas to be recovered belong to the same area group, while the other area to be recovered belongs to another area group, then the target area can be determined as all other data areas and / or parity areas and the first hot standby area.

[0165] In the distributed RAID50 shown in Table 11, in stripe 1 of stripe group 1, the area to be recovered is... , , The data area and If the region is within the verification area, then the target region of stripe 1 in stripe group 1 is determined as... , , The data area where it is located The verification area and and The first hot standby zone is based on , , Recover using formula (2) Based on , , , , , , By jointly decoding using formulas (3), (4), and (5), the results are recovered. , , In strip 2 of strip group 1, the region to be recovered is... , , The data area and If the area is within the verification region, then the target region of stripe 2 in stripe group 1 is determined as... , , The data area where it is located The verification area and and The first hot standby zone is based on , , Recover using formula (2) Based on , , , , , , By jointly decoding using formulas (3), (4), and (5), the results are recovered. , and .

[0166] Table 11

[0167]

[0168] In the above scheme, when four storage devices fail simultaneously, the four areas to be recovered are all data areas and / or parity areas. Three of these areas belong to the same area group, while the remaining area belongs to another group. The target area is explicitly designated as all other data areas and / or parity areas, plus the first hot spare area. First, all other data and / or first parity information within the area group that has lost only one piece of data or first parity information is used to recover that data or first parity information. Then, the first and second parity information within the stripe are used for joint decoding, successfully recovering the three lost data and / or parity information within the other area group. This overcomes the limitation of traditional RAID50 architectures in handling more than two concurrent failures, significantly enhancing the system's tolerance to severe local failures. It ensures data recoverability and system continuity under extreme failure conditions, thereby significantly improving the system's fault tolerance and recovery performance, and ultimately enhancing the overall reliability and resilience of the storage solution.

[0169] This application also provides a RAID data processing method in one example, wherein determining the target area based on the at most four areas to be recovered includes:

[0170] If four storage devices are detected to have failed, and the two areas to be recovered in the stripe are data areas and / or verification areas, and the other two areas to be recovered are first hot standby areas and the two data areas and / or verification areas to be recovered belong to different area groups, then the target area is determined to be all other data areas and / or verification areas.

[0171] For example, as shown in Table 12, in the distributed RAID 50 array shown in Table 12, if storage devices 1, 4, 6, and 9 fail simultaneously, then the shaded areas in Table 12 represent the recoverable regions of each stripe. The recoverable region of stripe 1 in stripe group 2 is... , The data area and , The first hot standby area. The area to be restored for strip 2 in strip group 2 is... , The verification area and , It is located in the first hot standby zone.

[0172] When four storage devices fail simultaneously, and the two areas to be recovered in the stripe are data areas and / or parity areas, and the other two areas to be recovered are the first hot standby areas, and the two data areas and / or parity areas to be recovered belong to different area groups, then the target area can be determined to be all other data areas and / or parity areas.

[0173] In the distributed RAID50 shown in Table 12, in stripe 1 of stripe group 2, the area to be recovered is... , The data area and , If the location is in the first hot standby region, then the target region of stripe 1 in stripe group 2 is determined as... , , , The data area and and The verification area is located, first based on , , Recover using formula (1) Based on , , Recover using formula (2) In strip 2 of strip group 2, the region to be recovered is... , The verification area and , If the first hot standby region is located there, then the target region of stripe 2 in stripe group 2 is determined as follows. , , , , , The data area in question is based on... , , Recover using formula (1) Based on , , Recover using formula (2) .

[0174] Table 12

[0175]

[0176] In the above scheme, when four storage devices fail simultaneously, the two corresponding areas to be recovered are the data area and / or parity area, and the other two areas to be recovered are the first hot standby areas. Since the two data areas and / or parity areas to be recovered belong to different area groups, the target area is explicitly designated as all other data areas and / or parity areas. The lost data or first parity information within each of the two area groups that have only lost one piece of data or first parity information is recovered using all other data and / or first parity information. This overcomes the limitation of traditional RAID50 architecture in handling more than two concurrent failures, significantly enhancing the system's tolerance to severe local failures, ensuring data recoverability and system continuity under extreme failure conditions, thereby significantly improving the system's fault tolerance and recovery performance, and ultimately enhancing the overall reliability and resilience of the storage solution.

[0177] This application also provides a RAID data processing method in one example, wherein the stripe further includes a second hot spare area, and the method further includes:

[0178] The recovered data and / or the first verification information are stored in the first hot standby area and / or the second hot standby area.

[0179] If a storage device fails and the corresponding data information to be recovered and / or the first verification information in the stripe is recovered, if there is a hot standby area in the stripe that has not failed, the recovered data information and / or the first verification information can be temporarily stored in the hot standby area.

[0180] When there is only one faulty data area or verification area, the data information or first verification information can be stored in the first hot standby area after recovery.

[0181] When there are two faulty data areas and / or verification areas at the same time, after the data information and / or the first verification information are recovered, they can be stored in the two first hot standby areas respectively.

[0182] When there are three faulty data areas and / or verification areas at the same time, a second hot standby area needs to be set in the stripe. One second hot standby area is also set in each area group. After the data information and / or the first verification information are restored, if neither the first hot standby area nor the second hot standby area has a fault, they can be stored in one first hot standby area and two second hot standby areas respectively.

[0183] When there are three faulty data areas and / or verification areas at the same time, a second hot standby area needs to be set in the stripe. One second hot standby area is also set in each area group. After the data information and / or the first verification information are restored, if neither the first hot standby area nor the second hot standby area has a fault, they can be stored in two first hot standby areas and two second hot standby areas respectively.

[0184] It should be noted that when there is only one faulty data area or verification area, or two faulty data areas and / or verification areas at the same time, if a second hot standby area is set in the stripe, the recovered data information and / or the first verification information will be preferentially stored in the second hot standby area where the data is stored.

[0185] In the above solution, after recovering the data and / or the first verification information, it is stored in the hot backup area of ​​the corresponding stripe. This eliminates the need to wait for replacement or repair of faulty disks, significantly improving the timeliness of data recovery.

[0186] This application also provides a RAID data processing method in one example, the method further comprising:

[0187] Once the fault in the storage device is detected and resolved, all data information and / or the first verification information are read.

[0188] The second verification information in the first hot standby area is determined based on the data information and / or the first verification information.

[0189] After the faulty storage device is repaired or replaced, it is necessary to restore the faulty or overwritten second verification information.

[0190] Read all data information and / or the first verification information, and perform joint decoding using the above formulas (3) and (4) to redetermine the two second verification information.

[0191] After restoring the second verification information, if the corresponding first hot standby area is not covered by the data information and the first verification information, it can be stored in the corresponding first hot standby area.

[0192] If the corresponding first hot standby area is already covered by data information and first verification information, the data information and first verification information in the first hot standby area can be transferred back to the recovered data area or verification area first, and then the recovered second verification information can be stored in the first hot standby area. Alternatively, the first hot standby area can be modified into a data area or verification area based on the stored data information or first verification information, and then the recovered second verification information can be stored in the recovered data area or verification area. Finally, based on the stored second verification information, the recovered data area or verification area can be modified into the first hot standby area.

[0193] In the above scheme, after the faulty storage device is repaired, the recovery data temporarily stored in the first hot standby area can be migrated back to its original, recovered area. Then, the newly generated second verification information is written back to the corresponding first hot standby area. This improves the flexibility of data recovery while maintaining the overall structural integrity of the system after data recovery.

[0194] To implement the above RAID data processing method, such as Figure 2 As shown, an example of this application provides a RAID data processing apparatus, including:

[0195] Processing module 201 is used to detect up to four storage devices that have failed and to determine up to four regions to be recovered in the stripe, wherein the regions to be recovered are a data region, a verification region and / or a first hot standby region;

[0196] Calculation module 202 is used to determine a target area based on the up to four areas to be recovered, wherein the target area is a data area, a verification area and / or a first hot standby area other than the areas to be recovered;

[0197] The calculation module 202 is also used to read target information in the target area and recover data information and / or first verification information of the area to be recovered based on the target information.

[0198] The calculation module 202 is further configured to detect a storage device failure and determine the target area as other data areas and / or verification areas in the area group where the area to be recovered is located, provided that the corresponding area to be recovered in the strip is a data area or a verification area.

[0199] The calculation module 202 is further configured to detect that two storage devices have failed, and that the two regions to be recovered in the stripe are both data regions and / or verification regions and belong to the same region group, and to determine the target region as all other data regions and / or verification regions and the first hot standby region.

[0200] The calculation module 202 is further configured to detect that two storage devices have failed, and that the two regions to be recovered in the stripe are both data regions and / or verification regions and belong to different region groups, and to determine the target region as all other data regions and / or verification regions.

[0201] The calculation module 202 is further configured to detect that two storage devices have failed, and that one of the areas to be recovered in the stripe is a data area or a verification area, and the other area to be recovered is a first hot standby area, and to determine that the target area is another data area and / or verification area in the area group to which the data area or verification area to be recovered belongs.

[0202] The calculation module 202 is further configured to detect that three storage devices have failed, and that the three regions to be recovered in the stripe are all data regions and / or verification regions and belong to the same region group, and to determine the target region as all other data regions and / or verification regions and the first hot standby region.

[0203] The calculation module 202 is further configured to detect that three storage devices have failed, that the three regions to be recovered in the stripe are all data regions and / or verification regions, that two of them belong to the same region group and the other belongs to another region group, and to determine that the target region is all other data regions and / or verification regions and the first hot standby region.

[0204] The calculation module 202 is further configured to detect that three storage devices have failed, that two regions to be recovered in the stripe are data regions and / or verification regions, that another region to be recovered is a first hot standby region, and that the two data regions and / or verification regions to be recovered belong to the same region group, and to determine that the target region is all other data regions and / or verification regions and the first hot standby region.

[0205] The calculation module 202 is further configured to detect that three storage devices have failed, that two regions to be recovered in the stripe are data regions and / or verification regions, that another region to be recovered is a first hot standby region, and that the two data regions and / or verification regions to be recovered belong to different region groups, and to determine that the target region is all other data regions and / or verification regions.

[0206] The calculation module 202 is further configured to detect that three storage devices have failed, and that two of the corresponding areas to be recovered in the stripe are first hot standby areas, and the other area to be recovered is a data area or a verification area, and to determine that the target area is all other data areas and / or verification areas in the area group to which the data area or verification area to be recovered belongs.

[0207] The calculation module 202 is further configured to detect that four storage devices have failed, that the four regions to be recovered in the stripe are all data regions and / or verification regions, that three of them belong to the same region group and the other belongs to another region group, and to determine that the target region is all other data regions and / or verification regions and the first hot standby region.

[0208] The calculation module 202 is further configured to detect that four storage devices have failed, that two regions to be recovered in the stripe are data regions and / or verification regions, that the other two regions to be recovered are first hot standby regions and that the two data regions and / or verification regions to be recovered belong to different region groups, and to determine that the target region is all other data regions and / or verification regions.

[0209] This application also provides a chip, which includes a processor capable of executing the RAID data processing method provided in this application.

[0210] This application also provides an electronic device.

[0211] Figure 3 A schematic block diagram of an example electronic device 300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0212] like Figure 3 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of the device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0213] Multiple components in device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of monitors, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0214] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as RAID data processing methods. For example, in some embodiments, the RAID data processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the RAID data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform RAID data processing methods by any other suitable means (e.g., by means of firmware).

[0215] This application provides a computer-readable storage medium storing executable instructions, wherein a computer program is stored, the computer program being used to execute the RAID data processing method provided in this application.

[0216] This application provides a computer program product, which includes a computer program or instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program or instructions from the computer-readable storage medium and executes the computer program or instructions, causing the computer device to perform the RAID data processing method described above in this application.

[0217] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0218] In some embodiments, a computer program may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0219] As an example, a computer program may be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0220] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0221] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0222] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0223] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0224] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0225] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0226] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0227] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.

[0228] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A RAID data processing method, characterized in that, The method is applied to RAID 50, which includes at least one stripe group, the stripe group including multiple stripes, and the area in the stripe is divided into two area groups. Each area group includes at least two data areas, one parity area, and one first hot spare area. The data areas store data information, the parity area stores first parity information, and the first hot spare area stores second parity information. The first parity information is determined based on all data information in its area group, and the second parity information is determined based on all data information in its stripe and all first parity information in its stripe. The method includes: If up to four storage devices fail, determine up to four regions to be recovered in the stripe, wherein the regions to be recovered are a data region, a verification region, and / or a first hot standby region. Based on the different number and types of the areas to be recovered, a corresponding target area is determined. The target area is the data area, verification area and / or first hot standby area other than the areas to be recovered. Read the target information in the target area, and recover the data information and / or first verification information of the area to be recovered based on the target information.

2. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If a storage device malfunction is detected, and the corresponding area to be recovered in the stripe is a data area or a verification area, the target area is determined to be another data area and / or verification area in the area group where the area to be recovered is located.

3. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If two storage devices are detected to have failed, and the two areas to be recovered in the stripe are both data areas and / or verification areas and belong to the same area group, the target area is determined to be all other data areas and / or verification areas and the first hot standby area.

4. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If two storage devices are detected to be faulty, and the two areas to be recovered in the stripe are both data areas and / or verification areas and belong to different area groups, the target area is determined to be all other data areas and / or verification areas.

5. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If two storage devices are detected to have failed, and one of the areas to be recovered in the stripe is a data area or a verification area, and the other area to be recovered is a first hot standby area, then the target area is determined to be another data area and / or verification area in the area group to which the data area or verification area to be recovered belongs.

6. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If three storage devices are detected to have malfunctioned, and the three areas to be recovered in the stripe are all data areas and / or verification areas and belong to the same area group, the target area is determined to be all other data areas and / or verification areas and the first hot standby area.

7. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If three storage devices are detected to have failed, and the three areas to be recovered in the stripe are all data areas and / or verification areas, with two belonging to the same area group and the other to another area group, the target area is determined to be all other data areas and / or verification areas as well as the first hot standby area.

8. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If three storage devices are detected to have failed, and two of the corresponding areas to be recovered in the stripe are data areas and / or verification areas, and another area to be recovered is a first hot standby area, and the two data areas and / or verification areas to be recovered belong to the same area group, then the target area is determined to be all other data areas and / or verification areas and the first hot standby area.

9. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If three storage devices are detected to have failed, and the two areas to be recovered in the stripe are data areas and / or verification areas, and the other area to be recovered is the first hot standby area, and the two data areas and / or verification areas to be recovered belong to different area groups, then the target area is determined to be all other data areas and / or verification areas.

10. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: Three storage devices were detected to have failed, and the two areas to be recovered in the stripe were first hot standby areas, and the other area to be recovered was a data area or a verification area. The target area was determined to be all other data areas and / or verification areas in the area group to which the data area or verification area to be recovered belonged.

11. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: Four storage devices were detected to have malfunctioned; the four regions to be recovered in the stripe were all data regions and / or verification regions; three of them belonged to the same region group and the other to another region group; the target region was determined to be all other data regions and / or verification regions and the first hot standby region.

12. The method according to claim 1, characterized in that, The determination of the target region based on the at most four regions to be restored includes: If four storage devices are detected to have failed, and the two areas to be recovered in the stripe are data areas and / or verification areas, and the other two areas to be recovered are first hot standby areas and the two data areas and / or verification areas to be recovered belong to different area groups, then the target area is determined to be all other data areas and / or verification areas.

13. A RAID data processing device, characterized in that, An apparatus for use with RAID 50, wherein the RAID 50 includes at least one stripe group, the stripe group includes multiple stripes, and the area within the stripe is divided into two area groups. Each area group includes at least two data areas, one parity area, and one first hot spare area. The data areas store data information, the parity area stores first parity information, and the first hot spare area stores second parity information. The first parity information is determined based on all data information in its area group, and the second parity information is determined based on all data information in its stripe and all first parity information in its stripe. The apparatus includes: The processing module is used to detect up to four storage devices failing and determine up to four regions to be recovered in the stripe, wherein the regions to be recovered are a data region, a verification region and / or a first hot standby region; The calculation module is used to determine the corresponding target area based on the different number and types of the areas to be recovered. The target area is a data area, a verification area, and / or a first hot standby area other than the areas to be recovered. The calculation module is also used to read target information in the target area and recover data information and / or first verification information of the area to be recovered based on the target information.

14. A chip, characterized in that, The chip includes a processor capable of executing the RAID data processing method according to any one of claims 1 to 12.

15. An electronic device, characterized in that, The electronic device includes a chip, the chip including a processor, the processor being capable of executing the RAID data processing method according to any one of claims 1 to 12.

16. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the RAID data processing method according to any one of claims 1 to 12.

17. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement the RAID data processing method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Data recovery method, device and system, electronic equipment and storage medium

    CN116501553A

  • Data recovery method and device for RAID (redundant array of independent disks)

    CN118796538A