RAID data processing method and device, chip, electronic equipment, storage medium and computer program product

By dividing the data area and check area of ​​distributed RAID6 into two area groups to store local and global check information, the problems of large read volume and insufficient recovery performance in traditional RAID6 are solved, and efficient data recovery and fault tolerance are achieved.

CN120704958AActive Publication Date: 2025-09-26SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Patent Information

Application Number
CN202511212568.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-26
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Traditional distributed RAID6 has problems with large read volume and insufficient recovery performance during data recovery, especially when multiple storage devices fail, it is unable to effectively recover data.

Method used

The data area and the check area are divided into two area groups, and local check information and global check information are stored in each area group. Data is restored through the local check formula, which reduces the reading amount and improves the recovery efficiency.

Benefits of technology

It significantly reduces the I/O load of data recovery, improves recovery performance and fault tolerance, and can efficiently recover data when multiple storage devices fail.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704958A_ABST
    Figure CN120704958A_ABST
Patent Text Reader

Abstract

The invention provides an RAID (redundant array of independent disks) data processing method and device, a chip, electronic equipment, a storage medium and a computer program product, and is applied to RAID, the method comprises the following steps: detecting that at most three storage devices fail, determining at most three corresponding to-be-recovered areas in a stripe, and storing the to-be-recovered areas in the stripe; the to-be-recovered area is a data area, a verification area, a first hot standby area and / or a second hot standby area; determining a target area based on at most three areas to be recovered, wherein the target area is a data area, a verification area, a first hot standby area and / or a second hot standby area except the areas to be recovered; and target information in the target area is read, and the data information and / or the first verification information of the to-be-recovered area are / is recovered based on the target information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a RAID data processing method, device, chip, electronic device, storage medium, and computer program product. Background Art

[0002] In the field of RAID technology, traditional distributed RAID6 consists of multiple stripe groups, each of which includes multiple stripes, and each stripe spans multiple storage devices. When a storage device fails, all storage devices participate in the data recovery process, eliminating the write bottleneck of a single storage device. However, the volume of data blocks read from each storage device is extremely large, and disk read speed becomes a performance bottleneck for data recovery. Furthermore, traditional distributed RAID6 can only simultaneously recover a maximum of two failed disks within a stripe, resulting in insufficient recovery performance. Summary of the Invention

[0003] The present application provides a RAID data processing method, device, chip, electronic device, storage medium and computer program product.

[0004] On one hand, an embodiment of the present application provides a RAID data processing method, applied to a RAID, wherein the RAID includes at least one stripe group, the stripe group includes multiple stripes, the stripe includes at least two data areas, two parity areas, a first hot spare area, and a second hot spare area, the at least two data areas and two parity areas of the stripe are divided into two area groups, the data area stores data information, the parity area stores first parity information, the first hot spare area stores second parity information, and the second hot spare area stores third parity information, the second parity information is determined based on data information and / or the first parity information in any one area group, and the third parity information is determined based on all data information and the first parity information. The method includes: Detecting failures in at most three storage devices, and determining at most three corresponding areas to be recovered in the stripe, wherein the areas to be recovered are a data area, a check area, a first hot spare area, and / or a second hot spare area; Determining a target area based on the at most three areas to be restored, the target area being the data area, the check area, the first hot standby area, and / or the second hot standby area excluding the area to be restored; Target information in the target area is read, and data information and / or first verification information of the area to be restored are restored based on the target information.

[0005] The determining of the target area based on the at most three areas to be restored includes: A storage device failure is detected, and the corresponding area to be recovered in the stripe is determined to be a data area or a check area, and the target area is determined to be other data areas and / or check areas in the area group where the area to be recovered is located and the first hot spare area.

[0006] The determining of the target area based on the at most three areas to be restored includes: Failures of two storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, and the target area is determined to be all other data areas and / or check areas.

[0007] The method further comprises: It is determined that the two areas to be restored belong to different area groups, and the target areas are determined to be all other data areas and / or verification areas and the first hot standby area.

[0008] The determining of the target area based on the at most three areas to be restored includes: Failures of two storage devices are detected, and one corresponding area to be recovered in the stripe is determined to be a data area or a check area, the other area to be recovered is a first hot spare area, and the target areas are determined to be all other data areas and check areas.

[0009] The determining of the target area based on the at most three areas to be restored includes: Failures of two storage devices are detected, and it is determined that one corresponding area to be recovered in the stripe is a data area or a check area, and the other area to be recovered is a second hot spare area, and the target areas are determined to be other data areas and / or check areas in the area group where the data area or check area is located and the first hot spare area.

[0010] The determining of the target area based on the at most three areas to be restored includes: Failures of three storage devices are detected, and it is determined that the corresponding three areas to be recovered in the stripe are all data areas and / or verification areas, and the target area is determined to be all other data areas and / or verification areas except the three areas to be recovered and the second hot spare area.

[0011] The determining of the target area based on the at most three areas to be restored includes: Failures of three storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, the other area to be recovered is the first hot spare area, and the target areas are determined to be all other data areas and / or check areas.

[0012] The determining of the target area based on the at most three areas to be restored includes: Failures of three storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, the other area to be recovered is the second hot spare area, and the target areas are determined to be all other data areas and / or check areas and the first hot spare area.

[0013] The determining of the target area based on the at most three areas to be restored includes: Failures of three storage devices are detected, and a corresponding area to be recovered in the stripe is determined to be a data area and / or a check area, and the other two areas to be recovered are respectively a first hot standby area and a second hot standby area, and the target area is determined to be all other data areas and / or check areas.

[0014] The strip further includes a third hot standby area, and the method further includes: The restored data information and / or first verification information is stored in the first hot standby area, the second hot standby area and / or the third hot standby area.

[0015] The method further comprises: Upon detecting that the failure of the storage device is resolved, reading all data information and / or first verification information of any one of the area groups; and determining second verification information in the first hot standby area based on the data information and / or first verification information; And / or, reading all data information and / or first verification information; and determining third verification information in the second hot standby area based on the data information and / or first verification information.

[0016] Another aspect of the present invention provides a RAID data processing device, comprising: A processing module is configured to detect failures in at most three storage devices and determine at most three corresponding areas to be recovered in the stripe, wherein the areas to be recovered are a data area, a check area, a first hot spare area, and / or a second hot spare area; a calculation module, configured to determine a target area based on the at most three areas to be recovered, wherein the target area is a data area, a check area, a first hot standby area, and / or a second hot standby area excluding the area to be recovered; The calculation module is further configured to read target information in the target area, and restore the data information and / or first verification information of the area to be restored based on the target information.

[0017] Among them, the calculation module is also used to detect that a storage device has failed, and determine that the corresponding area to be recovered in the stripe is a data area or a verification area, and determine that the target area is other data areas and / or verification areas in the area group where the area to be recovered is located and the first hot spare area.

[0018] The calculation module is further configured to detect failures in two storage devices, determine that the two corresponding areas to be recovered in the stripe are both data areas and / or check areas, and determine that the target area is all other data areas and / or check areas.

[0019] Among them, the computing module is also used to detect failures in three storage devices, and determine that the corresponding three areas to be recovered in the stripe are all data areas and / or verification areas, and determine that the target area is all other data areas and / or verification areas except the three areas to be recovered and the second hot standby area.

[0020] Another aspect of an embodiment of the present application provides a chip, the chip including a processor, and the processor capable of executing the RAID data processing method.

[0021] Another aspect of an embodiment of the present application provides an electronic device, the electronic device includes a chip, the chip includes a processor, and the processor is capable of executing the RAID data processing method.

[0022] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the RAID data processing method.

[0023] On the other hand, an embodiment of the present application provides a computer program product, including a computer program or instructions, for causing a processor to execute and implement the RAID data processing method provided in the embodiment of the present application.

[0024] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present application are shown in an illustrative and non-limiting manner, in which: In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.

[0026] Figure 1A flowchart of a RAID data processing method according to an embodiment of the present application is shown; Figure 2 A schematic structural diagram of a RAID data processing device according to an embodiment of the present application is shown; Figure 3 A schematic diagram of the structure of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] In order to make the purpose, features, and advantages of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0028] The solution of this application is used in distributed RAID 6 (a RAID level).

[0029] Each stripe in a distributed RAID 6 system stores a copy of the first parity information in two parity areas. Both copies of the first parity information are encoded based on the data stored in the stripe's data area. When a drive error occurs, up to two faulty drives can be recovered simultaneously.

[0030] For example, a distributed RAID 6 is composed of k+2 storage devices. Each stripe includes two parity areas (areas storing first parity information), k-2 data areas (areas storing data information), and two hot spare areas (for temporarily storing recovered data information and / or first parity information). These areas belong to different storage devices. The two copies of first parity information are p and q. The two copies of first parity information, p and q, can be determined using the data information stored in the k-2 data areas and the following formulas (1) and (2).

[0031]

[0032] in, is the XOR operation, For the Data information in each data area, For the The location information of the storage device to which each area belongs, and It is the first verification information.

[0033] When a disk error occurs, if there is only one disk error, the data information in the non-error disk area, the first check information, and the above formula (1) can be used for decoding to recover the single disk error. If there are two disk errors, the data information in the non-error disk area, the first check information, the above formula (1) and formula (2) can be used for decoding to recover the data in the two disk errors. However, whether there is one disk error or two disk errors, when recovering the data in the disk error, it is necessary to read the data information and the first check information stored in all other areas of the stripe to recover the data in the disk error. The amount of data to be read is very large, and in the case of two disk errors, the joint decoding based on formula (1) and formula (2) requires a large amount of calculation, resulting in low efficiency and slow speed of data recovery. Moreover, when there are three disk errors, recovery is impossible.

[0034] For example, as shown in Table 1, Table 1 shows a distributed RAID6 that includes 8 storage devices, divided into 2 stripe groups, each stripe group includes 6 stripes and is written to disk in a left-handed non-aligned manner. Each stripe has 4 data areas for storing data information, 2 check areas for storing first verification information, and 2 hot spare areas. The hot spare area is used to temporarily store data recovered from a faulty disk when a faulty disk occurs in the RAID. When the storage device 5 in the RIAD fails, it is necessary to read the data shown in the shaded portion in Table 1, that is, the data information and / or first verification information of all areas except the area where the faulty disk is located and the hot spare area, to restore the data information and / or first verification information in the storage device 5.

[0035] Table 1

[0036] For another example, as shown in Table 2, Table 2 also shows a distributed RAID6 that includes 8 storage devices, divided into 2 stripe groups, each stripe group includes 6 stripes, and is written to disk in a left-handed non-aligned manner. Each stripe has 4 data areas for storing data information, 2 check areas for storing first check information, and 2 hot spare areas. The hot spare area is used to temporarily store data recovered from a faulty disk when a faulty disk occurs in the RAID. When storage device 1, storage device 3, and storage device 5 in the RAID fail, the RAID cannot recover the faulty disk.

[0037] Table 2

[0038] In order to break through the performance bottleneck of data recovery, improve the speed and efficiency of data recovery, improve the recovery performance of RAID, and at the same time realize the simultaneous recovery of up to three faulty disks in RAID6, an embodiment of the present application provides a RAID data processing method, which is applied to RAID, wherein the RAID includes at least one stripe group, the stripe group includes multiple stripes, the stripe includes at least two data areas, two check areas, a first hot spare area and a second hot spare area, the at least two data areas and two check areas of the stripe are divided into two area groups, the data area stores data information, the check area stores first check information, the first hot spare area stores second check information, and the second hot spare area stores third check information, the second check information is determined based on the data information and / or the first check information in any one area group, and the third check information is determined based on all data information and the first check information.

[0039] For example, as shown in Table 3, Table 3 shows that on the basis of distributed RAID6, all data areas and check areas of each stripe are divided into two area groups. For the convenience of description, the hot spare area for storing the second check information is described as the first hot spare area, and the hot spare area for storing the third check information is described as the second hot spare area. All data information and / or first check information in any area group are encoded to obtain second check information, and the obtained second check information is stored in the first hot spare area. Stripe 1 in stripe group 1, 、 and The strip is divided into regional groups 1, 、 and The strip is divided into area groups 2, and then based on 、 and The second check information of the strip is obtained by encoding using the following formula (3): , or based on 、 and The second check information of the strip is obtained by encoding using the following formula (4): The second verification information is stored in the first hot spare area of ​​the stripe. The subsequent stripes are the same as the above method and will not be described again here.

[0040]

[0041] in, is the XOR operation, For the Data information in each data area, is the first group of regions in the strip, is the second group of regions in the strip, , It is the second check information in the stripe.

[0042] Then based on 、 、 、 、 and ,as well as 、 、 、 、 and Location information of the storage device to which it belongs 、 、 、 、 and Encode by formula (5) to obtain the third check information of the stripe The third check information is stored in the second hot standby area of ​​the stripe. The subsequent stripes are the same as the above method and will not be described in detail here.

[0043]

[0044] in, is the XOR operation, For the Data information in each data area, For the The location information of the storage device to which each area belongs, and is the first verification information, It is the third verification information.

[0045] Table 3

[0046] like Figure 1 As shown, the method includes: Step 101: Detect failures in at most three storage devices, and determine at most three corresponding areas to be recovered in the stripe, where the areas to be recovered are a data area, a check area, a first hot spare area, and / or a second hot spare area.

[0047] On the basis of traditional RAID6, after the second and third parity information are obtained by further encoding in the above manner, the information in the corresponding to-be-recovered area of ​​each stripe in each storage device can be recovered when up to three storage devices fail.

[0048] The area to be restored may be a data area, a check area, a first hot standby area and / or a second hot standby area.

[0049] Step 102: determining a target area based on the at most three areas to be recovered, wherein the target area is a data area, a check area, a first hot standby area, and / or a second hot standby area excluding the area to be recovered.

[0050] Depending on the number and type of areas to be restored, a corresponding target area needs to be determined, and then the target information in the target area is read to restore the information in the area to be restored. The target area can be a data area outside the area to be restored, a check area, a first hot standby area, and / or a second hot standby area.

[0051] Step 103: Read the target information in the target area, and restore the data information and / or first verification information of the area to be restored based on the target information.

[0052] When only one storage device fails at the same time, the corresponding area to be recovered in the stripe may be the data area, the check area, the first hot spare area, or the second hot spare area.

[0053] If the area to be recovered is a data area or a check area, the target area may be determined to be all data areas and / or check areas in the area group where the area to be recovered is located, excluding the area to be recovered, and the first hot standby area. The target data in the target area is read and the information in the area to be recovered is recovered using the above formula (3) or formula (4).

[0054] If the area to be restored is the first hot spare area, no restoration is required. After the failure of the storage device is resolved or repaired, all data information and / or first verification information in any area group are read, and the second verification information is re-determined using the above formula (3) or formula (4) and stored in the restored first hot spare area.

[0055] If the area to be restored is the second hot spare area, no restoration is required. After the failure of the storage device is resolved or repaired, all data information and the first verification information are read, and the third verification information is re-determined using the above formula (5) and stored in the restored second hot spare area.

[0056] When two storage devices fail simultaneously, there are three possible scenarios: The first type is that the two corresponding areas to be recovered in the stripe are both data areas and / or check areas. In this case, a judgment can be made to determine whether the two areas to be recovered are in the same area group. If they are in the same area group, the target area is determined to be all data areas and check areas except the two areas to be recovered, and the data information and / or the first check information in the two areas to be recovered are restored using the above formulas (1) and (2). If they are not in the same area group, the target area is determined to be all data areas and check areas except the two areas to be recovered, as well as the first hot spare area, and the data information and / or the first check information in the two areas to be recovered are restored using the above formulas (3) and (4).

[0057] It should be pointed out that when the two areas to be recovered are not in the same area group, they can also be recovered by the recovery method in the same area group. However, although the recovery method when they are not in the same area group reads the second verification information in the first hot standby area more than the other recovery method, the recovery by the above formula (3) and formula (4) has lower computational complexity than the recovery by formula (1) and formula (2), that is, the calculation speed is faster and the efficiency is higher.

[0058] The second type is that one of the corresponding areas to be recovered in the stripe is the data area or the check area, and the other area to be recovered is the first hot spare area. In this case, since the second check information in the first hot spare area is lost, it cannot be recovered using formula (3) or formula (4). Therefore, the target area is determined to be all data areas and check areas except the area to be recovered, the target information in the target area is read, and the data information or the first check information in the area to be recovered is recovered using the above formula (1). The second check information in the first hot spare area can be read after the failure of the storage device is repaired, and the second check information can be re-determined using the above formula (3) or formula (4) and stored in the restored first hot spare area.

[0059] In the third method, one of the corresponding areas to be recovered in the stripe is a data area or a check area, and the other area to be recovered is a second hot spare area. The target area is determined to be all data areas and / or check areas in the area group where the data area or check area to be recovered is located, excluding the data area or check area, and the first hot spare area. The target information in the target area is read, and the data information or check information is recovered using the above formula (3) or formula (4). The third check information in the second hot spare area can be re-determined using the above formula (5) after the storage device fault is repaired by reading all the data information and / or the first check information and storing it in the restored second hot spare area.

[0060] When three storage devices fail simultaneously, four situations may occur: In the first method, the three corresponding areas to be recovered in the stripe are all data areas and / or parity areas. The target areas are determined to be all data areas and / or parity areas except the three areas to be recovered, as well as the second hot spare area. The target information in the target areas is read, and the data information and / or first parity information in the three areas to be recovered are recovered using the above formulas (1), (2), and (5).

[0061] In the second type, the corresponding two areas to be recovered in the stripe are data areas or check areas, and the other area to be recovered is the first hot standby area.

[0062] At this time, since the second check information in the first hot spare area is lost, it cannot be restored using formula (3) and formula (4). Therefore, the target area is determined to be all data areas and check areas except the area to be restored, the target information in the target area is read, and the data information or first check information in the area to be restored is restored using the above formula (1) and formula (2). After the failure of the storage device is repaired, the second check information in the first hot spare area can be read from all data information and / or first check information in any group of area groups, and the second check information can be re-determined using the above formula (3) or formula (4) and stored in the restored first hot spare area.

[0063] In the third type, the two corresponding areas to be recovered in the stripe are data areas or check areas, and the other area to be recovered is a second hot standby area.

[0064] At this time, a judgment can also be made first to determine whether the two areas to be restored are in the same area group. If they are in the same area group, the target area is determined to be all data areas and check areas except the two areas to be restored, and the data information and / or first check information in the two areas to be restored are restored using the above formulas (1) and (2). If they are not in the same area group, the target area is determined to be all data areas and check areas except the two areas to be restored and the first hot spare area, and the data information and / or first check information in the two areas to be restored are restored using the above formulas (3) and (4). The third check information in the second hot spare area can be read after the fault of the storage device is repaired, and the third check information can be re-determined using the above formula (5) and stored in the restored second hot spare area.

[0065] In the fourth type, a corresponding area to be recovered in the stripe is a data area or a check area, one area to be recovered is a first hot standby area, and another area to be recovered is a second hot standby area.

[0066] The target area is determined to be all data areas and / or check areas except the area to be restored, the target information in the target area is read, and the data information or first check information in the area to be restored is restored using formulas (1) and (2). After the failure of the storage device is repaired, all data areas and / or check areas in any group of area groups are read, and the second check information is re-determined using formula (3) or (4) and stored in the restored first hot spare area. All data information and / or first check information are read, and the third check information is re-determined using formula (5) and stored in the restored second hot spare area.

[0067] In the above scheme, by dividing the data area and check area in the traditional distributed RAID6 into two area groups, the local second check information is determined to be stored in the first hot spare area based on the data information and / or the first check information in the area group, and the global third check information is determined to be stored in the second hot spare area based on the entire data information and the first check information, thereby significantly improving the system's fault tolerance, recovery performance, and recovery efficiency. When the fault area is located in a certain area group, only the other data information and / or the first check information in the same group need to be read (rather than all the data information and the first check information of the entire stripe), and the local check formula (3) or formula (4) is used for direct recovery. This reduces the amount of data read when recovering a single data information or first check information from the entire stripe to the area group size, significantly reducing the I / O load and completely solving the performance bottleneck problem of full stripe reading in the traditional scheme. When two faulty regions occur simultaneously and belong to different region groups, by reading the data information and / or first check information of the two region groups and the second check information of the first hot spare region, each group of faulty regions can be independently recovered using formulas (3) and (4). Compared with the traditional RAID6 in which the data information and first check information of the entire stripe are decoded and recovered using formulas (1) and (2), the computational complexity is significantly reduced and the recovery efficiency is improved. When three faulty regions occur simultaneously (all data regions and / or check regions), all data information, first check information, and third check information can be read, and the three faulty regions can be recovered using formulas (1), (2), and (5), significantly improving the system's fault tolerance and recovery performance.

[0068] In an example of the present application, a RAID data processing method is further provided, wherein determining a target area based on the at most three areas to be recovered includes: A storage device failure is detected, and the corresponding area to be recovered in the stripe is determined to be a data area or a check area, and the target area is determined to be other data areas and / or check areas in the area group where the area to be recovered is located and the first hot spare area.

[0069] For example, as shown in Table 4, in the distributed RAID6 shown in Table 4, if storage device 4 fails, the shaded area in Table 4 is the area to be recovered for each stripe. The area to be recovered for stripe 1 in stripe group 1 is The data area where the stripe 2 in stripe group 1 is located is The area to be restored for stripe 1 in stripe group 2 is The data area where the stripe 2 in stripe group 2 is located. The verification area.

[0070] When only one storage device fails and the corresponding area to be recovered in the stripe is a data area or a check area, the target area can be determined to be all data areas and / or check areas in the area group where the area to be recovered is located except the area to be recovered and the first hot spare area.

[0071] In the distributed RAID6 shown in Table 4, stripe 1 in stripe group 1 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The data area where the target area of ​​stripe 1 in stripe group 1 is determined to be 、 The verification area and The first hot standby area where the stripe group 1 is located. Stripe 2 in stripe group 1 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The target area of ​​stripe 2 in stripe group 1 is determined to be The data area and The verification area and The first hot standby area where the stripe group 2 is located. Stripe 1 in stripe group 2 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The data area where the target area of ​​stripe 1 in stripe group 2 is determined to be 、 The verification area and The first hot standby area where the stripe 2 in stripe group 2 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The target area of ​​stripe 2 in stripe group 2 is determined to be The data area and The verification area and After reading the target information of the target area in the stripe, the data information or the first check information in the area to be recovered is recovered using formula (3) or formula (4) (determined based on the area group where the area to be recovered is located).

[0072] Table 4

[0073] In the above scheme, when a single storage device failure is detected and the area to be recovered is the data area or the check area, the target area can be reduced from the "all data areas and check areas" of the traditional RAID6 to "other data areas and / or check areas in the same area group and the first hot spare area" by using the second check information pre-determined by the information in the area group. Only the other data information and / or first check information in the same group (rather than all the data information and first check information of the entire stripe) and the second check information need to be read to directly recover the area using the local check formula (3) or formula (4). This reduces the amount of data read when recovering a single piece of data information or the first check information from the entire stripe to the area group size, significantly reducing the I / O load and completely resolving the performance bottleneck problem of full stripe reading in the traditional scheme.

[0074] In an example of the present application, a RAID data processing method is further provided, wherein determining a target area based on the at most three areas to be recovered includes: Failures of two storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, and the target area is determined to be all other data areas and / or check areas.

[0075] For example, as shown in Table 5, in the distributed RAID 6 shown in Table 5, if storage device 3 and storage device 4 fail at the same time, the shaded area in Table 5 is the area to be recovered for each stripe. The area to be recovered for stripe 1 in stripe group 1 is and The data area where the stripe 2 in stripe group 1 is located is The data area and The area to be restored for stripe 1 in stripe group 2 is and The data area where the stripe 2 in stripe group 2 is located. The data area and The verification area.

[0076] When two storage devices fail at the same time, and the corresponding areas to be recovered in the stripe are all data areas or check areas, the target area may be determined to be all data areas and / or check areas except the areas to be recovered.

[0077] In the distributed RAID6 shown in Table 5, the area to be recovered for stripe 1 in stripe group 1 is and The data area where the target area of ​​stripe 1 in stripe group 1 is determined to be 、 The data area and 、 The area to be restored for stripe 2 in stripe group 1 is The data area and The target area of ​​stripe 2 in stripe group 1 is determined to be 、 、 The data area and The area to be restored for stripe 1 in stripe group 2 is and The data area where the target area of ​​stripe 1 in stripe group 2 is determined to be 、 The data area and 、 The area to be restored for stripe 2 in stripe group 2 is The data area and The target area of ​​stripe 2 in stripe group 2 is determined to be 、 、 The data area and After reading the target information of the target area in the stripe, the data information and / or the first check information in the area to be recovered are recovered by performing joint decoding using formula (1) and formula (2).

[0078] Table 5

[0079] In an example of the present application, a RAID data processing method is further provided, the method further comprising: It is determined that the two areas to be restored belong to different area groups, and the target areas are determined to be all other data areas and / or verification areas and the first hot standby area.

[0080] When a failure is detected in two storage devices, it can be determined whether the corresponding areas to be recovered in the stripes of the two storage devices belong to the same area group. If the two corresponding areas to be recovered belong to the same area group, the target areas are determined to be all data areas and parity areas excluding the area to be recovered. If the two corresponding areas to be recovered belong to different area groups, the target areas are determined to be all data areas and parity areas excluding the area to be recovered, as well as the first hot spare area.

[0081] For example, in the distributed RAID6 shown in Table 5, stripe 1 in stripe group 1 will be 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are and The data area where the two areas to be recovered belong to different area groups, so the target area of ​​stripe 1 in stripe group 1 is determined to be 、 The data area and 、 The verification area and The first hot standby area where the stripe group 1 is located. Stripe 2 in stripe group 1 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The data area and The two areas to be recovered belong to the same area group, so the target area of ​​stripe 2 in stripe group 1 is determined to be 、 、 The data area and The parity area where the stripe is located. Stripe 1 in stripe group 2 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are and The data area where the two areas to be recovered belong to different area groups, so the target area of ​​stripe 1 in stripe group 2 is determined to be 、 The data area and 、 The verification area and The first hot standby area where the stripe 2 in stripe group 2 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The data area and The two areas to be recovered belong to the same area group, so the target area of ​​stripe 2 in stripe group 2 is determined to be 、 、 The data area and For stripe 1 in stripe group 1 and stripe 1 in stripe group 2, after reading the target information of the target area in the stripe, joint decoding is performed using formulas (3) and (4) to recover the data information and / or the first check information in the area to be recovered. For stripe 2 in stripe group 1 and stripe 2 in stripe group 2, after reading the target information of the target area in the stripe, joint decoding is performed using formulas (1) and (2) to recover the data information and / or the first check information in the area to be recovered.

[0082] In the above scheme, when it is detected that two storage devices have failed at the same time and the areas to be recovered are all data areas or check areas, the second check information pre-determined by the information in the area group is used to read the data in all data areas and check areas except the area to be recovered and the data in the first hot spare area, and the data information and / or the first check information in the area to be recovered are recovered. Although compared with the decoding and recovery of the data information and the first check information of the full stripe in the traditional RAID6 using formulas (1) and (2), the read data includes the second check information in the first hot spare area, the decoding and recovery using formulas (3) and (4) has lower computational complexity, thus significantly improving the computational speed and recovery efficiency.

[0083] In an example of the present application, a RAID data processing method is further provided, wherein determining a target area based on the at most three areas to be recovered includes: Failures of two storage devices are detected, and one corresponding area to be recovered in the stripe is determined to be a data area or a check area, the other area to be recovered is a first hot spare area, and the target areas are determined to be all other data areas and check areas.

[0084] When one of the two areas to be recovered is a data area or a check area, and the other is the first hot standby area, since the second check information in the first hot standby area is lost, recovery cannot be performed using formula (3) or formula (4). Therefore, the target area is determined to be all data areas and check areas except the area to be recovered.

[0085] For example, as shown in Table 6, in the distributed RAID 6 shown in Table 6, if storage device 3 and storage device 7 fail at the same time, the shaded area in Table 6 is the area to be recovered for each stripe. The area to be recovered for stripe 1 in stripe group 1 is The data area and The first hot standby area where the stripe 1 in stripe group 2 is located is The data area and The first hot standby zone.

[0086] In the distributed RAID6 shown in Table 6, the area to be recovered for stripe 1 in stripe group 1 is The data area and The first hot standby area is located, then the target area of ​​stripe 1 in stripe group 1 is determined to be 、 、 The data area and 、 The area to be restored for stripe 1 in stripe group 2 is The data area and The first hot standby area is located, then the target area of ​​stripe 1 in stripe group 2 is determined to be 、 、 The data area and 、 After reading the target information of the target area in the stripe, the data information and / or the first check information in the area to be recovered are recovered by performing joint decoding using formula (1) and formula (2).

[0087] Table 6

[0088] The second verification information in the first hot spare area can be determined by reading all data information and / or first verification information in any group of area groups after the failure of the storage device is repaired, and the second verification information is re-determined using the above formula (3) or formula (4) and stored in the restored first hot spare area.

[0089] In an example of the present application, a RAID data processing method is further provided, wherein determining a target area based on the at most three areas to be recovered includes: Failures of two storage devices are detected, and it is determined that one corresponding area to be recovered in the stripe is a data area or a check area, and the other area to be recovered is a second hot spare area, and the target areas are determined to be other data areas and / or check areas in the area group where the data area or check area is located and the first hot spare area.

[0090] When one of the two corresponding areas to be recovered is a data area or a check area, and the other is a second hot spare area, recovery can be performed using the second check information in the first hot spare area. Therefore, the target area is determined to be all data areas and / or check areas in the area group where the data area or check area to be recovered is located, except for the data area or check area, and the first hot spare area.

[0091] For example, as shown in Table 6, in the distributed RAID 6 shown in Table 6, if storage device 3 and storage device 7 fail at the same time, the shaded area in Table 6 is the area to be recovered for each stripe. The area to be recovered for stripe 2 in stripe group 1 is The data area and The second hot standby area. The area to be restored for stripe 2 in stripe group 2 is The data area and The second hot standby zone.

[0092] In the distributed RAID6 shown in Table 6, stripe 2 in stripe group 1 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The data area and The second hot standby area is located, then the target area of ​​stripe 2 in stripe group 1 is determined to be 、 The verification area and The first hot standby area where the stripe 2 in stripe group 2 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The data area and The second hot standby area is located, then the target area of ​​stripe 2 in stripe group 2 is determined to be 、 The verification area and After reading the target information of the target area in the stripe, joint decoding is performed using formula (3) or formula (4) to recover the data information and / or the first check information in the area to be recovered.

[0093] The third verification information in the second hot spare area can be re-determined by reading all data information and / or first verification information after the failure of the storage device is repaired, and the third verification information is re-determined by the above formula (5) and stored in the restored second hot spare area.

[0094] In the above scheme, when two storage device failures are detected at the same time, and one of the areas to be recovered is a data area or a check area and the other is a second hot spare area, the target area can also be reduced from the "all data areas and check areas" of traditional RAID6 to "other data areas and / or check areas in the same area group and the first hot spare area" by using the second check information pre-determined by the information in the area group. Only the other data information and / or first check information in the same group (rather than all the data information and first check information of the entire stripe) and the second check information need to be read, and the local check formula (3) or formula (4) can be used for direct recovery. This reduces the amount of data read when recovering a single piece of data information or first check information from the entire stripe to the area group size, significantly reducing the I / O load and completely solving the performance bottleneck problem of full stripe reading in traditional schemes.

[0095] In an example of the present application, a RAID data processing method is further provided, wherein determining a target area based on the at most three areas to be recovered includes: Failures of three storage devices are detected, and it is determined that the corresponding three areas to be recovered in the stripe are all data areas and / or verification areas, and the target area is determined to be all other data areas and / or verification areas except the three areas to be recovered and the second hot spare area.

[0096] When it is detected that three storage devices fail at the same time, and the corresponding three areas to be recovered in the stripe are all data areas and / or check areas, the target areas are determined to be all data areas and / or check areas except the three areas to be recovered, as well as the second hot spare area, the target information in the target areas is read, and the data information and / or the first check information in the three areas to be recovered are recovered using the above formulas (1), (2), and (5).

[0097] For example, as shown in Table 7, in the distributed RAID 6 shown in Table 7, if storage device 2, storage device 3, and storage device 5 fail at the same time, the shaded area in Table 7 is the area to be recovered for each stripe. The area to be recovered for stripe 1 in stripe group 1 is 、 The data area and The area to be restored for stripe 2 in stripe group 1 is 、 The data area and The area to be restored for stripe 1 in stripe group 2 is 、 The data area and The area to be restored for stripe 2 in stripe group 2 is 、 The data area and The verification area.

[0098] In the distributed RAID6 shown in Table 7, the area to be recovered for stripe 1 in stripe group 1 is 、 The data area and The target area of ​​stripe 1 in stripe group 1 is determined to be 、 The data area and The verification area and The second hot standby area where the stripe 2 in stripe group 1 is located is 、 The data area and The target area of ​​stripe 2 in stripe group 1 is determined to be 、 The data area and The verification area and The second hot standby area. The area to be restored for stripe 1 in stripe group 2 is 、 The data area and The target area of ​​stripe 1 in stripe group 2 is determined to be 、 The data area and The verification area and The second hot standby area. The area to be restored for stripe 2 in stripe group 2 is 、 The data area and The target area of ​​stripe 2 in stripe group 2 is determined to be 、 The data area and After reading the target information of the target area in the stripe, the data information and / or the first check information in the area to be recovered are recovered by performing joint decoding using formulas (1), (2) and (5).

[0099] Table 7

[0100] In the above solution, when three storage device failures are detected simultaneously and the areas to be recovered are all data areas or parity areas, the system uses third parity information pre-determined using all data information in the stripe and the first parity information to read data from all data areas and parity areas other than the area to be recovered, as well as the second hot spare area. The data information and / or first parity information in the three areas to be recovered are then restored. This enables simultaneous recovery of the three failed disks, significantly improving the system's fault tolerance and recovery performance.

[0101] In an example of the present application, a RAID data processing method is further provided, wherein determining a target area based on the at most three areas to be recovered includes: Failures of three storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, the other area to be recovered is the first hot spare area, and the target areas are determined to be all other data areas and / or check areas.

[0102] When three storage devices are detected to have failed simultaneously, and of the three corresponding areas to be recovered in the stripe, two are data areas and / or check areas, and the other is the first hot spare area. At this time, due to the loss of the second check information, recovery cannot be performed using formulas (3) and (4). Therefore, the target area is determined to be all data areas and / or check areas except the three areas to be recovered, the target information in the target area is read, and the two data information and / or first check information to be recovered are recovered using the above formulas (1) and (2).

[0103] For example, as shown in Table 8, in the distributed RAID 6 shown in Table 8, if storage device 3, storage device 4, and storage device 7 fail at the same time, the shaded area in Table 8 is the area to be recovered for each stripe. The area to be recovered for stripe 1 in stripe group 1 is 、 The data area and The first hot standby area where the stripe 1 in stripe group 2 is located is 、 The data area and The first hot standby zone.

[0104] In the distributed RAID6 shown in Table 8, the area to be recovered for stripe 1 in stripe group 1 is 、 The data area and The first hot standby area is located, then the target area of ​​stripe 1 in stripe group 1 is determined to be 、 The data area and 、 The area to be restored for stripe 1 in stripe group 2 is 、 The data area and The first hot standby area is located, then the target area of ​​stripe 1 in stripe group 2 is determined to be 、 The data area and 、 After reading the target information of the target area in the stripe, the data information and / or the first check information in the area to be recovered are recovered by performing joint decoding using formula (1) and formula (2).

[0105] Table 8

[0106] The second verification information in the first hot spare area can be determined by reading all data information and / or first verification information in any group of area groups after the failure of the storage device is repaired, and the second verification information is re-determined using the above formula (3) or formula (4) and stored in the restored first hot spare area.

[0107] In an example of the present application, a RAID data processing method is further provided, wherein determining a target area based on the at most three areas to be recovered includes: Failures of three storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, the other area to be recovered is the second hot spare area, and the target areas are determined to be all other data areas and / or check areas and the first hot spare area.

[0108] When three storage devices are detected to have failed simultaneously, and two of the three corresponding areas to be recovered in the stripe are data areas and / or parity areas, and one is a second hot spare area, it can be determined whether the two data areas and / or parity areas to be recovered and the corresponding areas to be recovered in the stripe belong to the same area group. If the two corresponding areas to be recovered belong to the same area group, the target areas are determined to be all data areas and parity areas excluding the areas to be recovered. If the two corresponding areas to be recovered belong to different area groups, the target areas are determined to be all data areas and parity areas excluding the areas to be recovered and the first hot spare area.

[0109] For example, as shown in Table 8, in the distributed RAID 6 shown in Table 8, if storage device 3, storage device 4, and storage device 7 fail at the same time, the shaded area in Table 8 is the area to be recovered for each stripe. The area to be recovered for stripe 2 in stripe group 1 is The data area and The verification area and The second hot standby area. The area to be restored for stripe 2 in stripe group 2 is The data area and The verification area and The second hot standby zone.

[0110] In the distributed RAID6 shown in Table 8, stripe 2 in stripe group 1 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The data area and The verification area and The second hot spare area where the data area and the check area belong to the same area group, then the target area of ​​stripe 2 in stripe group 1 is determined to be 、 、 The data area and After reading the target information of the target area in the stripe, the data information and / or the first check information in the area to be recovered are restored by performing joint decoding using formula (1) and formula (2). Stripe 2 in stripe group 2 will 、 and Divide into a group of regional groups, 、 and Divided into another group of regional groups, the areas to be restored are The data area and The verification area and The second hot spare area where the data area and the check area belong to different area groups, then the target area of ​​stripe 2 in stripe group 2 is determined to be 、 The data area and After reading the target information of the target area in the stripe, the data information and / or the first check information in the area to be recovered are recovered by decoding using formula (3) and formula (4).

[0111] The third verification information in the second hot spare area can be re-determined by reading all data information and / or first verification information after the failure of the storage device is repaired, and the third verification information is re-determined by the above formula (5) and stored in the restored second hot spare area.

[0112] In the above scheme, when it is detected that two storage devices fail at the same time, and one of the areas to be recovered is a data area or a check area and the other is a second hot spare area, the second check information determined in advance by the information in the area group can also be used to read all data areas and check areas except the area to be recovered and the data in the first hot spare area, and recover the data information and / or the first check information in the area to be recovered. Although compared with the traditional RAID6 in which the data information and the first check information of the entire stripe are decoded and recovered using formulas (1) and (2), the read data includes the second check information in the first hot spare area, the decoding and recovery using formulas (3) and (4) has lower computational complexity, thereby significantly improving the computational speed and recovery efficiency.

[0113] In an example of the present application, a RAID data processing method is further provided, wherein determining a target area based on the at most three areas to be recovered includes: Failures of three storage devices are detected, and a corresponding area to be recovered in the stripe is determined to be a data area and / or a check area, and the other two areas to be recovered are respectively a first hot standby area and a second hot standby area, and the target area is determined to be all other data areas and / or check areas.

[0114] When it is detected that three storage devices fail at the same time, and among the three corresponding areas to be recovered in the stripe, one is a data area or a check area, one is a first hot spare area, and another is a second hot spare area, the target area is determined to be all data areas and / or check areas except the three areas to be recovered, the target information in the target area is read, and the data information or the first check information in the data area or the check area is recovered using the above formula (1) and formula (2).

[0115] For example, as shown in Table 9, in the distributed RAID 6 shown in Table 9, if storage device 2, storage device 6, and storage device 7 fail at the same time, the shaded area in Table 9 is the area to be recovered for each stripe. The area to be recovered for stripe 2 in stripe group 1 is The data area, The first hot standby area and The second hot standby area. The area to be restored for stripe 2 in stripe group 2 is The data area, The first hot standby area and The second hot standby zone.

[0116] In the distributed RAID6 shown in Table 9, the area to be recovered for stripe 2 in stripe group 1 is The data area, The first hot standby area and The second hot standby area is located, then the target area of ​​stripe 2 in stripe group 1 is determined to be 、 、 The data area and 、 The area to be restored for stripe 2 in stripe group 2 is The data area, The first hot standby area and The second hot standby area is located, then the target area of ​​stripe 2 in stripe group 2 is determined to be 、 、 The data area and 、 After reading the target information of the target area in the stripe, the data information and / or the first check information in the area to be recovered are recovered by performing joint decoding using formula (1) and formula (2).

[0117] Table 9

[0118] The second check information in the first hot spare area and the third check information in the second hot spare area can be determined by reading all data information and / or first check information in any area group after the storage device is repaired, re-determining the second check information using the above formula (3) or formula (4), and storing it in the restored first hot spare area. Alternatively, all data information and / or first check information can be read, re-determining the third check information using the above formula (5), and storing it in the restored second hot spare area.

[0119] In an example of the present application, a RAID data processing method is further provided, wherein the stripe further includes a third hot spare area, and the method further includes: The restored data information and / or first verification information is stored in the first hot standby area, the second hot standby area and / or the third hot standby area.

[0120] When a storage device fails and the corresponding data information to be recovered and / or first verification information in the stripe is recovered, if there is a hot spare area in the stripe that has not failed, the recovered data information and / or first verification information can be temporarily stored in the hot spare area.

[0121] If only one data area or parity area fails, after the data information or first parity information is recovered, it can be stored in either the first or second hot spare area. Because the computational complexity required to recover the second parity information is less than that required to recover the third parity information, if neither the first nor the second hot spare area has failed, the first hot spare area is prioritized for storage.

[0122] When there are two faulty data areas and / or check areas at the same time, after the data information and / or first check information is restored, if neither the first hot standby area nor the second hot standby area has faults, it can be stored in the first hot standby area or the second hot standby area.

[0123] When there are three faulty data areas and / or check areas at the same time, a third hot spare area needs to be provided in the stripe. After the data information and / or the first check information are restored, if there are no faults in the first hot spare area, the second hot spare area and the third hot spare area, they can be stored in the first hot spare area, the second hot spare area and the third hot spare area.

[0124] It should be pointed out that in order to improve the system fault tolerance, multiple hot standby areas can be set up.

[0125] In the above solution, after the data information and / or the first verification information to be recovered are restored, they are stored in the hot spare area of ​​the corresponding stripe, eliminating the need to wait for the bad disk to be replaced or repaired, significantly improving the timeliness of data recovery.

[0126] In an example of the present application, a RAID data processing method is further provided, the method further comprising: Upon detecting that the failure of the storage device is resolved, reading all data information and / or first verification information of any one of the area groups; and determining second verification information in the first hot standby area based on the data information and / or first verification information; And / or, reading all data information and / or first verification information; and determining third verification information in the second hot standby area based on the data information and / or first verification information.

[0127] After the failed storage device is repaired or replaced, it is necessary to restore the second verification information and / or the third verification information that failed or was overwritten by the data information and the first verification information.

[0128] All data information and / or first verification information in any group of regional groups are read, and the second verification information is re-determined using the above formula (3) or formula (4).

[0129] All data information and / or first verification information are read, and the third verification information is re-determined using the above formula (5).

[0130] After the second verification information or the third verification information is restored, if the corresponding first hot standby area or the second hot standby area is not covered by the data information and the first verification information, the data information may be stored in the corresponding first hot standby area or the second hot standby area.

[0131] If the corresponding first hot standby area or second hot standby area has been overwritten by the data information and the first verification information, the data information and the first verification information in the first hot standby area or the second hot standby area can be first transferred back to the restored data area or verification area, and then the restored second verification information or the third verification information can be stored in the first hot standby area or the second hot standby area. Alternatively, the first hot standby area or the second hot standby area can be modified into a data area or verification area based on the stored data information or the first verification information, and then the restored second verification information or the third verification information can be stored in the restored data area or verification area, and then the restored data area or verification area can be modified into the first hot standby area or the second hot standby area based on the stored second verification information or the third verification information.

[0132] In the above solution, after the faulty storage device is rectified, the restored data temporarily stored in the first or second hot spare area can be migrated back to the restored original area. The regenerated second or third verification information can then be written back to the corresponding first or second hot spare area. This improves the flexibility of data recovery while maintaining the overall structural integrity of the system after data recovery.

[0133] In order to implement the above-mentioned RAID data processing method, as Figure 2 As shown, an example of the present application provides a RAID data processing device, including: Processing module 201 is configured to detect failures in at most three storage devices and determine at most three corresponding areas to be recovered in the stripe, wherein the areas to be recovered are the data area, the check area, the first hot spare area, and / or the second hot spare area; A calculation module 202 is configured to determine a target area based on the at most three areas to be recovered, where the target area is a data area, a check area, a first hot standby area, and / or a second hot standby area other than the area to be recovered; The calculation module 202 is further configured to read target information in the target area, and restore the data information and / or first verification information of the area to be restored based on the target information.

[0134] In which, the calculation module 202 is also used to detect a failure of a storage device, and determine that the corresponding area to be recovered in the stripe is a data area or a check area, and determine that the target area is other data areas and / or check areas in the area group where the area to be recovered is located and the first hot spare area.

[0135] The calculation module 202 is further configured to detect failures in two storage devices, determine that the two corresponding areas to be recovered in the stripe are both data areas and / or check areas, and determine that the target area is all other data areas and / or check areas.

[0136] The calculation module 202 is further configured to determine that the two areas to be recovered belong to different area groups, and to determine that the target areas are all other data areas and / or verification areas and the first hot standby area.

[0137] The calculation module 202 is further used to detect failures in two storage devices, and determine that one of the corresponding areas to be recovered in the stripe is a data area or a check area, the other area to be recovered is a first hot standby area, and the target area is determined to be all other data areas and check areas.

[0138] In which, the calculation module 202 is also used to detect that two storage devices have failed, and determine that one of the corresponding areas to be recovered in the stripe is a data area or a verification area, and the other area to be recovered is a second hot spare area, and determine that the target area is other data areas and / or verification areas in the area group where the data area or verification area is located and the first hot spare area.

[0139] In which, the calculation module 202 is also used to detect failures in three storage devices, and determine that the corresponding three areas to be recovered in the stripe are all data areas and / or verification areas, and determine that the target area is all other data areas and / or verification areas except the three areas to be recovered and the second hot standby area.

[0140] In which, the calculation module 202 is also used to detect failures in three storage devices, and determine that the corresponding two areas to be recovered in the stripe are both data areas and / or verification areas, the other area to be recovered is the first hot standby area, and the target area is determined to be all other data areas and / or verification areas.

[0141] In which, the calculation module 202 is also used to detect failures in three storage devices, and determine that the corresponding two areas to be recovered in the stripe are both data areas and / or verification areas, and the other area to be recovered is the second hot backup area, and determine that the target area is all other data areas and / or verification areas and the first hot backup area.

[0142] In which, the calculation module 202 is also used to detect failures in three storage devices, and determine that a corresponding area to be recovered in the stripe is the data area and / or the verification area, and the other two areas to be recovered are the first hot standby area and the second hot standby area, and the target area is determined to be all other data areas and / or verification areas.

[0143] The processing module 201 is further configured to store the restored data information and / or first verification information into the first hot standby area, the second hot standby area, and / or the third hot standby area.

[0144] The calculation module 202 is further configured to detect that the failure of the storage device is resolved, read all data information and / or first verification information of any one of the region groups, and determine second verification information in the first hot standby region based on the data information and / or first verification information. And / or, the calculation module 202 is further configured to read all data information and / or first verification information; and determine third verification information in the second hot standby area based on the data information and / or first verification information.

[0145] An embodiment of the present application further provides a chip, which includes a processor capable of executing the RAID data processing method provided in the embodiment of the present application.

[0146] An embodiment of the present application also provides an electronic device.

[0147] Figure 3 A schematic block diagram of an example electronic device that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0148] like Figure 3 As shown, electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. RAM 303 may also store various programs and data required for the operation of device 300. Computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.

[0149] Various components in device 300 are connected to I / O interface 305, including: an input unit 306, such as a keyboard, mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a magnetic disk, optical disk, etc.; and a communication unit 309, such as a network card, modem, wireless communication transceiver, etc. The communication unit 309 allows device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0150] The computing unit 301 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as the RAID data processing method. For example, in some embodiments, the RAID data processing method may be implemented as a computer software program tangibly embodied on a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the RAID data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform the RAID data processing method via any other suitable means (e.g., via firmware).

[0151] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein a computer program is stored therein, and the computer program is used to execute the RAID data processing method provided by the embodiment of the present application.

[0152] An embodiment of the present application provides a computer program product, comprising a computer program or instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program or instructions from the computer-readable storage medium and executes the computer program or instructions, causing the computer device to perform the RAID data processing method described above in the embodiment of the present application.

[0153] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.

[0154] In some embodiments, a computer program may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0155] By way of example, a computer program may be deployed to be executed on one computing device or on multiple computing devices at one site or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0156] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0157] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0158] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0159] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0160] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0161] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0162] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0163] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0164] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A RAID data processing method, characterized in that: A method for applying a RAID, wherein the RAID includes at least one stripe group, the stripe group includes multiple stripes, the stripe includes at least two data areas, two parity areas, a first hot spare area, and a second hot spare area, the at least two data areas and two parity areas of the stripe are divided into two area groups, the data area stores data information, the parity area stores first parity information, the first hot spare area stores second parity information, and the second hot spare area stores third parity information, the second parity information is determined based on data information and / or the first parity information in any one area group, and the third parity information is determined based on all data information and the first parity information. Detecting failures in at most three storage devices, and determining at most three corresponding areas to be recovered in the stripe, the areas to be recovered being a data area, a check area, a first hot spare area, and / or a second hot spare area; Determining a target area based on the at most three areas to be restored, the target area being the data area, the check area, the first hot standby area, and / or the second hot standby area excluding the area to be restored; Target information in the target area is read, and data information and / or first verification information of the area to be restored are restored based on the target information.

2. The method according to claim 1, characterized in that The determining of the target area based on the at most three areas to be restored includes: A storage device failure is detected, and the corresponding area to be recovered in the stripe is determined to be a data area or a check area, and the target area is determined to be other data areas and / or check areas in the area group where the area to be recovered is located and the first hot spare area.

3. The method according to claim 1, characterized in that The determining of the target area based on the at most three areas to be restored includes: Failures of two storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, and the target area is determined to be all other data areas and / or check areas.

4. The method according to claim 3, characterized in that The method further comprises: It is determined that the two areas to be restored belong to different area groups, and the target areas are determined to be all other data areas and / or verification areas and the first hot standby area.

5. The method according to claim 1, wherein The determining of the target area based on the at most three areas to be restored includes: Failures of two storage devices are detected, and one corresponding area to be recovered in the stripe is determined to be a data area or a check area, the other area to be recovered is a first hot spare area, and the target areas are determined to be all other data areas and check areas.

6. The method according to claim 1, characterized in that The determining of the target area based on the at most three areas to be restored includes: Failures of two storage devices are detected, and it is determined that one corresponding area to be recovered in the stripe is a data area or a check area, and the other area to be recovered is a second hot spare area, and the target areas are determined to be other data areas and / or check areas in the area group where the data area or check area is located and the first hot spare area.

7. The method according to claim 1, characterized in that The determining of the target area based on the at most three areas to be restored includes: Failures of three storage devices are detected, and it is determined that the corresponding three areas to be recovered in the stripe are all data areas and / or verification areas, and the target area is determined to be all other data areas and / or verification areas except the three areas to be recovered and the second hot spare area.

8. The method according to claim 1, characterized in that The determining of the target area based on the at most three areas to be restored includes: Failures of three storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, the other area to be recovered is the first hot spare area, and the target areas are determined to be all other data areas and / or check areas.

9. The method according to claim 1, characterized in that The determining of the target area based on the at most three areas to be restored includes: Failures of three storage devices are detected, and it is determined that the corresponding two areas to be recovered in the stripe are both data areas and / or check areas, the other area to be recovered is the second hot spare area, and the target areas are determined to be all other data areas and / or check areas and the first hot spare area.

10. The method according to claim 1, characterized in that The determining of the target area based on the at most three areas to be restored includes: Failures of three storage devices are detected, and a corresponding area to be recovered in the stripe is determined to be a data area and / or a check area, and the other two areas to be recovered are respectively a first hot standby area and a second hot standby area, and the target area is determined to be all other data areas and / or check areas.

11. The method according to any one of claims 1 to 7, characterized in that: The stripe further includes a third hot standby area, and the method further includes: The restored data information and / or first verification information is stored in the first hot standby area, the second hot standby area and / or the third hot standby area.

12. The method according to claim 11, characterized in that The method further comprises: Upon detecting that the failure of the storage device is resolved, reading all data information and / or first verification information of any one of the area groups; and determining second verification information in the first hot standby area based on the data information and / or first verification information; And / or, reading all data information and / or first verification information; and determining third verification information in the second hot standby area based on the data information and / or first verification information.

13. A RAID data processing device, characterized in that: The device comprises: A processing module is configured to detect failures in at most three storage devices and determine at most three corresponding areas to be recovered in the stripe, wherein the areas to be recovered are a data area, a check area, a first hot spare area, and / or a second hot spare area; a calculation module, configured to determine a target area based on the at most three areas to be recovered, wherein the target area is a data area, a check area, a first hot standby area, and / or a second hot standby area excluding the area to be recovered; The calculation module is further configured to read target information in the target area, and restore the data information and / or first verification information of the area to be restored based on the target information.

14. The device according to claim 13, characterized in that include: The calculation module is further configured to detect a failure of a storage device, determine that the corresponding area to be recovered in the stripe is a data area or a check area, and determine that the target area is other data areas and / or check areas in the area group where the area to be recovered is located and the first hot spare area.

15. The device according to claim 13, characterized in that include: The calculation module is further configured to detect failures in two storage devices, determine that the two corresponding areas to be recovered in the stripe are both data areas and / or check areas, and determine that the target area is all other data areas and / or check areas.

16. The device according to claim 13, characterized in that include: The computing module is further configured to detect failures in three storage devices, determine that the corresponding three areas to be recovered in the stripe are all data areas and / or verification areas, and determine that the target area is all other data areas and / or verification areas other than the three areas to be recovered and the second hot standby area.

17. A chip, characterized in that: The chip includes a processor, and the processor is capable of executing the RAID data processing method according to any one of claims 1 to 12.

18. An electronic device, characterized in that: The electronic device includes a chip, the chip includes a processor, and the processor can execute the RAID data processing method according to any one of claims 1 to 12.

19. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and the computer program is used to execute the RAID data processing method according to any one of claims 1 to 12.

20. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the RAID data processing method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Method and device for writing into disk array and computer readable storage medium

    CN110413205A

  • Bad block data recovery method and device, storage medium and electronic equipment

    CN111930552A

  • RAID (redundant array of independent disks) management method and device, RAID card and storage medium

    CN116149574A

  • Data storage method and device, medium and product

    CN118779146A

  • Data management method and device, computer equipment and storage medium

    CN119847425A

Cited By

  • RAID data processing method and device, chip, electronic equipment, storage medium and computer program product

    CN120723544A

  • Raid data processing method and device, chip, electronic equipment, storage medium and computer program product

    CN120723544B