A method and system for data recovery

By partitioning the disk data blocks hot and cold and using the exclusive OR operation of local and global verification codes, the data recovery process of RS erasure codes is optimized, the problem of slow recovery speed in the existing technology is solved, and more efficient data recovery is achieved.

CN115269258BActive Publication Date: 2025-07-11SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210889212.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-07-11
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

The existing RS erasure code has problems such as many redundant calculations and slow recovery speed during data recovery, especially in the case of large data blocks.

Method used

The data recovery method of hot and cold partitions is adopted to divide the disk data blocks into hot zones and cold zones. Local verification codes and global verification codes are used to generate local and global verification codes respectively. Extraor and inverse matrix operations are used for different block types to optimize the data recovery process.

Benefits of technology

Improve the efficiency and speed of data recovery, especially when the data block errors in hot zone data block errors, reduce redundant data reading, improve the recovery speed of a single error, and maintain the recovery ability under a large number of errors when the data block errors in cold zone data block errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115269258B_ABST
    Figure CN115269258B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, storage medium and device for data recovery. The method includes: performing hot and cold partitioning on disk data blocks to divide the disk data blocks into hot zone data blocks and cold zone data blocks; generating local check codes for the hot zone data blocks; generating global check codes for the cold zone data blocks. When an error occurs in one of the data in the hot zone data blocks, perform exclusive OR on the local check code and the data of other data blocks in the hot zone data blocks to recover the data of one of the hot zone data blocks. When an error occurs in one of the data in the cold zone data blocks, use the global check code and the data of other data blocks in the disk data blocks to recover the data of one of the cold zone data blocks. When errors occur in more than one of the disk data blocks, use the global check code and the data of other data blocks in the disk data blocks to recover more than one of the disk data blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage and recovery, and particularly relates to a method, a system, a storage medium and a device for data recovery. Background Art

[0002] In the face of the storage requirements of massive data, in order to improve the data reliability of the storage system and ensure that the data collection node can reconstruct the original file with a high probability, it is necessary to additionally store a certain amount of redundancy on the basis of storing the original data, so that in the case of partial node failures, the system can still run normally and the data collection node can still decode and recover the original file. At the same time, in order to maintain the reliability of the system, it is necessary to repair the failed nodes in a timely manner. Therefore, it is very important to design a good node repair mechanism.

[0003] Erasure Code belongs to a forward error correction technology in coding theory and was first applied in the communication field to solve problems such as loss and corruption in data transmission. Since the erasure code technology has achieved good results in preventing data loss, it has been introduced into the storage field. Erasure code can effectively reduce the storage overhead on the premise of ensuring the same reliability. Therefore, the erasure code technology is widely applied in major storage systems and data centers such as Microsoft's Azure, Facebook's F4, etc.

[0004] It can be known that the core concept of erasure code is to construct an invertible encoding matrix to generate check data, and its inverse matrix can be calculated to recover the original data. The common RS erasure code uses the Cauchy matrix or Vandermonde matrix introduced above. The advantage of this is that the obtained matrix is definitely invertible, any of its submatrices is also invertible, and the expansion of the matrix size is simple.

[0005] The calculation of the inverse matrix of the common RS erasure code adopts the Gaussian elimination method. This general solution is applicable to the inversion of any invertible matrix, but it is not optimized for the characteristics of matrix encoding. Therefore, although the calculation is regular, it will introduce a large amount of redundant operations. When storing k data blocks and adding r check data blocks, the probability of a single data block error to be recovered accounts for 99.75% (statistics of the 2007 Storage Technology Conference). And using Gaussian elimination requires (k + r)3 operations to obtain the required inverse matrix and then recover the corresponding data block.

[0006] As an RS erasure code of the MDS code type, with the standardization of its format and the generality of encoding and decoding, it has become the most widely used technology in the current erasure module. However, with the increase of data blocks, the RS-based erasure must read a large number of data blocks each time for decoding, resulting in a very slow decoding speed.

[0007] Therefore, in view of the problem, a better data recovery mode needs to be proposed to improve the efficiency and speed of data recovery. Summary of the Invention

[0008] In view of this, the object of the present invention is to provide an improved data recovery method, system, storage medium and device to improve the efficiency and speed of data recovery.

[0009] Based on the above object, on the one hand, the present invention provides a data recovery method, which includes the following steps:

[0010] Perform hot and cold partitioning on the disk data blocks to divide the disk data blocks into hot zone data blocks and cold zone data blocks;

[0011] For the hot zone data blocks, generate local check codes using Equation (1):

[0012] LP h = f lp (D h ) (1)

[0013] where D h is the hot zone data block, f lp is the exclusive OR operation on the data block, and LP h is the local check code;

[0014] For the cold zone data blocks, generate global check codes using the Vandermonde matrix of Equation (2) or the Cauchy matrix of Equation (3):

[0015]

[0016]

[0017] where k is the number of data blocks, r is the number of global check codes, D1~D k are all the disk data blocks including the cold zone data blocks and the hot zone data blocks, and P1~P r are the r global check codes;

[0018] When an error occurs in one of the data in the hot zone data blocks, use the local check code and the data of other data blocks in the hot zone data blocks to perform exclusive OR to recover the data in one of the hot zone data blocks;

[0019] When an error occurs in one of the data in the cold zone data blocks, use the global check code and the data of other data blocks in the disk data blocks to recover the data in one of the cold zone data blocks; and

[0020] When more than one data in the disk data block has an error, use the global check code and the data of other data blocks in the disk data block to recover more than one data in the disk data block.

[0021] In some embodiments of the data recovery method according to the present invention, a new disk is added to store the local check code generated for the hot zone data block.

[0022] In some embodiments of the data recovery method according to the present invention, the disk data blocks are partitioned into hot and cold zones according to the working type and / or read-write frequency of the disk data.

[0023] Among them, when partitioning the disk data blocks according to the read-write frequency of the disk data, the cold zone data blocks are data blocks with relatively low read-write frequencies, and the hot zone data blocks are data blocks with relatively high read-write frequencies.

[0024] In some embodiments of the data recovery method according to the present invention, the method for using the global check code and the data of other data blocks in the disk data block to recover one data in the cold zone data block includes:

[0025] When r + 1 ≥ h, perform recovery by performing an exclusive OR operation on all the global check codes, the local check codes, and other cold zone data blocks in the disk data block, where r is the number of global check codes and h is the number of hot zone data blocks; and

[0026] When r + 1 < h, use the global check code and the data of other data blocks in the disk data block to perform recovery.

[0027] In some embodiments of the data recovery method according to the present invention, the method further includes:

[0028] Group the disk data blocks of every two stripes into a group, and perform an exclusive OR operation on the local check code of one stripe disk data block and the second global check code of the other stripe disk data block.

[0029] When one data block in the hot zone data block has an error, perform recovery according to the following steps:

[0030] Take k - 1 data blocks and the first global check code on one stripe to perform error recovery on the data blocks on this one stripe;

[0031] Take the hot zone data block with no error in the data of the other stripe, the second global check code of this one stripe, and the data information that has been read on this one stripe, and perform an inverse operation based on f lp to recover the data.

[0032] On the other hand, the present invention also provides a data recovery system, which includes:

[0033] A partitioning module configured to perform hot and cold partitioning on disk data blocks to divide the disk data blocks into hot zone data blocks and cold zone data blocks;

[0034] A local checksum generation module configured to generate a local checksum for the hot zone data blocks using Equation (1):

[0035] LP h = f lp (D h ) (1)

[0036] where D h is the hot zone data block, f lp is the exclusive OR operation on the data block, and LP h is the local checksum;

[0037] A global checksum generation module configured to generate a global checksum for the cold zone data blocks using the Vandermonde matrix of Equation (2) or the Cauchy matrix of Equation (3):

[0038]

[0039]

[0040] where k is the number of data blocks, r is the number of global checksums, D1 to D k are all the disk data blocks including the cold zone data blocks and the hot zone data blocks, and P1 to P r are the r global checksums; and

[0041] A data recovery module configured to recover the data of the disk data block where data has an error,

[0042] when one of the data in the hot zone data blocks has an error, the data recovery module performs an exclusive OR operation on the local checksum and the data of other data blocks in the hot zone data blocks to recover the data of one of the hot zone data blocks;

[0043] when one of the data in the cold zone data blocks has an error, the data recovery module uses the global checksum and the data of other data blocks in the disk data blocks to recover the data of one of the cold zone data blocks; and

[0044] When more than one piece of data in the disk data block has an error, the data recovery module uses the global check code and the data of other data blocks in the disk data block to recover more than one piece of data in the disk data block.

[0045] In some embodiments of the data recovery system according to the present invention, the system further includes

[0046] a dedicated disk module configured to store the local check code generated for the hot zone data block.

[0047] In some embodiments of the data recovery system according to the present invention, the partitioning module performs cold and hot partitioning on the disk data blocks according to the working type and / or read / write frequency of the disk data,

[0048] wherein, when the partitioning module performs cold and hot partitioning on the disk data blocks according to the read / write frequency of the disk data, the cold zone data blocks are data blocks with relatively low read / write frequency, and the hot zone data blocks are data blocks with relatively high read / write frequency.

[0049] In some embodiments of the data recovery system according to the present invention, the method by which the data recovery module uses the global check code and the data of other data blocks in the disk data block to recover one piece of data in the cold zone data block includes:

[0050] When r + 1 ≥ h, the data recovery module performs recovery by performing an exclusive OR operation on all the global check codes, the local check codes, and other cold zone data blocks in the disk data block, where r is the number of global check codes and h is the number of hot zone data blocks; and

[0051] When r + 1 < h, the data recovery module uses the global check code and the data of other data blocks in the disk data block to perform recovery.

[0052] In some embodiments of the data recovery system according to the present invention, the system further includes:

[0053] a grouping module configured to group the disk data blocks of every two stripes into a group, and perform an exclusive OR operation on the local check code of one stripe disk data block and the second global check code of the other stripe disk data block,

[0054] When one data block in the hot zone data block has an error, the data recovery module performs recovery according to the following steps:

[0055] Take k - 1 data blocks and the first global check code of one stripe to perform error recovery on the data blocks of this one stripe;

[0056] Take the hot zone data blocks where no error occurs in the data of another stripe, the second global check code of the one stripe, and the data information that has been read on the one stripe, and perform an inverse operation based on f lp to recover the data.

[0057] In another aspect of the present invention, there is also provided a computer-readable storage medium storing computer program instructions, and when the computer program instructions are executed, the data recovery method according to any one of the above-mentioned aspects of the present invention is implemented.

[0058] In still another aspect of the present invention, there is also provided a computer device including a memory and a processor, where a computer program is stored in the memory, and when the computer program is executed by the processor, the data recovery method according to any one of the above-mentioned aspects of the present invention is executed.

[0059] The present invention has at least the following beneficial technical effects: The present invention proposes an improved solution for hot and cold data repair, and also considers the switching problem between hot data and cold data. When the repair speed requirement for hot data is extremely high, it can be recovered in a way that uses less additional redundant data to ensure the recovery speed; when the hot data gradually becomes cold, it switches to the redundancy ratio of normal RS, which improves the data recovery speed for single errors compared to traditional RS and ensures the recovery ability under a large number of errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other embodiments according to these drawings without creative efforts.

[0061] In the figure:

[0062] Figure 1 It shows a schematic block diagram of an embodiment of the data recovery method according to the present invention;

[0063] Figure 2 It shows an example of the hot and cold zone coding situation in the data recovery method according to the present invention;

[0064] Figure 3 It shows an example of the secondary recovery state in the data recovery method according to the present invention;

[0065] Figure 4 It shows a schematic block diagram of an embodiment of the data recovery system according to the present invention;

[0066] Figure 5Schematic diagram of the hardware structure of an embodiment of a computer device implementing a data recovery method according to the present invention;

[0067] Figure 6 Schematic diagram of the framework of an embodiment of a chip according to the present invention. Detailed implementation manners

[0068] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further describes the embodiments of the present invention in detail with reference to specific embodiments and the accompanying drawings.

[0069] It should be noted that all the expressions using "first" and "second" in the embodiments of the present invention are used to distinguish two non-identical entities or non-identical parameters with the same name. It can be seen that "first" and "second" are only for the convenience of expression and should not be construed as a limitation on the embodiments of the present invention. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units inherently includes other steps or units.

[0070] According to a first aspect of the present invention, a data recovery method 100 is provided. Figure 1 Schematic block diagram showing an embodiment of a data recovery method according to the present invention. In the embodiment as Figure 1 shown, the method includes:

[0071] Step S10: Perform hot and cold partitioning on disk data blocks to divide the disk data blocks into hot zone data blocks and cold zone data blocks;

[0072] Step S20: For the hot zone data blocks, generate local check codes using Equation (1):

[0073] LP h = f lp (D h ) (1)

[0074] where D h is the hot zone data block, f lp is the exclusive OR operation on the data block, and LP h is the local check code;

[0075] Step S30: For the cold zone data blocks, generate global check codes using the Vandermonde matrix of Equation (2) or the Cauchy matrix of Equation (3):

[0076]

[0077]

[0078] Among them, k is the number of data blocks, r is the number of global parity codes, D1 to D k are all disk data blocks including cold zone data blocks and hot zone data blocks, and P1 to P r are r global parity codes;

[0079] Step S40: When an error occurs in one of the data in the hot zone data blocks, perform exclusive OR on the local parity code and the data of the other data blocks in the hot zone data blocks to recover the data of one of the hot zone data blocks;

[0080] Step S50: When an error occurs in one of the data in the cold zone data blocks, use the global parity code and the data of the other data blocks in the disk data blocks to recover the data of one of the cold zone data blocks; and

[0081] Step S60: When more than one of the data in the disk data blocks has an error, use the global parity code and the data of the other data blocks in the disk data blocks to recover more than one of the data in the disk data blocks.

[0082] There are many types of erasure codes. In actual storage systems, the more common one is the RS code (Reed-Solomon Code) applied in a distributed environment. The RS code is related to two parameters k and r. Given two positive integers k and r, for the RS code, k is the number of data blocks and r is the number of global parity codes. And the way of encoding the r parity blocks based on the Vandermonde matrix or Cauchy matrix is called the RS erasure code encoded by the Vandermonde matrix or Cauchy matrix. The specific encoding processes are shown in the above formulas (2) and (3) respectively.

[0083] The upper k*k matrix corresponds to k original data blocks, and the lower r*k matrix corresponds to the encoding matrix. By multiplying with the original data D1 to D k , the newly added P1 to P r are the r parity data obtained by encoding. When any at most r of the data are in error or lost during transmission and need to be corrected, multiply the inverse matrix of the matrix corresponding to the remaining data with the data, and the original data blocks D1 to D k will be obtained (the derivation process will not be elaborated here), which will also be briefly referred to as "RS erasure" later.

[0084] Taking the decoding of the loss of D1 to D r data as an example, the process is as shown in the following formula (4):

[0085]

[0086] As introduced in the background, assuming there are k data, the encoding of the RS erasure that generates r parities can be summarized as:

[0087]

[0088] Here, f is the encoding method used by different RSs. As described above, if the Vandermonde algorithm is used at this time, then Correspondingly, if it is Cauchy, there are different fs as described above. Similarly, assuming it is a decoding scenario, the matrix composed of the corresponding fs is inverted. Assuming one data needs to be recovered for error, the corresponding relationship is: f1 -1 (D1, D2,..., D k-1 , P1) = D k For other error scenarios, it is similar.

[0089] As can be seen above, when using RS erasure decoding, k data blocks need to be read for any error. Multiple errors can be decoded in parallel using multiple decoding modules, but still k data blocks need to be read. Limited by the read and write speed of the current storage media (any one of HDD, SSD, etc. of the same type), when this speed k is large, the recovery speed will be extremely slow.

[0090] According to the above embodiments, the disk data blocks are partitioned into hot and cold areas to divide the disk data blocks into hot area data blocks and cold area data blocks. Let the hot area be D h , and the relative cold area be D c .

[0091] In Figure 2 the example shown, the disk data blocks are D1 to D6. Assuming that the hot area data blocks are D2 and D4 at this time, and the cold area data blocks are the remaining disks, then there are:

[0092]

[0093] According to the above embodiments, for the hot area data blocks, local check codes are generated using Equation (1). For example, when r = 2, then f lp The formula of is:

[0094]

[0095] Continuing with the above situation as an example, using f lp to encode D h , the obtained check code is denoted as LP h , and we can get:

[0096]

[0097] As shown above, when using f lp to encode D h , because f lpThe generated encoded information contains all data blocks. Since there are differences in the data blocks after cold and hot partitioning, when encoding, it can be considered that the data blocks in the cold area correspond to 0 in D h and in formula (8), it is finally expressed as only for D h for the operation, and then XOR to obtain the final output LP h .

[0098] As Figure 2 shown, D2 and D4 belong to the hot area data blocks. Therefore, when errors occur in this part of the data, there is a higher requirement for the repair speed. For any error, only LP h and the remaining hot area data blocks need to be XORed. Taking the above as an example, if D2 or D4 has an error, only D4 or D2 and LP h are needed. Compared with the erasure correction method using global RS for operation recovery, in the example of Figure 2 , 4 fewer data blocks need to be read.

[0099] Similarly, when any data in the cold area data block has an error, for recovery, only all the global checksums P1, P2, LP h and the error-free data in the remaining cold area data blocks are used for recovery.

[0100] When more than one cold area data block and / or hot area data block has an error, global checksums need to be used for recovery. The specific recovery method is as described in the RS erasure correction above.

[0101] In a preferred embodiment of the data recovery method according to the present invention, a new disk is added to store the local checksum generated for the hot area data blocks.

[0102] In a preferred embodiment of the data recovery method according to the present invention, the disk data blocks are partitioned into cold and hot areas according to the working type and / or read-write frequency of the disk data. When partitioning the disk data blocks according to the read-write frequency of the disk data, the cold area data blocks are the data blocks with relatively low read-write frequencies, while the hot area data blocks are the data blocks with relatively high read-write frequencies.

[0103] In an actual storage scenario, there is a distinction between hot and cold data. For example, files such as text usually belong to cold data, that is, their read and write frequencies are not very high; while corresponding media files usually belong to hot data, that is, their read and write frequencies are high. The distinction between hot and cold data is based on the relativity of the read and write frequencies of data in the storage array. Correspondingly, hot data has a higher probability of error because of its higher read and write probability, and because of its high read and write frequency, when any error occurs, the requirement for its recovery speed is also higher, and the opposite is true for cold data. Therefore, based on the above, even the hot data in a storage array will become cold data after a period of time due to different working conditions, and correspondingly, cold data will also become hot data.

[0104] In a preferred embodiment of the data recovery method according to the present invention, the method for recovering one of the data blocks in the cold area by using the global check code and the data of other data blocks in the disk data block includes: when r + 1 ≥ h, perform recovery by performing exclusive OR on all global check codes, local check codes, and other cold area data blocks in the disk data block, where r is the number of global check codes and h is the number of hot area data blocks; and when r + 1 < h, use the global check code and the data of other data blocks in the disk data block to perform recovery.

[0105] Continuing with the above example, when D1 has an error, there are two recovery methods as follows:

[0106]

[0107] The choice of the specific recovery method depends on:

[0108] r + 1 ≤ h (10)

[0109] Where r is the number of global check codes and h is the number of hot area data blocks divided. When the formula (10) is satisfied, select 1 in the formula (9) for cold area data error recovery, and vice versa, select 2 for recovery.

[0110] In a preferred embodiment of the data recovery method according to the present invention, the method further includes: grouping the disk data blocks of every two stripes into a group, and performing exclusive OR on the local check code of one stripe disk data block and the second global check code of the other stripe disk data block. When an error occurs in the data of one data block in the hot area data block, perform recovery according to the following steps: perform error recovery on the data blocks on one stripe by taking k - 1 data blocks and the first global check code on one stripe; and take the hot area data blocks with no error in the data on the other stripe, the second global check code of one stripe, and the data information that has been read on one stripe, and perform an inverse operation based on f lp to recover the data. The above steps will be abbreviated as "secondary recovery" hereinafter.

[0111] In a preferred embodiment of the data recovery method 100 described above, in order to achieve a relatively high data recovery speed for the hot zone, an additional disk can be provided.

[0112] However, when the hot and cold zone data blocks are partitioned and the data recovery speed of the hot zone data blocks is required to be higher than that of RS erasure correction but a redundant disk cannot be provided; or when the requirement for the data recovery speed of the hot zone data blocks decreases (the required read / write frequency decreases) after working in the above state for a period of time, the secondary recovery state according to the present invention can be entered. When entering the secondary recovery state, the redundancy ratio is the same as that of RS erasure correction, but the data recovery speed for the hot zone data blocks is still higher than that of RS erasure correction. The specific changes are as follows.

[0113] As Figure 3 shown, continuing with the previous example, when entering the secondary recovery state, an operation of fusing every two stripes is performed. The specific way of the operation is to perform an exclusive OR operation on the odd (or even) LP h directly with the even (or odd) P2. When the number of global parity codes is greater than 2 (r>2), LP h can be exclusive ORed with any global parity code P. However, based on the simplicity of the operation, it is still recommended to perform exclusive OR on P2~Pr.

[0114] Suppose any one of the hot zone data blocks (such as D2 and D4 described above) has a data error at this time. The recovery method is as follows:

[0115] For the even (or odd) stripe, K-1 data blocks and P1 are taken for error recovery of the data blocks on the even (or odd) stripe. The specific recovery method is to use the standard process of RS erasure correction;

[0116] Then, the hot zone data blocks with no data errors on the odd (or even) stripe, the P2 on the even (or odd) stripe, and the data information that has been read on the even (or odd) stripe are taken to perform an inverse operation recovery based on f lp .

[0117] At this time, in the case where any one of the hot zone data blocks has a data error, the data recovery read volume required for every two stripes is: k+h.

[0118] Compared with the data recovery of RS erasure correction that reads 2k data volume, a lot of data reading requirements can be omitted, and there is a speed advantage.

[0119] In addition, when more than one data in the disk data blocks has an error, the data recovery of RS erasure correction is entered.

[0120] The present invention proposes a solution for accelerating the recovery (degraded read) of hot zone data in the case of differentiating hot and cold data blocks based on specific usage scenarios of user data in a storage array. There are two implementation methods for the solution. One is the full-speed form, which can quickly recover data errors (degraded read) in any of the hot zone data blocks, and also has a certain recovery effect on data errors in any of the cold zone data blocks, which can reduce the increase in redundancy. When necessary, it can also be converted to the sub-high-speed scenario (sub-recovery). In the sub-high-speed scheme, compared with RS, no additional redundant data storage is required, but it can also have a certain effect of accelerating the degraded read function for hot zone data blocks.

[0121] In a second aspect of the present invention, a data recovery system 200 is also provided. Figure 4 FIG. shows a schematic block diagram of an embodiment of a data recovery system 200 according to the present invention. As Figure 4 shown, the system includes:

[0122] A partitioning module 210 configured to partition disk data blocks into hot and cold zones to divide the disk data blocks into hot zone data blocks and cold zone data blocks;

[0123] A local parity code generation module 220 configured to generate a local parity code for the hot zone data blocks using Equation (1):

[0124] LP h = f lp (D h ) (1)

[0125] where D h is the hot zone data block, f lp is the exclusive OR operation on the data block, and LP h is the local parity code;

[0126] A global parity code generation module 230 configured to generate a global parity code for the cold zone data blocks using the Vandermonde matrix of Equation (2) or the Cauchy matrix of Equation (3):

[0127]

[0128]

[0129] where k is the number of data blocks, r is the number of global parity codes, D1 to D k are all the disk data blocks including cold zone data blocks and hot zone data blocks, and P1 to P r are r global parity codes;

[0130] A data recovery module 240, which is configured to recover the data of the disk data block where data errors occur.

[0131] When an error occurs in the data of one of the hot zone data blocks, the data recovery module 240 performs an exclusive OR operation on the local check code and the data of other data blocks in the hot zone data block to recover the data of one of the hot zone data blocks; when an error occurs in the data of one of the cold zone data blocks, the data recovery module 240 uses the global check code and the data of other data blocks in the disk data block to recover the data of one of the cold zone data blocks; and when errors occur in more than one of the disk data blocks, the data recovery module 240 uses the global check code and the data of other data blocks in the disk data block to recover more than one of the disk data blocks.

[0132] In a preferred embodiment of the data recovery system according to the present invention, the system further includes a dedicated disk module, which is configured to store the local check codes generated for the hot zone data blocks.

[0133] In a preferred embodiment of the data recovery system according to the present invention, the partitioning module 210 partitions the disk data blocks into hot and cold zones according to the working type and / or read / write frequency of the disk data. When the partitioning module 210 partitions the disk data blocks into hot and cold zones according to the read / write frequency of the disk data, the cold zone data blocks are the data blocks with relatively lower read / write frequencies, while the hot zone data blocks are the data blocks with relatively higher read / write frequencies.

[0134] In a preferred embodiment of the data recovery system according to the present invention, the method by which the data recovery module 240 uses the global check code and the data of other data blocks in the disk data block to recover the data of one of the cold zone data blocks includes:

[0135] When r + 1 ≥ h, the data recovery module 240 performs recovery by performing an exclusive OR operation on all the global check codes, local check codes, and other cold zone data blocks in the disk data block, where r is the number of global check codes and h is the number of hot zone data blocks; and

[0136] When r + 1 < h, the data recovery module 240 uses the global check code and the data of other data blocks in the disk data block to perform recovery.

[0137] In a preferred embodiment of the data recovery system according to the present invention, the system further includes:

[0138] A grouping module, which is configured to group the disk data blocks of every two stripes, and perform an exclusive OR operation on the local check code of one stripe of disk data blocks and the second global check code of the other stripe of disk data blocks.

[0139] When an error occurs in the data of a data block in the hot zone data block, the data recovery module 240 performs recovery according to the following steps: taking k-1 data blocks and the first global checksum of a stripe to perform error recovery on the data blocks on the stripe; and taking the hot zone data blocks with error-free data on another stripe, the second global checksum of a stripe, and the data information that has been read on a stripe, and performing an inverse operation based on f lp to recover the data.

[0140] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is further provided. Figure 5 FIG. shows a schematic diagram of a computer-readable storage medium according to the data recovery method provided by the embodiments of the present invention. As Figure 5 shown, the computer-readable storage medium 300 stores computer program instructions 310, and the computer program instructions 310 can be executed by a processor. When the computer program instructions 310 are executed, the method of any of the above embodiments is implemented.

[0141] It should be understood that, without conflict, all the embodiments, features, and advantages described above for the data recovery method according to the present invention are equally applicable to the data recovery system and storage medium according to the present invention.

[0142] In a fourth aspect of the embodiments of the present invention, a computer device 400 is further provided, including a memory 420 and a processor 410. A computer program is stored in the memory, and when the computer program is executed by the processor, the method of any of the above embodiments is implemented.

[0143] As Figure 6 shown, it is a schematic diagram of the hardware structure of an embodiment of a computer device for executing the data recovery method provided by the present invention. Taking the computer device 400 as shown Figure 6 as an example, in the computer device, there is a processor 410 and a memory 420, and it may further include: an input device 430 and an output device 440. The processor 410, the memory 420, the input device 430, and the output device 440 can be connected through a bus or other means. Figure 6 Taking connection through a bus as an example. The input device 430 can receive input digital or character information and generate signal inputs related to data recovery. The output device 440 may include a display device such as a display screen.

[0144] The memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the resource monitoring method in the embodiments of the present application. The memory 420 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created for use in the resource monitoring method, etc. In addition, the memory 420 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 420 may optionally include a memory remotely disposed relative to the processor 410, and these remote memories can be connected to the local module through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0145] By running the non-volatile software programs, instructions, and modules stored in the memory 420, the processor 410 executes various functional applications and data processing of the server, that is, implements the resource monitoring method in the above method embodiments.

[0146] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, a general description has been given of the functions of the various illustrative components, blocks, modules, circuits, and steps. Whether this function is implemented as software or hardware depends on the specific application and the design constraints imposed on the overall system. The functions that those skilled in the art can implement in various ways for each specific application, but such implementation decisions should not be construed as causing a departure from the scope of the disclosure of the embodiments of the present invention.

[0147] Finally, it should be noted that the computer-readable storage medium (e.g., memory) of this article can be a volatile memory or a non-volatile memory, or can include both volatile memory and non-volatile memory. By way of example and not limitation, non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM), which can act as an external cache memory. By way of example and not limitation, RAM can be obtained in various forms, such as synchronous RAM (DRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The storage devices of the disclosed aspects are intended to include, but are not limited to, these and other suitable types of memory.

[0148] The various exemplary logic blocks, modules, and circuits described in connection with the disclosure herein can be implemented or executed using the following components designed to perform the functions herein: a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components. The general-purpose processor can be a microprocessor, but alternatively, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP and / or any other such configuration.

[0149] The above are the exemplary embodiments disclosed in the present invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments disclosed in the claims of the present invention. The functions, steps, and / or actions of the method claims according to the disclosed embodiments herein need not be performed in any particular order. In addition, although the elements disclosed in the embodiments of the present invention can be described or claimed in individual form, they can also be understood as plural unless explicitly limited to the singular.

[0150] It should be understood that, as used herein, unless the context clearly supports exceptions, the singular form "a" is also intended to include the plural form. It should also be understood that the "and / or" used herein refers to any and all possible combinations of one or more of the associated listed items. The serial numbers of the disclosed embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments.

[0151] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope (including the claims) disclosed by the embodiments of the present invention is limited to these examples; under the concept of the embodiments of the present invention, the technical features in the above embodiments or different embodiments can also be combined, and there are many other variations in different aspects of the embodiments of the present invention as above, which are not provided in detail for the sake of brevity. Therefore, any omissions, modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.

Claims

1. A method for data recovery, characterized in that, The method includes the following steps: Perform hot and cold partitioning on the disk data blocks to divide the disk data blocks into hot area data blocks and cold area data blocks; For the hot area data blocks, generate local check codes by using Equation (1); LP h = f lp (D h ) (1) Among them, D h is the hot zone data block, f lp is the exclusive OR operation on the data block, and LP h is the local check code; For the cold area data blocks, generate global check codes by using the Vandermonde matrix of Equation (2) or the Cauchy matrix of Equation (3); Among them, k is the number of data blocks, r is the number of global parity codes, D1 to D k are all the disk data blocks including the cold zone data blocks and the hot zone data blocks, and P1 to P r are r global parity codes; When an error occurs in the data of one of the hot area data blocks, perform exclusive OR on the local check code and the data of other data blocks in the hot area data blocks to recover the data of one of the hot area data blocks; When an error occurs in the data of one of the cold area data blocks, use the global check code and the data of other data blocks in the disk data blocks to recover the data of one of the cold area data blocks; and When errors occur in more than one of the disk data blocks, use the global check code and the data of other data blocks in the disk data blocks to recover more than one of the disk data blocks.

2. The method according to claim 1, wherein Add a new disk for storing the local check codes generated for the hot area data blocks.

3. The method according to claim 1, wherein Perform hot and cold partitioning on the disk data blocks according to the working type and / or read / write frequency of the disk data, wherein, when performing hot and cold partitioning on the disk data blocks according to the read / write frequency of the disk data, the cold area data blocks are data blocks with relatively low read / write frequencies, and the hot area data blocks are data blocks with relatively high read / write frequencies.

4. The method according to claim 1, wherein The method for using the global check code and the data of other data blocks in the disk data blocks to recover the data of one of the cold area data blocks includes: When r + 1 ≥ h, perform recovery by performing exclusive OR on all the global check codes, the local check codes, and other cold area data blocks in the disk data blocks, where r is the number of global check codes and h is the number of hot area data blocks; and When r + 1 < h, use the global check code and the data of other data blocks in the disk data blocks to perform recovery.

5. The method according to claim 1, wherein The method further includes: Group the disk data blocks of every two stripes into a group, and perform exclusive OR on the local check code of one stripe disk data block and the second global check code of the other stripe disk data block, When an error occurs in the data of one data block in the hot area data blocks, perform recovery according to the following steps: Take k - 1 data blocks and the first global check code on one stripe to perform error recovery on the data blocks on the one stripe; Take the hot zone data block for which there is no error in the data on another stripe, the second global checksum of the one stripe, and the data information that has been read on the one stripe, and perform an inverse operation based on f lp to recover the data.

6. A data recovery system, characterized in that, The method includes: A partitioning module configured to perform hot and cold partitioning on the disk data blocks to divide the disk data blocks into hot area data blocks and cold area data blocks; A local check code generation module configured to generate local check codes for the hot area data blocks by using Equation (1); LP h = f lp (D h ) (1) Among them, D h is the hot zone data block, f lp is the exclusive OR operation on the data block, and LP h is the local check code; A global check code generation module configured to generate global check codes for the cold area data blocks by using the Vandermonde matrix of Equation (2) or the Cauchy matrix of Equation (3); where k is the number of data blocks, r is the number of global parity codes, D1 to D k are all the disk data blocks including the cold region data blocks and the hot region data blocks, and P1 to P r are the r global parity codes; and A data recovery module configured to recover the data of the disk data blocks in which errors occur, When an error occurs in the data of one of the hot - zone data blocks, the data recovery module uses the local checksum and the data of other data blocks in the hot - zone data blocks to perform exclusive OR to recover the data of one of the hot - zone data blocks; When an error occurs in the data of one of the cold - zone data blocks, the data recovery module uses the global checksum and the data of other data blocks in the disk data blocks to recover the data of one of the cold - zone data blocks; and When errors occur in more than one of the disk data blocks, the data recovery module uses the global checksum and the data of other data blocks in the disk data blocks to recover more than one of the disk data blocks.

7. The system according to claim 6, wherein Further comprising: A dedicated disk module configured to store the local checksum generated for the hot - zone data blocks.

8. The system according to claim 6, wherein, The partitioning module partitions the disk data blocks into hot and cold zones according to the working type and / or read - write frequency of the disk data, wherein, when the partitioning module partitions the disk data blocks into hot and cold zones according to the read - write frequency of the disk data, the cold - zone data blocks are data blocks with relatively low read - write frequencies, and the hot - zone data blocks are data blocks with relatively high read - write frequencies.

9. The system according to claim 6, wherein, The method by which the data recovery module uses the global checksum and the data of other data blocks in the disk data blocks to recover the data of one of the cold - zone data blocks includes: When r + 1≥h, the data recovery module performs recovery by performing exclusive OR using all of the global checksums, the local checksums, and other cold - zone data blocks in the disk data blocks, where r is the number of global checksums and h is the number of hot - zone data blocks; and When r + 1<h, the data recovery module uses the global checksum and the data of other data blocks in the disk data blocks to perform recovery.

10. The system according to claim 6, characterized in that, Further comprising: A grouping module configured to group the disk data blocks of every two stripes into a group, and perform exclusive OR on the local checksum of one stripe of disk data blocks and the second global checksum of the other stripe of disk data blocks, When an error occurs in the data of one data block in the hot - zone data blocks, the data recovery module performs recovery according to the following steps: For one stripe, take k - 1 data blocks and the first global checksum to perform error recovery on the data blocks on that one stripe; Take the hot zone data block where no error occurs in the data of another stripe, the second global check code of the one stripe, and the data information that has been read on the one stripe, and perform an inverse operation based on f lp to recover the data.

Citation Information

Patent Citations

  • Adaptive local reconstruction code design method for hot data storage and cloud storage system

    CN112000278A