Data processing method, device, electronic device and storage medium

By dividing the original data shard into two parts and generating multiple verification shards, the problem of low data recovery efficiency under the existing erasure encoding method is solved, and more efficient data recovery is achieved.

CN115686937BActive Publication Date: 2025-08-29CHONGQING UNISINSIGHT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211295002.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2025-08-29
Estimated Expiration
2042-10-21

Smart Images

  • Figure CN115686937B_ABST
    Figure CN115686937B_ABST
Patent Text Reader

Abstract

The present invention provides a data processing method, device, electronic device, and storage medium, relating to the field of data storage. The method divides each original data slice into a first original slice and a second original slice, obtains a first check slice that satisfies a preset erasure ratio based on all the first original slices and the second original slice serving as a second target slice, and obtains a first check slice that satisfies the preset erasure ratio based on all the second original slices and the first original slice serving as the first target slice, thereby obtaining multiple check data slices that satisfy the preset erasure ratio. When recovering a damaged original data slice, only the first original slice or the second original slice of the undamaged original data slice and the first check slice or the second check slice of the undamaged check data slice need to be read, thereby reducing the amount of data required for the recovery process, thereby reducing the time spent on reading data and improving data recovery efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data storage, and in particular to a data processing method, device, electronic device and storage medium. Background Art

[0002] Erasure coding technology is widely used in distributed storage. For example, assuming the erasure ratio is K+P, when storing data, the data to be stored is divided into K copies of original data. By encoding the K copies of original data, P copies of verification data are obtained, and then the K copies of original data and P copies of verification data are written to different hard disks respectively to store data.

[0003] Since each piece of verification data is obtained by encoding K pieces of original data, when the data in any hard disk is damaged, if the damaged data is original data, it is necessary to read all the data from at least K other hard disks and then decode the damaged data. If the damaged data is verification data, it is also necessary to read all the original data from K hard disks and re-encode the damaged data.

[0004] Due to the existing erasure coding method, when recovering lost data, it takes a lot of time to read all the data in at least K hard disks, which seriously affects the low efficiency of data recovery. Summary of the Invention

[0005] In order to overcome the deficiencies of the prior art, the present invention provides a data processing solution, device, electronic device and storage medium.

[0006] The technical solution of the embodiment of the present invention can be implemented as follows:

[0007] In a first aspect, an embodiment of the present invention provides a data processing method applied to a storage device, wherein the storage device includes multiple hard disks, and the method includes:

[0008] Acquire a plurality of ordered original data shards, wherein the plurality of original data shards are obtained by dividing the data to be stored according to a preset erasure ratio, and each of the original data shards is divided into a first original shard and a second original shard;

[0009] Determine a first target shard and a second target shard according to the sequence number of each of the original data shards, wherein the first target shard is determined from all the first original shards, and the second target shard is determined from all the second original shards;

[0010] Obtaining a plurality of first verification fragments based on the plurality of first original fragments and the plurality of second target fragments;

[0011] Obtaining a plurality of second verification fragments based on the plurality of second original fragments and the plurality of first target fragments;

[0012] Obtaining a plurality of verification data fragments according to the plurality of first verification fragments and the plurality of second verification fragments, wherein each of the verification data fragments includes a first verification fragment and a second verification fragment, and the first verification fragment and the second verification fragment included in any two of the verification data fragments are different;

[0013] Each of the original data slices and each of the verification data slices are written into different hard disks respectively.

[0014] Optionally, the step of determining the first target shard and the second target shard according to the sequence number of each original data shard includes:

[0015] For each of the original data shards, if the sequence number of the original data shard is an odd number, the first original shard of the original data shards is used as the first target shard;

[0016] If the sequence number of the original data shard is an even number, the second original shard of the original data shard is used as the second target shard.

[0017] Optionally, the multiple first check slices include a first uncoupled check slice and at least one first coupled check slice, and the step of obtaining the multiple first check slices based on the multiple first original slices and the multiple second target slices includes:

[0018] Performing encoding processing on all the first original slices to generate a plurality of first encoded slices;

[0019] Determining a first uncoupled check slice and each first coded slice to be coupled from the plurality of first coded slices according to a generation order of each first coded slice;

[0020] Determining a coupling load of each first coded slice to be coupled according to a generation order of each first coded slice to be coupled and a sequence number of an original slice corresponding to each second target slice, wherein the coupling load of each first coded slice to be coupled includes at least one second target slice;

[0021] Each first coding slice to be coupled is coupled with a coupling load of each first coding slice to be coupled to obtain each first coupling check slice.

[0022] Optionally, the plurality of second parity slices include one second uncoupled parity slice and at least one second coupled parity slice, and the step of obtaining the plurality of second parity slices based on the plurality of second original slices and the plurality of first target slices includes:

[0023] performing encoding processing on all the second original fragments to generate a plurality of second encoded fragments;

[0024] Determining a second uncoupled check segment and each second coded segment to be coupled from the plurality of second coded segments according to a generation order of each second coded segment;

[0025] Determining a coupling load of each second coded slice to be coupled according to a generation order of each second coded slice to be coupled and a sequence number of an original slice corresponding to each first target slice, wherein the coupling load of each second coded slice to be coupled includes at least one first target slice;

[0026] Each second coding slice to be coupled is coupled with a coupling load of each second coding slice to be coupled to obtain each second coupling check slice.

[0027] Optionally, the storage device contains data to be recovered, the data to be recovered including a damaged original data slice, an undamaged original data slice, and an undamaged check data slice, and there is only one damaged check data slice. The method further includes:

[0028] If the first original shard of the damaged original data shard is the first target shard, restoring the first original shard and the second original shard of the damaged original data shard according to the second original shard of the undamaged original data shard and the second check shard of the undamaged check data shard;

[0029] If the second original slice of the damaged original data slice is the second target slice, the first original slice and the second original slice of the damaged original data slice are restored according to the first original slice of the undamaged original data slice and the first check slice of the undamaged check data slice.

[0030] Optionally, the storage device contains data to be recovered, the data to be recovered including damaged parity data shards and undamaged original data shards, the first parity shard of the damaged parity data shards is a first coupling parity shard, and the second parity shard of the damaged parity data shards is a second coupling parity shard, and the method further includes:

[0031] Determining a target first original slice and a target second original slice according to the sequence number of the damaged parity data slice, wherein the first original slice is determined from the first original slices of all the undamaged original data slices, and the target second original slice is determined from the second original slices of all the undamaged original data slices;

[0032] Obtaining a first checksum slice of the damaged checksum data slice according to the first original slices of all the undamaged original data slices and the target second original slice;

[0033] A second parity slice of the damaged parity data slice is obtained according to the second original slices of all the undamaged original data slices and the target first original slice.

[0034] Optionally, the storage device contains data to be recovered, the data to be recovered including damaged original data slices, damaged verification data slices, undamaged original data slices, and undamaged verification data slices, and the method further includes:

[0035] Obtaining a coding coupling matrix corresponding to the data to be recovered;

[0036] Deleting row elements corresponding to damaged original data slices and damaged check data slices in the coding coupling matrix to obtain a processed coding coupling matrix, and obtaining an inverse matrix of the processed coding coupling matrix;

[0037] Constructing a decoding vector using the first original slice and the second original slice of the undamaged original data slice and the first check slice and the second check slice of the undamaged check data slice;

[0038] The damaged original data slice and the damaged check data slice are restored based on the inverse matrix and the decoding vector.

[0039] In a second aspect, an embodiment of the present invention provides a data processing apparatus, applied to a storage device, wherein the storage device includes multiple hard disks, and the apparatus includes:

[0040] an acquisition module, configured to acquire a plurality of ordered original data shards, wherein the plurality of original data shards are obtained by dividing the data to be stored according to a preset erasure ratio, and each of the original data shards is divided into a first original shard and a second original shard;

[0041] Processing module for:

[0042] Determine a first target shard and a second target shard according to the sequence number of each of the original data shards, wherein the first target shard is determined from all the first original shards, and the second target shard is determined from all the second original shards;

[0043] Obtaining a plurality of first verification fragments based on the plurality of first original fragments and the plurality of second target fragments;

[0044] Obtaining a plurality of second verification fragments based on the plurality of second original fragments and the plurality of first target fragments;

[0045] Obtaining a plurality of verification data fragments according to the plurality of first verification fragments and the plurality of second verification fragments, wherein each of the verification data fragments includes a first verification fragment and a second verification fragment, and the first verification fragment and the second verification fragment included in any two of the verification data fragments are different;

[0046] A writing module is used to write each of the original data slices and each of the verification data slices into different hard disks respectively.

[0047] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the data processing method as described in the first aspect when executing the computer program.

[0048] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the data processing method as described in the first aspect.

[0049] Compared with the prior art, the embodiments of the present invention provide a data processing method, device, electronic device and storage medium. First, a plurality of ordered original data shards are obtained, wherein the plurality of original data shards are obtained by dividing the data to be stored according to a preset erasure ratio, and each original data shard is divided into a first original shard and a second original shard; then, according to the sequence number of each original data shard, a first target shard and a second target shard are determined, wherein the first target shard is determined from all the first original shards, and the second target shard is determined from all the second original shards; then Then, based on the multiple first original shards and the multiple second target shards, multiple first check shards are obtained; based on the multiple second original shards and the multiple first target shards, multiple second check shards are obtained; and then based on the multiple first check shards and the multiple second check shards, multiple check data shards are obtained, wherein each check data shard includes a first check shard and a second check shard, and the first check shard and the second check shard included in any two check data shards are different; finally, each of the original data shards and each of the check data shards are written to different hard disks respectively. In the embodiment of the present invention, by dividing each original data shard into a first original shard and a second original shard, a first check shard satisfying a preset erasure ratio is obtained based on all the first original shards and the second original shard as the second target shard, and a first check shard satisfying a preset erasure ratio is obtained based on all the second original shards and the first original shard as the first target shard, thereby obtaining a plurality of check data shards satisfying the preset erasure ratio. When recovering a damaged original data shard, only the first original shard or the second original shard of the undamaged original data shard and the first check shard or the second check shard of the undamaged check data shard need to be read, thereby reducing the amount of data required for the recovery process, thereby reducing the time spent on reading data, and improving data recovery efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0051] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present invention;

[0052] Figure 2 A schematic diagram of a data processing method provided by an embodiment of the present invention Figure 1 ;

[0053] Figure 3 An example diagram of a data processing process provided by an embodiment of the present invention;

[0054] Figure 4 A schematic diagram of a data processing method provided by an embodiment of the present invention Figure 2 ;

[0055] Figure 5 A schematic diagram of a data processing method provided by an embodiment of the present invention Figure 3 ;

[0056] Figure 6 A block diagram of functional units of a data processing device provided by an embodiment of the present invention;

[0057] Figure 7 A schematic block diagram of the structure of an electronic device provided by an embodiment of the present invention.

[0058] Icon: 100 - data processing device; 101 - acquisition module; 102 - processing module; 103 - writing module; 200 - electronic device; 210 - memory; 220 - processor. DETAILED DESCRIPTION

[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0060] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0061] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0062] In addition, the terms "first", "second", etc., if used, are merely used to distinguish and describe, and should not be understood as indicating or implying relative importance.

[0063] It should be noted that, in the absence of conflict, the features in the embodiments of the present invention may be combined with each other.

[0064] Existing distributed storage methods usually process the data to be stored according to a set correction and erasure ratio to obtain several original data shards and several verification data shards, and then store the several original data shards and several verification data shards on different hard disks.

[0065] For example, assuming that the erasure ratio is K+P, Figure 1 As shown, the data to be stored is processed according to the erasure ratio to obtain K original data fragments, namely t0, t1, t2, ..., t K-1 , and P check data shards, namely s0, s1, s2, ..., s P-1 Then, t0, t1, t2, ..., t k-1 Store them in hard disk 0, hard disk 1, hard disk 2, ..., hard disk K-1 respectively, and store s0, s1, s2, ..., s P-1 Stored respectively to hard disk K, hard disk K+1, hard disk K+2, ..., hard disk K+P-1.

[0066] because Figure 1 The P check data shards in the

[10] are obtained by encoding K original data shards. When the data in any hard disk is damaged, if the damaged data is an original data shard, it is necessary to read all the data from at least K other hard disks, and then decode the damaged original data shard. If the damaged data is a check data shard, it is also necessary to read all the original data shards from the K hard disks and re-encode them to obtain the damaged check data shard.

[0067] The larger the hard disk space, the more time it takes to read the data needed for recovery when recovering damaged data, which reduces data recovery efficiency.

[0068] In order to improve data recovery efficiency, an embodiment of the present invention provides a data processing method, which will be described in detail below.

[0069] Please refer to Figure 2 , the data processing method includes steps S101 to S106.

[0070] S101, obtaining multiple ordered original data shards.

[0071] The multiple original data shards are obtained by dividing the data to be stored according to a preset erasure ratio, and each original data shard is divided into a first original shard and a second original shard.

[0072] For example, assuming that the preset erasure ratio is 4+3, Figure 3As shown, the data to be stored is divided into four original data shards according to the preset erasure ratio, namely t0, t1, t2 and t3, wherein t0 includes the first original shard a0 and the second original shard b0, t1 includes the first original shard a1 and the second original shard b1, t2 includes the first original shard a2 and the second original shard b2, and t3 includes the first original shard a3 and the second original shard b3.

[0073] S102: Determine a first target shard and a second target shard according to the sequence number of each original data shard.

[0074] The first target shard is determined from all the first original shards, and the second target shard is determined from all the second original shards.

[0075] As a specific implementation, the implementation process of step S102 may be as follows:

[0076] For each original data shard, if the sequence number of the original data shard is an odd number, the first original shard of the original data shard is used as the first target shard.

[0077] If the sequence number of the original data shard is an even number, the second original shard of the original data shard is used as the second target shard.

[0078] For example, Figure 3 As shown, for the original data shard t0, its sequence number 0 is an even number, so the second original shard b0 of t0 is used as the second target shard, for the original data shard t1, its sequence number 1 is an odd number, so the first original shard a1 of t1 is used as the first target shard, for the original data shard t2, its sequence number 2 is an even number, so the second original shard b2 of t2 is used as the second target shard, for the original data shard t3, its sequence number 3 is an odd number, so the first original shard a3 of t3 is used as the first target shard, that is, the first target shard includes a1 and a3, and the second target shard includes b0 and b2.

[0079] S103: Obtain a plurality of first verification fragments based on the plurality of first original fragments and the plurality of second target fragments.

[0080] The multiple first check slices include a first uncoupled check slice and at least one first coupled check slice.

[0081] As a possible implementation manner, step S103 may include sub-steps S103 - 1 to S103 - 4 .

[0082] S103-1: Encode all first original fragments to generate multiple first encoded fragments.

[0083] For example, Figure 3 As shown, according to a preset erasure correction ratio, the first original fragments a0, a1, a2 and a3 are encoded to obtain first encoded fragments a4', a5' and a6'.

[0084] S103-2: Determine a first uncoupled check segment and each first coded segment to be coupled from a plurality of first coded segments according to the generation order of each first coded segment.

[0085] For example, Figure 3 As shown, the first coded segment a4' that is generated first in the order of generation can be used as the first uncoupled check segment a4, and a5' and a6' can be used as the first coded segments to be coupled.

[0086] S103 - 3 , determining a coupling load of each first coded slice to be coupled according to a generation order of each first coded slice to be coupled and a sequence number of an original slice corresponding to each second target slice.

[0087] The coupling load of each first coding slice to be coupled includes at least one second target slice.

[0088] For example, Figure 3 As shown, there are two second target slices, namely b0 and b2, and two first coding slices to be coupled, namely a5' and a6'. According to the generation order of a5' and a6' and the sequence numbers of b0 and b2, b0 is used as the coupling load of a5' and b2 is used as the coupling load of a6'.

[0089] It can be understood that the second target slice is evenly distributed to each first coding slice to be coupled as a coupling load. In the embodiment of the present invention, the difference in the number of coupling loads between any two first coding slices to be coupled does not exceed 1.

[0090] S103 - 4 , performing coupling processing on each first coding slice to be coupled and the coupling load of each first coding slice to be coupled to obtain each first coupling check slice.

[0091] The coupling process refers to the process of combining the vector composed of the first coded slice to be coupled and its coupled load with the identity element row matrix ie_matrix 1*n =[1,1,1,……,1], performing matrix operations on finite fields.

[0092] For example, Figure 3 As shown, the process of coupling a5' and b0 to obtain the first coupling check fragment a5 can be: the vector composed of a5' and b0 is (a5', b0), and the identity element row matrix is ​​ie_matrix 1*2=[1,1], and by performing matrix operations on the two in a finite field, we get a5.

[0093] Similarly, the process of coupling a6' and b2 to obtain the first coupling check fragment a6 can be: the vector formed by a6' and b2 is (a6', b2), and the identity element row matrix is ​​ie_matrix 1*2 =[1,1], and by performing matrix operations on the two in a finite field, we get a6.

[0094] S104: Obtain a plurality of second verification fragments based on the plurality of second original fragments and the plurality of first target fragments.

[0095] The plurality of second check slices include one second uncoupled check slice and at least one second coupled check slice.

[0096] As a possible implementation manner, step S104 may include sub-steps S104 - 1 to S104 - 4 .

[0097] S104-1: Encode all the second original fragments to generate multiple second encoded fragments.

[0098] For example, Figure 3 As shown, according to the preset erasure correction ratio, the second original fragments b0, b1, b2 and b3 are encoded to obtain second encoded fragments b4', b5' and b6'.

[0099] S104-2: Determine a second uncoupled check segment and each second coded segment to be coupled from a plurality of second coded segments according to the generation order of each second coded segment.

[0100] For example, Figure 3 As shown, the second coding segment b4' which is generated first can be used as the second uncoupled check segment b4, and b5' and b6' can be used as the second coding segments to be coupled.

[0101] S104-3: Determine the coupling load of each second coded slice to be coupled according to the generation order of each second coded slice to be coupled and the sequence number of the original slice corresponding to each first target slice.

[0102] The coupling load of each second coding slice to be coupled includes at least one first target slice.

[0103] For example, Figure 3As shown, there are two first target slices, namely a1 and a3, and two second coded slices to be coupled, namely b5' and b6'. According to the generation order of b5' and b6' and the sequence numbers of a1 and a3, a1 is used as the coupling load of b5' and a3 is used as the coupling load of b6'.

[0104] It can be understood that the first target slice is evenly distributed to each second coding slice to be coupled as a coupling load. In the embodiment of the present invention, the difference in the number of coupling loads between any two second coding slices to be coupled does not exceed 1.

[0105] S104-4: perform coupling processing on each second coding slice to be coupled and the coupling load of each second coding slice to be coupled to obtain each second coupling check slice.

[0106] The coupling process refers to the process of combining the vector composed of the second coded slice to be coupled and its coupled load with the identity element row matrix ie_matrix 1*n =[1,1,1,……,1], performing matrix operations on finite fields.

[0107] For example, Figure 3 As shown, the process of coupling b5' and a1 to obtain the first coupling check fragment b5 can be: the vector formed by b5' and a1 is (b5', a1), and the identity element row matrix is ​​ie_matrix 1*2 =[1,1], and by performing matrix operations on the two in a finite field, we get b5.

[0108] Similarly, the process of coupling b6' and a3 to obtain the first coupling check fragment a6 can be: the vector formed by b6' and a3 is (b6', a3), and the identity element row matrix is ​​ie_matrix 1*2 =[1,1], and by performing matrix operations on the two in a finite field, we get b6.

[0109] S105: Obtain multiple verification data fragments according to the multiple first verification fragments and the multiple second verification fragments.

[0110] Each verification data shard includes a first verification shard and a second verification shard, and the first verification shard and the second verification shard included in any two verification data shards are different.

[0111] For example, Figure 3As shown, the check data slice s0 includes a first uncoupled check slice a4 and a second uncoupled check slice b4, the check data slice s1 includes a first coupled check slice a5 and a second coupled check slice b5, and the check data slice s2 includes a first coupled check slice a6 and a second coupled check slice b6.

[0112] S106: Write each original data slice and each verification data slice into different hard disks respectively.

[0113] For example, Figure 3 As shown, t0, t1, t2 and t3 are stored in hard disk 0, hard disk 1, hard disk 2 and hard disk 3 respectively, and s0, s1 and s2 are stored in hard disk 4, hard disk 5 and hard disk 6 respectively.

[0114] For the case where an original data slice is damaged and needs to be rebuilt in the data stored in the storage device using the above method, the data processing method provided by the embodiment of the present invention further includes steps S201 to S202.

[0115] S201: If the first original shard of the damaged original data shard is the first target shard, the first original shard and the second original shard of the damaged original data shard are restored according to the second original shard of the undamaged original data shard and the second check shard of the undamaged check data shard.

[0116] The second check slice of the undamaged check data slice includes a second uncoupled check slice and a second coupled check slice obtained by using the first original slice of the damaged original data slice as a coupled load.

[0117] For example, assuming Figure 3 The original data slice t1 written to hard disk 1 is damaged, and the first original slice a1 of t1 is the first target slice. The second coupling check slice obtained when a1 is used as the coupling load is b5. Therefore, when restoring t1, it is necessary to read the second original slices b0, b2 and b3 of the undamaged original data slices t0, t2 and t3 from hard disk 0, hard disk 2 and hard disk 3, read the second non-coupling check slice b4 in the undamaged check data slice s0 from hard disk 4, and read the second coupling check slice b5 in the undamaged check data slice s1 from hard disk 5.

[0118] Then, b0, b2, b3, and b4 are used to decode b1. Next, b0, b1, b2, and b3 are encoded to obtain the second encoded fragment b5'. Finally, b5 and b5' are used to decouple a1, thus completing the recovery of t1.

[0119] S202: If the second original shard of the damaged original data shard is the second target shard, the first original shard and the second original shard of the damaged original data shard are restored according to the first original shard of the undamaged original data shard and the first check shard of the undamaged check data shard.

[0120] The first check slice of the undamaged check data slice includes a first uncoupled check slice and a first coupled check slice obtained by using the second original slice of the damaged original data slice as a coupled load.

[0121] For example, assuming Figure 3 The original data slice t0 written to hard disk 1 is damaged, and the second original slice b0 of t0 is the second target slice. The first coupling check slice obtained when b0 is used as the coupling load is a5. Therefore, when restoring t0, it is necessary to read the first original slices a1, a2 and a3 of the undamaged original data slices t1, t2 and t3 from hard disk 1, hard disk 2 and hard disk 3, read the first non-coupling check slice a4 in the undamaged check data slice s0 from hard disk 4, and read the first coupling check slice a5 in the undamaged check data slice s1 from hard disk 5.

[0122] Then, using a1, a2, a3, and a4, a0 is decoded. Next, a0, a1, a2, and a3 are encoded to obtain the first encoded slice a5'. Finally, using a5 and a5', b0 is decoupled, thus completing the recovery of t0.

[0123] When restoring t0, the prior art needs to read out all the undamaged original data slices t1, t2, and t3 and the undamaged verification data slice s0, while the data processing method provided by the embodiment of the present invention only needs to read out the first original slices a1, a2, and a3 of the undamaged original data slices t1, t2, and t3, the first uncoupled verification slice a4 in the undamaged verification data slice s0, and the first coupled verification slice a5 in the undamaged verification data slice s1. The data reading amount of the data processing method provided by the embodiment of the present invention is only 62.5% of that of the prior art. Accordingly, the IO bandwidth is reduced by 37.5%.

[0124] For any erasure ratio K+P, when there is only one damaged original data shard that needs to be rebuilt, the maximum amount of data to be read is (K+2) / 2K in the existing technology, and the IO bandwidth can be reduced by a maximum of (K-2) / 2K, with a limit of 50%.

[0125] For the case where one or more check data slices are damaged and need to be rebuilt in the data stored in the storage device using the above method, the data processing method provided by the embodiment of the present invention also includes the following steps: Figure 4 Steps S301 to S303 are shown.

[0126] S301: Determine a target first original slice and a target second original slice according to the sequence number of the damaged check data slice.

[0127] The first check slice of the damaged check data slice is a first coupling check slice, the second check slice of the damaged check data slice is a second coupling check slice, the target first original slice is determined from the first original slices of all undamaged original data slices, and the target second original slice is determined from the second original slices of all undamaged original data slices.

[0128] In the embodiment of the present invention, the first coupling check slice is obtained when the target second original slice is used as a coupling load, and the second coupling check slice is obtained when the target second original slice is used as a coupling load.

[0129] It can be understood that the sequence number of the damaged parity data slice and the sequence number of the original data slice corresponding to the target first original slice satisfy a first preset relationship, and the sequence number of the damaged parity data slice and the sequence number of the original data slice corresponding to the target second original slice satisfy a second preset relationship. The first preset relationship is determined when executing step S104-3, and the second preset relationship is determined when executing step S103-3.

[0130] For example, assuming Figure 3 The check data slice s1 written to the hard disk 5 is damaged, and its first check slice a5 is the first coupling check slice. a5 is obtained when the second original slice b0 is used as the coupling load. Therefore, b0 is the target second original slice.

[0131] Similarly, the second check slice b5 of s1 is the second coupling check slice. b5 is obtained when the first original slice a1 is used as the coupling load. Therefore, a1 is the target first original slice.

[0132] That is to say, when restoring s1, it is necessary to read the undamaged original data fragments t0, t1, t2 and t3 from hard disk 0, hard disk 1, hard disk 2 and hard disk 3, and use the second original fragment b0 of t0 as the target second original fragment, and the first original fragment a1 of t1 as the target first original fragment.

[0133] S302: Obtain a first checksum slice of the damaged checksum data slice according to the first original slices of all undamaged original data slices and the target second original slice.

[0134] For example, in the recovery Figure 3When s1 is in the process, the first original fragments a0, a1, a2 and a3 of t0, t1, t2 and t3 are encoded to obtain the first encoded fragment a5', and a5' is coupled with the target second original fragment b0 to obtain the first check fragment a5 of s1.

[0135] S303: Obtain a second checksum slice of the damaged checksum data slice according to the second original slices of all undamaged original data slices and the target first original slice.

[0136] For example, in the recovery Figure 3 When s1 is in, the second original fragments b0, b1, b2 and b3 of t0, t1, t2 and t3 are encoded to obtain the second encoded fragment a5', and a5' is coupled with the target second original fragment b0 to obtain the first check fragment a5 of s1.

[0137] It can be understood that in the case where the first check slice of the damaged check data slice is the first non-coupling check slice and the second check slice is the second non-coupling check slice, it is only necessary to encode the first original slice of all undamaged original data slices to obtain the first check slice of the damaged check data slice.

[0138] Similarly, it is only necessary to perform encoding processing on the second original slices of all undamaged original data slices to obtain the second verification slices of the damaged verification data slices.

[0139] For data stored in a storage device using the above method, if multiple original data slices are damaged, or if both original data slices and check data slices are damaged and need to be reconstructed, the data processing method provided by the embodiment of the present invention also includes the following steps: Figure 5 Steps S401 to S404 are shown.

[0140] S401: Obtain a coding coupling matrix corresponding to the data to be recovered.

[0141] For example, assuming Figure 3 The original data slice t0 written to hard disk 0 and the check data slice s0 written to hard disk 4 are damaged. At this time, the following encoding coupling matrix G needs to be obtained:

[0142]

[0143] It can be understood that the process of steps S103 to S104 above can be equivalent to the process of performing matrix operation on the coding coupling matrix G and the vector (a0, b0, a1, b1, a2, b2, a3, b3, a4', b4', a5', b5', a6', b6'), and the result of the matrix operation is the vector (a0, b0, a1, b1, a2, b2, a3, b3, a4, b4, a5, b5, a6, b6).

[0144] S402: Delete row elements corresponding to damaged original data slices and damaged check data slices in the coding coupling matrix to obtain a processed coding coupling matrix, and obtain an inverse matrix of the processed coding coupling matrix.

[0145] Exemplarily, the two rows of elements (row 1 and row 2) corresponding to the original data slice t0 in the coding coupling matrix G are deleted, and the two rows of elements (row 9 and row 10) corresponding to the check data slice s0 are deleted to obtain the following processed coding coupling matrix G1, and then the inverse matrix G2 of G1 is obtained.

[0146]

[0147] S403 : Construct a decoding vector by using the first original slice and the second original slice of the undamaged original data slice and the first check slice and the second check slice of the undamaged check data slice.

[0148] For example, Figure 3 The undamaged original data slices t1, t2 and t3 written to hard disk 1, hard disk 2 and hard disk 3 are read out, and the undamaged verification data slices s1 and s2 written to hard disk 5 and hard disk 6 are read out.

[0149] Using the first original fragments a1, a2 and a3 of t1, t2 and t3, the second original fragments b1, b2 and b3, the first check fragments a5 and a6 of s1 and s2, and the second check fragments b5 and b6, construct a decoding vector (a1, b1, a2, b2, a3, b3, a5, b5, a6, b6).

[0150] S404: Restore the damaged original data slices and the damaged check data slices based on the inverse matrix and the decoding vector.

[0151] Exemplarily, the inverse matrix G2 of the processed coding coupling matrix G1 is subjected to matrix operation on the decoding vector (a1, b1, a2, b2, a3, b3, a5, b5, a6, b6) over a finite field to recover t0 and s0.

[0152] Compared with the prior art, the beneficial effect of the embodiments of the present invention lies in that by dividing each original data shard into a first original shard and a second original shard, a first check shard that meets the preset erasure ratio is obtained according to all the first original shards and the second original shard as the second target shard, and a first check shard that meets the preset erasure ratio is obtained according to all the second original shards and the first original shard as the first target shard, thereby obtaining multiple check data shards that meet the preset erasure ratio, so that when recovering the damaged original data shard, only the first original shard or the second original shard of the undamaged original data shard and the first check shard or the second check shard of the undamaged check data shard need to be read, thereby reducing the amount of data required for the recovery process, thereby reducing the time spent on reading data and improving data recovery efficiency.

[0153] In order to execute the corresponding steps in the above method embodiment and various possible implementations, an implementation of the data processing device 100 is provided below.

[0154] Please refer to Figure 6 The data processing device 100 includes an acquisition module 101 , a processing module 102 and a writing module 103 .

[0155] An acquisition module 101 is configured to acquire a plurality of ordered original data shards, wherein the plurality of original data shards are obtained by dividing the data to be stored according to a preset erasure ratio, and each original data shard is divided into a first original shard and a second original shard;

[0156] The processing module 102 is used to determine a first target shard and a second target shard based on the sequence number of each original data shard, wherein the first target shard is determined from all first original shards, and the second target shard is determined from all second original shards; obtain a plurality of first check shards based on multiple first original shards and multiple second target shards; obtain a plurality of second check shards based on multiple second original shards and multiple first target shards; obtain a plurality of check data shards based on the multiple first check shards and multiple second check shards, wherein each check data shard includes a first check shard and a second check shard, and the first check shard and the second check shard included in any two check data shards are different.

[0157] The writing module 103 is used to write each original data slice and each verification data slice into different hard disks respectively.

[0158] Optionally, the processing module 102 is specifically used to, for each original data shard, if the serial number of the original data shard is an odd number, use the first original shard of the original data shard as the first target shard; if the serial number of the original data shard is an even number, use the second original shard of the original data shard as the second target shard.

[0159] Optionally, the multiple first check slices include a first non-coupling check slice and at least one first coupling check slice, and the processing module 102 is further specifically used to perform encoding processing on all the first original slices to generate multiple first coded slices; according to the generation order of each first coded slice, determine a first non-coupling check slice and each first coded slice to be coupled from the multiple first coded slices; according to the generation order of each first coded slice to be coupled and the serial number of the original slice corresponding to each second target slice, determine the coupling load of each first coded slice to be coupled, wherein the coupling load of each first coded slice to be coupled includes at least one second target slice; couple each first coded slice to be coupled with the coupling load of each first coded slice to be coupled to obtain each first coupling check slice.

[0160] Optionally, the multiple second check slices include a second non-coupling check slice and at least one second coupled check slice, and the processing module 102 is further specifically used to perform encoding processing on all the second original slices to generate multiple second coded slices; according to the generation order of each second coded slice, determine a second non-coupling check slice and each second coded slice to be coupled from the multiple second coded slices; according to the generation order of each second coded slice to be coupled and the serial number of the original slice corresponding to each first target slice, determine the coupling load of each second coded slice to be coupled, wherein the coupling load of each second coded slice to be coupled includes at least one first target slice; couple each second coded slice to be coupled with the coupling load of each second coded slice to be coupled to obtain each second coupled check slice.

[0161] Optionally, there is data to be recovered in the storage device, and the data to be recovered includes a damaged original data slice, an undamaged original data slice, and an undamaged verification data slice, and there is only one damaged verification data slice. The processing module 102 is further used to, if the first original slice of the damaged original data slice is the first target slice, recover the first original slice and the second original slice of the damaged original data slice according to the second original slice of the undamaged original data slice and the second verification slice of the undamaged verification data slice; if the second original slice of the damaged original data slice is the second target slice, recover the first original slice and the second original slice of the damaged original data slice according to the first original slice of the undamaged original data slice and the first verification slice of the undamaged verification data slice.

[0162] Optionally, there is data to be recovered in the storage device, and the data to be recovered includes damaged verification data slices and undamaged original data slices. The processing module 102 is further used to determine the target first original slice and the target second original slice according to the sequence number of the damaged verification data slice, wherein the first original slice is determined from the first original slice of all undamaged original data slices, and the target second original slice is determined from the second original slice of all undamaged original data slices; the first verification slice of the damaged verification data slice is obtained according to the first original slices and the second target slice of all undamaged original data slices; the second verification slice of the damaged verification data slice is obtained according to the second original slices and the first target slice of all undamaged original data slices.

[0163] Optionally, there is data to be recovered in the storage device, and the data to be recovered includes damaged original data slices, damaged verification data slices, undamaged original data slices, and undamaged verification data slices. The processing module 102 is further used to obtain a coding coupling matrix corresponding to the data to be recovered; delete the row elements corresponding to the damaged original data slices and the damaged verification data slices in the coding coupling matrix to obtain a processed coding coupling matrix, and obtain an inverse matrix of the processed coding coupling matrix; use the first original slice and the second original slice of the undamaged original data slice, and the first verification slice and the second verification slice of the undamaged verification data slice to construct a decoding vector; based on the inverse matrix and the decoding vector, restore the damaged original data slice and the damaged verification data slice.

[0164] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the data processing device 100 described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0165] Furthermore, the embodiment of the present invention also provides an electronic device 200, please refer to Figure 7 , the electronic device 200 may include a memory 210 and a processor 220 .

[0166] The processor 220 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the data processing method provided in the above method embodiment.

[0167] The memory 210 can be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 210 can exist independently and be connected to the processor 220 via a communication bus. The memory 210 can also be integrated with the processor 220. Among them, the memory 210 is used to store machine executable instructions for executing the scheme of the present application. The processor 220 is used to execute the machine executable instructions stored in the memory 210 to implement the above-mentioned method embodiment.

[0168] An embodiment of the present invention further provides a computer-readable storage medium containing a computer program. When the computer program is executed, it can be used to perform relevant operations in the data processing method provided by the above method embodiment.

[0169] In summary, the embodiments of the present invention provide a data processing method, apparatus, electronic device, and storage medium. First, a plurality of ordered original data shards are obtained, wherein the plurality of original data shards are obtained by dividing the data to be stored according to a preset erasure ratio, and each original data shard is divided into a first original shard and a second original shard; then, according to the sequence number of each original data shard, a first target shard and a second target shard are determined, wherein the first target shard is determined from all the first original shards, and the second target shard is determined from all the second original shards; then, Based on multiple first original shards and multiple second target shards, multiple first check shards are obtained; based on multiple second original shards and multiple first target shards, multiple second check shards are obtained; then based on the multiple first check shards and the multiple second check shards, multiple check data shards are obtained, wherein each check data shard includes a first check shard and a second check shard, and the first check shard and the second check shard included in any two check data shards are different; finally, each of the original data shards and each of the check data shards are written to different hard disks respectively. In the embodiment of the present invention, by dividing each original data shard into a first original shard and a second original shard, a first check shard satisfying a preset erasure ratio is obtained based on all the first original shards and the second original shard as the second target shard, and a first check shard satisfying a preset erasure ratio is obtained based on all the second original shards and the first original shard as the first target shard, thereby obtaining a plurality of check data shards satisfying the preset erasure ratio. When recovering a damaged original data shard, only the first original shard or the second original shard of the undamaged original data shard and the first check shard or the second check shard of the undamaged check data shard need to be read, thereby reducing the amount of data required for the recovery process, thereby reducing the time spent on reading data, and improving data recovery efficiency.

[0170] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A data processing method, characterized in that: Applied to a storage device comprising multiple hard disks, the method comprises: Acquire a plurality of ordered original data shards, wherein the plurality of original data shards are obtained by dividing the data to be stored according to a preset erasure ratio, and each of the original data shards is divided into a first original shard and a second original shard; For each of the original data shards, if the sequence number of the original data shard is an odd number, the first original shard of the original data shard is used as the first target shard; if the sequence number of the original data shard is an even number, the second original shard of the original data shard is used as the second target shard; Obtaining a plurality of first verification fragments based on the plurality of first original fragments and the plurality of second target fragments; Obtaining a plurality of second verification fragments based on the plurality of second original fragments and the plurality of first target fragments; Obtaining a plurality of verification data fragments according to the plurality of first verification fragments and the plurality of second verification fragments, wherein each of the verification data fragments includes a first verification fragment and a second verification fragment, and the first verification fragment and the second verification fragment included in any two of the verification data fragments are different; Writing each of the original data slices and each of the verification data slices into different hard disks respectively; The multiple first check slices include a first uncoupled check slice and at least one first coupled check slice, and the multiple second check slices include a second uncoupled check slice and at least one second coupled check slice; the steps of obtaining the multiple first check slices based on the multiple first original slices and the multiple second target slices; and obtaining the multiple second check slices based on the multiple second original slices and the multiple first target slices include: Performing encoding processing on all the first original slices and all the second original slices to generate a plurality of first encoded slices and a plurality of second encoded slices; Determining, from the plurality of first coded slices, a first uncoupling check slice and each first coded slice to be coupled according to the generation order of each first coded slice; determining, from the plurality of second coded slices, a second uncoupling check slice and each second coded slice to be coupled according to the generation order of each second coded slice; Determine, based on the generation order of each first coded slice to be coupled and the sequence number of the original slice corresponding to each second target slice, a coupling load of each first coded slice to be coupled; determine, based on the generation order of each second coded slice to be coupled and the sequence number of the original slice corresponding to each first target slice, a coupling load of each second coded slice to be coupled; wherein, the coupling load of each first coded slice to be coupled includes at least one second target slice, and the coupling load of each second coded slice to be coupled includes at least one first target slice; Each first coding slice to be coupled is coupled with a coupling load of each first coding slice to be coupled, and each second coding slice to be coupled is coupled with a coupling load of each second coding slice to be coupled, to obtain each first coupling check slice and each second coupling check slice.

2. The method according to claim 1, wherein The storage device contains data to be recovered, the data to be recovered including a damaged original data slice, an undamaged original data slice, and an undamaged check data slice, the damaged original data slice being one, and the method further comprising: If the first original shard of the damaged original data shard is the first target shard, restoring the first original shard and the second original shard of the damaged original data shard according to the second original shard of the undamaged original data shard and the second check shard of the undamaged check data shard; If the second original slice of the damaged original data slice is the second target slice, the first original slice and the second original slice of the damaged original data slice are restored according to the first original slice of the undamaged original data slice and the first check slice of the undamaged check data slice.

3. The method according to claim 1, wherein The storage device contains data to be recovered, the data to be recovered including damaged parity data slices and undamaged original data slices, the first parity slice of the damaged parity data slice is a first coupling parity slice, and the second parity slice of the damaged parity data slice is a second coupling parity slice. The method further includes: Determining a target first original slice and a target second original slice according to the sequence number of the damaged parity data slice, wherein the target first original slice is determined from the first original slices of all the undamaged original data slices, and the target second original slice is determined from the second original slices of all the undamaged original data slices; Obtaining a first checksum slice of the damaged checksum data slice according to the first original slices of all the undamaged original data slices and the target second original slice; A second parity slice of the damaged parity data slice is obtained according to the second original slices of all the undamaged original data slices and the target first original slice.

4. The method according to claim 1, wherein The storage device contains data to be recovered, the data to be recovered including damaged original data slices, damaged verification data slices, undamaged original data slices, and undamaged verification data slices, and the method further includes: Obtaining a coding coupling matrix corresponding to the data to be recovered; Deleting row elements corresponding to damaged original data slices and damaged check data slices in the coding coupling matrix to obtain a processed coding coupling matrix, and obtaining an inverse matrix of the processed coding coupling matrix; Constructing a decoding vector using the first original slice and the second original slice of the undamaged original data slice and the first check slice and the second check slice of the undamaged check data slice; The damaged original data slice and the damaged check data slice are restored based on the inverse matrix and the decoding vector.

5. A data processing device, characterized in that: Applied to a storage device comprising multiple hard disks, the apparatus comprises: an acquisition module, configured to acquire a plurality of ordered original data shards, wherein the plurality of original data shards are obtained by dividing the data to be stored according to a preset erasure ratio, and each of the original data shards is divided into a first original shard and a second original shard; Processing module for: For each of the original data shards, if the sequence number of the original data shard is an odd number, the first original shard of the original data shard is used as the first target shard; if the sequence number of the original data shard is an even number, the second original shard of the original data shard is used as the second target shard; Based on the plurality of first original slices and the plurality of second target slices, a plurality of first check slices are obtained; the plurality of first check slices include a first uncoupled check slice and at least one first coupled check slice; Based on the plurality of second original fragments and the plurality of first target fragments, a plurality of second parity fragments are obtained; the plurality of second parity fragments include a second uncoupled parity fragment and at least one second coupled parity fragment; Obtaining a plurality of verification data fragments according to the plurality of first verification fragments and the plurality of second verification fragments, wherein each of the verification data fragments includes a first verification fragment and a second verification fragment, and the first verification fragment and the second verification fragment included in any two of the verification data fragments are different; A writing module, configured to write each of the original data slices and each of the verification data slices into different hard disks respectively; The processing module is specifically configured to perform encoding processing on all the first original slices and all the second original slices to generate a plurality of first coded slices and a plurality of second coded slices; determine a first non-coupling check slice and each first coded slice to be coupled from the plurality of first coded slices according to the generation order of each first coded slice; determine a second non-coupling check slice and each second coded slice to be coupled from the plurality of second coded slices according to the generation order of each first coded slice to be coupled and the sequence number of the original slice corresponding to each second target slice, and determine the coupling order of each first coded slice to be coupled. coupling load; determining the coupling load of each second coded slice to be coupled according to the generation order of each second coded slice to be coupled and the serial number of the original slice corresponding to each first target slice; wherein the coupling load of each first coded slice to be coupled includes at least one second target slice, and the coupling load of each second coded slice to be coupled includes at least one first target slice; coupling each first coded slice to be coupled with the coupling load of each first coded slice to be coupled, and coupling each second coded slice to be coupled with the coupling load of each second coded slice to be coupled, to obtain each first coupling check slice and each second coupling check slice.

6. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the data processing method according to any one of claims 1 to 4 when executing the computer program.

7. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed by a processor, implements the data processing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Data processing method and device based on erasure codes

    CN111090540A

  • Systems and methods for ultra fast ECC with parity

    US20190384671A1