Data processing method, processing node, data processing system and storage medium

By generating multiple encoded data blocks, the problem of high repair bandwidth in large proportion of erasure codes is solved, and data repair with low repair bandwidth is achieved, and the reliability of data storage and concurrent read and write performance is improved.

WO2025145580A1PCT designated stage expired Publication Date: 2025-07-10ZTE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/109458
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2024-08-02
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The prior art has the problem of high repair bandwidth in large proportions of erasure coding, which leads to data repair occupies a large amount of network bandwidth and is difficult to effectively apply in storage networks.

Method used

Using a data processing method, multiple encoded data blocks are generated through encoding parameters and encoding matrix, including a first verification data block, a second verification data block and a third verification data block, to reduce the repair bandwidth, and ensure the data repair capability through the third verification data block in the event of a local failure.

Benefits of technology

Data repair with low repair bandwidth is achieved under wide stripe and low redundancy ratio, improving the reliability of data storage and concurrent read and write performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109458_10072025_PF_FP_ABST
    Figure CN2024109458_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a data processing method, a processing node, a data processing system and a storage medium. The method comprises: acquiring k source data blocks, k being a positive integer greater than 1; and encoding the k source data blocks on the basis of encoding parameters to obtain n encoded data blocks, n being a positive integer greater than k, wherein the encoding parameters comprise a first check quantity, a second check quantity, a third check quantity, an additional source sub-data block index table, the number of repetitions and an encoding matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, processing node, data processing system and storage medium Technical Field

[0001] The present application relates to the field of data processing technology, for example, to a data processing method, a processing node, a data processing system and a storage medium. Background Art

[0002] With the rapid development of information technology, the amount of data that needs to be stored is exploding. How to safely, reliably, and efficiently store and access this data is a question that commercial companies, and even entire nations, must consider. The amount of data to be stored will increase annually in the future. If a smaller erasure code ratio is used, redundant checksum data will occupy a large amount of storage space, leading to higher storage costs. Therefore, large-scale erasure code (EC) technology is an important research direction in future storage EC. Large-scale EC technology can use wider stripes and lower checksum ratios, thereby improving storage efficiency. However, directly increasing the EC ratio based on current erasure codes will result in a large amount of data repair bandwidth when repairing one or more source data nodes. That is, when one or more source data nodes are faulty, k times the amount of data from a single source data node needs to be transmitted across the network. Data repair consumes a large amount of data transmission network bandwidth, which is difficult to accept in storage networks.

[0003] Summary of the Invention

[0004] The present application provides a data processing method, a processing node, a data processing system and a storage medium.

[0005] An embodiment of the present application provides a data processing method, applied to a first processing node, comprising:

[0006] Get k source data blocks, where k is a positive integer greater than 1;

[0007] Encoding the k source data blocks according to the encoding parameters to obtain n encoded data blocks, where n is a positive integer greater than k;

[0008] The encoding parameters include: a first check number, a second check number, a third check number, an additional source sub-data block index table, a repetition number, and an encoding matrix.

[0009] The embodiment of the present application further provides a data processing method, which is applied to a second processing node and includes:

[0010] Obtaining the type and number of faulty coded data blocks and part or all of the data of at least k non-faulty coded data blocks generated by the first processing node, where k is a positive integer greater than 1;

[0011] Part or all of the data of the at least k non-faulty coded data blocks are processed according to the type and number of the faulty coded data blocks to obtain the faulty coded data blocks.

[0012] An embodiment of the present application further provides a processing node, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned data processing method when executing the program.

[0013] An embodiment of the present application also provides a data processing system, including: a storage node set, a first processing node, and a second processing node; the storage node set includes n storage nodes, and the n storage nodes include k source data storage nodes and nk verification data storage nodes, where n is a positive integer greater than k, and k is a positive integer greater than 1.

[0014] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned data processing method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] FIG1 is a schematic diagram of an implementation environment of a data processing method provided by an embodiment;

[0016] FIG2 is a schematic diagram of a storage node set provided by an embodiment;

[0017] FIG3 is a flow chart of a data processing method provided by an embodiment;

[0018] FIG4 is a schematic diagram of an array corresponding to n coded data blocks provided by an embodiment;

[0019] FIG5 is a schematic diagram of generating a first verification data block provided by an embodiment;

[0020] FIG6 is a schematic diagram of generating a third check data block when linear combination coefficients are obtained from a Vandermonde matrix according to an embodiment;

[0021] FIG7 is a schematic diagram of generating a third check data block when linear combination coefficients are obtained from a deformation matrix of a Cauchy matrix, provided by an embodiment;

[0022] FIG8 is a schematic diagram of generating a third check data block when another linear combination coefficient is obtained from a deformation matrix of a Cauchy matrix according to an embodiment;

[0023] FIG9 is a flow chart of another data processing method provided by an embodiment;

[0024] FIG10 is a schematic structural diagram of a data processing device provided by an embodiment;

[0025] FIG11 is a schematic structural diagram of another data processing device provided by an embodiment;

[0026] FIG12 is a schematic diagram of the hardware structure of a first processing node provided by an embodiment;

[0027] FIG13 is a schematic diagram of the hardware structure of a second processing node provided by an embodiment;

[0028] FIG14 is a schematic structural diagram of a data processing system provided by an embodiment. DETAILED DESCRIPTION

[0029] The present application is described below in conjunction with the accompanying drawings and embodiments. It will be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application. It should be noted that, unless there is a conflict, the embodiments and features within the embodiments of the present application may be combined with each other in any manner. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present application, not all structures.

[0030] It should be noted that although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown in the flowcharts. The terms "first," "second," and the like in the specification, claims, and drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0031] In data storage systems, redundancy is usually added to ensure data reliability and persistence. Currently, the commonly used methods include multiple copies and erasure codes.

[0032] Multi-copy technology involves making n copies of the same data, which are then distributed and stored across different nodes according to a specific strategy. These nodes include, but are not limited to, disks, solid-state drives, servers, and other storage devices. These nodes are independent. If up to n-1 nodes fail, causing data failure within them, the same data as in the failed nodes can still be retrieved from the remaining nodes, thus providing data protection. The advantage of using multi-copy technology is that it requires no complex calculations, making it simple to operate. The lack of complex multiplication operations also results in faster data read and write speeds. However, multi-copy technology has low storage space utilization, which will exponentially increase storage costs for future massive data storage. Therefore, multi-copy technology is typically used for hot data.

[0033] Erasure codes are a forward error correction technology primarily used to prevent packet loss in network transmission and are often used in storage systems to improve storage reliability. Compared to multiple replicas, erasure codes can achieve higher data reliability with less redundancy. However, the encoding and decoding of erasure codes is more complex, involving multiplication operations in a finite field, and therefore requires more computing resources. In storage systems, they are generally applied to warm data or data. The most widely used erasure codes in storage systems are RS (Reed-Solomon) erasure codes. However, in large-scale ECs, RS erasure codes have a large decoding workload and a high repair bandwidth. That is, repairing a single data node requires k times the amount of data transmitted across the network as that of the single data node. To address the problem of high repair bandwidth, the concepts of locally repairable codes (LRC) and regeneration codes (RGC) have been proposed. Locally repairable codes have good repair locality. When a single source data node fails, repair only needs to be performed in the local area where the source data node is located, reducing the repair bandwidth. However, locally repairable codes are not maximum distance separable codes (MDS codes). Compared with RS codes, they use more parity data storage space with the same repair capability. Regeneration codes have smaller repair bandwidth overhead than RS codes. When a single source data node fails, the remaining data nodes only need to transmit part of their own data to complete the repair of the source data node. However, regeneration codes also have the problems of not being able to set the number of source data nodes k too large, high decoding complexity, and difficulty in meeting the MDS characteristics. These problems make regeneration codes difficult to apply in engineering when meeting a large proportion of EC.

[0034] The embodiment of the present application provides a data processing method that can complete the repair of source data nodes using low repair bandwidth in wide stripes and low redundancy ratio conditions, thereby improving the reliability of data storage and concurrent read and write performance.

[0035] Figure 1 is a schematic diagram of an implementation environment for a data processing method according to an embodiment. As shown in Figure 1 , the implementation environment includes, but is not limited to, a first transmission node 101, a second transmission node 201, networks 102 and 202, a first processing node 10, a second processing node 20, and a storage node set 30, wherein the storage node set includes at least one storage node 301.

[0036] In one embodiment, the first transmission node 101 and the second transmission node 201 may be the same device or different devices. The first transmission node 101 and the second transmission node 201 may include but are not limited to: a server, a base station, a mobile phone, a personal computer, a laptop, a tablet computer, an access point (AP), or other devices with the function of transmitting and storing data.

[0037] In one embodiment, the relative position and number of the first transmission node 101 and the second transmission node 201 may vary in specific application scenarios. For example, regarding relative position, the first transmission node 101 can send data over the network, and the second transmission node 201 can receive data over the network; at the same time, the second transmission node 201 can send data over the network, and the first transmission node 101 can receive data over the network. Regarding number, the number of the first transmission node 101 and the second transmission node 201 can be one or more, and they are connected via a network.

[0038] In one embodiment, network 102 and network 202 may include, but are not limited to, a wired network, a wireless network, a cellular network, a local area network, the Internet, a wide area network, the World Wide Web, or other networks that transmit data.

[0039] In one embodiment, the first processing node 10 and the second processing node 20 can be the same processing node of the same device at different times. The first processing node 10 and the second processing node 20 can access and transmit data through the network, and process the data according to the configuration items. In one example, the configuration items include but are not limited to: the source storage node location of the accessed data, the destination storage node location of the transmitted data, the number of check data blocks generated by encoding the data blocks, the number of first check data blocks generated by encoding the data blocks, the number of second check data blocks generated by encoding the data blocks, the number of source data blocks, the size of the source data blocks, the number of storage nodes, the number and location of faulty storage nodes, data recovery strategies, or other parameters for data processing. In one example, data processing includes but is not limited to: encoding the source data blocks, allocating data blocks to storage nodes, accessing data blocks from storage nodes, recovering faulty data blocks, or other processing methods such as reading, writing, deleting, and modifying data blocks.

[0040] In one embodiment, the execution entity of the data processing method in this embodiment may include but is not limited to the first processing node 10 and the second processing node 20 in the embodiment shown in Figure 1, or technical personnel in this field can choose to set the corresponding execution entity according to the actual application scenario. This embodiment does not make any restrictions, but it should not be understood as a limitation on the embodiments of the present application.

[0041] In one embodiment, the storage node set 30 includes at least n storage nodes 301, each storage node 301 may include but is not limited to: RAM, ROM, Electrically Erasable Programmable Read Only Memory (EEPROM), flash memory or other memory technology, or other optical disk storage, magnetic cassettes, tapes, disk storage, hard disks, solid-state drives or other storage devices, or any other medium that can be used to store desired information and can be accessed by a computer.

[0042] Figure 2 is a schematic diagram of a storage node set provided by one embodiment. As shown in Figure 2, a storage node set 30 includes n storage nodes 301, including k source data storage nodes and m = nk check data storage nodes, where n is a positive integer greater than k, and k is a positive integer greater than 1. Each storage node includes p Chunk data blocks 302, where p is an integer greater than 0. Each Chunk data block includes w Page data blocks 303, where w is an integer greater than 0. Each stripe 304 includes e*n Page data blocks, with each e Page data blocks distributed across one storage node. That is, within each stripe, each storage node includes e Page data blocks, where e is an integer greater than 0.

[0043] In one embodiment, the k source data blocks obtained correspond to the data parts stored in the first k storage nodes in a stripe, and the k source data blocks are distributed in k different storage nodes; the m check data blocks obtained by encoding the k source data blocks according to the encoding parameters correspond to the data parts stored in the last m storage nodes in a stripe, and the m check data blocks are respectively stored in the remaining m different storage nodes.

[0044] In one embodiment, the portion of each storage node in each stripe containing e Page data blocks is regarded as a source data block, k source data blocks constitute the source data portion of a stripe, and m check data blocks obtained by encoding the k source data blocks constitute the check data portion of a stripe.

[0045] Figure 3 is a flowchart of a data processing method provided in one embodiment, which can be applied to a first processing node. The execution entity of this data processing method may include, but is not limited to, the first processing node 10 shown in Figure 1. Alternatively, those skilled in the art may select and set a corresponding execution entity based on the actual application scenario, and this embodiment does not impose any restrictions on this. As shown in Figure 3, the method provided in this embodiment includes steps 110 and 120.

[0046] In step 110 , k source data blocks are obtained, where k is a positive integer greater than 1.

[0047] In step 120, the k source data blocks are encoded according to the encoding parameters to obtain n encoded data blocks, where n is a positive integer greater than k.

[0048] In this embodiment, the encoding parameters include: a first check quantifier, a second check quantifier, a third check quantifier, an additional source sub-data block index table, a repetition count, and an encoding matrix.

[0049] The source data block can be selected and set according to a specific application scenario, for example, obtained from a file, etc., and this embodiment does not limit this.

[0050] In this embodiment, when a first processing node obtains k source data blocks, it encodes the k source data blocks according to at least the following parameters to obtain n coded data blocks: a first checksum, a second checksum, a third checksum, an index table of additional source sub-data blocks, a number of repetitions, and an encoding matrix. When one or more nodes fail, a second processing node can obtain the type and number of the failed coded data blocks, as well as partial or full data of at least k non-faulty coded data blocks generated by the first processing node, and process partial or full data of at least k non-faulty coded data blocks based on the type and number of the failed coded data blocks to obtain the failed coded data blocks. The presence of the first and second checksums enables the source data blocks to be repaired with a low repair bandwidth when one or more source data nodes fail. The presence of the third checksum enables the source data blocks to be repaired using the third checksum when the first and second checksums are unable to complete data repair. Specifically, the first and second checksums reduce the repair bandwidth, while the third checksum ensures repair capability.

[0051] The encoding method of this embodiment is applicable to wide stripe repair, and can use wider stripes to improve concurrent read and write performance. The lower redundancy ratio reduces the storage overhead of redundant check data, and has the characteristics of low repair bandwidth, high repair capability, and high concurrent read of wide EC stripes. In addition, when a single source data node fails, the erasure coding method of this application can complete data repair at the optimal repair bandwidth or a level close to the optimal repair bandwidth, and the r check data nodes can repair all r-1 data node (including source data nodes and check data nodes) type failures, as well as most of the r data node types failures.

[0052] In one embodiment, the n coded data blocks include k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks, where m1 is the first check number and m1=1, m2 is the second check number and is a positive integer, m3 is the third check number and is a positive integer, and n=k+m1+m2+m3;

[0053] The additional source sub-data index table is an index table with g rows and m2 columns, where g is the number of source sub-data blocks contained in each source data block;

[0054] The size of the encoding matrix is ​​m*k, where m is the total number of check data blocks, m=m1+m2+m3;

[0055] The number of repetitions is a positive integer;

[0056] The elements in the encoding matrix are elements under the Galois field GF(q), where q is an integer power of a prime number.

[0057] In this embodiment, the n coded data blocks include k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks, and n=k+m1+m2+m3. In addition, m1=1; m2 is an integer greater than 0; and m3 is an integer greater than 0. The additional source sub-data index table is an index table of size g*m2, which is used to indicate the position of the additional source sub-data blocks required to generate the second check data block; the number of repetitions s is an integer greater than 0, which is used to obtain the number of source sub-data blocks contained in each source data block; the coding matrix is ​​a matrix of size (m1+m2+m3)*k, which is used to obtain the linear combination coefficients for generating the first check data block, the second check data block, and the third check data block. The elements in the coding matrix are elements under the Galois field GF(q), and q is an integer power of a prime number.

[0058] FIG4 is a schematic diagram of an array corresponding to n coded data blocks provided by an embodiment. As an example, the obtained n coded data blocks include: k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks. The schematic diagram of the k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks in the array is shown in FIG4. In the data array 1200 of the n coded data packets obtained by encoding, there are k source data blocks 1201, and each source data block includes g source sub-data blocks 1205. In one example, a source sub-data block is a 0,0, ; The first check data block 1202 includes g check sub-data blocks 1206; among the m2 second check data blocks 1203, each second check data block includes g check sub-data blocks; among the m3 third check data blocks 1204, each third check data block includes g check sub-data blocks; wherein the size of each source sub-data block is equal to the size of each check sub-data block.

[0059] In one embodiment, the k source data blocks respectively correspond to the data parts stored in the first k storage nodes in a stripe; the k source data blocks are distributed in k different storage nodes; the m check data blocks obtained by encoding the k source data blocks according to the encoding parameters respectively correspond to the data parts stored in the last m storage nodes in a stripe; the m check data blocks are respectively stored in m different storage nodes other than the k different storage nodes.

[0060] In one embodiment, k source data blocks belong to a stripe, corresponding to the source data portion of a stripe, and the k source data blocks are represented as a=[a0, a1, ..., a k-1 ], each source data block is represented as a i , 0≤i≤k-1, each source data block contains T kilobytes, T is an integer greater than 0, and T includes but is not limited to one of the following: 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096.

[0061] In one embodiment, each source data block a i It also includes g source sub-data blocks, and the i-th source data block is represented by a i =[a 0,i ,a 1,i ,...,a g-1,i ], g is an integer greater than 1. Each source sub-data block includes E kilobytes, E is a real number greater than 0 and the product of which and 1024 is a positive integer, and E includes but is not limited to one of the following: 1 / 2, 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024.

[0062] In one embodiment, the number of rows in the array of n coded data blocks is determined by at least one of the following parameters: k, the first check number, and the second check number, wherein the number of rows in the array of n coded data blocks is equal to the number of source sub-data blocks contained in each source data block.

[0063] In one embodiment, the number of rows of the array is: Where s is the number of repetitions and is a positive integer.

[0064] As an example, the number of rows g of the array is determined according to the following rules: Where s represents the number of repetitions, and s is an integer greater than 0.

[0065] In one example, the number of rows g of the array is determined according to the rule It is determined that the first check number m1=1, the second check number m2=3, the number of source data blocks k=20, and the number of repetitions s=1, then the number of rows of the array g=1024 is obtained.

[0066] In another example, the number of rows g of the array is calculated according to the rule It is determined that the first check number m1=1, the second check number m2=3, the number of source data blocks k=20, and the number of repetitions s=2, then the number of rows of the array g=64.

[0067] In another example, the number of rows g of the array is calculated according to the rule It is determined that the first check number m1=1, the second check number m2=3, the number of source data blocks k=160, and the number of repetitions s=8, then the number of rows of the array g=1024 is obtained.

[0068] In another example, the number of rows g of the array is calculated according to the rule It is determined that the first check number m1=1, the second check number m2=3, the number of source data blocks k=150, and the number of repetitions s=8, then the number of rows of the array g=1024 is obtained.

[0069] In one embodiment, the encoding matrix satisfies at least one of the following:

[0070] There is a row with all element values ​​​​"1";

[0071] Any square matrix consisting of m rows and m columns is reversible under the Galois field GF(q).

[0072] In this embodiment, the size of the encoding matrix is ​​m*k, and the element values ​​in the encoding matrix are the element values ​​in the Galois field GF(q). The encoding matrix has at least one of the following characteristics: there is a row whose element values ​​are all "1", and any square matrix consisting of m rows and m columns is reversible under the Galois field GF(q). Among them, there is a row whose element value is "1" to generate the first check data block. Since each check sub-data block of the first check data block is the bitwise exclusive OR of the k source sub-data blocks in the current row, the coding coefficients of the k source sub-data blocks are all 1, that is, p i,0 =1*a i,0 +1*a i,1 +...+1*a i,k-1 The remaining (m-1)=m1+m2 rows of the encoding matrix are used to generate the second and third check data blocks. Any square matrix consisting of m rows and m columns of the encoding matrix is ​​reversible in the Galois field GF(q), thereby making the encoding method as close as possible to the MDS code; where q is an integer power of a prime number. In one example, q is an integer power of 2, and q includes but is not limited to one of the following: 256, 128, 64, 32, 16, 8, 4, and 2.

[0073] In one embodiment, the encoding matrix may include, but is not limited to: a Vandermonde matrix, a modified matrix of a Cauchy matrix, and a combination of an all-1 matrix and a Cauchy matrix.

[0074] In one example, when the encoding matrix is ​​a Vandermonde matrix, the encoding matrix G is expressed as:

[0075] Among them, α is a primitive element under the Galois field GF(q), i1,i2,...,i m are all integers greater than or equal to 0 and less than or equal to q-2, i1,i2,...,i m The values ​​in are different, and i1,i2,...,i m There is only one value in [i1,i2,...,i m ]=[0,1,...,m-1], m=4, k=4, q=256,

[0076] Then the encoding matrix G is:

[0077] In another example, when the encoding matrix is ​​a transformation matrix of the Cauchy matrix, the encoding matrix is ​​expressed as:

[0078] Where α is a primitive element under the Galois Field GF(q), “+” represents bitwise exclusive OR, and “ / ” represents the division operation under the Galois Field GF(q). m-1are all integers greater than or equal to k and less than or equal to q-2, j1, j2, ..., j m-1 In a specific example, [j1,j2,...,j m-1 ]=[k,k+1,...,k+m-2], m=4, k=4, q=256, then the encoding matrix G is:

[0079] In another different example, when the encoding matrix is ​​a combination of an all-1 matrix and a Cauchy matrix, the encoding matrix is ​​expressed as:

[0080] Where α is a primitive element under the Galois Field GF(q), “+” represents bitwise exclusive OR, and “ / ” represents the division operation under the Galois Field GF(q). m-1 are all integers greater than or equal to k and less than or equal to q-2, j1, j2, ..., j m-1 In a specific example, [j1,j2,...,j m-1 ]=[k,k+1,...,k+m-2], m=4, k=4, q=256, then the encoding matrix G is:

[0081] It should be noted that the row index of the encoding matrix in this application increases from subscript 0.

[0082] In one embodiment, each check sub-data block in the first check data block is obtained by bitwise exclusive ORing the source sub-data blocks in the row of the array of the n coded data blocks.

[0083] In one example, the i-th syndrome data block of the first parity data block is represented as p i,0 , the source sub-data blocks corresponding to the i-th check sub-data block of the first check data block in the i-th row are [a i,0 ,a i,1 ,...,a i,k-1 ], then the i-th syndrome data block of the first parity data block can be obtained according to the following formula: i,0 =a i,0 +a i,1 +...+a i,k-1 , where "+" represents bitwise exclusive OR.

[0084] In one example, as shown in FIG5 , there are k=4 source data blocks, each of which has g=4 source sub-data blocks. For ease of description, each source sub-data block is one byte in size, which should not be a limitation of the present application. The one byte is represented by a decimal number, and the sub-check data blocks in the first check data block are obtained by bitwise exclusive-ORing the source sub-data blocks in the current row. For example, the value of the 0th sub-check data block in the first check data block is 10+50+90+130=216, where "+" represents bitwise exclusive-OR.

[0085] In one embodiment, each check sub-data block in the second check data block is obtained by bitwise exclusive ORing a linear combination value of source sub-data blocks in a row of the first partial array in the array of the n coded data blocks and a linear combination value of source sub-data blocks in at least one other row in the second partial array.

[0086] In one example, the i-th syndrome data block of the j-th second parity data block is represented as p i,j , i is an integer and satisfies 0≤i≤g-1, j is an integer and satisfies 1≤j≤m2; the linear combination value of the source sub-data block in the i-th row of the first part of the array is expressed as p i,j,LC1 The linear combination value of the source sub-data block of at least one other row in the second partial array is expressed as p i,j,LC2 LC1 and LC2 represent the linear combination identifiers of different types. The above values ​​satisfy the following relationship: i,j =p i,j,LC1 +p i,j,LC2 Wherein "+" represents bitwise exclusive OR. In the embodiment of the present application, the source sub-data blocks in the other rows are collectively referred to as additional source sub-data blocks.

[0087] In one embodiment, a linear combination value of the source sub-data blocks in the row of the first partial array corresponding to the i-th check sub-data block of the j-th second check sub-data block is determined by the source sub-data block in the i-th row of the array of the n coding data blocks and the j-th row in the coding matrix.

[0088] In this embodiment, the linear combination value p of the source sub-data blocks in the row of the first partial array is i,j,LC1 Determined by the source sub-data block in the i-th row of the array and the j-th row in the encoding matrix, where i is an integer satisfying 0≤i≤g-1, and j is an integer satisfying 1≤j≤m2.

[0089] As an example, the source sub-data block of row i in the first partial array is represented as [a i,0 ,a i,1 ,...,a i,k-1 ], the element values ​​of the jth row in the encoding matrix are expressed as [c j,0 ,c j,1,...,c j,k-1 ], the linear combination value p of the source sub-data blocks in the row where the first part array of the i-th parity sub-data block of the j-th second parity data block is located i,j,LC1 Obtained according to but not limited to the following rules: i,j,LC1 =c j,0 *a i,0+ c j,1 *a i,1+ ...+c j,k-1 *a i,k-1 Among them, “+” represents bitwise exclusive OR, and “*” represents multiplication operation under Galois Field GF(q).

[0090] In a specific example, the number of source data blocks k=4, and the source sub-data block in the i=0th row of the first partial array is represented as [a 0,0 ,a 0,1 ,a 0,2 ,a i,3 ]=[10,50,90,130], the element values ​​of the j=1th row in the encoding matrix are expressed as [c j,0 ,c j,1 ,...,c j,k-1 ]=[166,70,187,123], then the linear combination value of the source sub-data blocks in the row where the first partial array of the i=0th parity sub-data block of the j=1th second parity data block is located is expressed as:

[0091] p i,j,LC1 =166*10+70*50+187*90+123*130=65.

[0092] In one embodiment, a linear combination value of the source sub-data blocks of at least one other row in the second partial array is determined according to the additional source sub-data block index table and the number of repetitions.

[0093] In this embodiment, the linear combination value p of the source sub-data blocks of at least one other row in the second partial array is i, j,LC2 The additional source sub-data block index table is used to indicate the index position of the additional source sub-data block required to generate the i-th parity sub-data block of the j-th second parity data block. The additional source sub-data block index table is an index array of size g*m², where g is the number of rows in the source data block array and m² is the number of second parity data blocks.

[0094] In one embodiment, each element in the additional source sub-data block index table is a position coordinate set, and the number of one-dimensional coordinates or two-dimensional coordinates in the position coordinate set is at most or When the additional source sub-data block index table includes all the additional source sub-data block index positions required to generate the second verification data block, the number of coordinates in the position coordinate set is at most When the additional source sub-data block index table includes some of the additional source sub-data block index positions required to generate the second verification data block, the number of coordinates in the position coordinate set is at most The index positions of all additional source sub-data blocks required to generate the second verification data block are obtained through the above partial coordinates and the number of repetitions.

[0095] In one embodiment, each element in the g*m² index array is a set of position coordinates, including but not limited to one-dimensional index coordinates and two-dimensional index coordinates. It will be appreciated that both one-dimensional and two-dimensional index coordinates represent the position of the additional element in the entire data array, and there is a one-to-one correspondence between the one-dimensional and two-dimensional index coordinates.

[0096] In one embodiment, each element in the additional source sub-data block index table is a position coordinate set, and the number Ns of one-dimensional coordinates or two-dimensional coordinates in the position coordinate set is determined by at least one of the following parameters, including but not limited to: the number of source data blocks k, the first check number m1, and the second check number m2. The number Ns of one-dimensional coordinates or two-dimensional coordinates in the position coordinate set is at most one of the following values:

[0097] In one embodiment, each element in the additional source sub-data block index table is a two-dimensional coordinate set;

[0098] In the case where the additional source sub-data block index table includes all the additional source sub-data block index positions required to generate the second verification data block, the number of coordinates in the two-dimensional coordinate set is at most

[0099] In the case where the additional source sub-data block index table includes some of the additional source sub-data block index positions required for generating the second verification data block, the number of coordinates in the two-dimensional coordinate set is at most

[0100] In one embodiment, the additional source sub-data block index table is determined according to k, the first check number, and the second check number.

[0101] As an example, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the additional source sub-data block index table includes some of the additional source sub-data block index positions required to generate the second verification data block, the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most The linear combination value p of the source sub-data blocks of at least one other row in the second partial array i,j,LC2 According to the additional source sub-data block index table and the number of repetitions, it can be obtained according to but not limited to the following rules:

[0102] It is understandable that the above additional source sub-data block index table only contains the index positions of some additional source sub-data blocks. Some additional source sub-data blocks represent the previous The additional source sub-data blocks of the source data block are repeated s times to obtain the index positions of all additional source sub-data blocks, which reduces storage by introducing a small amount of calculation.

[0103] In one embodiment, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the additional source sub-data block index table includes all the additional source sub-data block index positions required to generate the second verification data block, the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most The linear combination value p of the source sub-data blocks of at least one other row in the second partial array i,j,LC2 Determined based on the additional source sub-data block index table, it can be obtained according to but not limited to the following rules:

[0104] It is understandable that the above additional source sub-data block index table contains the index positions of all additional source sub-data blocks, which does not require additional calculations, but increases storage. Indicates bitwise exclusive OR, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the two-dimensional coordinate set includes all generated p i,j,LC2 When the additional source sub-data block index position is required, the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most When the two-dimensional coordinate set includes only part of the generated p i,j, LC2 When the additional source sub-data block index position is required, the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most It is necessary to further obtain the index positions of all additional source sub-data blocks through the repetition number s.

[0105] In one embodiment, the additional source sub-data block index table IndexTable is obtained based on but not limited to the number of source data blocks k, the number of first verification data blocks m1, and the number of second verification data blocks m2, and the elements in the tth column and Gi-th row of the source sub-data block index table are respectively the coordinate indexes of the i-th column and Gj-th row sub-data blocks in the array. i is an integer greater than or equal to 1 and less than or equal to ceil(k / s); Gi is a set of row index values ​​formed by dividing the integer value {1,2,...,g} into (m1+m2)ceil(i / (m1+m2)) sub-parts in sequence, and starting from the mod(i-1,m1+m2)+1th sub-part, reading one sub-part every (m1+m2)th sub-part; the tth element in the set GSet of the i-th source data block in the array is Gj, j belongs to the first integer set, the first integer set is [floor(i / (m1+m2))+1,ceil(i / m1+m2)]\i; t belongs to the second integer set, the second integer set contains integers greater than or equal to 1 and less than or equal to the number of integers in the first integer set; Gj is Sequentially divide the integer value {1,2,...,g} into (m1+m2)ceil(j / (m1+m2)) subparts, and sequentially read a set of row index values ​​consisting of a subpart every (m1+m2) subparts, starting from the mod(j-1,m1+m2)+1th subpart. t and j are the elements of the first and second integer sets with the same index size, respectively, and the elements in the first and second integer sets are sorted from small to large. g is the number of rows in the array. ceil() represents adjusting a value to the minimum integer not less than the value itself, floor() represents adjusting a value to the maximum integer not greater than the value itself, and mod() represents a modulo operation. []\i represents removing element i from a set.

[0106] In one example, k=6, m1+m2=3, s=1, g=(m1+m2) ceil(k ' / (m1+m2))=9, when i=2, Gi=G2={4,5,6}, for the second data block, [floor(i / (m1+m2))+1,ceil(i / m1+m2)]\i={1,2,3}\{2}={1,3}, therefore, the set of the i=2th source data block in the array GSet={G1,G2,G3}\{G2}={G1,G3}, the value of t is an integer greater than or equal to 1 and less than or equal to 2, and the above G1={1,2,3}, G3={7,8,9}. For the source sub-data block index table, the t=1 column and G2={4,5,6} row are the coordinate indexes of the i=2 column and G1={1,2,3} row sub-data blocks in the array; for the source sub-data block index table, the t=2 column and G2={4,5,6} row are the coordinate indexes of the i=2 column and G3={7,8,9} row sub-data blocks in the array.

[0107] In the following examples, k' is used to represent ceil(k / s). In one example, a feasible source sub-data block index table can be obtained according to, but not limited to, the following rules:

[0108] Here, cell() constructs a cell array, floor() adjusts a value to its maximum integer, and g represents the number of rows in the data array. SplitNum() is a function that takes the sum of g, the number of first check data blocks, m1, and m2. SplitNum() outputs a matrix of size (Layer*(m1+m2))*(g / (m1+m2)), where each element is greater than or equal to 1 and less than or equal to g.

[0109] Where GNum is a matrix with a total of (ceil(k' / (m1+m2)))*r rows. The i-th row of GNum is denoted as Gi. Gi is formed by sequentially dividing the integer values ​​{1, 2, ..., g} into (m1+m2)ceil(i / (m1+m2)) sub-parts, starting from the mod(i-1,m1+m2)+1th sub-part, and reading a sub-part every (m1+m2) sub-parts. In an example, k = 6, m1+m2 = 3, and g = (m1+m2). ceil(k ' / (m1+m2)) =9, for Gi=G2, divide 1 to 9 into 3 sub-parts in order, namely {1, 2, 3}, {4, 5, 6}, {7, 8, 9}, and read a sub-part every 3 sub-parts from the 2nd sub-part mod(i-1, m1+m2)+1=2nd sub-part, so the set is {4, 5, 6}, so G2={4, 5, 6}. In another example, k=6, r=3, g=(m1+m2)ceil(k ' / (m1+m2)) =9, for Gi=G4, 1 to 9 are divided into 9 sub-parts in order, namely {1}, {2}, {3}, {4}, {5}, {6}, {7}, {8}, {9}. Starting from the mod(i-1, r)+1=1th word part, a sub-part is read every 3 sub-parts, and the set formed is {1, 4, 7}. Therefore, G4={1, 4, 7}. In one example, a feasible GNum matrix can be obtained according to, but not limited to, the following rules:

[0110] In this example, ceil() adjusts a value to the minimum integer not less than the value itself, g represents the number of rows in the data array, T represents transpose, and TempMx(j:(m1+m2):end,:) extracts all columns of the matrix TempMx and constructs a new matrix by extracting every (m1+m2) rows from the jth row to the last row. TempSubMx(:) expands a two-dimensional matrix into a one-dimensional matrix by columns.

[0111] As an example, each element in the additional source sub-data block index table is a two-dimensional coordinate set. When the two-dimensional coordinate set only includes part of the generated p i,j,LC2 When the additional source sub-data block index position is required, the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most It is necessary to further obtain the index positions of all additional source sub-data blocks through the repetition number s.

[0112] In a specific example, the number of source data blocks k = 4, the number of repetitions s = 1, the first check number m1 = 1, the second check number m2 = 3, and the maximum number of two-dimensional coordinates is Ns = 1. A feasible additional source sub-data block index table is shown in Table 1. The coordinates in the table represent the coordinate positions of the corresponding additional source sub-data blocks. For example, when i = 1 and j = 0, the coordinate set is {(1, 0)}. Then, for generating the i = 0th check sub-block of the j = 1th second check data block, the linear combination value p of the source sub-data blocks in at least one other row in the second partial array is i,j,LC2 , p i,j, LC2 =a 1,0 .

[0113] Table 1 A feasible additional source sub-data block index table when k=4, s=1, m1=1, m2=3

[0114] As an example, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the two-dimensional coordinate set only includes part of the generated p i,j,LC2 When the additional source sub-data block index position is required, the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most It is necessary to further obtain the index positions of all additional source sub-data blocks through the number of repetitions s. In a specific example, the number of source data blocks k = 8, the number of repetitions s = 1, the first check number m1 = 1, the second check number m2 = 3, and the maximum number of two-dimensional coordinates is Ns = 2. A feasible additional source sub-data block index table is shown in Table 2. For generating the linear combination value p of the source sub-data blocks in at least one other row of the second partial array of the i = 2th check sub-block of the j = 1th second check data block, i,j,LC2 ,but in, Represents bitwise exclusive OR.

[0115] Table 2 A feasible additional source sub-data block index table when k=8, s=1, m1=1, m2=3

[0116] As an example, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the two-dimensional coordinate set only includes part of the generated p i,j,LC2 When the additional source sub-data block index position is required, the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most It is necessary to further obtain the index positions of all additional source sub-data blocks through the number of repetitions s. In a specific example, the number of source data blocks k = 70, the number of repetitions s = 10, the first check number m1 = 1, the second check number m2 = 3, and the maximum number of two-dimensional coordinates is Ns = 2. A feasible additional source sub-data block index table is shown in Table 3. For example, when i = 3 and j = 1, the coordinate set is {(7,0)}. The coordinates in this set only include the first If all the additional sub-data block indices are to be obtained for the additional source sub-data blocks of the source data blocks, the additional sub-data block row index corresponding to the j-th column with other indices ranging from 7 to 69 is the same as the additional source data block row index of the mod(j,7) column. For generating the linear combination value p of the source sub-data blocks in at least one other row in the second partial array of the i=3th check sub-block of the j=1th second check data block, i,j,LC2 ,but in, Represents bitwise exclusive OR.

[0117] Table 3 A feasible additional source sub-data block index table when k=70, s=10, m1=1, m2=3

[0118] As an example, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the two-dimensional coordinate set includes all generated p i,j,LC2 When the additional source sub-data block index position is required, the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most In a specific example, the number of source data blocks k=8, the number of repetitions s=2, the first check number m1=1, the second check number m2=3, and the maximum number of two-dimensional coordinates is Ns=2. A feasible additional source sub-data block index table is shown in Table 4. For example, the coordinate set when i=3 and j=1 is {(0,3); (0,7)}, and the coordinates in this set include all additional source sub-data block indexes. For generating the source sub-data block linear combination value p of at least one other row in the second partial array of the i=3th check sub-block of the j=1th second check data block, i,j,LC2 ,but in, Represents bitwise exclusive OR.

[0119] Table 4 A feasible additional source sub-data block index table when k=8, s=2, m1=1, m2=3

[0120] It can be understood that the two-dimensional index coordinates in the above examples all represent the position of the additional sub-data block in the entire data array. At the same time, the two-dimensional coordinates can be converted into one-dimensional index coordinate representation, that is, there is a one-to-one correspondence between the one-dimensional coordinates and the two-dimensional index coordinates. In one feasible example, the one-dimensional coordinate index corresponding to the two-dimensional coordinate index (i1, j1) of the additional sub-data block is (g*j1+i1). In another feasible example, the one-dimensional coordinate index corresponding to the two-dimensional coordinate index (i1, j1) is (k*i+j1). Among them, i1 is an integer greater than or equal to 0 and less than g, j1 is an integer greater than 0 and less than or equal to k, and g is the number of rows in the data array.

[0121] In one embodiment, each check sub-data block in the third check data block is obtained by a linear combination of source sub-data blocks in a row of the array of the n coded data blocks, and the linear combination coefficients are obtained from the coding matrix and are not all equal to one.

[0122] In this embodiment, each syndrome sub-data block in the third parity data block is obtained by a linear combination of the source sub-data blocks in the row of the array. The linear combination coefficients are obtained from the encoding matrix and are not all equal to one. In one example, the i-th syndrome sub-data block of the j-th third parity data block is represented by p i,m1+m2+j , the source sub-data blocks corresponding to the i-th row of the data array where the i-th check sub-data block of the j-th third check data block is located are [a i,0 ,a i,1 ,...,a i,k-1 ], the linear combination coefficients corresponding to the jth third check data block are [c j,0 ,c j,1 ,...,c j,k-1 ], then the i-th syndrome data block of the third parity data block can be obtained according to but not limited to the following rules: i,m1+m2+j =c j,0 *a i,0+ c j,1 *a i,1+ ...+c j,k-1 *a i,k-1 Wherein, "+" represents a bitwise exclusive OR, "*" represents a multiplication operation under the Galois Field GF(q), j is an integer greater than 0 and less than m3, i is an integer greater than or equal to 0 and less than g, and g is the number of rows in the data array. Based on the above embodiment, generating the linear combination coefficients of the j-th third check data block includes, but is not limited to, obtaining them from the last m3 rows of the following encoding matrices: a Vandermonde matrix, a modified matrix of the Cauchy matrix, and a combination of an all-one matrix and a Cauchy matrix.

[0123] As an example, as shown in FIG6 , there are k=4 source data blocks in total, each source data block has g=4 source sub-data blocks, for ease of description, the size of each source sub-data block is one byte, which should not be a limitation of the present application, the one byte is represented by a decimal number, the number of first check data blocks m1=1, the number of second check data blocks m2=2, the number of third check data blocks m3=1, when the linear combination coefficient for generating the third check data block is obtained from the Vandermonde matrix, a feasible linear combination coefficient is expressed as [1, 8, 64, 58], and the i-th check sub-data block of the third check data block can be obtained according to the following rule: p i,m1+m2+j =c j,0 *a i,0+ c j,1 *a i,1+ ...+c j,k-1 *a i,k-1For example, the value of the 0th sub-check data block of the third check data block is 1*10+8*50+64*90+58*130=188, and the value of the 3rd sub-check data block of the third check data block is 1*40+8*80+64*120+58*160=166, where "+" represents bitwise exclusive OR and "*" represents multiplication operation under the Galois Field GF(q).

[0124] As an example, as shown in FIG7 , there are k=4 source data blocks in total, each source data block has g=4 source sub-data blocks, the size of each source sub-data block is one byte, and the one byte is represented by a decimal number. The number of first check data blocks m1=1, the number of second check data blocks m2=2, and the number of third check data blocks m3=1. When the linear combination coefficients for generating the third check data block are obtained from the deformation matrix of the Cauchy matrix, a feasible linear combination coefficient is expressed as [210, 143, 245, 200]. The i-th check sub-data block of the third check data block can be obtained according to the following rule: p i,m1+m2+j =c j,0 *a i,0+ c j,1 *a i,1+ ...+c j,k-1 *a i,k-1 For example, the value of the 0th sub-check data block of the third check data block is 210*10+143*50+245*90+200*130=77, and the value of the 3rd sub-check data block of the third check data block is 210*40+143*80+245*120+200*160=113, where "+" represents bitwise exclusive OR and "*" represents multiplication operation under the Galois Field GF(q).

[0125] As an example, as shown in FIG8 , there are k=4 source data blocks in total, each source data block has g=4 source sub-data blocks, the size of each source sub-data block is one byte, and the one byte is represented by a decimal number. The number of first check data blocks m1=1, the number of second check data blocks m2=2, and the number of third check data blocks m3=1. When the linear combination coefficients for generating the third check data block are obtained from a combination of an all-one matrix and a Cauchy matrix, a feasible linear combination coefficient is expressed as [122, 186, 71, 167]. The i-th check sub-data block of the third check data block can be obtained according to the following rule: p i,m1+m2+j =c j,0 *a i,0+ c j,1 *a i,1+ ...+c j,k-1 *a i,k-1For example, the value of the 0th sub-check data block of the third check data block is 122*10+186*50+71*90+167*130=105, and the value of the 3rd sub-check data block of the third check data block is 122*40+186*80+71*120+167*160=252, where "+" represents bitwise exclusive OR and "*" represents multiplication operation under the Galois Field GF(q).

[0126] In one embodiment, the method further includes step 130: storing n coded data blocks in n different storage nodes. By storing n coded data blocks in n different storage nodes, when a single source data block storage node fails, at least 1 / (m1+m2) of the data can be obtained from the remaining source data block storage nodes, the first parity data block storage node, and the second parity data block storage node to complete the recovery of the single source data block storage node. Furthermore, when (m1+m2) storage nodes fail and data block recovery cannot be completed using only the first parity data block and the second parity data block, the third parity data block can be further used to complete the recovery of the (m1+m2) failed storage nodes.

[0127] It can be understood that after the first processing node obtains k source data blocks, it divides the k source data blocks into a source data array of size g*k, the first check data block is the linear combination of the current row of the array and the linear combination coefficients are all 1; the second check data block is obtained by bitwise XORing the linear combination value of the current row of the array and the linear combination value of at least one source sub-data block in other rows, wherein the linear combination coefficient value of the current row of the second check data block is not all 1, and the linear combination coefficient value of at least one source sub-data block in other rows of the second check data block is all 1; the third check data block is the linear combination value of the current row of the array and the linear combination coefficient is not all 1. Since the second check data block contains additional source sub-data blocks of other rows, when a single source data block fails, the single source data block can be repaired through partial data of the first check data block and the second check data block, thereby reducing the repair bandwidth when a single storage node fails. Furthermore, when (m1+m2) coded data blocks fail and data block repair cannot be completed based solely on the first and second parity data blocks, recovery can be performed using the third parity data block. Because each parity sub-block in the third parity data block is only related to the source sub-block in the current row of the array, the amount of data recovery computation can be reduced during the data repair process. The implementation method of this application not only has local repair capabilities, but also has high repair capabilities when multiple storage nodes fail.

[0128] Figure 9 is a flowchart of a data processing method provided in one embodiment, which can be applied to a second processing node. The execution entity of this data processing method may include, but is not limited to, the second processing node 20 shown in Figure 1. Alternatively, those skilled in the art may select and set a corresponding execution entity based on the actual application scenario, and this embodiment does not limit this. As shown in Figure 9, the method provided in this embodiment includes steps 210 and 220.

[0129] In step 210 , the type and number of faulty coded data blocks and part or all of the data of at least k non-faulty coded data blocks generated by the first processing node are obtained, where k is a positive integer greater than 1.

[0130] In step 220, part or all of the data of the at least k non-faulty coded data blocks are processed according to the type and number of the faulty coded data blocks to obtain the faulty coded data blocks.

[0131] The types of the fault code data blocks include at least one of the following: source data block, first check data block, second check data block, and third check data block. The number Ne of the fault code data blocks is an integer greater than or equal to 0 and less than or equal to n.

[0132] In one embodiment, the at least k non-faulty coded data blocks are obtained by encoding k source data blocks by the first processing node according to coding parameters; the coding parameters include: a first check number, a second check number, a third check number, an additional source sub-data block index table, a repetition number, and a coding matrix.

[0133] In this embodiment, the encoded data block is obtained by the first processing node encoding the k source data blocks according to at least the following parameters after obtaining k source data blocks: a first check number m1, a second check number m2, a third check number m3, an additional source sub-data block index table, a number of repetitions s, and an encoding matrix G.

[0134] In one embodiment, obtaining part or all of the data of at least k non-faulty coded data blocks generated by the first processing node includes:

[0135] When the type of the faulty coded data block includes at least one of a first check data block, a second check data block, and a third check data block, or when the number of faulty source data blocks is greater than or equal to the sum of the first check number and the second check number, acquiring all non-faulty coded data blocks generated by the first processing node;

[0136] When only source data blocks fail and the number of failed source data blocks is less than the sum of the first checksum and the second checksum, a portion of the unfaulty coded data blocks generated by the first processing node is obtained. In one example, k = 4, m1 = 1, m2 = 3, m3 = 1, the number of repetitions s = 1, and the number of array rows g = 4. When the j = 0th source data block a0 fails, the remaining source data blocks a1, a2, a3, the first checksum, and the i = 0th sub-data block of the second checksum are obtained to complete the failed data block recovery. The obtained data only includes the portion of the unfaulty coded data blocks generated by the first processing node.

[0137] In this embodiment, when the faulty coded data block type includes one of the first check data block, the second check data block, and the third check data block, or the number of faulty source data blocks Ne is greater than or equal to (m1+m2), all non-faulty coded data blocks generated by the first processing node are obtained; when only the source data block is faulty and the number of faulty source data blocks Ne is less than (m1+m2), some non-faulty coded data blocks generated by the first processing node are obtained.

[0138] In one embodiment, processing part or all of the data of the at least k non-faulty coded data blocks according to the type and number of the faulty coded data blocks includes at least one of the following:

[0139] When the number of fault code data blocks exceeds the sum of the first check number, the second check number and the third check number, determining that data of the fault code data block is lost;

[0140] When the number of faulty coded data blocks is greater than 1 and less than or equal to the sum of the first check number, the second check number, and the third check number, and the types of the faulty coded data blocks include source data blocks and check data blocks, all non-faulty coded data blocks are acquired, and the faulty source data blocks are first restored, and then the check data blocks are restored;

[0141] When the number of faulty coded data blocks is greater than 1 and less than or equal to the sum of the first check number, the second check number, and the third check number, and when the types of the faulty coded data blocks include only source data blocks, obtaining at least a first set proportion of non-faulty coded data blocks, and restoring the faulty source data blocks, where the first set proportion is the proportion of the number of faulty coded data blocks to the sum of the first check number and the second check number;

[0142] When the number of faulty coded data blocks is greater than or equal to 1 and less than or equal to the sum of the first check number, the second check number, and the third check number, and the type of the faulty coded data blocks includes only check data blocks, all non-faulty data blocks are acquired, and the faulty source data blocks are restored;

[0143] When the number of faulty coded data blocks is equal to 1 and the faulty coded data block type only includes source data blocks, a second set ratio of non-faulty source data blocks, first verification data blocks, and second verification data blocks are respectively obtained, and the faulty source data blocks are restored. The second set ratio is the ratio of 1 to the sum of the first verification number and the second verification number.

[0144] In this embodiment, processing part or all of the data of the at least k non-faulty coded data blocks according to the type and number of the faulty coded data blocks to obtain the faulty coded data blocks includes but is not limited to at least one of the following:

[0145] Step 2201: If the number Ne of faulty coded data blocks exceeds (m1+m2m+m3), the data of the faulty coded data blocks are lost and data recovery cannot be completed;

[0146] Step 2202: If the number Ne of faulty coded data blocks is greater than 1 and less than or equal to (m1+m2+m3), and the faulty coded data block types include both source data blocks and check data blocks, all non-faulty coded data blocks are obtained, and the faulty source data blocks are first restored, followed by the check data blocks.

[0147] Step 2203: If the number of faulty coded data blocks Ne is greater than 1 and less than or equal to (m1+m2+m3), and the faulty coded data block type only includes source data blocks, obtain the remaining non-faulty data blocks in a ratio of at least Ne / (m1+m2), and then restore the faulty source data blocks.

[0148] Step 2204: If the number Ne of faulty coded data blocks is greater than or equal to 1 and less than or equal to (m1+m2+m3), and the faulty coded data block type only includes check data blocks, obtain all non-faulty data blocks, and then restore the faulty source data blocks;

[0149] Step 2205: If the number Ne of faulty coded data blocks is equal to 1, and the faulty coded data block type only includes source data blocks, obtain data amounts of at least 1 / (m1+m2) of the remaining non-faulty source data blocks, first check data blocks, and second check data blocks, respectively, and then restore the faulty source data blocks;

[0150] It can be understood that there is no order relationship between steps 2201 to S2205, which only represent at least several data processing methods that exist according to the type and quantity of the first encoded data blocks.

[0151] The present application also provides a data processing device. FIG10 is a schematic diagram of the structure of a data processing device provided by an embodiment. As shown in FIG10 , the data processing device includes:

[0152] An acquisition module 310 is configured to acquire k source data blocks, where k is a positive integer greater than 1;

[0153] The encoding module 320 is configured to encode the k source data blocks according to encoding parameters to obtain n encoded data blocks, where n is a positive integer greater than k; the encoding parameters include: a first check number, a second check number, a third check number, an additional source sub-data block index table, a number of repetitions, and a coding matrix.

[0154] In one embodiment, the n coded data blocks include k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks, where m1 is the first check number and m1=1, m2 is the second check number and is a positive integer, m3 is the third check number and is a positive integer, and n=k+m1+m2+m3;

[0155] The additional source sub-data index table is an index table with g rows and m2 columns, where g is the number of source sub-data blocks contained in each source data block;

[0156] The size of the encoding matrix is ​​m*k, where m is the total number of check data blocks, m=m1+m2+m3;

[0157] The number of repetitions is a positive integer;

[0158] The elements in the encoding matrix are elements under the Galois field GF(q), where q is an integer power of a prime number.

[0159] In one embodiment, the k source data blocks respectively correspond to data portions stored by the first k storage nodes in a stripe; the k source data blocks are distributed in k different storage nodes;

[0160] The m check data blocks obtained by encoding the k source data blocks according to the encoding parameters respectively correspond to the data parts stored in the last m storage nodes in a stripe;

[0161] The m check data blocks are respectively stored in m different storage nodes except the k different storage nodes.

[0162] In one embodiment, the number of rows of the array of n coded data blocks is determined by at least one of the following parameters: k, the first check number, and the second check number.

[0163] The number of rows of the array of n coded data blocks is equal to the number of source sub-data blocks contained in each source data block.

[0164] In one embodiment, the number of rows of the array is: Where s is the number of repetitions and is a positive integer.

[0165] In one embodiment, the encoding matrix satisfies at least one of the following:

[0166] There is a row with all element values ​​​​"1";

[0167] Any square matrix consisting of m rows and m columns is reversible under the Galois field GF(q).

[0168] In one embodiment, each check sub-data block in the first check data block is obtained by bitwise exclusive ORing the source sub-data blocks in the row of the array of the n coded data blocks.

[0169] In one embodiment, each check sub-data block in the second check data block is obtained by bitwise exclusive ORing a linear combination value of source sub-data blocks in a row of the first partial array in the array of the n coded data blocks and a linear combination value of source sub-data blocks in at least one other row in the second partial array.

[0170] In one embodiment, a linear combination value of the source sub-data blocks in the row of the first partial array corresponding to the i-th check sub-data block of the j-th second check sub-data block is determined by the source sub-data block in the i-th row of the array of the n coding data blocks and the j-th row in the coding matrix.

[0171] In one embodiment, a linear combination value of the source sub-data blocks of at least one other row in the second partial array is determined according to the additional source sub-data block index table and the number of repetitions.

[0172] In one embodiment, each element in the additional source sub-data block index table is a position coordinate set, and the number of one-dimensional coordinates or two-dimensional coordinates in the position coordinate set is at most or

[0173] In one embodiment, each element in the additional source sub-data block index table is a two-dimensional coordinate set;

[0174] In the case where the additional source sub-data block index table includes all the additional source sub-data block index positions required to generate the second verification data block, the number of coordinates in the two-dimensional coordinate set is at most

[0175] In the case where the additional source sub-data block index table includes some of the additional source sub-data block index positions required for generating the second verification data block, the number of coordinates in the two-dimensional coordinate set is at most

[0176] In one embodiment, the additional source sub-data block index table is determined according to k, the first check number, and the second check number.

[0177] In one embodiment, the elements in the tth column and the Gith row of the additional source sub-data block index table are respectively the coordinate indexes of the ith column and the Gjth row sub-data blocks in the array of n coded data blocks;

[0178] i is an integer greater than or equal to 1 and less than or equal to ceil(k / s);

[0179] The Gi is a set of row index values ​​consisting of the integer value {1, 2, ..., g} divided into (m1+m2)ceil(i / (m1+m2)) subparts in sequence, and starting from the mod(i-1,m1+m2)+1th subpart, reading one subpart every (m1+m2)th subpart; the tth element in the set GSet of the i-th source data block in the array is Gj;

[0180] The j belongs to a first set of integers, which is [floor(i / (m1+m2))+1,ceil(i / m1+m2)]\i;

[0181] t belongs to a second set of integers, the second set of integers including integers greater than or equal to 1 and less than or equal to the number of integers in the first set of integers;

[0182] The Gj is a set of row index values ​​consisting of dividing the integer value {1, 2, ..., g} into (m1+m2)ceil(j / (m1+m2)) subparts in sequence, and starting from the mod(j-1,m1+m2)+1th subpart, reading one subpart every (m1+m2) subparts;

[0183] t and j are elements corresponding to the same index size in the first integer set and the second integer set, respectively, and the elements in the first integer set and the second integer set are sorted from small to large;

[0184] g is the number of rows in the array;

[0185] ceil() means adjusting a value to the smallest integer not less than the value itself, floor() means adjusting a value to the largest integer not greater than the value itself, and mod() means modulo operation;

[0186] []\i means removing element i from a set.

[0187] In one embodiment, each check sub-data block in the third check data block is obtained by a linear combination of the source sub-data blocks in the row of the array of the n coding data blocks, and the linear combination coefficients are obtained from the n coding matrices and the linear combination coefficients are not all equal to one.

[0188] The data processing device proposed in this embodiment and the data processing method proposed in the above embodiment belong to the same inventive concept. Technical details not fully described in this embodiment can be referred to any of the above embodiments, and this embodiment has the same beneficial effects as executing the data processing method.

[0189] The present application also provides a data processing device. FIG11 is a schematic diagram of the structure of a data processing device provided by an embodiment. As shown in FIG11, the data processing device includes:

[0190] An acquisition module 410 is configured to acquire the type and quantity of the faulty coded data blocks and part or all of the data of at least k non-faulty coded data blocks generated by the first processing node, where k is a positive integer greater than 1;

[0191] The processing module 420 is configured to process part or all of the data of the at least k non-faulty coded data blocks according to the type and number of the faulty coded data blocks to obtain the faulty coded data blocks.

[0192] In one embodiment, the at least k non-faulty coded data blocks are obtained by encoding k source data blocks by the first processing node according to the encoding parameters;

[0193] The encoding parameters include: a first check number, a second check number, a third check number, an additional source sub-data block index table, a repetition number, and an encoding matrix.

[0194] In one embodiment, the acquisition module 410 is configured to:

[0195] When the type of the faulty coded data block includes at least one of a first check data block, a second check data block, and a third check data block, or when the number of faulty source data blocks is greater than or equal to the sum of the first check number and the second check number, acquiring all non-faulty coded data blocks generated by the first processing node;

[0196] In a case where only source data blocks fail and the number of failed source data blocks is less than the sum of the first check number and the second check number, some non-faulty coded data blocks generated by the first processing node are obtained.

[0197] In one embodiment, the processing module 420 is configured to be at least one of the following:

[0198] When the number of fault code data blocks exceeds the sum of the first check number, the second check number and the third check number, determining that data of the fault code data block is lost;

[0199] When the number of faulty coded data blocks is greater than 1 and less than or equal to the sum of the first check number, the second check number, and the third check number, and the types of the faulty coded data blocks include source data blocks and check data blocks, all non-faulty coded data blocks are acquired, and the faulty source data blocks are first restored, and then the check data blocks are restored;

[0200] When the number of faulty coded data blocks is greater than 1 and less than or equal to the sum of the first check number, the second check number, and the third check number, and when the type of the faulty coded data blocks includes only source data blocks, obtaining at least a first set proportion of non-faulty coded data blocks, and restoring the faulty source data blocks, where the first set proportion is the proportion of the number of faulty coded data blocks to the sum of the first check number and the second check number;

[0201] When the number of faulty coded data blocks is greater than or equal to 1 and less than or equal to the sum of the first check number, the second check number, and the third check number, and the type of the faulty coded data blocks includes only check data blocks, all non-faulty data blocks are acquired, and the faulty source data blocks are restored;

[0202] When the number of faulty coded data blocks is equal to 1 and the faulty coded data block type only includes source data blocks, a second set ratio of non-faulty source data blocks, first verification data blocks, and second verification data blocks are respectively obtained, and the faulty source data blocks are restored. The second set ratio is the ratio of 1 to the sum of the first verification number and the second verification number.

[0203] The data processing device proposed in this embodiment and the data processing method proposed in the above embodiment belong to the same inventive concept. Technical details not fully described in this embodiment can be referred to any of the above embodiments, and this embodiment has the same beneficial effects as executing the data processing method.

[0204] An embodiment of the present application also provides a processing node. Figure 12 is a schematic diagram of the hardware structure of a processing node provided by an embodiment. As shown in Figure 12, the processing node provided by the present application includes a processor 510 and a memory 520; the processor 510 in the processing node can be one or more, and Figure 12 takes one processor 510 as an example; the memory 520 is configured to store one or more programs; the one or more programs are executed by the one or more processors 510, so that the one or more processors 510 implement the data processing method as described in the embodiment of the present application.

[0205] The processing node further includes: a communication device 530 , an input device 540 and an output device 550 .

[0206] The processor 510, memory 520, communication device 530, input device 540 and output device 550 in the processing node may be connected via a bus or other means. FIG12 takes the bus connection as an example.

[0207] The input device 540 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the processing node. The output device 550 may include a display device such as a display screen.

[0208] The communication device 530 may include a receiver and a transmitter. The communication device 530 is configured to perform information transmission and reception communication according to the control of the processor 510.

[0209] The memory 520, as a computer-readable storage medium, can be configured to store software programs, computer executable programs, and modules, such as program instructions / modules corresponding to the data processing method described in the embodiment of the present application (for example, the acquisition module 310 and the encoding 320 in the data processing device). The memory 520 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the processing node, etc. In addition, the memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 520 may further include a memory remotely arranged relative to the processor 510, and these remote memories may be connected to the processing node via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0210] An embodiment of the present application also provides a processing node. Figure 13 is a schematic diagram of the hardware structure of a processing node provided by an embodiment. As shown in Figure 13, the processing node provided by the present application includes a processor 610 and a memory 620; the processor 610 in the processing node can be one or more, and Figure 13 takes one processor 610 as an example; the memory 620 is configured to store one or more programs; the one or more programs are executed by the one or more processors 610, so that the one or more processors 610 implement the data processing method as described in the embodiment of the present application.

[0211] The processing node further includes: a communication device 630 , an input device 640 and an output device 660 .

[0212] The processor 610, memory 620, communication device 630, input device 640 and output device 660 in the processing node may be connected via a bus or other means. FIG13 takes the bus connection as an example.

[0213] The input device 640 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the processing node. The output device 660 may include a display device such as a display screen.

[0214] The communication device 630 may include a receiver and a transmitter. The communication device 630 is configured to perform information transmission and reception communication according to the control of the processor 610.

[0215] The memory 620, as a computer-readable storage medium, may be configured to store software programs, computer executable programs, and modules, such as program instructions / modules corresponding to the data processing method described in the embodiments of the present application (e.g., the acquisition module 410 and the processing module 420 in the data processing device). The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; and the data storage area may store data created according to the use of the processing node, etc. In addition, the memory 620 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 620 may further include a memory remotely arranged relative to the processor 610, and these remote memories may be connected to the processing node via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0216] The present application also provides a data processing system. Figure 14 is a schematic diagram of the structure of a data processing system provided by one embodiment. As shown in Figure 14, the system includes: a storage node set 30, a first processing node 10, and a second processing node 20; the storage node set 30 includes n storage nodes 31, and the n storage nodes 31 include k source data storage nodes 32 and nk check data storage nodes 33, where n is a positive integer greater than k, and k is a positive integer greater than 1. The data processing system of this embodiment can be used to implement the data processing method of any of the above embodiments.

[0217] The embodiment of the present application also provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the data processing method described in any one of the embodiments of the present application. The data processing method includes: obtaining k source data blocks, k is a positive integer greater than 1; encoding the k source data blocks according to the encoding parameters to obtain n encoded data blocks, n is a positive integer greater than k; the encoding parameters include: a first check number, a second check number, a third check number, an additional source sub-data block index table, a number of repetitions, and a coding matrix. Alternatively, the data processing method includes: obtaining the type and number of faulty encoded data blocks and part or all of the data of at least k non-faulty encoded data blocks generated by the first processing node, k is a positive integer greater than 1; processing part or all of the data of the at least k non-faulty encoded data blocks according to the type and number of the faulty encoded data blocks to obtain a faulty encoded data block.

[0218] The computer storage medium of the embodiment of the present application can adopt any combination of one or more computer-readable media.Computer-readable media can be computer-readable signal media or computer-readable storage media.Computer-readable storage media can be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above.More specific examples (non-exhaustive list) of computer-readable storage media include: electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM), flash memories, optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.Computer-readable storage media can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0219] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0220] The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wire, optical cable, radio frequency (RF), etc., or any suitable combination of the foregoing.

[0221] The computer program code for performing the operations of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0222] The above description is merely an exemplary embodiment of the present application and is not intended to limit the scope of protection of the present application.

[0223] It will be understood by those skilled in the art that the term user terminal covers any suitable type of wireless user equipment, such as a mobile phone, a portable data processor, a portable web browser or a vehicle-mounted mobile station.

[0224] In general, various embodiments of the present application may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, although the present application is not limited thereto.

[0225] Embodiments of the present application may be implemented by executing computer program instructions by a data processor of a mobile device, for example, in a processor entity, or by hardware, or by a combination of software and hardware. The computer program instructions may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages.

[0226] The block diagram of any logical flow in the drawings of this application may represent program steps, or may represent interconnected logical circuits, modules and functions, or may represent a combination of program steps and logical circuits, modules and functions. A computer program may be stored on a memory. The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as but not limited to read-only memory (ROM), random access memory (RAM), optical storage devices and systems (digital versatile discs (DVD) or compact disks (CD), etc.). Computer-readable media may include non-transitory storage media. The data processor may be of any type suitable for the local technical environment, such as but not limited to a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and a processor based on a multi-core processor architecture.

[0227] The above description of exemplary embodiments of the present application has been provided by way of exemplary and non-limiting examples. However, various modifications and adaptations of the above embodiments will be apparent to those skilled in the art, when considered in conjunction with the accompanying drawings and the appended claims, without departing from the scope of the present application. Therefore, the proper scope of the present application will be determined by reference to the appended claims.

Claims

1. A data processing method, applied to a first processing node, comprising: Obtaining k source data blocks, where k is a positive integer greater than 1; Encoding the k source data blocks according to encoding parameters to obtain n encoded data blocks, where n is a positive integer greater than k; The encoding parameters include: a first check quantity, a second check quantity, a third check quantity, an additional source sub-data block index table, a repetition number, and an encoding matrix.

2. The method according to claim 1, wherein, Among the n encoded data blocks, there are k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks, where m1 is the first check quantity and m1 = 1, m2 is the second check quantity and is a positive integer, m3 is the third check quantity and is a positive integer, and n = k + m1 + m2 + m3; The additional source sub-data index table is an index table with g rows and m2 columns, where g is the number of source sub-data blocks included in each source data block; The size of the encoding matrix is m * k, where m is the total number of check data blocks, and m = m1 + m2 + m3; The repetition number is a positive integer; The elements in the encoding matrix are elements in the Galois field GF(q), where q is an integer power of a prime number.

3. The method according to claim 2, wherein, The k source data blocks respectively correspond to the data parts stored in the first k storage nodes in a stripe; The k source data blocks are distributed in k different storage nodes; The m check data blocks obtained by encoding the k source data blocks according to the encoding parameters respectively correspond to the data parts stored in the last m storage nodes in a stripe; The m check data blocks are respectively stored in m different storage nodes other than the k different storage nodes.

4. The method according to claim 2, wherein, The number of rows of the array of the n encoded data blocks is determined by at least one of the following parameters: k, the first check quantity, the second check quantity; The number of rows of the array of the n encoded data blocks is equal to the number of source sub-data blocks included in each source data block.

5. The method according to claim 4, wherein The number of rows of the array is: where s is the number of repetitions and s is a positive integer.

6. The method according to claim 2, wherein The encoding matrix satisfies at least one of the following: There is a row with all element values being "1"; Any square matrix formed by m rows and m columns is invertible in the Galois field GF(q).

7. The method according to claim 2, wherein, Each check sub-data block in the first check data block is obtained by bitwise XOR of the source sub-data blocks in the row where the array of the n encoded data blocks is located.

8. The method according to claim 2, wherein Each check sub-data block in the second check data block is obtained by bitwise XOR of the linear combination value of the source sub-data blocks in the row where the first part of the array of the n encoded data blocks is located and the linear combination value of the source sub-data blocks in at least one other row in the second part of the array.

9. The method according to claim 8, wherein The linear combination value of the source sub-data blocks in the row where the first part of the array is located corresponding to the i-th check sub-data block of the j-th second check data block is determined by the source sub-data blocks in the i-th row of the array of the n encoded data blocks and the j-th row in the encoding matrix.

10. The method according to claim 8, wherein, The linear combination value of the source sub-data blocks in at least one other row in the second part of the array is determined according to the additional source sub-data block index table and the repetition number.

11. The method according to claim 2, wherein, Each element in the additional source sub-data block index table is a set of position coordinates, and the maximum number of one-dimensional coordinates or two-dimensional coordinates in the set of position coordinates is or 12. The method according to claim 11, wherein, Each element in the additional source sub-data block index table is a set of two-dimensional coordinates; When all the index positions of the additional source sub-data blocks required to generate the second check data block are included in the additional source sub-data block index table, the maximum number of coordinates in the two-dimensional coordinate set is When the additional source sub-data block index table includes the index positions of the additional source sub-data blocks required to partially generate the second check data block, the maximum number of coordinates in the two-dimensional coordinate set is 13. The method according to claim 2, wherein, The additional source sub-data block index table is determined according to k, m1, and m2.

14. The method according to claim 13, wherein, The t-th column and the Gi-th row of the additional source sub-data block index table respectively contain the coordinate indexes of the sub-data blocks in the i-th column and the Gj-th row of the array of the n coded data blocks; where i is an integer greater than or equal to 1 and less than or equal to ceil(k / s); The Gi is formed by sequentially dividing the set of integer values {1, 2,..., g} into (m1 + m2) ceil(i / (m1+m2)) sub - parts, and starting from the (mod(i - 1, m1 + m2)+1)-th sub - part in sequence, reading one sub - part every (m1 + m2) sub - parts to form a set of row index values; the t - th element in the set GSet of the i - th source data block in the array is Gj; j belongs to the first integer set, and the first integer set is [floor(i / (m1 + m2)) + 1, ceil(i / (m1 + m2))]\i; t belongs to the second integer set, and the second integer set contains integers greater than or equal to 1 and less than or equal to the number of integers in the first integer set; The Gj is a set of row index values formed by sequentially dividing the integer values {1, 2,..., g} into (m1 + m2) ceil(j / (m1+m2)) sub - parts, and starting from the (mod(j - 1, m1 + m2)+1)-th sub - part in order, reading one sub - part every (m1 + m2) sub - parts. t and j are respectively elements corresponding to the same index size in the first integer set and the second integer set to which they belong, and the elements in the first integer set and the second integer set are sorted in ascending order; g is the number of rows of the array; ceil() means adjusting a numerical value to the smallest integer not less than the numerical value itself, floor() means adjusting a numerical value to the largest integer not greater than the numerical value itself, and mod() means taking the modulo operation; []\i means removing the element i from a set.

15. The method according to claim 2, wherein, Each check sub-data block in the third check data block is obtained by a linear combination of the source sub-data blocks in the row where the array of the n coded data blocks is located, and the linear combination coefficients are obtained from the n coding matrices and the linear combination coefficients are not all equal to one.

16. A data processing method, applied to a second processing node, includes: Obtaining the types and quantities of the faulty coded data blocks and part or all of the data of at least k non-faulty coded data blocks generated by a first processing node, where k is a positive integer greater than 1; Processing part or all of the data of the at least k non-faulty coded data blocks according to the types and quantities of the faulty coded data blocks to obtain the faulty coded data blocks.

17. The method according to claim 16, wherein The at least k non-faulty coded data blocks are obtained by the first processing node encoding k source data blocks according to encoding parameters; The encoding parameters include: the first check quantity, the second check quantity, the third check quantity, the additional source sub-data block index table, the number of repetitions, and the coding matrix.

18. The method according to claim 17, wherein, The obtaining part or all of the data of at least k non-faulty coded data blocks generated by the first processing node includes: When the type of the faulty coded data blocks includes at least one of the first check data block, the second check data block, and the third check data block, or when the number of faulty source data blocks is greater than or equal to the sum of the first check quantity and the second check quantity, obtaining all the non-faulty coded data blocks generated by the first processing node; When only source data blocks are faulty and the number of faulty source data blocks is less than the sum of the first check quantity and the second check quantity, obtaining part of the non-faulty coded data blocks generated by the first processing node.

19. The method according to claim 17, wherein, The processing part or all of the data of the at least k non-faulty coded data blocks according to the types and quantities of the faulty coded data blocks includes at least one of the following: In the case where the number of the fault coding data blocks exceeds the sum of the first check number, the second check number, and the third check number, determine that data loss occurs in the fault coding data blocks; In the case where the number of the fault coding data blocks is greater than 1, less than or equal to the sum of the first check number, the second check number, and the third check number, and the types of the fault coding data blocks include source data blocks and check data blocks, obtain all the non-fault coding data blocks, first recover the fault source data blocks, and then recover the check data blocks; In the case where the number of the fault coding data blocks is greater than 1, less than or equal to the sum of the first check number, the second check number, and the third check number, and the types of the fault coding data blocks only include source data blocks, obtain at least a first set ratio of the non-fault coding data blocks, and recover the fault source data blocks, where the first set ratio is the ratio of the number of the fault coding data blocks in the sum of the first check number and the second check number; In the case where the number of the fault coding data blocks is greater than or equal to 1, less than or equal to the sum of the first check number, the second check number, and the third check number, and the types of the fault coding data blocks only include check data blocks, obtain all the non-fault data blocks, and recover the fault source data blocks; In the case where the number of the fault coding data blocks is equal to 1 and the types of the fault coding data blocks only include source data blocks, respectively obtain a second set ratio of the non-fault source data blocks, the first check data blocks, and the second check data blocks, and recover the fault source data blocks, where the second set ratio is the ratio of 1 in the sum of the first check number and the second check number.

20. A processing node, comprising: A memory, and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method according to any one of claims 1-15.

21. A second processing node, comprising: A memory, and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method according to any one of claims 16-19.

22. A data processing system, comprising: A set of storage nodes, a first processing node according to claim 20, and a second processing node according to claim 21; The set of storage nodes includes n storage nodes, the n storage nodes include k source data storage nodes and n-k check data storage nodes, n is a positive integer greater than k, and k is a positive integer greater than 1.

23. A computer-readable storage medium having a computer program stored thereon, wherein, When the program is executed by a processor, it implements the data processing method according to any one of claims 1-19.

Citation Information

Patent Citations

  • Data processing method and device based on erasure codes

    CN111090540A

  • Data storage method, system and equipment and medium

    CN114281270A

  • Data storage method, system and equipment and medium

    CN114996047A

  • Data processing method and device, electronic equipment and storage medium

    CN115421975A

  • Storage controller, data processing chip, and data processing method

    WO2018165943A1