Data processing method, processing node, data processing system and storage medium

By generating multiple encoded data blocks and distribute storage in storage nodes, large-scale erasure coding technology solves the problem of high bandwidth repair, and achieves data storage effects with low redundancy ratio and high concurrent reading and writing.

CN120256190APending Publication Date: 2025-07-04ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410012448.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing large-scale erasure coding technology requires the transmission of a large amount of data during data repair, resulting in high repair bandwidth and difficult to effectively apply in storage networks.

Method used

A data processing method is adopted to generate multiple encoded data blocks through encoding parameters, including a first verification data block, a second verification data block and a third verification data block, using these data blocks to reduce the repair bandwidth during data repair and distribute storage in the storage node to achieve low redundancy ratio and high concurrent read and write performance.

Benefits of technology

Data repair with low repair bandwidth is achieved in wide stripe situations, improving the reliability of data storage and concurrent read and write performance, and reducing the storage overhead of redundant verification data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256190A_ABST
    Figure CN120256190A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method, a processing node, a data processing system and a storage medium. The method comprises the following steps: acquiring k source data blocks, wherein k is a positive integer greater than 1; coding the k source data blocks according to coding parameters to obtain n coded data blocks, wherein n is a positive integer greater than k; the coding parameters comprise a first check number, a second check number, a third check number, an additional source sub-data block index table, a repetition number and a coding matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, for example, to a data processing method, a processing node, a data processing system, and a storage medium. Background Art

[0002] With the rapid development of information technology, the amount of data to be stored has shown an explosive growth. How to store and obtain this data safely, reliably, and efficiently is an issue that commercial companies and even a country must consider. The amount of data to be stored in the future will increase year by year. If a relatively small proportion of erasure codes is used, redundant check data will occupy a large amount of storage space, resulting in an increase in storage costs. Therefore, the large-proportion erasure code (EC) technology is an important research direction in future storage EC. The large-proportion EC technology can use wider stripes and a lower check ratio, thereby improving storage efficiency. However, if the EC ratio is directly increased based on the current erasure code, a large amount of data repair bandwidth will be generated when repairing one or more source data nodes. That is, when one or more source data nodes are in error, it is necessary to transmit k times the data volume of a single source data node in the network. The data repair occupying a large amount of data transmission network bandwidth is unacceptable in the storage network. Summary of the Invention

[0003] This application provides a data processing method, a processing node, a data processing system, and a storage medium.

[0004] An embodiment of this application provides a data processing method, which is applied to a first processing node and includes:

[0005] Obtain k source data blocks, where k is a positive integer greater than 1;

[0006] Encode the k source data blocks according to encoding parameters to obtain n encoded data blocks, where n is a positive integer greater than k;

[0007] The encoding parameters include: the first check quantity, the second check quantity, the third check quantity, an additional source sub-data block index table, the number of repetitions, and an encoding matrix.

[0008] An embodiment of this application also provides a data processing method, which is applied to a second processing node and includes:

[0009] Obtain the type and quantity of faulty encoded data blocks and part or all of the data of at least k non-faulty encoded data blocks generated by the first processing node, where k is a positive integer greater than 1;

[0010] Process part or all of the data of the at least k non-faulty encoded data blocks according to the type and quantity of the faulty encoded data blocks to obtain the faulty encoded data blocks.

[0011] The embodiment of the present application further provides a processing node, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the above data processing method is implemented.

[0012] The embodiment of the present application further provides a data processing system, including: a set of storage nodes, a first processing node, and a second processing node; the set of storage nodes includes n storage nodes, and the n storage nodes include k source data storage nodes and n - k parity data storage nodes, where n is a positive integer greater than k, and k is a positive integer greater than 1.

[0013] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above data processing method is implemented. Description of the Drawings

[0014] Figure 1 Schematic diagram of an implementation environment of a data processing method provided for an embodiment;

[0015] Figure 2 Schematic diagram of a set of storage nodes provided for an embodiment;

[0016] Figure 3 Flowchart of a data processing method provided for an embodiment;

[0017] Figure 4 Schematic diagram of an array corresponding to n encoded data blocks provided for an embodiment;

[0018] Figure 5 Schematic diagram of generating a first parity data block provided for an embodiment;

[0019] Figure 6 Schematic diagram of generating a third parity data block when the linear combination coefficients are obtained from a Vandermonde matrix provided for an embodiment;

[0020] Figure 7 Schematic diagram of generating a third parity data block when the linear combination coefficients are obtained from a deformed matrix of a Cauchy matrix provided for an embodiment;

[0021] Figure 8 Schematic diagram of another way of generating a third parity data block when the linear combination coefficients are obtained from a deformed matrix of a Cauchy matrix provided for an embodiment;

[0022] Figure 9 Flowchart of another data processing method provided for an embodiment;

[0023] Figure 10 Schematic diagram of the structure of a data processing device provided for an embodiment;

[0024] Figure 11 Schematic structural diagram of another data processing device provided for an embodiment;

[0025] Figure 12 Schematic hardware structure diagram of a first processing node provided for an embodiment;

[0026] Figure 13 Schematic hardware structure diagram of a second processing node provided for an embodiment;

[0027] Figure 14 Schematic structural diagram of a data processing system provided for an embodiment. Detailed implementation manners

[0028] The present application will be described below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily combined with each other. Additionally, it should be noted that, for the convenience of description, only the parts related to the present application rather than all the structures are shown in the drawings.

[0029] It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that in the flowchart. The terms "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0030] In a data storage system, data reliability and persistence are usually ensured by adding redundancy. Currently, the commonly used methods are: multiple copies and erasure codes.

[0031] The multiple - copy technology refers to copying the same data n times, and then distributing and storing the n copies of data in different nodes according to a certain strategy. The nodes include, but are not limited to, devices with storage functions such as disks, solid - state drives, servers, etc. These nodes are independent. When at most n - 1 nodes fail and the data in them fails, the same data as that in the failed nodes can still be obtained from the non - failed nodes, thereby providing data protection. The advantage of using the multiple - copy technology is that it does not require complex operations, is easy to operate, and because there is no complex multiplication operation, the data read - write speed is fast. However, the storage space utilization rate of the multiple - copy technology is low. For future massive data storage, the storage cost will increase exponentially. Therefore, multiple copies are usually used in hot data.

[0032] Erasure code is a forward error correction technique, mainly used in network transmission to avoid packet loss, and it is usually used in storage systems to improve storage reliability. Compared with multiple replicas, erasure code can achieve higher data reliability with less redundancy. However, the encoding and decoding of erasure code are relatively complex, involving multiplication operations in finite fields, so it requires more computing resources and is generally applied to warm data or cold data in storage systems. The most widely used erasure code in storage systems is the RS (Reed-Solomon) type. However, in large-scale EC, the decoding operation of RS-type erasure code is computationally intensive and the repair bandwidth is high, that is, when repairing a single data node, k times the amount of data of a single data node needs to be transmitted in the network. To solve the problem of high repair bandwidth, the concepts of locally repairable codes (LRC) and regeneration codes (RGC) have been proposed. Locally repairable codes have good repair locality. When a single source data node fails, only the local area where the source data node is located needs to be repaired, reducing the repair bandwidth. However, locally repairable codes are not maximum distance separable (MDS) codes. Compared with RS codes, more check data storage space is used under the same repair ability; regeneration codes have a smaller repair bandwidth overhead compared with RS codes. When a single source data node fails, only a part of the data of the remaining data nodes needs to be transmitted to complete the repair of the source data node. However, regeneration codes also have problems such as the number of source data nodes k cannot be set too large, high decoding complexity, and difficulty in meeting the MDS characteristics. The above problems make it difficult for regeneration codes to be applied in engineering when meeting large-scale EC.

[0033] The embodiment of the present application provides a data processing method, which can complete the repair of the source data node with low repair bandwidth under the conditions of wide stripes and low redundancy ratio, and improve the reliability and concurrent read / write performance of data storage.

[0034] Figure 1 It is a schematic diagram of the implementation environment of a data processing method provided by an embodiment. As Figure 1 shown, the implementation environment includes but is not limited to: the first transmission node 101, the second transmission node 201, the network 102 and the network 202, the first processing node 10, the second processing node 20, and the storage node set 30, and the storage node set contains at least one storage node 301.

[0035] In one embodiment, the first transmission node 101 and the second transmission node 201 can be the same device or different devices. The first transmission node 101 and the second transmission node 201 can include, but are not limited to: a server, a base station, a phone, a personal computer, a laptop, a tablet computer, an access point (AP), or other devices with data transmission and storage functions.

[0036] In one embodiment, the relative positions, quantities, etc. of the first transmission node 101 and the second transmission node 201 can be different in specific application scenarios. For example, regarding the relative positions: the first transmission node 101 can send data through the network, and the second transmission node 201 can receive data through the network; at the same time, the second transmission node 201 can send data through the network, and the first transmission node 101 can receive data through the network. Regarding the quantities: the quantities of the first transmission node 101 and the second transmission node 201 can include one or more, and they are connected through the network.

[0037] In one embodiment, the network 102 and the network 202 can include, but are not limited to: a wired network, a wireless network, a cellular network, a local area network, the Internet, a wide area network, the World Wide Web, or other networks for data transmission.

[0038] In one embodiment, the first processing node 10 and the second processing node 20 may be the same processing node of the same device at different times. The first processing node 10 and the second processing node 20 may access and transfer data through a network and process the data according to configuration items. In one example, the configuration items include, but are not limited to: the source storage node location for accessing data, the destination storage node location for transferring data, the number of parity data blocks generated by encoding data blocks, the number of first parity data blocks generated by encoding data blocks, the number of second parity data blocks generated by encoding data blocks, the number of second parity data blocks generated by encoding data blocks, the number of source data blocks, the size of source data blocks, the number of storage nodes, the number and location of faulty storage nodes, the data recovery strategy, or other parameters for data processing. In one example, data processing includes, but is not limited to: encoding source data blocks, allocating data blocks to storage nodes, accessing data blocks from storage nodes, recovering faulty data blocks, or other processing methods for data block reading, writing, deleting, modifying, etc.

[0039] In one embodiment, the execution subject of the data processing method in this embodiment may include, but is not limited to, Figure 1 the first processing node 10 and the second processing node 20 in the illustrated embodiment, or those skilled in the art may select and set the corresponding execution subject according to the actual application scenario, which is not limited in this embodiment, but should not be construed as a limitation to the embodiments of the present application.

[0040] In one embodiment, the storage node set 30 includes at least n storage nodes 301, and each storage node 301 may include, but is not limited to: RAM, ROM, EEPROM, flash memory or other memory technologies, or other optical disc storage, magnetic cassette, magnetic tape, disk storage, hard disk, solid state drive or other storage devices, or any other medium that can be used to store desired information and can be accessed by a computer.

[0041] Figure 2 A schematic diagram of a storage node set provided for one embodiment. As Figure 2 shown, the storage node set 30 contains n storage nodes 301, including k source data storage nodes and m = n - k parity data storage nodes, where n is a positive integer greater than k, and k is a positive integer greater than 1; each storage node contains p Chunk data blocks 302, where p is an integer greater than 0; each Chunk data block contains w Page data blocks 303, where w is an integer greater than 0; each stripe 304 contains e * n Page data blocks, and every e Page data blocks are distributed in one storage node, that is, in each stripe, each storage node contains e Page data blocks, and e is an integer greater than 0.

[0042] In one embodiment, the k source data blocks obtained respectively correspond to the data portions stored in the first k storage nodes in a stripe, and the k source data blocks are distributed among k different storage nodes; the m parity data blocks obtained by encoding the k source data blocks according to the encoding parameters respectively correspond to the data portions stored in the last m storage nodes in a stripe, and the m parity data blocks are respectively stored in the remaining m different storage nodes.

[0043] In one embodiment, the portion containing e Page data blocks in each storage node of each stripe is regarded as a source data block, and the k source data blocks constitute the source data portion of a stripe, and the m parity data blocks obtained by encoding the k source data blocks constitute the parity data portion of a stripe.

[0044] Figure 3 A flowchart of a data processing method provided for an embodiment, and this method can be applied to a first processing node. The execution subject of this data processing method may include, but is not limited to, Figure 1 the first processing node 10 shown as follows, or those skilled in the art can select and set the corresponding execution subject according to the actual application scenario, and this embodiment does not limit this. As Figure 3 shown, the method provided in this embodiment includes step 110 and step 120.

[0045] In step 110, k source data blocks are obtained, where k is a positive integer greater than 1.

[0046] In step 120, n encoded data blocks are obtained by encoding the k source data blocks according to the encoding parameters, where n is a positive integer greater than k.

[0047] In this embodiment, the encoding parameters include: the first parity quantity, the second parity quantity, the third parity quantity, the additional source sub-data block index table, the repetition times, and the encoding matrix.

[0048] The source data blocks can be selected and set according to the specific application scenario, such as obtaining from a file, etc., and this embodiment does not limit this.

[0049] In this embodiment, when the first processing node obtains k source data blocks, n encoded data blocks are obtained by encoding the k source data blocks according to at least the following parameters: the first check quantity, the second check quantity, the third check quantity, the additional source sub-data block index table, the repetition times, and the encoding matrix. When one or more nodes fail, the second processing node can obtain the types and quantities of the failed encoded data blocks, and part or all of the data of at least k unfailed encoded data blocks generated by the first processing node, and process part or all of the data of at least k unfailed encoded data blocks according to the types and quantities of the failed encoded data blocks to obtain the failed encoded data blocks. Among them, the existence of the first check data block and the second check data block enables the repair of the source data blocks to be completed with a low repair bandwidth when a single or multiple source data nodes fail. The existence of the third check data block enables the repair of the source data blocks to be completed using the third check data block when the first check data block and the second check data block cannot complete the data repair. That is, the repair bandwidth is reduced by the first check data block and the second check data block, and the repair ability is ensured by the third check data block.

[0050] The encoding method of this embodiment is applicable to wide stripe repair, can use wider stripes to improve concurrent read and write performance, and has a lower redundancy ratio to reduce the storage overhead of redundant check data, with the characteristics of low repair bandwidth, high repair ability, and high concurrent reading of wide EC stripes. In addition, when a single source data node fails, the erasure code method of this application can complete data repair at the optimal repair bandwidth or close to the optimal repair bandwidth level, and r check data nodes can repair all types of failures of r - 1 data nodes (including source data nodes and check data nodes), as well as most types of failures of r data nodes.

[0051] In one embodiment, the n encoded data blocks include k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks, where m1 is the first check quantity and m1 = 1, m2 is the second check quantity and is a positive integer, m3 is the third check quantity and is a positive integer, and n = k + m1 + m2 + m3;

[0052] The additional source sub-data index table is an index table with g rows and m2 columns, where g is the number of source sub-data blocks included in each source data block;

[0053] The size of the encoding matrix is m * k, where m is the total number of check data blocks, and m = m1 + m2 + m3;

[0054] The repetition times are a positive integer;

[0055] The elements in the encoding matrix are elements in the Galois field GF(q), where q is an integer power of a prime number.

[0056] In this embodiment, among the n encoded data blocks, there are k source data blocks, m1 first parity data blocks, m2 second parity data blocks, and m3 third parity data blocks, where n = k + m1 + m2 + m3. Additionally, m1 = 1; m2 is an integer greater than 0; m3 is an integer greater than 0. The additional source sub-data index table is an index table of size g * m2, which is used to indicate the positions of the additional source sub-data blocks required to generate the second parity data blocks; the repetition number s is an integer greater than 0, which is used to obtain the number of source sub-data blocks included in each source data block; the encoding matrix is a matrix of size (m1 + m2 + m3) * k, which is used to obtain the linear combination coefficients for generating the first parity data blocks, second parity data blocks, and third parity data blocks, and the elements in the matrix are elements in the Galois field GF(q), where q is an integer power of a prime number.

[0057] Figure 4 FIG. is a schematic diagram of an array corresponding to n encoded data blocks provided for an embodiment. As an example, the obtained n encoded data blocks include: k source data blocks, m1 first parity data blocks, m2 second parity data blocks, and m3 third parity data blocks. The schematic diagrams of the k source data blocks, m1 first parity data blocks, m2 second parity data blocks, and m3 third parity data blocks in the array are as Figure 4 shown. In the data array 1200 of the n encoded data packets obtained by encoding, there are k source data blocks 1201, and each source data block includes g source sub-data blocks 1205. In one example, a source sub-data block is a 0,0, ; the first parity data block 1202 includes g parity sub-data blocks 1206; among the m2 second parity data blocks 1203, each second parity data block includes g parity sub-data blocks; among the m3 third parity data blocks 1204, each third parity data block includes g parity sub-data blocks; wherein, the size of each source sub-data block is equal to the size of each parity sub-data block.

[0058] In one embodiment, the k source data blocks respectively correspond to the data parts stored in the first k storage nodes in a stripe; the k source data blocks are distributed in k different storage nodes; the m parity data blocks obtained by encoding the k source data blocks according to the encoding parameters respectively correspond to the data parts stored in the last m storage nodes in a stripe; the m parity data blocks are respectively stored in m different storage nodes other than the k different storage nodes.

[0059] In one embodiment, the k source data blocks belong to a stripe and correspond to the source data part of a stripe. The k source data blocks are represented as a = [a0, a1,..., a k-1 , and each source data block is represented as a i, where \(0\leq i\leq k - 1\), each source data block contains \(T\) kilobytes, \(T\) is an integer greater than \(0\), and \(T\) includes but is not limited to one of the following: \(4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096\).

[0060] In one embodiment, each source data block \(a\) i further includes \(g\) source sub - data blocks, and the \(i\) - th source data block is denoted as \(a\) i =\([a\) 0,i , \(a\) 1,i ,\(\cdots\), \(a\) g-1,i , where \(g\) is an integer greater than \(1\). Each source sub - data block contains \(E\) kilobytes, \(E\) is a positive real number and the product of \(E\) and \(1024\) is a positive integer, and \(E\) includes but is not limited to one of the following: \(\frac{1}{2}, 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024\).

[0061] In one embodiment, the number of rows of the array of the \(n\) encoded data blocks is determined by at least one of the following parameters: \(k\), the first check quantity, the second check quantity. Wherein, the number of rows of the array of the \(n\) encoded data blocks is equal to the number of source sub - data blocks contained in each source data block.

[0062] In one embodiment, the number of rows of the array is: where \(s\) is the number of repetitions, and \(s\) is a positive integer.

[0063] As an example, the number of rows \(g\) of the array is determined according to the following rule: where \(s\) represents the number of repetitions, and \(s\) is an integer greater than \(0\).

[0064] In one example, the number of rows \(g\) of the array is determined according to the rule with the first check quantity \(m1 = 1\), the second check quantity \(m2 = 3\), the number of source data blocks \(k = 20\), and the number of repetitions \(s = 1\), then the number of rows \(g = 1024\) of the array is obtained.

[0065] In another example, the number of rows \(g\) of the array is determined according to the rule with the first check quantity \(m1 = 1\), the second check quantity \(m2 = 3\), the number of source data blocks \(k = 20\), and the number of repetitions \(s = 2\), then the number of rows \(g = 64\) of the array is obtained.

[0066] In another example, the number of rows \(g\) of the array is determined according to the rule with the first check quantity \(m1 = 1\), the second check quantity \(m2 = 3\), the number of source data blocks \(k = 160\), and the number of repetitions \(s = 8\), then the number of rows \(g = 1024\) of the array is obtained.

[0067] In another example, the number of rows \(g\) of the array is determined according to the rule It is determined that the first verification quantity m1 = 1, the second verification quantity m2 = 3, the number of source data blocks k = 150, and the number of repetitions s = 8, then the number of rows g of the obtained array is 1024.

[0068] In one embodiment, the encoding matrix satisfies at least one of the following:

[0069] There is a row with all element values being "1";

[0070] The square matrix formed by any m rows and m columns is invertible in the Galois field GF(q).

[0071] In this embodiment, the size of the encoding matrix is m*k, and the element values in the encoding matrix are element values in the Galois field GF(q). The encoding matrix has at least one of the following characteristics: there is a row with all element values being "1", and the square matrix formed by any m rows and m columns is invertible in the Galois field GF(q). Among them, having a row with element values being "1" is for generating the first verification data block. Since each parity-check sub-data block of the first verification data block is the bitwise XOR of the k source sub-data blocks in the current row, the encoding coefficients of the k source sub-data blocks are all 1, that is, p i,0 = 1*a i,0 + 1*a i,1 +...+ 1*a i,k-1 , and the remaining (m - 1) = m1 + m2 rows of the encoding matrix are used to generate the second verification data block and the third verification data block. The square matrix formed by any m rows and m columns of the encoding matrix is invertible in the Galois field GF(q), so that the encoding method can approximate the MDS code as much as possible; where q is an integer power of a prime number. In one example, q is an integer power of 2, and q includes but is not limited to one of the following: 256, 128, 64, 32, 16, 8, 4, 2.

[0072] In one embodiment, the encoding matrix may include but is not limited to: a Vandermonde matrix, a deformed matrix of a Cauchy matrix, a combination of an all-1 matrix and a Cauchy matrix.

[0073] In one example, when the encoding matrix is a Vandermonde matrix, the encoding matrix G is expressed as:

[0074]

[0075] Among them, α is a primitive element in the Galois field GF(q), and i1, i2,..., i m are all integers greater than or equal to 0 and less than or equal to q - 2, and the values of i1, i2,..., i m are all different from each other, and there is exactly one value among i1, i2,..., i m that is equal to 0. In a specific example, [i1, i2,..., i m= [0, 1, ..., m - 1], m = 4, k = 4, q = 256, then the encoding matrix G is:

[0076]

[0077] In another example, when the step encoding matrix is a deformed matrix of the Cauchy matrix, the encoding matrix is expressed as:

[0078]

[0079] Where α is a primitive element in the Galois field GF(q), "+" represents bitwise exclusive OR, and " / " represents the division operation in the Galois field GF(q). j1, j2, ..., j m-1 are all integers greater than or equal to k and less than or equal to q - 2, and the values of j1, j2, ..., j m-1 are all different from each other. In a specific example, [j1, j2, ..., j m-1 = [k, k + 1, ..., k + m - 2], m = 4, k = 4, q = 256, then the encoding matrix G is:

[0080]

[0081] In another different example, when the step encoding matrix is a combination of an all - 1 matrix and a Cauchy matrix, the encoding matrix is expressed as:

[0082]

[0083] Where α is a primitive element in the Galois field GF(q), "+" represents bitwise exclusive OR, and " / " represents the division operation in the Galois field GF(q). j1, j2, ..., j m-1 are all integers greater than or equal to k and less than or equal to q - 2, and the values of j1, j2, ..., j m-1 are all different from each other. In a specific example, [j1, j2, ..., j m-1 = [k, k + 1, ..., k + m - 2], m = 4, k = 4, q = 256, then the encoding matrix G is:

[0084]

[0085] It should be noted that in this application, the row index of the encoding matrix starts from subscript 0 and increases.

[0086] In one embodiment, each parity - check sub - data block in the first parity - check data block is obtained by bitwise exclusive OR of the source sub - data blocks in the row where the array of the n encoding data blocks is located.

[0087] In one example, the i-th parity sub-data block of the first parity data block is denoted as p i,0 , and the source sub-data blocks corresponding to the i-th row where the i-th parity sub-data block of the first parity data block is located are [a i,0 , a i,1 ,..., a i,k-1 . Then, the i-th parity sub-data block of the first parity data block can be obtained according to the following formula: p i,0 = a i,0 + a i,1 +... + a i,k-1 , where "+" represents bitwise XOR.

[0088] In one example, as Figure 5 shown, there are a total of k = 4 source data blocks, and each source data block has g = 4 source sub-data blocks. For the sake of description, the size of each source sub-data block is one byte, which should not be a limitation of this application. The one byte is represented by a decimal number. The sub-parity data blocks in the first parity data block are obtained by bitwise XOR of the source sub-data blocks in the current row. For example, the value of the 0-th sub-parity data block in the first parity data block is 10 + 50 + 90 + 130 = 216, where "+" represents bitwise XOR.

[0089] In one embodiment, each parity sub-data block in the second parity data block is obtained by bitwise XOR of the linear combination value of the source sub-data blocks in the first part of the array in the n-by-array of coded data blocks and the linear combination value of the source sub-data blocks in at least one other row of the second part of the array.

[0090] In one example, the i-th parity sub-data block of the j-th second parity data block is denoted as p i,j , i is an integer and satisfies 0 ≤ i ≤ g - 1, j is an integer and satisfies 1 ≤ j ≤ m2; the linear combination value of the source sub-data blocks in the i-th row of the first part of the array is denoted as p i,j,LC1 ; the linear combination value of the source sub-data blocks in at least one other row of the second part is denoted as p i,j,LC2 ; LC1 and LC2 respectively represent the above corresponding linear combination identifiers of different types. The above values satisfy the following relationship: p i,j = p i,j,LC1 + p i,j,LC2 . Where "+" represents bitwise XOR. In the embodiments of this application, the source sub-data blocks in the above other rows are uniformly referred to as additional source sub-data blocks.

[0091] In one embodiment, the linear combination value of the source sub-data blocks in the i-th row of the first part of the array corresponding to the i-th parity sub-data block of the j-th second parity data block is determined by the source sub-data blocks in the i-th row of the n-by-array of coded data blocks and the j-th row in the coding matrix.

[0092] In this embodiment, the linear combination value p of the source sub-data blocks in the row where the first part of the array is located i,j,LC1 is determined by the source sub-data blocks in the i-th row of the array and the j-th row in the encoding matrix. Here, i is an integer satisfying 0 ≤ i ≤ g - 1, and j is an integer satisfying 1 ≤ j ≤ m2.

[0093] As an example, the source sub-data blocks in the i-th row of the array are represented as [a i,0 , a i,1 ,..., a i,k-1 , the element values in the j-th row of the encoding matrix are respectively represented as [c j,0 , c j,1 ,..., c j,k-1 , and the first part p of the i-th check sub-data block of the j-th second check data block is obtained according to but not limited to the following rule: p i,j,LC1 = c i,j,LC1 * a j,0 + c i,0+ * a j,1 +... + c i,1+ * a j,k-1 i,k-1 . Here, "+" represents bitwise exclusive OR, and "*" represents multiplication operation in the Galois field GF(q).

[0094]

[0094] In a specific example, the number of source data blocks k = 4, the source sub-data blocks in the i = 0-th row of the array are represented as [a 0,0 , a 0,1 , a 0,2 , a i,3 = [10, 50, 90, 130], the element values in the j = 1-th row of the encoding matrix are respectively represented as [c j,0 , c j,1 ,..., c j,k-1 = [166, 70, 187, 123], then the first part of the i = 0-th check sub-data block of the j = 1-th second check data block is represented as:

[0095] p i,j,LC1 = 166 * 10 + 70 * 50 + 187 * 90 + 123 * 130 = 65.

[0096] In one embodiment, the linear combination value of the source sub-data blocks in at least one other row of the second part of the array is determined according to the additional source sub-data block index table and the number of repetitions.

[0097] In this embodiment, the linear combination value p of the source sub-data blocks in at least one other row of the second part i,j,LC2It is determined by, but not limited to, the following parameters: an additional source sub-data block index table and the number of repetitions. The additional source sub-data block index table is used to indicate the index positions of the additional source sub-data blocks in the array when generating the i-th check sub-data block of the j-th second check data block. The additional source sub-data block index table is an index array of size g*m2, where g is the number of rows of the source data block array and m2 is the number of second check data blocks.

[0098] In one embodiment, each element in the additional source sub-data block index table is a set of position coordinates, and the maximum number of one-dimensional coordinates or two-dimensional coordinates in the set of position coordinates is or When the additional source sub-data block index table includes all the index positions of the additional source sub-data blocks required to generate the second check data block, the maximum number of coordinates in the coordinate set is When the additional source sub-data block index table includes partial index positions of the additional source sub-data blocks required to generate the second check data block, the maximum number of coordinates in the coordinate set is All the index positions of the additional source sub-data blocks required to generate the second check data block are obtained through the above partial coordinates and the number of repetitions.

[0099] In one embodiment, each element in the index array of size g*m2 is a set of position coordinates, and the position coordinates include, but are not limited to: one-dimensional index coordinates and two-dimensional index coordinates. It can be understood that whether it is one-dimensional index coordinates or two-dimensional index coordinates, they both represent the positions of additional elements in the entire data array, and there is a one-to-one correspondence between one-dimensional index coordinates and two-dimensional index coordinates.

[0100] In one embodiment, each element in the additional source sub-data block index table is a set of position coordinates, and the number Ns of one-dimensional coordinates or two-dimensional coordinates in the set of position coordinates is determined by, but not limited to, at least one of the following parameters: the number k of source data blocks, the number m1 of first checks, and the number m2 of second checks. The number Ns of one-dimensional coordinates or two-dimensional coordinates in the set of position coordinates is at most one of the following values:

[0101] In one embodiment, each element in the additional source sub-data block index table is a set of two-dimensional coordinates;

[0102] In the case where the additional source sub-data block index table includes all the index positions of the additional source sub-data blocks required to generate the second check data block, the maximum number of coordinates in the coordinate set is

[0103] When the additional source sub - data block index table includes the index positions of some additional source sub - data blocks required to generate the second check data block, the maximum number of coordinates in the coordinate set is

[0104] In one embodiment, the additional source sub - data block index table is determined according to k, the first check quantity, and the second check quantity.

[0105] As an example, each element in the additional source sub - data block index table IndexTable is a two - dimensional coordinate set IJtable0. When the additional source sub - data block index table includes the index positions of some additional source sub - data blocks required to generate the second check data block, the maximum number of two - dimensional coordinates in the two - dimensional coordinate set Ns is The linear combination value p of the source sub - data blocks in at least one other row of the second part i,j,LC2 Determined according to the additional source sub - data block index table and the repetition times, can be obtained according to but not limited to the following rules:

[0106]

[0107] It can be understood that only some index positions of the additional source sub - data blocks are included in the above - mentioned additional source sub - data block index table. Some additional source sub - data blocks refer to the additional source sub - data blocks of the first previous source data blocks. All the index positions of the additional source sub - data blocks are obtained through the repetition times s. This reduces the storage by introducing a small amount of calculation.

[0108] In one embodiment, each element in the additional source sub - data block index table IndexTable is a two - dimensional coordinate set IJtable0. When the additional source sub - data block index table includes all the index positions of the additional source sub - data blocks required to generate the second check data block, the maximum number of two - dimensional coordinates in the two - dimensional coordinate set Ns is The linear combination value p of the source sub - data blocks in at least one other row of the second part i,j,LC2 Determined according to the additional source sub - data block index table, can be obtained according to but not limited to the following rules:

[0109]

[0110] It can be understood that all the index positions of the additional source sub - data blocks are included in the above - mentioned additional source sub - data block index table, which requires no additional calculation but increases the storage. The above represents exclusive - OR operation. Each element in the additional source sub - data block index table IndexTable is a two - dimensional coordinate set IJtable0. When the two - dimensional coordinate set includes all the coordinates for generating p i,j,LC2When obtaining the required additional source sub-data block index positions, the maximum number Ns of two-dimensional coordinates in the two-dimensional coordinate set is When the two-dimensional coordinate set only includes a partial generation p i,j,LC2 When obtaining the required additional source sub-data block index positions, the maximum number Ns of two-dimensional coordinates in the two-dimensional coordinate set is It is necessary to further obtain all the additional source sub-data block index positions through the repetition number s.

[0111] In one embodiment, the additional source sub-data block index table IndexTable is obtained based on, but not limited to, the number k of source data blocks, the number m1 of first check data blocks, and the number m2 of second check data blocks. The coordinates index of the sub-data block in the i-th column and Gj-th row in the array are respectively in the t-th column and Gi-th row of the source sub-data block index table. i is an integer greater than or equal to 1 and less than or equal to ceil(k / s); Gi is the set of row index values formed by dividing the set of integer values {1, 2,..., g} into (m1 + m2) ceil(i / (m1+m2)) sub-parts in sequence, and starting from the (mod(i - 1, m1 + m2) + 1)-th sub-part, reading one sub-part every (m1 + m2) sub-parts in sequence; the t-th element Gj in the set GSet of the i-th source data block in the array, where j belongs to the first integer set, and the first integer set is [floor(i / (m1 + m2)) + 1, ceil(i / m1 + m2)]\i; t belongs to the second integer set, and the second integer set contains integers greater than or equal to 1 and less than or equal to the number of integers in the first integer set; Gj is the set of row index values formed by dividing the set of integer values {1, 2,..., g} into (m1 + m2) ceil(j / (m1+m2)) sub-parts in sequence, and starting from the (mod(j - 1, m1 + m2) + 1)-th sub-part, reading one sub-part every (m1 + m2) sub-parts in sequence; t and j are respectively the elements corresponding to the same index size in the first integer set and the second integer set to which they belong, and the elements in the first integer set and the second integer set are sorted from small to large; g is the number of rows of the array; ceil() represents adjusting a numerical value to the smallest integer not less than the numerical value itself, floor() represents adjusting a numerical value to the largest integer not greater than the numerical value itself, and mod() represents the modulo operation; []\i represents removing the element i from a set.

[0112] In one example, k = 6, m1 + m2 = 3, s = 1, g = (m1 + m2) ceil(k ' / (m1+m2))= 9. When i = 2, Gi = G2 = {4, 5, 6}. For the second data block, [floor(i / (m1 + m2)) + 1, ceil(i / m1 + m2)] \ i = {1, 2, 3} \ {2} = {1, 3}. Therefore, the set GSet of the i = 2nd source data block of the array is GSet = {G1, G2, G3} \ {G2} = {G1, G3}. The value of t is an integer greater than or equal to 1 and less than or equal to 2. The above G1 = {1, 2, 3}, G3 = {7, 8, 9}. For the t = 1st column and the G2 = {4, 5, 6}th row of the source sub - data block index table, they are the coordinate indexes of the sub - data blocks in the i = 2nd column and the G1 = {1, 2, 3}th row of the array respectively; for the t = 2nd column and the G2 = {4, 5, 6}th row of the source sub - data block index table, they are the coordinate indexes of the sub - data blocks in the i = 2nd column and the G3 = {7, 8, 9}th row of the array respectively.

[0113] In the following examples, k’ is used to represent ceil(k / s). In one example, a feasible source sub - data block index table can be obtained according to but not limited to the following rules:

[0114]

[0115] Among them, cell() represents constructing a cell array, floor() represents adjusting a numerical value to the largest integer not greater than it, and g represents the number of rows of the data array. SplitNum() is a function of the sum of the number of rows g of the data array, the number of the first parity data blocks m1, and the number of the second parity data blocks m2. The SplitNum() outputs a matrix of size (Layer * (m1 + m2)) * (g / (m1 + m2)), and each element in the matrix is greater than or equal to 1 and less than or equal to g.

[0116] Among them, GNum is a matrix with (ceil(k' / (m1 + m2))) * r rows. The i - th row of GNum is denoted as Gi. Gi is formed by dividing the set of integer values {1, 2,..., g} into (m1 + m2) ceil(i / (m1+m2)) sub - parts in order, starting from the (mod(i - 1, m1 + m2) + 1)-th sub - part, and reading one sub - part every (m1 + m2) sub - parts. In one example, k = 6, m1 + m2 = 3, g = (m1 + m2) ceil(k ' / (m1+m2)) = 9. For Gi = G2, divide 1 - 9 into 3 sub - parts in order, which are {1, 2, 3}, {4, 5, 6}, {7, 8, 9} respectively. The set formed by starting from the (mod(i - 1, m1 + m2) + 1) = 2nd sub - part and reading one sub - part every 3 sub - parts is {4, 5, 6}. Therefore, G2 = {4, 5, 6}. In another example, k = 6, r = 3, g = (m1 + m2)ceil(k ' / (m1+m2)) = 9. For Gi = G4, divide the numbers from 1 to 9 into 9 sub - parts in order, which are {1}, {2}, {3}, {4}, {5}, {6}, {7}, {8}, {9}. Starting from the (mod(i - 1, r)+1) = 1 - st sub - part, the set obtained by reading one sub - part every 3 sub - parts is {1, 4, 7}. So G4 = {1, 4, 7}. In an example, a feasible GNum matrix can be obtained according to but not limited to the following rules:

[0117]

[0118]

[0119] where ceil() represents adjusting a numerical value to the smallest integer not less than the value itself, g represents the number of rows of the data array, T represents transpose, TempMx(j:(m1 + m2):end, :) represents extracting all columns of the matrix TempMx and extracting one row every (m1 + m2) rows from the j - th row to the last row to form a new matrix. TempSubMx(:) represents expanding a two - dimensional matrix into a one - dimensional matrix by columns.

[0120] As an example, each element in the additional source sub - data block index table is a set of two - dimensional coordinates. When the set of two - dimensional coordinates only includes some of the additional source sub - data block index positions required to generate p i,j,LC2 the number of two - dimensional coordinates Ns in the set of two - dimensional coordinates is at most and all additional source sub - data block index positions need to be further obtained through the repetition number s.

[0121] In a specific example, the number of source data blocks k = 4, the repetition number s = 1, the first check number m1 = 1, the second check number m2 = 3, and the maximum number of two - dimensional coordinates is Ns = 1. A feasible additional source sub - data block index table is shown in Table 1. The coordinates in the table represent the coordinate positions of the corresponding additional source sub - data blocks. For example, when i = 1 and j = 0, the set of coordinates is {(1, 0)}. Then for the second part p i,j,LC2 , p i,j,LC2 = a 1,0 .

[0122] Table 1 A feasible additional source sub - data block index table when k = 4, s = 1, m1 = 1, m2 = 3

[0123]

[0124] As an example, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the two-dimensional coordinate set only includes some of the additional source sub-data block index positions required for generating p i,j,LC2 the maximum number Ns of two-dimensional coordinates in the two-dimensional coordinate set is All additional source sub-data block index positions need to be further obtained through the repetition times. In a specific example, the number of source data blocks k = 8, the repetition times s = 1, the first check number m1 = 1, the second check number m2 = 3, and the maximum number of two-dimensional coordinates is Ns = 2. A feasible additional source sub-data block index table is shown in Table 2. For the second part p i,j,LC2 of the i = 2nd check sub-block for generating the j = 1st second check data block where denotes bitwise exclusive OR.

[0125] Table 2 A feasible additional source sub-data block index table when k = 8, s = 1, m1 = 1, m2 = 3

[0126]

[0127]

[0128] As an example, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the two-dimensional coordinate set only includes some of the additional source sub-data block index positions required for generating p i,j,LC2 the maximum number Ns of two-dimensional coordinates in the two-dimensional coordinate set is All additional source sub-data block index positions need to be further obtained through the repetition times. In a specific example, the number of source data blocks k = 70, the repetition times s = 10, the first check number m1 = 1, the second check number m2 = 3, and the maximum number of two-dimensional coordinates is Ns = 2. A feasible additional source sub-data block index table is shown in Table 3. For example, when i = 3 and j = 1, the coordinate set is {(7,0)}. The coordinates in this set only include the additional source sub-data blocks of the first several source data blocks. If all additional sub-data block indexes are to be obtained, according to the j-th column with other indexes from 7 to 69, the corresponding additional sub-data block row indexes are the same as those of the additional source data block row indexes in the mod(j,7) column. For the second part p i,j,LC2 of the i = 3rd check sub-block for generating the j = 1st second check data block where denotes bitwise exclusive OR.

[0129] Table 3 A feasible additional source sub-data block index table when k = 70, s = 10, m1 = 1, and m2 = 3

[0130]

[0131]

[0132] As an example, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the two-dimensional coordinate set includes all the additional source sub-data block index positions required to generate p i,j,LC2 The maximum number Ns of two-dimensional coordinates in the two-dimensional coordinate set is at most In a specific example, the number of source data blocks k = 8, the repetition number s = 2, the first check number m1 = 1, the second check number m2 = 3, and the maximum number of two-dimensional coordinates is Ns = 2. A feasible additional source sub-data block index table is shown in Table 4. For example, when i = 3 and j = 1, the coordinate set is {(0, 3); (0, 7)}. The coordinates in this set include all the additional source sub-data block indexes. For the second part p of the i = 3rd check sub-block for generating the j = 1st second check data block i,j,LC2 , then wherein represents bitwise exclusive OR.

[0133] Table 4 A feasible additional source sub-data block index table when k = 8, s = 2, m1 = 1, and m2 = 3

[0134]

[0135] It can be understood that in the above example, the two-dimensional index coordinates all represent the positions of the additional sub-data blocks in the entire data array. At the same time, the two-dimensional coordinates can be converted into one-dimensional index coordinates for representation, that is, there is a one-to-one correspondence between the one-dimensional coordinates and the two-dimensional index coordinates. In a feasible example, the one-dimensional coordinate index corresponding to the two-dimensional coordinate index (i1, j1) of the additional sub-data block is (g * j1 + i1). In another feasible example, the one-dimensional coordinate index corresponding to the two-dimensional coordinate index (i1, j1) is (k * i + j1). Wherein, i1 is an integer greater than or equal to 0 and less than g, j1 is an integer greater than 0 and less than or equal to k, and g is the number of rows of the data array.

[0136] In one embodiment, each check sub-data block in the third check data block is obtained by a linear combination of the source sub-data blocks in the row where the array of the n encoded data blocks is located. The linear combination coefficients are obtained from the encoding matrix and the linear combination coefficients are not all equal to one.

[0137] In this embodiment, each parity sub-data block in the third parity data block is obtained by a linear combination of source sub-data blocks in the row where the array is located. The linear combination coefficients are obtained from the coding matrix and the linear combination coefficients are not all equal to one. In one example, the i-th parity sub-data block of the j-th third parity data block is denoted as p i,m1+m2+j , and the source sub-data blocks corresponding to the i-th row of the data array where the i-th parity sub-data block of the j-th third parity data block is located are [a i,0 , a i,1 ,..., a i,k-1 , and the generated linear combination coefficients corresponding to the j-th third parity data block are [c j,0 , c j,1 ,..., c j,k-1 . Then, the i-th parity sub-data block of the third parity data block can be obtained according to, but not limited to, the following rules: p i,m1+m2+j = c j,0 * a i,0 + c j,1 * a i,1 +... + c j,k-1 * a i,k-1 . Among them, "+" represents bitwise exclusive OR, and "*" represents multiplication operation in the Galois field GF(q). j is an integer greater than 0 and less than m3, i is an integer greater than or equal to 0 and less than g, and g is the number of rows of the data array. Based on the above embodiment, the linear combination coefficients for generating the j-th third parity data block include, but are not limited to, being obtained from the last m3 rows of the following coding matrices: Vandermonde matrix, deformed matrix of Cauchy matrix, combination of all-ones matrix and Cauchy matrix.

[0138] As an example, as Figure 6 shows, there are a total of k = 4 source data blocks, and each source data block has g = 4 source sub-data blocks. For ease of description, the size of each source sub-data block is one byte, which should not be a limitation of this application. The one byte is represented by a decimal number. The number of first parity data blocks m1 = 1, the number of second parity data blocks m2 = 2, and the number of third parity data blocks m3 = 1. When the linear combination coefficients for generating the third parity data block are obtained from the Vandermonde matrix, a feasible set of linear combination coefficients is [1, 8, 64, 58]. The i-th parity sub-data block of the third parity data block can be obtained according to the following rules: p i,m1+m2+j = c j,0 * a i,0 + c j,1 * a i,1 +... + c j,k-1 * a i,k-1For example, the value of the 0th sub-check data block of the third check data block is 1 * 10 + 8 * 50 + 64 * 90 + 58 * 130 = 188, and the value of the 3rd sub-check data block of the third check data block is 1 * 40 + 8 * 80 + 64 * 120 + 58 * 160 = 166, where "+" represents bitwise exclusive OR and "*" represents multiplication operation in the Galois field GF(q).

[0139] As an example, as Figure 7 shown, there are k = 4 source data blocks in total, each source data block has g = 4 source sub-data blocks, the size of each source sub-data block is one byte, and the one byte is represented by a decimal number. The number of the first check data blocks m1 = 1, the number of the second check data blocks m2 = 2, and the number of the third check data blocks m3 = 1. When the linear combination coefficients for generating the third check data block are obtained from the deformed matrix of the Cauchy matrix, a feasible set of linear combination coefficients is represented as [210, 143, 245, 200]. The i-th check sub-data block of the third check data block can be obtained according to the following rule: p i,m1+m2+j = c j,0 * a i,0 + c j,1 * a i,1 +... + c j,k-1 * a i,k-1 For example, the value of the 0th sub-check data block of the third check data block is 210 * 10 + 143 * 50 + 245 * 90 + 200 * 130 = 77, and the value of the 3rd sub-check data block of the third check data block is 210 * 40 + 143 * 80 + 245 * 120 + 200 * 160 = 113, where "+" represents bitwise exclusive OR and "*" represents multiplication operation in the Galois field GF(q).

[0140] As an example, as Figure 8 shown, there are k = 4 source data blocks in total, each source data block has g = 4 source sub-data blocks, the size of each source sub-data block is one byte, and the one byte is represented by a decimal number. The number of the first check data blocks m1 = 1, the number of the second check data blocks m2 = 2, and the number of the third check data blocks m3 = 1. When the linear combination coefficients for generating the third check data block are obtained from the combination of the all-ones matrix and the Cauchy matrix, a feasible set of linear combination coefficients is represented as [122, 186, 71, 167]. The i-th check sub-data block of the third check data block can be obtained according to the following rule: p i,m1+m2+j = c j,0 * a i,0 + c j,1 * a i,1 +... + c j,k-1 * a i,k-1For example, the value of the 0th sub-check data block of the third check data block is 122 * 10 + 186 * 50 + 71 * 90 + 167 * 130 = 105, and the value of the 3rd sub-check data block of the third check data block is 122 * 40 + 186 * 80 + 71 * 120 + 167 * 160 = 252, where "+" represents bitwise exclusive OR and "*" represents multiplication operation in the Galois field GF(q).

[0141] In one embodiment, the method further includes step 130: storing the n encoded data blocks in n different storage nodes. By storing the n encoded data blocks in n different storage nodes, when a single source data block storage node fails, at least 1 / (m1 + m2) of the data volume can be obtained from the remaining source data block storage nodes, the first check data block storage node, and the second check data block storage node respectively to complete the recovery of the single source data block storage node. And when (m1 + m2) storage nodes fail and the data block recovery cannot be completed only through the first check data block and the second check data block, the third check data block can be further used to complete the recovery of the (m1 + m2) failed storage nodes.

[0142] It can be understood that after the first processing node obtains the k source data blocks, the k source data blocks are divided into a source data array with a size of g * k. The first check data block is a linear combination of the current row of the array and the linear combination coefficients are all 1; the second check data block is the bitwise exclusive OR of the linear combination value of the current row of the array and the linear combination value of at least one source sub-data block in other rows. Among them, the linear combination coefficient value of the current row of the second check data block is not all 1, and the linear combination coefficient value of at least one source sub-data block in other rows of the second check data block is all 1; the third check data block is the linear combination value of the current row of the array and the linear combination coefficients are not all 1. Since the second check data block contains additional source sub-data blocks in other rows, when a single source data block fails, the single source data block can be repaired through part of the data of the first check data block and the second check data block, thereby reducing the repair bandwidth when a single storage node fails. And when (m1 + m2) encoded data blocks fail and the data block repair cannot be completed only based on the first check data block and the second check data block, the third check data block can be used for recovery. Since each check sub-data block in the third check data block is only related to the source sub-data blocks in the current row of the array, the amount of data recovery operations can be reduced during the data repair process. The implementation method of the present application not only has the ability of local repair, but also has high repair ability when multiple storage nodes fail.

[0143] Figure 9 It is a flowchart of a data processing method provided for an embodiment. This method can be applied to the second processing node. The execution subject of this data processing method can include but is not limited toFigure 1 The second processing node 20 shown, or those skilled in the art can select and set the corresponding execution entity according to the actual application scenario, and this embodiment does not limit this. For example, Figure 9 As shown, the method provided in this embodiment includes step 210 and step 220.

[0144] In step 210, obtain the type and quantity of the fault coding data block and part or all of the data of at least k non-faulty coding data blocks generated by the first processing node, where k is a positive integer greater than 1.

[0145] In step 220, process part or all of the data of the at least k non-faulty coding data blocks according to the type and quantity of the fault coding data block to obtain the faulty coding data block.

[0146] The type of the fault coding data block includes at least one of the following: source data block, first check data block, second check data block, third check data block. The quantity Ne of the fault coding data block is an integer greater than or equal to 0 and less than or equal to n.

[0147] In one embodiment, the at least k non-faulty coding data blocks are obtained by the first processing node encoding k source data blocks according to the coding parameters; the coding parameters include: the first check quantity, the second check quantity, the third check quantity, the additional source sub-data block index table, the repetition times, and the coding matrix.

[0148] In this embodiment, the coding data block is obtained by the first processing node encoding k source data blocks according to at least the following parameters after obtaining the k source data blocks: the first check quantity m1, the second check quantity m2, the third check quantity m3, the additional source sub-data block index table, the repetition times s, and the coding matrix G.

[0149] In one embodiment, obtaining part or all of the data of at least k non-faulty coding data blocks generated by the first processing node includes:

[0150] When the type of the fault coding data block includes at least one of the first check data block, the second check data block, and the third check data block, or when the quantity of the fault source data blocks is greater than or equal to the sum of the first check quantity and the real-time second check quantity, obtain all the non-faulty coding data blocks generated by the first processing node;

[0151] In the case where only the source data blocks fail and the number of failed source data blocks is less than the sum of the first check quantity and the second check quantity, obtain the partially non-failed encoded data blocks generated by the first processing node. In one example, k = 4, m1 = 1, m2 = 3, m3 = 1, the repetition number s = 1, and the number of rows g of the array = 4. When the j = 0th source data block a0 fails, obtaining the remaining source data blocks a1, a2, a3, the first check data block, and the i = 0th sub-data block of the second check data block can complete the recovery of the failed data block, and the obtained data is only the partially non-failed encoded data blocks generated by the first processing node.

[0152] In this embodiment, when the type of the failed encoded data block includes one of the first check data block, the second check data block, and the third check data block, or the number of failed source data blocks Ne is greater than or equal to (m1 + m2), obtain all the non-failed encoded data blocks generated by the first processing node; when only the source data blocks fail and the number of failed source data blocks Ne is less than (m1 + m2), obtain the partially non-failed encoded data blocks generated by the first processing node.

[0153] In one embodiment, process some or all of the data of the at least k non-failed encoded data blocks according to the type and quantity of the failed encoded data blocks, including at least one of the following:

[0154] In the case where the number of failed encoded data blocks exceeds the sum of the first check quantity, the second check quantity, and the third check quantity, determine the data loss of the failed encoded data blocks;

[0155] In the case where the number of failed encoded data blocks is greater than 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the type of the failed encoded data blocks includes source data blocks and check data blocks, obtain all the non-failed encoded data blocks, and first recover the failed source data blocks, and then recover the check data blocks;

[0156] In the case where the number of failed encoded data blocks is greater than 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the type of the failed encoded data blocks only includes source data blocks, obtain at least the first set ratio of the non-failed encoded data blocks, and recover the failed source data blocks, where the first set ratio is the ratio of the number of failed encoded data blocks in the sum of the first check quantity and the second check quantity;

[0157] In the case where the number of failed encoded data blocks is greater than or equal to 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the type of the failed encoded data blocks only includes check data blocks, obtain all the non-failed data blocks, and recover the failed source data blocks;

[0158] When the number of fault coding data blocks is equal to 1 and the fault coding data block type only includes source data blocks, respectively obtain a second set proportion of non-faulty source data blocks, a first check data block, and a second check data block, and recover the faulty source data block, where the second set proportion is the proportion of 1 in the sum of the first check quantity and the second check quantity.

[0159] In this embodiment, according to the type and quantity of the fault coding data blocks, process some or all of the data of the at least k non-faulty coding data blocks to obtain faulty coding data blocks, including but not limited to at least one of the following:

[0160] Step 2201: If the number Ne of the fault coding data blocks exceeds (m1 + m2m + m3), the data of the fault coding data blocks is lost and data recovery cannot be completed;

[0161] Step 2202: If the number Ne of the fault coding data blocks is greater than 1 and less than or equal to (m1 + m2 + m3), and the fault coding data block type includes both source data blocks and check data blocks, obtain all non-faulty coding data blocks, and first recover the faulty source data blocks, and then recover the check data blocks;

[0162] Step 2203: If the number Ne of the fault coding data blocks is greater than 1 and less than or equal to (m1 + m2 + m3), and the fault coding data block type only includes source data blocks, obtain at least the data volume of the proportion of Ne / (m1 + m2) of the remaining non-faulty data blocks, and then recover the faulty source data blocks;

[0163] Step 2204: If the number Ne of the fault coding data blocks is greater than or equal to 1 and less than or equal to (m1 + m2 + m3), and the fault coding data block type only includes check data blocks, obtain all non-faulty data blocks, and then recover the faulty source data blocks;

[0164] Step 2205: If the number Ne of the fault coding data blocks is equal to 1, and the fault coding data block type only includes source data blocks, obtain at least the data volume of the proportion of 1 / (m1 + m2) of the remaining non-faulty source data blocks, the first check data block, and the second check data block respectively, and then recover the faulty source data blocks;

[0165] It can be understood that there is no order relationship among steps 2201 to S2205, and they only represent at least several data processing methods according to the type and quantity of the first coding data blocks.

[0166] The embodiment of the present application also provides a data processing device. Figure 10 It is a schematic structural diagram of a data processing device provided for an embodiment. As Figure 10As shown, the data processing device includes:

[0167] An acquisition module 310, configured to acquire k source data blocks, where k is a positive integer greater than 1;

[0168] An encoding module 320, configured to encode the k source data blocks according to encoding parameters to obtain n encoded data blocks, where n is a positive integer greater than k; the encoding parameters include: the first check quantity, the second check quantity, the third check quantity, an additional source sub-data block index table, the number of repetitions, and an encoding matrix.

[0169] In one embodiment, the n encoded data blocks include k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks, where m1 is the first check quantity and m1 = 1, m2 is the second check quantity and is a positive integer, m3 is the third check quantity and is a positive integer, and n = k + m1 + m2 + m3;

[0170] The additional source sub-data index table is an index table with g rows and m2 columns, where g is the number of source sub-data blocks included in each source data block;

[0171] The size of the encoding matrix is m * k, where m is the total number of check data blocks, and m = m1 + m2 + m3;

[0172] The number of repetitions is a positive integer;

[0173] The elements in the encoding matrix are elements in the Galois field GF(q), where q is an integer power of a prime number.

[0174] In one embodiment, the k source data blocks respectively correspond to the data parts stored in the first k storage nodes in a stripe; the k source data blocks are distributed in k different storage nodes;

[0175] The m check data blocks obtained by encoding the k source data blocks according to the encoding parameters respectively correspond to the data parts stored in the last m storage nodes in a stripe;

[0176] The m check data blocks are respectively stored in m different storage nodes other than the k different storage nodes.

[0177] In one embodiment, the number of rows of the array of the n encoded data blocks is determined by at least one of the following parameters: k, the first check quantity, the second check quantity.

[0178]

[0179] In one embodiment, the encoding matrix satisfies at least one of the following:

[0180] There is a row of element values all being "1".

[0181] Any square matrix formed by any m rows and m columns is invertible in the Galois field GF(q).

[0182] In one embodiment, each syndrome data block in the first check data block is obtained by bitwise XOR of the source sub-data blocks in the row where the array of the n coded data blocks is located.

[0183] In one embodiment, each syndrome data block in the second check data block is obtained by bitwise XOR of the linear combination value of the source sub-data blocks in the row where the first part of the array in the array of the n coded data blocks is located and the linear combination value of the source sub-data blocks in at least one other row of the second part of the array.

[0184] In one embodiment, the linear combination value of the source sub-data blocks in the row where the first part of the array corresponding to the i-th syndrome data block in the j-th second check data block is located is determined by the source sub-data blocks in the i-th row of the array of the n coded data blocks and the j-th row in the coding matrix.

[0185] In one embodiment, the linear combination value of the source sub-data blocks in at least one other row of the second part of the array is determined according to the additional source sub-data block index table and the number of repetitions.

[0186] In one embodiment, each element in the additional source sub-data block index table is a set of position coordinates, and the maximum number of one-dimensional coordinates or two-dimensional coordinates in the set of position coordinates is or

[0187] In one embodiment, each element in the additional source sub-data block index table is a set of two-dimensional coordinates;

[0188] When the additional source sub-data block index table includes all the index positions of the additional source sub-data blocks required to generate the second check data block, the maximum number of coordinates in the coordinate set is

[0189] When the additional source sub-data block index table includes partial index positions of the additional source sub-data blocks required to generate the second check data block, the maximum number of coordinates in the coordinate set is

[0190] In one embodiment, the additional source sub-data block index table is determined according to k, the first check quantity, and the second check quantity.

[0191] In one embodiment, the t-th column and the Gi-th row in the additional source sub-data block index table are respectively the coordinate indexes of the sub-data blocks in the i-th column and the Gj-th row in the array of the n coded data blocks;

[0192] where \(i\) is an integer greater than or equal to 1 and less than or equal to \(\lceil k / s\rceil\);

[0193] where \(G_i\) is a set of row index values formed by sequentially dividing the integer values \(\{1, 2, \ldots, g\}\) into \((m_1 + m_2)\) sub - parts, and starting from the \((\text{mod}(i - 1,m_1 + m_2)+1)\) - th sub - part in order, reading one sub - part every \((m_1 + m_2)\) sub - parts; the \(t\) - th element in the set \(GSet\) of the \(i\) - th source data block of the array is \(G_j\); ceil(i / (m1+m2)) where \(G_i\) is a set of row index values formed by sequentially dividing the integer values \(\{1, 2, \ldots, g\}\) into \((m_1 + m_2)\) sub - parts, and starting from the \((\text{mod}(i - 1,m_1 + m_2)+1)\) - th sub - part in order, reading one sub - part every \((m_1 + m_2)\) sub - parts; the \(t\) - th element in the set \(GSet\) of the \(i\) - th source data block of the array is \(G_j\);

[0194] where \(j\) belongs to the first set of integers, and the first set of integers is \([\lfloor i / (m_1 + m_2)\rfloor+1,\lceil i / m_1 + m_2\rceil]\setminus i\);

[0195] where \(t\) belongs to the second set of integers, and the second set of integers contains integers greater than or equal to 1 and less than or equal to the number of integers in the first set of integers;

[0196] where \(G_j\) is a set of row index values formed by sequentially dividing the integer values \(\{1, 2, \ldots, g\}\) into \((m_1 + m_2)\) sub - parts, and starting from the \((\text{mod}(j - 1,m_1 + m_2)+1)\) - th sub - part in order, reading one sub - part every \((m_1 + m_2)\) sub - parts; ceil(j / (m1+m2)) where \(G_j\) is a set of row index values formed by sequentially dividing the integer values \(\{1, 2, \ldots, g\}\) into \((m_1 + m_2)\) sub - parts, and starting from the \((\text{mod}(j - 1,m_1 + m_2)+1)\) - th sub - part in order, reading one sub - part every \((m_1 + m_2)\) sub - parts;

[0197] where \(t\) and \(j\) are elements corresponding to the same index size in the first set of integers and the second set of integers respectively, and the elements in the first set of integers and the second set of integers are sorted in ascending order;

[0198] where \(g\) is the number of rows of the array;

[0199] \(\lceil\rceil\) represents adjusting a numerical value to the smallest integer not less than the numerical value itself, \(\lfloor\rfloor\) represents adjusting a numerical value to the largest integer not greater than the numerical value itself, and \(\text{mod}(\ )\) represents the modulo operation;

[0200] \(\setminus i\) means removing the element \(i\) from a set.

[0201] In one embodiment, each parity sub - data block in the third parity data block is obtained by a linear combination of the source sub - data blocks in the row where the array of the \(n\) encoded data blocks is located, and the linear combination coefficients are obtained from the encoding matrix and the linear combination coefficients are not all equal to one.

[0202] The data processing device proposed in this embodiment and the data processing method proposed in the above embodiment belong to the same inventive concept. For technical details not described in detail in this embodiment, reference can be made to any of the above embodiments, and this embodiment has the same beneficial effects as the execution of the data processing method.

[0203] An embodiment of the present application further provides a data processing device. Figure 11 As shown in the structural schematic diagram of a data processing device provided for an embodiment. Figure 11 As shown, the data processing device includes:

[0204] An acquisition module 410, configured to acquire the type and quantity of a fault coding data block and part or all of the data of at least k non-fault coding data blocks generated by a first processing node, where k is a positive integer greater than 1;

[0205] A processing module 420, configured to process part or all of the data of the at least k non-fault coding data blocks according to the type and quantity of the fault coding data block to obtain a fault coding data block.

[0206] In one embodiment, the at least k non-fault coding data blocks are obtained by the first processing node encoding k source data blocks according to encoding parameters;

[0207] The encoding parameters include: a first check quantity, a second check quantity, a third check quantity, an additional source sub-data block index table, a repetition number, and an encoding matrix.

[0208] In one embodiment, the acquisition module 410 is configured to:

[0209] In the case where the type of the fault coding data block includes at least one of a first check data block, a second check data block, and a third check data block, or in the case where the number of fault source data blocks is greater than or equal to the sum of the first check quantity and the real-time second check quantity, acquire all non-fault coding data blocks generated by the first processing node;

[0210] In the case where only source data blocks are faulty and the number of fault source data blocks is less than the sum of the first check quantity and the second check quantity, acquire part of the non-fault coding data blocks generated by the first processing node.

[0211] In one embodiment, the processing module 420 is configured to at least one of the following:

[0212] In the case where the number of fault coding data blocks exceeds the sum of the first check quantity, the real-time second check quantity, and the third check quantity, determine that the data of the fault coding data block is lost;

[0213] When the number of faulty coded data blocks is greater than 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the types of faulty coded data blocks include source data blocks and check data blocks, obtain all non-faulty coded data blocks, first recover the faulty source data blocks, and then recover the check data blocks;

[0214] When the number of faulty coded data blocks is greater than 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the types of faulty coded data blocks only include source data blocks, obtain at least a first set proportion of non-faulty coded data blocks and recover the faulty source data blocks, where the first set proportion is the proportion of the number of faulty coded data blocks in the sum of the first check quantity and the second check quantity;

[0215] When the number of faulty coded data blocks is greater than or equal to 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the types of faulty coded data blocks only include check data blocks, obtain all non-faulty data blocks and recover the faulty source data blocks;

[0216] When the number of faulty coded data blocks is equal to 1 and the types of faulty coded data blocks only include source data blocks, obtain a second set proportion of non-faulty source data blocks, the first check data blocks, and the second check data blocks respectively, and recover the faulty source data blocks, where the second set proportion is the proportion of 1 in the sum of the first check quantity and the second check quantity.

[0217] The data processing device proposed in this embodiment and the data processing method proposed in the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to in any of the above embodiments, and this embodiment has the same beneficial effects as the execution of the data processing method.

[0218] An embodiment of the present application also provides a processing node, Figure 12 As shown in the schematic diagram of the hardware structure of a processing node provided for an embodiment, Figure 12 The processing node provided by the present application includes a processor 510 and a memory 520; the processor 510 in the processing node can be one or more, Figure 12 Taking one processor 510 as an example; the memory 520 is configured to store one or more programs; the one or more programs are executed by the one or more processors 510, so that the one or more processors 510 implement the data processing method as described in the embodiments of the present application.

[0219] The processing node further includes: a communication device 530, an input device 540, and an output device 550.

[0220] The processor 510, memory 520, communication device 530, input device 540, and output device 550 in the processing node can be connected via a bus or other means. Figure 12 Take connection via a bus as an example.

[0221] The input device 540 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function control of the processing node. The output device 550 can include display devices such as a display screen.

[0222] The communication device 530 can include a receiver and a transmitter. The communication device 530 is configured to perform information transceiver communication under the control of the processor 510.

[0223] The memory 520, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data processing method described in the embodiments of the present application (for example, the acquisition module 310 and encoding 320 in the data processing device). The memory 520 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the processing node. In addition, the memory 520 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 520 can further include a memory remotely set relative to the processor 510, and these remote memories can be connected to the processing node through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0224] The embodiments of the present application also provide a processing node. Figure 13 As shown in the schematic diagram of the hardware structure of a processing node provided in an embodiment, Figure 13 the processing node provided by the present application includes a processor 610 and a memory 620; the processor 610 in the processing node can be one or more. Figure 13 Take one processor 610 as an example; the memory 620 is configured to store one or more programs; the one or more programs are executed by the one or more processors 610, so that the one or more processors 610 implement the data processing method described in the embodiments of the present application.

[0225] The processing node further includes: a communication device 630, an input device 640, and an output device 660.

[0226] The processor 610, memory 620, communication device 630, input device 640, and output device 660 in the processing node can be connected by a bus or other means. Figure 13 Taking the connection through the bus as an example.

[0227] The input device 640 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the processing node. The output device 660 may include display devices such as a display screen.

[0228] The communication device 630 may include a receiver and a transmitter. The communication device 630 is configured to perform information transceiver communication under the control of the processor 610.

[0229] The memory 620, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data processing method described in the embodiments of the present application (for example, the acquisition module 410 and the processing module 420 in the data processing device). The memory 620 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the processing node. In addition, the memory 620 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 620 may further include a memory remotely set relative to the processor 610, and these remote memories can be connected to the processing node through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0230] The embodiments of the present application further provide a data processing system. Figure 14 As shown in the structural schematic diagram of a data processing system provided for an embodiment, such as Figure 14 shown, the system includes: a storage node set 30, a first processing node 10, and a second processing node 20; the storage node set 30 includes n storage nodes 31, and the n storage nodes 31 include k source data storage nodes 32 and n - k parity data storage nodes 33, where n is a positive integer greater than k, and k is a positive integer greater than 1. The data processing system of this embodiment can be used to implement the data processing method of any of the above embodiments.

[0231] An embodiment of the present application further provides a storage medium storing a computer program, which when executed by a processor implements any one of the data processing methods in the embodiments of the present application. The data processing method includes: obtaining k source data blocks, where k is a positive integer greater than 1; encoding the k source data blocks according to encoding parameters to obtain n encoded data blocks, where n is a positive integer greater than k; the encoding parameters include: the first check quantity, the second check quantity, the third check quantity, an additional source sub-data block index table, the number of repetitions, and an encoding matrix. Alternatively, the data processing method includes: obtaining the type and quantity of faulty encoded data blocks and some or all of the data of at least k non-faulty encoded data blocks generated by a first processing node, where k is a positive integer greater than 1; processing some or all of the data of the at least k non-faulty encoded data blocks according to the type and quantity of the faulty encoded data blocks to obtain the faulty encoded data blocks.

[0232] The computer storage medium of the embodiments of the present application may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0233] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to: an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0234] The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber cable, radio frequency (RF), etc., or any suitable combination of the above.

[0235] The computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0236] As described above, these are only exemplary embodiments of the present application and are not intended to limit the protection scope of the present application.

[0237] Those skilled in the art should understand that the term user terminal covers any suitable type of wireless user device, such as a mobile phone, a portable data processing device, a portable network browser, or an in-vehicle mobile station.

[0238] Generally speaking, various embodiments of the present application can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices, although the present application is not limited thereto.

[0239] Embodiments of the present application can be implemented by a data processor of a mobile device executing computer program instructions, for example, in a processor entity, or by hardware, or by a combination of software and hardware. The computer program instructions can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages.

[0240] Any block diagram of a logical process in the accompanying drawings of the present application may represent a program step, or may represent interconnected logic circuits, modules, and functions, or may represent a combination of program steps and logic circuits, modules, and functions. A computer program may be stored in a memory. The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as but not limited to read-only memory (ROM), random access memory (RAM), optical memory devices and systems (such as digital video disc (DVD) or compact disk (CD), etc.). The computer-readable medium may include a non-transitory storage medium. The data processor may be of any type suitable for the local technical environment, such as but not limited to general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and processors based on multi-core processor architectures.

[0241] By way of illustrative and non-limiting examples, detailed descriptions of exemplary embodiments of the present application have been provided above. However, various modifications and adjustments to the above embodiments will be apparent to those skilled in the art when considered in conjunction with the accompanying drawings and the claims, without departing from the scope of the present application. Therefore, the proper scope of the present application will be determined according to the claims.

Claims

1. A data processing method, applied to a first processing node, characterized in that Including: Obtain k source data blocks, where k is a positive integer greater than 1; Encode the k source data blocks according to encoding parameters to obtain n encoded data blocks, where n is a positive integer greater than k; The encoding parameters include: the first check quantity, the second check quantity, the third check quantity, an additional source sub-data block index table, the number of repetitions, and an encoding matrix.

2. The method according to claim 1, wherein Among the n encoded data blocks, there are k source data blocks, m1 first check data blocks, m2 second check data blocks, and m3 third check data blocks, where m1 is the first check quantity and m1 = 1, m2 is the second check quantity and is a positive integer, m3 is the third check quantity and is a positive integer, and n = k + m1 + m2 + m3; The additional source sub-data index table is an index table with g rows and m2 columns, where g is the number of source sub-data blocks included in each source data block; The size of the encoding matrix is m * k, where m is the total number of check data blocks, and m = m1 + m2 + m3; The number of repetitions is a positive integer; The elements in the encoding matrix are elements in the Galois field GF(q), where q is an integer power of a prime number.

3. The method according to claim 2, wherein The k source data blocks respectively correspond to the data parts stored in the first k storage nodes in a stripe; The k source data blocks are distributed in k different storage nodes; The m check data blocks obtained by encoding the k source data blocks according to the encoding parameters respectively correspond to the data parts stored in the last m storage nodes in a stripe; The m check data blocks are respectively stored in m different storage nodes other than the k different storage nodes.

4. The method according to claim 2, characterized in that, The number of rows of the array of the n encoded data blocks is determined by at least one of the following parameters: k, the first check quantity, the second check quantity; The number of rows of the array of the n encoded data blocks is equal to the number of source sub-data blocks included in each source data block.

5. The method according to claim 4, wherein The number of rows of the array is: where s is the number of repetitions and s is a positive integer.

6. The method according to claim 2, wherein The encoding matrix satisfies at least one of the following: There is a row with all element values being "1"; Any square matrix formed by m rows and m columns is invertible in the Galois field GF(q).

7. The method according to claim 2, characterized in that Each check sub-data block in the first check data block is obtained by bitwise XOR of the source sub-data blocks in the row where the array of the n encoded data blocks is located.

8. The method according to claim 2, wherein Each check sub-data block in the second check data block is obtained by bitwise XOR of the linear combination value of the source sub-data blocks in the row where the first part of the array of the n encoded data blocks is located and the linear combination value of the source sub-data blocks in at least one other row of the second part of the array.

9. The method according to claim 8, wherein The linear combination value of the source sub-data blocks in the row where the first part of the array is located corresponding to the i-th check sub-data block of the j-th second check data block is determined by the source sub-data blocks in the i-th row of the array of the n encoded data blocks and the j-th row in the encoding matrix.

10. The method according to claim 8, wherein The linear combination value of the source sub-data blocks in at least one other row of the second part of the array is determined according to the additional source sub-data block index table and the number of repetitions.

11. The method according to claim 2, wherein Each element in the additional source sub-data block index table is a set of position coordinates, and the maximum number of one-dimensional coordinates or two-dimensional coordinates in the set of position coordinates is or 12. The method according to claim 11, wherein Each element in the additional source sub-data block index table is a set of two-dimensional coordinates; When all the index positions of the additional source sub-data blocks required to generate the second check data block are included in the additional source sub-data block index table, the maximum number of coordinates in the coordinate set is When the additional source sub-data block index table includes the index positions of the additional source sub-data blocks required to partially generate the second check data block, the maximum number of coordinates in the coordinate set is 13. The method according to claim 2, characterized in that, The additional source sub-data block index table is determined according to k, m1, and m2.

14. The method according to claim 13, characterized in that, The coordinates index of the sub-data block in the Gj-th row and the i-th column of the array of n coded data blocks are respectively in the t-th column and the Gi-th row of the additional source sub-data block index table; where i is an integer greater than or equal to 1 and less than or equal to ceil(k / s); The Gi is a set of row index values formed by sequentially dividing the integer values {1, 2,..., g} into (m1 + m2) ceil(i / (m1+m2)) sub - parts, and starting from the (mod(i - 1, m1 + m2)+1)-th sub - part in order, reading one sub - part every (m1 + m2) sub - parts; the t - th element in the set GSet of the i - th source data block of the array is Gj; j belongs to the first integer set, and the first integer set is [floor(i / (m1 + m2)) + 1, ceil(i / (m1 + m2))]\i; t belongs to the second integer set, and the second integer set contains integers greater than or equal to 1 and less than or equal to the number of integers in the first integer set; The Gj is a set of row index values formed by sequentially dividing the integer values {1, 2,..., g} into (m1 + m2) ceil(j / (m1+m2)) sub - parts, starting from the (mod(j - 1, m1 + m2)+1)-th sub - part in order, and reading one sub - part every (m1 + m2) sub - parts t and j are respectively elements corresponding to the same index size in the first integer set and the second integer set to which they belong, and the elements in the first integer set and the second integer set are sorted from small to large; g is the number of rows of the array; ceil() means adjusting a numerical value to the smallest integer not less than the numerical value itself, floor() means adjusting a numerical value to the largest integer not greater than the numerical value itself, and mod() means taking the modulo operation; []\i means removing the element i from a set.

15. The method according to claim 2, wherein Each check sub-data block in the third check data block is obtained by a linear combination of the source sub-data blocks in the row where the array of the n coded data blocks is located, and the linear combination coefficients are obtained from the coding matrix and the linear combination coefficients are not all equal to one.

16. A data processing method, applied to a second processing node, characterized in that, Including: Obtaining the type and quantity of the faulty coded data blocks and part or all of the data of at least k non-faulty coded data blocks generated by the first processing node, where k is a positive integer greater than 1; Processing part or all of the data of the at least k non-faulty coded data blocks according to the type and quantity of the faulty coded data blocks to obtain the faulty coded data blocks.

17. The method according to claim 16, wherein The at least k non-faulty coded data blocks are obtained by the first processing node encoding k source data blocks according to the encoding parameters; The encoding parameters include: the first check quantity, the second check quantity, the third check quantity, the additional source sub-data block index table, the repetition times, and the coding matrix.

18. The method according to claim 17, wherein Obtaining part or all of the data of at least k non-faulty coded data blocks generated by the first processing node includes: When the type of the faulty coded data blocks includes at least one of the first check data block, the second check data block, and the third check data block, or when the number of faulty source data blocks is greater than or equal to the sum of the first check quantity and the real-time second check quantity, obtaining all non-faulty coded data blocks generated by the first processing node; When only the source data blocks are faulty and the number of faulty source data blocks is less than the sum of the first check quantity and the second check quantity, obtaining part of the non-faulty coded data blocks generated by the first processing node.

19. The method according to claim 17, characterized in that Processing part or all of the data of the at least k non-faulty coded data blocks according to the type and quantity of the faulty coded data blocks includes at least one of the following: When the number of faulty coded data blocks exceeds the sum of the first check quantity, the real-time second check quantity, and the third check quantity, determining that the data of the faulty coded data blocks is lost; When the number of faulty coded data blocks is greater than 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the types of faulty coded data blocks include source data blocks and check data blocks, all non-faulty coded data blocks are obtained, and the faulty source data blocks are first recovered, and then the check data blocks are recovered; When the number of faulty coded data blocks is greater than 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the types of faulty coded data blocks only include source data blocks, at least a first set ratio of non-faulty coded data blocks is obtained, and the faulty source data blocks are recovered, where the first set ratio is the ratio of the number of faulty coded data blocks in the sum of the first check quantity and the second check quantity; When the number of faulty coded data blocks is greater than or equal to 1, less than or equal to the sum of the first check quantity, the second check quantity, and the third check quantity, and the types of faulty coded data blocks only include check data blocks, all non-faulty data blocks are obtained, and the faulty source data blocks are recovered; When the number of faulty coded data blocks is equal to 1 and the types of faulty coded data blocks only include source data blocks, a second set ratio of non-faulty source data blocks, first check data blocks, and second check data blocks are respectively obtained, and the faulty source data blocks are recovered, where the second set ratio is the ratio of 1 in the sum of the first check quantity and the second check quantity.

20. A processing node, characterized in that, Comprising: A memory, and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method according to any one of claims 1-15.

21. A second processing node, characterized in that, Comprising: A memory, and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method according to any one of claims 16-19.

22. A data processing system, characterized in that, Comprising: A set of storage nodes, a first processing node according to claim 20, and a second processing node according to claim 21; The set of storage nodes contains n storage nodes, and the n storage nodes include k source data storage nodes and n-k check data storage nodes, where n is a positive integer greater than k, and k is a positive integer greater than 1.

23. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the data processing method according to any one of claims 1-19.