Code construction method, coding method, node repair method and data processing device
By constructing multi-row matrix blocks and check matrix blocks, the read skip problem and finite domain scale requirements in the node repair process in distributed storage systems are solved, low bandwidth repair and low complexity encoding and decoding are realized, and the applicability and efficiency of the system are improved.
Patent Information
- Application Number
- CN202410114392.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, distributed storage systems have serious read skip problems and limited domain scale requirements during node repair, which cannot meet the actual application requirements, resulting in high repair bandwidth and high encoding and codec complexity.
Construct a multi-row matrix block, including n a×a initial matrices sorted by columns. The matrix block contains coefficients on the diagonal line and a/2 coupling coefficients. The coupling coefficients are located in continuous a/2 rows and different columns. The check matrix is obtained by combining multi-row matrix blocks, which is used for encoding and repair, reducing the node repair bandwidth and reducing the codec complexity.
Effectively reduce node repair bandwidth, avoid reading skip problems, and use small-scale finite domains to construct long codewords, improve application adaptability, reduce encoding and decoding complexity, and meet large-scale encoding and storage needs.
Smart Images

Figure CN120377934A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of storage technologies, and in particular, to a method for constructing a code, an encoding method, a node repair method, and a data processing device. Background Art
[0002] With the development of technologies in fields such as the Internet, communication, artificial intelligence, and machine learning, data has grown exponentially. To meet the increasing data storage requirements and solve the problem of efficient data access in large-scale and high-concurrency scenarios, a distributed storage system has been proposed. The distributed storage system includes multiple interconnected nodes (such as servers), and data is stored in different nodes. Since multiple nodes are usually distributed in different regions and are independent of each other, data loss caused by various natural disasters or emergencies can be effectively reduced. Nodes in the distributed storage system are prone to temporary or permanent failures due to reasons such as disk failures, power outages, network disconnections, and server crashes. When a certain node fails, it is necessary to download data from surviving nodes to restore the data in the failed node, and then store the restored data in a new replacement node to ensure data integrity and the reliability of the storage system.
[0003] Related technologies have proposed an elastic transformation encoding and node repair method based on Reed-Solomon (RS) codes. In this method, the symbols stored in a node are divided into multiple square arrays, and the symbols symmetric about the diagonal in the square array are coupled to obtain the final codeword. When performing node repair, downloading mutually coupled data from surviving nodes can restore some of the original data, and then using the original data to restore the original codeword to obtain some of the data in the failed node. Finally, combining the coupled data of other surviving nodes to restore the other data of the failed node.
[0004] However, in related technologies, when the number of sub-packets is relatively large, the skip-reading problem during data download in the node repair process is relatively serious, thus unable to meet the actual application requirements. And to ensure that the final codeword satisfies the maximum distance separable (MDS) property, the coupling coefficient during data coupling must exist in a large finite field, so a relatively large scale of the finite field is required, but the actual application cannot meet the scale requirements of the finite field. Summary of the Invention
[0005] The present application provides a method for constructing a code, an encoding method, a node repair method, and a data processing device, which solve the problems in related technologies that the serious skip-reading problem cannot meet the actual application requirements and the actual application cannot meet the scale requirements of the finite field, and can effectively reduce the repair bandwidth of nodes and reduce the encoding and decoding complexity without skip-reading.
[0006] In a first aspect, the present application provides a method for constructing a code. The method includes: constructing multiple rows of matrix blocks, where a matrix block includes n initial matrices of a×a sorted by columns. The initial matrix in the first matrix block includes coefficients on the diagonal and a / 2 coupling coefficients. The a / 2 coupling coefficients are located in consecutive a / 2 rows and the columns where the a / 2 coupling coefficients are located are different. The first matrix block belongs to the multiple rows of matrix blocks, and the initial matrix in a non-first matrix block is a diagonal matrix; combining the multiple rows of matrix blocks to obtain a parity-check matrix, which is used to encode and repair the original data block; where the first matrix block includes initial matrices in three modes. For the initial matrix in any one of the three modes, the a / 2 coupling coefficients are located in rows r to r + a / 2 - 1, and in the initial matrices of the remaining two modes, only non-zero elements exist in columns r to r + a / 2 - 1 in rows r to r + a / 2 - 1, and r ≤ a / 2 + 1.
[0007] Where n represents the code length and a represents the number of sub-packets. The number of rows of the matrix block is the number of parity-check blocks. The value of a can be a non-zero power of 2. The n initial matrices in the matrix block correspond one-to-one with n nodes, and the multiple rows of matrix blocks correspond one-to-one with multiple parity-check blocks.
[0008] The coefficients on the diagonal in a single initial matrix can be the same or different, and the values of the a / 2 coupling coefficients in a single initial matrix in the first matrix block can be the same or different. Any two initial matrices in the first matrix block are different.
[0009] The beneficial effect is that the positions of the coupling coefficients are explicitly constructed, eliminating the need to calculate the positions of the coupling coefficients by searching nodes, reducing the complexity of encoding and decoding. During the node repair process, all data (i.e., the aforementioned symbols) need to be downloaded from the surviving nodes with the same mode as the initial matrix corresponding to the failed node, and only consecutive half of the data needs to be downloaded from other surviving nodes. The number of nodes corresponding to the initial matrices of the same mode in the embodiments of the present application is small, thereby effectively reducing the repair bandwidth of the nodes, and the downloaded data is all consecutive without skip-reading problems. In addition, the coefficients on the diagonal of the n initial matrices in the first matrix block can be reused, that is, there are multiple initial matrices with the same coefficients on the diagonal. This can not only reduce the number of multiplication calculations in encoding and decoding (including encoding and decoding) when using the parity-check matrix subsequently, that is, reduce the complexity of encoding and decoding. Moreover, it can construct longer codewords using a relatively small finite field, effectively improving the application adaptability of the method and meeting the large-scale encoding storage requirements.
[0010] In a possible implementation, the initial matrices of the three modes include: a / 2 coupling coefficients are distributed in the upper right a / 2×a / 2 square matrix of the initial matrix, a / 2 coupling coefficients are distributed in the lower left a / 2×a / 2 square matrix of the initial matrix, and among a / 2 coupling coefficients, a / 4 coupling coefficients are distributed in the upper left a / 2×a / 2 square matrix of the initial matrix and the remaining a / 4 coupling coefficients are distributed in the lower right a / 2×a / 2 square matrix of the initial matrix.
[0011] In a possible implementation, the values of a / 2 coupling coefficients of n initial matrices in the first matrix block are the same.
[0012] The beneficial effect is that when encoding and decoding, for the first matrix block, the elements on the diagonal can be encoded and decoded first, and then the coupling coefficients can be additionally encoded and decoded. The coupling coefficients of the n initial matrices in the first matrix are the same value, so the intermediate results of the previous encoding and decoding can be utilized, and finally, only a few simple XORs and multiplications are required to complete the encoding and decoding, further reducing the encoding and decoding complexity and lowering the computational overhead.
[0013] In a possible implementation, the initial matrix in the second matrix block is an identity matrix, and the second matrix block belongs to a multi-row matrix block.
[0014] The beneficial effect is that it can improve the degraded read performance during encoding and decoding.
[0015] In a possible implementation, the modes of the initial matrices with the same result of taking the remainder of the permutation serial number modulo 3 in the first matrix block are the same, and the modes of the initial matrices with different results of taking the remainder of the permutation serial number modulo 3 are different.
[0016] In a second aspect, the present application provides a node repair method, which includes: determining the check equation corresponding to each surviving node among n nodes based on a check matrix, where the check matrix is constructed by the method according to any item in the first aspect; restoring the data in the failed nodes among the n nodes according to the check equation corresponding to each surviving node and the data stored in the surviving nodes; where the n nodes correspond one-to-one to the n initial matrices in the matrix block, a / 2 coupling coefficients of the initial matrix corresponding to the failed node in the first matrix block are located in the r-th row to the r+a / 2−1-th row, and the check equation corresponding to the j-th node among the n nodes includes: the r-th row to the r+a / 2−1-th row of the j-th initial matrix in the first matrix block, and the r-th row to the r+a / 2−1-th row of the j-th initial matrix in any non-first matrix block, 1≤j≤n, r≤a / 2+1.
[0017] n nodes correspond one-to-one with n initial matrices in the matrix blocks of the parity-check matrix, and the i-th node corresponds to the n-th initial matrix in the matrix block. When a certain node fails, the data stored in the failed node can be recovered by a parity-check equations through the regeneration method (each surviving node transfers part or all of the data).
[0018] In a third aspect, the present application provides an encoding method, which includes: obtaining a parity-check matrix constructed by the method according to any item in the first aspect; encoding the original data block based on the parity-check matrix to obtain n encoded data blocks.
[0019] In a fourth aspect, the present application provides a data processing device, which includes: a construction module for constructing multiple rows of matrix blocks, where the matrix blocks include n a×a initial matrices sorted by columns. The initial matrices in the first matrix block include coefficients on the diagonal and a / 2 coupling coefficients, and the a / 2 coupling coefficients are located in consecutive a / 2 rows and the columns where the a / 2 coupling coefficients are located are different. The first matrix block belongs to the multiple rows of matrix blocks, and the initial matrices in non-first matrix blocks are diagonal matrices; a combination module for combining the multiple rows of matrix blocks to obtain a parity-check matrix, which is used for encoding and repairing the original data block; where the first matrix block includes initial matrices in three modes, and the a / 2 coupling coefficients in the initial matrix in any one of the three modes are located in the r-th row to the r+a / 2−1-th row, and only non-zero elements exist in the r-th column to the r+a / 2−1-th column in the r-th row to the r+a / 2−1-th row of the initial matrices in the remaining two modes, r≤a / 2+1.
[0020] In a possible implementation, the initial matrices in the three modes include: a / 2 coupling coefficients are distributed in the upper-right a / 2×a / 2 square matrix of the initial matrix, a / 2 coupling coefficients are distributed in the lower-left a / 2×a / 2 square matrix of the initial matrix, and among the a / 2 coupling coefficients, a / 4 coupling coefficients are distributed in the upper-left a / 2×a / 2 square matrix of the initial matrix and the remaining a / 4 coupling coefficients are distributed in the lower-right a / 2×a / 2 square matrix of the initial matrix.
[0021] In a possible implementation, the values of the a / 2 coupling coefficients of the n initial matrices in the first matrix block are the same.
[0022] In a possible implementation, the initial matrix in the second matrix block is an identity matrix, and the second matrix block belongs to the multiple rows of matrix blocks.
[0023] In a possible implementation, the initial matrices with the same result of taking the modulo of the permutation serial number by 3 in the first matrix block have the same mode, and the initial matrices with different results of taking the modulo of the permutation serial number by 3 have different modes.
[0024] Fifth aspect, the present application provides a data processing device, which includes: a processing module, configured to determine a check equation corresponding to each surviving node among n nodes based on a check matrix, where the check matrix is constructed according to the method of any one of the first aspect; a repair module, configured to restore the data in the failed node among the n nodes according to the check equation corresponding to each surviving node and the data stored in the surviving nodes; wherein, the n nodes correspond one-to-one to the n initial matrices in the matrix block, a / 2 coupling coefficients of the initial matrix corresponding to the failed node in the first matrix block are located in the r-th row to the r + a / 2 - 1-th row, and the check equation corresponding to the j-th node among the n nodes includes: the r-th row to the r + a / 2 - 1-th row of the j-th initial matrix in the first matrix block, and the r-th row to the r + a / 2 - 1-th row of the j-th initial matrix in any non-first matrix block, 1 ≤ j ≤ n, r ≤ a / 2 + 1.
[0025] Sixth aspect, the present application provides a data processing device, which includes: an acquisition module, configured to acquire a check matrix, where the check matrix is constructed according to the method of any one of the first aspect; an encoding module, configured to encode an original data block based on the check matrix to obtain n encoded data blocks.
[0026] Seventh aspect, the present application provides a data processing device, which includes: one or more processors; a memory, configured to store one or more computer programs or instructions; when the one or more computer programs or instructions are executed by the one or more processors, the one or more processors implement the method of any one of the first aspect, or implement the method of any one of the second aspect, or implement the method of any one of the third aspect.
[0027] Eighth aspect, the present application provides a data processing device, including a processor, configured to execute the method of any one of the first aspect, or execute the method of any one of the second aspect, or execute the method of any one of the third aspect.
[0028] Ninth aspect, the present application provides a data processing device, which includes: a processing circuit and an interface circuit; wherein, the interface circuit is configured to be coupled to a memory external to the data processing device and provide a communication interface for the processing circuit to access the memory; the processing circuit is configured to execute program instructions in the memory to implement the method of any one of the first aspect, or implement the method of any one of the second aspect, or implement the method of any one of the third aspect.
[0029] In the specific implementation process, the data processing device may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be transistors, gate circuits, flip-flops, and various logic circuits, etc. The input signal received by the input circuit may be received and input by, for example, but not limited to, a receiver. The signal output by the output circuit may be output to, for example, but not limited to, a transmitter and transmitted by the transmitter. Moreover, the input circuit and the output circuit may be the same circuit, which serves as the input circuit and the output circuit at different times respectively. The embodiments of the present application do not limit the specific implementation manners of the processor and various circuits.
[0030] In a tenth aspect, the present application provides a computer-readable storage medium. Program code is stored in the computer-readable storage medium. When the program code is executed by a processor, the method described in any one of the first aspect to the third aspect is implemented.
[0031] In an eleventh aspect, the present application provides a chip, including: at least one processor. The at least one processor is used to execute the method described in any one of the first aspect to the third aspect.
[0032] Optionally, the chip further includes a memory. The at least one processor is used to execute the code in the memory. When the at least one processor executes the code, the chip implements the method described in any one of the first aspect to the third aspect.
[0033] Optionally, the above-mentioned chip may also be an integrated circuit.
[0034] In a twelfth aspect, the present application provides a computer program product containing instructions. When it runs on a computer, the computer implements the method described in any one of the first aspect to the third aspect. Description of the Drawings
[0035] Figure 1 It is a schematic diagram of an erasure code provided by an embodiment of the present application;
[0036] Figure 2 It is a schematic diagram of the construction of an optimal read minimum storage regeneration code provided by an embodiment of the present application;
[0037] Figure 3 It is a schematic diagram of the architecture of a distributed storage system provided by an embodiment of the present application;
[0038] Figure 4 It is a schematic diagram of the flowchart of a method for constructing a code provided by an embodiment of the present application;
[0039] Figure 5 It is a schematic diagram of the flowchart of an encoding method provided by an embodiment of the present application;
[0040] Figure 6Schematic flowchart of a node repair method provided by an embodiment of the present application;
[0041] Figure 7 Block diagram of a data processing device provided by an embodiment of the present application;
[0042] Figure 8 Block diagram of another data processing device provided by an embodiment of the present application;
[0043] Figure 9 Block diagram of yet another data processing device provided by an embodiment of the present application;
[0044] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0045] Figure 11 Schematic structural diagram of a data processing device provided by an embodiment of the present application. Detailed implementation manners
[0046] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.
[0047] Terms such as "first" and "second" in the description, claims and drawings of the present application are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a series of steps or units included. A method, system, product or device is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0048] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0049] Compared with the servers in centralized storage, the frequency of node failure events in a distributed storage system is relatively high, and in more than 90% of the cases, it is single-node failure, where single-node failure means that only one node fails within a fixed period of time. Therefore, it is extremely important to reduce the repair cost of node failure events while storing data efficiently and reliably.
[0050] To ensure the reliability of distributed storage, the system will encode the original data before storing it, and redundant data will be generated after encoding. As an encoding scheme, erasure codes can effectively reduce the storage overhead caused by redundant data. In erasure codes, the original data is divided into k parts, and the k parts of the original data are encoded through a parity-check matrix to obtain n data blocks. The number of parity blocks (i.e., redundant data) is n - k = m, and the system can support up to m nodes failing simultaneously. Erasure codes satisfy the MDS property, and the data in any k surviving nodes can be used to recover the data in the failed nodes.
[0051] Taking the MDS code of [4, 2] as an example, n = 4 and k = 2. Please refer to Figure 1 , Figure 1 which is a schematic diagram of an erasure code provided by an embodiment of this application, Figure 1 shows 4 nodes (node 1 to node 4). The original data includes data block 1 and data block 2, and data block 1 and data block 2 are stored in node 1 and node 2 respectively. After encoding by the [4, 2] erasure code, the computing node obtains parity block 1 and parity block 2, and parity block 1 and parity block 2 are stored in node 3 and node 4 respectively. At this time, the system supports up to 2 nodes failing simultaneously, and the storage overhead is 200%. Assuming that node 2 fails, the computing node can recover the data in node 2 from the data in any two surviving nodes. Figure 1 Taking the computing node recovering the data in node 2 from the data in node 1 and node 3 as an example for illustration.
[0052] A regenerated code is a special erasure code that can select an optimal trade-off between storage overhead and repair bandwidth. The repair bandwidth refers to the total amount of data downloaded from surviving nodes during the repair process of a failed node. The original data of size B is encoded by a code of [n, k, d; α, β, B] to obtain nα symbols, and the nα symbols are stored in n different nodes respectively, with each node storing α symbols. When a certain node fails, β symbols can be downloaded from any d surviving nodes respectively to recover the data in the failed node.
[0053] Among them, the storage overhead is nα / B, and the repair bandwidth is dβ. The value of d is at least k and at most n - 1. The relationship between the storage overhead and the repair bandwidth satisfies the following formula:
[0054]
[0055] The aforementioned formula is also called the cut-set bound, and the code with parameters falling on the cut-set bound is called a regenerated code. The two endpoints of this curve correspond to the lowest storage overhead and the lowest repair bandwidth respectively. The regenerated code with the lowest storage overhead is called the minimum storage regenerating code (MSR code), and the regenerated code with the lowest repair bandwidth is called the minimum bandwidth regenerating code (MBR code).
[0056] Related technologies provide various construction methods for MBR codes and MSR codes with all parameters. Related technology one provides a minimum storage regenerating code with optimal reading. For a regenerated code with a code length of n and a dimension of k, this method can construct an optimal reading regenerated code type with a sub-packet number level of . Taking the regenerated code of [4, 2; 4] as an example, n = 4, k = 2, and α = 4.
[0057] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the construction of a minimum storage regenerating code with optimal reading provided by an embodiment of this application, Figure 2 showing 4 nodes (node a, node b, node p, and node q). The original data includes data block a and data block b. Data block a is divided into 4 symbols (a1, a2, a3, a4) and stored in node a, and data block b is divided into 4 symbols (b1, b2, b3, b4) and stored in node b. Figure 2 shows the parity-check matrix. The computing nodes encode based on this parity-check matrix through the erasure code of [4, 2; 4] to obtain parity-check block p and parity-check block q. Parity-check block p includes 4 symbols (p1, p2, p3, p4) and is stored in node p, and parity-check block q includes 4 symbols (q1, q2, q3, q4) and is stored in node q.
[0058] Among them, p1 = 3a2 + 3a3 + 3a4 + 3b1 + b3, p2 = 3a4 + 3b1 + 3b2 + 3b3 + b4, p3 = 4a1 + 2a2 + 3a3 + 3a4 + b1 + b3, p4 = 4a2 + 3a4 + 2b1 + b2 + 3b3 + b4.
[0059] q1 = 4a1 + 2a2 + 2a3 + 2a4 + b1 + 4b3, q2 = 4a2 + 2a4 + 2b1 + b2 + 2b3 + 4b4, q3 = a1 + 3a2 + a3 + 2a4 + 4b1 + 4b3, d4 = a2 + a4 + 3b1 + 4b2 + 2b3 + 3b4.
[0060] When node a fails, download b1 and b2 from node b, download p1 and p3 from node p, and download q1 and q3 from node q, then the data in node a can be restored through the following calculation method:
[0061] a1 = 4p1 + 4q1 + 4b1
[0062] a2 = 2p1 + 4p3 + q1 + 4b1
[0063] a3 = 4p3 + 4q3 + 4b3
[0064] a4 = 2p3 + 4q1 + q3 + 4b1 + 4b3
[0065] Similarly, as Figure 2 shown, when node b fails, download a2 and a4 from node a, download p2 and p4 from node p, and download q2 and q4 from node q, then the data in node b can be restored. When node p fails, download a1 and a2 from node a, download b1 and b2 from node b, and download q1 and q2 from node q, then the data in node p can be restored. When node q fails, download a3 and a4 from node a, download b3 and b4 from node b, and download p3 and p4 from node p, then the data in node q can be restored.
[0066] It can be seen from this that serious skip-reading problems will occur when downloading data from surviving nodes during the process of repairing node a and node b. That is, when repairing a failed node, it is necessary to read data that is not sequentially placed from the surviving nodes, which will seriously affect the disk throughput. For example, in a serial advanced technology attachment (SATA) hard disk, the throughput of non-sequential reading of a 4-kilobyte (KB) disk block is only 1000 KB / s, while the throughput of sequential reading can reach more than 100 MB / s. And in order to make the repair bandwidth reach the aforementioned cut-set bound, the number of sub-packets needs to be exponential. The excessive number of sub-packets will lead to a higher complexity in the encoding and node repair processes, making it impossible to be applied to an actual distributed storage system.
[0067] Related technology two provides a Piggybacking framework that couples the encoded codewords to obtain the final codewords. Taking the [6, 3] RS code as an example, and the number of sub-packets being 2, take two [6, 3] codewords (a1, a2, a3, p1, p2, p3) and (b1, b2, b3, q1, q2, q3). Couple the original data of the first codeword to the parity symbols of the second codeword. The data stored in nodes 1 to 6 is as shown in Table 1 below:
[0068] Table 1
[0069] Node 1 <![CDATA[a1]]> <![CDATA[b1]]> Node 2 <![CDATA[a2]]> <![CDATA[b2]]> Node 3 <![CDATA[a3]]> <![CDATA[b3 <!-- 6 -->]]> Node 4 <![CDATA[p1]]> <![CDATA[q1]]> Node 5 <![CDATA[p2]]> <![CDATA[q2 + a1 + a2]]> Node 6 <![CDATA[p3]]> <![CDATA[q3 + a3]]>
[0070] When node 1 fails, download b2, b3, and q1, and then use the MDS property of the RS code to recover b1, q2, and q3. Then download q2 + a1 + a2 and a2 in node 5 to recover a1, thereby recovering the data in node 1. However, in related technology two, the repair bandwidth is large, and as the number of sub-packets increases, more and more serious skip-reading problems will occur during the repair process of some nodes, making it impossible to be applied to an actual distributed storage system.
[0071] Related technology three is the elastic change method described above. When the number of sub-packets is less than or equal to n - k, taking two sub-packets as an example, take the codewords (a1, a2, p1, p2) and (b1, b2, q1, q2) in two [4, 2] RS codes. The four symbols in the codewords are respectively stored in nodes 1 to 4. Divide node 1 and node 2 into a group, and node 3 and node 4 into a group. Couple b1 in node 1 with a2 in node 2, and couple p2 in node 3 with q1 in node 2. The data stored in nodes 1 to 4 is as shown in Table 2 below:
[0072] Table 2
[0073] Node 1 <![CDATA[a1]]> <![CDATA[a2 + γb1]]> Node 2 <![CDATA[a2+b1]]> <![CDATA[b2]]> Node 3 <![CDATA[p1]]> <![CDATA[p2 + γq1]]> Node 4 <![CDATA[p2+q1]]> <![CDATA[q2]]>
[0074] When node 1 fails, download p1 and p2 + γq1 from node 3 and p2 + q1 from node 4 to recover p1 and p2. a1 and a2 can be recovered through the MDS property of the RS code. Then download a2 + b1 from node 2, and combine a2 + b1 and the recovered a2 to recover a2 + γb1, thereby recovering the data in node 1.
[0075] When the number of sub-packets is greater than n - k, elastic transformation is performed in an iterative manner. Taking 4 sub-packets as an example, the elastic transformation is divided into two steps. First, perform elastic transformation on node 1 and node 2. Then, the first two nodes use the direct sum method to transform to generate two additional symbols, and the last two nodes use the elastic transformation method. Finally, the data stored in nodes 1 to 4 is shown in Table 3 as follows:
[0076] Table 3
[0077] Node 1 <![CDATA[a1]]> <![CDATA[a2 + γb1]]> <![CDATA[c1]]> <![CDATA[c2+γd1]]> Node 2 <![CDATA[a2+b1]]> <![CDATA[b2]]> <![CDATA[c2+d1]]> <![CDATA[d2]]> Node 3 <![CDATA[p1]]> <![CDATA[p2+γq1]]> <![CDATA[p2 + γr1]]> <![CDATA[q2+γt1]]> Node 4 <![CDATA[p2+r1]]> <![CDATA[q2+t1]]> <![CDATA[r2]]> <![CDATA[t2]]>
[0078] When node 1 fails, download p1, p2 + γr1 from node 3, p2 + r1 and r2 from node 4 to recover p1, p2, r1 and r2. a1, a2, c1 and c2 can be recovered through the MDS property of the RS code. Then download a2 + b1 and c2 + d1 from node 2, thereby recovering the data in node 1. As can be seen from the foregoing process, there is still a skip-reading problem in Related Art Three, and there is a high demand for the scale of the finite field, which cannot be applied to an actual distributed storage system.
[0079] The embodiment of the present application provides a method for constructing a code. Compared with the foregoing related technologies, it effectively reduces the repair bandwidth and does not generate a skip-reading problem during the node repair process. This method can be applied to a distributed storage system, and the constructed parity-check matrix can be used in the encoding process and node repair process of the distributed storage system. The encoding process can be, for example, an erasure code encoding process.
[0080] Exemplarily, please refer to Figure 3 , Figure 3 which is a schematic diagram of the architecture of a distributed storage system provided by the embodiment of the present application, Figure 3 showing the encoding process of the [6, 4] erasure code of the distributed storage system. As Figure 3As shown, first, the file is divided into data chunks to obtain the original data chunks F1, F2, F3, and F4. Then, erasure codes are used to encode the foregoing four original data chunks to obtain parity data chunks (also known as redundant data chunks) P1 and P2. Next, F1, F2, F3, F4, P1, and P2 are written into different nodes for disk storage respectively. When data loss occurs due to node failure, valid original data chunks and / or parity chunks are downloaded from surviving nodes for node repair (also known as erasure code decoding) to recover the lost data. Figure 3 Taking the loss of F1 and F2 as an example, the lost data chunks (including F1' and F2') are recovered by downloading F3, F4, P1, and P2, and then the recovered data chunks are rewritten into the repair node to ensure the storage reliability of the distributed storage system.
[0081] Please refer to Figure 4 , Figure 4 is a schematic flowchart of a method for constructing a code provided by an embodiment of the present application. The code may include erasure codes, for example, may include MDS array codes. The parity check matrix is constructed by this method, and the parity check matrix is used to encode the original data (such as erasure code encoding) to obtain parity data in the encoding layer of the distributed storage system, and to repair the data in the failed node (also known as decoding). This method can be applied to a distributed storage system, specifically, can be applied to computing nodes in the distributed storage system, and the parity check matrix can be stored in the computing nodes. The computing nodes can be encoding / decoding computing nodes, for example, can be server devices based on software computing or chips based on hardware computing, etc., and the embodiments of the present application do not limit this. As Figure 4 shown, this method may include the following process:
[0082] 101. Construct multiple rows of matrix blocks. The matrix blocks include n a×a initial matrices sorted by columns. The initial matrix in the first matrix block includes coefficients on the diagonal and a / 2 coupling coefficients. The a / 2 coupling coefficients are located in consecutive a / 2 rows and the columns where the a / 2 coupling coefficients are located are different. The first matrix block belongs to the multiple rows of matrix blocks. The initial matrices in non-first matrix blocks are diagonal matrices. Among them, the first matrix block includes initial matrices in three modes. In any one of the three modes, the a / 2 coupling coefficients in the initial matrix are located in rows r to r + a / 2 - 1, and in the remaining two modes of the initial matrices, only non-zero elements exist in columns r to r + a / 2 - 1 in rows r to r + a / 2 - 1, where r ≤ a / 2 + 1.
[0083] In the a / 2 rows where the a / 2 coupling coefficients of each mode of the initial matrix are located, the column where any coupling coefficient is located is different from the column where the coefficient on the diagonal is located.
[0084] When encoding, it is necessary to pre-configure encoding parameters, which include code length, dimension, and the number of sub-packets, etc. The code length represents the number of original data blocks and parity blocks after encoding (i.e., the number of occupied nodes), the dimension represents the number of original data blocks, and the number of sub-packets represents the number of symbols included in each data block or parity block. The encoding parameters can be pre-configured by the user or the distributed storage system. In this process 101, n represents the code length and a represents the number of sub-packets. The number of rows of the matrix block is the number of parity blocks. The value of a can be a non-zero power of 2. The n initial matrices in the matrix block correspond one-to-one with n nodes, and multiple rows of matrix blocks correspond one-to-one with multiple parity blocks.
[0085] The coefficients on the diagonal in a single initial matrix can be the same or different. The values of a / 2 coupling coefficients in a single initial matrix in the first matrix block can be the same or different. The coefficients on the diagonal can be selected from a preset number of values, and the coupling coefficients are non-zero indeterminates.
[0086] Any two initial matrices in the first matrix block are different. Specifically, when the coefficients on the diagonal of two initial matrices in the first matrix block are the same, the patterns of the two initial matrices are different and the values of a / 2 coupling coefficients are the same, or the patterns of the two initial matrices are different and the values of a / 2 coupling coefficients are different, or the patterns of the two initial matrices are the same and the values of a / 2 coupling coefficients are different. When the coefficients on the diagonal of two initial matrices in the first matrix block are different, the patterns of the two initial matrices are the same and the values of a / 2 coupling coefficients are the same, or the patterns of the two initial matrices are the same and the values of a / 2 coupling coefficients are different, or the patterns of the two initial matrices are different and the values of a / 2 coupling coefficients are the same, or the patterns of the two initial matrices are different and the values of a / 2 coupling coefficients are different.
[0087] Due to the existence of coupling coefficients in the first matrix block, the coefficients on the diagonal of the n initial matrices can be reused, that is, there are multiple initial matrices with the same coefficients on the diagonal. In this way, when performing encoding and decoding (including encoding and decoding) using the parity matrix later, the encoding and decoding complexity can be reduced.
[0088] Exemplarily, the values of a / 2 coupling coefficients of the n initial matrices in the first matrix block can be all the same. When performing encoding and decoding, for the first matrix block, first, the elements on the diagonal can be encoded and decoded, and then the coupling coefficients can be additionally encoded and decoded. The coupling coefficients of the n initial matrices in the first matrix are the same value, so the intermediate results of the previous encoding and decoding can be utilized, and finally, only a few simple XOR and multiplications are required to complete the encoding and decoding, further reducing the encoding and decoding complexity and lowering the computational overhead.
[0089] In the embodiments of the present application, the initial matrix in the second matrix block belonging to the multi-row matrix block is an identity matrix. The identity matrix can improve the degraded read performance during encoding and decoding.
[0090] If the coefficients on the diagonal of a single initial matrix are the same, in all matrix blocks except the second matrix block in the multi-row matrix block, the elements on the diagonal of the s-th initial matrix increase exponentially according to the number of rows of the matrix block where they are located, where s is a positive integer less than or equal to n. At this time, all matrix blocks except the second matrix block can be regarded as Vandermonde matrices. For example, assume that the number of rows of the matrix block is 4, the second matrix block is the first-row matrix block, and the first matrix block is the second-row matrix block. If the elements on the diagonal of the s-th initial matrix in the second-row matrix block are all γ, then the elements on the diagonal of the s-th initial matrix in the third-row matrix block are all γ 2 , and the elements on the diagonal of the s-th initial matrix in the fourth-row matrix block are all γ 3 . In this way, other matrix blocks except the second matrix block are all pure diagonal matrices after removing the coupling coefficients. Therefore, during encoding and decoding, the encoding and decoding acceleration algorithm of erasure codes can be used for acceleration to improve the encoding and decoding efficiency. The encoding and decoding acceleration algorithm can include the fast Fourier acceleration algorithm, etc.
[0091] Exemplarily, the pattern of the initial matrix can be determined based on the arrangement serial number of the initial matrix in the first matrix block. For example, the arrangement serial number of the initial matrix can be modulo 3. Initial matrices with the same result of taking modulo 3 of the arrangement serial number have the same pattern, and initial matrices with different results of taking modulo 3 of the arrangement serial number have different patterns. Or, every three initial matrices in the first matrix block can be grouped as a set, and the initial matrices in a set respectively adopt three different patterns. If the number of initial matrices in the first matrix block is not a multiple of 3, the patterns of the remaining one or two initial matrices can be randomly selected. Or, the pattern of the initial matrix can be arbitrarily selected as long as any two initial matrices are different. The embodiments of the present application do not make any limitations in this regard.
[0092] The initial matrices of the three patterns can include: Pattern 1: a / 2 coupling coefficients are distributed in the upper right a / 2×a / 2 square matrix of the initial matrix; Pattern 2: a / 2 coupling coefficients are distributed in the lower left a / 2×a / 2 square matrix of the initial matrix; Pattern 3: Among a / 2 coupling coefficients, a / 4 coupling coefficients are distributed in the upper left a / 2×a / 2 square matrix of the initial matrix and the remaining a / 4 coupling coefficients are distributed in the lower right a / 2×a / 2 square matrix of the initial matrix.
[0093] Exemplarily, assume that it is necessary to construct a two-redundancy array code with 4 sub-packets (i.e., a = 4, n - k = 2), and the number of rows of the matrix block is n - k = 2. The initial matrices of the three modes included in the first matrix block may include: Mode 1: 2 coupling coefficients are distributed in the upper right 2×2 square matrix of the initial matrix; Mode 2: 2 coupling coefficients are distributed in the lower left 2×2 square matrix of the initial matrix; Mode 3: Among the 2 coupling coefficients, 1 coupling coefficient is distributed in the upper left 2×2 square matrix of the initial matrix and the remaining 1 coupling coefficient is distributed in the lower right 2×2 square matrix of the initial matrix. And in the 2 rows where the 2 coupling coefficients are located in the initial matrix of each mode, the column where any coupling coefficient is located is different from the column where the coefficient on the diagonal is located.
[0094] The 2-row matrix block can be as follows:
[0095]
[0096] Among them, H i corresponds to a check data block, and H 0,i and H 1,i The matrix formed corresponds to node i.
[0097] Assume that the initial matrices in the first row of the matrix block are all identity matrices, then the initial matrices in the first row of the matrix block can be as follows, 0 ≤ i ≤ n - 1:
[0098]
[0099] Assume that the second row of the matrix block is the first matrix block. Determine the mode of the initial matrix according to the arrangement serial number of the initial matrix in the first matrix block. Exemplarily, when i ≡ 0 (mod 3), the initial matrix can adopt the aforementioned Mode 1, that is, 2 coupling coefficients are distributed in the upper right 2×2 square matrix of the initial matrix. The initial matrix can be as follows, the 2 coupling coefficients are located in the 1st row and 4th column and the 2nd row and 3rd column respectively, 0 ≤ i ≤ n - 1, * represents the coupling coefficient:
[0100]
[0101] When i ≡ 1 (mod 3), the initial matrix can adopt the aforementioned Mode 2, that is, 2 coupling coefficients are distributed in the lower left 2×2 square matrix of the initial matrix. The initial matrix can be as follows, the 2 coupling coefficients are located in the 3rd row and 2nd column and the 4th row and 1st column respectively:
[0102]
[0103] When \(i\equiv2\pmod{3}\), the initial matrix can adopt the aforementioned Pattern 3, that is, one of the two coupling coefficients is distributed in the upper left \(2\times2\) square matrix of the initial matrix and the other coupling coefficient is distributed in the lower right \(2\times2\) square matrix of the initial matrix. The initial matrix can be as follows, and the two coupling coefficients are located in the first column of the second row and the fourth column of the third row respectively:
[0104]
[0105] In the initial matrices of the aforementioned three patterns, it can be seen that in each initial matrix, the two coupling coefficients are located in two consecutive rows, and the columns where the two coupling coefficients are located are different. And in the two rows where the two coupling coefficients are located, for any coupling coefficient, the column where it is located is different from the column where the coefficient on the diagonal is located. In the initial matrix of Pattern 1, the two coupling coefficients are located in the first row to the second row, then in the initial matrices of Pattern 2 and Pattern 3, there are only non-zero elements in the first column to the second column in the first row to the second row. In the initial matrix of Pattern 2, the two coupling coefficients are located in the third row to the fourth row, then in the initial matrices of Pattern 1 and Pattern 3, there are only non-zero elements in the third column to the fourth column in the third row to the fourth row. In the initial matrix of Pattern 3, the two coupling coefficients are located in the second row to the third row, then in the initial matrices of Pattern 1 and Pattern 2, there are only non-zero elements in the second column to the third column in the second row to the third row.
[0106] Exemplarily, assume that it is required to construct a triple-redundancy array code with the number of sub-packets being 8 (i.e., \(a = 8\), \(n - k = 3\)), and the number of rows of the matrix block is \(n - k = 3\). The initial matrices of the three patterns included in the first matrix block can include: Pattern 1: Four coupling coefficients are distributed in the upper right \(4\times4\) square matrix of the initial matrix; Pattern 2: Four coupling coefficients are distributed in the lower left \(4\times4\) square matrix of the initial matrix; Pattern 3: Two of the four coupling coefficients are distributed in the upper left \(4\times4\) square matrix of the initial matrix and the other two coupling coefficients are distributed in the lower right \(4\times4\) square matrix of the initial matrix. And in the four rows where the four coupling coefficients are located, for any coupling coefficient, the column where it is located is different from the column where the coefficient on the diagonal is located.
[0107] The \(3 -\)row matrix block can be as follows:
[0108]
[0109] Among them, \(H\) i corresponds to a check data block, and the matrix composed of \(H\) 0,i , \(H\) 1,i and \(H\) 2,i corresponds to node \(i\).
[0110] Assume that the initial matrices in the first row matrix block are all identity matrices, then the initial matrices in the first row matrix block can be as follows, \(0\leq i\leq n - 1\):
[0111]
[0112] Assume that the second row matrix block is the first matrix block, and determine the pattern of the initial matrix according to the arrangement serial number of the initial matrix in the first matrix block. For example, when i ≡ 0 (mod 3), the initial matrix can adopt the aforementioned pattern 1, that is, 4 coupling coefficients are distributed in the upper right 4×4 square matrix of the initial matrix. The initial matrix can be as follows. The 4 coupling coefficients are respectively located in the 1st row and 8th column, the 2nd row and 7th column, the 3rd row and 6th column, and the 4th row and 5th column, 0 ≤ i ≤ n - 1, * represents the coupling coefficient:
[0113]
[0114] When i ≡ 1 (mod 3), the initial matrix can adopt the aforementioned pattern 2, that is, 4 coupling coefficients are distributed in the lower left 4×4 square matrix of the initial matrix. The initial matrix can be as follows. The 4 coupling coefficients are respectively located in the 5th row and 4th column, the 6th row and 3rd column, the 7th row and 2nd column, and the 8th row and 1st column:
[0115]
[0116] When i ≡ 2 (mod 3), the initial matrix can adopt the aforementioned pattern 3, that is, 2 of the 4 coupling coefficients are distributed in the upper left 4×4 square matrix of the initial matrix and the remaining 2 coupling coefficients are distributed in the lower right 4×4 square matrix of the initial matrix. The initial matrix can be as follows. The 4 coupling coefficients are respectively located in the 3rd row and 1st column, the 4th row and 2nd column, the 5th row and 7th column, and the 6th row and 8th column:
[0117]
[0118] Or when i ≡ 0 (mod 3), the initial matrix can be as follows. The 4 coupling coefficients are respectively located in the 1st row and 7th column, the 2nd row and 8th column, the 3rd row and 5th column, and the 4th row and 6th column:
[0119]
[0120] When i ≡ 1 (mod 3), the initial matrix can be as follows. The 4 coupling coefficients are respectively located in the 5th row and 3rd column, the 6th row and 4th column, the 7th row and 1st column, and the 8th row and 2nd column:
[0121]
[0122] When i ≡ 2 (mod 3), the initial matrix can be as follows. The 4 coupling coefficients are respectively located in the 3rd row and 2nd column, the 4th row and 1st column, the 5th row and 8th column, and the 6th row and 7th column:
[0123]
[0124] It should be noted that the patterns corresponding to the above modular results and the positions of the coupling coefficients are all for illustrative purposes. When i ≡ 0 (mod 3), the initial matrix can also adopt Pattern 2 or Pattern 3. When i ≡ 1 (mod 3), the initial matrix can also adopt Pattern 1 or Pattern 3. The position of the coupling coefficient can also be other positions, as long as the conditions of the initial matrix of each pattern described above are satisfied. The embodiments of the present application do not make limitations in this regard.
[0125] 102. Combine multiple rows of matrix blocks to obtain a parity-check matrix, which is used to encode and repair the original data block.
[0126] The computing node combines multiple rows of matrix blocks in sequence and determines the values of the coupling coefficients to ensure that the MDS property of the codeword holds, thereby obtaining the parity-check matrix.
[0127] Referring to the example in the foregoing process 101, when the number of rows of the matrix block is 2, the parity-check matrix When the number of rows of the matrix block is 3, the parity-check matrix
[0128] From the form of the initial matrix in the foregoing process 101, it can be seen that there is a solution for the coupling coefficient in the finite field in the parity-check matrix such that the code satisfies the MDS property, that is, the matrix formed by any two columns of the initial matrix in the parity-check matrix (for example or ) is invertible.
[0129] The following further illustrates the parity-check matrix with specific numerical values.
[0130] Example 1: 4 sub-packets, [6, 4] regeneration code type, that is, k = 4, n = 6, the redundancy is n - k = 2, and the coupling coefficients are all 2. The following parity-check matrix is constructed in the finite field GF(4):
[0131]
[0132]
[0133]
[0134] In the foregoing two-row matrix block, H0 is the second matrix block, and the initial matrices H 0,i included therein are all identity matrices, where 1 ≤ i ≤ 6. H1 is the first matrix block, and H 1,1 and H 1,2 both adopt Pattern 1. The coefficient on the diagonal of H 1,1 is 0, and the coefficient on the diagonal of H 1,2 is 1. H 1,3 and H 1,4Both adopt Mode 2, H 1,3 The coefficients on the diagonal of are 0, H 1,4 The coefficients on the diagonal of are 1. H 1,5 and H 1,6 Both adopt Mode 3, H 1,5 The coefficients on the diagonal of are 2, H 1,6 The coefficients on the diagonal of are 3.
[0135] It can be seen from H1 that the coefficients on the diagonal of the initial matrices with different modes can be the same, that is, the values of the coefficients on the diagonal can be reused. Therefore, the existence of the coupling coefficient enables the construction of longer codewords using a finite field with a smaller scale. For example, in the first embodiment, codewords of length 6 can be constructed from GF(4), and the encoding and decoding complexity is relatively low. And the 8×8 matrix composed of any two columns of the initial matrix is invertible. Therefore, the MDS property of the codewords obtained based on this parity-check matrix holds.
[0136] Embodiment 2: 4 sub-packets, [24, 22] regenerating code type, that is, k = 22, n = 24, the redundancy is n - k = 2, and the coupling coefficients are all 8. The following parity-check matrix is constructed in the finite field GF(16):
[0137]
[0138] H0 is as follows:
[0139]
[0140] H1 is as follows:
[0141]
[0142] In the above two rows of matrix blocks, H0 is the second matrix block, and the included initial matrix H 0,i are all identity matrices, 1 ≤ i ≤ 24. H1 is the first matrix block, and H 1,1 to H 1,8 all adopt Mode 1, and H 1,1 to H 1,8 The coefficients on the diagonal are 0 to 7 respectively. H 1,9 to H 1,16 all adopt Mode 2, and H 1,9 to H 1,16 The coefficients on the diagonal are 0 to 7 respectively. H 1,17 to H 1,24 all adopt Mode 3, and H 1,17 to H 1,24 The coefficients on the diagonal are 8 to 15 respectively.
[0143] As can be seen from H1, the coefficients on the diagonal of the initial matrices with different patterns can be the same, that is, the values of the coefficients on the diagonal can be reused. Therefore, the existence of the coupling coefficient enables the construction of longer codewords using a finite field with a smaller scale. For example, in the second embodiment, a codeword of length 24 can be constructed from GF(16). Moreover, the 8×8 matrix formed by any two columns of the initial matrices is invertible, so the MDS property of the codewords obtained based on this parity-check matrix holds.
[0144] Embodiment 3: 4 sub-packets, [384, 382] regenerated code type, that is, k = 382, n = 384, the redundancy is n - k = 2, and the coupling coefficients are all 128. The following parity-check matrix is constructed in the finite field GF(256):
[0145]
[0146]
[0147] H1 is as follows:
[0148]
[0149] In the above two rows of matrix blocks, H0 is the second matrix block, and the initial matrix H it includes 0,i are all identity matrices, where 1 ≤ i ≤ 384. H1 is the first matrix block, and H 1,1 to H 1,128 all adopt pattern 1, and the coefficients on the diagonal of H 1,1 to H 1,128 are 0 to 127 respectively. H 1,129 to H 1,256 all adopt pattern 2, and the coefficients on the diagonal of H 1,129 to H 1,256 are 0 to 127 respectively. H 1,257 to H 1,384 all adopt pattern 3, and the coefficients on the diagonal of H 1,257 to H 1,384 are 128 to 255 respectively.
[0150] As can be seen from H1, the coefficients on the diagonal of the initial matrices with different patterns can be the same, that is, the values of the coefficients on the diagonal can be reused. Therefore, the existence of the coupling coefficient enables the construction of longer codewords using a finite field with a smaller scale. For example, in Embodiment 3, a codeword of length 384 can be constructed from GF(256). Moreover, the 8×8 matrix formed by any two columns of the initial matrices is invertible, so the MDS property of the codewords obtained based on this parity-check matrix holds.
[0151] In summary, for the code construction method provided in the embodiments of the present application, first, a multi-line matrix block is constructed, and then the multi-line matrix blocks are combined to obtain a parity-check matrix. Among them, the matrix block includes n initial matrices of a×a arranged in columns. The initial matrix in the first matrix block includes coefficients on the diagonal and a / 2 coupling coefficients. The a / 2 coupling coefficients are located in a continuous a / 2 rows and the columns where the a / 2 coupling coefficients are located are different. The first matrix block belongs to the multi-line matrix block, and the initial matrix in the non-first matrix block is a diagonal matrix. Among them, the first matrix block includes initial matrices of three modes. In any one of the three modes, the a / 2 coupling coefficients in the initial matrix are located in the r-th row to the (r + a / 2 - 1)-th row, and in the remaining two modes of the initial matrix, only non-zero elements exist in the r-th column to the (r + a / 2 - 1)-th column in the r-th row to the (r + a / 2 - 1)-th row, where r ≤ a / 2 + 1. This method specifies the positions of the coupling coefficients in the initial matrices of the three modes, that is, except for the coefficients on the diagonal in the initial matrix of the first matrix block, the positions of the non-zero elements are all determined. It can be seen from the explicit construction of the positions of the coupling coefficients that the matrix formed by any two columns of the initial matrices is invertible. Therefore, the MDS property of the codeword constructed based on the parity-check matrix holds. It can be seen from this that in the embodiments of the present application, there is no need to calculate the position of the coupling coefficient by the node, which reduces the complexity of encoding and decoding. During the node repair process, all the data (i.e., the aforementioned symbols) need to be downloaded from the surviving nodes with the same mode as the initial matrix corresponding to the failed node, and only continuous half of the data needs to be downloaded from other surviving nodes. Therefore, the more types of the initial matrices, the more surviving nodes from which only continuous half of the data needs to be downloaded. The embodiments of the present application provide three types of initial matrices. Therefore, the number of nodes corresponding to the same type of initial matrix is small, which can effectively reduce the repair bandwidth of the nodes, and the downloaded data is continuous, without the problem of skipping reading.
[0152] In addition, due to the existence of the coupling coefficients in the first matrix block, the coefficients on the diagonals of the n initial matrices can be reused, that is, the coefficients on the diagonals of multiple initial matrices are the same. This can not only perform encoding and decoding (including encoding and decoding) using the parity-check matrix subsequently, but also first XOR the components corresponding to the positions using the same finite field elements and then perform multiplication, thereby reducing the number of multiplication calculations in encoding and decoding, that is, reducing the encoding and decoding complexity. Moreover, it can construct longer codewords using a smaller finite field, effectively improving the application adaptability of this method and meeting the large-scale coding storage requirements. For example, in a codeword with 4 packets and 2 redundancies, the embodiments of the present application can construct a codeword with a length of 1.5q + 1 in the finite field GF(q). In summary, the code construction method provided in the embodiments of the present application has high application value.
[0153] The embodiments of the present application provide an encoding method, which encodes using the parity-check matrix constructed as described above. Exemplarily, please refer to Figure 5 ,Figure 5 The flowchart of a coding method provided by an embodiment of the present application. This method can be applied to a distributed storage system, specifically to a computing node in the distributed storage system. The computing node can be a coding computing node, such as a server device based on software computing or a chip based on hardware computing, etc. The embodiment of the present application does not limit this. As Figure 5 shown, this method can include the following processes:
[0154] 201. Obtain a parity-check matrix.
[0155] The parity-check matrix is constructed through the foregoing process and can be stored in the computing node. The relevant description of the parity-check matrix can refer to the foregoing explanation, and the embodiment of the present application will not elaborate here.
[0156] 202. Encode the original data blocks based on the parity-check matrix to obtain n encoded data blocks.
[0157] The computing node can calculate the parity data based on the following formula:
[0158] 0 = H × c T = H × [m, p] T = H × [m1, m2, …, m k , p1, …, p n-k T
[0159] where H represents the parity-check matrix, H = [H1, H2, …, H n , c represents the codeword (i.e., n encoded data blocks), m i (1 ≤ i ≤ k) represents the original data blocks, m i includes a symbols: [m i,1 , m i,2 , … m i,a . p j (1 ≤ j ≤ n - k) represents the n - k parity data blocks obtained by encoding, p j includes a symbols: [p j,1 , p j,2 , … p j,a .
[0160] From the foregoing formula, the following formula can be obtained:
[0161] H × c T = [H1, H2, …, H k × [m1, m2, …, m k T + [H k+1 , …, H n × [p1, …, pn-k T = 0
[0162] Since [H k+1 , …, H n is an invertible matrix, the check data blocks can be obtained as follows:
[0163] [p1, …, p n-k T = [H k+1 , …, H n -1 × [H1, H2, …, H k × [m1, m2, …, m k T
[0164] The n encoded data blocks include the aforementioned k original data blocks [m1, m2, …, m k and the n - k check data blocks [p1, …, p n-k .
[0165] Taking Example 1 in the aforementioned code construction method as an example, c = [m1, m2, m3, m4, p1, p2], m i (1 ≤ i ≤ 4) represents the original data block, which includes 4 symbols. p j (1 ≤ j ≤ 2) represents the check data block, which includes 4 symbols. The check matrix H = [H1, H2, H3, H4, H5, H6], H i (1 ≤ i ≤ 6) is an 8×4 matrix, c satisfies the following formula:
[0166] H × c T = [H1, H2, H3, H4] × [m1, m2, m3, m4] T + [H5, H6] × [p1, p2] T = 0
[0167] Since [H5, H6] is an invertible matrix, the check data blocks can be obtained as follows:
[0168] [p1, p2] T = [H5, H6] -1 × [H1, H2, H3, H4] × [m1, m2, m3, m4] T
[0169] Taking Example 2 in the aforementioned code construction method as an example, c = [m1, m2, m3, …, m 22 , p1, p2], m i (1 ≤ i ≤ 22) represents the original data block, which includes 4 symbols. pj (1 ≤ j ≤ 2) represents the parity check data blocks, which include 4 symbols. The parity check matrix H = [H1, H2, H3, …, H 22 , H 23 , H 24 , H i (1 ≤ i ≤ 24) is an 8×4 matrix, c satisfies the following formula:
[0170] H × c T = [H1, H2, H3, …, H 22 × [m1, m2, m3, …, m 22 T + [H 23 , H 24 × [p1, p2] T = 0
[0171] Since [H 23 , H 24 is an invertible matrix, the parity check data blocks can be obtained as follows:
[0172] [p1, p2] T = [H 23 , H 24 -1 × [H1, H2, H3, …, H 22 × [m1, m2, m3, …, m 22 T
[0173] Taking Example 3 in the construction method of the foregoing code as an example, c = [m1, m2, m3, …, m 382 , p1, p2], m i (1 ≤ i ≤ 382) represents the original data blocks, which include 4 symbols. p j (1 ≤ j ≤ 2) represents the parity check data blocks, which include 4 symbols. The parity check matrix H = [H1, H2, H3, …, H 382 , H 383 , H 384 , H i (1 ≤ i ≤ 384) is an 8×4 matrix, c satisfies the following formula:
[0174] H × c T = [H1, H2, H3, …, H 382 × [m1, m2, m3, …, m 382 T + [H 383 , H 384 × [p1, p2] T = 0
[0175] Since [H 383 , H 384 is an invertible matrix, the check data block can be obtained as follows:
[0176] [p1, p2] T = [H 383 , H 384 -1 × [H1, H2, H3,..., H 382 × [m1, m2, m3,..., m 382 T
[0177] The embodiment of the present application provides a node repair method, which uses the check matrix constructed above for node repair. Exemplarily, please refer to Figure 6 , Figure 6 which is a schematic flow chart of a node repair method provided by the embodiment of the present application. This method can be applied to a distributed storage system, specifically to a computing node in a distributed storage system. The computing node can be a decoding computing node, for example, it can be a server device based on software computing or a chip based on hardware computing, etc. The embodiment of the present application does not make any limitations in this regard. As Figure 6 shown, this method can include the following processes:
[0178] 301. Determine the check equation corresponding to each surviving node among the n nodes based on the check matrix, where the n nodes correspond one-to-one to the n initial matrices in the matrix block, a / 2 coupling coefficients of the initial matrix corresponding to the failed node in the first matrix block are located in the r-th row to the r + a / 2 - 1-th row, and the check equation corresponding to the j-th node among the n nodes includes: the r-th row to the r + a / 2 - 1-th row of the j-th initial matrix in the first matrix block, and the r-th row to the r + a / 2 - 1-th row of the j-th initial matrix in any non-first matrix block, 1 ≤ j ≤ n, r ≤ a / 2 + 1.
[0179] The n nodes correspond one-to-one to the n initial matrices in the matrix block of the check matrix, and the i-th node corresponds to the n-th initial matrix in the matrix block. When a certain node fails, the data stored in the failed node can be restored by a check equation through the regeneration method (each surviving node transmits part or all of the data).
[0180] Exemplarily, assume that the codeword is a double-redundancy array code with 4 sub-packets, and the data block c j stored in the j-th (1 ≤ j ≤ n) node is j,1 = [c j,2 , c j,3 , c j,4 . The failed node is the $i$-th node, and the initial matrix corresponding to the failed node is the $i$-th initial matrix in the first matrix block. Any initial matrix in a non-first matrix block is an identity matrix.
[0181] Referring to the foregoing process 301, when the failed node is the $i$-th node and $i\equiv0\ (\text{mod}\ 3)$, the two coupling coefficients of the initial matrix corresponding to the $i$-th node are located in the first row to the second row, i.e., $r = 1$. At this time, the parity-check equations corresponding to the node include: the first row to the second row of the initial matrix corresponding to the node in the first matrix block and the first row to the second row of the initial matrix corresponding to the node in any non-first matrix block. The repair formula is as follows:
[0182]
[0183] where, is the parity-check equation corresponding to the $j$-th node when $j\equiv0\ (\text{mod}\ 3)$, is the parity-check equation corresponding to the $j$-th node when $j\equiv1\ (\text{mod}\ 3)$, is the parity-check equation corresponding to the $j$-th node when $j\equiv2\ (\text{mod}\ 3)$.
[0184] The repair process of the node is to solve the four symbols $[c$ j,1 , $c$ j,2 , $c$ j,3 , $c$ j,4 in the failed node by using the above equations, and it is necessary to download the data involved in the above equations from the surviving nodes. The number of rows of the parity-check equation is equal to the number of unknowns, that is, equal to the number of sub-packets, so the number of rows of the parity-check equation is 4.
[0185] When downloading data from the surviving nodes, for the surviving nodes with the same pattern as the initial matrix corresponding to the failed node, all symbols need to be downloaded. For the surviving nodes with a different pattern from the initial matrix corresponding to the failed node, only half of the consecutive symbols need to be downloaded. For the $j$-th surviving node and $j\equiv0\ (\text{mod}\ 3)$, the pattern of the initial matrix corresponding to this surviving node is the same as the pattern of the initial matrix corresponding to the failed node, so all symbols need to be downloaded from this node, that is, download $[c$ j,1 , $c$ j,2 , $c$ j,3 , $c$ j,4 . For the $j$-th node and $j\equiv1\ (\text{mod}\ 3)$ or $j\equiv2\ (\text{mod}\ 3)$, only the first column and the second column of the parity-check equation corresponding to this surviving node have non-zero elements, so only the first symbol and the second symbol need to be downloaded from this node, that is, download $[c$ j,1 , $c$ j,2 .
[0186] When the failed node is the \(i\)-th node and \(i\equiv1\pmod{3}\), the two coupling coefficients of the initial matrix corresponding to the \(i\)-th node are located in the 3rd to 4th rows, i.e., \(r = 3\). At this time, the parity-check equations corresponding to the node include: the 3rd to 4th rows of the initial matrix corresponding to the node in the first matrix block and the 3rd to 4th rows of the initial matrix corresponding to the node in any non-first matrix block. The repair formula is as follows:
[0187]
[0188] where, is the parity-check equation corresponding to the \(j\)-th node when \(j\equiv0\pmod{3}\), is the parity-check equation corresponding to the \(j\)-th node when \(j\equiv1\pmod{3}\), is the parity-check equation corresponding to the \(j\)-th node when \(j\equiv2\pmod{3}\).
[0189] When downloading data from the surviving nodes, for the \(j\)-th surviving node and \(j\equiv0\pmod{3}\) or \(j\equiv2\pmod{3}\), only the 3rd and 4th columns of the parity-check equation corresponding to this surviving node have non-zero elements. Therefore, only the 3rd symbol and the 4th symbol need to be downloaded from this node, i.e., download \([c j,3 , c j,4 \). For the \(j\)-th node and \(j\equiv1\pmod{3}\), the pattern of the initial matrix corresponding to this surviving node is the same as the pattern of the initial matrix corresponding to the failed node. Therefore, all symbols need to be downloaded from this node, i.e., download \([c j,1 , c j,2 , c j,3 , c j,4 \).
[0190] When the failed node is the \(i\)-th node and \(i\equiv2\pmod{3}\), the two coupling coefficients of the initial matrix corresponding to the \(i\)-th node are located in the 2nd to 3rd rows, i.e., \(r = 2\). At this time, the parity-check equations corresponding to the node include: the 2nd to 3rd rows of the initial matrix corresponding to the node in the first matrix block and the 2nd to 3rd rows of the initial matrix corresponding to the node in any non-first matrix block. The repair formula is as follows:
[0191]
[0192] where, is the parity-check equation corresponding to the \(j\)-th node when \(j\equiv0\pmod{3}\), is the parity-check equation corresponding to the \(j\)-th node when \(j\equiv1\pmod{3}\), is the parity-check equation corresponding to the \(j\)-th node when \(j\equiv2\pmod{3}\).
[0193] 302. Recover the data in the failed nodes among the n nodes according to the check equations corresponding to each surviving node and the data stored in the surviving nodes.
[0194] The computing node downloads the data involved in the equations from each surviving node based on the check equation set described in the foregoing process 301, so as to calculate the data in the failed nodes.
[0195] When downloading data from the surviving nodes, for the j-th surviving node where j ≡ 0 (mod 3) or j ≡ 1 (mod 3), only the second column and the third column of the check equation corresponding to this surviving node have non-zero elements. Therefore, only the second symbol and the third symbol need to be downloaded from this node, that is, download [c j,2 , c j,3 . For the j-th node where j ≡ 2 (mod 3), the pattern of the initial matrix corresponding to this surviving node is the same as the pattern of the initial matrix corresponding to the failed node. Therefore, all symbols need to be downloaded from this node, that is, download [c j,1 , c j,2 , c j,3 , c j,4 .
[0196] The node repair process of the triple-redundancy array code with 8 sub-packets can refer to the foregoing process, and the embodiments of the present application will not elaborate herein.
[0197] Taking the check matrices provided in the foregoing Embodiment 1 to Embodiment 3 as examples, the node repair process will be further described below.
[0198] Corresponding to Embodiment 1, if the first node (or the second node) fails, the check equation corresponding to the i-th node includes: the first to second rows in H 0,i and the first to second rows in H 1,i . Based on the foregoing linear equation set, when the first node fails, only all the data needs to be downloaded from the second node, and the first symbol and the second symbol need to be downloaded from the other surviving nodes except the second node. All the data in the first node can be recovered through the foregoing linear equation set. When the second node fails, only all the data needs to be downloaded from the first node, and the first symbol and the second symbol need to be downloaded from the other surviving nodes except the first node. All the data in the second node can be recovered through the foregoing linear equation set.
[0199] If the third node (or the fourth node) fails, the check equation corresponding to the i-th node includes: the third to fourth rows in H 0,i and H 1,iLines 3 to 4. Based on the foregoing linear equations, when the third node fails, it is only necessary to download all the data from the fourth node, download the third symbol and the fourth symbol from the other surviving nodes except the fourth node, and all the data in the third node can be recovered through the foregoing linear equations. When the fourth node fails, it is only necessary to download all the data from the third node, download the third symbol and the fourth symbol from the other surviving nodes except the third node, and all the data in the fourth node can be recovered through the foregoing linear equations.
[0200] If the fifth node (or the sixth node) fails, the check equation corresponding to the i-th node includes: H 0,i Lines 2 to 3 in 1,i Lines 2 to 3. Based on the foregoing linear equations, when the fifth node fails, it is only necessary to download all the data from the sixth node, download the second symbol and three symbols from the other surviving nodes except the sixth node, and all the data in the fifth node can be recovered through the foregoing linear equations. When the sixth node fails, it is only necessary to download all the data from the fifth node, download the second symbol and the third symbol from the other surviving nodes except the fifth node, and all the data in the sixth node can be recovered through the foregoing linear equations.
[0201] Corresponding to Example 2, if the first node fails, the check equation corresponding to the i-th node includes: H 0,i Lines 1 to 2 in 1,i Lines 1 to 2. Based on the foregoing linear equations, it is only necessary to download all the data from the second to the eighth nodes, download the first symbol and the second symbol from the other surviving nodes except the second to the eighth nodes, and all the data in the first node can be recovered through the foregoing linear equations. The download situation of the data when the second to the eighth nodes fail can refer to the foregoing analysis, and the embodiments of the present application will not be elaborated here.
[0202] If the ninth node fails, the check equation corresponding to the i-th node includes: H 0,i Lines 3 to 4 in 1,i Lines 3 to 4. Based on the foregoing linear equations, it is only necessary to download all the data from the tenth to the sixteenth nodes, download the third symbol and the fourth symbol from the other surviving nodes except the tenth to the sixteenth nodes, and all the data in the ninth node can be recovered through the foregoing linear equations. The download situation of the data when the tenth to the sixteenth nodes fail can refer to the foregoing analysis, and the embodiments of the present application will not be elaborated here.
[0203] If the seventeenth node fails, the check equation corresponding to the i-th node includes: H0,i lines 2 to 3 in and H 1,i lines 2 to 3 in. Based on the foregoing linear equations, it can be seen that only all data needs to be downloaded from the 18th to the 24th nodes, and the second symbol and three symbols need to be downloaded from the other surviving nodes except the 18th to the 24th nodes. All data in the 17th node can be restored through the foregoing linear equations. The download situation of the data when the 18th to the 24th nodes fail can refer to the foregoing analysis, and the embodiments of the present application will not elaborate here.
[0204] Corresponding to Embodiment 3, if the first node fails, the check equation corresponding to the i-th node includes: H 0,i lines 1 to 2 in and H 1,i lines 1 to 2 in. Based on the foregoing linear equations, it can be seen that only all data needs to be downloaded from the 2nd to the 128th nodes, and the first symbol and the second symbol need to be downloaded from the other surviving nodes except the 2nd to the 128th nodes. All data in the first node can be restored through the foregoing linear equations. The download situation of the data when the 2nd to the 128th nodes fail can refer to the foregoing analysis, and the embodiments of the present application will not elaborate here.
[0205] If the 129th node fails, the check equation corresponding to the i-th node includes: H 0,i lines 3 to 4 in and H 1,i lines 3 to 4 in. Based on the foregoing linear equations, it can be seen that only all data needs to be downloaded from the 130th to the 256th nodes, and the third symbol and the fourth symbol need to be downloaded from the other surviving nodes except the 130th to the 256th nodes. All data in the 129th node can be restored through the foregoing linear equations. The download situation of the data when the 130th to the 256th nodes fail can refer to the foregoing analysis, and the embodiments of the present application will not elaborate here.
[0206] If the 257th node fails, the check equation corresponding to the i-th node includes: H 0,i lines 2 to 3 in and H 1,i lines 2 to 3 in. Based on the foregoing linear equations, it can be seen that only all data needs to be downloaded from the 258th to the 384th nodes, and the second symbol and three symbols need to be downloaded from the other surviving nodes except the 258th to the 384th nodes. All data in the 257th node can be restored through the foregoing linear equations. The download situation of the data when the 258th to the 384th nodes fail can refer to the foregoing analysis, and the embodiments of the present application will not elaborate here.
[0207] In summary, for the node repair method provided by the embodiments of the present application, first, check equations corresponding to each surviving node among the n nodes are determined based on the check matrix, and then, according to the check equations corresponding to each surviving node and the data stored in the surviving nodes, the data in the failed nodes among the n nodes is recovered. Among them, the n nodes correspond one by one to the n initial matrices in the matrix block. a / 2 coupling coefficients of the initial matrix corresponding to the failed node in the first matrix block are located in the r-th row to the (r + a / 2 - 1)-th row. The check equation corresponding to the j-th node among the n nodes includes: the r-th row to the (r + a / 2 - 1)-th row of the j-th initial matrix in the first matrix block, and the r-th row to the (r + a / 2 - 1)-th row of the j-th initial matrix in any non-first matrix block, where 1 ≤ j ≤ n and r ≤ a / 2 + 1. According to the properties of the foregoing check matrix, in the embodiments of the present application, there is no need to calculate the positions of the node search coupling coefficients, reducing the complexity of node repair. During the node repair process, all the data (i.e., the foregoing symbols) needs to be downloaded from the surviving nodes with the same pattern as the initial matrix corresponding to the failed node, and only consecutive half of the data needs to be downloaded from other surviving nodes. Therefore, the more types of initial matrices there are, the fewer surviving nodes from which only consecutive half of the data needs to be downloaded. The embodiments of the present application provide three types of initial matrices. Therefore, the number of nodes corresponding to the same type of initial matrix is small, which can effectively reduce the repair bandwidth of the nodes, and the downloaded data is all consecutive, without the problem of skipping reading. The embodiments of the present application can reach the theoretical lower bound of the repair bandwidth under the condition of no skipping reading.
[0208] In addition, due to the existence of coupling coefficients in the first matrix block, the coefficients on the diagonal of the n initial matrices can be reused, that is, there are coefficients on the diagonal of multiple initial matrices that are the same. This can not only, during node repair, first perform exclusive OR on the components corresponding to the positions using the same finite field element and then perform multiplication, thereby reducing the number of multiplication calculations in node repair, that is, reducing the complexity of node repair. Moreover, it can construct longer codewords using a smaller finite field, effectively improving the application adaptability of this method and meeting the large-scale coding storage requirements.
[0209] The sequence of the method provided by the embodiments of the present application can be adjusted appropriately, and the process can also be increased or decreased accordingly according to the situation. Any person skilled in the art can easily think of a changed method within the technical scope disclosed by the present application, and it should be covered by the protection scope of the present application. The embodiments of the present application do not make any limitations in this regard.
[0210] The above mainly introduced the method for constructing codes, the encoding method, and the node repair method provided by the embodiments of the present application from the perspective of devices. It can be understood that, in order to implement the above functions, the computing nodes for executing the above methods include the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0211] Embodiments of the present application can divide the functional modules of the computing node according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into a processing subsystem. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical functional division, and there can be other division methods in actual implementation.
[0212] Figure 7 It is a block diagram of a data processing device provided by an embodiment of the present application. When each functional module is divided corresponding to each function, the data processing device 400 may include: a construction module 401 and a combination module 402. Exemplarily, the data processing device may be a computing node, or a chip or other combined devices, components, etc. with the functions of the above data processing device. The functions of each module of the device are as follows:
[0213] The construction module is used to construct multiple rows of matrix blocks. The matrix blocks include n a×a initial matrices sorted by columns. The initial matrices in the first matrix block include coefficients on the diagonal and a / 2 coupling coefficients. The a / 2 coupling coefficients are located in consecutive a / 2 rows and the columns where the a / 2 coupling coefficients are located are different. The first matrix block belongs to the multiple rows of matrix blocks, and the initial matrices in non-first matrix blocks are diagonal matrices;
[0214] The combination module is used to combine multiple rows of matrix blocks to obtain a parity-check matrix, and the parity-check matrix is used to encode and repair the original data block;
[0215] Among them, the first matrix block includes initial matrices of three modes. For the initial matrix of any one of the three modes, a / 2 coupling coefficients are located in the r-th row to the (r + a / 2 - 1)-th row. In the initial matrices of the remaining two modes, only non-zero elements exist in the r-th column to the (r + a / 2 - 1)-th column in the r-th row to the (r + a / 2 - 1)-th row, where r ≤ a / 2 + 1.
[0216] Combined with the above solution, the initial matrices of the three modes include: a / 2 coupling coefficients are distributed in the upper-right a / 2×a / 2 square matrix of the initial matrix, a / 2 coupling coefficients are distributed in the lower-left a / 2×a / 2 square matrix of the initial matrix, and among the a / 2 coupling coefficients, a / 4 coupling coefficients are distributed in the upper-left a / 2×a / 2 square matrix of the initial matrix and the remaining a / 4 coupling coefficients are distributed in the lower-right a / 2×a / 2 square matrix of the initial matrix.
[0217] Combined with the above solution, the values of a / 2 coupling coefficients in the n initial matrices in the first matrix block are the same.
[0218] Combined with the above solution, the initial matrix in the second matrix block is an identity matrix, and the second matrix block belongs to a multi-row matrix block.
[0219] Combined with the above solution, the initial matrices with the same result of taking the remainder of the permutation serial number modulo 3 in the first matrix block have the same mode, and the initial matrices with different results of taking the remainder of the permutation serial number modulo 3 have different modes.
[0220] Figure 8 As shown in the block diagram of another data processing device provided by the embodiments of the present application, when each functional module is divided according to the corresponding functions, the data processing device 500 may include: a processing module 501 and a repair module 502. Exemplarily, the data processing device may be a computing node, or a chip therein, or other combined devices, components, etc. with the functions of the above data processing device. The functions of each module of the device are as follows:
[0221] The processing module is configured to determine a check equation corresponding to each surviving node among the n nodes based on a check matrix, and the check matrix is constructed by the method described in any of the foregoing embodiments;
[0222] The repair module is configured to restore the data in the failed nodes among the n nodes according to the check equation corresponding to each surviving node and the data stored in the surviving nodes;
[0223] Among them, n nodes correspond to n initial matrices in the matrix block one by one. a / 2 coupling coefficients of the initial matrix corresponding to the failed node in the first matrix block are located in the r-th row to the (r + a / 2 - 1)-th row. The check equation corresponding to the j-th node among the n nodes includes: the r-th row to the (r + a / 2 - 1)-th row of the j-th initial matrix in the first matrix block, and the r-th row to the (r + a / 2 - 1)-th row of the j-th initial matrix in any non-first matrix block, where 1 ≤ j ≤ n and r ≤ a / 2 + 1.
[0224] Figure 9 As shown in the block diagram of another data processing device provided in the embodiments of the present application. When each functional module is divided according to the corresponding functions, the data processing device 600 may include: an acquisition module 601 and an encoding module 602. Exemplarily, the data processing device may be a computing node, or a chip therein, or other combined devices, components, etc. having the functions of the above data processing device. The functions of each module of the device are as follows:
[0225] The acquisition module is configured to acquire a check matrix, and the check matrix is constructed by the method described in any of the foregoing embodiments;
[0226] The encoding module is configured to encode the original data block based on the check matrix to obtain n encoded data blocks.
[0227] Figure 10 As shown in the structural schematic diagram of an electronic device provided in the embodiments of the present application. The electronic device 700 may be a chip or a functional module in a computing node. As Figure 10 shown, the electronic device 700 includes a processor 701, a transceiver 702, and a communication line 703.
[0228] Among them, the processor 701 is configured to execute any step in the method embodiments as Figures 4 to 6 shown, and when executing processes such as acquiring the original data block, etc., it can optionally call the transceiver 702 and the communication line 703 to complete the corresponding operations.
[0229] Further, the electronic device 700 may further include a memory 704. Among them, the processor 701, the memory 704, and the transceiver 702 may be connected through the communication line 703.
[0230] The transceiver 702 is configured to communicate with other devices or other communication networks. The other communication networks may be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The transceiver 702 may be a module, a circuit, a transceiver, or any device capable of realizing data processing.
[0231] The transceiver 702 is mainly used for sending and receiving data, etc., and may include a transmitter and a receiver for sending and receiving data, etc. respectively; operations other than sending and receiving data, etc. are implemented by the processor, such as constructing multi-line matrix blocks, etc.
[0232] The communication line 703 is used to transmit information between the components included in the electronic device 700.
[0233] In one design, the processor can be regarded as a logic circuit, and the transceiver can be regarded as an interface circuit.
[0234] The memory 704 is used to store instructions. Among them, the instructions can be computer programs.
[0235] It should be noted that the memory 704 can exist independently of the processor 701 or can be integrated with the processor 701. The memory 704 can be used to store instructions or program codes or some data, etc. The memory 704 can be located inside the electronic device 700 or outside the electronic device 700, without limitation. The processor 701 is used to execute the instructions stored in the memory 704 to implement the method provided in the above embodiments of the present application.
[0236] In one example, the processor 701 can include one or more processors, such as Figure 10 the processor 0 and the processor 1 in
[0237] As an optional implementation manner, the electronic device 700 includes multiple processors. For example, in addition to Figure 10 the processor 701 in
[0238] As an optional implementation manner, the electronic device 700 further includes an output device 705 and an input device 706. Exemplarily, the input device 706 is a device such as a keyboard, a mouse, a microphone, or a joystick, and the output device 705 is a device such as a display screen or a speaker.
[0239] It should be noted that the electronic device 700 can be a chip system or a device with a Figure 10 similar structure in Figure 10 The actions, terms, etc. involved between the embodiments of the present application can be referred to each other without limitation. The message names or parameter names in the messages exchanged between the devices in the embodiments of the present application are only examples, and other names can also be adopted in specific implementations without limitation. In addition, Figure 10 the shown composition structure does not constitute a limitation on the electronic device 700. In addition to Figure 10More or fewer components as shown, or combining certain components, or different component arrangements.
[0240] The processor and transceiver described in this application can be implemented on an integrated circuit (IC), analog IC, radio frequency integrated circuit, mixed-signal IC, application specific integrated circuit (ASIC), printed circuit board (PCB), electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (NMOS), P-type metal oxide semiconductor (PMOS), bipolar junction transistor (BJT), BiCMOS, silicon germanium (SiGe), gallium arsenide (GaAs), etc.
[0241] Figure 11 It is a schematic structural diagram of a data processing device provided for an embodiment of this application. The data processing device can be applied to the scenarios shown in the above method embodiments. For the convenience of description, Figure 11 only the main components of the data processing device are shown, including a processor, a memory, a control circuit, and an input / output device. The processor is mainly used to process communication protocols and communication data, execute software programs, and process the data of software programs. The memory is mainly used to store software programs and data. The control circuit is mainly used for power supply and the transmission of various electrical signals. The input / output device is mainly used to receive data input by users and output data to users.
[0242] When the data processing device is a computing node, the control circuit may be a motherboard, the memory includes storage media with storage functions such as hard disks, RAM, and ROM, the processor may include a baseband processor and a central processing unit. The baseband processor is mainly used to process communication protocols and communication data, and the central processing unit is mainly used to control the entire data processing device, execute software programs, and process the data of software programs. The input / output device includes a display screen, a keyboard, a mouse, etc.; the control circuit may further include or be connected to a transceiver circuit or a transceiver, for example: a network cable interface, etc., for sending or receiving data or signals, for example, for data transmission and communication with other devices. Further, an antenna may also be included for data transceiver, for data / request transmission with other devices.
[0243] According to the method provided by the embodiments of the present application, the present application also provides a computer program product, which includes computer program code. When the computer program code runs on a computer, it causes the computer to execute any of the methods described in the embodiments of the present application.
[0244] The embodiments of the present application also provide a computer-readable storage medium. All or part of the processes in the above method embodiments can be executed by a computer or a device with data processing capabilities by executing computer programs or instructions to control relevant hardware to complete. The computer program or the set of instructions can be stored in the above computer-readable storage medium. When the computer program or the set of instructions is executed, it may include the processes of the above method embodiments. The computer-readable storage medium may be an internal storage unit of the computing node in any of the foregoing embodiments, such as the hard disk or memory of the computing node. The above computer-readable storage medium may also be an external storage device of the above computing node, such as a plug-in hard disk equipped on the above computing node, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the above computer-readable storage medium may also include both the internal storage unit of the above computing node and the external storage device. The above computer-readable storage medium is used to store the above computer programs or instructions and other programs and data required by the above computing node. The above computer-readable storage medium may also be used to temporarily store data that has been output or is about to be output.
[0245] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0246] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the device described above can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein.
[0247] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0248] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0249] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0250] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0251] As described above, this is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all such changes or substitutions should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims. The above mainly introduces the code construction method, encoding method, and node repair method provided by the embodiments of the present application from the perspective of the device. It can be understood that in order for the device to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
Claims
1. A method for constructing a code, characterized in that, The method includes: Constructing a multi-line matrix block, where the matrix block includes n initial matrices of a×a sorted by column. The initial matrices in the first matrix block include coefficients on the diagonal and a / 2 coupling coefficients. The a / 2 coupling coefficients are located in consecutive a / 2 rows and the columns where the a / 2 coupling coefficients are located are different. The first matrix block belongs to the multi-line matrix block, and the initial matrices in non-first matrix blocks are diagonal matrices; Combining the multi-line matrix block to obtain a parity-check matrix, which is used to encode and repair the original data block; Among them, the first matrix block includes initial matrices of three modes. For any one of the three modes, the a / 2 coupling coefficients in the initial matrix are located in rows r to r + a / 2 - 1, and in the initial matrices of the remaining two modes, only non-zero elements exist in columns r to r + a / 2 - 1 in rows r to r + a / 2 - 1, where r ≤ a / 2 + 1.
2. The method according to claim 1, wherein The initial matrices of the three modes include: the a / 2 coupling coefficients are distributed in the upper-right a / 2×a / 2 square matrix of the initial matrix, the a / 2 coupling coefficients are distributed in the lower-left a / 2×a / 2 square matrix of the initial matrix, and a / 4 of the a / 2 coupling coefficients are distributed in the upper-left a / 2×a / 2 square matrix of the initial matrix and the remaining a / 4 coupling coefficients are distributed in the lower-right a / 2×a / 2 square matrix of the initial matrix.
3. The method according to claim 1 or 2, wherein The values of the a / 2 coupling coefficients of the n initial matrices in the first matrix block are the same.
4. The method according to any one of claims 1 to 3, wherein The initial matrix in the second matrix block is an identity matrix, and the second matrix block belongs to the multi-line matrix block.
5. The method according to any one of claims 1 to 4, wherein The initial matrices with the same result of taking the modulus of the permutation serial number by 3 in the first matrix block have the same mode, and the initial matrices with different results of taking the modulus of the permutation serial number by 3 have different modes.
6. A node repair method, characterized in that, The method includes: Determining a parity-check equation corresponding to each surviving node among n nodes based on the parity-check matrix, where the parity-check matrix is constructed according to the method described in any one of claims 1 to 5; Restoring the data in the failed nodes among the n nodes according to the parity-check equation corresponding to each surviving node and the data stored in the surviving nodes; Among them, the n nodes correspond to the n initial matrices in the matrix block one by one. The a / 2 coupling coefficients of the initial matrix corresponding to the failed node in the first matrix block are located in the r-th row to the (r + a / 2 - 1)-th row. The check equation corresponding to the j-th node among the n nodes includes: the r-th row to the (r + a / 2 - 1)-th row of the j-th initial matrix in the first matrix block, and the r-th row to the (r + a / 2 - 1)-th row of any initial matrix other than the j-th initial matrix in the first matrix block, where 1 ≤ j ≤ n and r ≤ a / 2 + 1.
7. A coding method, characterized in that, The method includes: Obtaining a check matrix, where the check matrix is constructed by the method according to any one of claims 1 to 5; Encoding the original data block based on the check matrix to obtain n encoded data blocks.
8. A data processing device, characterized in that, The apparatus includes: One or more processors; A memory for storing one or more computer programs or instructions; When the one or more computer programs or instructions are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5, or implement the method according to claim 6, or implement the method according to claim 7.
9. A data processing device, characterized in that, The apparatus includes: A processing circuit and an interface circuit; Among them, the interface circuit is used to couple with a memory external to the data processing device and provide a communication interface for the processing circuit to access the memory; The processing circuit is used to execute the program instructions in the memory to implement the method according to any one of claims 1 to 5, or implement the method according to claim 6, or implement the method according to claim 7.
10. A computer-readable storage medium, characterized in that Program code is stored in the computer-readable storage medium, and when the program code is executed by a processor, the method according to any one of claims 1 to 7 is implemented.