Data verification method and apparatus, data recovery method and apparatus, device, and medium

By using sparse binary array codes based on finite domains in distributed storage systems, a finite domain matrix with maximum distance separability is constructed, which solves the problem of high data checksum recovery complexity in the prior art and achieves efficient encoding and codec performance.

WO2025118811A1PCT designated stage expired Publication Date: 2025-06-12HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/123680
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-10-09
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently perform data checksum recovery in distributed storage systems, especially when node failure or data errors, traditional methods require complex matrix inversion operations, resulting in high computational complexity.

Method used

The sparse binary array code based on finite domains is adopted, and the finite domain matrix with the maximum distance separability is constructed and converted into a sparse binary check matrix is ​​supported, which supports pure XOR method of high-performance encoding and decoding to avoid matrix inversion operations.

Benefits of technology

It realizes efficient data checksum recovery, reduces computational complexity, improves encoding and codec performance, and can reach or approach the theoretical limit in some cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024123680_12062025_PF_FP_ABST
    Figure CN2024123680_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to a data verification method and apparatus, a data recovery method and apparatus, a device, and a computer readable storage medium. The data verification method comprises: acquiring a set of data units each having a plurality of subpackets and a corresponding verification matrix, and using the verification matrix to determine verification units for the set of data units. The data recovery method comprises: acquiring a set of data units including a data unit to be recovered, a set of corresponding verification units, and a verification matrix; and using the verification matrix to determine the data unit to be recovered. In particular, the verification matrix used in the embodiments of the present application is a binary matrix representation of a finite field matrix. The verification matrix has the characteristic of sparse rules derived from a corresponding finite field element matrix thereof. Using such a matrix to perform verification encoding and decoding achieves high-performance small-subpacket pure XOR encoding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and medium for data verification and data recovery

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 7, 2023, with application number 202311691201.6, and invention name “Methods, devices, equipment and media for data verification and data recovery”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] Embodiments of the present application generally relate to the field of data security, and more specifically, to methods, devices, equipment, and computer-readable storage media for data verification and data recovery. Background Art

[0003] With the development of computer technology, various industries are becoming increasingly dependent on digital information, and the amount of electronic data generated and stored is rapidly increasing. Due to its high storage costs and low scalability, traditional centralized storage can no longer meet this growing data storage demand. In recent years, distributed storage systems have gained widespread attention and application due to their cost-effectiveness, scalability, and high reliability.

[0004] In a distributed storage system, data is stored across multiple servers (called nodes) interconnected by a network. Nodes may be based on inexpensive servers that are susceptible to temporary or permanent failures due to disk failures, power outages, network outages, and server downtime, posing challenges to data storage reliability. To ensure data reliability, redundant data is required in the storage system. Furthermore, if a node fails or its data becomes unavailable for other reasons (e.g., due to storage errors), the storage system can use the redundant data to restore the data on the corresponding node.

[0005] Summary of the Invention

[0006] This application provides a scheme for data verification. The check matrix used in this scheme is converted from a regular and usually sparse matrix of elements in a finite field. The finite field matrix is ​​constructed based on the structure of the codeword to be encoded and has the Maximum Distance Separable (MDS) property. Utilizing such an array code can improve encoding and decoding efficiency.

[0007] In a first aspect, a method for data verification is provided. The method includes obtaining a set of data units, each data unit in the set of data units having multiple subpackets; and obtaining a check matrix, the check matrix being a binary matrix representation of a finite field matrix, each element in the finite field matrix being from a specific finite field. In other words, the check matrix can be divided into a plurality of submatrices of the same order, wherein each submatrix is ​​a binary matrix representation of elements at corresponding positions in the finite field matrix. The method also includes determining a set of check units based on the check matrix and the set of data units. The check matrix used in the method can support high-performance check encoding in a pure XOR manner. In addition, decoding data encoded in this manner can avoid complex matrix inversion operations.

[0008] In some embodiments of the first aspect, the finite field matrix includes a first part, wherein the first part uses a matrix of the same order formed by elements in the two-expansion field as a construction unit, wherein the product of the order of the construction unit and the degree of the two-expansion field is equal to the number of subpackets of the data unit in a group of data units, and the number of construction units included in the first part in the row direction and the column direction is respectively equal to one of the following: the number of data units in the group and the number of check units in the group. Thus, a code pattern construction method for the finite field matrix is ​​provided. The binary check matrix converted from the finite field matrix constructed in this manner can support high-performance check coding using a pure XOR method. In addition, decoding data encoded using the resulting binary check matrix can avoid complex matrix inversion operations.

[0009] In some embodiments of the first aspect, the construction unit is a first-order matrix or a sparse matrix of elements in the two extended fields, and the construction unit in the first row of construction units in the first part is an identity matrix. The finite field matrix constructed in this manner can be converted into a relatively sparse binary check matrix, so that the complexity of encoding and decoding using the binary check matrix is ​​reduced.

[0010] In some embodiments of the first aspect, the same row of construction units in the first part includes multiple pairs of construction units, wherein each pair of construction units is equal, or each pair of construction units differs by a permutation matrix after being converted into a binary matrix representation. Thus, conditions for constructing a finite field matrix used in the embodiments of the present application are provided, and the finite field constructed in this way can be converted into a regular sparse high-performance binary check matrix.

[0011] In some embodiments of the first aspect, the number of subpackets of each data unit is less than or equal to 4, and the same row of construction units in the first portion includes multiple pairs of construction units, wherein the values ​​of corresponding positions of each pair of construction units are the same or sum to 1. Thus, simplified conditions are provided for constructing a finite field matrix in the case of small subpackets, and the finite field matrix constructed in this way can be converted into a regular sparse high-performance binary check matrix.

[0012] In some embodiments of the first aspect, each m-order submatrix of the finite field matrix is ​​full rank, where m is equal to the number of rows of the finite field matrix. Thus, a constraint is provided for constructing a finite field matrix that satisfies the MDS property, so that a binary check matrix derived therefrom has the MDS property.

[0013] In some embodiments of the first aspect, determining a group of check units includes: determining, based on values ​​of a pair of submatrices and a group of data units in a check matrix, intermediate values ​​of the group of check units corresponding to the pair of submatrices, the pair of submatrices corresponding to a pair of adjacent construction units in a row direction of the finite field matrix; and determining values ​​of the group of check units based on the intermediate values. This reduces the number of computational steps in the encoding process for determining the check units, thereby reducing computational overhead.

[0014] In some embodiments of the first aspect, the finite field is a two-dimensional field GF (2 2 ). By using this finite field to construct a check matrix, the encoding and decoding performance of the obtained check matrix can reach the theoretical limit in some cases.

[0015] In some embodiments of the first aspect, the method further includes writing the set of data units to a set of data nodes; and writing the set of check units to a set of check nodes. Thus, the method can be applied to a storage system including multiple data nodes to provide redundant protection with high computing performance for the data therein.

[0016] In a second aspect, a method for data recovery is provided. The method comprises: obtaining a set of data units and a corresponding set of check units, the set of check units being determined based on a set of data units and a check matrix, the set of data units including the data units to be recovered; obtaining the check matrix, the check matrix being a binary matrix representation of a finite field matrix, each element of the finite field matrix being from a specific finite field; and determining the value of the data unit to be recovered based on the set of data units, the set of check units, and the check matrix. The check matrix used in this method can support high-performance check encoding using a pure XOR scheme. Therefore, using the check matrix for corresponding decoding in this method can also achieve high performance and avoid complex matrix inversion operations.

[0017] In some embodiments of the second aspect, the finite field matrix includes a first part, wherein the first part uses a matrix of the same order formed by elements in two expansion fields as a construction unit, wherein the product of the order of the construction unit and the number of times the two expansion fields are equal to the number of subpackets of the data units in the group of data units, and the number of construction units included in the row direction and the column direction of the first part are respectively equal to one of the following: the number of the group of data units and the number of the group of check units. Thus, a code pattern construction method for a finite field matrix is ​​provided. The binary check matrix converted from the finite field matrix constructed in this manner can support high-performance check coding in a pure XOR manner. Therefore, corresponding decoding using the obtained binary check matrix can also achieve high performance and avoid complex matrix inversion operations.

[0018] In some embodiments of the second aspect, determining the value of the data unit to be recovered includes: determining, based on the parity check matrix, a calculation path for calculating the value of the data unit to be recovered and an intermediate value generated in the calculation path; and determining the value of the data unit to be recovered based on the calculation path and the intermediate value. This eliminates matrix inversion operations during the decoding process to recover the original data, reduces the complexity of the steps, and thus reduces computational overhead.

[0019] In a third aspect, a device for data verification is provided. The device includes: a data acquisition module configured to acquire a set of data units, each data unit in the set of data units having multiple subpackets; a matrix acquisition module configured to acquire a check matrix, the check matrix being a binary matrix representation of a finite field matrix, each element in the finite field matrix being from a specific finite field; and a determination module configured to determine a set of check units based on the check matrix and the set of data units.

[0020] In some embodiments of the third aspect, the finite field matrix includes a first part, wherein the first part uses a matrix of the same order formed by elements in two extended fields as a construction unit, wherein the product of the order of the construction unit and the degree of the two extended fields is equal to the number of subpackaging of data units in a group of data units, and the number of construction units included in the row direction and the column direction of the first part is respectively equal to one of the following: the number of data units in the group and the number of check units in the group.

[0021] In some embodiments of the third aspect, the construction unit is a first-order matrix or a sparse matrix of elements in the two extended fields, and the construction unit in the first row of construction units in the first part is an identity matrix.

[0022] In some embodiments of the third aspect, the same row of construction units in the first part includes multiple pairs of construction units, wherein each pair of construction units is equal, or each pair of construction units differs by a permutation matrix after each pair is converted into a binary matrix representation.

[0023] In some embodiments of the third aspect, the number of subpackets of each data unit is less than or equal to 4, and the same row of construction units in the first part includes multiple pairs of construction units, wherein the values ​​of corresponding positions of each pair of construction units are the same or the sum is 1.

[0024] In some embodiments of the third aspect, each m-order submatrix of the finite field matrix is ​​full rank, where m is equal to the number of rows of the finite field matrix.

[0025] In some embodiments of the third aspect, the determination module includes: an intermediate value determination module, configured to determine, based on values ​​of a pair of sub-matrices and a group of data units in the check matrix, intermediate values ​​of a group of check units corresponding to a pair of sub-matrices, the pair of sub-matrices corresponding to a pair of construction units adjacent in a row direction of the finite field matrix; and an intermediate value multiplexing module, configured to determine values ​​of the group of check units based on the intermediate values.

[0026] In some embodiments of the third aspect, the finite field is a two-dimensional field GF (2 2 ).

[0027] In some embodiments of the third aspect, the apparatus further comprises: a data unit writing module configured to write the group of data units to a group of data nodes; and a check unit writing module configured to write the group of check units to a group of check nodes.

[0028] In a fourth aspect, a device for data recovery is provided. The device includes: a unit acquisition module configured to acquire a group of data units and a corresponding group of check units, the group of check units being determined based on the group of data units and a check matrix, the group of data units including the data unit to be recovered; a matrix acquisition module configured to acquire the check matrix, the check matrix being a binary matrix representation of a finite field matrix, each element of the finite field matrix being from a specific finite field; and a determination module configured to determine a value of the data unit to be recovered based on the group of data units, the check units, and the check matrix.

[0029] In some embodiments of the fourth aspect, the finite field matrix includes a first part, the finite field matrix includes the first part, the first part uses a same-order matrix composed of elements in the two extended fields as a construction unit, wherein the product of the order of the construction unit and the order of the two extended fields is equal to the number of subpackaging of the data units in the group of data units, and the number of construction units included in the row direction and the column direction of the first part is respectively equal to one of the following: the number of the group of data units and the number of the group of check units.

[0030] In some embodiments of the fourth aspect, the determination module includes: an intermediate value determination module, configured to determine, based on a check matrix, a calculation path for calculating the value of the data unit to be recovered, and an intermediate value generated in the calculation path; and an intermediate value multiplexing module, configured to determine the value of the data unit to be recovered based on the calculation path and the intermediate value.

[0031] In a fifth aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores instructions that, when executed by the processor, cause the electronic device to perform the actions of the method according to the first aspect or any embodiment thereof, or the actions of the method according to the second aspect or any embodiment thereof.

[0032] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions that, when executed by an electronic device, cause the electronic device to perform the actions of the method according to the first aspect or any embodiment thereof, or the actions of the method according to the second aspect or any embodiment thereof.

[0033] In a seventh aspect, a computer program product is provided. The computer program product is tangibly stored on a computer-readable medium and includes computer-executable instructions that, when executed, implement the actions of the method according to the first aspect or any embodiment thereof, or the actions of the method according to the second aspect or any embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and other features, advantages and aspects of the embodiments of the present application will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0035] FIG1 is a schematic diagram showing an example environment in which various embodiments of the present application can be implemented;

[0036] FIG2 shows a flow chart of an example method for data verification according to some embodiments of the present application;

[0037] FIG3 shows a flowchart of an example method for data recovery according to some embodiments of the present application;

[0038] FIG4 shows an exemplary schematic diagram according to some embodiments of the present application, wherein an exemplary structural form of a coefficient matrix used in a decoding process is illustrated;

[0039] FIG5 is a schematic diagram illustrating encoding and decoding of erasure codes in an example storage system according to some embodiments of the present application;

[0040] FIG6 shows a schematic diagram of the construction of an example binary array code according to some embodiments of the present application;

[0041] FIG7 shows a schematic diagram of the construction of another example binary array code according to some embodiments of the present application;

[0042] FIG8 shows a schematic diagram of the construction of another example binary array code according to some embodiments of the present application;

[0043] FIG9 shows a schematic block diagram of an apparatus for data verification according to some embodiments of the present application;

[0044] FIG10 shows a schematic block diagram of an apparatus for data recovery according to some embodiments of the present application; and

[0045] FIG11 shows a schematic block diagram of an example device that can be used to implement embodiments of the present application. DETAILED DESCRIPTION

[0046] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.

[0047] In the description of the embodiments of this application, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. can refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0048] To ensure data reliability and restore the corresponding data in the event of a node failure in a multi-node storage system, redundant data must be added to the storage system. A traditional method for generating redundancy is repetition coding. For example, two additional copies of each data block can be generated. These three copies of the data block can then be stored in three storage nodes. If a node fails, data can be restored by copying the same copy from any other node to the recovery node. However, this approach incurs significant additional storage resource overhead, resulting in wasted resources.

[0049] More storage systems use erasure codes to maintain data availability to avoid data loss due to storage node failure. EC code is a coding fault-tolerant technology that divides data into multiple parts and uses these parts to generate redundant check blocks. These parts are then stored in different locations. For example, when data is being downloaded from a storage system, the system can divide the original data into k original data blocks and apply a specific EC code scheme to these data blocks to generate r redundant check blocks. However, the system can write each block of the data block and the check block to a different storage node among n = k + r storage nodes for downloading.

[0050] Using the corresponding EC code scheme, the encoded data can be verified. In the event of data errors, disk damage, or node failures, valid data blocks and corresponding parity blocks can be read from intact disks or valid nodes to perform erasure code decoding. This allows the lost data blocks to be recovered and written to repair disks or repair nodes, maintaining the data reliability of the storage system.

[0051] MDS array code is a type of erasure code. The codeword encoded by the MDS array code consists of k+r units, where each unit has α bits of data or symbols. In the codeword, the k units are units from the k blocks of the original data, and the r units are check units calculated based on these original data units. Moreover, the array code satisfies the MDS property, so that any k units among the k+r units are sufficient to reconstruct all the k original data units. In other words, it can tolerate the failure of any r data units. For example, the MDS array code can be used in a redundant array of independent disks (RAID) with k data disks and r redundant check disks, where the data of the same unit of each codeword can be stored in the same disk, thereby allowing the failure of r disks in the RAID.

[0052] Conventional erasure coding schemes are encoded and decoded over finite fields, and the encoding and decoding process requires multiplication operations in finite fields. Finite fields are also called Galois Fields (GF). This paper uses GF (p m) is used to refer to a finite field with characteristic p and degree m in its prime subfield. The multiplication operation in a finite field is much more complicated than the exclusive OR (XOR) operation, which makes the computational task of the encoding process very heavy. In addition, the encoding process does not support incremental encoding, which is not friendly to servers with weak computing power. Some methods convert the finite field elements in the check matrix used into a binary matrix (for example, a 0-1 matrix) in an attempt to convert the multiplication operation in the encoding process into an exclusive OR operation. However, the matrix converted in this way is relatively dense, and the encoding and decoding efficiency is relatively low. In addition, these methods cannot avoid the matrix inversion operation during the decoding process.

[0053] Other schemes aim to construct k+2 pattern array codes using only XOR operations, which can tolerate up to two node failures and achieve the MDS property. These pure XOR MDS codes require fine-grained data slicing, with the total number of slices being called the number of packets. However, the number of packets in these schemes is approximately equal to the code length. A large number of packets leads to a sharp increase in the complexity of the encoding and decoding implementation, which makes industrial applications more complex.

[0054] To address the above and other issues, the present application provides a solution for data verification and recovery. This solution utilizes a sparse binary array code based on an extended field representation to determine the corresponding check units for a set of original data units, thereby encoding the original data units into codewords. In this solution, a check matrix is ​​converted from a regular and typically sparse matrix of elements in a finite field. The finite field matrix is ​​constructed based on the codeword structure and has the MDS property. Thus, the converted check matrix is ​​a binary matrix representation of the corresponding finite field matrix. In other words, the check matrix can be divided into multiple sub-matrices of the same order, where each sub-matrix is ​​a binary matrix representation of elements at corresponding positions in the finite field matrix. The check matrix has the sparse and regular properties derived from the corresponding finite field element matrix. Using this check matrix, pure XOR encoding and decoding of erasure codes can be performed, which has an optimized small subpacket number level, can reuse intermediate calculation results during encoding and decoding, and can avoid complex matrix inversion operations during decoding, thereby improving encoding and decoding efficiency. As an example, the data checksum recovery scheme of the present application can support, for example, a coding mode with mostly 2 redundancy (ie, k+2), so that its coding and decoding performance can reach or approach the theoretical limit in some cases.

[0055] The embodiments of the present application can be applied to storage systems to perform erasure coding on data to be stored, thereby improving the reliability of the storage system. For ease of explanation, the embodiments of the present application will be described below in the context of a storage system. However, it should be understood that the embodiments of the present application can also be applied to other application scenarios to encode codewords and perform data verification and recovery.

[0056] First, reference is made to Figure 1, which shows a schematic diagram of an example environment 100 in which multiple embodiments of the present application can be implemented. The example environment 100 includes data nodes 110-1 to 110-k (individually or collectively referred to as data nodes 110), and check nodes 120-1 to 120-r (individually or collectively referred to as check nodes 120). It should be understood that the data nodes 110 and check nodes 120 in Figure 1 are distinguished based on the data content they are configured to store. Among them, the data node 110 is intended to store original data, and the check node 120 is intended to store check data calculated based on the original data, thereby providing redundancy protection for the original data. The physical structure of each node in the data node 110 and the check node 120 may be the same, and the embodiments of the present application do not limit this. For example, the disks in the data node 110 and the check node 120 can be disks that together constitute a RAID.

[0057] Environment 110 also includes a computing device 130. Computing device 130 may be a control device connected to data nodes 110 and check nodes 120, and may perform various actions according to embodiments of the present application. Examples of computing device 130 include, but are not limited to, desktop computers, laptop computers, servers, and the like. Furthermore, although illustrated as a single entity, computing device 130 may also be implemented as a centralized device group, a distributed device, a cloud-based device, and the like.

[0058] In response to receiving a command to store data, computing device 130 may divide the data to be stored into k data blocks. For example, the value of k may be equal to the number of data nodes 110. As shown, computing device 130 may divide data 135 into data blocks 115-1 through 115-k. Furthermore, computing device 130 may perform checksum coding on the data blocks, for example, according to the methods of embodiments of the present application, to calculate r checksum blocks based on the k data blocks. For example, the value of r may be equal to the number of checksum nodes.

[0059] In some embodiments, computing device 130 may write a corresponding data block of k data blocks to each of data nodes 110, and write a corresponding check block of r check blocks to each of check nodes. For example, computing device 130 may determine check blocks 125-1 to 125-r based on data blocks 115-1 to 115-k, and write these data blocks and check blocks to the corresponding nodes, as shown in FIG1 .

[0060] The computing device 130 may write the data into corresponding locations of each node based on a predetermined specification or protocol, so that the computing device 130 can read and combine the corresponding data from the node when needed. In an embodiment of the present application, during encoding, the computing device 130 may divide each data block into multiple slices (i.e., so that each data unit includes multiple subpackets).

[0061] For example, for a subpacket number α, during encoding, the computing device 130 may obtain α bits of data from each of the k data blocks as a basic unit of encoding, and calculate r α-bit check units based on these k data units. When reading data, based on the check encoding method, the computing device 130 may use the r check units corresponding to the k data blocks read to determine whether the k data blocks are consistent with the stored original data.

[0062] In some cases, data may be erroneous during storage, or some of the data nodes 110 may fail, rendering some of the k data blocks of stored data unusable. In such cases, the computing device 130 can utilize the remaining data blocks to recover the lost data based on the parity encoding scheme used. For example, as described in greater detail below, the computing device 130 can recalculate the lost data based on an array code according to an embodiment of the present application.

[0063] When the array code used has the MDS property, the computing device 130 can recover the data on all data nodes 110 from the corresponding data from any k nodes among the data nodes 110 and the check nodes 120. In other words, the array code allows any r nodes in a group of nodes having r redundant nodes to fail.

[0064] It should be understood that the architecture and functionality of the example environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present application. Furthermore, other devices, systems, or components not shown may also exist in the example environment 100. For example, in some embodiments, device 130 may be a server of a cloud-based storage system, and environment 110 may also include a client device. The client device may send requests, such as data read and storage, to server 130, thereby triggering various actions described in the embodiments of the present application.

[0065] For example, in some embodiments, performing checksum encoding on data and recovering lost data can be performed by computing device 130 and another computing device (not shown) different from device 130, respectively. In addition, embodiments of the present application can also be applied to other environments with different and / or other functions. For illustrative purposes, embodiments of the present application will be described below in the context of actions performed by computing device 130.

[0066] Reference is now made to FIG. 2 , which illustrates a flow chart of an example method 200 for data verification according to some embodiments of the present application. Example method 200 may be performed, for example, by computing device 130 as shown in FIG. It should be understood that method 200 may also include additional actions not shown, and the scope of the present application is not limited in this respect. Method 200 is described in detail below in conjunction with example environment 100 of FIG. 1 .

[0067] At 210, a set of data units is obtained, each data unit in the set of data units having multiple subpackets. For example, for data nodes 130-1 to 130-k, computing device 130 may obtain a set of k data units, each data unit having multiple subpackets. For example, a basic unit to be encoded into a single codeword may have multiple bits of data or symbols.

[0068] In some embodiments, the computing device 130 may receive a command to store data, which instructs the computing device 130 to store specific data in the data nodes 110-1 to 110-k. In order to promote the security of the stored data, the computing device 130 may divide the data into k blocks so as to be stored separately on each of the data nodes 110-1 to 110-k. Furthermore, the computing device 130 may divide each part into the same number of slices to facilitate array code-based checksum coding. The total number of slices of each part is called the number of subpackets. For example, the computing device may slice the data into α subpackets. Based on such a structure, for example, the computing device 130 may obtain a group of k data units each having α bits or symbols as the data portion of the coding unit for checksum coding.

[0069] At 220, a check matrix is ​​obtained, which is a binary matrix representation of a finite field matrix, where each element in the finite field matrix comes from a specific finite field. For example, the finite field can be a binary extended field GF(2 n ). Two extended domain GF(2 n ) is a finite field of n-th extension of the prime field GF(2) (i.e., its degree over its prime subfield is n, and is also referred to as its degree n in this paper), which contains 2 n Such a check matrix can be divided into multiple sub-matrices of the same order, where each sub-matrix is ​​a binary matrix representation of the elements at the corresponding position in the finite field matrix.

[0070] For example, the computing device 130 may obtain a check matrix that can be divided into a plurality of submatrices of the same order, wherein each submatrix is ​​a 0-1 binary matrix representation (e.g., a 0-1 binary matrix representation) of an element at a corresponding position in the matrix of an element in a finite field. In some embodiments, the computing device 130 may load a check matrix stored as part of the checksum coding software at the beginning of the overall coding process. In other embodiments, the computing device 130 may access a check matrix stored in its coding chip.

[0071] As will be understood by those skilled in the art, the elements in a finite field can be represented by a binary matrix. 4 ), the element i can first be converted into a binary expression (i0,i1,i2,i3), where On this basis, i can be further converted into the following 0-1 binary matrix The matrix A is:

[0072] Therefore, the finite field GF(2 n ) can have a corresponding n-order binary matrix representation. In another example, the finite field GF(2 2 The corresponding binary matrix representation of each element in ) can be shown as follows:

[0073] Later we will discuss the finite field GF(2 2 ) to describe the specific construction of the check matrix in more detail. According to an embodiment of the present application, the binary matrix obtained by the computing device 130 for checking the code has sparseness, that is, the number of its 0 elements is much greater than the number of its non-0 elements (for example, the ratio between the two meets the threshold). And, based on the finite field GF(2 2 ), the binary check matrix can be exactly divided into a plurality of second-order (2×2) sub-matrices, where each sub-matrix corresponds to the binary matrix representation of one of the above elements. In other words, the check matrix can be converted into a finite field GF(2 2 ) is a matrix containing the elements in .

[0074] At 230, a set of check units is determined based on the check matrix and the set of data units. The number of subpackets for each check unit in the set of check units is the same as the number of subpackets for each unit in the set of data units. For example, for check nodes 120-1 to 120-r, computing device 130 may determine a set of r check units for the set of data units based on the set of k data units obtained at 210 and the check matrix obtained at 220. Just as each data unit has α bits of data or symbol, each of the determined check units also has α bits of data or symbol.

[0075] As will be described in more detail in the examples below, the check matrix contains information about how to encode the data block. Therefore, the computing device 130 can generate a check block for the data block based on the encoding information indicated in the check matrix. Let H represent the check matrix. For a code with data dimension k and redundancy r, the computing device 130 can calculate the check data based on the following relationship: H×[m,p] T =H×[m1,m2,…,m k ,p1,…,p r ] T =0

[0076] where m i (1≤i≤k) is the data unit in the original data block, m i =[m i,1 ,m i,2 ,…,m i,α ] contains α data bits or data symbols; p i (1≤i≤nk) is the check unit in the redundant data block generated by encoding, p i =[p i,1 ,p i,2 ,…,p i,α ] contains α check bits or check symbols.

[0077] In some embodiments, after determining the check blocks according to method 200, computing device 130 may download the encoded data to disk. For example, computing device 130 may write the original data block to data node 110 and the check block to check node 120. Ideally, computing device 130 may write one data block from the original data block to each data node and one check block from the check blocks to each check node in check node 120. This provides redundancy protection for the stored data. Subsequently, under the above ideal conditions, when up to r data nodes in data node 110 fail, computing device 130 may utilize the data on the surviving nodes in data node 110 and check node 120 to perform decoding based on the check matrix to recover the data lost on the failed node.

[0078] The check matrix used in method 200 has a sparse regularity characteristic derived from its corresponding finite field element matrix. Using such a matrix for check coding can achieve high-performance small subpacket pure XOR coding, and can avoid complex matrix inversion operations during subsequent decoding. For example, the data check scheme of the present application can support most 2-redundant (i.e., k+2) coding modes, so that its encoding and decoding performance can reach or approach the theoretical limit in some cases. The following will illustrate the check matrix construction and encoding and decoding process according to the embodiment of the present application with reference to more specific examples.

[0079] First, the construction of the check matrix according to the embodiment of the present application is described by taking the redundancy r=2 as an example. In this case, for an array code with an array dimension of α (equal to the number of subpackets) and a code length of n=k+2), the final check matrix H consists of two concatenated parts, as shown in the following formula: H=(H1|I 2α×2α )

[0080] Among them, I 2α×2α is a 2α×2α unit matrix of order 2α, I α×α is the α×α identity matrix, A i is a 0-1 binary matrix of α×α. That is, H1 can be regarded as a submatrix of order α as a construction unit, which has two (2) rows of construction units equal to the redundancy 2, where each construction unit of the first row is a unit matrix, and has k columns of construction units equal to the number of data units k.

[0081] For the 0-1 matrix H1, we can first construct a finite field GF(2 n ) Each element is GF(2 n ), and I is the same as B i Using the matrix representation of the elements in the finite field as described above, each matrix I can be expanded into the corresponding position I in H1 α×α matrix, and each B i Can be expanded into A at the corresponding position of H1 i .

[0082] That is, for 1≤i≤k),

[0083] Can be expanded into a 0-1 matrix

[0084] in And in the matrix A i Corresponding to According to the shape of the check matrix based on the codeword structure to be encoded, it is easy to know that each for 0-1 matrix, and the finite field GF(2 n ) on its prime subfield GF(2) satisfies the number n That is, the number of subpackets of the code is α=n*t.

[0085] The minimum number of subpackets supported by the current code type can be selected based on the code length. According to an embodiment of the present application, for 4 subpackets, the supported code length can be up to 9, and for 8 subpackets, the supported code length can be up to 22. Compared to other array codes where the number of subpackets needs to be close to the code length (for example, a code length of 22 may require 22 subpackets), the check matrix used in the embodiment of the present application can support encoding with a smaller number of subpackets when the codeword is long, thereby reducing computational complexity.

[0086] Furthermore, n can be selected to be divisible by the number of subpackets. For the purpose of computational simplicity, a lower n can be selected to experiment with constructing matrices of finite fields with MDS properties. For example, GF(2 2 ) is used as an example to describe some examples of this application. From these examples, it can be seen that the encoding and decoding performance of the check matrix constructed based on this domain can reach or approach the theoretical limit in some cases. After determining other parameters, the B to be constructed i The order t of can be derived accordingly. In some embodiments, it is possible that each B i Can include GF(2 2 ), namely B i is a first-order matrix, for example, as shown in the example described later with respect to FIG. 5 .

[0087] Continuing with the redundancy of 2, we take t = 2 as an example to illustrate the finite field matrix H 1′ For 1≤i≤k, we can sequentially add H 1′ Fill in the sparse finite field matrix B i , such as diagonal, anti-diagonal, upper and lower triangular, or anti-upper and lower triangular matrices. Then, H 1′ Can be expanded later into I 2α×2α The unit matrix of B is concatenated to obtain the representation of the check matrix H on the finite field. In this representation, B i Pairwise matching, that is, each pair of adjacent sparse submatrices B i The values ​​of the corresponding positions of H are the same or sum to 1. In other words, every two check sub-matrices are the same or differ by one unit matrix, which is applicable when the number of sub-packets does not exceed 4 (for example, 2 sub-packets, 4 sub-packets). In addition, each sub-matrix whose order of the sparse matrix is ​​equal to the number of rows of H is full rank, so that the MDS property can be satisfied.

[0088] The structure of the finite field matrix described above is sparse and highly regular, and the 0-1 parity check matrix converted from it also retains these properties. Using such a parity check matrix for encoding allows a large number of intermediate calculation results to be reused, significantly reducing encoding complexity. Furthermore, the efficiency of the corresponding decoding process used for data recovery can be improved.

[0089] Reference is now made to FIG3 , which illustrates a flow chart of an example method 300 for data recovery according to some embodiments of the present application. For example, example method 300 can be used in combination with example method 200 to recover damaged portions of data that has been checksum-coded using example method 200. Example method 300 can be performed, for example, by computing device 130 as shown in FIG1 . As another example, example method 200 and example method 300 can be performed by computing device 130 and another computing device, respectively. It should be understood that method 300 may also include additional actions not shown, and the scope of the present application is not limited in this respect. For illustrative purposes, method 300 is described in detail below in conjunction with example environment 100 of FIG1 .

[0090] At 310 , a group of data units and a corresponding group of check units are obtained, where the group of check units is determined based on a group of data units and a check matrix, and the group of data units includes data units to be recovered.

[0091] For example, computing device 130 may obtain a set of data units and a corresponding set of check units, where the set of check units is determined based on the set of data units and the check matrix, and the set of data units includes the data units to be recovered. For example, the set of check units may be determined by computing device 130 or another computing device using method 200 and written to data node 110 and check node 120. Computing device 130 may read and combine the corresponding data based on the distribution of the set of data units and corresponding check units in the nodes.

[0092] At 320 , a parity check matrix is ​​obtained, where the parity check matrix is ​​a binary matrix representation of a finite field matrix, each element of which is from a specific finite field. For example, the computing device 130 may obtain a parity check matrix corresponding to the data unit and parity check unit obtained at 310 .

[0093] In an embodiment of the present application, the matrix can be divided into a plurality of parity check matrices of the same order submatrices, wherein each submatrix is ​​a 0-1 binary matrix representation (e.g., a 0-1 binary matrix representation) of an element in a matrix corresponding to an element in a finite field, as described above with respect to method 200. In some embodiments, computing device 130 can load a parity check matrix stored as part of its parity check encoding software at the beginning of the overall decoding process. In other embodiments, computing device 130 can access a parity check matrix embedded in its decoding chip.

[0094] At 330, the value of the data unit to be recovered is determined based on the set of data units, the set of check units, and the check matrix. For example, computing device 130 may determine the value of the data unit to be recovered from the remaining data in the obtained unit based on the data unit, check unit, and check matrix obtained at 310 and 320, according to the decoding information indicated in the check matrix, as described in more detail below.

[0095] Similar to method 200, the check matrix used in method 200 has a sparse regularity characteristic derived from its corresponding finite field element matrix. When decoding the data encoded with it to recover the data, a large number of intermediate calculation results can be reused, and complex matrix inversion operations can be avoided, significantly reducing the complexity of decoding. For example, the data recovery scheme of the present application can support most 2-redundancy (i.e., k+2) encoding and decoding modes, so that its decoding performance can reach or approach the theoretical limit in some cases.

[0096] The following describes how the inversion operation is eliminated during the decoding process in conjunction with FIG4 . FIG4 shows an example schematic diagram 400 according to some embodiments of the present application, which illustrates an example structural form of a coefficient matrix for the decoding process. The example in FIG4 takes the failure of the first two data nodes as an example. In this case, according to the symbolic representation of the array code structure above, the decoding process for recovering the data of the first two nodes requires calculating However, due to the characteristics of the check matrix of the embodiment of the present application, the matrix inversion operation can be converted into a few simple elimination transformation operations.

[0097] For example, let the finite field matrices B1 and B2 corresponding to the 0-1 matrices A1 and A2 be upper and lower triangular matrices, respectively. According to the matrix construction described above, the corresponding positions of the matrices B1 and B2 are the same or the sum is 1. Then the expanded 0-1 matrix A1 can be in the form shown by reference numeral 410, and the expanded 0-1 matrix A2 can be in one of the several forms indicated by reference numeral 420. As shown in the figure, both matrices A1 and A2 are relatively sparse and regular. By reusing the intermediate results, the above coefficient matrix It can be transformed into a form as shown in one of the reference numerals 430 - 1 to 430 - 3 . Finally, by simple row and column permutation, the above matrix can be transformed into a ladder matrix, thereby eliminating the matrix inversion calculation of the decoding process.

[0098] In general, when constructing a check matrix according to an embodiment of the present application, a corresponding high-order finite field matrix can be first constructed. If it is determined under a finite field that the non-zero bit structure of the check matrix has the possibility of satisfying the MDS property, the finite field is traversed to determine the elements at each position to construct a finite field matrix, where the sub-matrices are matched pairwise as described above. The matrix of the finite field can then be converted into a binary check matrix using the matrix representation of the finite field. This binary check matrix can then be used in the encoding and decoding process to determine the check data and the recovered data.

[0099] In such an embodiment, similar to the finite field matrix corresponding to the above-mentioned check matrix H, the constructed finite field matrix includes a first part. The first part is a two-field extended field GF (2 n ) is a construction unit, where the product of the construction unit's order and the degree of the second extended field used is equal to the number of subpackets α for each data unit. The finite field matrix further includes a second part concatenated with the first part. The second part is the identity matrix, and its order is easily understood to be equal to the number of columns in the first part.

[0100] In some embodiments, the construction unit of the first part can be a first-order matrix (expressed as a single element in the resulting finite field matrix). In some embodiments, the construction unit of the first part can be a sparse submatrix of the same order with elements in the two-expansion field. Moreover, the construction unit in the first row of construction units in the first part is a unit matrix. The same row of construction units in the first part includes multiple pairs of construction units, wherein each pair of construction units is equal, or the units in each pair of construction units differ from each other by a permutation matrix after each conversion into a binary matrix representation. When the number of subpackaging is 4 or less, this characteristic can simplify the values ​​of the corresponding positions of each pair of construction units to be the same or to be 1. In addition, each m-order submatrix of the finite field matrix is ​​full rank, where m is equal to the number of rows of the finite field matrix.

[0101] Specifically, for the encoding of k data units plus r redundant units, the first part of the finite field matrix includes k construction units in the row direction and r construction units in the column direction, and each unit in the first row of units is the identity matrix. i In this way, the remaining r-1 rows can be filled with r-1 rows of construction units, so that these construction units and the submatrix that meets the conditions in the resulting finite field matrix are full rank, such as each m-order submatrix as described above. For example, when the redundancy is 3, two more rows of construction units will be filled under the unit matrix of the first row. In practice, simply converting the rows and columns of the check matrix having the above structure also falls within the scope of the check matrix according to the embodiment of the present application.

[0102] The embodiments of the present application can be applied in encoding / decoding computing devices of storage systems, for example, in server devices based on software computing, or in chips based on hardware computing. A computing device (for example, computing device 130) can use a coding check matrix according to an embodiment of the present application to encode the original data in the storage system to generate redundant check data. Figure 5 shows a schematic diagram 400 of encoding and decoding of erasure codes in a non-limiting example storage system, which schematically shows the changes of various data based on the original data during the encoding and decoding process. The embodiments of the present application can be applied in this storage system to encode and decode erasure codes. For the sake of illustration, Figure 4 will be described below in the context of the actions performed by the computing device 130 of Figure 1.

[0103] As a non-limiting example, the storage system in FIG5 may include 4 data nodes and 2 check nodes, which may be example implementations of the data nodes 110-1 to 110-k and the check nodes 120-1 to 120-r in FIG1 , respectively. To cope with disk deletion errors, when downloading the original data 505 to disk, the computing device 130 may partition the original data 505 into blocks, thereby generating original data blocks 515-1 to 515-4. Next, the computing device 130 may encode the original data blocks 515-1 to 515-4 using a method according to an embodiment of the present application, generating redundant data blocks 525-1 and 525-2. The computing device 130 may then write each of the original data blocks 515-1 to 515-4 and the redundant data blocks 525-1 and 525-2 to different storage nodes for downloading.

[0104] When a node fails due to reasons such as disk damage, the data stored on it will be lost. In this example, as shown in the figure 550, the original data blocks 515-1 and 515-2 are lost. In this case, the computing device 130 can read valid original data or redundant data from the surviving node and decode it according to the method of the embodiment of the application to generate recovery data blocks 535-1 and 535-2 corresponding to the lost original data blocks 515-1 and 515-2, thereby reconstructing the node. The computing device 130 can then write the recovery data blocks 535-1 and 535-2 to the corresponding repair nodes, thereby maintaining the data reliability of the storage system.

[0105] It should be understood that in the example of FIG5 , the number and order of data nodes, check nodes, and failed nodes are shown only to facilitate understanding, and the embodiments of the present application are not limited to the specific number and order of nodes. It should be understood that in a storage system with redundancy of r, such as that shown in FIG1 , an erasure coding scheme with MDS properties can tolerate the failure of any number of nodes from 1 to r.

[0106] Furthermore, in some embodiments, due to limitations such as hardware conditions, check nodes and data nodes may not correspond one-to-one with the data blocks and check blocks of a codeword. For example, the number of computing nodes may be less than k, and / or the number of check nodes may be less than r. In this case, after performing check encoding, computing device 130 may write the data blocks and check blocks to each node in the currently available optimized distribution. In some such embodiments, a node may store one or more data blocks for a certain set of data and one or more check blocks for another set of data. Thus, when recovering data, computing device 130 can retrieve the corresponding k+r blocks for decoding based on the distribution in which the data were stored. Thus, when the array code used has the MDS property, computing device 130 can recover any r blocks.

[0107] In some embodiments, during encoding, computing device 130 may determine an intermediate value of a check unit corresponding to a pair of submatrices and a set of data units in a check matrix based on the values ​​of the submatrices. The submatrix pair corresponds to a pair of adjacent construction units in the row direction of the corresponding finite field matrix. When determining the value of the check unit for the data unit, computing device 130 may reuse the intermediate values, thereby reducing computational overhead and improving encoding efficiency.

[0108] Correspondingly, computing device 130 may also reuse calculated intermediate values ​​during the decoding process to improve decoding efficiency. In some embodiments, when determining the value of a data unit to be recovered, computing device 130 may determine, based on a parity check matrix, a calculation path for calculating the value of the data unit to be recovered, as well as intermediate values ​​generated in the calculation path. Computing device 130 may then further determine the value of the data unit to be recovered based on the calculation path and the intermediate values, thereby recovering the lost data.

[0109] For example, when some of the data nodes 110 fail, the computing device 110 can read the surviving data of the encoded codeword from the surviving data nodes 110 and check nodes 130, and determine the codeword to be decoded by combining the sequence numbers corresponding to the surviving nodes and the failed nodes. It should be understood that the codeword mentioned herein refers to a logical concept including a set of data units and a set of check units. Depending on the specific implementation, during actual execution, the computing device 130 may not physically assemble the codeword structure, but instead directly calculate the data to be recovered based on the decoding calculation path.

[0110] The following describes three illustrative examples according to some embodiments of the present application in conjunction with Figures 6-8. In these examples, based on the structure of the codeword to be encoded (for example, determined based on the number of check nodes and redundant nodes in the storage system), the corresponding array code is determined according to the code pattern construction of the embodiments of the present application. These examples then further illustrate the encoding process of encoding data using the determined array code, and the decoding process of recovering lost data using the array code. The relevant actions in the examples of Figures 6-8 can be performed, for example, by the computing device 130 in Figure 1. Alternatively, in some embodiments, the encoding action and the decoding action can also be performed by different devices. For illustration purposes, the examples of Figures 6-8 will be described below in the context of the actions performed by the computing device 130.

[0111] FIG6 illustrates a schematic diagram 600 of an example binary array code according to some embodiments of the present application. This example binary array code is constructed for a 2-packet [4,2] encoding and exhibits the MDS property. Specifically, this binary array code is intended to encode a codeword structured as follows: a code length of 4, a data length of 2, and each data unit containing 2 bits of data or a symbol. For example, such an array code can be used in a storage system with two data nodes and two check nodes, and can recover all original data even if at most two nodes fail.

[0112] According to some embodiments of the present application, in order to construct the array code in the diagram 600, a regular corresponding finite field matrix may be first constructed. In this example, a smaller number of 2 packets has been determined based on the code length 4. On this basis, an attempt is made to use GF(2 2 ) constructs a finite field matrix that satisfies the MDS property and obtains a finite field matrix 610. According to the code pattern described above, the constructed finite field matrix 610 is composed of two parts 611 and 612, where part 612 is the identity matrix.

[0113] Furthermore, based on the constraints between the parameters defining the pattern structure, it is easy to understand that the construction unit of portion 611 is a first-order matrix (i.e., represented by a single element in the constructed matrix 610). Portion 611 has two columns of construction units corresponding to the data length of 2, and two rows of construction units corresponding to the redundancy of 2. Since the order of each construction unit is 1, portion 611 has two columns and two rows of finite field elements. Thus, after being concatenated with portion 612, which serves as the identity matrix, the resulting finite field matrix 610 has four columns and two rows of finite field elements.

[0114] Furthermore, the first row of the building blocks of the portion 611 is filled with 1, and the remaining second row is filled with GF(2 2) so that each 2×2 submatrix in the finite field matrix 610 is full rank. On this basis, according to the binary matrix representation of the finite field elements described above, each element in the finite field matrix 610 can be replaced with the corresponding binary matrix representation to obtain the final 0-1 check matrix. For ease of reading, GF(2 2 ) is again represented as follows:

[0115] Based on the conversion table, the finite field matrix 610 can be converted into a final binary check matrix 620. The computing device 130 can maintain the check matrix 620 in software or hardware form for encoding and decoding. Those skilled in the art will understand that during the encoding process, the computing device 130 can determine two corresponding check units from two 2-packet data units based on the matrix 620. For illustration, the first data unit is A = (a1, a2), the second data unit is B = (b1, b2), and the two check units to be determined are P = (p1, p2) and Q = (q1, q2) respectively. According to the check matrix 620, the value of the check unit can be determined as follows: p1=1a1+0a2+1b1+0b2=a1+b1; p2=0a1+1a2+0b1+1b2=a1+b2; q1=0a1+1a2+1b1+1b2=p2+b1; q2=1a1+1a2+1b1+0b2=p1+c2.

[0116] It should be understood that, unless otherwise indicated, the "+" sign in this document refers to an XOR operation. As shown in the above formula, the intermediate results p1 and p2 can be reused during the encoding process. Therefore, the actual encoding complexity becomes 4 XOR operations, requiring an average of 2 XOR operations per bit, reaching the theoretical limit.

[0117] In a storage system scenario, the aforementioned data units can be stored in corresponding nodes. The following example illustrates the decoding process using parity check matrix 620, assuming the nodes containing two data units A and B fail. In this scenario, computing device 130 can use parity check matrix 620 to calculate the values ​​of data units A and B. Corresponding to the aforementioned encoding process, the values ​​of the data units can be recovered as follows: b1 = q1 + p2; a2 = q2 + p1; b2 = p1 + a2; a1 = p2 + b1.

[0118] As shown in the above equation, the intermediate results a2 and b1 can be reused during the decoding process. Therefore, the actual decoding complexity is reduced to four XOR operations, requiring an average of two XOR operations per bit, which also reaches the theoretical limit. Furthermore, this decoding process does not require matrix inversion.

[0119] FIG7 shows a schematic diagram 700 illustrating the construction of another example binary array code according to some embodiments of the present application. The binary array code in this example is constructed for a 4-packet [6,4] encoding and exhibits the MDS property. Specifically, the binary array code can be used to encode a codeword structured as follows: a code length of 6, a data length of 4, and each data unit containing 4 bits of data or symbols. For example, such an array code can be used in a storage system with 4 data nodes and 2 check nodes, and can recover all original data even if up to two nodes fail.

[0120] According to some embodiments of the present application, in order to construct the array code in the diagram 700, a regular corresponding finite field matrix may be first constructed. In this example, the number of packets is determined to be 4 based on the code length of 6. On this basis, try GF(2 2 ) constructs a finite field matrix that satisfies the MDS property and obtains a finite field matrix 710. According to the code pattern described above, the constructed finite field matrix 710 is composed of two parts 711 and 712, where part 712 is the identity matrix.

[0121] Furthermore, based on the constraints between the parameters defining the pattern structure (as described above in conjunction with Figures 2 and 4 ), it is easy to understand that the construction units of portion 711 are a 2×2 matrix of finite field elements. Portion 711 has four columns of construction units corresponding to a data length of 4, and two rows of construction units corresponding to a redundancy of 2. Therefore, portion 711 has 8 columns and 4 rows of finite field elements. Thus, after being concatenated with portion 712, which serves as the identity matrix, the resulting finite field matrix 710 has 12 columns and 4 rows of finite field elements.

[0122] Furthermore, the first row of the construction unit of the portion 711 is filled with a second-order identity matrix, for example, as shown by reference numeral 713. The remaining second row is filled with GF(2 2 ) in a second-order sparse matrix (such as a diagonal, anti-diagonal, upper and lower triangular, or anti-upper and lower triangular matrix), for example, as shown by reference numeral 714. The filling B can be as described above. i The second-order sparse matrix of the row is filled in a manner such that these second-order sparse matrices match each other, wherein the elements at each corresponding position are the same or sum to 1, and each 4×4 submatrix in the finite field matrix 710 is full rank.

[0123] On this basis, according to GF(2 2 ) is represented by the following binary matrix. Since each element in matrix 710 can be replaced by the corresponding binary matrix representation, the final 0-1 check matrix is ​​obtained:

[0124] Thus, the finite field matrix 710 can be converted into a final binary check matrix 720. The computing device 130 can maintain the check matrix 720 in software or hardware form for encoding and decoding. Those skilled in the art will understand that during the encoding process, the computing device 130 can determine two corresponding check units from four 4-packet data units based on the matrix 720. For illustration, the data units are A = (a1, a2, a3, a4), B = (b1, b2, b3, b4), C = (c1, c2, c3, c4), and D = (d1, d2, d3, d4), and the two check units to be determined are P = (p1, p2, p3, p4), Q = (q1, q2, q3, q4). Similar to the example of Figure 6, according to the check matrix 720, the value of the check unit can be determined as follows: x1=a1+b1 x2=a2+b2 x3=a3+b3 x4=a4+b4 y1=c1+d1 y2=c2+d2 y3=c3+d3 y4=c4+d4 p1=x1+y1 p2=x2+y2 p3=x3+y3 p4=x4+y4 q1=x1+y2+a3+d1 q2=x2+y1+a4+c2 q3=x3+y4+b1+d3 q4=x4+y3+b2+c4.

[0125] As shown in the above equation, during the encoding process, due to the sparse and regular nature of parity check matrix 720, the intermediate calculation results x1, x2, x3, x4 and y1, y2, y3, y4 corresponding to the submatrices in parity check matrix 720 can be reused. Therefore, the actual encoding complexity becomes 24 XOR operations, requiring an average of 2 XOR operations per bit, reaching the theoretical limit.

[0126] In a storage system scenario, the above data units can be stored in corresponding nodes respectively. The following uses the failure of the nodes where two data units A and B are located as an example to illustrate the decoding process using the check matrix 720. In this case, the computing device 130 can use the check matrix 720 to calculate the values ​​of data units A and B. Corresponding to the above encoding process, the values ​​of the data units can be restored as follows: x1=c1+d1 x2=c2+d2 x3=c3+d3 x4=c4+d4 y1=p1+x1 y2=p2+x2 y3=p3+x3 y4=p4+x4 a3=q1+x2+y1+d1 a4=q2+x1+y2+c2 b1=q3+x4+y3+d3 b2=q4+x3+y4+c4 b3=y3+a3 b4=y4+a4 a1=y1+b1 a2=y2+b2

[0127] As shown in the above equation, during the decoding process, the intermediate calculation results x1, x2, x3, x4 and y1, y2, y3, y4, etc. can be reused. Therefore, the actual decoding complexity is reduced to 24 XOR operations, requiring an average of 2 XOR operations per bit, which also reaches the theoretical limit. Furthermore, this decoding process does not require matrix inversion operations.

[0128] FIG8 shows a schematic diagram 800 illustrating the construction of another example binary array code according to some embodiments of the present application. The binary array code in this example is constructed for a 4-packet [8,6] encoding and exhibits the MDS property. Specifically, the binary array code can be used to encode a codeword with a code length of 8, a data length of 6, and each data unit containing 4 bits of data or symbols. For example, such an array code can be used in a storage system with 6 data nodes and 2 check nodes, and can recover all original data even if up to two nodes fail.

[0129] According to some embodiments of the present application, in order to construct the array code in the diagram 800, a regular corresponding finite field matrix may be first constructed. In this example, a smaller number of 4 packets has been determined based on the code length of 8. On this basis, an attempt is made to use GF(2 2 ) constructs a finite field matrix that satisfies the MDS property and obtains a finite field matrix 810. According to the code pattern described above, the constructed finite field matrix 810 is composed of two parts 811 and 812, where part 812 is the identity matrix.

[0130] Furthermore, based on the constraints between the parameters defining the pattern structure, it's easy to understand that the construction units of portion 811 are a 2×2 matrix of finite field elements. Similar to the two examples above, portion 811 has six columns of construction units corresponding to a data length of 6, and two rows of construction units corresponding to a redundancy of 2. Therefore, portion 811 has 12 columns and 4 rows of finite field elements. Thus, after concatenating it with portion 812, which serves as the identity matrix, the resulting finite field matrix 810 has 16 columns and 4 rows of finite field elements.

[0131] Furthermore, the first row of the building blocks of part 811 is filled with the second-order identity matrix. The remaining second row is filled with GF(2 2 ) is a second-order sparse matrix of elements in ). Similarly, B can be filled as described above i The second-order matrices of the row are filled in a manner such that these second-order matrices correspond to each other, the elements in each corresponding position are the same or the sum is 1, and each 2×2 submatrix in the finite field matrix 810 is full rank.

[0132] On this basis, according to GF(2 2) is represented by the following binary matrix. Since each element in matrix 810 can be replaced by the corresponding binary matrix representation, the final 0-1 check matrix is ​​obtained:

[0133] Thus, the finite field matrix 810 can be converted into a final binary check matrix 820. The computing device 130 can maintain the check matrix 820 in software or hardware form for encoding and decoding. Those skilled in the art will understand that during the encoding process, the computing device 130 can determine two corresponding check units from six 4-packet data units based on the matrix 620. For illustration, the data units are A = (a1, a2, a3, a4), B = (b1, b2, b3, b4), C = (c1, c2, c3, c4), D = (d1, d2, d3, d4), E = (e1, e2, e3, e4), and F = (f1, f2, f3, f4), and the two check units to be determined are P = (p1, p2, p3, p4), and Q = (q1, q2, q3, q4). Similar to the two previous examples, according to the check matrix 820, the values ​​of the check units can be determined as follows: x1=a1+b1 x2=a2+b2 x3=a3+b3 x4=a4+b4 y1=c1+d1 y2=c2+d2 y3=c3+d3 y4=c4+d4 z1=e1+f1 z2=e2+f2 z3=e3+f3 z4=e4+f4 p1=x1+y1+z1 p2=x2+y2+z2 p3=x3+y3+z3 p4=x4+y4+z4 q1=x1+y2+z4+a3+d1+f3 q2=x2+y1+z3+a4+c2+e4 q3=x3+y4+z2+b1+d3+e1 q4=x4+y3+z1+b2+c4+f2.

[0134] As shown in the above equation, during the encoding process, due to the sparse and regular nature of parity check matrix 620, the intermediate calculation results x1, x2, x3, x4 and y1, y2, y3, y4 can be reused. Therefore, the actual encoding complexity is reduced to 40 XOR operations, requiring an average of 2 XOR operations per bit, reaching the theoretical limit.

[0135] In a storage system scenario, the aforementioned data units can be stored in corresponding nodes. The following uses the example of a node failure where two data units A and B reside to illustrate the decoding process using parity check matrix 820. In this case, computing device 130 can use parity check matrix 520 to calculate the values ​​of data units A and B. Corresponding to the above encoding process, the value of the data unit can be restored as follows: x1=c1+d1 x2=c2+d2 x3=c3+d3 x4=c4+d4 y1=e1+f1 y2=e2+f2 y3=e3+f3 y4=e4+f4 z1=p1+x1+y1 z2=p2+x2+y2 z3=p3+x3+y3 z4=p4+x4+y4 a3=q1+x2+y4+z1+d1+f3 a4=q2+x1+y3+z2+c2+e4 b1=q3+x4+y2+z3+d3+e1 b2=q4+x3+y1+z4+c4+f2 b3=z3+a3 b4=z4+a4 a1=z1+b1 a2=z2+b2

[0136] As shown in the above formula, during the decoding process, the intermediate calculation results x1, x2, x3, x4 and y1, y2, y3, y4 can be reused. Therefore, the actual decoding complexity becomes 40 XOR operations, requiring an average of 2 XOR operations per bit, which also reaches the theoretical limit. Moreover, this decoding process does not require a matrix inversion operation. In particular, of the 28 possible decoding scenarios using the parity check matrix in this example, at least 20 can achieve the theoretical limit of decoding complexity.

[0137] As illustrated in the examples above, the embodiments of this application provide a scheme for erasure coding using a novel pure XOR array code. This array code can be constructed by designing a sparse and regular matrix over a high-order finite field and converting it into a binary check matrix using the matrix representation of the finite field. This allows multiplication operations during the encoding and decoding process to be converted into XOR operations, reducing computational complexity.

[0138] Furthermore, the resulting binary check matrix retains its sparse and regular properties, including pairwise matching of identical or similar check sub-matrices. This allows for the reuse of intermediate calculation results during the encoding and decoding process to significantly reduce computational overhead, and in some embodiments, can reach the theoretical limit. For example, for identical check sub-matrices, reusing intermediate calculation results can completely eliminate this portion of the XOR computation overhead; for check sub-matrices that differ by one unit matrix, reusing intermediate calculation results can completely eliminate XOR operations outside the diagonal. Furthermore, the matrix inversion operation during the decoding process can be eliminated through row and column permutations, converted into a simple elimination operation, significantly reducing the complexity of the decoding steps.

[0139] It should be understood that the examples in Figures 6-8 are shown only as examples. Those skilled in the art will understand that the parameters of the code pattern construction according to the embodiments of the present application (for example, the shape of the check matrix and the selected finite field) and the calculation path of the codec can be adjusted based on the original data length, the check data length, the location of the fault data, etc. This can improve the calculation efficiency for specific application scenarios while meeting the codec requirements and hardware limitations.

[0140] Figure 9 shows a schematic block diagram of a data verification apparatus 900 according to some embodiments of the present application. Apparatus 900 may be implemented as or included in the computing device 130 of Figure 1. Apparatus 900 may include multiple modules for performing corresponding actions in method 200, for example.

[0141] As shown in the figure, apparatus 900 includes a data acquisition module 910, a matrix acquisition module 920, and a determination module 930. Data acquisition module 910 is configured to acquire a set of data units, each of which has multiple subpackets. Matrix acquisition module 920 is configured to acquire a check matrix, which is a binary matrix representation of a finite field matrix, each element of which comes from a specific finite field. Determination module 930 is configured to determine a set of check units based on the check matrix and the set of data units.

[0142] In some embodiments, the finite field matrix includes a first part, wherein the first part uses a matrix of the same order formed by elements in two extended fields as a construction unit, wherein the product of the order of the construction unit and the degree of the two extended fields is equal to the number of subpackaging of data units in a group of data units, and the number of construction units included in the row direction and the column direction of the first part is respectively equal to one of the following: the number of data units in the group and the number of check units in the group.

[0143] In some embodiments, the construction unit is a first-order matrix or a sparse matrix of elements in the two extended fields, and the construction unit in the first row of construction units in the first part is an identity matrix.

[0144] In some embodiments, the same row of construction units in the first portion includes multiple pairs of construction units, wherein each pair of construction units is equal, or each pair of construction units differs by a permutation matrix after being converted into binary matrix representations.

[0145] In some embodiments, the number of subpackets of each data unit is less than or equal to 4, and the same row of construction units in the first part includes multiple pairs of construction units, wherein the values ​​of corresponding positions of each pair of construction units are the same or sum to 1.

[0146] In some embodiments, each m-order submatrix of the finite field matrix is ​​full rank, where m is equal to the number of rows of the finite field matrix.

[0147] In some embodiments, the determination module 930 includes: an intermediate value determination module, configured to determine the intermediate values ​​of a group of check units corresponding to a pair of sub-matrices and a group of data units in the check matrix, wherein the pair of sub-matrices corresponds to a pair of construction units adjacent in the row direction of the finite field matrix; and an intermediate value multiplexing module, configured to determine the values ​​of a group of check units based on the intermediate values.

[0148] In some embodiments, the finite field is a two-dimensional field GF (2 2 ).

[0149] In some embodiments, the apparatus 900 further includes: a data unit writing module configured to write the group of data units to a group of data nodes; and a check unit writing module configured to write the group of check units to a group of check nodes.

[0150] FIG10 shows a schematic block diagram of an apparatus 1000 for data recovery according to some embodiments of the present application. Apparatus 1000 may be implemented as or included in computing device 130 of FIG1 . Apparatus 1000 may include multiple modules for performing, for example, corresponding actions in method 300. It should be understood that, in some embodiments, apparatus 1000 and apparatus 900 described above may be implemented or included in the same computing device or in different computing devices.

[0151] As shown in the figure, apparatus 1000 includes a unit acquisition module 1010, a matrix acquisition module 1020, and a determination module 1030. Unit acquisition module 1010 is configured to acquire a set of data units and a corresponding set of check units, the set of check units being determined based on the set of data units and a check matrix, the set of data units including the data units to be recovered. Matrix acquisition module 1020 is configured to acquire the check matrix, which is a binary matrix representation of a finite field matrix, each element of which comes from a specific finite field. Determination module 1030 is configured to determine the value of the data unit to be recovered based on the set of data units, the check units, and the check matrix.

[0152] In some embodiments, the finite field matrix includes a first part, the finite field matrix includes the first part, the first part uses a same-order matrix composed of elements in two extended fields as a construction unit, wherein the product of the order of the construction unit and the number of times of the two extended fields is equal to the number of subpackaging of the data units in the group of data units, and the number of construction units included in the row direction and the column direction of the first part is respectively equal to one of the following: the number of the group of data units and the number of the group of check units.

[0153] In some embodiments, the determination module includes: an intermediate value determination module, configured to determine, based on a check matrix, a calculation path for calculating the value of the data unit to be recovered, and an intermediate value generated in the calculation path; and an intermediate value multiplexing module, configured to determine the value of the data unit to be recovered based on the calculation path and the intermediate value.

[0154] Figure 11 shows a schematic block diagram of an example device 1100 that can be used to implement an embodiment of the present application. Device 1100 can be used to implement the function of the computing device 130 shown in Figure 1. As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to the computer program instructions stored in a random access memory (RAM) 1103 and / or a read-only memory (ROM) 1102 or the computer program instructions loaded into the RAM 1103 and / or ROM 1102 from a storage unit 1108. In RAM 1103 and / or ROM 1102, various programs and data required for the operation of device 1100 can also be stored. Computing unit 1101 and RAM 1103 and / or ROM 1102 are connected to each other by bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.

[0155] Various components in device 1100 are connected to I / O interface 1105, including an input unit 1106, such as a keyboard and mouse; an output unit 1107, such as various types of displays and speakers; a storage unit 1108, such as a magnetic disk and optical disk; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0156] The computing unit 1101 may be a variety of general and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a CPU, a GPU, various dedicated AI computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as method 200 and / or method 300. For example, in some embodiments, method 200 and / or method 300 may be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1100 via RAM and / or ROM and / or communication unit 1109. When the computer program is loaded into RAM and / or ROM and executed by the computing unit 1101, one or more steps of method 200 and / or method 300 described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to execute the method 200 and / or the method 300 in any other appropriate manner (eg, by means of firmware).

[0157] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a server or terminal, the process or function described in the embodiment of the present application is generated in whole or in part. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a server or terminal or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, and a tape, etc.), an optical medium (e.g., a digital video disk (DVD), etc.), or a semiconductor medium (e.g., a solid-state drive, etc.).

[0158] In addition, although adopting specific order to describe each operation, this should be understood as requiring such operation to be carried out with shown specific order or with sequential order, or requiring all illustrated operations to be carried out to obtain desired result.Under certain environment, multitasking and parallel processing may be advantageous.Similarly, although comprising some specific implementation details in the above discussion, these should not be interpreted as limiting the scope of the application.Some features described in the context of independent embodiment can also be implemented in a single implementation in combination.On the contrary, the various features described in the context of independent implementation also can be implemented in a plurality of implementations individually or in the mode of any suitable subcombination.

[0159] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for data verification, characterized in that: include: Acquire a group of data units, each data unit in the group of data units having a plurality of sub-packets; Obtain a check matrix, wherein the check matrix is ​​a binary matrix representation of a finite field matrix, and each element in the finite field matrix comes from a specific finite field; as well as A set of check elements is determined based on the check matrix and the set of data elements.

2. The method according to claim 1, characterized in that: The finite field matrix includes a first part, wherein the first part uses a matrix of the same order formed by elements in two extended fields as a construction unit, wherein The product of the order of the construction unit and the second extension number is equal to the number of subpackets of the data unit in the group of data units, and The number of construction units included in the row direction and the column direction of the first part is respectively equal to one of the following: the number of the group of data units and the number of the group of check units.

3. The method according to claim 2, characterized in that The construction unit is a first-order matrix or a sparse matrix of elements in the second extended field, and the construction unit in the first row of construction units in the first part is a unit matrix.

4. The method according to claim 3, characterized in that The same row of construction units in the first part includes a plurality of pairs of construction units, wherein each pair of construction units is equal or each pair of construction units differs by a permutation matrix after being respectively converted into binary matrix representations.

5. The method according to claim 3, characterized in that: The number of subpackets of each data unit is less than or equal to 4, and the same row of construction units in the first part includes multiple pairs of construction units, wherein the values ​​of corresponding positions of each pair of construction units are the same or the sum is 1.

6. The method according to claim 3, characterized in that Each m-order submatrix of the finite field matrix is ​​full rank, where m is equal to the number of rows of the finite field matrix.

7. The method according to claim 2, characterized in that: Determining the set of verification units includes: Determining, based on a pair of sub-matrices in the check matrix and the values ​​of the group of data units, an intermediate value of the group of check units corresponding to a pair of sub-matrices, the pair of sub-matrices corresponding to a pair of construction units adjacent in a row direction of the finite field matrix; and Based on the intermediate value, the value of the set of check cells is determined.

8. The method according to claim 1, wherein the finite field is a two-dimensional field GF (2 2 ).

9. The method according to claim 1, characterized in that: Also includes: Writing the set of data units to a set of data nodes; as well as The set of check units is written to a set of check nodes.

10. A method for data recovery, characterized in that: include: Acquire a group of data units and a corresponding group of check units, wherein the group of check units is determined based on the group of data units and a check matrix, and the group of data units includes data units to be restored; Acquire the check matrix, where the check matrix is ​​a binary matrix representation of a finite field matrix, and each element in the finite field matrix comes from a specific finite field; as well as Based on the group of data units, the group of check units and the check matrix, a value of the data unit to be restored is determined.

11. The method according to claim 10, characterized in that The finite field matrix includes a first part, wherein the first part uses a matrix of the same order formed by elements in two extended fields as a construction unit, wherein The product of the order of the construction unit and the second extension number is equal to the number of subpackets of the data unit in the group of data units, and The number of construction units included in the row direction and the column direction of the first part is respectively equal to one of the following: the number of the group of data units and the number of the group of check units.

12. The method according to claim 10, characterized in that Determining the value of the data unit to be restored includes: Based on the check matrix, determining a calculation path for calculating the value of the data unit to be restored, and an intermediate value generated in the calculation path; and Based on the calculation path and the intermediate value, a value of the data unit to be restored is determined.

13. A device for data verification, characterized in that: include: A data acquisition module is configured to acquire a group of data units, each data unit in the group of data units having a plurality of sub-packets; A matrix acquisition module is configured to acquire a check matrix, wherein the check matrix is ​​a binary matrix representation of a finite field matrix, and each element in the finite field matrix comes from a specific finite field; as well as The determination module is configured to determine a group of check units based on the check matrix and the group of data units.

14. A device for data recovery, characterized in that: include: a unit acquisition module, configured to acquire a group of data units and a corresponding group of check units, wherein the group of check units is determined based on the group of data units and the check matrix, and the group of data units includes the data units to be recovered; A matrix acquisition module is configured to acquire the check matrix, where the check matrix is ​​a binary matrix representation of a finite field matrix, and each element in the finite field matrix comes from a specific finite field; as well as The determination module is configured to determine the value of the to-be-recovered data unit based on the group of data units, the group of check units and the check matrix.

15. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein the memory stores instructions, and when the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 9 or 10 to 12.

16. A computer storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed by an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 9 or 10 to 12.

Citation Information

Patent Citations

  • Method, device, equipment and medium for data verification and data recovery

    CN120123144A

  • Data recovery method, system and device and readable storage medium

    CN111682874A

  • Distributed storage node error correction method and system

    CN115237662A

  • Encoding and decoding method and device for erasure codes and medium

    CN117149509A

  • Distributed storage system, method and apparatus

    US20200081778A1