Data coding method and related device
By using a check matrix or generator matrix represented by a polynomial ring in a distributed storage system and selecting irreducible polynomials for data encoding, the problem of high computational overhead in the prior art is solved and efficient data redundancy storage is achieved.
Patent Information
- Application Number
- CN202410294634.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-09-12
AI Technical Summary
Existing data redundancy technology is limited by code length in large distributed storage systems, making it difficult to meet large-scale encoding and storage requirements. In addition, the encoding and decoding process has high computational complexity, resulting in large computational overhead.
A check matrix or generator matrix represented by a polynomial ring is adopted. By selecting irreducible polynomials as the polynomials corresponding to the polynomial ring, data units of a finite field are transformed to perform an XOR operation, thereby reducing computational overhead.
It effectively reduces the computational overhead in the data encoding process, improves encoding and decoding efficiency, and meets the needs of large-scale encoding storage.
Smart Images

Figure CN120639102A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of storage technology, and in particular to a data encoding method and device. Background Art
[0002] Currently, distributed storage solutions are the primary method for storing massive amounts of data. This solution allows for efficient data access and computation across tens of thousands of distributed data storage nodes. Distributed storage not only offers high data scalability but also reduces overhead by managing and maintaining data across multiple distributed disk arrays.
[0003] However, in current large-scale distributed storage systems, storage node failures are common. To prevent data loss, storage systems typically use data redundancy technology to encode and store stored data, thereby restoring data protection in failed nodes.
[0004] Current data redundancy technology is limited by code length and cannot meet the needs of large-scale coding and storage. In addition, current data redundancy technology requires a large number of finite field multiplication and addition operations during the data encoding and decoding process, which increases the computational complexity of the encoding, thereby reducing data encoding efficiency and resulting in high computational overhead. Summary of the Invention
[0005] The present application provides a data encoding method and apparatus thereof, wherein a check matrix or a generator matrix is constructed using a polynomial that meets a preset condition, thereby reducing computational overhead during encoding.
[0006] In a first aspect, the present application provides a data encoding method, comprising: obtaining a target matrix, the target matrix being a check matrix or a generator matrix, the target matrix being a finite field 2 a A binary matrix is represented on a polynomial ring in which the data units are represented, and the polynomial corresponding to the polynomial ring is an irreducible polynomial with a highest degree not greater than a; according to the target matrix, the data block is encoded to obtain an encoding result; wherein the encoding result is used as a redundant data block of the data block; or, the encoding result is a recovery result of the data block of the node.
[0007] The test matrix or generator matrix can be represented by a polynomial ring, wherein the test matrix or generator matrix can include multiple data units, the data units can be elements in the matrix, the data units can be finite fields 2 a The data unit in the polynomial ring can be a matrix (that is, the target matrix in the embodiment of the present application), and the target matrix is the finite field 2 aThe data unit (belonging to the check matrix or generator matrix) is represented by a binary matrix on the polynomial ring (a matrix whose elements are represented by binary is a binary matrix). In the existing technology, for the finite field 2 a The polynomial ring corresponding to the data unit converted into is x a +x a-1 +…+x 2 +x+1=0 (the selected polynomial will affect the size of the polynomial ring), the highest degree of the polynomial is a, the coefficient of each term is 1, and the degree of each term is reduced from a until the degree becomes 0. In this polynomial (x a +x a-1 +…+x 2 +x+1=0) is a reducible polynomial, it is impossible to convert the finite field 2 a The data unit is converted to x a +x a-1 +…+x 2 +x+1=0. However, for a polynomial with the highest degree a (each term has a coefficient of 1) (excluding x a +x a-1 +…+x 2 +x+1=0) often includes one or more irreducible polynomials. In the embodiment of the present application, an irreducible polynomial can be selected as the polynomial corresponding to the polynomial ring, so that the finite field 2 a The data unit can be converted into a corresponding polynomial ring for representation, and then, an XOR operation equivalent to a multiplication operation can be performed on the data represented by the polynomial ring, thereby reducing the computational overhead.
[0008] In one possible implementation, for the finite field 2 a For a data unit, a polynomial can be selected from (one or more) irreducible polynomials with the highest degree not greater than a as the polynomial corresponding to the polynomial ring. For example, the polynomial with the smallest order among (one or more) irreducible polynomials with the highest degree not greater than a can be selected. Here, the order is defined as: the polynomial can divide one or more expressions of the form xb+1. The order of the polynomial is the number of the one or more expressions of the form x b The smallest value of b in the expression of +1.
[0009] For example, the polynomial that divides x 3 +1 and x 4 +1, then the order of the polynomial is 3.
[0010] In one possible implementation, the target matrix includes a submatrix for representing a first data unit, the data block includes a second data unit, the second data unit is represented by a first vector, and the first vector is a binary vector; encoding the data block according to the target matrix includes: performing an XOR operation on the submatrix and the first vector to obtain a second vector, and the second vector is a binary vector; dividing a polynomial with binary elements in the second vector as coefficients by a polynomial corresponding to the polynomial ring and taking the remainder to obtain an operation result equivalent to the product result of the first data unit and the second data unit.
[0011] In one possible implementation, a check matrix or a generator matrix may include multiple data units (i.e., elements in a matrix), wherein a target matrix is obtained by representing each data unit through a polynomial ring. Taking one data unit (a first data unit) among the multiple data units as an example, the target matrix may include a submatrix for representing the first data unit. When determining the submatrix of the first data unit, the binary vector corresponding to the first data unit may be used as a column of the submatrix, and then the column may be cyclically shifted to obtain other columns, thereby obtaining a submatrix of the first data unit, wherein the number of rows of the submatrix is the order of the polynomial corresponding to the polynomial ring.
[0012] In one possible implementation, when determining the submatrix of the first data unit, the binary vector corresponding to the first data unit can be used as the column of the submatrix (for example, each element in the binary representation is filled into the column of the submatrix from top to bottom, or each element after filling is transformed, that is, the "0" element becomes "1" and the "1" becomes "0"), and a certain transformation can be performed on the column (the transformation is a homomorphic transformation, that is, it will not affect the results of subsequent calculations). The transformation can be to change the "0" element of the column to "1" or to change "1" to "0".
[0013] In one possible implementation, the submatrix includes a target column, where the target column is obtained by transforming some binary elements after padding the binary representation of the first data unit. The target column includes X rows, and the transformation includes: when an element of the target row in M rows following the X rows is 1, elements of rows before the target row in the X rows are transformed accordingly based on the polynomial.
[0014] The target column includes X rows, wherein it can be determined whether each element in the last M rows of the X rows is 1 or 0, wherein when the element is 1, the elements of the previous rows are transformed accordingly based on the polynomial, for example, the polynomial is x 4 +x+1=0, before the transformation, x 4Set the elements of the row corresponding to x to 1, and set the elements of the row corresponding to 1 to 1. When the first row of the next M rows is 1, the polynomial is x. 5 +x 2 +x=0, and further, x 2 The corresponding row will be transformed, and the row corresponding to x will be transformed.
[0015] Here, X can be the value of the order b.
[0016] In a possible implementation, M is the difference between X and a.
[0017] In one possible implementation, when the number of "1" elements included in the polynomial ring is smaller and the number of "0" elements is larger, the computational overhead of the subsequent XOR operation will be lower. Therefore, when performing a homomorphic transformation on the submatrix, the goal can be set as: to pass the homomorphic transformation and minimize the number of 1 elements included in the submatrix.
[0018] Specifically, the target column is obtained by performing the transformation on some binary elements after filling the binary representation of the first data unit so as to minimize the number of elements with 1 included in the submatrix used to represent the first data unit.
[0019] In one possible implementation, after determining the submatrix of the elements in the finite field, it is necessary to select the elements used to constitute the target matrix and the positions of the submatrices corresponding to the elements in the target matrix from the elements in the finite field. For the target matrix, including fewer "1" elements will reduce the computing power overhead of the subsequent XOR operation, and the selection of the elements in the finite field and the positions of the submatrices corresponding to the elements in the target matrix needs to meet the following requirements: the optimal compromise between storage overhead and fault tolerance can be achieved (that is, the maximum distance separable (MDS) property is satisfied). Specifically, in one possible implementation, the target matrix includes multiple submatrices for representing multiple data units, and the positions of the multiple data units and the corresponding submatrices in the target matrix are based on the following constraints from the finite field 2 a Selected in: Minimize the number of elements of 1 included in the target matrix while ensuring that the target matrix has the maximum distance separable mds property.
[0020] In a second aspect, the present application provides a data encoding device, comprising:
[0021] The acquisition module is used to obtain a target matrix, wherein the target matrix is a check matrix or a generator matrix, and the target matrix is obtained by converting a finite field 2 aA binary matrix represented by the data units on a polynomial ring, wherein the polynomial corresponding to the polynomial ring is an irreducible polynomial with a highest degree not greater than a;
[0022] The encoding module is used to encode the data block according to the target matrix to obtain an encoding result; wherein,
[0023] The encoding result is used as a redundant data block of the data block; or, the encoding result is a recovery result of the data block of the node.
[0024] In one possible implementation, x a +x a-1 +…+x 2 +x+1=0 is a reducible polynomial, and the polynomial corresponding to the polynomial ring is an irreducible polynomial whose highest degree is not greater than a.
[0025] In a possible implementation, the data block and the encoding result are stored on different nodes.
[0026] In a possible implementation, the polynomial corresponding to the polynomial ring is the polynomial with the smallest order among the irreducible polynomials whose highest degree is not greater than a, wherein the polynomial can divide one or more polynomials of the form x b +1, the order of the polynomial is the one or more forms of x b The smallest value of b in the expression of +1.
[0027] In a possible implementation, the target matrix includes a submatrix for representing a first data unit, the data block includes a second data unit, the second data unit is represented by a first vector, and the first vector is a binary vector;
[0028] The processing module is specifically used to:
[0029] Performing an XOR operation on the submatrix and the first vector to obtain a second vector, where the second vector is a binary vector;
[0030] A polynomial having binary elements in the second vector as coefficients is divided by a polynomial corresponding to the polynomial ring and a remainder is taken to obtain an operation result equivalent to a product result of the first data unit and the second data unit.
[0031] In a possible implementation, the target matrix includes a submatrix for representing the first data unit, and the number of rows of the submatrix is the order of the polynomial corresponding to the polynomial ring.
[0032] In one possible implementation, the target column is obtained by filling the binary representation of the first data unit and then performing the transformation on some binary elements so that the number of elements of 1 included in the submatrix used to represent the first data unit is minimized, and the transformation includes changing 0 to 1, or changing 1 to 0.
[0033] In one possible implementation, the submatrix includes a target column, where the target column is obtained by transforming some binary elements after padding the binary representation of the first data unit. The target column includes X rows, and the transformation includes: when an element of the target row in M rows following the X rows is 1, elements of rows before the target row in the X rows are transformed accordingly based on the polynomial.
[0034] In a possible implementation, M is the difference between X and a.
[0035] In a possible implementation, the target column is obtained by filling the binary representation of the first data unit and performing the transformation on some binary elements so as to minimize the number of elements with 1 included in the submatrix used to represent the first data unit.
[0036] In a possible implementation, the target matrix includes a plurality of sub-matrices for representing a plurality of data units, and the positions of the plurality of data units and the corresponding sub-matrices in the target matrix are based on the following constraints from the finite field 2 a Selected from:
[0037] Under the condition that the target matrix has the maximum distance separable (mds) property, the number of elements of 1 included in the target matrix is minimized.
[0038] In a third aspect, the present application provides a computing device comprising a memory and a processor, wherein the processor executes computer instructions stored in the memory, so that the computing device executes the method described in any possible implementation manner in the first aspect.
[0039] In a fourth aspect, the present application provides a storage system comprising any possible data processing device as described in the second aspect.
[0040] In a fifth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computing device, enables the computing device to execute the method described in any possible implementation of the first aspect.
[0041] In a sixth aspect, the present application provides a computer program product, which, when the program code contained in the computer program product is executed by a computer device, implements the method described in any possible implementation method of the first aspect of the present application.
[0042] Since the devices provided in this application can be used to execute the steps of any possible implementation method in the first aspect, the technical effects that can be obtained by the devices in this application can refer to the technical effects obtained by the aforementioned methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a schematic diagram of an exemplary application architecture;
[0044] Figure 2 This is a schematic diagram of an exemplary application architecture;
[0045] Figure 3 is a schematic diagram illustrating an exemplary data repair process;
[0046] Figure 4 is a schematic diagram of an exemplary data encoding method;
[0047] Figures 5 to 12 is one of the schematic diagrams of an exemplary check matrix;
[0048] Figure 13 A schematic diagram of the structure of a device provided in an embodiment of the present application;
[0049] Figure 14 A schematic diagram of the structure of a chip provided in an embodiment of the present application;
[0050] Figure 15 A schematic structural diagram of a data encoding device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0052] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0053] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first target object" and "second target object" are used to distinguish different objects, rather than to describe a specific order of objects.
[0054] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0055] In the description of the embodiments of this application, unless otherwise specified, "multiple" means two or more. For example, "multiple processing units" means two or more processing units; "multiple systems" means two or more systems.
[0056] In the era of big data, data-based applications rely on massive amounts of data to acquire knowledge and effectively perform operations such as data classification and retrieval. For example, artificial intelligence systems, such as deep learning, rely on massive amounts of data for parameter updates and feature learning. To effectively store and maintain this massive amount of data, traditional centralized storage solutions suffer from limited scalability, limited ability to locate and retrieve key data, and high costs associated with maintaining centralized data management mechanisms within data centers.
[0057] Therefore, distributed storage solutions are currently the primary approach for storing massive amounts of data. This solution allows for efficient data access and computation across tens of thousands of distributed data storage nodes. Distributed storage not only offers high data scalability but also reduces overhead by managing and maintaining data across multiple distributed disk arrays.
[0058] However, in current large-scale distributed storage systems, storage node failures are common. To prevent data loss, storage systems typically employ data redundancy technologies to protect stored data. Currently, data redundancy technologies primarily include replica storage and erasure coding.
[0059] Replica storage technology primarily replicates original data into multiple copies and stores these copies on different nodes in the system. For example, triple-copy storage technology replicates original data three times and stores each copy on three nodes in the storage system. One node stores the original data, while two nodes store the replicas. This ensures data validity as long as all nodes are not damaged. However, triple-copy storage technology only offers a storage efficiency of one-third, resulting in low data storage efficiency. In replica storage technology, if data is stored in n copies, the storage efficiency is only one-nth, resulting in high system storage overhead and inability to meet the storage requirements of large-scale storage solutions.
[0060] Erasure code storage technology can divide the original data into k equal parts of original data blocks, and then obtain m redundant data blocks through [n, k] erasure code encoding. Finally, n (here n = m + k) parts of data (including original data blocks and redundant data blocks) are stored on n storage nodes of the system respectively.
[0061] Before introducing the technical solution of this application, the following technical terms are first explained:
[0062] The dimension k is used to represent the number of original data blocks after the original data is divided into blocks when encoding the original data.
[0063] Redundancy is used to indicate the number of redundant data blocks generated when encoding the original data, or the number of check nodes used to store redundant data blocks.
[0064] The code length n represents the number of data blocks in the encoded data generated after encoding the original data (e.g., the number of original data blocks (i.e., the dimension k mentioned above) + the number of redundant data blocks (i.e., the redundancy mentioned above)). Furthermore, the code length is equal to the number of nodes in which the encoded data is stored. The number of redundant data blocks (or the number of check nodes, or the redundancy) is (nk).
[0065] The array dimension m is used to represent the number of data fragments (or data subpackets) contained in each of the k original data blocks.
[0066] The locality r is used to represent the number of nodes required to repair the original data when there are damaged storage nodes in the distributed storage system used to store the encoded data.
[0067] Erasure codes (EC) are a key technology for achieving data reliability in data storage. Erasure codes encode valid data blocks to generate redundant check blocks. If any data block is damaged, the check block and other data blocks can be read to reconstruct (or decode) the data block, thereby improving the reliability of the data block. This application provides a data processing method and device based on erasure codes, which can implement encoding or decoding of erasure codes based on a check matrix (or generator matrix) to reliably store data.
[0068] This application can be applied to but not limited to Figure 1 The application scenario shown. Figure 1The application scenario shown includes an application server and a storage node, and the application server and the storage node can be connected via a communication network. A storage controller and one or more memories are provided in the storage node. This application does not limit the type of memory. For example, the memory can be a memory in a storage server or a storage array or a hard disk or a storage controller. Among them, the hard disk includes a solid state disk (SSD) or a mechanical hard disk (HDD) or a hybrid hard disk.
[0069] The application scenarios to which this application is applicable may be various, and this application does not specifically limit this. For example, the application scenarios to which this application is applicable may include Figure 1 More or fewer application servers or storage nodes. Storage nodes may include Figure 1 More or less storage, the storage node can be a storage node in a centralized storage system or a distributed storage system.
[0070] Below Figure 1 Taking the application scenario shown as an example, possible examples of the data processing method of this application are introduced.
[0071] refer to Figure 2 , the storage controller can receive the data to be written sent by the application server and encode it based on the erasure code. Take the K+M ratio of 6+4 as an example (that is, the code length n is 10 and the redundancy is 4), where "6+4" here can be understood as there are 6 copies of the data to be written (excluding the check blocks, that is, redundant data), and the 6 copies of data need to be stored on 6 memories (or, can be described as storage nodes). Based on the 6 copies of the data to be written, 4 check blocks can be obtained. The storage controller can divide the data to be written into 6 data blocks (D1, D2, ..., D6), encode the 6 data blocks according to the check matrix, and obtain 4 check blocks (P1, P2, P3, P4). Afterwards, the storage controller can store the 6 data blocks and 4 check blocks in 10 memories respectively.
[0072] This application can be applied to distributed storage systems to perform erasure coding on storage system data to improve storage system reliability. The current mainstream approach is to store large files in blocks and manage them in a disk distributed storage manner. Figure 2This paper presents a distributed storage system architecture. To mitigate disk deletion errors, during data downloading from the storage system, the original file is partitioned into raw data blocks (D1-D6). Erasure coding is then applied to these blocks to generate redundant data blocks (P1-P4). Finally, the original data blocks (D1-D6) and redundant data blocks (P1-P4) are written to different storage nodes for downloading. In the event of disk corruption or node failure, valid raw data or redundant checksum blocks are read from intact disks or valid nodes, decoded using erasure coding, and the lost data blocks are recovered and written to repair disks or repair nodes, thus maintaining data reliability in the distributed storage system.
[0073] Reference Figure 3 , Figure 3 This is a schematic diagram of an application architecture of an embodiment of the present application, and Figure 2 The difference is, Figure 3 Taking the K+M ratio of 4+2 as an example, the present application can be applied to a distributed storage system to perform erasure coding on the storage system data to provide storage system reliability. In order to deal with disk deletion errors, in the process of downloading data from the storage system, the original file is divided into data blocks to generate original data blocks (F1-F4), and the erasure coding scheme is applied to encode the original data blocks to generate redundant data blocks (P1, P2). Finally, the original data blocks (F1-F4) and the redundant data blocks (P1, P2) are written to different storage nodes for downloading. When a disk is damaged or a node fails, the valid original data or redundant checksum is read from the undamaged disk or valid node to perform erasure coding decoding, and the lost data blocks (F1', F2') are recovered and written to the repair disk or repair node, thereby maintaining the data reliability of the distributed storage system.
[0074] When calculating redundant data blocks based on original data blocks (that is, data to be written), or recovering data based on redundant data blocks (or original data blocks), multiplication operations between data units are required (specifically: multiplication operations between check matrices or generating matrices and original data blocks or redundant data blocks). In order to reduce the computational overhead of operations and improve the performance of encoding and decoding, multiplication operations are often converted into XOR operations. For example, multiplication operations can be converted into XOR operations through matrix representation of finite fields. This conversion is a dense transformation method, and the number of XOR operations after conversion is large, and the encoding and decoding performance cannot be significantly improved. In order to improve the encoding and decoding performance and reduce the number of XOR calculations, it is possible to convert the operation to a ring (i.e., a polynomial ring). In this case, it is necessary to transform the finite field elements of the check matrix or the generator matrix (the finite field elements in the embodiment of the present application can also be referred to as data units) into a 0-1 matrix on the ring (for example, into a column on the matrix). The other columns on the matrix can be obtained by circular shift, and then the ring matrix used to represent the check matrix or the generator matrix is obtained. The calculation result is then obtained by performing an XOR operation on the ring matrix and the binary vector used to represent the original data block or the redundant data block.
[0075] In the existing technology, the ring transformation technology is to transform the finite field GF(2 m ) is converted to a polynomial x m +x m-1 +…+x+1 polynomial ring, however, the transformation that transforms the elements of a finite field into a ring matrix does not exist for any finite field. For example, for the finite field GF(2 m ), when the polynomial x m +x m-1 +…+x+1 is irreducible (when a polynomial can be decomposed into the product of multiple polynomials, it can be considered that the polynomial is reducible), GF(2 m ) on the ring transformation (convert to polynomial x m +x m-1 +…+x+1) is feasible, and when the polynomial x m +x m-1 When +…+x+1 is reducible, GF(2 m ) on the ring transformation (convert to polynomial x m +x m-1 +…+x+1) is infeasible, for example GF(256).
[0076] Figure 4 The encoding process of the storage controller based on the erasure code is shown as an example. Figure 4 The encoding process shown may include steps S401 to S402.
[0077] S401, obtaining a target matrix, wherein the target matrix is a check matrix or a generator matrix, and the target matrix is a matrix that converts a finite field 2 a A binary matrix is represented by the data units on a polynomial ring, and the polynomial corresponding to the polynomial ring is an irreducible polynomial with a highest degree not greater than a.
[0078] In one possible implementation, when encoding the data to be written (which may be referred to as the original data block in the embodiment of the present application), the data to be written can be obtained and encoded to obtain multiple redundant data blocks, wherein the data to be written can be divided into multiple data blocks, and the embodiment of the present application does not limit the division method of the data to be written.
[0079] In one possible implementation, when recovering data on a storage node, data on other nodes (nodes that have not lost data) can be obtained (either original data blocks or redundant data blocks), and the obtained data can be encoded to obtain the data on the storage node where the data needs to be recovered.
[0080] Among them, the above-mentioned encoded objects (original data blocks or redundant data blocks) are collectively referred to as data blocks in the embodiments of the present application for the convenience of description.
[0081] In a possible implementation, the data block may be encoded according to a check matrix or a generator matrix.
[0082] For example, a data block (original data block) can be encoded according to a check matrix or a generator matrix to obtain a redundant data block. Specifically, for an original data block with a code length of n and a dimension of k, during the data encoding process, the redundant data block can be calculated based on the following relationship:
[0083] H×[m, p] T =H×[m1,m2,...,m k ,p1,...,p n-k ] T =0;
[0084] where m i (1≤i≤k) is the data unit in the original data block, m i =[m i,1 , m i,2 ,...,m i,α ] contains α data bits or data symbols; p i (1≤i≤nk) is the redundant data block generated by encoding, p i =[p i,1 , p i,2 ,...,p i,α ] contains α check bits or check symbols.
[0085] The above encoding process includes multiplication operations between data units (for example, multiplication operations between data units of the original data block and data units of the matrix (check matrix or generator matrix).
[0086] The test matrix or generator matrix can be represented by a polynomial ring, wherein the test matrix or generator matrix can include multiple data units, the data units can be elements in the matrix, the data units can be finite fields 2 a The data unit in the polynomial ring can be a matrix (that is, the target matrix in the embodiment of the present application), and the target matrix is the finite field 2 a The data unit (belonging to the check matrix or generator matrix) is represented by a binary matrix on the polynomial ring (a matrix whose elements are represented by binary is a binary matrix). In the existing technology, for the finite field 2 a The polynomial ring corresponding to the data unit converted into is x a +x a-1 +...+x 2 +x+1=0 (the selected polynomial will affect the size of the polynomial ring), the highest degree of the polynomial is a, the coefficient of each term is 1, and the degree of each term is reduced from a until the degree becomes 0. In this polynomial (x a +x a-1 +...+x 2 +x+1=0) is a reducible polynomial, it is impossible to convert the finite field 2 a The data unit is converted to x a +x a-1 +...+x 2 +x+1=0. However, for a polynomial with the highest degree a (each term has a coefficient of 1) (excluding x a +x a-1 +...+x 2 +x+1=0) often includes one or more irreducible polynomials. In the embodiment of the present application, an irreducible polynomial can be selected as the polynomial corresponding to the polynomial ring, so that the finite field 2 a The data unit can be converted into a corresponding polynomial ring for representation, and then, an XOR operation equivalent to a multiplication operation can be performed on the data represented by the polynomial ring, thereby reducing the computational overhead.
[0087] Here, taking the check matrix H as an example (the generator matrix and the check matrix are similar, and the description of the check matrix can be referred to), an illustration of the form of the check matrix H is as follows:
[0088]
[0089] Among them, h i,j∈GF(2 α )=GF(2)[x] / (p(x)), which is a polynomial of degree less than α on GF(2), where p(x) is an irreducible polynomial of degree α on GF(2). A suitable irreducible polynomial p(x) can be selected to obtain the polynomial ring F2[x] / (x ord(p(x)) +1).
[0090] In one possible implementation, for the finite field 2 a For a data unit, a polynomial can be selected from (one or more) irreducible polynomials whose highest degree is not greater than the a as the polynomial corresponding to the polynomial ring. For example, the polynomial with the smallest order among (one or more) irreducible polynomials whose highest degree is not greater than the a can be selected, where the order is defined as: the polynomial can divide one or more expressions of the form xb+1, and the order of the polynomial is the smallest value of b among the one or more expressions of the form xb+1.
[0091] For example, the polynomial that divides x 3 +1 and x 4 +1, the order of the polynomial is 3. In the embodiment of the present application, the order of the polynomial can be described as ord(p(x)).
[0092] The choice of polynomial affects the structure of the polynomial ring, the determination process, and the inverse transformation process (that is, from the polynomial ring back to the finite field). The following are introduced respectively:
[0093] 1. The choice of polynomial affects the size of the polynomial ring.
[0094] In one possible implementation, a check matrix or a generator matrix may include multiple data units (i.e., elements in a matrix), wherein a target matrix is obtained by representing each data unit through a polynomial ring. Taking one data unit (a first data unit) among the multiple data units as an example, the target matrix may include a submatrix for representing the first data unit. When determining the submatrix of the first data unit, the binary vector corresponding to the first data unit may be used as a column of the submatrix, and then the column may be cyclically shifted to obtain other columns, thereby obtaining the submatrix of the first data unit, wherein the number of rows of the submatrix is the order of the polynomial corresponding to the polynomial ring.
[0095] For example, if the order of the polynomial ring is 7, then the number of rows of the submatrix is 7.
[0096] 2. The choice of polynomial will affect the homomorphic transformation process of the polynomial ring.
[0097] Still taking the first data unit as an example, in a possible implementation, when determining the submatrix of the first data unit, the binary vector corresponding to the first data unit can be used as the column of the submatrix (for example, each element in the binary representation is filled into the column of the submatrix from top to bottom, or each element after filling is transformed, that is, the "0" element becomes "1", and the "1" becomes "0"), and for the column, a certain transformation can be performed (the transformation is a homomorphic transformation, that is, it will not affect the results of subsequent calculations), and the transformation can be to change the "0" element of the column to "1", or to change "1" to "0".
[0098] In one possible implementation, the submatrix includes a target column, the target column is obtained by transforming some binary elements after filling the binary representation of the first data unit, and the some elements are selected from the last M rows of the target column. The target column includes X rows, wherein it can be determined whether each element in the last M rows of the X rows is 1 or 0, wherein when the element is 1, the elements of the previous rows are transformed accordingly based on the polynomial, for example, the polynomial is x 4 +x+1=0, before the transformation, x 4 Set the elements of the row corresponding to x to 1, and set the elements of the row corresponding to 1 to 1. When the first row of the next M rows is 1, the polynomial is x. 5 +x 2 +x=0, and further, x 2 The corresponding row will be transformed, and the row corresponding to x will be transformed.
[0099] In one possible implementation, when the number of "1" elements included in the polynomial ring is smaller and the number of "0" elements is larger, the computational overhead of the subsequent XOR operation will be lower. Therefore, when performing a homomorphic transformation on the submatrix, the goal can be set as: to pass the homomorphic transformation and minimize the number of 1 elements included in the submatrix.
[0100] Specifically, the target column is obtained by filling the binary representation of the first data unit and then performing the transformation on some binary elements so as to minimize the number of elements with 1 included in the submatrix used to represent the first data unit.
[0101] It should be understood that the submatrix of each data unit included in the target matrix can be obtained by performing homomorphic transformation with the above-mentioned objectives as the optimization direction, so that the target matrix as a whole can include fewer "1" elements and more "0" elements.
[0102] For example, for each element in a finite field, the optimal homomorphic transformation for each element can be determined:
[0103]
[0104] Subject to: l0, l1,..., l ord(p(x))-1-α ∈{0, 1};
[0105] in, It represents the image of the finite field element a(x), that is, the polynomial ring of the finite field element a(x). wt() represents the weight of the image, that is, the number of 1 elements included in the polynomial ring. Min represents minimization. Subject to: l0, l1, ..., l ord(p(x))-1-α ∈{0, 1} means that the first M(ord(p(x))-1-α) rows on the column can be transformed.
[0106] Through the above method, the embodiment of the present application corresponds all elements to a simplest homomorphic transformation (finite field to polynomial ring), and the weight of the transformed image is the smallest.
[0107] In addition, after determining the submatrix of the elements in the finite field, it is necessary to select the elements used to constitute the target matrix and the positions of the submatrices corresponding to the elements in the target matrix from the elements in the finite field. For the target matrix, including fewer "1" elements will reduce the computing power overhead of the subsequent XOR operation, and the selection of the elements in the finite field and the positions of the submatrices corresponding to the elements in the target matrix must meet the following requirements: the optimal compromise between storage overhead and fault tolerance can be achieved (i.e., the maximum distance separable (MDS) property can be satisfied). Specifically, in one possible implementation, the target matrix includes multiple submatrices for representing multiple data units, and the positions of the multiple data units and the corresponding submatrices in the target matrix are selected from the finite field 2a based on the following constraints: while ensuring that the target matrix has the maximum distance separable (MDS) property, the number of 1 elements included in the target matrix is minimized.
[0108] For example, taking the check matrix as an example, for the matrix where h′ i,j (x)∈GF(2)[x] / (x ord(p(x)) +1);
[0109] satisfy:
[0110]
[0111] The check matrix H can be constructed:
[0112]
[0113] Among them, mod means modulo.
[0114] In this way, the submatrix corresponding to the elements of the finite field is used to fill in the generated matrix or the check matrix. The matrix satisfies the MDS property and minimizes the weight corresponding to the matrix, thereby minimizing the encoding and decoding complexity and reducing computing power overhead.
[0115] S402. Encode the data block according to the target matrix to obtain an encoding result; wherein the encoding result is used as a redundant data block of the data block; or the encoding result is a recovery result of the data block of the node.
[0116] Among them, when encoding the data block according to the target matrix, multiplication operations between data units can be performed. Since the target matrix is represented by a polynomial ring, the multiplication operations between data units can be achieved through exclusive OR operations between the polynomial ring and the binary vector, and the inverse transformation of the ring.
[0117] In one possible implementation, the target matrix includes a submatrix for representing a first data unit, the data block includes a second data unit, the second data unit is represented by a first vector, the first vector is a binary vector, the submatrix and the first vector can be XORed to obtain a second vector, the second vector is a binary vector, and then the polynomial with the binary elements in the second vector as coefficients and the polynomial corresponding to the polynomial ring can be divided and the remainder is taken to obtain an operation result equivalent to the product result of the first data unit and the second data unit. The above-mentioned division and remainder process is the process of ring inverse transformation, that is, converting the data represented by the ring into a finite field.
[0118] For example, the second vector obtained by performing an XOR operation on the sub-matrix and the first vector is [00000100], and the polynomial with the binary elements in the second vector as coefficients is x 6 =0, the polynomial ring corresponds to the polynomial x 4 +x+1=0, the result obtained by dividing the polynomial with the binary elements in the second vector as coefficients by the polynomial corresponding to the polynomial ring and taking the remainder is
[0110] .
[0119] The embodiment of the present application proposes a high-performance array code encoding and decoding technology based on homomorphic computing. Considering the construction of the array code (i.e., the design of the check matrix), a 0-1 matrix is obtained through homomorphic transformation, and the multiplication of the finite field is converted into an XOR operation on a polynomial ring. Finally, the inverse transformation of the ring is used to obtain the calculation result on the finite field, thereby reducing the encoding and decoding complexity and reaching the theoretical limit.
[0120] Here are a few specific examples:
[0121] Take a binary code with dimension k=6 and code length 8 as an example.
[0122] For finite fields 2 4 (i.e., the elements of GF(16)), the polynomial chosen is p(x)=x 4 +x+1, first, a homomorphic transformation is performed on each finite field element to minimize the weight of each image. Some results are shown in Table 1:
[0123] Table 1
[0124]
[0125] Therefore, the check matrix constructed on the finite field is as follows:
[0126] 1 1 1 1 1 1 1 1 2 4 8 3 6 1
[0127] The matrix above can be converted into a 0-1 matrix through the matrix representation of the finite field, so that the multiplication operation in the encoding and decoding process is converted into an XOR operation, and the check matrix becomes Figure 5 The matrix shown.
[0128] However, the matrix here is relatively dense and the computational complexity is still very high. In order to simplify the operation, the above matrix operation process can be divided into two steps. The first step is the check matrix operation on the ring, and the second step is to convert the ring operation result into the inverse transformation matrix of the finite field operation result.
[0129] For the calculation of the check matrix on the polynomial ring, the check matrix can be calculated as Figure 6 shown.
[0130] Therefore, this step of encoding calculation only requires (48-4-9)÷4=8.75 XOR times.
[0131] When performing the ring inverse transformation, the inverse transformation matrix is as follows Figure 7 shown.
[0132] Because the first redundancy is an identity transformation, no inverse operation is required. Therefore, this step requires only (2-1) × (15-4) ÷ 4 = 2.75 XOR operations. Ultimately, the encoding process requires only 11.5 XOR operations, reaching the theoretical limit. Furthermore, the encoding process is incremental, meaning that each block is read and calculated.
[0133] Take a binary code with dimension k=15 and code length 17 as an example.
[0134] For finite fields 2 4 (i.e., the elements of GF(16)), the polynomial chosen is p(x)=x 4+x+1, first, a homomorphic transformation is performed on each finite field element to minimize the weight of each image. Some results are shown in Table 2:
[0135] Table 2
[0136]
[0137] Therefore, the check matrix constructed on the finite field is as follows:
[0138] 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 2 4 8 3 6 12 11 5 10 7 14 15 13 9 1
[0139] The matrix above can be converted into a 0-1 matrix through the matrix representation of the finite field, so that the multiplication operation in the encoding and decoding process is converted into an XOR operation, and the check matrix becomes Figure 8 The matrix shown.
[0140] However, the matrix here is relatively dense and the computational complexity is still very high. In order to simplify the operation, the above matrix operation process can be divided into two steps. The first step is the check matrix operation on the ring, and the second step is to convert the ring operation result into the inverse transformation matrix of the finite field operation result.
[0141] For the calculation of the check matrix on the polynomial ring, the check matrix can be calculated as Figure 9 shown.
[0142] Therefore, this step of encoding calculation only requires (120-4-15)÷4=25.25 XOR times.
[0143] When performing the ring inverse transformation, the inverse transformation matrix is as follows Figure 10 shown.
[0144] Because the first redundancy is an identity transformation, no inverse operation is required. Therefore, this step requires only (2-1) × (15-4) ÷ 4 = 2.75 XOR operations. Ultimately, the encoding process requires only 11.5 XOR operations, reaching the theoretical limit. Furthermore, the encoding process is incremental, meaning that each block is read and calculated.
[0145] Take a binary code with dimension k=20 and code length 23 as an example.
[0146] For the elements of the finite field 28 (i.e. GF(256)), the polynomial chosen is p(x) = x 8 +x 5 +x 4 +x 3 +1, first, a homomorphic transformation is performed on each finite field element to minimize the weight of each image. Some results are shown in Table 3:
[0147] Table 3
[0148]
[0149] Therefore, the check matrix constructed on the finite field is as follows:
[0150] 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 2 4 8 16 32 64 128 57 114 228 241 219 143 39 78 156 3 6 12 1 1 4 16 64 57 228 219 39 156 2 8 32 128 114 241 143 78 5 20 80 1
[0151] The above matrix can be converted into a 0-1 matrix through the matrix representation of the finite field, so that the multiplication operation in the encoding and decoding process is converted into an XOR operation. However, the matrix is relatively dense and the computational complexity is still very high. In order to simplify the operation, the above matrix operation process can be divided into two steps. The first step is the check matrix operation on the ring, and the second step is to convert the ring operation result into the inverse transformation matrix of the finite field operation result.
[0152] For the calculation of the check matrix on the polynomial ring, the check matrix can be calculated as Figure 11 shown.
[0153] Therefore, this step of encoding calculation only requires (528-8-17-17)÷8=60.75 XORs.
[0154] When performing the ring inverse transformation, the inverse transformation matrix is as follows Figure 12 shown.
[0155] Because the first redundancy is an identity transformation, no inverse operation is required. Therefore, this step requires only (3-1) × (48-8) ÷ 8 = 10 XOR operations. Ultimately, the encoding process requires only 70.75 XOR operations, approaching the theoretical limit. Furthermore, the encoding process is incremental, meaning that each block is read and calculated.
[0156] In one possible implementation, the present application provides a data encoding device. The data encoding device includes one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from a memory and send the signals to the processors, the signals including computer instructions stored in the memory; when the processors execute the computer instructions, the processors can implement the data encoding method described in any of the above implementations.
[0157] The effects of the code generating device of this embodiment are similar to the effects of the data encoding methods of the above embodiments, and will not be described in detail here.
[0158] The following describes a device provided by an embodiment of the present application. Figure 13 As shown:
[0159] Figure 13 This is a structural diagram of a data encoding device provided in an embodiment of the present application. Figure 13 As shown, the device 500 may include: a processor 501, a transceiver 505, and optionally a memory 502.
[0160] The transceiver 505 may be referred to as a transceiver unit, a transceiver, or a transceiver circuit, etc., and is configured to implement transceiver functions. The transceiver 505 may include a receiver and a transmitter. The receiver may be referred to as a receiver or a receiving circuit, etc., and is configured to implement a receiving function; the transmitter may be referred to as a transmitter or a transmitting circuit, etc., and is configured to implement a transmitting function.
[0161] The memory 502 may store a computer program or software code or instruction 504, which may also be referred to as firmware. The processor 501 may control the MAC layer and the PHY layer by running the computer program or software code or instruction 503 therein, or by calling the computer program or software code or instruction 504 stored in the memory 502, to implement the data encoding method, encoding method, or decoding method provided in each embodiment of the present application. The processor 501 may be a central processing unit (CPU), and the memory 502 may be, for example, a read-only memory (ROM) or a random access memory (RAM).
[0162] The processor 501 and transceiver 505 described in this application can be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit RFIC, a mixed signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc.
[0163] The above-mentioned device 500 may further include an antenna 506. The modules included in the data encoding device 500 are only for illustration and are not limited in this application.
[0164] The structure of the data encoding device, encoding device, and decoding device may not be affected by Figure 9 The data encoding device, encoding device, and decoding device may be independent devices or may be part of a larger device. For example, the data encoding device may be implemented in the following manner:
[0165] (1) An independent integrated circuit IC, or chip, or chip system or subsystem; (2) A collection of one or more ICs, optionally including a storage component for storing data and instructions; (3) A module that can be embedded in other devices; (4) In-vehicle equipment, etc.; (5) Others, etc.
[0166] For the case where the data encoding device, encoding device, and decoding device are implemented in the form of a chip or a chip system, see Figure 14 Schematic diagram of the chip structure shown. Figure 14 The chip shown includes a processor 601 and an interface 602. The number of processors 601 can be one or more, and the number of interfaces 602 can be multiple. Optionally, the chip or chip system can include a memory 603.
[0167] Among them, all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0168] Based on the same technical concept, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program includes at least one section of code, and the at least one section of code can be executed by a computer to control the computer to implement the above method embodiment.
[0169] Based on the same technical concept, an embodiment of the present application also provides a computer program, which, when executed, is used to implement the above method embodiment.
[0170] The program may be stored in whole or in part on a storage medium packaged with the processor, or may be stored in whole or in part on a memory not packaged with the processor.
[0171] Based on the same technical concept, an embodiment of the present application further provides a chip including a processor. The processor can implement the above method embodiment.
[0172] Reference Figure 15 , Figure 15 The present invention provides a schematic diagram of the structure of a data encoding device, wherein the device includes:
[0173] The acquisition module 1501 is used to obtain a target matrix, which is a check matrix or a generator matrix. The target matrix is obtained by converting the finite field 2 a A binary matrix represented by the data units on a polynomial ring, wherein the polynomial corresponding to the polynomial ring is an irreducible polynomial with a highest degree not greater than a;
[0174] The description of the acquisition module 1501 may refer to the introduction of step S401 in the above embodiment, and the similarities are not repeated here.
[0175] The encoding module 1502 is used to encode the data block according to the target matrix to obtain an encoding result; wherein,
[0176] The encoding result is used as a redundant data block of the data block; or, the encoding result is a recovery result of the data block of the node.
[0177] In one possible implementation, x a +x a-1 +…+x 2 +x+1=0 is a reducible polynomial, and the polynomial corresponding to the polynomial ring is an irreducible polynomial whose highest degree is not greater than a.
[0178] In a possible implementation, the data block and the encoding result are stored on different nodes.
[0179] In a possible implementation, the polynomial corresponding to the polynomial ring is the polynomial with the smallest order among the irreducible polynomials whose highest degree is not greater than a, wherein the polynomial can divide one or more polynomials of the form x b +1, the order of the polynomial is the one or more forms of x b The smallest value of b in the expression of +1.
[0180] In a possible implementation, the target matrix includes a submatrix for representing a first data unit, the data block includes a second data unit, the second data unit is represented by a first vector, and the first vector is a binary vector;
[0181] The processing module is specifically used to:
[0182] Performing an XOR operation on the submatrix and the first vector to obtain a second vector, where the second vector is a binary vector;
[0183] A polynomial having binary elements in the second vector as coefficients is divided by a polynomial corresponding to the polynomial ring and a remainder is taken to obtain an operation result equivalent to a product result of the first data unit and the second data unit.
[0184] In a possible implementation, the target matrix includes a submatrix for representing the first data unit, and the number of rows of the submatrix is the order of the polynomial corresponding to the polynomial ring.
[0185] In one possible implementation, the target column is obtained by filling the binary representation of the first data unit and then performing the transformation on some binary elements so that the number of elements of 1 included in the submatrix used to represent the first data unit is minimized, and the transformation includes changing 0 to 1, or changing 1 to 0.
[0186] In one possible implementation, the submatrix includes a target column, where the target column is obtained by transforming some binary elements after padding the binary representation of the first data unit. The target column includes X rows, and the transformation includes: when an element of the target row in M rows following the X rows is 1, elements of rows before the target row in the X rows are transformed accordingly based on the polynomial.
[0187] In a possible implementation, M is the difference between X and a.
[0188] In a possible implementation, the target column is obtained by filling the binary representation of the first data unit and performing the transformation on some binary elements so as to minimize the number of elements with 1 included in the submatrix used to represent the first data unit.
[0189] In a possible implementation, the target matrix includes a plurality of sub-matrices for representing a plurality of data units, and the positions of the plurality of data units and the corresponding sub-matrices in the target matrix are based on the following constraints from the finite field 2 a Selected from:
[0190] Under the condition that the target matrix has the maximum distance separable (mds) property, the number of elements of 1 included in the target matrix is minimized.
[0191] The steps of the method or algorithm described in conjunction with the disclosure of the embodiments of the present application can be implemented in a hardware manner, or can be implemented by a processor executing a software instruction. The software instruction can be composed of corresponding software modules, and the software module can be stored in a random access memory (Random Access Memory, RAM), a flash memory, a read-only memory (Read Only Memory, ROM), an erasable programmable read-only memory (Erasable Programmable ROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM), a register, a hard disk, a mobile hard disk, a read-only compact disc (CD-ROM) or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and can write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0192] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0193] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A data encoding method, characterized in that: include: Get the target matrix, which is a check matrix or a generator matrix. The target matrix is a finite field 2 a A binary matrix represented by the data units on a polynomial ring, wherein the polynomial corresponding to the polynomial ring is an irreducible polynomial with a highest degree not greater than a; According to the target matrix, the data block is encoded to obtain an encoding result; wherein, The encoding result is used as a redundant data block of the data block; or, the encoding result is a recovery result of the data block of the node.
2. The method according to claim 1, characterized in that x a +x a-1 +…+x 2 +x+1=0 is a reducible polynomial, and the polynomial corresponding to the polynomial ring is an irreducible polynomial whose highest degree is not greater than a.
3. The method according to claim 1 or 2, characterized in that The data block and the encoding result are stored on different nodes.
4. The method according to any one of claims 1 to 3, characterized in that: The polynomial corresponding to the polynomial ring is the polynomial with the smallest order among the irreducible polynomials whose highest degree is not greater than a, wherein the polynomial can divide one or more forms of x b +1, the order of the polynomial is the one or more forms of x b The smallest value of b in the expression of +1.
5. The method according to any one of claims 1 to 4, characterized in that: The target matrix includes a submatrix for representing a first data unit, the data block includes a second data unit, the second data unit is represented by a first vector, and the first vector is a binary vector; The encoding of the data block according to the target matrix includes: Performing an XOR operation on the submatrix and the first vector to obtain a second vector, where the second vector is a binary vector; A polynomial having binary elements in the second vector as coefficients is divided by a polynomial corresponding to the polynomial ring and a remainder is taken to obtain an operation result equivalent to a product result of the first data unit and the second data unit.
6. The method according to any one of claims 1 to 5, characterized in that: The target matrix includes a submatrix for representing the first data unit, and the number of rows of the submatrix is the order of the polynomial corresponding to the polynomial ring.
7. The method according to any one of claims 1 to 6, characterized in that: The target column is obtained by filling the binary representation of the first data unit and then performing the transformation on some binary elements so that the number of elements of 1 included in the submatrix used to represent the first data unit is minimized, and the transformation includes changing 0 to 1, or changing 1 to 0.
8. The method according to any one of claims 1 to 7, characterized in that: The submatrix includes a target column, where the target column is obtained by transforming some binary elements after filling the binary representation of the first data unit. The target column includes X rows, and the transformation includes: when an element of the target row in the M rows after the X rows is 1, elements of rows before the target row in the X rows are correspondingly transformed based on the polynomial.
9. The method according to claim 8, characterized in that The M is the difference between X and a.
10. The method according to claim 8 or 9, characterized in that The target column is obtained by performing the transformation on some binary elements after filling the binary representation of the first data unit so as to minimize the number of elements with 1 included in the submatrix used to represent the first data unit.
11. The method according to any one of claims 1 to 10, characterized in that: The target matrix comprises a plurality of sub-matrices for representing a plurality of data units, wherein the positions of the plurality of data units and the corresponding sub-matrices in the target matrix are based on the following constraints from the finite field 2 a Selected from: Under the condition that the target matrix has the maximum distance separable (mds) property, the number of elements of 1 included in the target matrix is minimized.
12. A data encoding device, characterized in that: include: The acquisition module is used to obtain a target matrix, wherein the target matrix is a check matrix or a generator matrix, and the target matrix is obtained by converting a finite field 2 a A binary matrix represented by the data units on a polynomial ring, wherein the polynomial corresponding to the polynomial ring is an irreducible polynomial with a highest degree not greater than a; The encoding module is used to encode the data block according to the target matrix to obtain an encoding result; wherein, The encoding result is used as a redundant data block of the data block; or, the encoding result is a recovery result of the data block of the node.
13. The device according to claim 12, characterized in that x a +x a-1 +…+x 2 +x+1=0 is a reducible polynomial, and the polynomial corresponding to the polynomial ring is an irreducible polynomial whose highest degree is not greater than a.
14. The device according to claim 12 or 13, characterized in that The data block and the encoding result are stored on different nodes.
15. The device according to any one of claims 12 to 14, characterized in that The polynomial corresponding to the polynomial ring is the polynomial with the smallest order among the irreducible polynomials whose highest degree is not greater than a, wherein the polynomial can divide one or more forms of x b +1, the order of the polynomial is the one or more forms of x b The smallest value of b in the expression of +1.
16. The device according to any one of claims 12 to 15, characterized in that The target matrix includes a submatrix for representing a first data unit, the data block includes a second data unit, the second data unit is represented by a first vector, and the first vector is a binary vector; The processing module is specifically used to: Performing an XOR operation on the submatrix and the first vector to obtain a second vector, where the second vector is a binary vector; A polynomial having binary elements in the second vector as coefficients is divided by a polynomial corresponding to the polynomial ring and a remainder is taken to obtain an operation result equivalent to a product result of the first data unit and the second data unit.
17. The device according to any one of claims 12 to 16, characterized in that The target matrix includes a submatrix for representing the first data unit, and the number of rows of the submatrix is the order of the polynomial corresponding to the polynomial ring.
18. The device according to any one of claims 12 to 17, characterized in that The target column is obtained by filling the binary representation of the first data unit and then performing the transformation on some binary elements so that the number of elements of 1 included in the submatrix used to represent the first data unit is minimized, and the transformation includes changing 0 to 1, or changing 1 to 0.
19. The device according to any one of claims 12 to 18, characterized in that The submatrix includes a target column, where the target column is obtained by transforming some binary elements after filling the binary representation of the first data unit. The target column includes X rows, and the transformation includes: when an element of the target row in the M rows after the X rows is 1, elements of rows before the target row in the X rows are correspondingly transformed based on the polynomial.
20. The device according to claim 19, characterized in that The M is the difference between X and a.
21. The device according to claim 19 or 20, characterized in that The target column is obtained by performing the transformation on some binary elements after filling the binary representation of the first data unit so as to minimize the number of elements with 1 included in the submatrix used to represent the first data unit.
22. The device according to any one of claims 12 to 21, characterized in that The target matrix comprises a plurality of sub-matrices for representing a plurality of data units, wherein the positions of the plurality of data units and the corresponding sub-matrices in the target matrix are based on the following constraints from the finite field 2 a Selected from: Under the condition that the target matrix has the maximum distance separable (mds) property, the number of elements of 1 included in the target matrix is minimized.
23. A computing device, characterized in that The computing device includes a memory and a processor, and the processor executes computer instructions stored in the memory, so that the computing device performs the method according to any one of claims 1 to 11.
24. A storage system, characterized in that: A device comprising the apparatus of any one of claims 12 to 22.
25. A computer-readable storage medium, characterized in that The method comprises instructions which, when executed on a computing device, cause the computing device to perform the method according to any one of claims 1 to 11.