Data processing method and device and storage medium
By using the first coding matrix to encode k source data blocks to generate k source data blocks and m check data blocks, the problems of wide-band EC technology in repair bandwidth and decoding complexity are solved, and efficient data storage is achieved.
Patent Information
- Application Number
- CN202410288351.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing multi-copy technology and narrow-stripe erasure codes are inefficient in data storage, and wide-stripe EC technology has problems with repair bandwidth and decoding complexity, making it difficult to effectively improve storage efficiency.
By using the first coding matrix to encode k source data blocks, k source data blocks and m check data blocks are generated, thereby reducing the amount of redundant data and improving data storage efficiency.
On the basis of ensuring data reliability, using as few check data blocks as possible reduces the amount of redundant data and improves data storage efficiency.
Smart Images

Figure CN120639100A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular to a data processing method, device, and storage medium. Background Art
[0002] With the rapid development of technologies such as high-speed fixed-line broadband and mobile internet, the amount of data generated daily is exploding. How to store this data securely, reliably, and efficiently has become an important research topic.
[0003] Currently, the most commonly used data storage technologies are multi-replication and narrow-stripe erasure coding (EC). While these two storage technologies alleviate data growth issues to a certain extent, they generate redundant checksum data and consume a large amount of storage space, resulting in high storage costs and low efficiency. Therefore, compared to multi-replication and narrow-stripe EC, wide-stripe EC, which generates less redundant checksum data, is becoming the future storage trend. Proper use of wide-stripe EC can effectively alleviate data growth issues and improve storage efficiency.
[0004] Wide-stripe EC technology has been used in related technologies to store data. However, due to technical limitations, wide-stripe EC has not fully utilized its storage capabilities, resulting in limited improvements in storage efficiency. Therefore, how to use wide-stripe EC technology to store and process data and improve data storage efficiency is an urgent problem to be solved. Summary of the Invention
[0005] The embodiments of the present disclosure provide a data processing method, device, and storage medium for improving data storage efficiency.
[0006] In a first aspect, a data processing method is provided, comprising:
[0007] Obtaining a first encoding matrix;
[0008] Based on the first coding matrix, k source data blocks are processed to obtain n coded data blocks; wherein the n coded data blocks include: k source data blocks and m check data blocks, k is greater than 1, n is greater than k, and k, n, are all positive integers, and m is equal to nk.
[0009] In a second aspect, a data processing device is provided, comprising:
[0010] An acquisition module, configured to acquire a first encoding matrix;
[0011] A processing module is used to process k source data blocks based on a first coding matrix to obtain n coded data blocks; wherein the n coded data blocks include: k source data blocks and m check data blocks, k is greater than 1, n is greater than k, and k, n, are all positive integers, and m is equal to nk.
[0012] In a third aspect, another data processing device is provided, comprising a processor, which implements the data processing method of the first aspect when executing a computer program.
[0013] In a fourth aspect, a computer-readable storage medium is provided, the computer-readable storage medium including computer instructions; wherein, when the computer instructions are executed, the data processing method of the first aspect mentioned above is implemented.
[0014] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to implement the data processing method of the first aspect.
[0015] In the disclosed embodiment, k source data blocks are encoded using a first encoding matrix to obtain k source data blocks and m check data blocks. While ensuring the reliability of the data content, as few check data blocks as possible are used, thereby reducing the amount of redundant data that needs to be stored and improving data storage efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present disclosure, the following briefly introduces the drawings required for use in some embodiments of the present disclosure. Obviously, the drawings described below are only drawings of some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0017] Figure 1 A schematic diagram of the architecture of a data processing system provided in an embodiment of the present disclosure;
[0018] Figure 2 A schematic diagram of the structure of a storage node set provided in an embodiment of the present disclosure;
[0019] Figure 3 A flowchart of a data processing method provided in an embodiment of the present disclosure;
[0020] Figure 4 A flowchart of another data processing method provided in an embodiment of the present disclosure;
[0021] Figure 5 A flowchart of another data processing method provided in an embodiment of the present disclosure;
[0022] Figure 6 A schematic diagram of the structure of a source data block and a check data block provided in an embodiment of the present disclosure;
[0023] Figure 7 A flowchart of another data processing method provided in an embodiment of the present disclosure;
[0024] Figure 8 A flowchart of another data processing method provided in an embodiment of the present disclosure;
[0025] Figure 9 A schematic structural diagram of a data processing device provided in an embodiment of the present disclosure;
[0026] Figure 10 A schematic structural diagram of another data processing device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present disclosure.
[0028] In the description of the present disclosure, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more. Words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not limit them to be necessarily different.
[0029] It should be noted that in this disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this disclosure as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0030] In data transmission systems, redundant checksums are often used to ensure data reliability and durability. Currently, commonly used data storage technologies include multiple copies and erasure codes.
[0031] Multi-copy technology involves making n copies of the same data, then distributing and storing these copies across different nodes according to a specific storage strategy. Nodes can be disks, solid-state drives, servers, or other storage devices. Each node is independent. If up to n-1 nodes fail, causing data loss, the same data stored on the failed nodes can still be retrieved from surviving nodes, thus providing data protection. The advantage of multi-copy technology is that it requires no complex computations, making it easy to operate. Furthermore, this lack of computation allows for faster data read and write speeds. However, multi-copy technology suffers from low storage space utilization, resulting in low storage efficiency. Given the current volume of data generated, multi-copy technology cannot meet future data storage needs. Therefore, multi-copy technology is typically used in hot data storage applications that require high availability and fast access.
[0032] Erasure coding is a forward error correction technology primarily used to prevent packet loss during network transmission. Erasure coding is commonly used in storage systems to improve data storage reliability. Compared to multiple replication technologies, erasure coding can achieve higher data reliability with less redundancy. However, the encoding and decoding of erasure codes is complex, involving multiplication operations in a finite field and requiring significant computing resources. Erasure codes are typically applied to warm or general data. Erasure codes are categorized into narrow-stripe EC and wide-stripe EC. Narrow-stripe EC involves fewer data blocks and is suitable for small-scale storage systems, but cannot meet future data storage needs.
[0033] Wide-stripe EC is not only applicable to large-scale storage systems, but also generates less redundant checksum data than multi-copy and narrow-stripe EC technologies. Therefore, wide-stripe EC technology is the future storage trend.
[0034] The most commonly used erasure code in storage systems is the Reed-Solomon (RS) erasure code. However, the decoding computational complexity of RS erasure codes in wide-strip EC is large, and the repair bandwidth is high. For example, if the stripe width is directly increased based on the currently used erasure code, and the number of source data blocks is increased to k based on the RS erasure code of the Intel Storage Acceleration Library (ISA-L), a large amount of data repair bandwidth will be generated when repairing one or more source data nodes. This means that when an error occurs in one or more source data nodes, k times the amount of data of a single source data block needs to be transmitted across the network, and the large amount of network bandwidth occupied by data repair is unacceptable.
[0035] To address the issue of high repair bandwidth, local repair concepts such as locally repairable codes (LRC) and regeneration codes (RGC) have been proposed. Locally repairable codes offer excellent repair locality. When a single source data node fails, repair only needs to be performed in the area where the source data node is located, reducing the repair bandwidth. However, locally repairable codes are not maximum distance separable codes (MDS). Compared to RS-type erasure codes, they occupy more storage space for parity data while maintaining the same repair capability. Regeneration codes have lower repair bandwidth overhead than RS-type erasure codes. When a single source data node fails, the remaining data nodes only need to transmit a portion of their own data to complete the repair of the source data node. However, when regeneration codes are applied to wide-strip EC, they suffer from high decoding complexity and inferior repair capability to RS codes, making their application in wide-strip EC difficult.
[0036] For example, if the idea of local repair code is directly used for encoding, although the repair bandwidth can be reduced, the redundancy rate is high and the storage efficiency cannot be effectively improved.
[0037] Based on this, the present disclosure provides a data processing method that encodes k source data blocks using a first encoding matrix to obtain k source data blocks and m check data blocks. While ensuring the reliability of the data content, the method uses as few check data blocks as possible, thereby reducing the amount of redundant data that needs to be stored and improving data storage efficiency.
[0038] The data processing method provided by the present disclosure can be applied to Figure 1 In the data processing system shown, Figure 1 A schematic diagram of the architecture of a data processing system provided by an embodiment of the present disclosure is shown.
[0039] like Figure 1 As shown, the data processing system 100 includes a first transmission node 110 , a second transmission node 170 , networks 120 and 160 , a first processing node 130 , a second processing node 150 , and a storage node set 140 , wherein the storage node set includes at least one storage node 141 .
[0040] In some embodiments, the first transmission node and the second transmission node can be the same device or different devices. The first transmission node 110 and the second transmission node 170 can include but are not limited to: a server, a base station, a mobile phone, a personal computer, a laptop, a tablet computer, an access point (AP), or other devices with the function of transmitting and storing data.
[0041] In some embodiments, the relative position, quantity, etc. of the first transmission node 110 and the second transmission node 170 may vary in specific application scenarios. The relative position refers to the positional relationship of the transmission nodes in the data processing system 100 or other network topology, such as the connection method between the transmission nodes and the direction of data transmission. For example, the first transmission node 110 can send data via the network 120, and the second transmission node 170 can receive data via the network 160. Furthermore, the second transmission node 170 can send data via the network 160, and the first transmission node 110 can receive data via the network 120. The data transmission system may include one or more first transmission nodes 110 and second transmission nodes 120, each of which is connected via the network 120 and the network 160.
[0042] In some embodiments, network 120 and network 160 can be any of the following: wired network, wireless network, cellular network, local area network, Internet, wide area network, World Wide Web, or other networks that transmit data.
[0043] In some embodiments, the first processing node 130 and the second processing node 150 can be the same processing node of the same device at different times. The first processing node 130 and the second processing node 150 can access and transmit data via a network and process the data according to configuration items. Exemplarily, the configuration items include at least one of the following: the location of the storage node for accessing the source data block, the location of the destination storage node for transmitting data, the number of check data blocks generated by encoding the source data block, the number of second check data blocks generated by encoding the data block, the number of source data blocks, the size of the source data block, the number of storage nodes, the number of columns in the second encoding matrix, the number and location of failed storage nodes, a data recovery strategy, or other parameters used for data processing. Processing data according to the configuration items includes at least one of the following: determining a second encoding matrix based on the number k of source data blocks, the number n of encoded data blocks, and the number K of columns of the second encoding matrix; generating a first encoding matrix based on the second encoding matrix; processing k source data blocks based on the first encoding matrix to obtain n encoded data blocks; allocating data blocks to storage nodes; accessing data blocks from storage nodes; recovering faulty data blocks; or other processing methods such as reading, writing, deleting, and modifying data blocks.
[0044] In some embodiments, the storage node set 140 includes at least n storage nodes 141, each storage node 141 can be any of the following: RAM, ROM, EEPROM, flash memory or other memory technology, or other optical disk storage, magnetic cassettes, tapes, disk storage, hard disks, solid-state drives or other storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0045] In some embodiments, as Figure 2 As shown, it is a structural diagram of a storage node set 1400 provided by an embodiment of the present disclosure. Storage node 1401 represents a storage node set, which includes n storage nodes, including k source data block storage nodes and m check data block storage nodes, k is a positive integer greater than 1, n is a positive integer greater than k, and nk>2, m=nk; 1402 represents a Chunk data block, each storage node contains p Chunk data blocks, p is an integer greater than 0; 1403 represents a Page data block, each Chunk data block contains w Page data blocks, w is an integer greater than 0; 1404 represents a stripe, each stripe contains e*n Page data blocks, and each e Page data blocks are distributed in a storage node, that is, in each stripe, each storage node contains e Page data blocks, and e is an integer greater than 0. It can be understood that in the present disclosure, Chunk data blocks and Page data blocks only represent the inclusion relationship between large data blocks and small data blocks, and do not specifically refer to other meanings.
[0046] In some embodiments, k source data blocks respectively correspond to the data parts stored in the first k storage nodes in a stripe, and the k source data blocks are distributed in k different storage nodes; m check data blocks obtained by encoding the k source data blocks according to the encoding matrix respectively correspond to the data parts stored in the last m storage nodes in a stripe, and the m check data blocks are respectively stored in the remaining m different storage nodes.
[0047] In some embodiments, the portion of each storage node of each stripe containing e Page data blocks can be regarded as a source data block, k source data blocks constitute the source data portion of a stripe, and m check data blocks obtained by encoding the k source data blocks constitute the check data portion of a stripe.
[0048] The data processing method provided by the embodiment of the present disclosure can be applied to Figure 1 Alternatively, the corresponding execution subject may be selected according to the actual application scenario, which is not limited in the embodiment of the present disclosure.
[0049] It should be noted that Figure 1 This is just an illustrative framework diagram. Figure 1 The number of devices included in the Figure 1 In addition to the devices shown, a data processing system may include other devices.
[0050] Figure 3 A flow chart of a data processing method provided by the present disclosure is shown in FIG. Figure 3 As shown, the following steps are included:
[0051] S101: Obtain a first encoding matrix.
[0052] In some embodiments, as Figure 4 As shown, the first encoding matrix is determined according to the following method:
[0053] S201: Determine a second coding matrix according to the number k of source data blocks, the number n of coded data blocks, and the number K of columns of the second coding matrix.
[0054] The dimension of the second encoding matrix is (nk)*K, K is greater than k, and K is a positive integer.
[0055] In some embodiments, the number of columns K of the second coding matrix is greater than or equal to the first number k+N1 and less than or equal to the second number k+N2; wherein N1 is the number of unrecovered coded data blocks when the first k columns of the second coding matrix are used as the coding matrix to recover m-1 coded data blocks; and N2 is qn, where q is the size of the Galois Field GF(q). q is an integer power of a prime number, and q can be 256 or other values, which are not limited in this disclosure.
[0056] For example, when k=160, n=165, K=240, q=256, and m-1=4, and the first k columns of the second coding matrix are used as the coding matrix to recover four coding data blocks, and the number of unrecovered coding data blocks N1=4, since the value of K is required to be greater than or equal to the first number 164 and less than or equal to the second number 251, K can be 240. N1 can be determined by traversing the failure conditions of all m-1 coding data blocks.
[0057] As another example, when k=160, n=165, K=165, and q=256, and the first k columns of the second coding matrix are used as the coding matrix to restore m-1 coded data blocks, and the number of unrestored coded data blocks N1=7, the value of K does not meet the requirement of being greater than or equal to the first number 167 and less than or equal to the second number 251. Therefore, K cannot be 165, and the value of K needs to be increased to meet the K value requirement.
[0058] In some embodiments, the number of columns K of the second encoding matrix is determined according to the number k of source data blocks.
[0059] Exemplarily, the number of columns K∈[K1,K2] of the second coding matrix is represented by the number of columns K of the second coding matrix belonging to any integer in the above set, and the sizes of K1 and K2 are determined according to k. Among them, when k=160, n=165, q=256, m-1=4, K1 is an integer greater than 163 and less than 200, and K2 is an integer greater than 220 and less than 251. For example, when the number of source data blocks k is 160, the number of columns K∈[167,253] of the second coding matrix, K can be 240. The purpose of limiting the range of K is to select as few columns of the second coding matrix as possible while satisfying the repair capability, thereby reducing the amount of calculation.
[0060] In some embodiments, a square matrix consisting of any (nk) rows and (nk) columns in the second coding matrix is reversible in the Galois field GF(q), and there is one row in the second coding matrix whose elements are all 1.
[0061] In some embodiments, the elements in the second coding matrix are the calculation results of the first element and the second element corresponding to the element in the Galois field GF(q); wherein the first element is determined based on the number of columns of the second coding matrix and the number of columns to which the element belongs; the second element is determined based on the number of columns of the second coding matrix, the number of columns to which the element belongs, and the number of rows to which the element belongs.
[0062] As a possible implementation, the element in the i-th row and j-th column of the second coding matrix can be the result of dividing the corresponding first element by the second element in the Galois field GF(q). The second element corresponding to the element in the i-th row and j-th column of the second encoding matrix is i is an integer greater than or equal to 0 and less than or equal to m-1, and j is an integer greater than or equal to 0 and less than or equal to K-1. Represents bitwise exclusive OR. The second encoding matrix is represented by GO and can be expressed as follows:
[0063]
[0064] For example, taking the Galois field size q=256, the number of columns of the second coding matrix K=240, and the number of check data blocks m=5 as an example, the elements in the second coding matrix are the first elements corresponding to the Galois field GF(q) Divide by the second element For example, the second encoding matrix can be expressed as follows:
[0065]
[0066] In some embodiments, elements of two adjacent rows from the 2nd row to the mth row in the second encoding matrix support swapping.
[0067] For example, when the second row and the third row in the second coding matrix are swapped, the second coding matrix can be expressed as follows:
[0068]
[0069] In some embodiments, each of the n coded data blocks has a unique index representation, which is used to represent the relative position of each coded data block in a slice.
[0070] Exemplarily, the number of coded data blocks is n, the value range of the index identifier of the coded data block is an integer greater than or equal to 1 and less than or equal to n, and the index identifier of each coded data block is different. Alternatively, the number of coded data blocks is n, the value range of the index identifier of the coded data is an integer greater than or equal to 0 and less than or equal to n-1, and the index identifier of each coded data block is different.
[0071] S202: Generate a first encoding matrix based on the second encoding matrix.
[0072] The dimension of the first encoding matrix is (nk)*k.
[0073] In some embodiments, as Figure 5 As shown, based on the second coding matrix, generating the first coding matrix can be specifically implemented as follows:
[0074] S301: Obtain the first k columns of the second coding matrix as the index of the coding data block that has not been completely restored when the coding matrix restores m-1 coding data blocks.
[0075] For example, the first k columns of the second coding matrix can be used as the coding matrix to recover the m-1 coding data blocks, and the index of the coding data block that has not been recovered can be obtained. When the first k columns of the second coding matrix are used as the coding matrix, the coding matrix G can be expressed as follows:
[0076]
[0077] In another exemplary embodiment, the number of columns of the second coding matrix K = 240, the number of source data blocks k = 160, the number of coded data blocks n = 165, the number of check data blocks m = nk = 5, and the dimension size of the second coding matrix is 5*240. When the first 160 columns of the second coding matrix are used as the coding matrix to restore 4 coded data blocks, all failure cases that are not fully restored are: {45, 90, 91, 106}th coded data blocks have failures at the same time, {17, 63, 118, 142}th coded data blocks have failures at the same time, {83, 95, 122, 156}th coded data blocks have failures at the same time, and {88, 110, 147, 151}th coded data blocks have failures at the same time. In the above set, each element represents a coded data block with an index identifier of that value. For example, the index identifier of the 45th coded data block is 45, and the index identifier is an integer with a minimum value equal to 1 and a maximum value equal to n. By traversing the m-1 coding data block failure situations, the failure situation that has not been completely recovered can be obtained, and then the index of the coding data block that has not been completely recovered can be obtained.
[0078] S302: Replace the L column elements corresponding to the index of the unrecovered coded data block in the second coding matrix with any L column elements from the k+1th column to the Kth column in the second coding matrix.
[0079] Exemplarily, after obtaining the index identifiers of the unrecovered coded data blocks in the second coding matrix, the index identifiers constitute an index set S={s0, s1, ..., s L-1}, the length of the index set S is L. For example, the number of columns of the second coding matrix is K = 240, the number of source data blocks is k = 160, the number of coded data blocks is n = 165, the number of check data blocks is m = nk = 5, and the dimension of the second coding matrix is 5 * 240. When the first 160 columns of the second coding matrix are used as the coding matrix to restore four coded data blocks, all the fault conditions that are not fully restored are: {45, 90, 91, 106}th coded data blocks have faults at the same time, {17, 63, 118, 142}th coded data blocks have faults at the same time, {83, 95, 122, 156}th coded data blocks have faults at the same time, and {88, 110, 147, 151}th coded data blocks have faults at the same time. For example, the coded data block index identifiers corresponding to the data blocks that have not been fully recovered are all the third positions of the m-1 faulty coded data block index identifiers, and the index set S = {91, 118, 122, 147}, and the length of the index set S is 4.
[0080] It should be noted that, in the case where m-1 coded data blocks are faulty at the same time, each coded data block supports a maximum of P faulty sub-data blocks to be recovered simultaneously, where P is an integer less than or equal to m-1 and greater than or equal to 1. Therefore, m-1 coded data blocks support a maximum of P*(m-1) faulty sub-data blocks to be recovered simultaneously. And recovering P*(m-1) faulty sub-data blocks requires at least P*(m-1) check sub-data blocks, that is, at least P*(m-1) linearly independent check equations are required. During the recovery process, there are check equations in the first part that can correctly recover P*(m-1) sub-data blocks, and there are check equations in the second part that cannot correctly recover P*(m-1) sub-data blocks. Among them, the check equations in the first part and the check equations in the second part do not contain each other. For the second part of the verification equation that cannot recover or correctly recover P*(m-1) sub-data blocks, if there is at least one sub-data block Z that is linearly related to at most P*(m-1)-1 other sub-data blocks, then the index identifier value of the encoded data block where sub-data block Z is located is the index identifier of the encoded data block corresponding to the data block that has not been fully recovered.
[0081] Exemplarily, the number of columns of the second coding matrix K = 240, the number of source data blocks k = 160, the number of coded data blocks n = 165, each coded data block includes 1024 sub-data blocks, and m-1 = 4 coded data blocks are simultaneously faulty. When the faulty coded data blocks are {45, 90, 91, 106}, each of the above 4 coded data blocks supports a maximum of 4 faulty sub-data blocks for joint recovery. If the {1, 65}th faulty coded sub-data blocks of each coded data block are jointly recovered, then there are 2*(m-1) = 8 faulty coded sub-data blocks that are jointly recovered, and the at least 8 linear equations formed are all linearly independent and can complete the recovery of 2*(m-1) sub-data blocks. The linear equations involved in the above recovery process belong to the verification equations of the first part. Therefore, in the above case, there is no index of the coded data block that has not been fully recovered.
[0082] In another exemplary embodiment, the number of columns of the second coding matrix is K=240, the number of source data blocks is k=160, the number of coded data blocks is n=165, each coded data block includes 1024 sub-data blocks, and m-1=4 coded data blocks are simultaneously faulty. When the faulty coded data blocks are {45, 90, 91, 106}, each of the above 4 coded data blocks supports a maximum of 4 faulty sub-data blocks for joint recovery. If the {17th, 33rd, 81st, 97th} faulty coded sub-data blocks of each coded data block are jointly recovered, then there are (m-1)*(m-1)=16 faulty coded sub-data blocks that are jointly recovered. Among the at least 16 linear equations formed, the 81st sub-data block of the 91st coded data block is linearly related to the other 15 sub-data blocks. The above linear equations belong to the verification equation of the second part, and the index value of the coded data block where the sub-data block is located is 91. Therefore, in the above case, the coded data block index corresponding to the coded data block that has not been fully recovered is identified as 91.
[0083] As a possible implementation manner, the L column elements corresponding to the index of the unrecovered coded data block in the second coding matrix are replaced by the L column elements from the k+1th column to the k+Lth column in the second coding matrix.
[0084] For example, the index in the second coding matrix is the index from column s0 to column s1 in the index set. L-1The elements of the columns are replaced with the elements of the L columns from k+1 to k+L. For example, the second coding matrix has K = 240 columns, k = 160 source data blocks, n = 165 coded data blocks, m = nk = 5 check data blocks, and the dimension of the second coding matrix is 5 * 240. Using the first 160 columns of the second coding matrix as the coding matrix, all failure scenarios where any four coded data block failures cannot be recovered are: {45, 90, 91, 106}th coded data blocks are simultaneously faulty, {17, 63, 118, 142}th coded data blocks are simultaneously faulty, {83, 95, 122, 156}th coded data blocks are simultaneously faulty, and {88, 110, 147, 151}th coded data blocks are simultaneously faulty. The index set S = {91, 118, 122, 147} of the indices of the coded data blocks that have not been recovered is composed of the length of the index set L = 4. Therefore, the elements in columns {91, 118, 122, 147} in the second coding matrix may be replaced with elements in four columns from columns 161 to 165 in the second coding matrix to obtain a replaced second coding matrix.
[0085] S303: Use the 1st to kth columns in the replaced second coding matrix as the first coding matrix.
[0086] Exemplarily, the first to kth columns of the replaced second coding matrix are used as the first coding matrix, and the size of the first coding matrix is m*k. For example, taking m=5, K=240, and k=160 as an example, the second coding matrix after replacement is 5*240, and the first to 160th columns of the replaced second coding matrix are used as the first coding matrix, and the size of the first coding matrix is 5*160. For example, the elements in the first coding matrix are as follows:
[0087] First row: all 1 elements;
[0088] Second line:
[0089]
[0090] Third line:
[0091]
[0092] Fourth line:
[0093]
[0094] Line 5:
[0095]
[0096] In some embodiments, any square matrix consisting of m rows and m columns in the first encoding matrix is reversible in the Galois Field GF(q).
[0097] It should be noted that the first coding matrix is generated based on the first to kth columns of the second coding matrix, and each element in the second coding matrix is the first element under the Galois field GF(q) Divide by the second element The calculation result is as follows. Take the second coding matrix with K = 240 columns, k = 160 source data blocks, n = 165 coding data blocks, and m = nk = 5 check data blocks as an example. The dimension of the second coding data block is 5*240. When the first k = 160 columns of the second coding matrix are used as the coding matrix to recover 4 coding data blocks, all the failure conditions are: {45, 90, 91, 106}th coding data blocks are simultaneously faulty, {17, 63, 118, 142}th coding data blocks are simultaneously faulty, {83, 95, 122, 156}th coding data blocks are simultaneously faulty, and {88, 110, 147, 151}th coding data blocks are simultaneously faulty. The index set S = {91, 118, 122, 147} of the indices of the coding data blocks that have not been recovered is composed of, and the length of the index set L = 4. Therefore, the elements in columns {91, 118, 122, 147} of the second coding matrix can be replaced with the elements in columns 161 to 164 of the second coding matrix to obtain the replaced second coding matrix. The elements in columns 1 to 160 of the replaced second coding matrix are used as the second coding matrix, which is reversible in the Galois Field GF(q).
[0098] In some embodiments, the first coding matrix is used to repair at least any m-1 coded data blocks.
[0099] S102: Process k source data blocks based on the first coding matrix to obtain n coded data blocks.
[0100] The n coded data blocks include: k source data blocks and m check data blocks, k is greater than 1, n is greater than k, and k, n, are both positive integers, and m is equal to nk.
[0101] In some embodiments, the m check data blocks include at least one of the following: 1 first check data block, m-2 second check data blocks, and 1 third check data block.
[0102] For example, Figure 6 The schematic diagram of the structure of k source data blocks and m check data blocks is shown. Figure 6 The relationship shown in the figure is that the maximum number of sub-data blocks g contained in each source data block, the first check data block, the second check data block, and the third check data block is s is an integer greater than 0 and less than k. In the present disclosure, s may also be referred to as the number of repetitions. The n coded data blocks include: k source data blocks and m check data blocks. Figure 6 1300 represents the data array of n coded data packets; 1301 represents k source data blocks, each source data block includes g source sub-data blocks. For example, 1305 represents one of the source sub-data blocks a 0,0 1302 represents a first check data block, which includes g check sub-data blocks. For example, 1306 represents one of the check sub-data blocks p 0,0 1303 represents m-2 second parity data blocks, each of which includes g parity sub-data blocks. 1304 represents third parity data blocks, each of which includes g parity sub-data blocks. The size of each source sub-data block is equal to the size of each parity sub-data block.
[0103] In some embodiments, as Figure 7 As shown, based on the first coding matrix, k source data blocks are processed to obtain n coded data blocks, including at least one of the following:
[0104] S401 . Encode k source data blocks based on the elements of the first row of a first encoding matrix to generate a first verification data block.
[0105] In some embodiments, the syndrome data block in the first syndrome data block is obtained by linearly combining source sub-data blocks corresponding to the syndrome data block, and coefficients of the linear combination are obtained based on the first row of the first coding matrix.
[0106] Exemplarily, each syndrome data block in the first parity data block is obtained by linear combination of the corresponding source sub-data blocks, the linear combination coefficients are obtained from the first row of the first encoding matrix, and all the linear combination coefficients are equal to 1. The i-th syndrome data block of the first parity data block is p i,0 , the source sub-data blocks corresponding to the i-th check sub-data block of the first check data block in the i-th row are [a i,0 ,a i,1 ,...,a i,k-1 The i-th syndrome data block of the first parity data block can be determined according to the following formula: i,0 =a i,0 ⊕a i,1 ⊕...⊕a i,k-1, where “⊕” represents bitwise XOR. Take k=4, g=4 as an example. The size of each source sub-data block is one byte, which is represented by a decimal number. The sub-check data block in the first check data block is obtained by bitwise XOR of the source sub-data blocks in the current row. For example, if the first source sub-data blocks of the four source data blocks are 10, 50, 90, and 130 respectively, then the value p of the first sub-check data block of the first check data block is 0,0 It is 10⊕5⊕90⊕130=216.
[0107] S402 : Encode k source data blocks based on elements from the 2nd row to the m-1th row of the first encoding matrix to generate a second verification data block.
[0108] In some embodiments, the check sub-data block in the second check data block is determined based on the linear combination value of the source sub-data blocks with a first source sub-data block index in the k source data blocks and the linear combination value of the source sub-data blocks with a second source sub-data block index in the k source data blocks; wherein the first index is the check sub-data block index corresponding to the c-th check sub-data block in the second check data block; the second index is any check sub-data block index in the second check data block except the first index; and c is a positive integer.
[0109] The second index is any check sub-data block index in the second check data block except the first index, and belongs to the set [1, g]\{c}, where g represents the number of coding sub-data blocks in each coding data block.
[0110] Exemplarily, each syndrome data block in the second parity data block is determined by a bitwise exclusive OR of a linear combination value of source sub-data blocks indexed as a first index in the k source data blocks and a linear combination value of at least one source sub-data block indexed as a second index in the k source data blocks. The i+1th syndrome data block p in the jth second parity data block is i,j The linear combination value p of the source sub-data block with the first index in the k source data blocks i,j,LC1 , and the linear combination value p of at least one source sub-data block with the second index in the k source data blocks i,j,LC2 The bitwise XOR is determined. And the following relationship is satisfied: p i,j =p i,j,LC1 ⊕p i,j,LC2 Wherein, i is an integer greater than or equal to 0 and less than or equal to g-1, and j is an integer greater than or equal to 1 and less than or equal to m-1. It can be understood that the first syndrome data block of the first second parity data block is p 0,1For example, if g = 1024, the first index value corresponding to the first syndrome sub-data block in the second syndrome data block is 1, and the second index value corresponding to the first syndrome sub-data block is an integer index value belonging to the set [2, 1024]. In the present disclosure, the source sub-data block indexed by the second index may also be referred to as an additional source sub-data block.
[0111] It can be understood that for the i+1th syndrome data block p in the jth second syndrome data block, i,j , the linear combination value p of the source sub-data block with the first index in the k source data blocks i,j,LC1 Determined by the source sub-data block in the i+1th row of the k source data blocks and the j+1th row in the encoding matrix, where i is an integer satisfying 0≤i≤g-1, and j is an integer satisfying 1≤j≤m-1.
[0112] For example, the i+1th source sub-data block of k source data blocks is represented as [a i,0 ,a i,1 ,...,a i,k-1 ], the element values of the j+1th row in the encoding matrix are expressed as [c j,0 ,c j,1 ,...,c j,k-1 ], p of the i+1th syndrome data block of the jth second parity data block i,j,LC1 It can be determined according to the following formula: i,j,LC1 =c j,0 *a i,0 ⊕c j,1 *a i,1 ⊕...⊕c j,k-1 *a i,k-1 Among them, “⊕” represents bitwise exclusive OR, and “*” represents multiplication operation under Galois Field GF(q). For example, the number of source data blocks k = 4, and the first source sub-data block of k source data blocks is represented as [a 0,0 ,a 0,1 ,a 0,2 ,a i,3 ]=[10,50,90,130], the element values of the j+1=2th row in the encoding matrix are expressed as [c j,0 ,c j,1 ,...,c j,k-1 ]=[166,70,187,123], then the first part p of the i=1th syndrome data block of the j=1th second syndrome data block i,j,LC1 =166*10+70*50+187*90+123*130=65.
[0113] It can be understood that for the i+1th syndrome data block p in the jth second syndrome data block, i,j, the linear combination value p of at least one source sub-data block with the second index of the k source data blocks is i,j,LC2 The determination is based on at least one of the following: an additional source sub-data block index table and a repetition count s. The additional source sub-data block index table is used to indicate the index positions in the array of additional source sub-data blocks required to generate the (i+1)th parity sub-data block of the jth second parity data block. The additional source sub-data block index table is an index array of size g*(m-2), where g is the number of rows in the source data block array and m-2 is the number of second parity data blocks.
[0114] Exemplarily, each element in the additional source sub-data block index table is a set of position coordinates. The position coordinates can be at least one of the following: a one-dimensional index coordinate or a two-dimensional index coordinate. It will be understood that both the one-dimensional index coordinate and the two-dimensional index coordinate represent the position of the additional element in the entire data array, and there is a one-to-one correspondence between the one-dimensional index coordinate and the two-dimensional index coordinate.
[0115] In another exemplary embodiment, each element in the additional source sub-data block index table is a position coordinate set. The number Ns of one-dimensional coordinates or two-dimensional coordinates in the position coordinate set is determined according to at least one of the following: the number k of source data blocks and the number m of check data blocks. The maximum value of the number Ns of one-dimensional coordinates or two-dimensional coordinates in the position coordinate set is Wherein, s is the number of repetitions, and s is a positive integer.
[0116] In another example, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. The number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most In the case of the second part, the linear combination value p of the source sub-data block of at least one other row i,j,LC2 It can be determined according to at least the following rules:
[0117]
[0118] in, Indicates bitwise exclusive OR, each element in the additional source sub-data block index table IndexTable is a two-dimensional coordinate set IJtable0. When the number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most When the two-dimensional coordinate set includes part of the generated p i,j,LC2 The index position of the required additional source sub-data block needs to be further obtained by repeating the number of times s. The above additional source sub-data block index table contains the index positions of some additional source sub-data blocks, and some additional source sub-data blocks represent the previous The additional source sub-data blocks of the source data block are repeated s times to obtain the index positions of all additional source sub-data blocks, which reduces storage by introducing a small amount of calculation.
[0119] Exemplarily, the additional source sub-data block index table IndexTable is determined according to at least one of the following: the number of source data blocks k, the number of second check data blocks m. The t-th column and Gi-th row of the source sub-data block index table are the coordinate indexes of the i-th column and Gj-th row of the array. Wherein, i is an integer greater than or equal to 1 and less than or equal to ceil(k / s). Gi is constructed in the following manner: the integer value {1, 2, ..., g} is divided into (m-1) ceil(i / (m-1)) sub-parts, and starting from the mod(i-1,m-1)+1th sub-part in sequence, read a sub-part every (m-1) sub-parts, and form a row index value set of the read sub-parts, which is Gi. The t-th element in the set GSet of the i-th source data block of the array is Gj, j belongs to the first integer set, and the first integer set is [floor(i / (m-1))+1,ceil(i / m-1)]\{i}. t belongs to the second integer set, and the second integer set contains integers greater than or equal to 1 and less than or equal to the number of integers in the first integer set. Gj is constructed in the following way: the integer value {1,2,...,g} is divided into (m-1) ceil(j / (m-1)) sub-parts, and starting from the mod(j-1,m-1)+1th sub-part in sequence, read a sub-part every (m-1) sub-parts, and form a row index value set of the read sub-parts, which is Gj. t and j are the elements corresponding to the same index size in the first integer set and the second integer set respectively. The elements in the first integer set and the second integer set are sorted from small to large. g is the number of rows in the array, ceil() means adjusting a value to the minimum integer not less than the value itself, floor() means adjusting a value to the maximum integer not greater than the value itself, mod() means modulo operation, and []\{i} means removing element i from a set.
[0120] As another example, let k' represent ceil(k / s), k=6, m-1=3, s=1, g=(m-1) ceil(k' / (m-1))=9 as an example. When i=2, Gi=G2={4,5,6}. For the second data block, [floor(i / (m-1))+1,ceil(i / (m-1))]\{i}={1,2,3}\{2}={1,3}. Therefore, the set of the second source data block in the array GSet={G1,G2,G3}\{G2}={G1,G3}. The value of t is an integer greater than or equal to 1 and less than or equal to 2. Above G1={1,2,3}, G3={7,8,9}. For the source sub-data block index table, the t=1 column and G2={4,5,6} row are the coordinate indexes of the sub-data blocks in the i=2 column and G1={1,2,3} row in the array respectively. For the source sub-data block index table, the t=2nd column and G2={4, 5, 6}th row are the coordinate indexes of the sub-data blocks in the array in the 2nd column and G3={7, 8, 9}th row respectively.
[0121] In another exemplary embodiment, k' represents ceil(k / s). The source sub-data block index table can be determined according to at least one of the following rules:
[0122]
[0123] Among them, cell() means building a cell array, floor() means adjusting a value to the maximum integer not greater than it, and g represents the number of rows in the data array. SplitNum() is a function related to the number of rows in the data array g and the number of check data blocks m. SplitNum() outputs a matrix of size (Layer*(m-1))*(g / (m-1)), where each element in the matrix is greater than or equal to 1 and less than or equal to g. GNum is a matrix with a total of (ceil(k' / (m-1)))*(m-1) rows. The i-th row of GNum is denoted as Gi, and Gi is constructed in the following way: the integer values {1,2,...,g} are divided into (m-1) in sequence. ceil(i / (m-1)) sub-parts, starting from the mod(i-1,m-1)+1th sub-part in sequence, read a sub-part every (m-1) sub-parts, and form Gi with the read sub-parts.
[0124] For example, k=6, m-1=3, g=(m-1) ceil(k' / (m-1)) For example, if Gi = G2, 1 to 9 are divided into 3 sub-parts in order, namely {1, 2, 3}, {4, 5, 6}, and {7, 8, 9}. Starting from the mod(i-1, m-1)+1=2th sub-part, a sub-part is read every 3 sub-parts, and the set formed is {4, 5, 6}, that is, G2 = {4, 5, 6}.
[0125] Another example, with k=6, r=3, g=(m-1) ceil(k' / (m-1))For example, if Gi = G4, 1 to 9 are divided into 9 sub-parts in order, namely {1}, {2}, {3}, {4}, {5}, {6}, {7}, {8}, {9}. Starting from the mod(i-1,m-1)+1=1 sub-part, a sub-part is read every 3 sub-parts, and the set formed is {1,4,7}, that is, G4 = {1,4,7}.
[0126] In another exemplary embodiment, the GNum matrix may be determined according to at least one of the following rules:
[0127]
[0128] Here, ceil() represents adjusting a value to the minimum integer not less than the value itself, g represents the number of rows in the data array, T represents transpose, and TempMx(j:(m-1):end,:) is a new matrix determined by extracting all columns of the matrix TempMx and every (m-1) rows from the jth row to the last row, forming a new matrix with these extracted columns. This new matrix is TempMx(j:(m-1):end,:). TempSubMx(:) is a one-dimensional matrix determined by expanding a two-dimensional matrix column by column to form a one-dimensional matrix. This one-dimensional matrix is TempSubMx(:).
[0129] For example, each element in the additional source sub-data block index table is a two-dimensional coordinate set. The number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most When , the two-dimensional coordinate set includes the second index source sub-data block position required for partially generating pi, j, LC2, and further obtains all the additional source sub-data block index positions through the number of repetitions s. For example, taking the number of source data blocks k = 8, the number of repetitions s = 1, the number of check data blocks m = 5, and the maximum number of two-dimensional coordinates Ns = 2 as an example, the corresponding additional source sub-data block index table is shown in Table 1. For example, to generate pi, j, LC2 of the third check sub-block of the first second check data block, then pi, j, LC2 = a6,
[0130] Table 1
[0131]
[0132] In another exemplary embodiment, each element in the additional source sub-data block index table is a two-dimensional coordinate set. The number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most When , the two-dimensional coordinate set includes the second index source sub-data block positions required for partially generating pi, j, LC2, and further obtains all the additional source sub-data block index positions through the number of repetitions s. For example, taking the number of source data blocks k = 70, the number of repetitions s = 10, the number of check data blocks m = 5, and the maximum number of two-dimensional coordinates Ns = 2 as an example, the corresponding additional source sub-data block index table is shown in Table 2. For example, to generate pi, j, LC2 of the fourth check sub-block of the first second check data block, then pi, j, LC2 = a7,
[0133] Table 2
[0134]
[0135] In another exemplary embodiment, each element in the additional source sub-data block index table is a two-dimensional coordinate set. The number of two-dimensional coordinates Ns in the two-dimensional coordinate set is at most When the two-dimensional coordinate set includes part of the generated p i,j,LC2 The required second index source sub-data block position is further obtained by repetition number s to obtain the index positions of all additional source sub-data blocks. For example, taking the number of source data blocks k = 160, the number of repetitions s = 8, the number of check data blocks m = 5, and the maximum number of two-dimensional coordinates Ns = 5 as an example, the corresponding partial additional source sub-data block index table is shown in Table 3.
[0136] Table 3
[0137]
[0138] It can be understood that the two-dimensional index coordinates in the above examples all represent the position of the additional sub-data block in the entire data array, and the two-dimensional coordinates can be converted into one-dimensional index coordinate representation, that is, there is a one-to-one correspondence between the one-dimensional coordinate and the two-dimensional index coordinate. Exemplarily, the one-dimensional coordinate index corresponding to the two-dimensional coordinate index (i1, j1) of the additional sub-data block is (g*j1+i1). Another exemplary example is that the one-dimensional coordinate index corresponding to the two-dimensional coordinate index (i1, j1) is (k*i1+j1). Wherein, i1 is an integer greater than or equal to 0 and less than g, j1 is an integer greater than 0 and less than or equal to k, and g is the number of rows of the data array.
[0139] S403 . Encode k source data blocks based on the elements of the mth row of the first encoding matrix to generate a third verification data block.
[0140] The syndrome data block in the third syndrome data block is obtained by linearly combining the source sub-data blocks corresponding to the syndrome data block, and the coefficients of the linear combination are obtained based on the mth row of the first coding matrix.
[0141] Exemplarily, each syndrome data block in the third parity data block is obtained by a linear combination of the corresponding source sub-data blocks, and the linear combination coefficients are obtained from the mth row of the encoding matrix. The i-th syndrome data block of the third parity data block is represented by p i,m-1 , the source sub-data blocks corresponding to the i-th row of the data array where the i-th check sub-data block of the third check data block is located are [a i,0 ,a i,1 ,...,a i,k-1 ], the linear combination coefficients corresponding to the third check data block are [c j,0 ,c j,1 ,...,c j,k-1 The linear combination coefficients are obtained from the mth row of the encoding matrix. The i-th syndrome data block of the third check data block can be determined according to at least one of the following rules: i,m-1 =c j,0 *a i,0 ⊕c j,1 *a i,1 ⊕...⊕c j,k-1 *a i,k-1 . Among them, "⊕" represents bitwise exclusive OR, "*" represents the multiplication operation under the Galois field GF(q), i is an integer greater than or equal to 0 and less than g, and g is the number of rows of the data array. For example, take k=4 and g=4 as an example. The size of each source sub-data block is one byte, which is represented by a decimal number. When the linear combination coefficient is obtained from the mth row of the encoding matrix. Taking the above linear combination coefficient as [1,8,64,58], and the first source sub-data block of the four source data blocks being 10, 50, 90, and 130 respectively, the i-th check sub-data block of the third check data block can be determined according to the following rules: p i,m-1 =c j,0 *a i,0 ⊕c j,1 *a i,1 ⊕...⊕c j,k-1 *a i,k-1 =1*10⊕8*50⊕64*90⊕58*130=188. “⊕” represents bitwise exclusive OR, and “*” represents multiplication operation under Galois Field GF(q).
[0142] It should be noted that the m check data blocks include 1 first check data block, m-2 second check data blocks, and 1 third check data block. The presence of the first check data block and the second check data block enables the repair of the source data block to be completed with a lower repair bandwidth when a single or multiple source data nodes fail. The presence of the third check data block enables the repair of the source data block to be completed using the third check data block when the first check data block and the second check data block are unable to complete data repair. By reducing the repair bandwidth through the first check data block and the second check data block, and ensuring the repair capability through the third check data block, the characteristics of high repair capability and low repair bandwidth are guaranteed.
[0143] In this way, by encoding k source data blocks using the first encoding matrix, k source data blocks and m check data blocks are obtained. Using as few check data blocks as possible reduces the amount of redundant data that needs to be stored while ensuring the reliability of the data content, thereby improving data storage efficiency.
[0144] In some embodiments, Figure 8 A flow chart of another data processing method provided by the present disclosure is shown as follows: Figure 8 As shown, the following steps are included:
[0145] S501: Receive an encoded data packet.
[0146] The coded data packet is generated by encoding k source data blocks according to the first coding matrix.
[0147] S502: Process the coded data packet based on the first coding matrix to restore data of the faulty coded data block.
[0148] The first coding matrix can be calculated as follows: Figure 4 The embodiments shown are certain and the present disclosure will not be described in detail here.
[0149] Exemplarily, the fault code data block includes at least one of the following: a source data block, a first check data block, a second check data block, and a third check data block. The number Ne of the fault code data blocks is an integer greater than or equal to 0 and less than or equal to n.
[0150] As a possible implementation, restoring data of a faulty coded data block includes at least one of the following:
[0151] When the number of faulty coded data blocks Ne is greater than 1 and less than or equal to m, and the faulty coded data blocks include source data blocks and check data blocks, all non-faulty coded data blocks are acquired, and the faulty source data blocks are restored first, followed by the check data blocks.
[0152] When the number Ne of faulty coded data blocks is greater than 1 and less than or equal to m, and the faulty coded data blocks include only source data blocks, obtain a data volume of at least Ne / (m-1) proportion of the remaining non-faulty data blocks, and then restore the faulty source data blocks;
[0153] When the number Ne of faulty coded data blocks is greater than or equal to 1 and less than or equal to m, and the faulty coded data blocks only include check data blocks, all non-faulty source data blocks are acquired, and then the faulty check data blocks are restored;
[0154] When the number Ne of faulty coded data blocks is equal to 1 and the faulty coded data blocks only include source data blocks, obtain the data amounts of the remaining non-faulty source data blocks, the first verification data blocks, and the second verification data blocks in a ratio of 1 / (m-1) respectively, and then restore the faulty source data blocks.
[0155] It should be noted that when the number Ne of faulty coded data blocks exceeds m, the data representing the faulty coded data blocks is lost and data recovery cannot be completed.
[0156] It is understandable that, in order to implement the above functions, the data processing device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments of the present disclosure, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present disclosure.
[0157] The embodiment of the present disclosure can divide the data processing device into functional modules according to the above-mentioned method embodiment. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one functional module. The above-mentioned integrated module can be implemented in the form of hardware or software. It should be noted that the division of modules in the embodiment of the present disclosure is schematic and is only a logical function division. There may be other division methods in actual implementation. The following is an example of dividing each functional module corresponding to each function.
[0158] Figure 9 Schematic diagram of a data processing device provided by an embodiment of the present disclosure. The data processing device 90 can execute the data processing method provided by the above method embodiment. Figure 9 As shown, the data processing device 90 includes an acquisition module 901 and a processing module 902 .
[0159] An acquisition module 901 is configured to acquire a first encoding matrix;
[0160] The processing module 902 is used to process k source data blocks based on the first coding matrix to obtain n coded data blocks; wherein the n coded data blocks include: k source data blocks and m check data blocks, k is greater than 1, n is greater than k, and k, n, are all positive integers, and m is equal to nk.
[0161] In some embodiments, the first coding matrix is determined according to the following method: determining the second coding matrix according to the number k of source data blocks, the number n of coded data blocks, and the number K of columns of the second coding matrix; wherein the dimension size of the second coding matrix is (nk)*K, K is greater than k, and K is a positive integer; based on the second coding matrix, generating the first coding matrix, the dimension size of the first coding matrix is (nk)*k.
[0162] In some embodiments, the number of columns K of the second coding matrix is greater than or equal to the first number k+N1, and less than or equal to the second number k+N2; wherein N1 is the number of coded data blocks that are not fully recovered when the first k columns of the second coding matrix are used as the coding matrix to recover m-1 coded data blocks; N2 is qn, where q is the size of the Galois field GF(q).
[0163] In some embodiments, the number of columns K of the second encoding matrix is determined according to the number k of source data blocks.
[0164] In some embodiments, a square matrix consisting of any (nk) rows and (nk) columns in the second coding matrix is reversible in the Galois field GF(q), and there is one row in the second coding matrix whose elements are all 1.
[0165] In some embodiments, the elements in the second coding matrix are the calculation results of the first element and the second element corresponding to the element in the Galois field GF(q); wherein the first element is determined based on the number of columns of the second coding matrix and the number of columns to which the element belongs; the second element is determined based on the number of columns of the second coding matrix, the number of columns to which the element belongs, and the number of rows to which the element belongs.
[0166] In some embodiments, elements of two adjacent rows from the 2nd row to the mth row in the second encoding matrix support swapping.
[0167] In some embodiments, generating a first coding matrix based on a second coding matrix includes: obtaining the first k columns of the second coding matrix as the index of the coding data block that has not been completely recovered when the coding matrix recovers m-1 coding data blocks; replacing the L column elements corresponding to the index of the coding data block that has not been completely recovered in the second coding matrix with any L column elements from the k+1th column to the Kth column in the second coding matrix, where L is a positive integer and L is less than the difference between K and k; and using the 1st to kth columns of the replaced second coding matrix as the first coding matrix.
[0168] In some embodiments, replacing L column elements corresponding to the index of the unrecovered coded data block in the second coding matrix with any L column elements from the k+1th column to the Kth column in the second coding matrix includes: replacing L column elements corresponding to the index of the unrecovered coded data block in the second coding matrix with L column elements from the k+1th column to the k+Lth column in the second coding matrix.
[0169] In some embodiments, any square matrix consisting of m rows and m columns in the first encoding matrix is reversible in the Galois Field GF(q).
[0170] In some embodiments, the first coding matrix is used to repair at least any m-1 coded data blocks.
[0171] In some embodiments, the m check data blocks include at least one of the following: 1 first check data block, m-2 second check data blocks, and 1 third check data block.
[0172] In some embodiments, the processing module 902 is specifically used to: encode k source data blocks based on the elements of the 1st row of the first coding matrix to generate a first verification data block; encode k source data blocks based on the elements of the 2nd to m-1th rows of the first coding matrix to generate a second verification data block; encode k source data blocks based on the elements of the mth row of the first coding matrix to generate a third verification data block.
[0173] In some embodiments, the syndrome data block in the first syndrome data block is obtained by linearly combining source sub-data blocks corresponding to the syndrome data block, and coefficients of the linear combination are obtained based on the first row of the first coding matrix.
[0174] In some embodiments, the check sub-data block in the second check data block is determined based on the linear combination value of the source sub-data blocks with a first source sub-data block index in the k source data blocks and the linear combination value of the source sub-data blocks with a second source sub-data block index in the k source data blocks; wherein the first index is the check sub-data block index corresponding to the c-th check sub-data block in the second check data block; the second index is any check sub-data block index in the second check data block except the first index; and c is a positive integer.
[0175] In some embodiments, the syndrome data block in the third syndrome data block is obtained by linearly combining source sub-data blocks corresponding to the syndrome data block, and coefficients of the linear combination are obtained based on the mth row of the first coding matrix.
[0176] In the case of implementing the functions of the above-mentioned integrated modules in the form of hardware, the embodiments of the present disclosure provide another possible structure of the data processing device involved in the above-mentioned embodiments. Figure 10 As shown, the data processing device 200 includes: a processor 2002 and a bus 2004. Optionally, the data processing device 200 may further include a memory 2001; and optionally, the data processing device 200 may further include a communication interface 2003.
[0177] Processor 2002 may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of this disclosure. Processor 2002 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of this disclosure. Processor 2002 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.
[0178] The communication interface 2003 is used to connect to other devices via a communication network, such as Ethernet, wireless access network, wireless local area network (WLAN), etc.
[0179] The memory 2001 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0180] As a possible implementation, the memory 2001 can exist independently of the processor 2002. The memory 2001 can be connected to the processor 2002 via a bus 2004 to store instructions or program codes. When the processor 2002 calls and executes the instructions or program codes stored in the memory 2001, the data processing method provided in the embodiments of the present disclosure can be implemented.
[0181] In another possible implementation, the memory 2001 may also be integrated with the processor 2002 .
[0182] The bus 2004 may be an extended industry standard architecture (EISA) bus, etc. The bus 2004 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0183] Some embodiments of the present disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium), which stores computer program instructions. When the computer program instructions are executed on a computer, the computer executes a data processing method as described in any of the above embodiments.
[0184] Exemplarily, the above-mentioned computer-readable storage media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes, etc.), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memories (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0185] An embodiment of the present disclosure provides a computer program product comprising instructions. When the computer program product is run on a computer, the computer is enabled to execute the data processing method described in any one of the above embodiments.
[0186] The above is only a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or replacements within the technical scope disclosed in the present disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A data processing method, characterized in that: The method comprises: Obtaining a first encoding matrix; Based on the first coding matrix, k source data blocks are processed to obtain n coded data blocks; wherein the n coded data blocks include: k source data blocks and m check data blocks, k is greater than 1, n is greater than k, and k, n, are all positive integers, and m is equal to nk.
2. The method according to claim 1, characterized in that The first encoding matrix is determined according to the following method: Determine a second encoding matrix according to the number k of source data blocks, the number n of encoding data blocks, and the number K of columns of the second encoding matrix; wherein the dimension of the second encoding matrix is (nk)*K, K is greater than k, and K is a positive integer; Based on the second encoding matrix, a first encoding matrix is generated, where the dimension of the first encoding matrix is (nk)*k.
3. The method according to claim 2, characterized in that The number of columns K of the second coding matrix is greater than or equal to the first number k+N1, and less than or equal to the second number k+N2; wherein N1 is the number of coded data blocks that are not recovered when the first k columns of the second coding matrix are used as the coding matrix to recover m-1 coded data blocks; N2 is qn, where q is the size of the Galois field GF(q).
4. The method according to claim 2, characterized in that The number of columns K of the second encoding matrix is determined according to the number k of the source data blocks.
5. The method according to claim 2, characterized in that A square matrix consisting of any (nk) rows and (nk) columns in the second coding matrix is reversible in the Galois field GF(q), and there is one row in the second coding matrix whose elements are all 1.
6. The method according to claim 5, characterized in that The elements in the second coding matrix are the calculation results of the first element and the second element corresponding to the element in the Galois field GF(q); wherein the first element is determined based on the number of columns of the second coding matrix and the number of columns to which the element belongs; the second element is determined based on the number of columns of the second coding matrix, the number of columns to which the element belongs, and the number of rows to which the element belongs.
7. The method according to claim 6, characterized in that The elements of two adjacent rows from the 2nd row to the mth row in the second encoding matrix support interchange.
8. The method according to claim 2, characterized in that The generating the first encoding matrix based on the second encoding matrix includes: Obtaining the first k columns of the second coding matrix as the index of the coding data block that has not been completely restored when restoring m-1 coding data blocks by the coding matrix; Replacing L column elements corresponding to the index of the unrecovered coded data block in the second coding matrix with any L column elements from the k+1th column to the Kth column in the second coding matrix, where L is a positive integer and L is less than the difference between K and k; The first to kth columns in the replaced second coding matrix are used as the first coding matrix.
9. The method according to claim 8, characterized in that The replacing the L column elements corresponding to the index of the unrecovered coded data block in the second coding matrix with any L column elements from the k+1th column to the Kth column in the second coding matrix includes: The L column elements corresponding to the index of the unrecovered coded data block in the second coding matrix are replaced by the L column elements from the k+1th column to the k+Lth column in the second coding matrix.
10. The method according to claim 1, characterized in that Any square matrix consisting of m rows and m columns in the first encoding matrix is reversible in the Galois field GF(q).
11. The method according to claim 1, wherein The first coding matrix is used to repair at least any m-1 coded data blocks.
12. The method according to claim 1, characterized in that The m check data blocks include at least one of the following: 1 first check data block, m-2 second check data blocks, and 1 third check data block.
13. The method according to claim 12, characterized in that The step of processing the k source data blocks based on the first encoding matrix to obtain n encoded data blocks includes at least one of the following: Encoding the k source data blocks based on the elements of the first row of the first encoding matrix to generate the first verification data block; Encode the k source data blocks based on elements from the 2nd row to the m-1th row of the first encoding matrix to generate the second verification data block; The k source data blocks are encoded based on the elements of the mth row of the first encoding matrix to generate the third verification data block.
14. The method according to claim 13, characterized in that The syndrome data block in the first syndrome data block is obtained by linearly combining source sub-data blocks corresponding to the syndrome data block, and the coefficients of the linear combination are obtained based on the first row of the first encoding matrix.
15. The method according to claim 13, characterized in that The check sub-data block in the second check data block is determined based on a linear combination value of source sub-data blocks whose source sub-data block index is a first index in the k source data blocks and a linear combination value of source sub-data blocks whose source sub-data block index is a second index in the k source data blocks; wherein the first index is the check sub-data block index corresponding to the c-th check sub-data block in the second check data block; the second index is any check sub-data block index in the second check data block except the first index; and c is a positive integer.
16. The method according to claim 13, characterized in that The syndrome data block in the third syndrome data block is obtained by linearly combining source sub-data blocks corresponding to the syndrome data block, and coefficients of the linear combination are obtained based on the mth row of the first coding matrix.
17. A data processing device, characterized in that: The method comprises a processor, wherein when the processor executes the computer program, the processor implements the data processing method according to any one of claims 1 to 16.
18. A computer-readable storage medium, characterized in that The computer-readable storage medium includes computer instructions; wherein, when the computer instructions are executed, the data processing method according to any one of claims 1 to 16 is implemented.