Data processing method and related equipment

By constructing a check matrix with low density, the problem of large amount of encoding and codec operations in erasure coding technology is solved, and efficient encoding and codec throughput is achieved.

CN120336073APending Publication Date: 2025-07-18HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410070788.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing erasure coding technology, the high density check matrix leads to large amount of encoding and codec operations, low throughput, and poor encoding and codec performance.

Method used

Construct a check matrix with a lower density. By obtaining the base element group and using the base elements to construct the check matrix, the number of XOR calculations is reduced and the codec throughput is improved.

Benefits of technology

Reduces codec overhead, improves codec throughput, and improves codec performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336073A_ABST
    Figure CN120336073A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method which comprises the steps that a base element group is obtained, the base element group comprises a plurality of base elements, the number of 1 in binary representation of the plurality of base elements is not 0 and is smaller than a first threshold value, a check matrix is constructed according to the plurality of base elements, and the check matrix is used for encoding a data block to generate a check block. In the method, the number of 1 in the binary representation of the base element is smaller than the first threshold value, and the density of the constructed check matrix is relatively low, so that the number of times of XOR calculation can be reduced in a coding process of generating a check block by using the check matrix or a decoding process of recovering incomplete data by using the check matrix, the coding and decoding overhead is reduced, and the decoding efficiency is improved. And the coding and decoding throughput rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technology, and in particular, to a data processing method, apparatus, computing device, and computer program product. Background Art

[0002] With the continuous development of computer technology, the amount of data has increased accordingly, and data storage has become particularly important. Erasure code (EC) technology can achieve reliable data storage. Specifically, erasure code technology uses a parity-check matrix (also referred to as a coding matrix) to encode multiple data blocks into at least one parity block, and stores the multiple data blocks and at least one parity block on different storage nodes. When data is damaged or lost, the damaged or lost data can be decoded and restored through the complete data in the multiple data blocks and at least one parity block.

[0003] Considering that the erasure code needs to have the maximum distance separable (MDS) property, the parity-check matrix usually has a high density. However, during the encoding or decoding process, a parity-check matrix with a high density can lead to a large amount of computation and a low throughput, thereby reducing the encoding and decoding performance. Summary of the Invention

[0004] This application provides a data processing method, which can construct a parity-check matrix with a lower density, reduce the encoding and decoding overhead, and improve the encoding and decoding throughput rate. This application also provides a data processing apparatus, a computing device, and a computer program product corresponding to the method.

[0005] In a first aspect, this application provides a data processing method, which can be executed by a data processing apparatus. The data processing apparatus can be a software apparatus, and the software apparatus can be deployed in a computing device. The computing device executes the program code of the software apparatus to execute the data processing method of this application. In some examples, the data processing apparatus can be a hardware apparatus. For example, the data processing apparatus can be a computing device with data processing functions such as constructing a parity-check matrix. When the above hardware apparatus runs, it executes the data processing method of this application.

[0006] Specifically, the data processing apparatus can obtain a group of basis elements, where the group of basis elements includes multiple basis elements, and the number of 1s in the binary representation of the multiple basis elements is not 0 and less than a first threshold. Then, the data processing apparatus can construct a parity-check matrix according to the multiple basis elements, and the parity-check matrix can be used to encode data blocks to generate parity blocks.

[0007] In this method, by obtaining a base element group and constructing a parity-check matrix using multiple base elements in the base element group, since the number of 1s in the binary representation of the base element is less than a first threshold, the constructed parity-check matrix has a low density. Thus, in the encoding process of generating parity-check blocks using the parity-check matrix or the decoding process of recovering incomplete data using the parity-check matrix, the number of XOR calculations can be reduced, the encoding and decoding overhead can be lowered, and the encoding and decoding throughput can be improved.

[0008] In some possible implementation manners, the data processing device may obtain the number of data blocks and the number of parity-check blocks, and determine a base element group corresponding to the number of data blocks and the number of parity-check blocks.

[0009] In this method, the base element group may be related to the number of data blocks and the number of parity-check blocks. For different numbers of data blocks and parity-check blocks, a base element group including different base elements is determined to meet the actual encoding and decoding requirements of users.

[0010] In some possible implementation manners, the data processing device may determine a base element group according to the number of data blocks, the number of parity-check blocks, and the field coefficient of the Galois field. Wherein, the field coefficient may indicate the number of elements included in the Galois field, and multiple base elements may be composed of the elements of the Galois field.

[0011] In this method, the field coefficient of the Galois field and the number of base elements in the base element group may show a positive variation relationship. By jointly determining the base element group in combination with the field coefficient of the Galois field, it is avoided that the base element group includes too many base elements, and the encoding and decoding performance is improved.

[0012] In some possible implementation manners, the data processing device may determine an encoding coefficient according to the number of data blocks and the number of parity-check blocks. The encoding coefficient may be used to measure the proportional relationship between the number of data blocks and the number of parity-check blocks. Then, the data processing device may determine a base element group according to the encoding coefficient and the field coefficient of the Galois field.

[0013] In this method, the encoding coefficient may be used to measure whether the data processing device needs to encode long codes or short codes. Considering that it is difficult to meet the encoding requirements for long codes when the number of base elements in the base element group is small, therefore, the base element group is determined through the encoding coefficient and the field coefficient of the Galois field to meet different encoding requirements for long codes or short codes.

[0014] In some possible implementation manners, the number of check blocks may be r. When the coding coefficient and the field coefficient of the Galois field satisfy the first condition, the data processing device may determine the first set of basis elements as the set of basis elements. When the coding coefficient and the field coefficient of the Galois field do not satisfy the first condition, the data processing device may determine the second set of basis elements as the set of basis elements. Among them, the first condition may indicate that by using the basis elements with the number of 1s in the binary representation being 1, a parity-check matrix in which any r×r submatrix is invertible can be constructed. The first set of basis elements includes multiple basis elements with the number of 1s in the binary representation being 1, and the second set of basis elements includes more basis elements than the first set of basis elements.

[0015] This method can construct a parity-check matrix by using the basis elements in the set of basis elements under different coding requirements, improve the construction efficiency of the parity-check matrix, avoid resource waste caused by too many basis elements in the set of basis elements, and at the same time, improve the flexibility and scalability of constructing the parity-check matrix.

[0016] In some possible implementation manners, for the first position of the parity-check matrix, the data processing device may use any one of the multiple basis elements as the element at the first position of the parity-check matrix, where the first position may be any position of the parity-check matrix.

[0017] This method samples multiple basis elements in the set of basis elements to determine the elements at each position of the parity-check matrix, thereby constructing a parity-check matrix with a lower density and more coefficients.

[0018] In some possible implementation manners, the number of check blocks may be r. The multiple basis elements may include 1, and the parity-check matrix may satisfy the following conditions: there is a row in the parity-check matrix where all elements are 1. Except for the row where all elements are 1, the number of identical elements in the other rows of the parity-check matrix is not greater than r - 1.

[0019] This method constructs a parity-check matrix by referring to the structure of the Vandermonde matrix. By constructing a row of elements in the parity-check matrix as 1, the number of XOR operations can be further reduced in the subsequent encoding and decoding processes, the overhead of XOR calculations can be reduced, and the throughput efficiency of encoding and decoding can be improved.

[0020] In some possible implementation manners, the number of check blocks may be r. After constructing the parity-check matrix, the data processing device may also obtain the r×r submatrix of the parity-check matrix. When any r×r submatrix is non-invertible, the data processing device may reconstruct the parity-check matrix according to the multiple basis elements in the set of basis elements.

[0021] This method ensures that the parity-check matrix has the MDS property by checking the r×r submatrix of the parity-check matrix, enabling the parity-check matrix to be used as RS coding for subsequent encoding and decoding to meet the application requirements.

[0022] In some possible implementation manners, the data processing device may be deployed in a cloud storage system to encode and decode the stored data in the cloud storage system, so as to improve the reliability of the cloud storage system.

[0023] In some possible implementation manners, the data processing device may provide basic support for the read / write capabilities of the cloud storage system, connect other modules in the cloud storage system (such as a cluster management device, a service management device, and a storage node), and play a connecting role.

[0024] In a second aspect, the present application provides a data processing device, and the device includes:

[0025] An obtaining module, configured to obtain a base element group, where the base element group includes a plurality of base elements, and the number of 1s in the binary representation of the plurality of base elements is not 0 and is less than a first threshold;

[0026] A constructing module, configured to construct a parity-check matrix according to the plurality of base elements, where the parity-check matrix is used to encode a data block to generate a check block.

[0027] In some possible implementation manners, the obtaining module is specifically configured to:

[0028] Obtain the number of the data blocks and the number of the check blocks;

[0029] Determine a base element group corresponding to the number of the data blocks and the number of the check blocks.

[0030] In some possible implementation manners, the obtaining module is specifically configured to:

[0031] Determine a base element group according to the number of the data blocks, the number of the check blocks, and the field coefficient of a Galois field, where the field coefficient indicates the number of elements included in the Galois field, and the plurality of base elements are composed of the elements of the Galois field.

[0032] In some possible implementation manners, the obtaining module is specifically configured to:

[0033] Determine an encoding coefficient according to the number of the data blocks and the number of the check blocks, where the encoding coefficient is used to measure the proportional relationship between the number of the data blocks and the number of the check blocks;

[0034] Determine a base element group according to the encoding coefficient and the field coefficient of the Galois field.

[0035] In some possible implementation manners, the number of the check blocks is r, and the obtaining module is specifically configured to:

[0036] When the coding coefficient and the field coefficient of the Galois field satisfy a first condition, determine a first set of basis elements as the set of basis elements. The first set of basis elements includes a plurality of basis elements with the number of 1s in the binary representation being 1. The first condition indicates that by using the basis elements with the number of 1s in the binary representation being 1, a parity-check matrix in which any r×r submatrix is invertible can be constructed;

[0037] When the coding coefficient and the field coefficient of the Galois field do not satisfy the first condition, determine a second set of basis elements as the set of basis elements. The number of basis elements included in the second set of basis elements is greater than the number of basis elements included in the first set of basis elements.

[0038] In some possible implementation manners, the constructing module is specifically configured to:

[0039] For a first position of the parity-check matrix, use any one of the plurality of basis elements as the element at the first position of the parity-check matrix. The first position is any position of the parity-check matrix.

[0040] In some possible implementation manners, the number of the parity-check blocks is r, the plurality of basis elements include 1, and the parity-check matrix satisfies the following conditions:

[0041] There is a row in the parity-check matrix where all elements are 1;

[0042] Except for the row where all elements are 1, the number of identical elements in the other rows of the parity-check matrix is not greater than r - 1.

[0043] In some possible implementation manners, the number of the parity-check blocks is r, and the device further includes a checking module. The checking module is configured to:

[0044] Obtain an r×r submatrix of the parity-check matrix;

[0045] When there is any non-invertible r×r submatrix, reconstruct the parity-check matrix according to the plurality of basis elements.

[0046] In a third aspect, the present application provides a computing device. The computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is configured to execute instructions stored in the at least one memory, so that the computing device executes the data processing method as described in the first aspect or any one of the implementation manners of the first aspect.

[0047] Fourth aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is configured to execute instructions stored in the at least one memory, so that the computing device cluster executes the data processing method as described in the first aspect or any implementation manner of the first aspect.

[0048] Fifth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions direct a computing device or a computing device cluster to execute the data processing method as described in the first aspect or any implementation manner of the first aspect.

[0049] Sixth aspect, the present application provides a computer program product including instructions, which, when running on a computing device or a computing device cluster, cause the computing device or the computing device cluster to execute the data processing method as described in the first aspect or any implementation manner of the first aspect.

[0050] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below.

[0052] Figure 1 It is a schematic diagram of the architecture of a data processing device provided by an embodiment of the present application;

[0053] Figure 2 It is a schematic diagram of an application scenario of a cloud storage system provided by an embodiment of the present application;

[0054] Figure 3 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0055] Figure 4 It is a schematic flowchart of a process for obtaining a base element group provided by an embodiment of the present application;

[0056] Figure 5 It is a schematic flowchart of a process for constructing a parity check matrix provided by an embodiment of the present application;

[0057] Figures 6A to 6E It is a schematic diagram of a parity check matrix provided by an embodiment of the present application;

[0058] Figure 7 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0059] Figure 8 This is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application. Detailed implementation manners

[0060] The terms "first" and "second" in the embodiments of the present application are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined by "first" and "second" may explicitly or implicitly include one or more of such features.

[0061] First, some technical terms related to the present application are introduced.

[0062] The erasure code (EC) technology is a data protection method. Specifically, the erasure code technology divides the original data into multiple data blocks, and encodes the multiple data blocks using a parity-check matrix (also referred to as an encoding matrix) to generate at least one parity block (also referred to as a redundant data block), thereby realizing data redundancy. The erasure code technology can be applied to a distributed storage system. After generating at least one parity block, the multiple data blocks and at least one parity block can be stored on different storage nodes in the distributed storage system. In this way, when the original data is damaged or lost, the damaged or lost data can be decoded and recovered through the complete data in the multiple data blocks and at least one parity block, thereby realizing data protection.

[0063] Generally, the erasure code needs to have the maximum distance separable (MDS) property. For k data blocks and r parity blocks generated using a parity-check matrix, if any k blocks among the k data blocks and r parity blocks can be decoded to recover the original data, then the parity-check matrix has the MDS property. Among them, k can be a natural number greater than 1, and r can be a natural number greater than 0.

[0064] The Reed-Solomon (RS) code is an erasure code with the MDS property. The RS encoding corresponding to k data blocks and r parity blocks can be denoted as RS(k, r). In RS(k, r), the relevant operations in the encoding process can be performed in the Galois field GF(2 w ), where w can be a natural number greater than 0. In some examples, considering that in GF(2 8 ), each element in RS(k, r) can be represented by 8 bits (i.e., 1 byte), so encoding in GF(2 8 ) can improve the operation efficiency of the computing device.

[0065] A Galois field, also known as a finite field, refers to a field that contains a finite number of elements. A Galois field can be represented as GF(q), where q is the order of the Galois field, and the order of the Galois field is equal to the number of elements in the Galois field. For example, GF(2 w ) contains 2 w elements.

[0066] In RS coding, common parity-check matrices can include Vandermonde matrices and Cauchy matrices. The first row of the Vandermonde matrix is all 1s. Among the positions other than the first row, the element in the i-th row and j-th column is defined as The element in the i-th row and j-th column of the Cauchy matrix is defined as A i,j = 1 / (x i + y j ), where x and y are two sets with distinct elements.

[0067] The process of encoding multiple data blocks using a parity-check matrix can be understood as a process of matrix multiplication between the parity-check matrix and a matrix composed of multiple data blocks. For k data blocks and r parity-check blocks, the parity-check matrix can be an r×k matrix. At this time, the process of generating parity-check blocks can be:

[0068]

[0069] where A i,j represents the elements in the parity-check matrix, D j represents the data blocks, P i represents the parity-check blocks, and there is

[0070] During the encoding process, the elements in the parity-check matrix can be multiplied with the data blocks over GF(2 w ), and then the multiple multiplication results are added over GF(2 w ) to generate the parity-check blocks. Taking encoding over GF(2 8 ) as an example for illustration, the elements corresponding to each element in the parity-check matrix over GF(2 8 ) can be represented as an 8×8 bit matrix over GF(2). Among them, the elements on the main diagonal of the bit matrix are 0 or 1, and the elements in other positions are all 0. For example, when the element in the parity-check matrix is 7, the bit matrix over GF(2) can be:

[0071]

[0072] Since the addition over GF(2) can be understood as an exclusive OR operation, the sparsity of the parity-check matrix can affect the encoding and decoding performance. To improve the encoding and decoding performance, the industry usually uses algorithms to reduce the density of the parity-check matrix.

[0073] The genetic algorithm (GA) can search for Cauchy matrices with low density by simulating the laws of natural evolution. Specifically, a batch of (x, y) corresponding to Cauchy matrices are randomly initialized as feasible solutions first. Then, the number of 1s in the binary representation of the elements in the Cauchy matrix corresponding to the feasible solutions (x, y) is calculated. Next, two feasible solutions are selected as parent chromosomes and crossed. During the crossing process, elements common to the parent chromosomes are preferentially selected, and then other elements in the parent chromosomes are randomly selected to obtain the offspring chromosomes. Then, the offspring chromosomes are mutated by changing one of the elements to a random element, and the number of 1s in the binary representation of the elements in the mutated offspring chromosomes is calculated. This process is repeated until the number of cycles is reached.

[0074] Through the above genetic algorithm, better (x, y) can be searched, thereby generating Cauchy matrices with lower density. However, the genetic algorithm has randomness, resulting in poor search effects. Moreover, the search time is long and the search efficiency is low. At the same time, the above method can only search for (x, y) of Cauchy matrices, with a small search space, making it difficult to further search for a parity-check matrix with lower density and MDS properties.

[0075] In some other methods for reducing the density of parity-check matrices using algorithms, Cauchy-like matrices can be used to replace Cauchy matrices. Different from Cauchy matrices which are determined by (x, y), Cauchy-like matrices are determined by (x, y, r, s). The element in the i-th row and j-th column of the Cauchy-like matrix is defined as A i,j =(r i ×s j ) / (x i +y j ). Further, the search for (x, y, r, s) of the Cauchy-like matrix is transformed into a mixed-integer programming problem. By solving (x, y, r, s) of the Cauchy-like matrix, the number of 1s in the binary representation of the elements in the generated Cauchy-like matrix is minimized.

[0076] In the above method, although the Cauchy-like matrix has two more search dimensions r and s compared to the Cauchy matrix, it is still restricted by the Cauchy matrix, with a small search space, making it difficult to search for a parity-check matrix with lower density and MDS properties.

[0077] In view of this, the present application provides a data processing method. This method can be executed by a data processing device. Among them, the data processing device can be a software device, and this software device can be deployed in a computing device cluster. The computing device cluster executes the program code of the software device, thereby executing the data processing method of the present application. In some examples, the data processing device can be a hardware device. For example, the data processing device can be a computing device cluster with data processing functions such as constructing a parity check matrix. When the above hardware device runs, it executes the data processing method of the present application.

[0078] Specifically, the data processing device can obtain a set of basis elements. Among them, the set of basis elements includes multiple basis elements, and the number of 1s in the binary representation of the multiple basis elements is not 0 and less than a first threshold. Then, the data processing device can construct a parity check matrix according to the multiple basis elements, and this parity check matrix can be used to encode a data block to generate a parity block.

[0079] In this method, by obtaining a set of basis elements and constructing a parity check matrix using the multiple basis elements in the set of basis elements, since the number of 1s in the binary representation of the basis elements is less than the first threshold, the constructed parity check matrix has a low density. Furthermore, in the encoding process of generating a parity block using the parity check matrix or the decoding process of recovering incomplete data using the parity check matrix, the number of XOR calculations can be reduced, the encoding and decoding overhead can be reduced, and the encoding and decoding throughput can be improved.

[0080] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the present application will be introduced below with reference to the accompanying drawings.

[0081] See Figure 1 In the architecture diagram of the data processing device shown, the data processing device 10 can be deployed in the cloud storage system 100 to encode and decode the stored data in the cloud storage system 100 to improve the reliability of the cloud storage system.

[0082] The data processing device 10 can include an acquisition module 101 and a construction module 102. Specifically, the acquisition module 101 is used to obtain a set of basis elements. Among them, the set of basis elements includes multiple basis elements, and the number of 1s in the binary representation of the multiple basis elements is not 0 and less than a first threshold. In some possible implementation manners, the number of data blocks and the number of parity blocks can be obtained first, and then the set of basis elements corresponding to the number of data blocks and the number of parity blocks can be determined. In this way, the set of basis elements is determined according to the actual encoding requirements.

[0083] The construction module 102 is used to construct a check matrix according to a plurality of basis elements. In a specific implementation, for the first position of the check matrix, the construction module 102 can use any one of the plurality of basis elements as the element at the first position of the check matrix, where the first position can be any position of the check matrix. In this way, a check matrix is constructed using the plurality of basis elements included in the basis element group, so that the number of 1s in the binary representation of the elements of the check matrix is small.

[0084] The check matrix can be used to encode data blocks to generate check blocks. Among them, the check blocks can be used to recover the incomplete data in the data blocks when the data blocks are incomplete. For example, when the number of data blocks is k and the number of check blocks is r, the check matrix can be an r×k matrix. By performing matrix multiplication on the check matrix and the matrix composed of k data blocks, r check blocks can be generated. Since the number of 1s in the binary representation of the elements in the check matrix is small, the number of XOR operations is small during the encoding process, thereby improving the throughput efficiency of encoding.

[0085] In some embodiments, the data processing device 10 may further include a verification module 103. Specifically, after the check matrix is constructed, the verification module 103 is used to obtain the r×r submatrix of the check matrix, and when any r×r submatrix is irreversible, the check matrix is reconstructed according to a plurality of basis elements. In this way, it is ensured that the check matrix satisfies the MDS property.

[0086] As Figure 1 shown, the cloud storage system 100 may further include a scheduling device 20, a cluster management device 30, a service management device 40, and a storage node 50, which will be introduced separately below.

[0087] The scheduling device 20 is used to schedule read / write tasks in the cloud storage system 100. The scheduling device 20 can manage the task queue corresponding to the read / write tasks using a scheduler, or can improve the scheduling performance using a connection pool. The cluster management device 30 is used to manage the clusters in the cloud storage system 100. Specifically, cluster management can involve resource management, device management, node management (such as adding nodes, replacing nodes, etc.), shared lease management, data migration management, and storage space management. The service management device 40 is used to manage the storage services of the cloud storage system 100. Specifically, service management can involve protocol management, cache management, file system management, service quality management, metadata management, and index management. The storage node 50 is used to store the data in the cloud storage system 100. Specifically, the storage node 50 can include a solid state disk (SSD) and a hard disk drive (HDD).

[0088] Based on the above description, in the cloud storage system 100, the data processing device 10 provides fundamental support for the read / write capabilities of the cloud storage system 100, connects the cluster management device 30, the service management device 40, and the storage nodes 50, playing a bridging role. Moreover, during the encoding and decoding process, the encoding and decoding performance is improved, enhancing the performance of the cloud storage system 100.

[0089] The application scenarios of the above cloud storage system 100 will be described below with reference to the accompanying drawings.

[0090] Referring to Figure 2 the schematic diagram of the application scenario of the cloud storage system shown, the cloud storage system 100 can be a distributed storage system. The cloud storage system 100 includes multiple storage servers, for example, including multiple of the above storage nodes 50. The storage nodes 50 can be deployed in a cloud environment, such as in a public cloud, a private cloud, or a hybrid cloud.

[0091] The cloud storage system 100 can be connected to the terminal 200 or the server 300 through a communication network, such as through Ethernet or InfiniBand. The terminal 200 or the server 300 can send the data to be stored (i.e., the original data) to the cloud storage system 100. The cloud storage system 100 receives the original data and performs erasure code-based encoding on the original data. After encoding, the cloud storage system 100 writes the data blocks and parity blocks into multiple storage nodes 50 to achieve distributed data storage.

[0092] Next, the cases where the data storage requesters are the terminal 200 and the server 300 will be described separately. In some embodiments, the terminal 200 can be a device operated by a user, and the user can be a user of an object storage service (OBS), a cloud hard disk, or a cloud database. When the user has a data storage requirement, the user can send the original data to the cloud storage system 100 through the terminal 200, and the cloud storage system 100 processes the original data through encoding and other operations to achieve data storage.

[0093] In other embodiments, the server 300 can be an application server. For example, the server 300 can be a video application server. During the operation of a video application, a large amount of data (such as operation logs, operation records, etc.) can be generated. The video application server can send the above data as the original data to the cloud storage system 100, and the cloud storage system 100 processes the original data through encoding and other operations to achieve the storage of big data.

[0094] For another example, the server 300 may be an enterprise application server. The enterprise application server, also known as the enterprise data center, is used to manage various types of enterprise data. The enterprise application server may send the enterprise data as the original data to the cloud storage system 100, and the cloud storage system 100 performs processing such as encoding on the original data to implement the storage of the enterprise data and meet the backup and archiving requirements of enterprise applications.

[0095] Based on the above description, it can be known that the cloud storage system can store different types of data such as hot data and low-frequency access data. By deploying the data processing device provided in this application in the cloud storage system, the encoding and decoding performance can be improved during the encoding and decoding process of the original data, and the security and stability of the data can be guaranteed.

[0096] Based on Figure 1 the data processing device 10 shown, this application also provides a data processing method. The data processing method of this application will be introduced below in conjunction with embodiments.

[0097] See Figure 3 the flowchart of the data processing method shown. The method includes the following steps:

[0098] S301: The data processing device 10 obtains a base element group.

[0099] The base element group includes multiple base elements. In the embodiments of this application, a base element can be understood as the smallest unit for constructing a parity-check matrix. In other words, during the subsequent process of constructing the parity-check matrix, the data processing device 10 can use multiple base elements to construct the parity-check matrix.

[0100] The number of 1s in the binary representation of a base element can be understood as the order of the base element. For example, in GF(2 8 ), when the base element is 1, the binary representation is 00000001, and the number of 1s is 1. Therefore, the base element 1 can be called a first-order base element. For another example, in GF(2 8 ), when the base element is 7, the binary representation is 00000111, and the number of 1s is 3. Therefore, the base element 7 can be called a third-order base element.

[0101] In the embodiments of this application, the number of 1s in the binary representations of multiple base elements is not 0 and less than a first threshold. Considering that a base element can be an element in GF(2 w ), and the elements in GF(2 w ) usually do not include 0, therefore, the number of 1s in the binary representation of the base element in the embodiments of this application is not 0.

[0102] The first threshold can be set according to actual requirements. In some embodiments, the first threshold can be 3. In other words, the number of 1s in the binary representation of multiple basis elements can be 1 or 2, and the multiple basis elements can be first-order basis elements or second-order basis elements. From the above description, it can be known that the basis elements in the embodiments of the present application can be understood as power-of-2 elements and elements obtained by adding powers-of-2 with a number less than the first threshold (which can also be called quasi power-of-2).

[0103] Taking the first threshold being 3 as an example for illustration, in GF(2 w ), the basis elements with the number of 1s in the binary representation being 1 can include 1, 2, 4, ……, 2 w-1 , and the basis elements with the number of 1s in the binary representation being 2 can include 1, 2, 4, ……, 2 w-1 , 2 a +2 b , where both a and b are natural numbers, and a < w, b < w.

[0104] Considering that the fewer the number of basis elements included in the basis element group, the higher the subsequent construction efficiency of the parity check matrix. In some possible implementation manners, the basis element group can be related to the number of data blocks and the number of parity check blocks, so as to avoid including too many basis elements in the basis element group.

[0105] In specific implementation, the data processing device 10 can first obtain the number of data blocks and the number of parity check blocks, and determine the basis element group corresponding to the number of data blocks and the number of parity check blocks.

[0106] Among them, the number of data blocks usually refers to the number of divisions when dividing the original data, and the number of parity check blocks usually refers to the number of redundant data blocks to be generated. The number of data blocks and the number of parity check blocks can be determined according to the actual requirements of the user. For example, when the amount of original data to be stored by the user is large, the number of data blocks can be large.

[0107] When the user uses the cloud storage system to store the original data, the cloud storage system can recommend the number of data blocks and the number of parity check blocks to the user. For example, for the original data with a small amount of data, the number of data blocks can be 8, and the number of parity check blocks can be 4. Another example is that for the original data with a large amount of data, the number of data blocks can be 128, and the number of parity check blocks can be 4. The number of data blocks and the number of parity check blocks can also be other numbers, and the embodiments of the present application do not limit this.

[0108] By obtaining the number of data blocks and the number of parity check blocks, the data processing device 10 can determine the corresponding basis element group. In this way, for different numbers of data blocks and parity check blocks, a basis element group including different basis elements is determined, improving the subsequent efficiency of constructing the parity check matrix.

[0109] In some embodiments, the data processing apparatus 10 may determine a set of basis elements according to the number of data blocks, the number of check blocks, and the field coefficients of the Galois field. Among them, the field coefficients of the Galois field may indicate the number of elements included in the Galois field. In other words, the field coefficients of the Galois field may be related to the order of the Galois field. For example, when the Galois field is GF(2 w ), the field coefficients of the Galois field may be w, or w - 1, or other coefficient expressions that can indicate the number of elements included in the Galois field. The embodiments of the present application do not limit this.

[0110] In the embodiments of the present application, multiple basis elements may be composed of elements of the Galois field. It can be understood that when encoding on GF(2 w ), the elements in the parity-check matrix are elements on GF(2 w ). Therefore, the basis elements in the set of basis elements may also be elements on GF(2 w ). For example, when encoding on GF(2 8 ), multiple basis elements may be elements from 0 to 2 8 - 1 in which the number of 1s in the binary representation is not 0 and less than the first threshold. Therefore, the field coefficients of the Galois field and the number of basis elements in the set of basis elements show a positive variation relationship, that is, the larger the field coefficients of the Galois field, the more basis elements included in the set of basis elements, and the smaller the field coefficients of the Galois field, the fewer basis elements included in the set of basis elements.

[0111] Referring to Figure 4 the schematic flowchart of a process for obtaining a set of basis elements shown, the data processing apparatus 10 may determine encoding coefficients according to the number of data blocks and the number of check blocks, and thus determine the set of basis elements according to the encoding coefficients and the field coefficients of the Galois field.

[0112] Among them, the encoding coefficients may be used to measure the proportional relationship between the number of data blocks and the number of check blocks. For example, when the number of data blocks is k and the number of check blocks is r, the encoding coefficients may be k / r, or k / (r - 1), or other coefficient expressions that can measure the proportional relationship between k and r. The embodiments of the present application do not limit this.

[0113] Further, since the check blocks can be generated by matrix multiplication of the parity-check matrix and the matrix composed of data blocks, the size of the parity-check matrix is related to the number of data blocks and the number of check blocks. Specifically, when the number of data blocks is k and the number of check blocks is r, the parity-check matrix may be an r×k matrix.

[0114] It can be understood that when the coding coefficient is large, it indicates that the difference between the number of data blocks and the number of parity blocks is large. When the coding coefficient is small, it indicates that the difference between the number of data blocks and the number of parity blocks is small. Therefore, when the coding is large, it can be understood that the data processing device 10 needs to perform coding for long codes subsequently. When the coding coefficient is small, it can be understood that the data processing device 10 needs to perform coding for short codes subsequently.

[0115] Based on the above description, it can be known that the coding coefficient can be used to measure whether the data processing device 10 needs to perform coding for long codes or short codes. The field coefficient of the Galois field can be used to indicate the number of basis elements included in the basis element group. When coding for long codes is required, but the number of basis elements in the basis element group is small, it is difficult for the number of basis elements in the basis element group to meet the requirements of the coding code length. It is difficult for the data processing device 10 to construct a parity check matrix that satisfies the MDS property using the limited basis elements in the basis element group. Therefore, the data processing device 10 can determine the specific basis elements included in the basis element group according to the coding coefficient and the field coefficient of the Galois field.

[0116] In specific implementation, when the coding coefficient and the field coefficient of the Galois field satisfy the first condition, the data processing device 10 can determine the first basis element group as the basis element group. When the coding coefficient and the field coefficient of the Galois field do not satisfy the first condition, the data processing device 10 can determine the second basis element group as the basis element group. Among them, the first basis element group includes a plurality of basis elements with the number of 1s in the binary representation being 1, and the second basis element group includes more basis elements than the first basis element group. For example, the second basis element group can include a plurality of elements with the number of 1s in the binary representation being 1, and a plurality of elements with the number of 1s in the binary representation being 2.

[0117] That is to say, when the coding coefficient and the field coefficient of the Galois field satisfy the first condition, the basis element group can only include power-of-2 elements. Otherwise, the basis element group can include elements with a number greater than the number of power-of-2 elements, such as power-of-2 elements and quasi-power-of-2 elements. Taking GF(2 8 ) as an example, when the coding coefficient and the field coefficient of the Galois field satisfy the first condition, the basis elements can be 1, 2, 4, 8, 16, 32, 64, and 128. When the coding coefficient and the field coefficient of the Galois field do not satisfy the first condition, the basis elements in the basis element group can be 1, 2, 4, 8, 16, 32, 64, 128, and 2 c +2 d , where c and d are both natural numbers, and c < 8, d < 8.

[0118] The first condition can indicate that a basis element with the number of 1s in its binary representation being 1 can be used to construct a parity-check matrix in which any r×r submatrix is invertible. It can be understood that the invertibility of any r×r submatrix of the parity-check matrix indicates that the parity-check matrix has the MDS property. Therefore, when the coding coefficients and the field coefficients of the Galois field satisfy the first condition, indicating that the basis elements in the basis element group only include power-of-2 elements, the requirements for the coding length can be met. The data processing device 10 can construct a parity-check matrix with the MDS property only using power-of-2 basis elements. At this time, the number of 1s in the binary representation of multiple basis elements is 1. By reducing the number of basis elements in the basis element group, the construction efficiency of the subsequent parity-check matrix can be improved.

[0119] In some embodiments, the coding coefficient can be k / (r - 1), the field coefficient of the Galois field can be w - 1, and the first condition can be k / (r - 1) < w - 1. To facilitate determining whether the coding coefficient and the field coefficient of the Galois field satisfy the first condition, after taking the integer part of the coding coefficient, the size relationship can be compared with the field coefficient. The first condition can also have other forms, and the embodiments of the present application do not limit this.

[0120] Based on the number of data blocks and the number of parity-check blocks, the data processing device 10 can determine the corresponding basis element group. In this way, subsequently, the data processing device 10 can construct a parity-check matrix using multiple basis elements under different coding requirements, improving the construction efficiency of the parity-check matrix, avoiding resource waste caused by an excessive number of basis elements in the basis element group, and at the same time, improving the flexibility and scalability of constructing the parity-check matrix.

[0121] S302: The data processing device 10 constructs a parity-check matrix according to multiple basis elements.

[0122] After obtaining the basis element group, the data processing device 10 can use multiple basis elements to construct a parity-check matrix, which can be used to encode data blocks to generate parity-check blocks. For example, matrix multiplication can be performed on the parity-check matrix and the matrix composed of data blocks to generate parity-check blocks. In this way, when the data blocks are incomplete, the incomplete data in the data blocks can be restored using the parity-check matrix. For example, decoding is performed using the parity-check matrix, the complete data in the data blocks, and the complete parity-checks in the parity-check blocks to restore the incomplete data and ensure data reliability.

[0123] During the process of constructing the parity-check matrix, for the first position of the parity-check matrix, the data processing device 10 can use any one of the multiple basis elements as the element at the first position of the parity-check matrix. The first position can be any position of the parity-check matrix.

[0124] In the embodiments of the present application, the data processing device 10 may determine the size of the parity check matrix according to the number of data blocks and the number of parity check blocks. For example, when the number of data blocks is k and the number of parity check blocks is r, the size of the parity check matrix is r×k. Further, the data processing device 10 may traverse each position of the parity check matrix. For example, the data processing device 10 may traverse each position of the parity check matrix in the order from left to right and from top to bottom. The first position may be the current access position of the data processing device 10, such as the first row and the first column, the first row and the second column. When all the positions in the first row are accessed, the access position may become the second row until all the positions of the parity check matrix are traversed. The data processing device 10 may also traverse each position of the parity check matrix in a different access order, and the embodiments of the present application do not limit this.

[0125] For any position of the parity check matrix, the data processing device 10 may select one basis element from multiple basis elements as the element at the first position of the parity check matrix. For example, the multiple basis elements are 1, 2, 4, 8, 16, 32, 64, and 128, and the first position of the parity check matrix is the second row and the third column. The data processing device 10 may select 8 as the element at the second row and the third column in the parity check matrix.

[0126] In specific implementation, the data processing device 10 may sample the multiple basis elements to determine the element at the first position of the parity check matrix. The embodiments of the present application do not limit the sampling method. The data processing device 10 may use different sampling methods such as the Monte Carlo algorithm, the genetic algorithm, and the simulated annealing algorithm for sampling.

[0127] For any position of the parity check matrix, the data processing device 10 may construct the parity check matrix by using multiple basis elements. In this way, the number of 1s in the binary representation of the elements in the parity check matrix is not 0 and is less than the first threshold, and the constructed parity check matrix has a low density and is relatively sparse.

[0128] In some possible implementation manners, considering that the Vandermonde matrix has the characteristic that the first row is all 1s, the data processing device 10 may construct the parity check matrix by referring to the structure of the Vandermonde matrix. Refer to Figure 5 As shown in the schematic flowchart of a process for constructing a parity check matrix, after obtaining the basis element group, the data processing device 10 may construct one row of elements of the parity check matrix as 1s. For example, the data processing device 10 may construct the elements in the first row of the parity check matrix as 1s. Then, for the first position of the other rows of the parity check matrix except the row that is all 1s, the data processing device 10 may determine whether the number of identical elements in the other row elements is not greater than r - 1, and when the number of identical elements in the other row elements is not greater than r - 1, use any one of the multiple basis elements as the element at the first position of the parity check matrix, and further construct the parity check matrix.

[0129] It can be understood that when the number of identical elements in other row elements is greater than r - 1, since there is a row in the parity-check matrix where all elements are 1, the row of elements all being 1 is linearly dependent on the row of elements with the number of identical elements greater than r - 1. At this time, the parity-check matrix does not have the MDS property. Therefore, when the number of identical elements in other row elements is greater than r - 1, the data processing device 10 can end the construction process.

[0130] In some examples, the constructed parity-check matrix can be expressed as:

[0131]

[0132] By constructing a row of elements in the parity-check matrix as 1, the number of XOR operations can be further reduced in subsequent encoding and decoding processes, the overhead of XOR calculations can be reduced, and the throughput efficiency of encoding and decoding can be improved.

[0133] Furthermore, after the data processing device 10 constructs the parity-check matrix, the invertibility of the r-order submatrix of the parity-check matrix can also be verified. In specific implementation, the data processing device 10 can obtain the r×r submatrix of the parity-check matrix. When any r×r submatrix is non-invertible, the data processing device 10 can reconstruct the parity-check matrix according to multiple basis elements.

[0134] In this way, by verifying the r×r submatrix of the parity-check matrix, it is ensured that the parity-check matrix has the MDS property, enabling the parity-check matrix to be used as RS coding for subsequent encoding and decoding to meet application requirements.

[0135] In the embodiments of the present application, the number of 1s in the binary representation of the elements in the parity-check matrix is small, and the parity-check matrix has the property of low density. Therefore, a simple encoding and decoding process can be achieved, the complexity of encoding and decoding can be reduced, the number of XOR operations in the calculation process can be reduced, and a high encoding and decoding benefit can be obtained. Below, the encoding and decoding performance of the parity-check matrix constructed by using the data processing method provided in the embodiments of the present application will be introduced in combination with specific examples.

[0136] In some examples, encoding and decoding are performed over GF(2 8 ), the number of data blocks is 8, the number of parity blocks is 3, and the constructed parity-check matrix is as follows:

[0137]

[0138] Compared with encoding and decoding using other parity-check matrices, the comparison results of encoding and decoding performance are shown in Table 1.

[0139] As can be seen from Table 1, since the density of the parity-check matrix in the embodiments of the present application is relatively low, the number of XOR calculations during the encoding and decoding processes is small, the encoding and decoding processes are simple, and the computational overhead can be effectively saved. Compared with Cauchy matrices, Vandermonde matrices, and optimized Cauchy matrices, the encoding performance has been significantly improved. The encoding speed has increased by 50.93% compared with Cauchy matrices, by 25.46% compared with Vandermonde matrices, by 20.93% compared with Cauchy matrices optimized using genetic algorithms, and by 17.95% compared with Cauchy-like matrices.

[0140] Table 1

[0141] Parity-check matrix Encoding speed (MB / S) Decoding speed (MB / S) Cauchy matrix 10136.01 10536.30 Cauchy-like matrix 13962.04 11765.98 Vandermonde matrix 12628.21 10912.61 Cauchy matrix optimized by genetic algorithm 13045.91 11650.32 Matrix with all elements being prime numbers 13300.71 11424.63 Parity-check matrix in the embodiments of the present application 15863.60 11209.36

[0142] In some other examples, encoding and decoding are performed over GF(2 8 ), the number of data blocks is 100, the number of parity-check blocks is 4, and the constructed parity-check matrix is as shown in Figures 6A to 6E . Among them, Figure 6A are the elements of the 1st to 20th columns of the parity-check matrix, Figure 6B are the elements of the 21st to 40th columns of the parity-check matrix, Figure 6C are the elements of the 41st to 60th columns of the parity-check matrix, Figure 6D are the elements of the 61st to 80th columns of the parity-check matrix, Figure 6E are the elements of the 81st to 100th columns of the parity-check matrix.

[0143] Compared with encoding and decoding using other parity-check matrices, when encoding and decoding are performed in Framework A, the comparison results of the encoding and decoding performance are shown in Table 2.

[0144] Table 2

[0145] Parity-check matrix Encoding speed (MB / S) Decoding speed (MB / S) Cauchy matrix 4404.53 4642.25 Vandermonde matrix 5551.21 5271.23 Parity-check matrix in the embodiments of the present application 6798.71 5391.25

[0146] When encoding and decoding are performed in Framework B, the comparison results of the encoding and decoding performance are shown in Table 3.

[0147] Table 3

[0148] Parity-check matrix Encoding speed (MB / S) Decoding speed (MB / S) Cauchy matrix 10484.02 10924.47 Vandermonde matrix 11630.97 11036.31 Parity-check matrix in the embodiments of the present application 13308.17 11266.97

[0149] As can be seen from Table 2 and Table 3, by using the data processing method in the embodiments of the present application, a parity-check matrix with a relatively low density can be constructed for long codes. When encoding and decoding are performed under different frameworks, the encoding and decoding processes are simple, and the computational overhead can be effectively saved. Compared with Cauchy matrices and Vandermonde matrices, the encoding performance has been significantly improved. In Framework A, the encoding speed has increased by 54.36% compared with Cauchy matrices and by 22.48% compared with Vandermonde matrices. In Framework B, the encoding speed has increased by 26.93% compared with Cauchy matrices.

[0150] Based on the data processing method of the foregoing embodiments, an embodiment of the present application further provides a data processing device 10 as described above. The data processing device 10 will be introduced below with reference to the accompanying drawings.

[0151] See Figure 1 The structural schematic diagram of the data processing device 10 shown. The data processing device 10 includes:

[0152] An acquisition module 101, configured to acquire a base element group, where the base element group includes a plurality of base elements, and the number of 1s in the binary representation of the plurality of base elements is not 0 and less than a first threshold;

[0153] A construction module 102, configured to construct a parity check matrix according to the plurality of base elements, where the parity check matrix is used to encode a data block to generate a check block.

[0154] In some possible implementation manners, the acquisition module 101 is specifically configured to:

[0155] Acquire the number of data blocks and the number of check blocks;

[0156] Determine a base element group corresponding to the number of data blocks and the number of check blocks.

[0157] In some possible implementation manners, the acquisition module 101 is specifically configured to:

[0158] Determine a base element group according to the number of data blocks, the number of check blocks, and the field coefficient of the Galois field, where the field coefficient indicates the number of elements included in the Galois field, and the plurality of base elements are composed of the elements of the Galois field.

[0159] In some possible implementation manners, the acquisition module 101 is specifically configured to:

[0160] Determine an encoding coefficient according to the number of data blocks and the number of check blocks, where the encoding coefficient is used to measure the proportional relationship between the number of data blocks and the number of check blocks;

[0161] Determine a base element group according to the encoding coefficient and the field coefficient of the Galois field.

[0162] In some possible implementation manners, the number of check blocks is r, and the acquisition module 101 is specifically configured to:

[0163] When the encoding coefficient and the field coefficient of the Galois field satisfy a first condition, determine the first base element group as the base element group, where the first base element group includes a plurality of base elements with the number of 1s in the binary representation being 1, and the first condition indicates that a parity check matrix in which any r×r submatrix is invertible can be constructed by using the base elements with the number of 1s in the binary representation being 1;

[0164] When the coding coefficient does not satisfy the first condition with the field coefficient of the Galois field, the second set of basis elements is determined as the set of basis elements, where the number of basis elements included in the second set of basis elements is greater than the number of basis elements included in the first set of basis elements.

[0165] In some possible implementation manners, the constructing module 102 is specifically configured to:

[0166] For the first position of the parity-check matrix, any one of the multiple basis elements is used as the element at the first position of the parity-check matrix, where the first position is any position of the parity-check matrix.

[0167] In some possible implementation manners, the number of parity-check blocks is r, the multiple basis elements include 1, and the parity-check matrix satisfies the following conditions:

[0168] There is a row in the parity-check matrix where all elements are 1;

[0169] Except for the row where all elements are 1, the number of identical elements in the other rows of the parity-check matrix is not greater than r - 1.

[0170] In some possible implementation manners, the number of parity-check blocks is r, and the apparatus 10 further includes a checking module 103, and the checking module 103 is configured to:

[0171] Obtain an r×r submatrix of the parity-check matrix;

[0172] When there is any non-invertible r×r submatrix, reconstruct the parity-check matrix according to the multiple basis elements.

[0173] This application further provides a computing device. As Figure 7 shown, the computing device 700 includes a processor 710, a memory 720, a communication interface 730, and a bus 740. The processor 710, the memory 720, and the communication interface 730 communicate with each other through the bus 740. The bus 740 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 730 is used to communicate with the outside, for example, to receive data processing tasks.

[0174] Among them, the processor 710 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits. The processor 710 may also be an integrated circuit chip with signal processing capabilities. In the implementation process, the functions of each module in the data processing device can be completed by the integrated logic circuit in the hardware of the processor 710 or instructions in the form of software. The processor 710 may also be a general-purpose processor, a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 720, and the processor 710 reads the information in the memory 720 and combines its hardware to complete some or all of the functions in the data processing device.

[0175] The memory 720 may include a volatile memory, such as a random access memory (RAM). The memory 720 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, an HDD, or an SSD.

[0176] The memory 720 stores executable code, and the processor 710 executes the executable code to execute the method executed by the foregoing data processing device.

[0177] Specifically, in the case of Figure 1 implementing the acquisition module 101, the construction module 102, and the verification module 103 described in the embodiment shown Figure 3 in the embodiment shown, and Figure 1When the acquisition module 101, the construction module 102, and the verification module 103 described in the illustrated embodiment are implemented by software, the acquisition module 101, the construction module 102, and the verification module 103 execute Figure 3 The software or program code required for the functions of the steps in the illustrated embodiment is stored in the memory 720. The interaction between the acquisition module 101 and other devices is implemented through the communication interface 730. The processor 710 is used to execute the instructions in the memory 720 to implement the method executed by the data processing device.

[0178] Figure 8 A schematic structural diagram of a computing device cluster is shown. Among them, Figure 8 The illustrated computing device cluster 80 includes multiple computing devices 700. Each computing device 700 includes a processor 710, a memory 720, a communication interface 730, and a bus 740. Among them, the processor 710, the memory 720, and the communication interface 730 are communicatively connected to each other through the bus 740. The above data processing device can be distributedly deployed on multiple computing devices 700 in the computing device cluster 80.

[0179] The processor 710 can be a CPU, a GPU, an ASIC, or one or more integrated circuits. The processor 710 can also be an integrated circuit chip with signal processing capabilities. During implementation, some functions of the data processing device can be completed by the integrated logic circuit in the processor 710 or instructions in software form. The processor 710 can also be a DSP, an FPGA, a general-purpose processor, other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute some of the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly implemented by the hardware decoding processor, or implemented by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 720. In each computing device 700, the processor 710 reads the information in the memory 720 and can complete some functions of the data processing device in combination with its hardware.

[0180] The memory 720 may include a ROM, a RAM, a static storage device, a dynamic storage device, a hard disk (such as an SSD, an HDD), etc. The memory 720 may store program codes, for example, partial or all program codes for implementing the acquisition module 101, partial or all program codes for implementing the construction module 102, and partial or all program codes for implementing the verification module 103, etc. For each computing device 700, when the program codes stored in the memory 720 are executed by the processor 710, the processor 710 executes part of the methods performed by the data processing device based on the communication interface 730. For example, some of the computing devices 700 may be used to execute the methods performed by the above-mentioned acquisition module 101, and some other computing devices 700 are used to execute the methods performed by the above-mentioned construction module 102 and verification module 103. The memory 720 may also store data, such as intermediate data or result data generated by the processor 710 during the execution process, for example, the above-mentioned acquired base element group and the constructed check matrix, etc.

[0181] The communication interface 730 in each computing device 700 is used for external communication, such as interacting with other computing devices 700, etc.

[0182] The bus 740 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. For the sake of simplicity of representation, Figure 8 the bus 740 in each computing device 700 is only represented by a thick line, but it does not mean that there is only one bus or one type of bus.

[0183] A communication path is established among the above-mentioned multiple computing devices 700 through a communication network to implement the functions of the data processing device. Any computing device may be a computing device (such as a server) in a cloud environment, or a computing device in an edge environment, or a terminal device.

[0184] In addition, an embodiment of the present application also provides a computer-readable storage medium, in which instructions are stored. When the instructions are run on one or more computing devices, the one or more computing devices execute the methods performed by the respective modules of the data processing device in the above-mentioned embodiment.

[0185] In addition, an embodiment of the present application also provides a computer program product. When the computer program product is executed by one or more computing devices, the one or more computing devices execute any of the methods in the foregoing data processing methods. The computer program product may be a software installation package. In the case where any of the foregoing data processing methods needs to be used, the computer program product may be downloaded and executed on a computer.

[0186] In addition, it should be noted that the system embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the system embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0187] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, in more cases for this application, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0188] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0189] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. A data processing method, characterized in that, The method includes: Obtaining a base element group, where the base element group includes a plurality of base elements, and the number of 1s in the binary representation of the plurality of base elements is non-zero and less than a first threshold; Constructing a parity-check matrix according to the plurality of base elements, where the parity-check matrix is used to encode a data block to generate a check block.

2. The method according to claim 1, wherein The obtaining of the base element group includes: Obtaining the number of the data blocks and the number of the check blocks; Determining a base element group corresponding to the number of the data blocks and the number of the check blocks.

3. The method according to claim 2, wherein The determining of the base element group corresponding to the number of the data blocks and the number of the check blocks includes: Determining a base element group according to the number of the data blocks, the number of the check blocks, and the field coefficient of a Galois field, where the field coefficient indicates the number of elements included in the Galois field, and the plurality of base elements are composed of the elements of the Galois field.

4. The method according to claim 3, characterized in that, The determining of the base element group according to the number of the data blocks, the number of the check blocks, and the field coefficient of the Galois field includes: Determining an encoding coefficient according to the number of the data blocks and the number of the check blocks, where the encoding coefficient is used to measure the proportional relationship between the number of the data blocks and the number of the check blocks; Determining a base element group according to the encoding coefficient and the field coefficient of the Galois field.

5. The method according to claim 4, wherein When the number of the check blocks is r, the determining of the base element group according to the encoding coefficient and the field coefficient of the Galois field includes: When the encoding coefficient and the field coefficient of the Galois field satisfy a first condition, determining a first base element group as the base element group, where the first base element group includes a plurality of base elements with the number of 1s in the binary representation being 1, and the first condition indicates that a parity-check matrix in which any r×r submatrix is invertible can be constructed by using the base elements with the number of 1s in the binary representation being 1; When the encoding coefficient and the field coefficient of the Galois field do not satisfy the first condition, determining a second base element group as the base element group, where the number of base elements included in the second base element group is greater than the number of base elements included in the first base element group.

6. The method according to any one of claims 1 to 5, characterized in that, The constructing of the parity-check matrix according to the plurality of base elements includes: For a first position of the parity-check matrix, taking any one of the plurality of base elements as the element at the first position of the parity-check matrix, where the first position is any position of the parity-check matrix.

7. The method according to any one of claims 1 to 6, characterized in that, When the number of the check blocks is r and the plurality of base elements include 1, the parity-check matrix satisfies the following conditions: There is a row in the parity-check matrix where all elements are 1; Except for the row where all elements are 1, the number of identical elements in the other row elements of the parity-check matrix is not greater than r - 1.

8. The method according to any one of claims 1 to 7, characterized in that, When the number of the check blocks is r, after the constructing of the parity-check matrix, the method further includes: Obtaining an r×r submatrix of the parity-check matrix; When any of the r×r submatrices is non-invertible, reconstructing the parity-check matrix according to the plurality of base elements.

9. A data processing device, characterized in that, The apparatus includes: An obtaining module, configured to obtain a base element group, where the base element group includes a plurality of base elements, and the number of 1s in the binary representation of the plurality of base elements is non-zero and less than a first threshold; A construction module, configured to construct a parity-check matrix according to the multiple basis elements, where the parity-check matrix is used to encode a data block to generate a check block.

10. The device according to claim 9, characterized in that Specifically, the obtaining module is configured to: Obtain the number of the data blocks and the number of the check blocks; Determine a basis element group corresponding to the number of the data blocks and the number of the check blocks.

11. The device according to claim 10, characterized in that, Specifically, the obtaining module is configured to: Determine a basis element group according to the number of the data blocks, the number of the check blocks, and the field coefficient of a Galois field, where the field coefficient indicates the number of elements included in the Galois field, and the multiple basis elements are composed of the elements of the Galois field.

12. The device according to claim 11, wherein Specifically, the obtaining module is configured to: Determine an encoding coefficient according to the number of the data blocks and the number of the check blocks, where the encoding coefficient is used to measure the proportional relationship between the number of the data blocks and the number of the check blocks; Determine a basis element group according to the encoding coefficient and the field coefficient of the Galois field.

13. The device according to claim 12, characterized in that, The number of the check blocks is r, and specifically, the obtaining module is configured to: When the encoding coefficient and the field coefficient of the Galois field satisfy a first condition, determine a first basis element group as the basis element group, where the first basis element group includes multiple basis elements with the number of 1s in the binary representation being 1, and the first condition indicates that a parity-check matrix in which any r×r submatrix is invertible can be constructed by using the basis elements with the number of 1s in the binary representation being 1; When the encoding coefficient and the field coefficient of the Galois field do not satisfy the first condition, determine a second basis element group as the basis element group, where the number of basis elements included in the second basis element group is greater than the number of basis elements included in the first basis element group.

14. The device according to any one of claims 9 to 13, characterized in that Specifically, the construction module is configured to: For a first position of the parity-check matrix, use any one of the multiple basis elements as the element at the first position of the parity-check matrix, where the first position is any position of the parity-check matrix.

15. The device according to any one of claims 9 to 14, characterized in that, The number of the check blocks is r, the multiple basis elements include 1, and the parity-check matrix satisfies the following conditions: There is a row in the parity-check matrix where all elements are 1; Except for the row where all elements are 1, the number of identical elements in the other row elements of the parity-check matrix is not greater than r−1.

16. The device according to any one of claims 9 to 15, characterized in that The number of the check blocks is r, and the apparatus further includes a verification module, where the verification module is configured to: Obtain an r×r submatrix of the parity-check matrix; When any of the r×r submatrices is non-invertible, reconstruct the parity-check matrix according to the multiple basis elements.

17. A computing device, characterized in that, The computing device includes at least one processor and at least one memory, where computer-readable instructions are stored in the at least one memory; the at least one processor executes the computer-readable instructions so that the computing device executes the method according to any one of claims 1 to 8.

18. A computer program product, characterized in that, Include computer-readable instructions; the computer-readable instructions are used to implement the method according to any one of claims 1 to 8.