Low-memory coding method and device based on sub-blocking
By dividing the data block and the verification block into sub-strips, and using the calculation unit to calculate the verification code according to the encoding matrix segments, the problems of high peak memory overhead and long encoding time in the encoding process in the prior art are solved, and lower memory requirements and shorter encoding time are achieved.
Patent Information
- Application Number
- CN202510117829.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
AI Technical Summary
The existing erasure coding encoding technology has high peak memory overhead and a long encoding time during the encoding process, resulting in high repair costs.
The low memory encoding method based on subblocking is adopted to divide the data blocks and the verification blocks into sub-strips, and the verification encoding is calculated in segments according to the encoding matrix according to the encoding matrix to reduce the peak memory demand and encoding time.
It effectively reduces the peak memory requirement during the encoding calculation process, shortens the encoding time, and significantly improves the overall performance and response speed of the system through parallel computing.
Smart Images

Figure CN120011131A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer storage technology, and more specifically, to a low-memory encoding method and device based on sub-blocking. Background Art
[0002] Distributed storage systems allow enterprises to process large amounts of data across multiple storage nodes. These systems are usually built with inexpensive commodity servers and are therefore prone to failure. To address this issue, enterprises usually increase the reliability of the system by increasing data redundancy. Erasure coding is widely adopted because it can provide the same fault tolerance at a lower storage cost. This technology works by splitting the data into multiple data blocks and generating check blocks through specific matrix operations. These data blocks and check blocks together constitute a data stripe.
[0003] Although erasure coding provides an economical redundancy solution, it also brings high repair costs. Excessive repair bandwidth requirements may interfere with the normal service of the storage system and may cause the repair time to be extended, thus leaving the system in a degraded state for a long time and increasing the risk of data loss.
[0004] To address this problem, there is already an erasure coding scheme with low repair bandwidth in the prior art. However, since this scheme involves the operation of embedding the information of the check sub-blocks into each other, it has a higher peak memory overhead and a longer encoding time.
[0005] Therefore, how to reduce the peak memory during the encoding process and shorten the encoding time is a technical problem that needs to be solved urgently. Summary of the invention
[0006] In view of the defects of the prior art, the purpose of the present application is to provide a low-memory encoding method and device based on sub-blocking, aiming to solve the problems of high peak memory overhead and long encoding time in the prior art.
[0007] To achieve the above objectives, in a first aspect, the present application provides a low-memory encoding method based on sub-blocking, comprising: Will include data blocks and The stripes of parity blocks are divided into sub-strips, each sub-strip includes sub-blocks and Sub-check blocks, ; Constructing a coding matrix according to the number of data blocks and the number of check blocks, and calculating the check codes of the original data sub-blocks in sections in units of sub-blocks according to the coding matrix using a calculation unit; The checksum code of each original data sub-block is calculated repeatedly until the calculation of the checksum information of all original data sub-blocks is completed to obtain a final calculation result.
[0008] Optionally, the data blocks and The stripes of parity blocks are divided into sub-strips, including: Divide each data block into Sub-data blocks; Divide each check block into Sub-check blocks; Any sub-stripe is determined according to any sub-data block in each data block and any sub-check block in each check block.
[0009] Optionally, the calculating unit calculating the check code of the original data sub-block in segments in units of sub-blocks according to the coding matrix includes: Take any original data sub-block o as the current data sub-block; Determining corresponding coefficients of a coding matrix of the current data sub-block; Determine one of the check information of the current data sub-block according to the corresponding coefficient, and merge the check information into the buffer of the final calculation result through Galois field addition and subtraction operation; The encoding operation of the original data sub-block is repeatedly performed until all the check sub-blocks corresponding to the m coefficients of the original data sub-block are calculated and added to the final result.
[0010] Optionally, determining one piece of check information of the current data sub-block according to the corresponding coefficient, and merging the check information into the buffer of the final calculation result through Galois field addition and subtraction operations, includes: Calculating check information of the original data sub-block by Galois Field multiplication according to the corresponding coefficients, and using the check information as an intermediate calculation result; Determining a target position of the correlation check of the original data sub-block in a coding matrix; The intermediate calculation results are calculated by Galois field addition or subtraction, and the intermediate calculation results are merged into the buffer of the final calculation result according to the target position.
[0011] Optionally, the repeatedly calculating the check code of each original data sub-block until the calculation of the check information of all original data sub-blocks is completed to obtain a final calculation result includes: Traversing each original data sub-block, and allocating the original data sub-block to cache and computing operations performed in idle computing units; Merging the check data of the original data sub-block into the final result buffer; After completing the traversal of all original data sub-blocks, it is determined that all check information of all original data sub-blocks is calculated and merged into the corresponding area of the final result buffer; The stored data in the final result buffer is used as the final calculation result.
[0012] Optionally, the number of the computing units is set to p, where p≥1; The final result buffer is determined according to k, m, and o; Each computing unit allocates a memory buffer of size o as a raw data buffer, and the computing unit stores one raw data sub-block in one calculation.
[0013] Optionally, the comprising data blocks and The stripes of the parity blocks are obtained as follows: Divide the original data into data blocks; Based on the erasure code encoding method, Data blocks are encoded to generate A check block; According to the data blocks and the A check block is obtained to obtain the stripe.
[0014] In a second aspect, the present application further provides a low-memory encoding device based on sub-blocking, comprising: Get the module that contains data blocks and The stripes of parity blocks are divided into sub-strips, each sub-strip includes sub-blocks and Sub-check blocks, ; A segmented coding module, used for constructing a coding matrix according to the number of data blocks and the number of check blocks, and calculating the check codes of the original data sub-blocks in segments according to the coding matrix and in units of sub-blocks by using a calculation unit; The repetitive coding module is used to repeatedly calculate the check code of each original data sub-block until the calculation of the check information of all original data sub-blocks is completed to obtain the final calculation result.
[0015] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0017] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0018] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0019] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the prior art: (1) The present application forms a stripe with k data blocks and m check blocks, and performs sub-striping using erasure coding, thereby effectively reducing the peak memory demand in the encoding calculation process. Sub-striping divides the data into smaller units, so that each computing unit only needs to focus on a limited number of sub-data blocks and sub-check blocks during processing, thereby reducing the peak memory demand in the encoding calculation process and shortening the encoding time.
[0020] (2) The present application performs parallel calculations through multiple computing units, which can significantly shorten the encoding time. Through parallel processing, multiple sub-strips can perform encoding calculations simultaneously, thereby improving the overall performance and response speed of the system and ensuring that high efficiency and reliability can be maintained during large-scale data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of a low-memory encoding method based on sub-blocking provided in an embodiment of the present application; Figure 2 is a schematic diagram of an example of a coding matrix provided in an embodiment of the present application; Figure 3 It is one of the structural schematic diagrams of a low-memory encoding device based on sub-blocking provided in an embodiment of the present application; Figure 4 This is a second structural diagram of a low-memory encoding device based on sub-blocking provided in an embodiment of the present application; Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.
[0024] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of the objects. For example, a first response message and a second response message are used to distinguish different response messages rather than to describe a specific order of the response messages.
[0025] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0026] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.
[0027] In a distributed storage system, a parameter is The erasure code is , that is, the size is The original data D is divided into equal Data blocks , each data block contains Then, it is encoded according to certain rules to generate By adding this total The encoding blocks are placed separately in different storage nodes of the distributed storage system to ensure the original data Can be arbitrarily ( ) blocks are recovered. If , then the erasure code is called Maximum Distance Separable (MDS).
[0028] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.
[0029] Reference Figure 1 , the present application provides a low memory encoding method based on sub-blocking, comprising: S101. will include data blocks and The stripes of parity blocks are divided into sub-strips, each sub-strip includes sub-blocks and Sub-check blocks, ; S102. Construct a coding matrix according to the number of data blocks and the number of check blocks, and use a calculation unit to calculate the check code of the original data sub-block in segments in units of sub-blocks according to the coding matrix; S103. Repeat the calculation of the check code of each original data sub-block until the calculation of the check information of all original data sub-blocks is completed to obtain the final calculation result.
[0030] First, through S101, the stripe is divided into multiple sub-stripes. Specifically, the stripe contains k data blocks and m check blocks, which are evenly divided into t sub-stripes, and the number of t is equal to m. This means that each sub-strip will contain k sub-data blocks and m sub-check blocks. The stripe is refined by division, and each sub-strip can be made an independent processing unit, which is convenient for subsequent encoding calculations.
[0031] The segmentation method of the embodiment of the present application enhances the flexibility of the system, so that the number and structure of sub-strips can be flexibly adjusted when processing data sets of different sizes. By organizing data and check information in the form of sub-strips, the data and check relationship within each sub-strip can be concretized to provide support for the subsequent encoding process.
[0032] S102 constructs a coding matrix according to the number of data blocks and check blocks in the stripe, which is specifically embodied in: constructing a data block matrix according to the number of data blocks, and constructing a check block matrix according to the number of check blocks.
[0033] It should be noted that in the embodiment of the present application, striping is a data storage technology that divides continuous data into data blocks of the same size, and then writes each piece of data to different disks in the array. In erasure coding, a stripe contains data blocks and Pass The check block is generated by encoding the data blocks. and are all integers greater than 1.
[0034] For a stripe obtained by erasure coding, it has data blocks and Check blocks, assuming that the size of the data block and the check block are both O'. It is required that the data block and the check block can be evenly divided into Each portion is o in size, that is For a strip ( ) coding blocks, and divide these blocks equally into sub-blocks, then generate sub-strips, each of which contains sub-blocks and Sub-check blocks. Among them, The value of can be The values of are the same.
[0035] Reference Figure 2 , Figure 2 It is a schematic diagram of the coding matrix of an embodiment of the present application.
[0036] The erasure code parameters are , which can be understood as the number of data blocks k is 6, and the number of check blocks is 3, so it is divided into 3 sub-bands, that is, the data block matrix has 6 rows and 3 columns, the check block matrix has 3 rows and 3 columns, and the encoding matrix has 9 rows and 3 columns. , is the set value; Furthermore, the checksum code of each original data sub-block is systematically recalculated until the comprehensive calculation of the checksum information of all data sub-blocks can be completed. The recalculation can not only improve the accuracy of the code, but also enhance the fault tolerance of the system under the influence of potential uncertain factors and ensure data integrity. The complete checksum information of all original data sub-blocks is obtained through the final calculation result.
[0037] Optionally, the data blocks and The stripes of parity blocks are divided into sub-strips, including: Divide each data block into Sub-data blocks; Divide each check block into Sub-check blocks; Any sub-stripe is determined according to any sub-data block in each data block and any sub-check block in each check block.
[0038] Any sub-stripe is determined according to any sub-data block in each data block and any sub-check block in each check block.
[0039] In the specific implementation, for a data blocks and Assume that the size of each data block and each check block is O', and each data block and each check block can be divided into share.
[0040] By dividing each data block into sub-data blocks, and divide each check block into sub-check blocks, get sub-blocks and sub-check blocks. The size of each sub-block is .
[0041] for Any of the sub-strips can be Sub-data blocks can be selected arbitrarily sub-blocks, and from The number of sub-check blocks can be selected arbitrarily Sub-check blocks are obtained.
[0042] Optionally, the calculating unit calculating the check code of the original data sub-block in segments in units of sub-blocks according to the coding matrix includes: Take any original data sub-block o as the current data sub-block; Determining corresponding coefficients of a coding matrix of the current data sub-block; Determine one of the check information of the current data sub-block according to the corresponding coefficient, and merge the check information into the buffer of the final calculation result through Galois field addition and subtraction operation; The encoding operation of the original data sub-block is repeatedly performed until all the check sub-blocks corresponding to the m coefficients of the original data sub-block are calculated and added to the final result.
[0043] Specifically, the encoding process of the embodiment of the present application is as follows: 1. Select one of a group of original data sub-blocks as the current data sub-block.
[0044] 2. Find the correlation coefficient of the current data sub-block in the coding matrix. The coding matrix contains the contribution coefficient of each original data sub-block to different check blocks, which will be used for subsequent check information calculation.
[0045] 3. Using the determined coefficients, calculate the information of each relevant check block. This process involves combining the current data sub-block with the corresponding coefficients to generate a check information for each check block.
[0046] 4. Merge the calculated checksum information into the final result buffer. This process uses addition and subtraction operations to update the checksum information into the corresponding checksum block.
[0047] 5. Continue to select other original data sub-blocks and perform the above steps in sequence until all check blocks are calculated and added to the final result buffer. Each time a new data sub-block is selected, the process of selecting, determining coefficients, calculating check information, and merging results is repeated to ensure that all check information is updated.
[0048] Further, determining one of the check information of the current data sub-block according to the corresponding coefficient, and merging the check information into the buffer of the final calculation result through Galois field addition and subtraction operations, includes: Calculating check information of the original data sub-block by Galois Field multiplication according to the corresponding coefficients, and using the check information as an intermediate calculation result; Determining a target position of the correlation check of the original data sub-block in a coding matrix; The intermediate calculation results are calculated by Galois field addition or subtraction, and the intermediate calculation results are merged into the buffer of the final calculation result according to the target position.
[0049] Specifically, in this embodiment, Galois Field multiplication is used to calculate the check information. The calculation multiplies the current data sub-block by the correlation coefficient, and the generated result will be used as the intermediate calculation result of this check block.
[0050] Then, it is necessary to determine the target position of the check information in the encoding matrix. The intermediate calculation results are merged according to the target position. Through the addition or subtraction operation of the Galois Field, the intermediate results are effectively merged into the final calculation result buffer according to the target position, ensuring that the information in the check block is accurately updated, providing a reliable basis for data recovery. Repeat the above process until the check information of all original data sub-blocks is calculated and merged into the final result buffer.
[0051] Specifically, m calculation operations are performed according to the m coefficients in the coding matrix of the check block; for each calculation operation, Galois Field multiplication is performed on the original data sub-block to obtain an intermediate result of the check data.
[0052] A sub-data block of size o is allocated to an idle computing unit and stored in a corresponding original data buffer of size o in the computing unit.
[0053] The original data sub-block is calculated m times, and each calculation corresponds to a coefficient in the coding matrix. Specifically, for each calculation, the original data is multiplied by a coefficient according to the Galois Field multiplication, and the coefficient is determined by the sequence number of the data sub-block and the corresponding function of the coding matrix to generate an intermediate result; then the intermediate result is merged into the final result buffer according to the Galois Field addition or subtraction, and the type of calculation performed (addition or subtraction) and the area merged into the final result buffer are determined by the sequence number of the data sub-block and the corresponding function of the coding matrix.
[0054] Optionally, the repeatedly calculating the check code of each original data sub-block until the calculation of the check information of all original data sub-blocks is completed to obtain a final calculation result includes: Traversing each original data sub-block, and allocating the original data sub-block to cache and computing operations performed in idle computing units; Merging the check data of the original data sub-block into the final result buffer; After completing the traversal of all original data sub-blocks, it is determined that all check information of all original data sub-blocks is calculated and merged into the corresponding area of the final result buffer; The stored data in the final result buffer is used as the final calculation result.
[0055] In some embodiments, the above operation is performed for each original data sub-block until all coding coefficients of the original data sub-block are calculated; then the final generated verification information is stored in the final result verification buffer, specifically including: Traverse each original data sub-block, assign the sub-block to an idle computing unit to perform the above-mentioned caching and calculation operations, and merge the verification data of the original data sub-block into the final result buffer. If there is no idle computing unit, wait. After completing the traversal of all sub-blocks, all verification information of all sub-blocks is calculated and merged into the corresponding area of the final result buffer. At this time, the final verification result is stored in the final result buffer.
[0056] Optionally, the number of the computing units is set to p, where p≥1; The final result buffer is determined according to k, m, and o; Each computing unit allocates a memory buffer of size o as a raw data buffer, and the computing unit stores one raw data sub-block in one calculation.
[0057] In this embodiment, p computing units (p≥1) are allocated, and one computing unit corresponds to a CPU core and a certain amount of memory in hardware, and corresponds to a thread that executes computing tasks and a certain amount of memory in the operating system.
[0058] Allocate m*t*o size memory as the final result buffer, which is exactly the size of all the verification data. After the entire encoding step is completed, the final verification calculation result is stored in the buffer.
[0059] Each computing unit is allocated a memory buffer of size o (a total of p*o memory) as the original data buffer, and each computing unit can store exactly one data sub-block in one calculation.
[0060] It should be noted that one computing unit can independently complete the computing task of one original data sub-block, and multiple computing units can complete multiple computing tasks in parallel at the same time, specifically including: The calculation tasks of each original data sub-block are not related to each other, and multiple calculation units can be used to calculate the calculation tasks of multiple original data sub-blocks at the same time. By increasing the number of calculation units for parallel calculation, the encoding calculation speed is accelerated.
[0061] For example, assume that a 27MB stripe is divided into 3 sub-stripes, and each sub-strip has 6 sub-data blocks and 3 sub-check blocks, and each sub-block is 1MB in size. Assuming that 3 computing units are allocated, each computing unit occupies a CPU core for most of the time, and corresponds to a thread on the operating system to complete the computing task. At the same time, 9MB of memory is allocated as the final result buffer to store the last calculated check block data. In addition, each computing unit is allocated a 1MB memory buffer (3MB in total) as the original data buffer to store an original sub-data block.
[0062] Optionally, the comprising data blocks and The stripes of the parity blocks are obtained as follows: Divide the original data into data blocks; Based on the erasure code encoding method, Data blocks are encoded to generate A check block; According to the data blocks and the A check block is obtained to obtain the stripe.
[0063] Continue to see Figure 2 In the embodiment of the present application, the erasure code parameter is . And the three check equations of the basic code RS(9, 6) used are:
[0064]
[0065]
[0066] Assume that the third sub-block of the first sub-strip is calculated The calculation task is first assigned to an idle computing unit. The original data is read into the original data buffer of the computing unit. There are three calibration coefficients corresponding to the data, respectively .
[0067] according to Figure 2 The check matrix, (Right now ) is stored in the first sub-check block of the first sub-stripe and the first sub-check block of the second sub-stripe. Multiply by the coefficient 1 through Galois Field multiplication to generate an intermediate result Then, the intermediate result The data area corresponding to the first sub-check block of the first sub-strip in the final result buffer is calculated by Galois Field addition; the intermediate result The data area corresponding to the first sub-check block of the second sub-strip in the final result buffer is calculated by Galois Field subtraction.
[0068] Similarly, according to Figure 2 The check matrix, (Right now ) is stored in the second sub-check block of the first sub-stripe and the second sub-check block of the second sub-stripe. Multiply by 4 through Galois Field multiplication to generate an intermediate result Then, the intermediate result The data area corresponding to the second sub-check block of the first sub-strip in the final result buffer is calculated by Galois Field addition; the intermediate result The data area corresponding to the second sub-check block of the second sub-strip in the final result buffer is calculated by Galois Field subtraction.
[0069] Similarly, according to Figure 2 The check matrix, (Right now ) is stored in the third sub-check block of the first sub-stripe and the third sub-check block of the second sub-stripe. Multiply by 9 through Galois Field multiplication to generate an intermediate result Then, the intermediate result The data area corresponding to the third sub-check block of the first sub-strip in the final result buffer is calculated by Galois Field addition; the intermediate result The data area corresponding to the third sub-check block of the second sub-strip in the final result buffer is calculated by Galois Field subtraction.
[0070] In some embodiments, the above cache and calculation operations are performed on each original data sub-block until all checks are calculated and added to the final result buffer, at which point the final check result is stored in the final result buffer.
[0071] This encoding method can effectively reduce the peak memory size during the encoding process, while allowing the encoding process to be accelerated by repeatedly setting computing units and parallel computing.
[0072] by Figure 2 Taking the corresponding encoding process as an example, compared with the RS overall encoding scheme in the prior art, when encoding, three computing units are used to perform encoding in parallel with sub-blocks as the granularity, and the encoding of three sub-blocks can be calculated at the same time, which can achieve up to 3 times the encoding speed. At the same time, since the overall encoding requires all the complete data to be read into the memory, in the above example, the classic RS overall encoding scheme requires a total of 27MB of memory (for storing all original data and verification data), while the low-memory encoding method based on sub-blocking can only read the sub-block being calculated, requiring a total of 12MB of memory, significantly reducing memory overhead.
[0073] Reference Figure 3 The present application also provides a low-memory encoding device based on sub-blocking, comprising: The acquisition module 310 is used to include data blocks and The stripes of parity blocks are divided into sub-strips, each sub-strip includes sub-blocks and Sub-check blocks, ; A segmented coding module 320, configured to construct a coding matrix according to the number of data blocks and the number of check blocks, and to calculate the check codes of the original data sub-blocks in segments according to the coding matrix in units of sub-blocks using a calculation unit; The repetitive coding module 330 is used to repeatedly calculate the check code of each original data sub-block until the calculation of the check information of all original data sub-blocks is completed to obtain a final calculation result.
[0074] Optionally, the data blocks and The stripes of parity blocks are divided into sub-strips, including: Divide each data block into Sub-data blocks; Divide each check block into Sub-check blocks; Any sub-stripe is determined according to any sub-data block in each data block and any sub-check block in each check block.
[0075] Optionally, the calculating unit calculating the check code of the original data sub-block in segments in units of sub-blocks according to the coding matrix includes: Take any original data sub-block o as the current data sub-block; Determining corresponding coefficients of a coding matrix of the current data sub-block; Determine one of the check information of the current data sub-block according to the corresponding coefficient, and merge the check information into the buffer of the final calculation result through Galois field addition and subtraction operation; The encoding operation of the original data sub-block is repeatedly performed until all the check sub-blocks corresponding to the m coefficients of the original data sub-block are calculated and added to the final result.
[0076] Optionally, determining one piece of check information of the current data sub-block according to the corresponding coefficient, and merging the check information into the buffer of the final calculation result through Galois field addition and subtraction operations, includes: Calculating check information of the original data sub-block by Galois Field multiplication according to the corresponding coefficients, and using the check information as an intermediate calculation result; Determining a target position of the correlation check of the original data sub-block in a coding matrix; The intermediate calculation results are calculated by Galois field addition or subtraction, and the intermediate calculation results are merged into the buffer of the final calculation result according to the target position.
[0077] Optionally, the repeatedly calculating the check code of each original data sub-block until the calculation of the check information of all original data sub-blocks is completed to obtain a final calculation result includes: Traversing each original data sub-block, and allocating the original data sub-block to cache and computing operations performed in idle computing units; Merging the check data of the original data sub-block into the final result buffer; After completing the traversal of all original data sub-blocks, it is determined that all check information of all original data sub-blocks is calculated and merged into the corresponding area of the final result buffer; The stored data in the final result buffer is used as the final calculation result.
[0078] Optionally, the number of the computing units is set to p, where p≥1; The final result buffer is determined according to k, m, and o; Each computing unit allocates a memory buffer of size o as a raw data buffer, and the computing unit stores one raw data sub-block in one calculation.
[0079] Optionally, the comprising data blocks and The stripes of the parity blocks are obtained as follows: Divide the original data into data blocks; Based on the erasure code encoding method, Data blocks are encoded to generate A check block; According to the data blocks and the A check block is obtained to obtain the stripe.
[0080] Reference Figure 4 , Figure 4 is a schematic diagram of the device structure of the embodiment of the present application, and the original data is , , the calculation unit passes and Perform calculations to obtain the final calculation results.
[0081] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method, which will not be repeated here.
[0082] Reference Figure 5 Based on the method in the above embodiment, the embodiment of the present application provides an electronic device, which may include: a processor (Processor) 510, a communication interface (Communications Interface) 520, a memory (Memory) 530 and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logic instructions in the memory 530 to execute the method in the above embodiment.
[0083] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.
[0084] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.
[0085] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0086] It is understandable that the processor in the embodiment of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0087] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0088] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
[0089] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0090] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A low memory encoding method based on sub-blocking, characterized in that: include: Will include data blocks and The stripes of parity blocks are divided into sub-strips, each sub-strip includes sub-blocks and Sub-check blocks, ; Constructing a coding matrix according to the number of data blocks and the number of check blocks, and calculating the check codes of the original data sub-blocks in sections in units of sub-blocks according to the coding matrix using a calculation unit; The checksum code of each original data sub-block is calculated repeatedly until the calculation of the checksum information of all original data sub-blocks is completed to obtain a final calculation result.
2. The low memory encoding method based on sub-blocking according to claim 1, characterized in that: The will include data blocks and The stripes of parity blocks are divided into sub-strips, including: Divide each data block into Sub-data blocks; Divide each check block into Sub-check blocks; Any sub-stripe is determined according to any sub-data block in each data block and any sub-check block in each check block.
3. The low memory encoding method based on sub-blocking according to claim 1, characterized in that: The calculating unit calculates the check code of the original data sub-block in segments in units of sub-blocks according to the coding matrix, including: Take any original data sub-block o as the current data sub-block; Determining corresponding coefficients of a coding matrix of the current data sub-block; Determine one of the check information of the current data sub-block according to the corresponding coefficient, and merge the check information into the buffer of the final calculation result through Galois field addition and subtraction operation; The encoding operation of the original data sub-block is repeatedly performed until all the check sub-blocks corresponding to the m coefficients of the original data sub-block are calculated and added to the final result.
4. The low memory encoding method based on sub-blocking according to claim 3, characterized in that: The step of determining one of the check information of the current data sub-block according to the corresponding coefficient, and merging the check information into the buffer of the final calculation result through Galois field addition and subtraction operations, comprises: Calculating check information of the original data sub-block by Galois Field multiplication according to the corresponding coefficients, and using the check information as an intermediate calculation result; Determining a target position of the correlation check of the original data sub-block in a coding matrix; The intermediate calculation results are calculated by Galois field addition or subtraction, and the intermediate calculation results are merged into the buffer of the final calculation result according to the target position.
5. The low memory encoding method based on sub-blocking according to claim 1, characterized in that: The repeatedly calculating the check code of each original data sub-block until the calculation of the check information of all original data sub-blocks is completed to obtain the final calculation result includes: Traversing each original data sub-block, and allocating the original data sub-block to cache and computing operations performed in idle computing units; Merging the check data of the original data sub-block into the final result buffer; After completing the traversal of all original data sub-blocks, it is determined that all check information of all original data sub-blocks is calculated and merged into the corresponding area of the final result buffer; The stored data in the final result buffer is used as the final calculation result.
6. The low memory encoding method based on sub-blocking according to claim 5, characterized in that: The number of the computing units is set to p, where p≥1; The final result buffer is determined according to k, m, and o; Each computing unit allocates a memory buffer of size o as a raw data buffer, and the computing unit stores one raw data sub-block in one calculation.
7. The low memory encoding method based on sub-blocking according to claim 1, characterized in that: The inclusion data blocks and The stripes of the parity blocks are obtained as follows: Divide the original data into data blocks; Based on the erasure code encoding method, Data blocks are encoded to generate A check block; According to the data blocks and the A check block is obtained to obtain the stripe.
8. A low memory encoding device based on sub-blocking, characterized in that: include: Get the module that contains data blocks and The stripes of parity blocks are divided into sub-strips, each sub-strip includes sub-blocks and Sub-check blocks, ; A segmented coding module, used for constructing a coding matrix according to the number of data blocks and the number of check blocks, and calculating the check codes of the original data sub-blocks in segments according to the coding matrix and in units of sub-blocks by using a calculation unit; The repetitive coding module is used to repeatedly calculate the check code of each original data sub-block until the calculation of the check information of all original data sub-blocks is completed to obtain the final calculation result.
9. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor is caused to execute the method according to any one of claims 1 to 7.