Storage system, data storage method, device, storage medium, and program product

By bypassing CPU cache in the data storage process by using non-temporal instruction sets and directly transmitting data between registers, the problem of memory channel bandwidth loss in erasure coded storage is solved, and the IO bandwidth performance and hardware resource utilization of the storage system are improved.

CN120196290BActive Publication Date: 2025-07-25JINAN INSPUR DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510686385.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-25
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

When the prior art stores data based on erasure coding, there is additional loss in memory channel bandwidth level, resulting in a degradation of overall IO bandwidth performance of the storage system and insufficient utilization of hardware computing resources.

Method used

The non-temporal instruction set is used to directly write the data block to be stored from the original memory location to the target memory location through the register, and bypass the CPU cache during the storage of the erasure code calculation results and write it directly to the result memory to avoid additional memory reading operations.

Benefits of technology

It effectively improves the upper limit of external IO bandwidth performance of the storage system, improves the utilization rate of hardware computing resources, and reduces the time when computing devices wait for data supply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196290B_ABST
    Figure CN120196290B_ABST
Patent Text Reader

Abstract

The present invention discloses a storage system, a data storage method, a device, a storage medium, and a program product, which relate to the field of storage technologies. Among them, the method includes invoking a non-temporal instruction set, writing each data block to be stored of the data to be stored from the original memory location to the target memory location respectively through a scattered data register, and writing each data block to be stored at the target memory location to an intermediate data register. Performing erasure code calculation on each data block to be stored in the intermediate data register, writing the erasure code calculation result to a result register by invoking the non-temporal instruction set, and writing the erasure code calculation result in the result register to a result memory. The present invention can solve the problem of additional loss at the memory channel bandwidth level in the related art, effectively improve the upper limit of the external IO bandwidth performance of the storage system, and improve the utilization rate of hardware computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technology, and in particular to a storage system, a data storage method, a device, a storage medium, and a program product. Background Art

[0002] With the continuous growth of the amount of data to be stored and the increasing requirements for data security, the storage system stores data based on the erasure code method.

[0003] When the related technology stores data based on the erasure code, there is an additional loss at the memory channel bandwidth level, which reduces the upper limit of the overall IO (Input / Output) bandwidth performance of the storage system, and the utilization rate of hardware computing resources is insufficient, which is not conducive to the rapid reading and writing of large-scale data. Summary of the Invention

[0004] The present invention provides a storage system, a data storage method, an electronic device, a computer-readable storage medium, and a computer program product, which effectively improve the upper limit of the external IO bandwidth performance of the storage system and improve the utilization rate of hardware computing resources.

[0005] To solve the above technical problems, the present invention provides the following technical solutions:

[0006] On the one hand, the present invention provides a data storage method, including:

[0007] Invoking a non-temporal instruction set, writing each data block to be stored of the data to be stored into a target memory location from an original memory location through a scattered data register respectively;

[0008] Invoking a non-temporal instruction set to write each data block to be stored at the target memory location into an intermediate data register;

[0009] Performing erasure code calculation on each data block to be stored in the intermediate data register, and writing the erasure code calculation result into a result register by invoking a non-temporal instruction set;

[0010] Invoking a non-temporal instruction set to write the erasure code calculation result in the result register into a result memory.

[0011] The present invention also provides an electronic device, including a memory and a processor, and the processor is used to implement the steps of any one of the above data storage methods when executing a computer program stored in the memory.

[0012] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of any one of the above data storage methods when being executed by a processor.

[0013] The present invention also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of any of the above data storage methods.

[0014] Finally, the present invention also provides a storage system, at least including a storage device, a processor, a scattered data register, an intermediate data register, and a result register;

[0015] When receiving service data sent by a service client, the service data is stored as data to be stored at an original memory location. The processor is used to implement the steps of any of the above data storage methods when executing a computer program stored in the memory, so as to write the data to be stored into a result memory of the storage device through the scattered data register, the intermediate data register, and the result register.

[0016] The advantages of the technical solution provided by the present invention are as follows: during the process of scattering and placing the data to be stored, a non-temporal instruction set is used to write from the original memory location to the target memory location through registers, avoiding the way of passing through the central processing unit cache, resulting in cache pollution of the central processing unit. It can eliminate the additional memory read operations generated during the storage process, and save 1 / 3 of the bandwidth access to the memory channel. Further, during the process of saving the erasure correction result, a non-temporal instruction set is used to write to the result memory through registers, and it can be directly written to the memory without passing through the central processing unit cache, saving 1 / 2 of the memory access bandwidth. Thus, it effectively eliminates the problem of high memory access bandwidth consumption caused during the erasure correction calculation process, avoids the storage system prematurely triggering the performance limit of the hardware architecture, can effectively improve the external IO bandwidth limit of the storage system, that is, the IO bandwidth limit presented to the user, can transmit more data per unit time, enable the user to obtain higher performance, effectively save the waiting time of the computing device for data supply, and effectively improve the utilization rate of computing resources. In addition, the present invention also provides corresponding electronic devices, computer-readable storage media, computer program products, and storage systems for the data storage method, further making the method more practical, and the electronic devices, computer-readable storage media, computer program products, and storage systems have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of a data storage method provided by the present invention;

[0019] Figure 2 Schematic diagram for storing and re - placing the data to be stored in a scattered manner;

[0020] Figure 3 Schematic diagram for storing the calculation result of the erasure code;

[0021] Figure 4 Structural framework diagram under an exemplary embodiment of the data storage device provided by the present invention;

[0022] Figure 5 Structural diagram of an exemplary embodiment of the storage device provided by the present invention. Detailed implementation manners

[0023] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Among them, terms such as "first" and "second" in the specification and the above - mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non - exclusive inclusion. The term "exemplary" means "serving as an example, embodiment or illustration". Any embodiment described herein as "exemplary" does not necessarily need to be construed as superior to or better than other embodiments.

[0024] In order to improve the reliability of data storage in the storage system and reduce the storage cost, in the related art, the storage system stores data based on erasure codes, that is, non - hardware offloading. This storage method does not combine well with the hardware architecture, such as the CPU (Central Processing Unit) and memory, resulting in additional losses at the memory channel bandwidth level during the erasure calculation process. The storage system will prematurely reach the performance limit of the hardware architecture, reducing the overall IO bandwidth performance limit of the storage system, and the performance obtained by users is not high.

[0025] In view of this, in the process of scattering and placing the data to be stored, the present invention uses a non - temporal instruction set to write from the original memory location to the target memory location through a register. In the process of saving the erasure result, the non - temporal instruction set is used to write to the result memory through a register. Compared with the related art in the process of storing the data to be stored based on erasure codes, by bypassing the cache of the CPU for data storage operations, two target memory read operations can be omitted, and the bandwidth access to the memory channel can be saved by 1 / 3 + 1 / 2. Avoiding the storage system from prematurely reaching the performance limit of the hardware architecture can effectively improve the external IO bandwidth limit of the storage system, enabling users to obtain higher performance.

[0026] After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. First, please refer to Figure 1 , Figure 1 which is a schematic flowchart of a data storage method provided in this embodiment. This embodiment may include the following content:

[0027] S101: Call the non-temporal instruction set, and write each data block to be stored of the data to be stored into the target memory location from the original memory location through the scattered data register respectively.

[0028] In this embodiment, the non-temporal instruction set includes non-temporal instructions for implementing various operations. The non-temporal instruction set can use, for example, the NT (non-temporal) instruction set provided by the CPU. Those skilled in the art can also perform corresponding instruction addition, instruction deletion, or instruction modification on the existing non-temporal instruction set according to the actual application scenario, which does not affect the implementation of the present invention. When the present invention performs data loading and data operations on the data to be stored, it uses the instructions in the non-temporal instruction set that implement the corresponding operations, and does not use the general instruction set, so as to eliminate the "extra memory read bandwidth loss introduced by saving the calculation result of the erasure code" during the erasure process, and then convert the saved hardware resources into additional performance improvement.

[0029] Among them, the data to be stored is the service data sent by the service end to the storage system, and this service data needs to be stored in the storage system. Each data block to be stored is a data segment of the data to be stored, and each data block to be stored does not overlap with each other and constitutes the complete data to be stored. Among them, the scattered data register can use the SIMD (Single Instruction Multiple Data) register of the CPU or any general register, which is used to temporarily store each data block to be stored read from the original memory location. For the convenience of description, it is defined as the scattered data register. Among them, the target memory location is the location where the original data to be stored is scattered and stored. The target memory location and the original memory location are two different memory locations, and they can be located on the same storage device or on different storage devices, which does not affect the implementation of the present invention.

[0030] S102: Call the non-temporal instruction set to write each data block to be stored at the target memory location into the intermediate data register.

[0031] Among them, the intermediate data register can use the SIMD register of the CPU or any general register, which is used to temporarily store each data block to be stored read from the target memory location. For the convenience of description, it is defined as the intermediate data register.

[0032] S103: Perform erasure code calculation on each data block to be stored in the intermediate data register, and write the erasure code calculation result into the result register by invoking the non-temporal instruction set.

[0033] In this step, erasure code calculation refers to the process of generating additional redundant segments based on each data block to be stored divided from the data to be stored. The total number of data blocks after erasure = the original data blocks to be stored + the newly generated parity data blocks. That is to say, the erasure code calculation result includes at least the original data blocks to be stored and the newly generated parity data blocks. The result register can adopt the SIMD register of the CPU or any general-purpose register, which is used to temporarily store the erasure code calculation result. For the convenience of description, it is defined as the result register.

[0034] S104: Invoke the non-temporal instruction set to write the erasure code calculation result in the result register into the result memory.

[0035] When in the previous step, k original segments, that is, each data block to be stored, are used to calculate m new segments, that is, parity blocks, and stored in the result register, the non-temporal instructions for implementing data storage operations in the non-temporal instruction set can be used to store each data block to be stored and each parity block in the result register to a specified location. The result memory includes storage areas for storing each data block to be stored and each parity block. These storage areas can be located on the same storage device or on different devices. To improve the fault tolerance of the storage system, each parity block can be scattered and stored on different storage devices.

[0036] In the technical solution provided in this embodiment, during the process of scattering and placing the data to be stored, the non-temporal instruction set is used to write from the original memory location to the target memory location through the register, avoiding the way of passing through the CPU's cache, which may cause CPU cache pollution. It can eliminate the additional memory read operations generated during the storage process, and save 1 / 3 of the bandwidth access to the memory channel. Further, during the process of saving the erasure result, the non-temporal instruction set is used to write into the result memory through the register, which can be directly written into the memory without passing through the CPU's cache, saving 1 / 2 of the memory access bandwidth. Thus, it can effectively eliminate the problem of high memory access bandwidth consumption caused during the erasure code calculation process, avoid the storage system prematurely triggering the performance limit of the hardware architecture, and can effectively improve the external IO bandwidth limit of the storage system, that is, the IO bandwidth limit presented to the user. It can transmit more data per unit time, enabling users to obtain higher performance, effectively saving the waiting time of the computing device for data supply, and effectively improving the utilization rate of computing resources.

[0037] The above embodiments do not make any limitations on how to obtain each data block to be stored. Based on the above embodiments, the present invention also provides multiple implementation manners, which may include the following:

[0038] As an exemplary implementation, when the storage system receives a user's data block parameter setting request, the data block parameter setting request will carry at least the occupied space capacity value of the data block, or the address storing the occupied space capacity value of the data block. The occupied space capacity value of the data block refers to the size of each data block to be stored. According to the position where the occupied space capacity value of the data block is carried in the parsed data block parameter setting request, this information is extracted from the corresponding position of the data block parameter setting request, thus obtaining the occupied space capacity value of the data block by parsing the data block parameter setting request. When the storage system receives the data to be stored sent by the user side, the data to be stored is segmented into corresponding numbers of data blocks to be stored according to the occupied space capacity value of the data block. The occupied space capacity value of the data block can be, for example, 4k or 8k or any value smaller than the size of the data to be stored. For example, if the data to be stored is 1M and the occupied space capacity value of the data block is 4k, the 1M data to be stored is segmented into 256 4K data blocks to be stored.

[0039] As an exemplary implementation, when the storage system receives a data block parameter setting request sent by the user, the data block parameter setting request will carry at least the data length value of the data block, or the address storing the data length value of the data block. The data length value of the data block refers to the data length of each data block to be stored. According to the position where the data length value of the data block is carried in the parsed data block parameter setting request, this information is extracted from the corresponding position of the data block parameter setting request, thus obtaining the data length value of the data block by parsing the data block parameter setting request. Starting from the starting position of the data to be stored, data with the corresponding length is sequentially read as the current data block to be stored according to the data length value, and the current data block to be stored is written from the original memory location to the target memory location through the scatter data register. For example, if the starting position of the data to be stored in the original memory location is A and the data length value is 4 bytes, starting from A, 4 bytes of data are read as the first data block to be stored, and then taking the end position of the first data block to be stored as the starting position, the 1M data to be stored is segmented into 256 4K data blocks to be stored according to the read length.

[0040] Furthermore, in order to reduce memory operations and improve the upper limit of the external IO bandwidth performance of the storage system, the parsed data length value or the occupied space capacity value of the data block can be called, and the non - temporal store instruction is used to replace the general store instruction to first load the data length value or the occupied space capacity value of the data block into the parameter register. When it is detected that the user sends the data to be stored, the data length value or the occupied space capacity value of the data block is read from the parameter register by calling the non - temporal load instruction.

[0041] In the above embodiments, there is no limitation on how to write each data block to be stored to the target memory location through the scattered data register. Based on the above embodiments, the present invention further provides an exemplary implementation method, which may include the following content:

[0042] Call a non-temporal data loading instruction to write each data block to be stored from the original memory location to the scattered data register. For example, the non-temporal data loading instruction uses the non-temporal load instruction. Correspondingly, call the non-temporal load instruction to replace the general load instruction, and directly load each data block to be stored from the original memory location to the scattered data register, such as the SIMD or general register of the CPU, so as to avoid the CPU cache pollution caused by the original method passing through the CPU cache. Call a non-temporal data storage instruction to write each data block to be stored from the scattered data register to the corresponding target memory location. For example, the non-temporal data storage instruction uses the non-temporal store. Correspondingly, call the non-temporal store instruction to replace the general store instruction, and directly write the data from the scattered data register, such as the SIMD or general register of the CPU, to the target memory location. When the prior art uses a general instruction to copy, an additional cache filling operation of "reading from the target memory to the CPU cache" will be generated during the store process. That is to say, compared with the prior art, the additional memory read operation generated during the store process can be eliminated.

[0043] In this embodiment, as Figure 2As shown, the related technology takes a 1M business data block as the data to be stored, divides it into multiple data blocks according to 4k or 8k or N k, and then stores it to the target storage location (i.e., K(2)) in the way of memory copy or direct address reference. The access bandwidth to memory in the prior art is: 1 time read of the original memory + 1 time read of the target memory + 1 time write of the target memory = 3M. For the memory copy method of the related technology, the processing process of the IO block size of 1M issued by the service layer in this embodiment is as follows: for the 1M data to be stored, taking 4k as a unit, gradually loop as follows: for each 4k, taking the size of the SIMD register or general register of the CPU as a unit, gradually loop as follows: call the non-temporal load instruction to directly load the 4k data block from the original memory location into the SIMD or general register of the CPU. Call the non-temporal store instruction to directly write the 4k data from the SIMD or general register of the CPU to the target memory location. Continuously repeat the above process until all 4k data blocks to be stored are written from the original storage location to the target memory location. After replacing the general memory copy with the NT instruction of the CPU, the memory access bandwidth is: 1 time read of the original memory + 1 time write of the target memory = 2M. By comparing the two, it can be seen that this embodiment saves 1 / 3 of the bandwidth access to the memory channel compared with the memory copy method of the prior art, thus effectively increasing the upper limit of the external (i.e., presented to the user) IO bandwidth of the storage system, transmitting more data per unit time, enabling computing devices such as CPUs and graphics processors not to wait for data supply, and improving the utilization rate of computing resources. For example, when training a neural network model, the IO bottleneck may cause the utilization rate of the graphics processor to be less than 30%. Through the technical solution of the present invention, it can be increased to more than 80%.

[0044] To further improve the fault tolerance, the target memory location may include a first target memory and a second target memory, and the number of data blocks to be stored written in the first target memory and the second target memory is the same. Correspondingly, a non-temporal instruction set is called to write each data block to be stored from the original memory location to a scattered data register, and randomly write each data block to be stored stored in the scattered data register to the first target memory and the second target memory. Exemplarily, a first non-temporal instruction is called to read a first data block to be stored from the original memory location and write it to the scattered data register; a second non-temporal instruction is called to write the first data block to be stored from the scattered data register to the first target memory; a first non-temporal instruction is called to read a second data block to be stored from the original memory location and write it to the scattered data register; a second non-temporal instruction is called to write the second data block to be stored from the scattered data register to the second target memory. Among them, the first non-temporal instruction may be a data loading instruction, and the second non-temporal instruction may be a data storing instruction. For example, call the non temporal load instruction to directly load each data block to be stored from the original memory location to the scattered data register, and call the non temporal store instruction to write a part of the data blocks to be stored to the first target memory and the remaining part of the data blocks to be stored to the second target memory.

[0045] To further improve the fault tolerance, based on the above embodiment, the intermediate data register may include a first register and a second register; the result memory includes at least a first result memory located in the first storage device and a second result memory located in the second storage device. Call the non-temporal data loading instruction to write the data blocks to be stored in the first target memory to the first register and the data blocks to be stored in the second target memory to the second register respectively. The first register and the second register may adopt the SIMD register of the CPU or any general register. Call the non-temporal data storing instruction to randomly store each check block read from the result register to the first result memory and the second result memory.

[0046] For example, as Figure 3As shown, the related art reads the above-mentioned 4K data blocks to be stored horizontally, calculates the check block M, and directly writes it into the result memory M(1) through ordinary "store" semantic instructions. The access bandwidth to the memory is: reading the target memory once + writing the target memory once = 2 times. In this embodiment, taking 4k as a unit, for each 4k data block to be stored in the first target memory and the second target memory, for each 4k, taking the size of the CPU's SIMD register or general register as a unit, the progressive loop is as follows: for the first target memory and the second target memory respectively, use the non-temporal instruction to load the data into the CPU's SIMD register. Compared with the ordinary load instruction, it can avoid re-polluting the CPU cache during the load process. Perform erasure calculation on the data in the first register and the second register and generate 1 erasure result, and store it in the result register. Store the data in the result register into the result memory through the non-temporal instruction. Repeat the above process until all 4k data blocks to be stored and their corresponding check blocks are written into the result memory. To improve the fault tolerance, the check block can be stored in different locations, such as disks, storage nodes, etc. Similarly, as mentioned above, this non-temporal store operation can avoid the extra memory read operation generated by the original ordinary store. And using the NT instruction to bypass the CPU cache and directly write to the memory, the access bandwidth to the memory is: writing the target memory once, and the corresponding memory access bandwidth is saved by 1 / 2.

[0047] For the above data scattering stage and data storage stage, each stage can act alone and reduce a part of the memory amplification. If both act simultaneously, there can be a further improvement, and the superimposed quantization effect is as follows: The total memory access bandwidth saved: data to be stored * (1 / 3) + erasure code calculation result * (1 / 2). Taking k = 2, m = 1 as an example, the data to be stored is 1M, and the erasure code calculation result is 0.5M, and 1M * (1 / 3) + 0.5M * 0.5 = 58% can be saved; taking k = 4, m = 1 as an example, the data to be stored is 1M, and the erasure code calculation result is 0.25M, and 1M * (1 / 3) + 0.25M * 0.5 = 45% can be saved. The consumption of hardware resources by the saved memory access bandwidth will ultimately be converted into an improvement in IO bandwidth performance.

[0048] The above embodiment does not make any limitation on how to perform erasure code calculation on each data block to be stored in the intermediate data register. Based on the above embodiment, the present invention also gives an exemplary erasure code calculation method, which may include the following content:

[0049] Construct a generation matrix according to the number k of data blocks to be stored and the number m of parity blocks, where the generation matrix satisfies that any k*k submatrix is invertible; determine the symbol vectors of each data block to be stored according to each data block to be stored in the first register and the second register and the finite field; generate each parity block according to the symbol vectors of each data block to be stored and the generation matrix.

[0050] Among them, before calculating the erasure code, the respective numbers of data blocks to be stored and parity blocks can be set first. For example, the number of data blocks to be stored is k, and the number of parity blocks, that is, the number of redundant blocks, is m. The total number of storage blocks n = k + m. Construct a generation matrix G according to the numbers of data blocks to be stored and parity blocks. The first k rows of the generation matrix are the data blocks to be stored, and the last m rows can be the parity blocks, satisfying that any k*k submatrix is invertible, such as using a Cauchy matrix, a Vandermonde matrix, etc. Then select a finite field, such as , GF represents the finite field, and each byte is used as an element of the finite field for operations to ensure efficient operations and meet the data scale.

[0051] After the data to be stored is sliced into multiple data blocks, each data block is further divided into multiple symbols. For example, each symbol is 1 byte corresponding to , and the number of symbols of all data blocks needs to be aligned. Pad with zeros when insufficient. For each symbol position s, such as the 1st byte, the 2nd byte, etc., perform the following operations: Extract the s-th symbol from k data blocks to be stored and construct a vector. For ease of description, it is defined as the symbol vector. The symbol vectors of all data blocks to be stored can be expressed as , and each parity block is calculated by the product of the last m rows of the generation matrix and the symbol vector. For example, , , the first row of the generation matrix is , the symbol vector is , then the parity block p1 is , represents matrix multiplication.

[0052] After storing k original data blocks and m parity data blocks in the above embodiments, when less than or equal to m blocks, data blocks or parity blocks are abnormal, such as lost, the lost data blocks can be calculated backward through the remaining data blocks. Similarly, the present invention provides a data recovery process based on a non-temporal instruction set, thereby further improving the data storage security and stability of the storage system, which may include the following content:

[0053] When there is an abnormality in the first data block to be stored or the first parity block in the result memory, call the non-temporal data loading instruction to write the generation matrix of the result memory, the target data block to be stored and / or the target parity block in the normal state to the data to be recovered register; according to the generation matrix and the surviving data blocks, recover the first data block to be stored or the first parity block to obtain a recovered block; call the non-temporal data loading instruction to write the recovered block to the data recovery register; call the non-temporal data storage instruction to write the recovered block in the data recovery register back to the corresponding position in the result memory.

[0054] In this embodiment, the data to be recovered register and the data recovery register can be SIMD registers or any general-purpose register. The data to be recovered register stores the data required for data recovery, and the data recovery register stores the recovered data blocks. For the sake of distinction, they are defined separately. The first data block to be stored is any data block to be stored in the result memory in the above embodiment, and the first parity block is any parity block stored in the result memory in the above embodiment. The abnormality includes the complete loss of the data block or partial data loss. The surviving data blocks refer to the data blocks that are not lost and are complete in the result memory. For the sake of description, they are defined as the target data block to be stored and the target parity block. The target data block to be stored and / or the target parity block together constitute the surviving data blocks. That is to say, the surviving data blocks can only include the data blocks to be stored, or only include the parity blocks, or both, as long as the total number of target data blocks to be stored and the total number of target parity blocks are the same as the total number of data blocks to be stored. The generation matrix is the generation matrix in the process of generating the parity block, and the recovered block is the recovered data block. If the first data block to be stored is abnormal, the recovered block is the recovered first data block to be stored. If the first parity block is abnormal, the recovered block is the recovered first parity block. Exemplarily, the recovery process can be: select the data corresponding to the number of rows of the surviving data blocks from the generation matrix to form a target sub-matrix, perform an inverse operation on the target sub-matrix to obtain a recovered sub-matrix; based on the recovered sub-matrix, perform a recovery process on the surviving symbol vector corresponding to the surviving data blocks to obtain a recovered block. In this embodiment, select k rows corresponding to the surviving blocks from the generation matrix G to form a sub-matrix , calculate , for each symbol position s, construct a surviving symbol vector, recover the original data symbol through the inverse matrix, and recombine the recovery results of all symbol positions into the original data block.

[0055] As can be seen from the above, in this embodiment, the original data is converted into encoded data with redundancy through the above steps, ensuring complete recovery even when part of the data is lost and improving the data security of the storage system.

[0056] The present invention also provides a corresponding device for the data storage method, further making the method more practical. Among them, the device can be described from the perspective of functional modules and the perspective of hardware respectively. The data storage device provided by the present invention will be introduced below. This device is used to implement the data storage method provided by the present invention. In this embodiment, the data storage device may include or be divided into one or more program modules. These one or more program modules are stored in a storage medium and executed by one or more processors to complete the data storage method disclosed in Embodiment 1. The program modules referred to in this embodiment refer to a series of computer program instruction segments that can complete specific functions, and are more suitable for describing the execution process of the data storage device in the storage medium than the program itself. The following description will specifically introduce the functions of each program module in this embodiment. The data storage device described below can be correspondingly referred to the data storage method described above.

[0057] From the perspective of functional modules, referring to Figure 4 , Figure 4 is a structural diagram of the data storage device provided in this embodiment in a specific implementation manner. The device may include:

[0058] The original data scattering module 401 is used to call a non-temporal instruction set to write each data block to be stored of the data to be stored from the original memory location to the target memory location through a scattered data register.

[0059] The erasure code calculation module 402 is used to call a non-temporal instruction set to write each data block to be stored at the target memory location to the intermediate data register; perform erasure code calculation on each data block to be stored in the intermediate data register, and write the erasure code calculation result to the result register by calling a non-temporal instruction set.

[0060] The data writing module 403 is used to call a non-temporal instruction set to write the erasure code calculation result in the result register to the result memory.

[0061] Exemplarily, in some implementation manners of this embodiment, the above-mentioned original data scattering module 401 may further be used to: call a non-temporal data loading instruction to write each data block to be stored from the original memory location to the scattered data register; call a non-temporal data storage instruction to write each data block to be stored from the scattered data register to the corresponding target memory location.

[0062] Exemplarily, in some other embodiments of this embodiment, the above-mentioned original data scattering module 401 may further be configured to: the target memory locations include a first target memory and a second target memory, call a non-temporal instruction set, write each data block to be stored from the original memory location into a scattered data register, and randomly write each data block to be stored stored in the scattered data register into the first target memory and the second target memory; wherein, the number of data blocks to be stored written into the first target memory and the second target memory is the same.

[0063] Exemplarily, in some other embodiments of this embodiment, the above-mentioned original data scattering module 401 may further be configured to: the target memory locations include a first target memory and a second target memory, call a first non-temporal instruction, read a first data block to be stored from the original memory location and write it into the scattered data register; call a second non-temporal instruction, write the first data block to be stored from the scattered data register into the first target memory; call a first non-temporal instruction, read a second data block to be stored from the original memory location and write it into the scattered data register; call a second non-temporal instruction, write the second data block to be stored from the scattered data register into the second target memory.

[0064] Exemplarily, in some other embodiments of this embodiment, the above-mentioned erasure calculation module 402 may further be configured to: the target memory locations include a first target memory and a second target memory, and the intermediate data register includes a first register and a second register; call a non-temporal data loading instruction, and write the data blocks to be stored in the first target memory into the first register respectively, and write the data blocks to be stored in the second target memory into the second register.

[0065] As an exemplary implementation manner of the above embodiment, the above-mentioned erasure calculation module 402 may further be configured to: construct a generator matrix according to the number k of data blocks to be stored and the number m of parity blocks, and the generator matrix satisfies that any k*k submatrix is invertible; determine the symbol vectors of each data block to be stored according to each data block to be stored in the first register and the second register and the finite field; generate each parity block according to the symbol vectors of each data block to be stored and the generator matrix.

[0066] Exemplarily, in some other embodiments of this embodiment, the above-mentioned data writing module 403 may further be configured to: the result memory at least includes a first result memory located in a first storage device and a second result memory located in a second storage device, call a non-temporal data storage instruction, and randomly store each parity block read from the result register into the first result memory and the second result memory.

[0067] Exemplarily, in some other embodiments of this embodiment, the above-mentioned original data scattering module 401 may further be configured to: when receiving a data block parameter setting request, obtain the data block occupied space capacity value by parsing the data block parameter setting request; when receiving data to be stored, split the data to be stored into corresponding numbers of data blocks to be stored according to the data block occupied space capacity value.

[0068] Exemplarily, in some other embodiments of this embodiment, the above-mentioned original data scattering module 401 may further be configured to: when receiving a data block parameter setting request, obtain the data block length value by parsing the data block parameter setting request; starting from the starting position of the data to be stored, sequentially read data of corresponding lengths as the current data block to be stored according to the data block length value, and write the current data block to be stored from the original memory location to the target memory location through the scattered data register.

[0069] Exemplarily, in some other embodiments of this embodiment, the above-mentioned device may further include a data recovery module, which is configured to: when there is an abnormality in the first data block to be stored or the first check block in the result memory, call the non-temporal data loading instruction to write the generation matrix, the target data blocks to be stored and / or the target check blocks in the normal state in the result memory to the data to be recovered register; where the total number of target data blocks to be stored and the total number of target check blocks are the same as the total number of data blocks to be stored, as surviving data blocks; according to the generation matrix and the surviving data blocks, recover the first data block to be stored or the first check block to obtain a recovery block; call the non-temporal data loading instruction to write the recovery block to the data recovery register; call the non-temporal data storage instruction to write the recovery block in the data recovery register back to the corresponding position in the result memory.

[0070] As an exemplary implementation manner of the above embodiment, the above-mentioned data recovery module may further be configured to: select the data corresponding to the rows of the surviving data blocks from the generation matrix to form a target sub-matrix, perform an inverse operation on the target sub-matrix to obtain a recovery sub-matrix; based on the recovery sub-matrix, perform a recovery process on the surviving symbol vector corresponding to the surviving data blocks to obtain a recovery block.

[0071] The data storage device mentioned above is described from the perspective of functional modules. Further, the present invention also provides an electronic device, which is described from the perspective of hardware. The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned data storage method embodiments.

[0072] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any of the above data storage method embodiments when running.

[0073] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), external hard drives, magnetic disks or optical discs that can store computer programs.

[0074] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above data storage method embodiments.

[0075] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above data storage method embodiments.

[0076] Finally, the present invention also provides a storage system. Please refer to Figure 5 , which may at least include a storage device 501, a processor 502, a scattered data register 503, an intermediate data register 504, and a result register 505; when receiving service data sent by a service client, storing the service data as data to be stored at an original memory location, and when the processor 502 executes a computer program, it implements the steps of the data storage method described in any of the above embodiments, and writes the data to be stored into the result memory of the storage device 501 through the scattered data register 503, the intermediate data register 504, and the result register 505.

[0077] As can be seen from the above, in the storage system of this embodiment, a non-temporal instruction set is used to replace the general instruction set to eliminate the additional memory read bandwidth loss introduced by saving the calculation result of the erasure code during the erasure process. By reducing the memory read / write amplification introduced during the erasure calculation, the saved hardware resources are converted into additional performance improvements, improving the IO bandwidth performance of the storage system for the service layer.

[0078] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0079] The above has introduced in detail a data storage method, an electronic device, a computer-readable storage medium, a computer program product, and a storage system provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A data storage method, characterized in that, Including: Call a non-temporal instruction set to write each data block to be stored from the original memory location to the target memory location through a scatter data register; Call the non-temporal instruction set to write each data block to be stored at the target memory location to an intermediate data register; Perform erasure code calculation on each data block to be stored in the intermediate data register, and write the erasure code calculation result to a result register by calling the non-temporal instruction set; Call the non-temporal instruction set to write the erasure code calculation result in the result register to a result memory.

2. The data storage method according to claim 1, wherein Call a non-temporal instruction set to write each data block to be stored from the original memory location to the target memory location through a scatter data register, including: Call a non-temporal data load instruction to write each data block to be stored from the original memory location to a scatter data register; Call a non-temporal data store instruction to write each data block to be stored from the scatter data register to the corresponding target memory location.

3. The data storage method according to claim 1, wherein The target memory location includes a first target memory and a second target memory. Call a non-temporal instruction set to write each data block to be stored from the original memory location to the target memory location through a scatter data register, including: Call a non-temporal instruction set to write each data block to be stored from the original memory location to the scatter data register, and randomly write each data block stored in the scatter data register to the first target memory and the second target memory; Wherein, the number of data blocks to be stored written to the first target memory and the second target memory is the same.

4. The data storage method according to claim 1, wherein The target memory location includes a first target memory and a second target memory. Call a non-temporal instruction set to write each data block to be stored from the original memory location to the target memory location through a scatter data register, including: Call a first non-temporal instruction to read a first data block to be stored from the original memory location and write it to a scatter data register; Call a second non-temporal instruction to write the first data block to be stored from the scatter data register to the first target memory; Call the first non-temporal instruction to read a second data block to be stored from the original memory location and write it to the scatter data register; Call the second non-temporal instruction to write the second data block to be stored from the scatter data register to the second target memory.

5. The data storage method according to claim 1, wherein The target memory location includes a first target memory and a second target memory, and the intermediate data register includes a first register and a second register; Call the non-temporal instruction set to write each data block to be stored at the target memory location to an intermediate data register, including: Call a non-temporal data load instruction to write the data blocks to be stored in the first target memory to the first register and write the data blocks to be stored in the second target memory to the second register respectively.

6. The data storage method according to claim 5, wherein Perform erasure code calculation on each data block to be stored in the intermediate data register, including: Construct a generator matrix according to the number k of data blocks to be stored and the number m of parity blocks, and the generator matrix satisfies that any k*k submatrix is invertible; Determine the symbol vectors of each data block to be stored according to the data blocks to be stored and the finite field of the first register and the second register; Generate each check block according to the symbol vectors of each data block to be stored and the generation matrix.

7. The data storage method according to claim 1, characterized in that The result memory at least includes a first result memory located in a first storage device and a second result memory located in a second storage device. Invoking the non-temporal instruction set to write the erasure code calculation result of the result register to the result memory includes: Invoking a non-temporal data storage instruction to randomly store each check block read from the result register to the first result memory and the second result memory.

8. The data storage method according to claim 1, wherein Before writing each data block to be stored of the data to be stored to the target memory location from the original memory location through the scattered data register respectively, it includes: When receiving a data block parameter setting request, obtain the data block occupied space capacity value by parsing the data block parameter setting request; When receiving the data to be stored, split the data to be stored into corresponding numbers of data blocks to be stored according to the data block occupied space capacity value.

9. The data storage method according to claim 1, characterized in that, Writing each data block to be stored of the data to be stored to the target memory location from the original memory location through the scattered data register respectively, includes: When receiving a data block parameter setting request, obtain the data block length value by parsing the data block parameter setting request; Starting from the starting position of the data to be stored, sequentially read data of the corresponding length as the current data block to be stored according to the data block length value, and write the current data block to be stored to the target memory location from the original memory location through the scattered data register.

10. The data storage method according to any one of claims 1 to 9, characterized in that After invoking the non-temporal instruction set to write the erasure code calculation result of the result register to the result memory, it further includes: When there is an abnormality in a first data block to be stored or a first check block in the result memory, invoking a non-temporal data loading instruction to write the generation matrix of the result memory, the target data block to be stored and / or the target check block in a normal state to the data to be recovered register; wherein the total number of target data blocks to be stored and the total number of target check blocks are the same as the total number of data blocks to be stored, as surviving data blocks; Recover the first data block to be stored or the first check block according to the generation matrix and the surviving data blocks to obtain a recovered block; Invoking a non-temporal data loading instruction to write the recovered block to the data recovery register; Invoking a non-temporal data storage instruction to write the recovered block in the data recovery register back to the corresponding position in the result memory.

11. The data storage method according to claim 10, wherein Recovering the first data block to be stored or the first check block according to the generation matrix and the surviving data blocks includes: Select data corresponding to the number of rows of the surviving data blocks from the generation matrix to form a target submatrix, and perform an inverse operation on the target submatrix to obtain a recovered submatrix; Based on the recovered submatrix, perform a recovery process on the surviving symbol vectors corresponding to the surviving data blocks to obtain a recovered block.

12. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the data storage method according to any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the data storage method described in any one of claims 1 to 11 are implemented.

14. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the data storage method described in any one of claims 1 to 11 are implemented.

15. A storage system, characterized in that, It at least includes a storage device, a processor, a scattered data register, an intermediate data register, and a result register; When receiving service data sent by a service client, the service data is stored as data to be stored at an original memory location, and when the processor executes the computer program, the steps of the data storage method described in any one of claims 1 to 11 are implemented to write the data to be stored into a result memory of the storage device through the scattered data register, the intermediate data register, and the result register.

Citation Information

Patent Citations

  • Register-friendly efficient XOR erasure code coding method

    CN115934409A

  • Exclusive OR calculation method and device of storage system and product

    CN118394565A