System and method for hybrid data reliability for object storage devices
A stateless hybrid reliability manager with pluggable mechanisms addresses data reliability and efficiency challenges in key-value storage by adapting to varying data sizes and access frequencies, ensuring efficient storage and repair of key-value pairs.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2018-12-10
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional data storage systems, particularly those using solid-state drives (SSDs), face challenges in ensuring data reliability and efficiency for key-value storage devices due to varying object sizes and unstructured formats, necessitating novel mechanisms that address space efficiency and fast access times.
A stateless hybrid reliability manager is implemented to manage key-value storage devices using pluggable reliability mechanisms such as object replication, K-object (k, r) erasure coding packing, single-object (k, r) erasure coding splitting, K-object (k, r, d) regeneration coding packing, and single-object (k, r, d) regeneration coding splitting, which enable efficient storage, retrieval, and repair of key-value pairs of different sizes.
The solution provides improved data reliability and storage efficiency by ensuring single-key repair procedures and maintaining fast access times, even in the event of device failures, through a hybrid reliability mechanism that adapts to varying data sizes and access frequencies.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AREA
[0001] One or more aspects of embodiments of the present disclosure relate generally to data storage systems and more specifically to a method for selecting a reliability mechanism for reliably storing key-value data in a key-value reliability system comprising a plurality of key-value storage devices. BACKGROUND
[0002] Data compliance mechanisms such as erasure coding can be used to overcome data loss due to data corruption and storage device malfunctions in many installations that have multiple storage devices.
[0003] Conventional solid-state drives (SSDs) typically use only a block interface and can provide data reliability through a redundant arrangement of independent disks (i.e., RAID) by ensuring encoding or replication. As object formats become variable in size and unstructured, there is a need for efficient data conversion between object-level and block-level interfaces. Furthermore, it is desirable to ensure data reliability while maintaining space efficiency and fast access times.
[0004] Techniques such as RAID have been well-studied for traditional block storage devices. However, relatively new key-value storage devices may have different interfaces and storage semantics compared to traditional block devices. Consequently, many new key-value storage devices may benefit from novel data reliability mechanisms that are designed for or implemented on key-value data and key-value storage devices.
[0005] From US Patent 8,504,535 B1, various embodiments of an erasure coding scheme and a redundant replication storage scheme in a data storage system are known. Data objects that are larger than a certain size threshold and are accessed less frequently than a certain access threshold are stored in an erasure coding scheme, while data objects that are smaller than a certain size threshold or are accessed more frequently than a certain access threshold are stored in a redundant replication storage scheme.
[0006] From US patent 8,949,180 B1, a method for replicating a key-value pair is known, comprising the following: intercepting a command to update a key-value pair in a key-value pair database, wherein the key-value database contains metadata of a virtual disk; sending an updated key-value pair to a data backup device; receiving an acknowledgment that the data backup device has received the updated key-value pair; and updating the key-value pair in the key-value database after receiving the acknowledgment. SUMMARY
[0007] The invention is specified in the main claim and dependent claims 11 and 15. Further developments of the invention are specified in the dependent claims.
[0008] The embodiments described herein provide improvements for the field of storage, since the admissibility mechanisms of the embodiments are each capable of a single key repair procedure which enables a virtual device management layer to repair and copy all of the keys present in the failed storage device for a new control device. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] These and / or other aspects will be made clearer and more easily recognized from the following description of the embodiments, together with the accompanying drawings, in which: Fig. 1 is a block diagram which represents a key-value reliability system for storing key-value data based on a selected reliability mechanism according to an embodiment of the present disclosure; Fig. 2 is a flowchart which represents a selection of a reliability mechanism to be used by a key-value reliability system based on a size limit corresponding to the size of data of a key-value pair, according to an embodiment of the present disclosure; Fig. 3 is a block diagram representing a group of KV storage devices configured to store key value data according to a reliability mechanism of K-object (k, r) erasure coding or multi-object “packing” using conventional erasure coding according to an embodiment of the present disclosure; Fig. 4 is a block diagram which represents a storage of value objects and parity objects in accordance with the reliability mechanism of K-object (k, r) erasure coding or multi-object “packing” using conventional erasure coding according to an embodiment of the present disclosure; Fig. 5 is a block diagram representing a group of KV storage devices configured to store key value data according to a reliability mechanism of single-object (k, r) erasure coding or “splitting” using conventional erasure coding according to an embodiment of the present disclosure; and Fig. 6 is a block diagram representing a group of KV storage devices configured to store key value data according to a reliability mechanism of single-object (k, r, d) erasure coding or “splitting” using regeneration erasure coding according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0010] Features of the inventive concept and of methods for achieving it can be more easily understood by reference to the following detailed description of embodiments and the accompanying drawings. Embodiments are described in more detail below with reference to the accompanying drawings. However, the described embodiments can be implemented in various different forms and should not be considered limited to those illustrated herein. Rather, these embodiments are provided as examples so that this disclosure will be conscientious and complete and will fully convey the aspects and features of the present inventive concept to those skilled in the art.Consequently, processes, elements, and techniques that are not necessary for a person skilled in the art to fully understand the aspects and features of the present inventive concept may not be described. Unless otherwise specified, identical reference numerals denote identical elements across the accompanying drawings and the description, and therefore, descriptions of such elements will not be repeated. Furthermore, parts not related to the description of embodiments may not be shown in order to clarify the description. In the drawings, the relative sizes of elements, layers, and regions may be exaggerated for clarity.
[0011] If a particular embodiment can be implemented in different ways, a specific process sequence can be carried out differently from the described sequence. For example, two processes described sequentially can be carried out essentially at the same time or in a sequence opposite to the one described.
[0012] The electronic or electrical devices and / or all other relevant devices or components according to embodiments of the present disclosure described herein can be implemented using any suitable hardware, firmware (for example, an application-specific integrated circuit), software, or a combination of software, firmware, and hardware. For example, the various components of these devices can be formed on an integrated circuit (IC) chip or on separate IC chips. Furthermore, the various components of these devices can be implemented on a flexible printed circuit film, a tape carrier package (TCP), a printed circuit board (PCB), or on a substrate.Furthermore, the various components of these devices can be a process or a thread running on one or more processors in one or more computing devices, which execute computer program instructions and interact with other system components to perform the various functionalities described herein. The computer program instructions are stored in memory, which can be implemented in a computing device using a standard storage device such as random access memory (RAM). The computer program instructions can also be stored in other non-perishable, computer-readable media such as a CD-ROM, a flash drive, or the like.Likewise, a person skilled in the art should recognize that the functionality of different calculating devices can be combined or integrated into a single calculating device, or that the functionality of a particular calculating device can be distributed over one or more calculating devices, without deviating from the concept and scope of the embodiment of the present disclosure.
[0013] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as generally understood by a person skilled in the art to whose field the present inventive concept belongs. It shall further be understood that terms such as those defined in conventionally used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant field and / or the present description, and should not be interpreted in an idealized or overly formal sense, unless expressly defined herein.
[0014] As described below, embodiments of the present disclosure provide a method for reliably storing key-value data in a key-value reliability system comprising a plurality of key-value (KV) storage devices grouped into a logical unit. Furthermore, embodiments of the present disclosure provide a stateless hybrid reliability manager for managing the drives and for managing the storage of key-value (KV) pairs, which relies on several pluggable reliability mechanisms / techniques / implementations comprising: object replication; K-object (k, r) erasure coding pack; single-object (k, r) erasure coding split; K-object (k, r, d) regeneration coding pack; single-object (k, r, d) regeneration coding split.
[0015] The described embodiments enable improved storage (for example, storing key-value data in key-value storage devices) because a stateless hybrid reliability manager, which relies on multiple pluggable reliability mechanisms, can manage the devices and control the storage of KV pairs, and because the disclosed methods for selecting the reliability mechanism ensure efficient storage, retrieval, and repair of KV pairs of different sizes.
[0016] Fig. Figure 1 is a block diagram representing a key-value reliability system for storing key-value data based on a selected reliability mechanism according to an embodiment of the present disclosure.
[0017] Referring to Fig. As mentioned above, many new key-value (KV) storage devices / memory devices / drives / KV SSDs 130 may benefit from new data reliability mechanisms that are tailored to or used for key-value data and KV storage devices 130. Accordingly, a hybrid key-value reliability system for such KV storage devices 130 may include a stateless hybrid reliability manager / virtual device manager layer / virtual device management layer 120 configured to exercise a hybrid reliability mechanism to manage the KV storage devices 130 and to control the storage of KV pairs therein according to one or more pluggable reliability mechanisms.It should be noted that, although solid-state drives (SSDs) are generally used to refer to the KV storage devices described herein, other storage devices may be used in accordance with embodiments of this disclosure. The design and operations of the virtual device management layer 120 according to embodiments of this disclosure are described below.
[0018] In the present embodiment, the virtual device management layer 120 can enable a method for reliably storing key-value data / a KV pair 170 in a key-value reliability system. The key-value reliability system can comprise a plurality of KV storage devices 130, which are grouped together into a logical unit. The logical unit can be referred to as a reliability group 140.
[0019] The KV storage devices 130 of the permission group 140 can store respective blocks or segments of erasure-encoded data and / or replicated data, which can correspond to the key-value data 170. The KV storage devices 130 in the reliability group 140 expose a single virtual device 110, to which key-value operations are routed via the virtual device management layer 120.
[0020] Virtual device 110 can have a stateless hybrid reliability manager as the virtual device management layer 120. This means that the virtual device management layer 120 can operate in a stateless manner (that is, without the need to maintain any key-value-to-device mapping).
[0021] Therefore, the virtual device 110 can store the key-value data 170 across N KV storage devices 130, where N is an integer (for example, KV-SSDs 130-1, 130-2, 130-3, 130-4 ... 130-N), and can store the key-value data 170 in the KV storage devices 130 via the virtual device management layer 120. This means that the virtual device management layer 120 can manage the KV storage devices 130 and control the storage of KV pairs within them.
[0022] In other embodiments, the key-value reliability system can also include a cache to optionally store data and / or metadata associated with keys of the key-value data 170, thereby increasing operating speed. The reliability mechanisms can append a value to a key-value pair to contain the metadata corresponding to that pair. This means that both a key and a value with information corresponding to a metadata identifier "MetaID" can be appended to store additional metadata specific to the key-value pair. The metadata can include a checksum, a reliability mechanism identifier to identify the reliability mechanisms used to store the data, an erasure coding identifier, object sizes, placement of parity group elements, etc.
[0023] In other embodiments, the key-value reliability system can also include Bloom filters corresponding to the reliability mechanisms described below. These Bloom filters can store keys that are stored using the corresponding reliability mechanism, thus supporting the key-value reliability system in read operations. Consequently, one or more Bloom filters or caches of the key-value reliability system can enable rapid testing of keys against existing reliability mechanisms.
[0024] Each of the reliability mechanisms described herein can first store the first copy or block or segment of the KV pair corresponding to the key-value data 170, using the same hash function on the key modulo the number of corresponding KV storage devices 130. That is to say, for each of the pluggable reliability mechanisms, the reliability mechanism can store at least the first copy / block or segment using the same key as the user key.
[0025] As discussed above, embodiments of the present disclosure provide several pluggable reliability mechanisms for ensuring reliable storage of key-value data 170 in the plurality of KV storage devices 130. Accordingly, the virtual device management layer 120 can rely on the reliability mechanisms and can determine which of the reliability mechanisms are to be used.
[0026] The reliability mechanisms can be based on rules such as value limits set during the setup of the virtual device 110 and / or object read / write frequency. Consequently, the virtual device management layer 120 can select a suitable reliability mechanism based on the system's specific rules.
[0027] Five reliability mechanisms of embodiments of the present disclosure, how the reliability mechanisms work, and when the reliability mechanisms can be appropriately used and selected by the virtual device management layer 120 are described below. Such reliability mechanisms may include techniques that can be referred to as object replication, K-object (k, r) erasure coding packing, single-object (k, r) erasure coding splitting, K-object (k, r, d) regeneration coding packing, and single-object (k, r, d) regeneration coding splitting.
[0028] Fig. 2 is a flowchart 200 which represents a selection of a reliability mechanism to be used by a key value reliability system based on a size limit which corresponds to a size of data of a KV pair, according to an embodiment of the present disclosure.
[0029] Referring to Fig. 2 can determine the value size of data (for example, key value data 170) for no fewer than all of the supported reliability mechanisms that are based on a size limit (for example, the five reliability mechanisms mentioned above), and can determine whether the value size is less than a given limit t. i , which corresponds to the respective reliability mechanism and can select the first reliability mechanism in which the value limit condition is met.
[0030] For example, in S210, the virtual device management layer 120 can receive a number "n" of supported size limit-based reliability mechanisms, where n is an integer. In S220, the virtual device management layer 120 can simply check each of the reliability mechanisms, one at a time, in order from one to n. In S230, the virtual device management layer 120, in its check of each of the supported reliability mechanisms, can determine whether the value of the data is less than the limit t. i , which corresponds to the respective reliability mechanism.
[0031] In S240, if a reliability mechanism is found which has a limit value t iIf the value of the data is greater than or equal to the value of the data, the virtual device management layer 120 selects this reliability mechanism to use. For S250, this can be done either when determining which reliability mechanism to use for S240, or when determining that none of the n reliability mechanisms has a limit t. i has which is suitable for recording the value size in a final iteration of S220, the virtual device management layer 120 completes its determination of which reliability mechanism to use.
[0032] In the present embodiment, “n” can be five, corresponding to the five different reliability mechanisms of the embodiments described herein. For relatively small key values (that is, where the value size is relatively small), the virtual device management layer 120 can select the object replication reliability mechanism for use. For slightly larger key values, the virtual device management layer 120 can select the packing and then splitting reliability mechanisms (for example, in that order), using conventional erasure coding for each. For even larger key values, however, the virtual device management layer can select packing and then splitting, using regeneration erasure coding instead of conventional erasure coding.
[0033] In embodiments of the present disclosure, the selection of a reliability mechanism for use can be based on one or more of the object's size, the object's throughput requirements, the read / write temperature of a corresponding key-value pair, the underlying coding capabilities of the majority of kV storage devices, and / or a detection of whether a key is hot or cold. For example, "hot" keys can utilize the object replication reliability mechanism regardless of their value size, while "cold" keys can be applied to one of the reliability mechanisms of the erasure-coded schemes depending on their value size. As another example, the decision to use the object replication reliability mechanism can be based on both the size and the write temperature.Therefore, a limit value corresponding to an object image read / write frequency can be found in flowchart 200 of the . Fig. 2 instead of the limit value, which corresponds to a quantity, can be used to determine a reliability mechanism.
[0034] The respective operations of the five reliability mechanisms are discussed below.
[0035] Referring back to Fig. 1. As mentioned above, if the value size of the key-value pair is relatively small, the reliability mechanism of "object replication" may be suitable for selection by the virtual device management layer 120. Object replication can be applied per object (for example, per key-value pair / key-value data 170). Although the object replication reliability mechanism may have a high memory overhead, it also has low read and restore costs, which is why it may be suitable for very small value sizes.
[0036] The reliability mechanism of object replication can also be suitable for key values (for example, key value data 170) that have frequent updates, and can therefore be selected based on read and write frequency.
[0037] During object replication, the object / key-value data 170 is replicated to one or more additional KV storage devices 130 whenever a write occurs. A primary copy of the key-value data 170 can be located on any of the KV storage devices 130, which is indicated by a hash of the key modulo N. Replicas of the primary copy of the key-value data 170 can be located on immediately adjacent KV storage devices 130 or on consecutive storage devices 130 in a circular pattern.
[0038] The virtual device management layer 120 or a user can decide how many replicas of the key-value data 170 are created. For example, a distributed system using the virtual device management layer 120 can select three-way replication and make it the default. However, a user of the system can configure the number of replicas of the object to be more or less than the selected default.
[0039] Therefore, and for example when three-way replication is used, if a primary KV storage device 130-2 has a primary copy of the data (for example, the key-value data 170), the virtual device management layer 120 can store replicas of the primary copy of the data on subsequent replica KV storage devices 130-3 and 130-4, with all copies of the data being identical. That is, the copies of the data can be stored on the two (or more) immediately following KV storage devices 130-3 and 130-4 (for example, in a circular fashion) after the KV storage device 130-2, which has the primary copy of the data.
[0040] The copies of the data can be stored under the same key name / user key as the primary KV storage device 130-2, as well as in the replica KV storage devices 130-3 and 130-4. All copies of the data can have a checksum and an identifier to indicate the key-value data 170 that have been replicated.
[0041] Accordingly, if a specific KV storage device 130 fails (for example, if KV storage device 130-3 fails), a recovery mechanism using key name recovery on KV storage devices 130 immediately before and after the failed KV storage device 130 (for example, KV storage devices 130-2 and 130-4 immediately before and after KV storage device 130-3) ensures that a replicated key is also recovered.
[0042] To summarize the object replication reliability mechanism, the virtual device management layer 120 can receive a key value datum 170 and hash a key object to determine which KV storage device 130 to use for storing replicas of the key object. The virtual device management layer 120 can then write updated values under the same user key name (for example, with an appropriate MetaID field) to the selected KV storage devices 130 (for example, selected KV storage devices 130-2, 130-3, and 130-4).
[0043] Fig. Figure 3 is a block diagram representing a group of KV storage devices configured to store key value data according to a reliability mechanism of K-object (k, r) erasure coding or multi-object “packing” using conventional erasure coding according to an embodiment of the present disclosure.
[0044] Referring to Fig. 3. The reliability mechanism of packing, which uses conventional erasure coding, can be selected for data with small values that are not suitable for being divided into blocks or segments (for example, for the purpose of better data throughput). For example, the reliability mechanism of packing using conventional erasure coding can be selected by the virtual device management layer 120 for data with values larger than those that would lead to the selection of the previously described reliability mechanism of object replication, but which are still relatively small.
[0045] Packing using conventional erasure coding can be configured with conventional (k, r) maximum distance separable (MDS) erasure coding and can be used with any systemic MDS code. As an example, the erasure code could be a (4,2) Reed-Solomon code by default, since a (4,2) Reed-Solomon code is relatively well-studied, and fast implementation libraries corresponding to it are readily available.
[0046] When using the reliability mechanism of packing using conventional erasure coding, k keys / key objects 350 are selected from the queues of k different KV storage devices 330, which are part of the same parity group / erasure coding group 340, and erasure-coded to be packed, where k is an integer.
[0047] For example, the virtual device management layer 120 can maintain a buffer of recently written key objects 350 for each KV storage device 330 (for example, each KV storage device 130 of reliability group 140 of the Fig. 1) maintain to enable the virtual device management layer 120 to select k key objects 350 from k different KV storage devices 330 to be erasure-encoded in order to pack the k key objects 350 corresponding to the KV pairs.
[0048] In the present example, the virtual device management layer 120 selects four key objects 350x, 350y, 350b and 350c from four different KV storage devices 330-1, 330-3, 330-4 and 330-N respectively (that is, in the present example k = four).
[0049] Fig. Figure 4 represents a block diagram which provides a storage of value objects and parity objects in accordance with the reliability mechanism of K-object (k, r) erasure coding or multi-object “packing” using conventional erasure coding according to an embodiment of the present disclosure.
[0050] Referring to the Fig. 3 and Fig. 4. The key objects 350 are again combined into a (hash of the key modulo n) ten KV storage device 330 is placed. This means that for each key object 350, a respective hash of the key modulo n can be performed, which can be sent to the queue of this specific KV storage device 330. In the present example, the key i 350-i hashed and placed in KV-SSD 1 330-1, key j 350-j is hashed and placed in KV-SSD 2 330-2 and key. k350-k is hashed and placed in KV-SSD 4 330-4.
[0051] The user value length / value size 462 of the respective value object 450, which are stored, are the same as those that were written. However, for the purpose of consistency, in order to enable erasure coding, the user values / value objects 450 are all considered to have the same size by applying "zero" padding / virtual zeros / virtual zero padding 464 to them.This means that since the respective user value sizes 462 of different value objects 450 can vary (that is, the value objects 450 can have varying lengths / are variable-length key values), by implementing a virtual zero-padding procedure 464 (that is, by padding the value objects 450 with virtual zero-padding zeros 464 for coding purposes, while avoiding an actual rewrite of the data representing the value objects 450 to include the padded zeros), the parity objects 460 are able to be the same size as the largest value object(s) 470 in the parity group 340. Therefore, in the present example, value objects 450 "Val x", "Val y", and "Val b" are padded with virtual zeros that are not currently stored in any of the KV storage devices 330, in order to appear to be the same size. be like valuable item 470 “Val c”.After that, the parity objects 460 can be calculated.
[0052] After encoding the k key objects 350, the virtual device management layer 120 can calculate r parity objects 460 from the k values / k value objects 450 corresponding to the k key objects 350, where r is an integer, where k+r = N, and where N is the number of KV storage devices 330 in parity group 340 (for example, the N KV storage devices 130 of reliability group 140 of the Fig. 1).
[0053] The virtual device management layer 120 can store the r parity objects 460 in r remaining distinct KV storage devices 330 in parity group 340 (that is, in r KV storage devices 330, which are separate from the k KV storage devices 330, which have the queues from which the k key objects 350 are selected and erasure-encoded). Therefore, each of the k key objects 350 and r parity objects 460 can be stored in one of the N KV storage devices 330, and the corresponding data is evenly distributed in each of the N KV storage devices 330 of parity group 340.
[0054] Although reading and writing are relatively direct for the reliability mechanism of packing using conventional erasure coding, restoring and recalculating parity can be much simpler. For restoring and recalculating parity (for example, in the case of an update), in order to provide knowledge about which key objects 350 are grouped together in the same parity group 340, thus enabling parity calculation (for example, which key objects 350 are in an erasure code group 340), information concerning the groupings of the key objects 350, along with the current value size 462 of each value object 450 (that is, the value size 462 without the virtual zero padding 464), can be stored as a metadata object in each of the KV storage devices 330 (for example, the KV storage devices 130 of the Fig. 1) be stored. Accordingly, in the present embodiment, additional metadata can be used to identify the key objects 350 (for example, the key objects 350 which are in reliability group 140 of the Fig. 1 are placed) to store, wherein the original length of each of the value objects 450 corresponds to the key objects 350, and the KV storage devices 130 in the order of encoding the key objects 350.
[0055] For example, the metadata object value can have a field that displays all of the key objects 350 of the reliability group 140, and also displays the value sizes 462 of the value objects 450, and can have another field that displays the parity object keys (that is, the value objects 450 including the zeros of the virtual zero padding 464), the value sizes 462 of the parity objects 460, and the device IDs for identifying the corresponding r KV storage device 330 in which the r parity objects 460 are stored.
[0056] The data can be stored using the user key. The metadata can be stored in an internal key, which is generated using the user key and a MetaID indicator designated "Metadata". Furthermore, the value parameters 462 can be stored in the metadata to provide knowledge about the location of the virtual zero padding 464 (that is, where the zeros are added) for accurate reproduction when the value objects 450 are generated by determining where the value objects 450 end and the zeros of the virtual zero padding 464 begin.
[0057] Should one of the KV storage devices 330 fail, since both the data and the metadata can be stored in the same KV storage device 330, both the data and the metadata could potentially be lost, making recovery impossible. However, to avoid such a scenario, the metadata object value can be replicated using an "object replication machine" of the virtual device management layer 120, which is capable of implementing the aforementioned object replication reliability mechanism on the metadata object value.
[0058] Additionally, since the metadata object value is the same for all objects in a reliability group 140, if the KV storage device 330 supports object linking, the same metadata object value can be linked to multiple key names that are generally located in the same KV storage device 330. Furthermore, if batch writing is supported, object values can be processed together in batches for improved throughput.
[0059] To summarize the reliability mechanism of packing using conventional erasure coding according to the present embodiment, the virtual device management layer 120 can select k recently stored key objects 350 from k different KV storage devices 330 via a buffer. The virtual device management layer 120 can then retrieve the value objects 450 corresponding to the respective key objects 350 (other than a largest value object 470 of the parity group 440) and pad them with a virtual zero pad 464 to give the value objects 450 the same size (for example, the size of the largest value object 470). The virtual device management layer 120 can then use an MDS code process to generate r parity objects 460 from the k key objects 350.The virtual device management layer 120 can then write the r parity objects 460 to r KV storage devices 350, which are different from the k KV storage devices 330, from which the key objects 350 of the N KV storage devices 330 were selected, where k+r equals N. The virtual device management layer 120 can then create a metadata object representing the above information. Finally, the virtual device management layer 120 can write the key objects 350 and parity objects 460 to NKV storage devices 330 (similar to a replication machine, for example) with keys that, apart from the user key and a metadata identifier, are formed.
[0060] Fig. Figure 5 is a block diagram depicting a group of KV storage devices configured to store key value data according to a reliability mechanism of single-object(k, r) erasure coding or splitting using conventional erasure coding according to an embodiment of the present disclosure.
[0061] Referring to Fig. For values whose magnitudes are greater than those suitable for the object replication and packing reliability mechanisms described above using conventional erasure coding, the virtual device management layer 120 can select the single-object (k, r) erasure coding or "splitting" reliability mechanism using conventional erasure coding. The splitting reliability mechanism using conventional erasure coding is a per-object / KV-pair reliability mechanism that may be suitable for a KV value / object 570 with a relatively large magnitude and will have good throughput if the KV value 570 is split into k splits / segments / values / objects 550 of the same size.
[0062] According to one embodiment, after splitting the KV value 570, the virtual device management layer 120 can calculate a checksum for each of the k objects 550. The virtual device management layer 120 can then insert metadata before each of the k objects 550.
[0063] A partitioning using conventional erasure coding can involve splitting the KV value 570 into several smaller objects 550, and then distributing the several smaller objects 550 of the KV value 570 across k successive storage devices 530. Consequently, the size of the k objects 550 of equal size can be supported by each of the underlying KV storage devices 530.
[0064] When partitioning using conventional erasure coding is employed, the virtual device management layer 120 can also add r parity values / objects 560, which are generated using a systemic MDS code (for example, the virtual device management layer 120 can be configured with conventional (k, r) MDS erasure coding such as a (4,2) Reed-Solomon code as the default code). Then, in a manner similar to the reliability mechanism of packing using conventional erasure coding described above, the virtual device management layer 120 can write the k objects 550 and the r parity objects 560 to N KV storage devices 530 (k+r = N).
[0065] Therefore, the virtual device management layer 120 can divide a relatively large KV value 570 into k objects 550, can calculate and add r parity objects 560, and can store the k objects 550 and r parity objects 560 in k+r KV storage devices 530.
[0066] By using the reliability mechanism of partitioning using conventional erasure coding, after hashing a key 580 corresponding to the KV value 570, the virtual device management layer 120 can assign a primary KV storage device 530a (for example, KV-SSD 2 in the example shown in Fig. 5 is shown) to determine the storage of a corresponding object (for example, a first of the k objects 550, D1 in the example shown in Fig. (as shown in Figure 5, stored in a hash mark zero). Then the k+r objects 550 and 560 can be written under the same user key name to one of the primary KV storage devices 530a and N-1, respectively, to successive KV storage devices 530. That is, in the example shown in Fig. As shown in Figure 5, a first of the k objects 550 can be written to the primary KV storage device 530a “KV-SSD 2”, and the rest of the k objects 550 together with the r parity objects 560 are written sequentially in a circular manner to KV storage devices 530 “KV-SSD 3” to “KV-SSD N” and “KV-SSD 1” (for example, in a manner similar to that described above with respect to the reliability mechanism of object replication).
[0067] To summarize the reliability mechanism of partitioning using conventional erasure coding, the virtual device management layer 120 can partition a relatively large KV value 570 into k objects of equal size 550. The virtual device management layer 120 can then use an MDS coding process to generate r parity objects 560 for the k objects 550. The virtual device management layer 120 can then hash the key corresponding to the KV value 570 to determine a primary KV storage device 530a in which to place the object.The virtual device management layer 120 can then store the k+r objects 550, 560 under the same user key name, which may have a suitable MetaID field generated by the virtual device management layer 120, and which corresponds to the primary KV storage device 530a and N-1 successive KV storage devices 530 in a circular manner.
[0068] Referring back to the Fig. 3 and Fig. 4 According to another embodiment, the virtual device management layer 120 can implement the reliability mechanism of K-object (k, r, d) erasure coding or multi-object “packing” using regeneration erasure coding (for example, in accordance with flowchart 200 of the Fig. 2) select. The present reliability mechanism is similar to the previously described reliability mechanism of packing using conventional erasure coding in that the virtual device management layer packs 120 k objects into k KV storage devices. However, the packing using regeneration erasure coding j uses (k, r, d) regeneration codes instead of conventional (k, r) erasure codes. Consequently, the Fig. 3 and Fig. 4. General reference is made to the present embodiment.
[0069] Therefore, regeneration erasure packing can be used when regeneration codes are appropriate, but not when splitting the objects is not appropriate / when it is more appropriate to keep the objects intact. Regeneration erasure packing can be appropriate for value sizes larger than those used for the previously described object replication reliability mechanisms and for packing and splitting using conventional erasure coding. Regeneration erasure packing can be used when reading multiple subpacks of an object does not result in lower performance than reading the entire object. Regeneration erasure packing can also be appropriate when the underlying KV storage devices (for example, KV storage devices 130 of the Fig. 1 or KV storage devices 330 of the Fig. 3) Regeneration code-aware KV storage devices are those that are able to assist during a repair / recovery process.
[0070] Fig. Figure 6 is a block diagram representing a group of KV storage devices configured to store key value data according to a reliability mechanism of single-object (k, r, d) erasure coding or “splitting” using regeneration erasure coding according to an embodiment of the present disclosure.
[0071] Referring to Fig. 6. The present reliability mechanism of the virtual device management layer allows it to operate in a manner similar to partitioning using conventional erasure coding, as described in Fig. Figure 4 shows, except that (k, r, d) regeneration codes are used instead of conventional (k, r) MDS erasure coding. As with packing using regeneration erasure coding, the present reliability mechanism may be suitable if the underlying KV storage devices are regeneration code-aware KV storage devices that assist during a repair / reconstruction.
[0072] Splitting using regeneration erasure coding may be suitable if an object 670 has a value size that is larger than the objects that correspond to the reliability mechanisms described above, and if reading several subpackets 690 of k splits 680 of the object 670 does not result in lower performance than reading all the splits 680 (as is done, for example, with the split reliability mechanism using conventional erasure coding).
[0073] The reliability mechanism of partitioning using regeneration erasure coding is a per-object (KV-pair) mechanism which may be suitable for objects / KV values 670 which have very large value sizes, which will have a suitable throughput even if the object 670 is partitioned into k objects / partitions 650 of the same size, and if the partitions 650 are further virtually divided into a number of subpackets 690 (for example, four subpackets 690 per partition 650 in the present example), and if reading several subpackets 690 from an object 670 has a better throughput than reading the entire object 670, where the value size is supported by all underlying KV storage devices 630.
[0074] Similar to splitting and using conventional erasure coding, as in Fig.As shown in Figure 4, the virtual device management layer 120 of the present reliability mechanism adds r parity objects 660 using a systemic regeneration code and can write the k partitions 650 and the r parity objects 660 to N KV storage devices 630 (k+r = N). However, each of the r parity objects 660 can be partitioned into a number of parity subpackages 692 (for example, a number equal to the number of subpackages 690 per partition / k object 650). Unlike partitioning using conventional erasure coding, the standard code in the present embodiment can be a (4, 2, 5) zigzag code.
[0075] To summarize the reliability mechanism of partitioning using regeneration erasure coding, the virtual device management layer 120 can partition a large KV value 670 into k objects 650 of equal size. The virtual device management layer 120 can then partition each of the k objects 650 into m subpackets 690 of equal size, where m is an integer. The virtual device management layer 120 can then use a regeneration coding process to generate r parity objects 660 for the k objects 650, and each of the r parity objects 660 can be partitioned into m parity subpackets 692 of equal size. The virtual device management layer 120 can then hash the key corresponding to the KV value 670 to determine a primary KV storage device 630a in which to place the object.The virtual device management layer 120 can then write the k+r objects 650, 660, which have m subpackages 690, 692, for each under the same user key name, which may have a suitable MetaID field, which is generated by the virtual device management layer 120, and which corresponds to the primary KV storage device 630a and N-1 successive KV storage devices 630 in a circular manner.
[0076] As described above, a virtual device management layer can select a suitable reliability mechanism from a group of reliability mechanisms for storing data based on one or more characteristics of the data. Therefore, the embodiments described herein provide improvements for the field of memory storage, since the reliability mechanisms described are each capable of performing a single key repair procedure. If an entire storage device fails, the virtual device management layer of the embodiments of this disclosure can repair all the keys present in the failed storage device and copy them to a new storage device.The virtual device management layer can achieve a repair and copy of all the keys by iterating over all the keys present in the storage devices adjacent to the failed storage device in the reliability group, and by performing per-key repairs on the keys that the reliability mechanism determines were on the failed storage device.
[0077] The embodiments described herein further provide improvements for the area of memory storage, since very large KV pairs with value sizes that are larger than those supported by the underlying reliability mechanisms (for example, in accordance with the underlying memory device size restrictions) are explicitly split into multiple KV pairs by the reliability manager, and since the reliability mechanisms store the number of splits and split number information together with the metadata stored in the values.
Claims
[1] Method for storing data in a key-value reliability system comprising one or more storage devices which are grouped into a reliability group (140) as a single logical unit and which are managed by a virtual device management layer (120), wherein the method comprises: Determining that the data meets a limit value corresponding to a reliability mechanism for storing the data, wherein the reliability mechanism features object replication; a use of the reliability mechanism specified by a reliability mechanism identifier contained in metadata that identifies the reliability mechanism; and Storing the data according to the reliability mechanism by: selecting a KV value; a calculation of a hash to hash a key which corresponds to the selected KV value; a determination of a subset of storage devices of one or more storage devices for storing a replica of a key object corresponding to the KV value; and Writing an updated value, which corresponds to the KV value, to the subset of storage devices under the same user key name. [2] The method of claim 1, wherein the limit value is based on one or more of the following: an object size of the data; a throughput analysis of the data; a read / write temperature of the data; and an underlying erasure coding capability of one or more storage devices. [3] Method according to claim 1, further comprising the use of a Bloom filter or a cache for testing the data for the reliability mechanism. [4] Method according to claim 1, further comprising inserting the metadata with a key corresponding to the data for recording the reliability mechanism, wherein the metadata comprises a checksum for the one or more storage devices storing the data, an object size of a value of the data stored in the one or more storage devices storing the data, and a location of an element of a parity group of the one or more storage devices to indicate which of the one or more storage devices are storing the data. [5] The method of claim 1, wherein the reliability mechanism comprises packing, and wherein storing the data comprises: a selection of one or more key objects, each of which is stored in an equal number of first corresponding one or more storage devices of the one or more storage devices of the reliability group (140); a retrieval of one or more value objects that correspond to one or more key objects; a padding of a virtual zero at one end of one or more value objects which does not have a largest value of one or more value objects, in order to equalize a virtual value of one or more value objects; generating one or more parity objects from one or more key objects; a writing of one or more key objects to the first corresponding one or more storage devices; and a writing of one or more parity objects to an equal number of second corresponding one or more storage devices of the one or more storage devices, wherein the second corresponding one or more storage devices are different from the first corresponding one or more storage devices, where a number of one or more key objects plus a number of one or more parity objects equals a number of one or more storage devices. [6] Method according to claim 5, wherein the reliability mechanism comprises packing using erasure coding, and wherein the one or more storage devices are configured with maximum distance separable (MDS) erasure coding. [7] Method according to claim 5, wherein the reliability mechanism comprises packing using regeneration erasure coding, and wherein the one or more storage devices are configured with regeneration erasure coding. [8] The method of claim 1, wherein the reliability mechanism comprises a splitting, and wherein a storage of the data comprises: a division of the KV value into one or more objects of the same size; a creation of one or more parity objects from the one or more objects of the same size; Determining a primary device among one or more storage devices into which the KV value is to be placed, based on the hash; and a writing of one or more value objects to an equal number of first corresponding one or more storage devices of the one or more storage devices, and a writing of one or more parity objects to an equal number of second corresponding one or more storage devices of the one or more storage devices, wherein the second corresponding one or more storage devices are different from the first corresponding one or more storage devices, where a number of one or more value objects plus a number of one or more parity objects equals a number of one or more storage devices. [9] Method according to claim 8, wherein the reliability mechanism comprises partitioning using erasure coding, and wherein the one or more storage devices are configured with maximum distance separable (MDS) erasure coding. [10] Method according to claim 8, wherein the reliability mechanism comprises splitting using regeneration erasure coding, wherein one or more storage devices are configured with regeneration erasure coding, and wherein storing the data further comprises using regeneration erasure coding to split each of the one or more objects of the same size into one or more subpackages, and splitting the one or more parity objects into one or more parity subpackages. [11] Data reliability system for storing data based on a reliability mechanism, wherein the data reliability system comprises: one or more storage devices configured as a virtual device using stateless data protection; and a virtual device management layer (120) configured to manage the one or more storage devices as the virtual device for storing data in the one or more storage devices according to a reliability mechanism that includes object replication, wherein the virtual device management layer (120) is configured to: to determine that the data meets a threshold value that corresponds to a reliability mechanism for storing the data; to use the reliability mechanism specified by a reliability mechanism identifier contained in metadata that identifies the reliability mechanism; and to store the data according to the reliability mechanism by: selecting a KV value; a calculation of a hash to hash a key which corresponds to the selected KV value; a determination of a subset of storage devices of one or more storage devices for storing a replica of a key object corresponding to the KV value; and Writing an updated value, which corresponds to the KV value, to the subset of storage devices under the same user key name. [12] Data reliability system according to claim 11, wherein the reliability mechanism comprises packing, and wherein the virtual device management layer (120) is configured to store the data by: a selection of one or more key objects which are stored in an equal number of corresponding first storage devices of the one or more storage devices; a retrieval of one or more value objects that correspond to one or more key objects; a filling of a virtual zero at one end of one or more value objects which does not have a largest value size of one or more value objects, in order to equalize a virtual value size of one or more value objects; generating one or more parity objects from one or more key objects; a writing of one or more key objects to the first corresponding one or more storage devices; and a writing of one or more parity objects to an equal number of second corresponding one or more storage devices of the one or more storage devices, wherein the second corresponding one or more storage devices are different from the first corresponding one or more storage devices, where a number of one or more key objects plus a number of one or more parity objects equals a number of one or more storage devices. [13] Data reliability system according to claim 11, wherein the reliability mechanism comprises a split, and wherein the virtual device management layer (120) is configured to store the data by: a division of the KV value into one or more objects of the same size; a creation of one or more parity objects from the one or more objects of the same size; Determining a primary device among one or more storage devices in which the KV value is to be placed, based on the hash; and a writing of one or more value objects to an equal number of first corresponding one or more storage devices of the one or more storage devices, and a writing of one or more parity objects to an equal number of second corresponding one or more storage devices in successive order and starting with the primary device, wherein the second corresponding one or more storage devices are different from the first corresponding one or more storage devices, where a number of one or more value objects plus a number of one or more parity objects equals a number of one or more storage devices. [14] Data reliability system according to claim 13, wherein the reliability mechanism comprises splitting using regeneration erasure coding, wherein one or more storage devices are configured with regeneration erasure coding, and wherein the virtual device management layer (120) is further configured to store the data by using regeneration erasure coding to split each of the one or more objects of the same size into one or more subpackages, and to split the one or more parity objects into one or more parity subpackages. [15] Non-transitory, computer-readable medium comprising computer code which, when executed on a processor, implements a method for storing data in a key-value reliability system comprising one or more storage devices which are grouped into a reliability group (140) as a single logical unit and which are managed by a virtual device management layer (120), wherein the method comprises: Determining that the data meets a limit value which corresponds to an admissibility mechanism for storing the data, wherein the reliability mechanism features object replication; a use of the reliability mechanism specified by a reliability mechanism identifier contained in metadata that identifies the reliability mechanism; and storing the data according to the reliability mechanism by selecting a kV value; a calculation of a hash to hash a key which corresponds to the selected KV value; a determination of a subset of storage devices of one or more storage devices for storing a replica of a key object corresponding to the KV value; and Writing an updated value, which corresponds to the KV value, to the subset of storage devices under the same user key name. [16] Non-transistoric, computer-readable medium according to claim 15, wherein the reliability mechanism comprises packing, and wherein storing the data comprises: a selection of one or more key objects which are stored in an equal number of first corresponding one or more storage devices of the one or more storage devices of the reliability group (140); a retrieval of one or more value objects that correspond to one or more key objects; a padding of a virtual zero at one end of one or more value objects which does not have a largest value size of one or more value objects, in order to equalize a virtual value size of one or more value objects; generating one or more parity objects from one or more key objects; a writing of one or more key objects to the first corresponding one or more storage devices; and a writing of one or more parity objects to an equal number of second corresponding one or more storage devices of the one or more storage devices, wherein the second corresponding one or more storage devices are different from the first corresponding one or more storage devices, where a number of one or more key objects plus a number of one or more parity objects equals a number of one or which is one of several storage devices. [17] Non-transitory, computer-readable medium according to claim 15, wherein the reliability mechanism comprises a splitting, and wherein storing the data comprises: a division of the KV value into one or more objects of the same size; a creation of one or more parity objects from one or more objects of the same size; Determining a primary device among one or more storage devices into which the KV value is to be placed, based on the hash; and a writing of one or more value objects to an equal number of first corresponding one or more storage devices of the one or more storage devices, and a writing of one or more parity objects to an equal number of second corresponding one or more storage devices of the one or more storage devices in successive order and starting with the primary device, wherein the second corresponding one or more storage devices are different from the first corresponding one or more multiple storage devices, where a number of one or more value objects plus a number of one or more parity objects equals a number of one or more storage devices.
Citation Information
Patent Citations
Erasure coding and redundant replication
US8504535B1
Replicating key-value pairs in a continuous data protection system
US8949180B1